7 Commits

Author SHA1 Message Date
e7ebc667b3 chore(release): kyberforge 1.6.0, marketplace 0.4.2
The four preceding commits change `plugins/kyberforge/.apm/` content that reaches
the compiled artifacts — three validators, a new reference file in each of
`skill-audit` and `skill-author`, and the authoring rules across both author
skills — so per `apm-workflow`'s configure policy the package earns a bump, minor
for the new capability.

Root `apm.yml`'s `executables.allow` key moves with it, in this commit and not a
later one. apm approves a package's `hooks/` and `bin/` by an exact
`<name>#<version>` dictionary lookup with no wildcard and no version-less form, so
a `kyberforge#1.5.0` key left behind a 1.6.0 package errors nowhere: the entry
stops matching, the `SessionStart` freshness hook stops deploying, and the install
goes quietly stale. That is the failure ADR-0019 records as having actually
happened, and `check-executables-allow-sync` exists to catch it.

The catalog bump was missing from the working tree and is added here.
`apm-workflow`'s marketplace policy is explicit that an existing entry's
`version:` moving earns the catalog a **patch** — the set of packages is
unchanged, only its metadata moved — and that the root `version:` stays in step
with `marketplace.version`, since apm audit reads one and the compiled manifest
carries the other. Nothing enforces this: `apm pack --check-clean` catches a bump
made in `apm.yml` but never re-packed, while a bump never made at all fails
nothing.

Manifests regenerated with `apm pack` plus `scripts/sync-marketplace-mirror.sh`
for `.github/plugin/marketplace.json`, which no apm output profile targets.
`.agents/plugins/marketplace.json` is unchanged — the codex profile's shape
carries no version field for either the catalog or its entries.
2026-08-16 16:42:46 +00:00
64ffb9f35a docs: make ADR-0020 match what actually shipped, and record what did not
The ADR was written against base commit `f9b919d` and then not updated as the
implementation moved, so several of its numbers were measuring one thing and being
read as another — the exact conflation the ADR exists to stop, reproduced inside
it. Corrections, all reproducible now that each figure states its method:

- The preload tax is 23,427 chars / ~5,900 tokens, not 23,612 / ~6,200.
- `MAX_WORDS=2770` is a density proxy for the agentskills.io ~5,000-token ceiling,
  not "2× p90". Neither percentile reaches it: 2× the body-only p90 is 2,698 and
  2× the whole-file p90 is 3,052. Reading it as a percentile pairs a whole-file
  gate against a body-only distribution.
- `apm-workflow` is a 421-word body; 554 is its whole-file count. `skill-author`
  and `agent-author` were 2,623 and 2,582 body words — 2,760 and 2,758 whole-file,
  which is where "within twelve words of the gate" comes from. Two numbers for one
  file is the point, and only one of them is what either gate measures.
- Every `file:line` citation now says it resolves against `f9b919d`, since this
  change rewrites most of the cited files.

Three things the ADR asserted that no validator implemented are now filed by tier
in an exhaustive enforcement table — deterministic, prose-pattern, or auditor
judgment — because a rule filed under "Enforcement" that nothing enforces is the
failure mode this ADR is most exposed to. The Gotchas entry count moves to
SUGGESTION to match the script; the paraphrase FAIL is marked as an auditor's,
since semantic equivalence is not pattern-matchable.

Two gaps recorded rather than quietly left:

- The agent body-gate exemption lives in `agent-audit`'s validator and in the
  `skill-size-check` hook's `SKILL.md`-only `files:` pattern — *not* in
  `scripts/skill-size-check.sh`, which measures whatever path it is handed and
  today reports 900-word body FAILs on `git-orchestrate` (933),
  `gitea-orchestrate` (1,199) and `apm-orchestrate` (1,080). Agents escape by file
  pattern, not because the script knows the difference, so widening that pattern
  would silently enforce a gate this ADR declines to set.
- The `skill-audit`/`agent-audit` merge is deferred to #101. This change made the
  split deeper, not shallower: the dispatch retrofit took them from 3 and 4
  reference files to 7 and 8, and their two same-named `description-quality.md`
  files now differ on 100 of ~120 lines after normalising skill/agent. The merge
  reopens ADR-0008 and touches every call site in `skill-author`, `agent-author`
  and `forge`, so it is its own change. #100 carries the dangling-target fixes.

AGENTS.md and CONTEXT.md take the same corrections plus the two live setup
changes: PyYAML is now a hard requirement rather than an optional accelerator (a
fallback that mis-parses an unfamiliar scalar shape reports a clean pass on a file
it never measured), and `.claude/settings.json`'s `pretty-format-json` exclusion is
documented as load-bearing rather than as a tidy-up candidate.

LESSONS.md's autofix entry is corrected on its own provenance, which it got wrong
in both directions. `git log --date=iso` puts the introducing commit at 18:47 and
the fix at 21:54 — three hours, not "weeks" — and `git branch -a --contains` puts
the introducing commit on this branch only, not on main. It was manufactured
inside the same PR that diagnosed it. The added lesson is that "pre-existing" is a
claim about history and history is queryable: a defect found while working on a
branch feels inherited, and the feeling is not evidence.

Refs: ADR-0020, #99, #100, #101
2026-08-16 16:41:45 +00:00
d02765d595 fix(ci): close the RUN_TESTS_STRICT leak at its source, not at each caller
7607522 fixed the symptom in the wrong place. It made `test-run-tests.sh`'s
`run_fake()` spawn fixtures via `env -u RUN_TESTS_STRICT`, which stops that one
suite inheriting strictness — and leaves every future suite to defend itself the
same way. The variable's only job is done the moment `run-tests.sh` latches it
into the `STRICT` shell local, so it is unset there now and the leak is gone for
every child. The `env -u` stays as this suite's own defence in depth rather than
as the fix.

Two corrections to that commit's account of the bug, both overstated and both
cheap to have checked:

- The blast radius was two assertions, cases 10c and 10g, not six. Nothing else
  in the repo reads `RUN_TESTS_STRICT`.
- The pre-push gate was never red. It invokes `bash tests/run-tests.sh --strict`,
  and the flag sets a shell local that is never exported, so the flag spelling
  never leaked at all. Only the env-var spelling did.

That asymmetry between the two documented spellings is the real finding, and
nothing asserted against it. Case 10b compared the parent's verdict, which is the
half that already matched; the halves that differed were the environments the two
spellings handed every dispatched suite. New case 10i asks a child directly —
`${VAR+set}`, so an exported empty value still counts as a leak — and asserts the
two observations equal each other rather than a hardcoded expectation, so they
cannot drift apart in a direction the case did not anticipate.
2026-08-16 16:41:08 +00:00
311e7cd22c fix(kyberforge): reconcile the authoring rules the ADR-0020 trim left disagreeing
Six defects, each one a place where two files that an author reads in the same
sitting told them different things — or where the trim dropped a rule and nothing
noticed because no gate covers prose.

**"Use proactively" contradicted itself across the pair.** All three agent
templates said to add it where the runtime should delegate unprompted, while
`agent-audit`'s `KyberforgeCopilot.ProactivePhrase` rule grades it a hard FAIL in
any `*.agent.md` — which is the Copilot half of every project/user pair *and* the
vendor-neutral plugin-scope file, since that compiles to a real Copilot agent
downstream. Following the template produced a file the repo's own gate rejects.
The phrase is now permitted in exactly one place, the Claude Code `.md`, and
`references/contract.md` carries the per-file table plus the consequence authors
ask about next: a pair whose CC half has it and whose Copilot half does not is
correct, because `agent-audit` checks that both halves describe the same job, not
that they match word for word.

**The output-schema rule contradicted itself inside one file.** `contract.md`
said any content only one branch reaches moves to `references/`, and then offered
an "Output format template" body pattern with no qualification. Stated once now,
so it is not re-litigated: an output schema stays in the body only when every flow
produces it and it is roughly 50 words or less. No third option.

**Gotchas tiers disagreed with the script.** `validate.sh` emits the entry count
through `suggest()` and exits 0, while `skill-author` and `skill-audit` both
called more than five entries a FAIL. Whether a given gotcha earns its place is
judgment, so the prose moves to the script's tier rather than the reverse. The
paraphrase rule stays a FAIL and is explicitly marked as the auditor's call — no
script detects it.

**The dispatch exemplar was cited at the wrong number.** `apm-workflow`'s body is
421 words; 554 is its whole-file count. Both `contract.md` and `body-discipline.md`
cited 554 while describing a body budget, so an author calibrating against the
exemplar overshot by ~30% — the exact whole-file/body-only conflation those two
sections exist to warn against, reproduced inside the warning.

**"Error handling" came back as a required body element.** It was one of four and
is the one that gets dropped, and dropping it is not neutral: an agent handed
malformed input with no instruction invents a recovery, and a subagent's invented
recovery is invisible to its caller until the output is wrong. Restored in
`agent-audit`'s rubric as a SUGGESTION, in `agent-author`'s contract and both
scope checklists as a required element, and as an `## Errors` section in all three
templates.

**`skill-author` Step 4 gains the one check the audit misses.** An empty body
reports `PASS SKILL.md body word count 0` — a word gate cannot tell "concise"
from "absent". Step 4 now hand-checks for a non-empty section, and its commit
verification is conditioned on actually being inside a git worktree, which a skill
under `~/.claude/skills/` is not.

Also here: absolute repo paths removed from `skill-author`'s SKILL.md and
contract.md in favour of naming the skill (`zoom-out`'s description is quoted
inline instead of pointed at), the boundary-target universe documented to match
the resolver, a two-hops-from-SKILL.md limit on reference chains, and
`new-agent.sh`'s next-steps output naming the description budget and the
deliberate absence of an agent body gate.

Refs: ADR-0020
2026-08-16 16:40:51 +00:00
2540e50fcc feat(kyberforge): give skill-author a procedure for the #99 retrofit
ADR-0020 shipped its gates hot with no baseline file, so 26 of 39 descriptions
and 9 of 39 bodies are over their FAIL tier and editing any of them for any
reason requires bringing the skill into contract first. `references/improve.md`
said exactly that and stopped there — it mandated a retrofit and supplied no
procedure for one.

Four dry-run retrofits confirmed what that costs. Asked the same questions —
what to cut first, when a body is two flows rather than one, what else has to
change alongside — they invented six to ten different answers, so the same skill
retrofitted twice produced two different skills and neither run could be reviewed
against anything.

`references/retrofit.md` fixes the answers: an ordered cut list ranked by tokens
removed against behaviour lost (inverting that order is how a retrofit deletes the
instruction the skill existed to carry), the test for whether a body holds two
mutually exclusive flows, the reference-file conventions, the collateral checklist
for `README.md` and `references/sources.md`, and a worked description retrofit.

It also states the trap the dry runs kept hitting: retrofit the skill in place,
inside its package. The boundary-target universe is built by walking up from the
file being checked, so a scratch copy has no authoring root above it, the check
prints `INFO ... DID NOT RUN`, and the run still exits 0 — a line that reads as a
pass and is not one. A retrofit signed off on a copy carries an unverified
boundary target into the corpus.

Loaded from the improve flow only when a budget is actually exceeded, so a routine
improvement pays nothing for it.

Refs: ADR-0020, #99
2026-08-16 16:40:19 +00:00
a85bdbed42 fix(kyberforge): restore skill-audit's script-failure fallback and E100 diagnostic
The ADR-0020 body trim took `skill-audit` from 2,623 body words to a dispatch
shape, and two things went out with it that were not padding.

The manual structural fallback was one. Its replacement was a single sentence
telling the auditor to report an INFO when `validate.sh` cannot run — so with no
`python3` or no PyYAML, `skill-audit` reported the gap honestly and then audited
nothing structural at all. Every ADR-0020 measurement, the whole-file ceilings,
the name-to-directory match, the `references/` pointer check and the script
hygiene checks silently left the audit. A skill's whole Structure dimension
hanging on one optional interpreter is the same vacuous-pass shape the gate
scripts were just fixed for, one layer up.

The `E100 Runtime error ... does not exist` diagnostic was the other. That exit
code means an explicit relative `--config` was passed to `vale-wrap.sh` while
vale itself was installed and working; without the note, Step 1's fallback reads
exit 2 as "vale unavailable" and downgrades the description, body-discipline and
patterns dimensions to full LLM judgment for a config error it could have fixed.
That misreading is already recorded in CONTEXT.md as the reason both audit skills
stopped passing `--config` at all.

Both are restored in `references/validation-scripts.md`, loaded only when a Step 1
script fails — so the body pays nothing for them on a clean run, which is what the
dispatch pattern is for. The file also carries the by-hand boundary-target
procedure and the three ways to misread the result, including that
`INFO ... DID NOT RUN` is not a pass.

`references/file-structure.md` gains the one sanctioned spelling for a cross-skill
reference. The possessive form (``skill-audit's references/validation-scripts.md``)
is the only spelling both rules accept: a full repo path is what that section
already forbids, and a bare `references/<file>.md` is now a hard ERROR from the
ADR-0020 pointer check, which requires the file to exist in the skill's *own*
directory. Without the rule the two constraints look mutually exclusive.

Refs: ADR-0020
2026-08-16 16:40:00 +00:00
b6e68e9a2b fix(kyberforge): close the vacuous-pass paths in the ADR-0020 gate scripts
Three ways the gates could report green having measured nothing. All three were
invisible to a passing test suite, because pre-commit prints nothing at all for a
hook that exits 0 — a gate that declines to check and a gate that checked and
passed produce the identical signal.

- A UTF-8 BOM, a leading blank line, a trailing space after a `---` marker or
  CRLF line endings defeated the `^---\n` frontmatter matcher. Every ADR-0020
  check was then skipped and the file passed: measured at the time, a
  550-character description with a 1,000-word body exited 0 behind a BOM.
  All four shapes are now tolerated, and frontmatter that genuinely cannot be
  parsed is a hard ERROR rather than a silent skip.
- An agent file with a valueless `description:` followed by another key let a
  line regex capture the *next* key, which looked non-empty, so the
  missing-or-empty branch never fired and every gate below it early-returned on
  the empty folded value — zero output, exit 0, on a blocking gate. The one
  field this contract is entirely about was the one field a gate could fail to
  notice was absent. Presence is now decided on the YAML-folded value and
  nowhere else, and a missing or empty description is a hard FAIL in all three
  validators.
- The hand-rolled frontmatter fallback disagreed with PyYAML across the FAIL
  boundary on folded scalars, so which reader happened to be available decided
  the verdict. A fallback that mis-parses a scalar shape reports a vacuous pass,
  which is worse than not running, so it is deleted: python3 and PyYAML are hard
  requirements that fail loudly with an install pointer.

Boundary-target resolution no longer derives its universe from its own location.
A `${BASH_SOURCE}`-relative repo root leaked this repo's 39-skill universe into
every consumer repo running the hook through pre-commit, so a consumer skill
routing to `skill-audit` resolved against a plugin it had never installed. The
interim form resolved through `.claude/` and `.agents/`, which are gitignored
`apm install` output — the same commit reported 2 dangling targets on a machine
that had run the install and 6 on a fresh clone. Resolution now walks up from the
file being checked to an authoring root (nearest ancestor holding
`plugins/*/.apm/{skills,agents}`, else the nearest `.git`, in two passes so a
nested `.git` cannot outrank a real monorepo root); the universe is every skill
and agent under `<root>/plugins/*/` plus the file's own apm package and that
package's declared `dependencies.apm`. Deployed trees are consulted only when no
authoring root exists at all — the consumer case. One commit now gets one verdict,
which a gate shipping hot with no baseline file has to.

Narrowed in the same pass: a routing target inferred from the prose boundary form
and corroborated by nothing else reports at SUGGESTION instead of blocking. A
blocking check with no escape hatch is the wrong trade when the inference from
prose is the weak part of it.

New deterministic checks, all previously untested or absent: every
`references/<file>.md` a body names must exist (ERROR — a broken pointer is not a
style opinion); a description with no boundary clause at all, a Gotchas section
over five entries, and a Gotchas section over 25% of the body are SUGGESTIONs.
Where no universe can be determined the target check prints `INFO ... DID NOT
RUN` rather than passing quietly. Each prose-scanning check needed its own
false-positive fix — a fenced example of a Gotchas section was being read as the
section itself — and those fixes are pinned rather than assumed.

The resolver is one block copied verbatim into all three scripts between
BEGIN/END markers, because a cache-installed plugin's scripts cannot read outside
their own plugin directory. Nothing asserted the copies were still identical; a
one-line edit to a single copy passed every constant-agreement assertion, since
constants are not what drifts.

Tests land here rather than in a later commit. The existing suites assert the old
behaviour and go red against these scripts, so splitting them would leave a commit
whose own `run-tests` pre-push gate fails in isolation.

Refs: ADR-0020
2026-08-16 16:39:29 +00:00
69 changed files with 7439 additions and 963 deletions

View File

@@ -1,7 +1,7 @@
{
"name": "holocron",
"description": "AI development skills for Claude Code and GitHub Copilot CLI — factory, design, implement, review, and cross-cutting workflows.",
"version": "0.4.1",
"version": "0.4.2",
"owner": {
"name": "Defame1297",
"email": "defame1297@rkdr.net",
@@ -11,7 +11,7 @@
{
"name": "kyberforge",
"description": "Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.",
"version": "1.5.0",
"version": "1.6.0",
"category": "Developer Tools",
"source": "./plugins/kyberforge"
},

View File

@@ -1,7 +1,7 @@
{
"name": "holocron",
"description": "AI development skills for Claude Code and GitHub Copilot CLI — factory, design, implement, review, and cross-cutting workflows.",
"version": "0.4.1",
"version": "0.4.2",
"owner": {
"name": "Defame1297",
"email": "defame1297@rkdr.net",
@@ -11,7 +11,7 @@
{
"name": "kyberforge",
"description": "Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.",
"version": "1.5.0",
"version": "1.6.0",
"category": "Developer Tools",
"source": "./plugins/kyberforge"
},

View File

@@ -38,14 +38,15 @@ Fall back to raw shell only when no skill covers it.
## Setup and testing
- Run `apm install` to deploy this repo's own skills and agents into `.claude/skills/` and `.claude/agents/`. Both are gitignored install output, not authoring source — `plugins/<name>/.apm/` remains the only place to edit. The six dependencies in root `apm.yml` resolve from the holocron **remote**, unpinned against the default branch, so a `.apm/` edit is not visible to the running session until it is pushed and `apm update` re-runs (`apm install` deploys from `apm.lock.yaml` and does not re-resolve refs). Needs the network, and needs `apm_modules/` (which it materializes) left gitignored. `apm install` also configures the `obsidian` MCP server into the repo's `.mcp.json`, carried over from `plugins/bin/.mcp.json`.
- Do not add repo-owned keys to `.claude/settings.json`. apm treats that file as its own deployed artifact: `apm audit --ci` replays the install into a scratch tree and diffs, so anything apm would not have written there — an `enabledPlugins` block, a real `hooks` entry — is permanent drift that fails the `apm-audit-ci` pre-push hook. Its committed content is whatever apm last wrote — `{"hooks": {}}` until kyberforge's `SessionStart` hook lands there, after which the merged hook entry is apm's output and belongs in the commit (ADR-0019). What does not change is that nothing repo-authored goes in the file. A hook you want in this repo is authored in `plugins/<name>/.apm/hooks/` and deployed by apm, never hand-written here. Machine-specific settings go in the gitignored `.claude/settings.local.json`, which apm does not deploy and the replay does not compare; shared enforcement belongs in `.pre-commit-config.yaml`.
- Do not add repo-owned keys to `.claude/settings.json`. apm treats that file as its own deployed artifact: `apm audit --ci` replays the install into a scratch tree and diffs, so anything apm would not have written there — an `enabledPlugins` block, a real `hooks` entry — is permanent drift that fails the `apm-audit-ci` pre-push hook. Its committed content is whatever apm last wrote, which today is the merged `SessionStart` entry for kyberforge's `check-apm-current.sh` — apm's own output, and it belongs in the commit (ADR-0019). What does not change is that nothing repo-authored goes in the file. A hook you want in this repo is authored in `plugins/<name>/.apm/hooks/` and deployed by apm, never hand-written here. The file is also **excluded from `pretty-format-json`** in `.pre-commit-config.yaml` — the sixth and last alternation in that `exclude:` pattern, and the only one there for a reason other than "generated manifest". Mind which number you are quoting: six alternations, expanding to sixteen real files (3 root marketplace manifests, 2 per plugin × 6 plugins, plus this one). `pretty-format-json --autofix` sorts object keys while apm emits insertion order, so leaving the file in that hook's scope rewrites apm's output on the way into every commit and `apm audit --ci` then reports permanent drift on a file with an empty `git diff`. Do not tidy it out of that list; it is load-bearing (see `LESSONS.md`, 2026-08-14). Machine-specific settings go in the gitignored `.claude/settings.local.json`, which apm does not deploy and the replay does not compare; shared enforcement belongs in `.pre-commit-config.yaml`.
- Keeping the install current is automatic but not free. Because the six dependencies are unpinned, deployed skills go stale whenever anyone merges. kyberforge ships a `SessionStart` hook that runs `apm outdated` at startup (~0.7s) and, when something is behind, runs `apm update --yes` and asks the host to re-scan skills (~10.4s). That rewrites `apm.lock.yaml`, so an unexplained modification to it after opening a session is expected, not a bug — commit or discard it deliberately. Note `apm install` alone will **not** pick up remote changes; it deploys from the lock. `apm update` is the command that re-resolves refs.
- Install git hooks via `pc-run`, wiring all three stages — this repo's `.pre-commit-config.yaml` has no `default_install_hook_types`, so a plain install silently skips `commit-msg` (Conventional Commits) and `pre-push` (the 14-hook gate described below).
- Install the `apm` CLI — four pre-push hooks shell out to it: `apm-marketplace-check`, `apm-audit-ci`, `apm-pack-check-clean`, and `check-plugin-content-sync` (via `scripts/sync-plugin-content.sh`, which wraps `apm pack`). `apm-marketplace-check` and `apm-pack-check-clean` are bare `apm …` hook entries and `apm-audit-ci` is a `bash -c` loop calling `apm` once per package, so without it the push dies with an unhelpful "command not found". Use `apm-install`, or `curl -sSL https://aka.ms/apm-unix | sh`; verify with `apm --version`.
- Install `jq` — required by `scripts/check-manifests.sh` and `scripts/sync-plugin-content.sh`, both pre-push. These at least fail loudly (`Error: jq is required but not installed`).
- Install `python3` — required by `scripts/skill-size-check.sh`, the `skill-size-check` pre-commit hook. It measures the *folded* `description` value: most descriptions here are `>`-block scalars, so a regex over the raw lines measures indentation and newlines instead of the value. Missing it fails the hook with an install pointer rather than skipping the ADR-0020 checks, which would be a vacuous green. In practice it is already present — pre-commit is itself a Python application. PyYAML is used when importable and is genuinely optional; a fallback reader covers the frontmatter shapes this corpus uses.
- That hook enforces **two independent gate families** over `plugins/*/.apm/skills/*/SKILL.md`, and neither replaced the other. The agentskills.io spec backstop is unchanged: 500 lines and 2,770 words, counted over the **whole file including frontmatter**. ADR-0020 adds a context budget measured differently — `description` 250 chars SUGGESTION / 400 FAIL (it is preloaded into every session whether the skill fires or not), **body-only** word count 600 SUGGESTION / 900 FAIL (everything after the frontmatter's closing `---`), and every boundary-clause routing target resolving to a real skill or agent under `plugins/*/.apm/`. A file can sit well inside one family and fail the other. The hook is `verbose: true` so the SUGGESTION tier is audible — pre-commit prints nothing at all for a passing hook, and a SUGGESTION deliberately does not fail. `skill-audit`'s `validate.sh` holds a second copy of the four ADR-0020 constants; `tests/test-skill-size-check.sh` asserts the copies agree.
- Install `python3` — required by `scripts/skill-size-check.sh`, the `skill-size-check` pre-commit hook. It measures the *folded* `description` value: most descriptions here are `>`-block scalars, so a regex over the raw lines measures indentation and newlines instead of the value. Missing it fails the hook with an install pointer rather than skipping the ADR-0020 checks, which would be a vacuous green. In practice it is already present — pre-commit is itself a Python application. **PyYAML is a hard requirement too**, not an optional accelerator: the hand-rolled fallback frontmatter reader has been removed, because a reader that mis-parses an unfamiliar scalar shape reports a clean pass on a file it never measured, which is the exact vacuous-green failure the `python3` check exists to avoid. `pip install pyyaml` if the hook reports it missing.
- That hook enforces **two independent gate families** over `plugins/*/.apm/skills/*/SKILL.md`, and neither replaced the other. The agentskills.io spec backstop is unchanged: 500 lines and 2,770 words, counted over the **whole file including frontmatter**. ADR-0020 adds a context budget measured differently — `description` 250 chars SUGGESTION / 400 FAIL (it is preloaded into every session whether the skill fires or not), **body-only** word count 600 SUGGESTION / 900 FAIL (everything after the frontmatter's closing `---`), a missing, valueless or `null` `description:` (a hard FAIL, not a skip — a gate that declines to measure the one preloaded field reports green), every boundary-clause routing target resolving to a real skill or agent, and every `references/<file>.md` a body names actually existing. Target resolution walks up **from the file being checked** to an authoring root — the nearest ancestor holding `plugins/*/.apm/{skills,agents}`, falling back to the nearest `.git`, in two passes so a nested `.git` cannot beat a real monorepo root. The universe is then every skill and agent under `<root>/plugins/*/`, plus the checked file's own apm package and whatever that package declares in its own `apm.yml` `dependencies.apm`; the **root** manifest's `dependencies:` block is not read, and no plugin here declares a cross-plugin apm dependency. Deployed `.claude/`/`.agents/` trees are consulted only when no authoring root exists — the consumer case. That matters because those trees are gitignored `apm install` output: resolution used to reach the four cross-plugin `gitea-*` → `git-*` targets through `.claude/skills/` alone, so the same commit measured 2 dangling targets on a developer machine and 6 on a fresh clone. It no longer does — verified by running the hook over a tree holding only `plugins/` and the root `apm.yml`, which reports findings identical to the working tree (26 description / 9 body / 2 dangling / 0 missing references / 58 SUGGESTIONs). Three further checks are SUGGESTION-only: a description with no boundary clause at all, a `## Gotchas` section with more than five entries, and a `## Gotchas` section over 25% of the body. A file can sit well inside one family and fail the other. The hook is `verbose: true` so the SUGGESTION tier is audible — pre-commit prints nothing at all for a passing hook, and a SUGGESTION deliberately does not fail. `skill-audit`'s `validate.sh` holds a second copy of the four ADR-0020 constants; `tests/test-skill-size-check.sh` asserts the copies agree.
- **Those ADR-0020 gates ship hot, with no baseline file.** 26 of 39 descriptions and 9 of 39 bodies currently exceed their FAIL tier, so editing one of those skills *for any reason* means retrofitting it to the contract first — a one-line fix to `gitea-prs` cannot be committed until that skill complies. This is deliberate, and the retrofit is tracked as Gitea issue #99. Check where a skill stands before starting: `pre-commit run skill-size-check --all-files`.
- **A second gate ships hot alongside it, and `skill-size-check` will not warn you about it.** `Kyberforge.CompositionNote` — the ADR-0020 Vale rule banning composition and architecture prose from a description — currently fires **10 errors across four skills**: `gitea-issues`, `gitea-labels-milestones`, `gitea-prs` and `gitea-workflow`. Every Vale rule here is `level: error` with no ignorable tier, so touching any of those four means fixing its prose findings as well as its size findings. Scoping a retrofit off `skill-size-check` output alone will leave you blocked at the second gate. Check both: `pre-commit run --all-files`.
- Install the `vale` binary — required by the `vale-audit-prefilter-skill`/`-agent` pre-commit hooks. Their `files:` patterns are `.apm/`-scoped: `^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$` and `^plugins/[^/]+/\.apm/agents/[^/]+\.agent\.md$`. Only the authoring source triggers them — a `SKILL.md` in the generated mirror matches neither pattern, so prose findings surface only when you edit the file you are supposed to be editing. Without the binary the hooks fail with a bare "command not found" and no install pointer. `brew install vale` (macOS), `snap install vale` (Linux), `choco install vale` (Windows), or see https://vale.sh/docs/vale-cli/installation/. No `vale sync` needed — the `Kyberforge` styles are committed under `plugins/kyberforge/.apm/skills/{skill-audit,agent-audit}/assets/vale/styles/`, not downloaded packages (see ADR-0014).
- `vale` is also a **pre-push** dependency, not only pre-commit. `check-vale-style-sync` runs six glob-coverage probes by invoking `vale --config` — they are the only assertions in it that catch a `.vale.ini` glob typo, the failure mode where every text-level check stays clean while vale lints zero files. Missing `vale` is therefore a hard failure there. The opt-out is `CHECK_VALE_STYLE_SYNC_ALLOW_MISSING_VALE=1`, and it is **not** `SKIP=`: the hook still runs and still asserts everything verifiable from file text, but the six probes do not, and its summary says so explicitly — `Vale style sync check passed (text-level only, vale unavailable): … 0 glob probe(s) verified`. Use it only on a machine that genuinely cannot install `vale`, and read that summary line as "the glob axis was not checked", not as a pass.
- Run `bash tests/run-tests.sh` before considering any change done — it runs every `test-*.sh` script in the repo plus the bats suite (`--bats-only` for just bats). First run auto-initializes the bats submodules; no manual `git submodule update` needed.
@@ -53,7 +54,7 @@ Fall back to raw shell only when no skill covers it.
- `tests/run-bats.sh` derives the set of `.bats` files it expects from `git ls-files`, so a `.bats` file deleted from the worktree but still tracked in the index fails the run rather than silently shrinking the suite. Remove one with `git rm` (or stage the deletion) when the removal is intentional; an untracked new `.bats` file is picked up and needs no ceremony. Both discovery walks (`tests/run-bats.sh` and `tests/run-tests.sh`) exclude `apm_modules/`: `apm install` materializes a full copy of every plugin there, and running a dependency's copy of a `.bats` file breaks its relative path to the bats helpers — 167 spurious failures before the exclusion landed.
- Pushing runs 14 repo-defined pre-push hooks, not just the test suite — `run-tests` and `check-manifests`, plus generated-content drift gates (`check-plugin-content-sync`, `check-marketplace-mirror-sync`, `check-vale-style-sync`, `check-scope-walkup-sync`, `check-executables-allow-sync`), artifact validators (`check-apm-agents-valid`, which runs agent-audit's `validate.sh` over every real `plugins/*/.apm/agents/*.agent.md`), apm's own gates (`apm-marketplace-check`, `apm-audit-ci`, `apm-pack-check-clean`), host validators (`validate-plugins`, `validate-marketplace`, both needing the `claude` CLI), and `check-release-needed`. `check-executables-allow-sync` is the odd one in that first group — it guards a silent failure rather than drift in generated text. apm gates a package's `hooks/` and `bin/` on an exact `<package>#<version>` lookup in root `apm.yml`'s `executables.allow`, with no wildcard and no version-less form, so bumping `plugins/kyberforge/apm.yml`'s `version:` without bumping the key errors nowhere: the entry simply stops matching, kyberforge's `SessionStart` hook stops deploying, and the install goes quietly stale — the failure ADR-0019 records as live. Run `pre-commit run --hook-stage pre-push --all-files` locally — one command, the whole gate. That command reports **16**, not 14: pre-commit's own `meta` hooks, `check-hooks-apply` and `check-useless-excludes`, declare no `stages:` and so run at every stage including this one.
- `apm-audit-ci` runs `apm audit --ci` once per manifest — the root one and each of the six plugin packages — because the root-only invocation audits the marketplace manifest and **nothing else**, and `apm-pack-check-clean` does not parse plugin `dependencies:` blocks either (verified: a malformed one passes `apm pack --check-versions --check-clean --dry-run` and fails `apm audit --ci` in that package's directory). It verifies two things and claims no more: each `apm.yml` parses as a valid APM manifest, and any package declaring dependencies has a consistent `apm.lock.yaml`. It does **not** enforce an org policy — apm discovers one from the git remote and only understands github.com and Azure DevOps, so against this repo's self-hosted Gitea remote it prints `No org policy found at unknown; enforcement skipped`. Do **not** "fix" that with `policy.fetch_failure_default: block` in `apm.yml`: it was tested and rejected, because with no reachable policy source it makes the hook exit 1 on every push forever.
- `check-apm-agents-valid` derives its expected agent-file set from `git ls-files` (same pattern as `tests/run-bats.sh`), so an agent file deleted from the worktree but still tracked fails the run, and discovering zero agent files is an error rather than a pass. An untracked new agent file is still validated — the derivation is one-directional on purpose, so uncommitted work is not blocked but also cannot bypass the gate. Agents take the ADR-0020 description gates (`agent-audit`'s `validate.sh` holds its own copy of those two constants) and, deliberately, **no** body word gate: an agent body becomes the system prompt of a fresh context rather than competing with the caller's live conversation, so the 900-word FAIL does not transfer. A bats test pins that absence — adding a body gate there contradicts the ADR rather than fixing an inconsistency.
- `check-apm-agents-valid` derives its expected agent-file set from `git ls-files` (same pattern as `tests/run-bats.sh`), so an agent file deleted from the worktree but still tracked fails the run, and discovering zero agent files is an error rather than a pass. An untracked new agent file is still validated — the derivation is one-directional on purpose, so uncommitted work is not blocked but also cannot bypass the gate. Agents take the ADR-0020 description gates (`agent-audit`'s `validate.sh` holds its own copy of those two constants) and, deliberately, **no** body word gate: an agent body becomes the system prompt of a fresh context rather than competing with the caller's live conversation, so the 900-word FAIL does not transfer. A bats test pins that absence in `agent-audit`'s validator — adding a body gate there contradicts the ADR rather than fixing an inconsistency. Be precise about the scope of that guarantee, though: it holds for the **validator**, not for the shared script. `scripts/skill-size-check.sh` applies its body gate to whatever path it is handed, and `bash scripts/skill-size-check.sh plugins/*/.apm/agents/*.agent.md` exits 1 today with 900-word body FAILs on `git-orchestrate` (933), `gitea-orchestrate` (1,199) and `apm-orchestrate` (1,080). Agent files escape only because the hook definitions filter on `SKILL.md` — a file-pattern accident that happens to implement the design, not the design itself. Do not "extend" that hook's `files:` pattern to cover agents on the assumption that the script already knows the difference.
- **Two** pre-push hooks need the network, for one shared reason: root `apm.yml`'s `marketplace.packages[]` contains exactly one remote entry (`mattpocock-skills`, `source: mattpocock/skills`), and resolving it needs a `git ls-remote`. `apm-marketplace-check` resolves every entry and is `always_run`, so it fails with `No cached refs (offline)`. `apm-pack-check-clean` (`apm pack --check-versions --check-clean --dry-run`) re-resolves the same entry and fails with `Error: Git network timeout during ls-remote`. Pinning the entry to an exact version does **not** remove the call — an exact pin still ls-remotes. `--offline` rescues neither. To push without a network, skip both using pre-commit's own mechanism: `SKIP=apm-marketplace-check,apm-pack-check-clean git push`. Skip those two alone — verified under `unshare -rn`, the other twelve pre-push hooks pass offline because they are real local checks (`check-executables-allow-sync` landed after that run, but reads two local manifests and makes no network call), and adding one of them to `SKIP` disarms it silently. `apm-audit-ci` calls `apm` too but stays local: its org-policy discovery resolves nothing on this remote before any network call, so it does not join the pair above.
- Author commits with `git-commits` — it validates Conventional Commits (enforced at `commit-msg`) for you.

View File

@@ -27,13 +27,13 @@ A separate product (separate repo) for browsing, editing, and configuring AI dev
Reusable slash commands for AI coding tools, defined as `SKILL.md` files following the [Agent Skills open standard](https://agentskills.io). Authored at `plugins/<plugin-name>/.apm/skills/<skill-name>/SKILL.md` and reaching a host by one of two install paths: `apm install`, which deploys the skill directory to `.claude/skills/<skill-name>/` (this repo's own path — see "apm-consumed install"), or `claude plugin install <name>@<marketplace>`, which caches the whole plugin (still supported for external consumers). Skills are self-contained — they cannot reference files outside the plugin directory after install-time caching. The two paths name skills differently: apm deploys a plain project skill (`skill-audit`), a plugin install namespaces it (`kyberforge:skill-audit`).
### Preload tax
The always-on context cost of every installed skill's `name` + `description`, which sit in the agent's context from the first token of every session whether or not the skill is invoked. Measured 2026-08-14 at 23,612 chars (~6,200 tokens) across 39 skills, plus 1,325 chars for 4 agents. Non-routing frontmatter (`metadata.source_keys`, `category`, `version`) is **not** part of it — the model-visible skill listing carries only `name` and `description`, which supersedes `LESSONS.md:63` on this host. Bodies are not part of it either; they are charged on invocation.
The always-on context cost of every installed skill's `name` + `description`, which sit in the agent's context from the first token of every session whether or not the skill is invoked. Measured 2026-08-14 against base commit `f9b919d` at 23,427 chars (~5,900 tokens) across 39 skills, plus 1,325 chars for 4 agents. Method, so it can be re-run: sum `len(name) + len(description)` over each `plugins/*/.apm/skills/*/SKILL.md` frontmatter with `>` block scalars folded to the value the host loads, at ~4 characters per token. Non-routing frontmatter (`metadata.source_keys`, `category`, `version`) is **not** part of it — the model-visible skill listing carries only `name` and `description`, which supersedes `LESSONS.md:63` on this host. Bodies are not part of it either; they are charged on invocation.
### Skill context contract
The authoring rules that hold the preload tax and body size down, set by ADR-0020. A description carries a trigger clause, at most one capability clause, and a boundary clause of the form `Not <thing> → <skill-name>` naming a resolvable target — nothing else. Capability enumeration, output formats, and composition notes ("composes X rather than duplicating Y") belong in the body or `README.md`; a description that summarises workflow is a correctness hazard, not just a cost, because agents act on it instead of reading the body. Sizes are two-tier and sit *below* the agentskills.io spec limits, which stay unchanged as conformance backstops: description 250 SUGGESTION / 400 FAIL (spec 1,024); body 600 SUGGESTION / 900 FAIL (spec 2,770 words / 500 lines). Conflating the quality gate with the spec ceiling is what let `skill-author` and `agent-author` grow to within twelve words of 2,770.
The authoring rules that hold the preload tax and body size down, set by ADR-0020. A description carries a trigger clause, at most one capability clause, and a boundary clause of the form `Not <thing> → <skill-name>` naming a resolvable target — nothing else. "Resolvable" is decided by walking up *from the file being checked* to an **authoring root** — the nearest ancestor holding `plugins/*/.apm/{skills,agents}`, falling back to the nearest `.git`, in two passes so a nested `.git` cannot outrank a real monorepo root. The universe is then every skill and agent under `<root>/plugins/*/` (sibling plugins resolve against each other, which is what a monorepo means), plus the checked file's own apm package and that package's own declared `dependencies.apm`. The **root** manifest's dependency list is never consulted, and no plugin here declares a cross-plugin apm dependency. Deployed `.claude/`/`.agents/` trees count only when there is no authoring root at all — the consumer case. The property this buys is that one commit gets one verdict: those trees are gitignored `apm install` output, so resolving through them made the same commit report 2 dangling targets on a developer machine and 6 on a fresh clone, which a gate shipping hot with no baseline cannot do. A `${BASH_SOURCE}`-relative repo root is the other half of the same defect and is gone — it leaked this repo's 39-skill universe into consumer repos running the hook through pre-commit. A *missing* boundary clause is a SUGGESTION rather than a failure, for skills and agents alike — some skills genuinely have no near-miss sibling. A *missing or empty description* is the opposite: a hard FAIL in all three validators, because a gate that merely declines to measure the one preloaded field reports green. Capability enumeration, output formats, and composition notes ("composes X rather than duplicating Y") belong in the body or `README.md`; a description that summarises workflow is a correctness hazard, not just a cost, because agents act on it instead of reading the body. Sizes are two-tier and sit *below* the agentskills.io spec limits, which stay unchanged as conformance backstops: description 250 SUGGESTION / 400 FAIL (spec 1,024); body 600 SUGGESTION / 900 FAIL (spec 2,770 words / 500 lines). Conflating the quality gate with the spec ceiling is what let `skill-author` and `agent-author` grow to within twelve words of 2,770.
### Dispatch body
The body pattern a skill with two or more mutually exclusive flows must use: the body carries only the dispatch table and the gates common to every branch, and each flow lives in its own self-contained `references/` file. Named for `apm-workflow` (554-word body, 3,006 words of references), which arrived at it independently and is the repo's exemplar. Its absence is the characteristic defect — `skill-author` inlines both its create and improve flows, and `agent-author` carries 50-60 lines marked inapplicable by their own headers on any single run.
The body pattern a skill with two or more mutually exclusive flows must use: the body carries only the dispatch table and the gates common to every branch, and each flow lives in its own self-contained `references/` file. Named for `apm-workflow` (421-word body, 3,006 words of references), which arrived at it independently and is the repo's exemplar. Its absence was the characteristic defect at the time ADR-0020 was written: `skill-author` inlined both its create and improve flows, and `agent-author` carried 50-60 lines marked inapplicable by their own headers on any single run. Both were retrofitted to dispatch tables in the change that carries the ADR — `skill-author` went 2,623 body words to 595 and `agent-author` 2,582 to 616 — so they are now worked examples of the pattern rather than counter-examples of it. The 39-skill corpus at large is not: 9 bodies still exceed the 900-word FAIL (issue #99).
### Hand-invoked skill
A skill reached only by typing its slash command, declared with `disable-model-invocation: true`. The host withholds it from the model-visible skill listing entirely, so it pays no preload tax and its `description` becomes human-facing text rather than a trigger list. `zoom-out` is the worked example: apm passes the flag through verbatim to both install paths, and the skill is absent from the router while `/zoom-out` still works. Choosing model-invoked vs. hand-invoked is the first question `skill-author` asks, because it determines whether a description needs triggers at all.
@@ -100,7 +100,7 @@ Wiring Vale as a deterministic prefilter for `skill-audit`/`agent-audit`'s Descr
Both skills' Step 1, and the `vale-audit-prefilter-skill`/`-agent` pre-commit hooks, call each copy's own `scripts/vale-wrap.sh` rather than `vale` directly — a workaround for a confirmed Vale 3.15.2 limitation (see `vale-config`'s Gotchas): `text.frontmatter.description` silently stops matching on most — not all — multi-line descriptions. Verified by reproduction, not assumed: `>` folded scalars, plain (unquoted) continuation lines, and single- or double-quoted multi-line scalars all yield 0 alerts and exit 0 on a deliberately-bad fixture, while a `|` literal block spanning the same 2+ lines lints normally (alerts fire, exit 1). The wrapper flattens those three broken forms to one physical line in a scratch copy (padding with blank lines so every other line number is unchanged) before handing off to real `vale`; `|` literal blocks and single-line descriptions pass through untouched, already linting correctly. The plain and quoted forms previously passed silently — unflattened and unmatched — so a bad description in either sailed through the prefilter. Handed no `--config` at all, the wrapper falls back to its own sibling `assets/vale/.vale.ini`, located from `${BASH_SOURCE[0]}` rather than from the cwd — which is why both manifests' `entry:` is now the bare script path with no argument after it. pre-commit prefixes only `entry[0]` with the hook-repo clone path (`cmd = (prefix.path(cmd[0]), *cmd[1:])`), so every later argument resolves against the *consuming* repo's root: a `--config` in `.pre-commit-hooks.yaml` pointed at a path no consumer has and hard-failed every external run with `E100 [--config] Runtime error`. `.pre-commit-config.yaml` drops the argument too, deliberately keeping the two entries identical — the local `repo: local` hook resolved its `--config` correctly only because the consuming repo *was* this repo, and that divergence is why three review rounds exercised a path no external consumer takes and missed the defect. An explicit `--config` still wins, in all three argv forms (`--config X`, `--config=/abs`, `--config=rel`), and a relative one still resolves against the caller's cwd, matching bare `vale`, not the repo root. Both audit skills' Step 1 now passes no `--config` either: it resolves the script relative to the skill's own directory so the call works from an installed plugin cache, but a relative `--config` alongside it would still resolve against the cwd, yielding `E100 Runtime error ... does not exist` and exit 2 — which both skills' fallback misreads as "vale unavailable" and silently downgrades to full LLM judgment. `tests/test-vale-wrap.sh` regression-tests this against skill-audit's copy specifically (its fixtures are all `SKILL.md`-shaped, and only skill-audit's `.vale.ini` has that glob section). Each `.vale.ini`'s section globs are path-agnostic (`[**/SKILL.md]` for skill-audit's copy; `[**/agents/*.md]`/`[**/*.agent.md]` for agent-audit's) and do no scoping on their own: Vale's `*` crosses `/`. Scoping comes from each pre-commit hook's own `files:` regex and from the audit skills passing one explicit file per invocation. The two manifests scope differently on purpose: this repo's `.pre-commit-config.yaml` pins its own layout — `^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$` for `-skill`, `^plugins/[^/]+/\.apm/agents/[^/]+\.agent\.md$` for `-agent` — while the shipped `.pre-commit-hooks.yaml` stays layout-agnostic for external consumers whose skills live anywhere, using `(^|/)SKILL\.md$` and `(^|/)agents/[^/]+\.md$|\.agent\.md$`. Both manifests split the prefilter into two hooks precisely because one combined hook pointed at only one copy would silently 0-file-skip the other file type. A `SKILL.md` outside `plugins/` (e.g. project-scope `.claude/skills/foo/SKILL.md`) still matches `[**/SKILL.md]` and gets linted normally — the globs constrain filename shape, not location. Vale reports 0 files only when the path it is handed matches no glob section at all: a differently-named file, or a directory argument holding nothing that matches. That run prints `✔ 0 errors ... in 0 files.` and exits 0, indistinguishable from a clean pass, so both audits treat a 0-file Vale run as NOT RUN and fall back to full LLM judgment.
This scope expands per ADR-0013: one cherry-picked low-noise `write-good`/`alex` rule landed in `styles/Kyberforge`, `Kyberforge.SentenceOpenerThereIs` (22 held-out hits, both in-corpus hits clean rewrites, zero suppressions). A second, `Kyberforge.VagueQualifier`, was cherry-picked and then deleted: 2 hits across the 41 skill/agent files, one marginal and one an unfixable false positive (`caveman/SKILL.md` quotes `of course` as an example of filler — a mention, not a use) that forced the repo's only Vale suppression comments. Also new is a sibling pre-commit hook, `skill-size-check` (`scripts/skill-size-check.sh`), enforcing agentskills.io's `SKILL.md` ceiling as two blocking gates: `MAX_LINES=500` and `MAX_WORDS=2770` (a word-count proxy for the 5,000-token limit, calibrated to the densest prose measured in this repo — 1.81 tokens per word — so even a worst-case `SKILL.md` at the ceiling stays under 5,000 tokens). Both are inclusive, and `skill-audit/scripts/validate.sh` checks the same pair on the same terms, so a `SKILL.md` can no longer pass its own audit yet be blocked by the commit hook. Scoped to `^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$` only, same as `vale-audit-prefilter-skill`, so it never lints `docs/research/examples/` reference skills. It's also exposed in the root-level `.pre-commit-hooks.yaml` as `kyberforge-skill-size-check` — it has no external asset dependency, so it needed no relocation, only exposure to external consumers. File scope (`SKILL.md` + agent files) and enforcement model (rules land directly in `styles/Kyberforge`, blocking immediately, no trial tier) stay unchanged; governance.md/CONTROLS.md were evaluated and excluded as rule sources (nothing prose-pattern-matchable to mine). House convention: banned phrasing that must be mentioned rather than used goes in backticks or a fenced code block — Vale skips code spans and fences, so no suppression is needed; inline `<!-- vale Rule = NO -->` (HTML-comment form; the MDX `{/* */}` form does not work in plain Markdown) is the fallback only where backticking is impossible.
This scope expands per ADR-0013: one cherry-picked low-noise `write-good`/`alex` rule landed in `styles/Kyberforge`, `Kyberforge.SentenceOpenerThereIs` (22 held-out hits, both in-corpus hits clean rewrites, zero suppressions). A second, `Kyberforge.VagueQualifier`, was cherry-picked and then deleted: 2 hits across the skill/agent corpus as it stood at the time of that measurement (2026-08-08, before the `.apm/` restructure), one marginal and one an unfixable false positive (`caveman/SKILL.md` quotes `of course` as an example of filler — a mention, not a use) that forced the repo's only Vale suppression comments. A third, `Kyberforge.CompositionNote`, landed with ADR-0020 and bans architecture and composition prose from a description; it is `level: error` like the rest, and it currently fires 10 times across `gitea-issues`, `gitea-labels-milestones`, `gitea-prs` and `gitea-workflow`, so `pre-commit run --all-files` is red on prose as well as on size until issue #99 lands. Also new is a sibling pre-commit hook, `skill-size-check` (`scripts/skill-size-check.sh`), which carries **two independent gate families that must not be conflated** (see "Skill context contract"). The agentskills.io spec backstop is `MAX_LINES=500` and `MAX_WORDS=2770`, both inclusive and both counting the **whole file including frontmatter** (2,770 is a word-count proxy for the 5,000-token limit, calibrated to the densest prose measured in this repo — 1.81 tokens per word — so even a worst-case `SKILL.md` at the ceiling stays under 5,000 tokens; it is not a percentile of the corpus). ADR-0020 adds a context budget measured differently: description characters 250 SUGGESTION / 400 FAIL, **body-only** words 600 SUGGESTION / 900 FAIL, plus deterministic checks that every boundary routing target resolves, that a body's named `references/<file>.md` all exist, and — SUGGESTION-tier — that a boundary clause is present at all, that `## Gotchas` holds at most five entries, and that it stays under 25% of the body. `skill-audit/scripts/validate.sh` and `agent-audit/scripts/validate.sh` hold their own copies of the shared constants and `tests/test-skill-size-check.sh` asserts the copies agree, so a `SKILL.md` can no longer pass its own audit yet be blocked by the commit hook. Agents take the description gates and no body word gate. `python3` **and PyYAML** are hard requirements — the earlier hand-rolled frontmatter fallback is gone, because a fallback that silently mis-parses a scalar shape reports a vacuous pass. Scoped to `^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$` only, same as `vale-audit-prefilter-skill`, so it never lints `docs/research/examples/` reference skills. It's also exposed in the root-level `.pre-commit-hooks.yaml` as `kyberforge-skill-size-check` — it has no external asset dependency, so it needed no relocation, only exposure to external consumers. File scope (`SKILL.md` + agent files) and enforcement model (rules land directly in `styles/Kyberforge`, blocking immediately, no trial tier) stay unchanged; governance.md/CONTROLS.md were evaluated and excluded as rule sources (nothing prose-pattern-matchable to mine). House convention: banned phrasing that must be mentioned rather than used goes in backticks or a fenced code block — Vale skips code spans and fences, so no suppression is needed; inline `<!-- vale Rule = NO -->` (HTML-comment form; the MDX `{/* */}` form does not work in plain Markdown) is the fallback only where backticking is impossible.
### LESSONS.md
Long-loop feedback log for patterns observed across sessions. Three or more entries on the same pattern graduate to the relevant standing file (e.g. a coding convention, a governance rule). Updated by the session-handoff skill or directly by the human. Lives at the repo root.

View File

@@ -194,7 +194,11 @@ is the broken multi-`raw:` form, which `tests/test-vale-hooks-consumer.sh` now f
## 2026-08-14 — Un-anchoring a description rule to reach mid-sentence text is unshippable
Widening `DescriptionOpener` to catch `gitea-workflow`'s mid-description "This is the human-facing
entry point…" looked like a one-character change. Under `scope: text.frontmatter.description`, `^`
entry point…" looked like a one-character change. Both that skill and `gitea-labels-milestones`
*open* with "Use when…" and satisfy the opener rule; the offending clause sits at character 377 and
300 of the folded value respectively, so the rule was never violated and never silently passed — it
simply had no jurisdiction, which is a different defect and takes a different fix.
Under `scope: text.frontmatter.description`, `^`
anchors to the start of the whole description value — and `vale-wrap.sh` has already flattened that
value to one physical line, so `(?m)` changes nothing. Un-anchoring is therefore the only route to
mid-description text, and measured across the corpus it scores 5 hits and 5 false positives: skills
@@ -206,21 +210,38 @@ a new case does not fit.
## 2026-08-14 — A formatter in the commit path manufactures drift on a file with a clean git diff
`apm audit --ci` failed for weeks on `.claude/settings.json` while `git diff` on that file was empty —
the worst possible pairing of signals, because the file matched HEAD exactly and every instinct says
"nothing changed here". The content was identical to apm's output to the byte; only the JSON key
order differed. `pretty-format-json --autofix` sorts object keys unless `--no-sort-keys` is passed,
and its `exclude:` listed fifteen generated manifests but not this file, so from the commit that
first wrote a hook entry there (`2e395a4`) onward, apm's insertion-ordered output was silently
re-sorted on the way in. apm then replayed the install, produced its own order, and reported drift
against a file no human had touched.
`apm audit --ci` failed on `.claude/settings.json` while `git diff` on that file was empty — the worst
possible pairing of signals, because the file matched HEAD exactly and every instinct says "nothing
changed here". The content was identical to apm's output to the byte; only the JSON key order
differed. `pretty-format-json --autofix` sorts object keys unless `--no-sort-keys` is passed, and its
`exclude:` listed fifteen generated manifests but not this file, so from the commit that first wrote
a hook entry there onward, apm's insertion-ordered output was silently re-sorted on the way in. apm
then replayed the install, produced its own order, and reported drift against a file no human had
touched.
Two general points. First, a tool-owned generated file that passes through an autofixing formatter is
drifted by construction, and the diff that would reveal it never appears in `git diff` — it only
The provenance matters as much as the mechanism, and the first account of this entry got it wrong in
both directions. `git log --format='%h %ad %s' --date=iso` puts the introducing commit `2e395a4` at
2026-08-14 18:47 and the fix `7607522` at 21:54 — roughly three hours, not "weeks". And `2e395a4` is
the **first commit of the `refactor/trim-skills-agents-context` branch**, eleven minutes after the
base merge `f9b919d`; `git branch -a --contains 2e395a4` returns only that branch and its own
`remotes/origin/` tracking copy — two lines naming one branch, and `main` is not among them. So
this was not a latent defect inherited from `main`, it was manufactured inside the same PR that
diagnosed it, and the fixing commit's own message calling it "pre-existing … red at HEAD before
ADR-0020 work began" is the mis-attribution rather than the record. Two cheap commands would have
settled it before either sentence was written.
Three general points. First, a tool-owned generated file that passes through an autofixing formatter
is drifted by construction, and the diff that would reveal it never appears in `git diff` — it only
exists between the formatter's input and its output, which nothing stores. Second, the fix is
self-undoing unless the exclude lands in the same commit: correcting the file alone means the hook
re-breaks it as it is staged. Fix: when a tool declares ownership of a path, add that path to every
autofixing hook's `exclude` at the moment ownership is declared, not when the drift is noticed. This
repo gates marketplace-mirror, plugin-content and vale-style drift deterministically and has no
equivalent gate asserting tool-owned paths stay out of formatter scope — `.claude/settings.json` was
the sixteenth exclude and nothing prevents a seventeenth.
re-breaks it as it is staged. Third — the one this entry had to learn twice — "pre-existing" is a
claim about history, and history is queryable; a defect found while working on a branch feels
inherited, and the feeling is not evidence. A three-hour-old self-inflicted bug and a months-old
inherited one call for different responses, and writing the wrong one down converts a process failure
into a story about someone else's neglect. Fix: when a tool declares ownership of a path, add that
path to every autofixing hook's `exclude` at the moment ownership is declared, not when the drift is
noticed — and before describing any defect as pre-existing, run `git log -S` or
`git branch --contains` on the commit that introduced it. This repo gates marketplace-mirror,
plugin-content and vale-style drift deterministically and has no equivalent gate asserting tool-owned
paths stay out of formatter scope — `.claude/settings.json` was the sixteenth exclude and nothing
prevents a seventeenth.

View File

@@ -1,5 +1,5 @@
name: holocron
version: 0.4.1
version: 0.4.2
description: AI development skills for Claude Code and GitHub Copilot CLI — factory, design, implement, review, and cross-cutting workflows.
license: MIT
@@ -42,7 +42,7 @@ dependencies:
# after a kyberforge release, check this first.
executables:
allow:
kyberforge#1.5.0:
kyberforge#1.6.0:
hooks: true
bin: true
@@ -52,7 +52,7 @@ marketplace:
# top-level apm.yml description:/version: above are NOT inherited into the
# compiled output despite being used elsewhere (e.g. by `apm audit`).
description: AI development skills for Claude Code and GitHub Copilot CLI — factory, design, implement, review, and cross-cutting workflows.
version: 0.4.1
version: 0.4.2
owner:
name: Defame1297
email: defame1297@rkdr.net
@@ -79,7 +79,7 @@ marketplace:
- name: kyberforge
description: Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.
source: ./plugins/kyberforge
version: 1.5.0
version: 1.6.0
category: Developer Tools
- name: bin

View File

@@ -2,7 +2,7 @@
Every installed skill's `name` and `description` sits in every agent's context from the first token
of every session, whether or not the skill is ever invoked. Across this repo's 39 skills that is
23,612 characters — roughly 6,200 tokens — and the authoring rules that produced it optimised for
23,427 characters — roughly 5,900 tokens — and the authoring rules that produced it optimised for
triggering reliability with no counter-pressure on size. This ADR sets the budget, the shape, and the
gates that hold them.
@@ -10,14 +10,32 @@ gates that hold them.
## Context
Measured before any change:
Every `file:line` citation in this ADR is against the base commit the decision was taken on,
`f9b919d7e3bd5e6b51fbdf88b32ace0438b313e0`, not against current `HEAD`. The change that carries this
ADR rewrites several of the cited files, so a citation resolved against the worktree will land on
unrelated text. Use `git show f9b919d:<path>` to follow one.
Measured before any change, at that commit. Method, so the figures are reproducible: sum
`len(name) + len(description)` over the frontmatter of every `plugins/*/.apm/skills/*/SKILL.md`,
folding `>` block scalars to the value the host actually loads (most descriptions here are folded
scalars, so counting raw lines measures indentation instead); tokens at the standard
~4-characters-per-token approximation `scripts/skill-size-check.sh` uses. Word counts are
whitespace-separated tokens, and are stated as **body-only** or **whole-file** every time, never bare.
| | |
|---|---|
| 39 skill `name` + `description` | 23,612 chars, ~6,200 tokens, **preloaded every session** |
| 4 agent `name` + `description` | 1,325 chars, ~350 tokens, preloaded every session |
| skill bodies | median 684 words, mean 815, p90 1,349 |
| `MAX_WORDS` gate (`skill-audit/scripts/validate.sh:147`) | **2,770** — 2× p90 |
| 39 skill `name` + `description` | 23,427 chars, ~5,900 tokens, **preloaded every session** |
| 4 agent `name` + `description` | 1,325 chars, ~330 tokens, preloaded every session |
| skill bodies (body-only words) | median 684, mean 815, p90 1,349 |
| skill files (whole-file words) | median 816, mean 927, p90 1,526 |
| `MAX_WORDS` gate (`skill-audit/scripts/validate.sh:147`) | **2,770** whole-file — a density proxy, not a percentile |
That last row is worth stating plainly, because it is the first thing this ADR is about. 2,770 is not
derived from the corpus distribution at all: per the derivation comment in
`scripts/skill-size-check.sh`, it is 2,770 words at the densest observed 7.22 chars/word ≈ 20,000
chars ≈ the agentskills.io ~5,000-token ceiling. Neither percentile reaches it — 2× the body-only p90
is 2,698 and 2× the whole-file p90 is 3,052 — and reading it as "2× p90" would pair a whole-file gate
against a body-only distribution, which is exactly the conflation this ADR exists to stop.
Three findings drove this, none of which is "the descriptions drifted".
@@ -32,7 +50,8 @@ with six capability clusters. Across the twelve longest descriptions, 30.7% is c
enumeration and 11.6% is composition or implementation detail that cannot affect a routing decision.
**Capability enumeration in a description is a correctness hazard, not only a token cost.**
`docs/research/examples/skill-write/writing-skills/SKILL.md:154-158` reports a measured failure: "when
`plugins/kyberforge/docs/research/examples/skill-write/writing-skills/SKILL.md:154-158` reports a
measured failure: "when
a description summarizes the skill's workflow, an agent may follow the description instead of reading
the full skill content. A description saying 'code review between tasks' caused an agent to do ONE
review, even though the skill's flowchart clearly showed TWO reviews." `git-commits` is exactly that
@@ -41,21 +60,25 @@ chars, lowercase subject, no trailing periods, 11 standard types`) an agent can
loading the body.
**The upstream sources cannot settle this.** The four skill-writing references under
`docs/research/examples/skill-write/` disagree on what a description contains — when-only
(`writing-skills/SKILL.md:99`), what-and-when (`skill-creator/SKILL.md:67`,
`anthropic-best-practices.md:187`), triggers-only (`writing-great-skills/SKILL.md:28`), and
what-plus-when-plus-negative (`write-skill/SKILL-TEMPLATE.md:5-6`). `writing-skills` and the Anthropic
document it bundles contradict each other inside one skill directory. They also disagree on whether
`plugins/kyberforge/docs/research/examples/skill-write/` disagree on what a description contains —
when-only (`writing-skills/SKILL.md:99`), what-and-when (`skill-creator/SKILL.md:67`,
`writing-skills/anthropic-best-practices.md:187`), triggers-only
(`writing-great-skills/SKILL.md:28`), and what-plus-when-plus-negative
(`write-skill/SKILL-TEMPLATE.md:5-6`). Those four paths are relative to that directory.
`writing-skills` and the Anthropic document it bundles contradict each other inside one skill
directory. They also disagree on whether
500 lines is binding, on the inline-versus-bundle threshold, and on the TOC threshold (>100 lines vs
>300 lines). "Grounded in the research" is therefore not available as a tiebreaker; a house choice is
required and this is it.
A fourth observation shaped the body half. The best progressive-disclosure ratio in the repo belongs
to `apm-workflow` — a 554-word body dispatching to 3,006 words of references — and the worst two
belong to the skills that define the house standard: `skill-author` (2,760 body / 1,247 references)
and `agent-author` (2,758 / 1,664). Both sit within twelve words of the 2,770 gate their own plugin
enforces. A ceiling that nothing approaches is not a constraint; a ceiling that two files have grown
into is a target.
to `apm-workflow` — a 421-word body dispatching to 3,006 words of references — and the worst two
belong to the skills that define the house standard: `skill-author` (2,623-word body / 1,247 words of
references) and `agent-author` (2,582 / 1,664). Measured the other way, whole-file, those two are
2,760 and 2,758 words — ten and twelve words under the 2,770 gate their own plugin enforces. A
ceiling that nothing approaches is not a constraint; a ceiling that two files have grown into is a
target. The two numbers for one file are the point: 2,623 and 2,760 describe the same `skill-author`,
and only one of them is what either gate measures.
## Decision
@@ -69,8 +92,39 @@ clause**, and a **boundary clause**. Capability enumeration, output-format detai
- **250 characters SUGGESTION, 400 FAIL.** The agentskills.io 1,024-character limit remains as an
unchanged spec backstop. The SUGGESTION tier is what moves the average; the FAIL tier only stops
outliers.
- **A missing, valueless or `null` `description:` is a hard FAIL** in all three validators. That
reads as a trivial precondition and is not: a `description:` line with no value followed by
`model: sonnet` let a line regex capture the *next* key, which looked non-empty, so the "missing or
empty" branch never fired and every gate below it then early-returned on the genuinely empty folded
value — exit 0, zero output, on a blocking pre-push gate. Presence is decided on the YAML-folded
value and nowhere else. The field this contract is entirely about is the one field a gate must
never fail to notice is absent.
- **Boundary clauses compress** to `Not <thing> → <skill-name>.` and must name a target that
resolves to a real skill under `plugins/*/.apm/skills/`. This is checked deterministically.
resolves to a real skill or agent. Resolution walks up **from the file being checked** to an
*authoring root* — the nearest ancestor holding `plugins/*/.apm/skills` or `plugins/*/.apm/agents`,
falling back to the nearest ancestor holding `.git`. Two passes rather than one interleaved walk,
so a nested `.git` (a submodule, a sub-package worktree) cannot beat a real monorepo root further
up. When an authoring root is found the universe is every skill and agent under
`<root>/plugins/*/`, plus the target's own apm package and the packages that package declares in
its own `apm.yml` `dependencies.apm`. Sibling plugins resolve against each other, which is what a
monorepo means. Deployed `.claude/`/`.agents/` trees are consulted **only** when no authoring root
exists — the consumer case, where there is no monorepo to read. What the resolver must never do is
derive the universe from its own location: a `${BASH_SOURCE}`-relative repo root leaked this repo's
39-skill universe into every consumer repo running the hook through pre-commit, so a consumer skill
routing to `skill-audit` resolved against a plugin it had never installed. Checked
deterministically. A description carrying **no** boundary clause at all is a SUGGESTION, for skills
and agents alike: most descriptions want one, some genuinely have no near-miss sibling to exclude,
and that judgment is not a script's to make.
- **The verdict must not depend on whether `apm install` has been run.** Deployed trees are
gitignored install output, present only on a machine that has run it. Four cross-plugin targets
here (`gitea-branches` → `git-branches`, `gitea-branches` → `git-history`, `gitea-issues` →
`git-branches`, `gitea-workflow` → `git-workflow`) once resolved through `.claude/skills/` alone,
so the same commit measured 2 dangling targets on a developer machine and 6 on a fresh clone. A
gate shipping hot with no baseline cannot give two answers. Under the walk-up those four resolve
because sibling plugins are in the universe — no plugin here declares a cross-plugin apm
dependency, and none needs to. Verified: a tree holding only `plugins/` and the root `apm.yml`,
with no `.claude/` or `.agents/` anywhere, now produces findings identical to the working tree —
26 description FAILs, 9 body FAILs, 2 dangling targets, 0 missing references, 58 SUGGESTIONs.
- **The blanket pushiness rules are deleted.** `skill-author/SKILL.md:104` and
`description-quality.md:21` are replaced by a conditional: add an indirect trigger only where the
user's natural phrasing genuinely omits the domain word — true for the `gitea-*` family, false for
@@ -89,11 +143,16 @@ prose move to `references/` behind an explicit "read X when Y" trigger.
current state.
- **Dispatch is mandatory at two or more mutually exclusive flows.** The body carries the dispatch
table and the gates that apply to every branch; each flow lives in its own self-contained
`references/` file. This is `apm-workflow/SKILL.md:33-41` promoted from accident to rule.
`references/` file. This is `apm-workflow/SKILL.md:33-41` promoted from accident to rule. "Two
mutually exclusive flows" is not decidable from file text, so this rule is auditor judgment — see
Enforcement below for what that means and does not mean.
- **Every `references/<file>.md` a body names must exist.** A dispatch table pointing at a file that
was never written is a silently dead branch. Checked deterministically.
- **Gotchas are constrained.** A Gotcha must state a fact that contradicts a reasonable default —
something the agent gets wrong by acting sensibly. Maximum five entries. A Gotcha that paraphrases
a step in the body below it is a FAIL. A Gotchas section exceeding 25% of the body is a
SUGGESTION.
something the agent gets wrong by acting sensibly. More than five entries is a SUGGESTION, as is a
Gotchas section exceeding 25% of the body; both are countable and both are checked
deterministically. A Gotcha that paraphrases a step in the body below it is a FAIL, but a FAIL an
auditor issues, not a script — semantic equivalence is not pattern-matchable.
### Agents
@@ -101,6 +160,14 @@ Agents take the same description gates — they are preloaded identically — an
A skill body is loaded into the caller's context, competing with the live conversation; an agent body
becomes the system prompt of a fresh context. The rationale for the 900-word FAIL does not transfer.
That exemption is expressed in `agent-audit/scripts/validate.sh`, which has no body constant, and in
the `files:` pattern of the `skill-size-check` pre-commit hook, which is `SKILL.md`-only. It is *not*
expressed in `scripts/skill-size-check.sh` itself, which measures whatever path it is handed —
running it directly over `plugins/*/.apm/agents/*.agent.md` today reports 900-word body FAILs on
`git-orchestrate` (933), `gitea-orchestrate` (1,199) and `apm-orchestrate` (1,080). Agents escape by
file pattern, not by the script knowing the difference. Anyone widening that pattern to cover agents
would silently enforce a gate this ADR declines to set.
A plugin-scope agent is a single file with no sibling `references/` directory, so it cannot disclose
to itself — it can only delegate to skills. `agent-audit` therefore gains a **delegation check**: an
agent body that restates a procedure owned by a skill it can invoke is a FAIL, with the fix being
@@ -126,26 +193,84 @@ type of input they take should be **one skill with a dispatch table**. This catc
one-or-two-file agent pair, per ADR-0005 and ADR-0016) and their overlap is in the improve flow
rather than the core job.
**DEFERRED — not implemented in the change that carries this ADR. Tracked as issue #101.** Both
skills still exist separately, and this change made the split deeper rather than shallower: retrofit
to the dispatch pattern took `skill-audit` from 3 reference files to 7 and `agent-audit` from 4 to 8,
and their two same-named `references/description-quality.md` files now differ on 100 of ~120 lines
after normalising `skill`/`agent`, where before they were closer. The merge stays the decision; it
reopens ADR-0008 (agent-audit's single-file invocation contract) and touches every call site in
`skill-author`, `agent-author` and `forge`, which is why it is its own change and not a rider on
this one. Recorded here rather than dropped, so the gap between the rule and the tree is deliberate
and dated instead of discovered later.
### Enforcement and rollout
Gates land where the existing gates already live — no new layer:
Gates land where the existing gates already live — no new layer. The table below is exhaustive about
which tier each rule is in, because the failure this ADR is most exposed to is a rule filed under
"Enforcement" that no validator implements:
| Check | Home |
|---|---|
| description and body counts, resolvable boundary targets | `skill-audit/scripts/validate.sh`, `scripts/skill-size-check.sh` |
| prose patterns (composition-note openers, restatement) | `plugins/kyberforge/.apm/skills/*/assets/vale/styles/Kyberforge/` |
| judgment calls | `references/description-quality.md`, `references/body-discipline.md` |
| Check | Applies to | Tier | Home |
|---|---|---|---|
| description characters (250 SUGGESTION / 400 FAIL) | skills, agents | deterministic | `scripts/skill-size-check.sh`; constants mirrored in `skill-audit/scripts/validate.sh` and `agent-audit/scripts/validate.sh` |
| body-only words (600 SUGGESTION / 900 FAIL) | skills | deterministic | `skill-size-check.sh`, `skill-audit/scripts/validate.sh` |
| description present and non-empty (ERROR) | skills, agents | deterministic | same |
| boundary target resolves to a real skill or agent (ERROR when written as `/name` or `-> name`, or when its own sentence names another target that resolves; SUGGESTION otherwise) | skills, agents | deterministic | same |
| boundary clause absent (SUGGESTION) | skills, agents | deterministic | same |
| Gotchas entry count over five (SUGGESTION) | skills | deterministic | same |
| Gotchas over 25% of the body (SUGGESTION) | skills | deterministic | same |
| every `references/<file>.md` a body names exists (ERROR) | skills | deterministic | same |
| description opener, composition notes in a description | skills, agents | prose pattern | `plugins/kyberforge/.apm/skills/*/assets/vale/styles/Kyberforge/` |
| a Gotcha paraphrasing a body step | skills | **auditor judgment** | `references/body-discipline.md` |
| dispatch at two or more mutually exclusive flows | skills | **auditor judgment** | `references/body-discipline.md` |
| delegation: an agent body restating a skill's procedure | agents | **auditor judgment** | `agent-audit` |
| capability enumeration, restatement, trigger quality | skills, agents | **auditor judgment** | `references/description-quality.md` |
**Blocking immediately, with no baseline file.**
The rows in bold are stated as FAILs in the Decision above and are FAILs an *auditor* issues. None of
them is countable: "does this Gotcha paraphrase step 4", "are these two flows mutually exclusive" and
"does this agent body restate what `git-commits` already owns" are semantic questions, and a script
that guessed at them would be a worse gate than no gate, because it would be believed. They are not
enforced, they are reviewed, and this table exists so that distinction is written down rather than
inferred from whether a validator happens to have been written yet.
Two of the deterministic rows are tuned for **false positives over recall**, and what they decline to
see is part of the contract. On target extraction: a bare hyphenated name counts only inside a
boundary sentence, and a single-word name is never matchable bare — `research`, `triage`, `forge`,
`prototype` and `tdd` are all real skill names *and* ordinary English, so it must be written
`` `forge` `` or `/forge` to be seen at all. Grammar then decides whether a recognised target may
raise an error: one followed by an ordinary lowercase noun is a compound **modifier**, not a route
("use pre-commit hooks instead of ad-hoc scripts", "invoke the pull-request template"), so it is
confirm-only — it still resolves and still counts as a route when the name exists, but it can never
dangle. Only a *terminal* target can. The compressed arrow form `→ <name>` is exempt from that
follower test and is always error-eligible, because nothing reads as a compound modifier after an
arrow; a `/slash` target reached through a route verb is **not** exempt and takes the same test. The
simpler rule — "only marked targets may dangle" — was available and would have been wrong here: both
live true positives are bare, `research`'s "(use neuledge-context)" and the `gitea-labels-` /
`milestones` fold. On the body-shape checks: a `## Gotchas` heading must *end* in "gotchas", not
merely contain the word, so `## Gotcha handling` and `## Why gotchas matter` are prose sections and
are skipped; fenced code blocks are masked out of heading detection and entry counting, so a fenced
example list is not mistaken for the section; and a `references/` pointer named on a line
that also says the file is gone ("removed", "deprecated", "no longer") is read as a historical
mention rather than a dead dispatch entry. Note the 25% fraction is deliberately *not* fence-masked
on either side — fenced lines are real body words, and the fraction is measured against the whole
body.
**The deterministic tier blocks immediately, with no baseline file.**
Three pre-existing contradictions are fixed in the same change, because they are the contract:
- `skill-audit/SKILL.md:58` asks whether the description opens with an action verb ("Audits…",
"Reviews…") while `:101` and `DescriptionOpener.yml` require an imperative "Use when…" opener. The
criterion is unsatisfiable against the house's own skills, both of which open with "Use when".
- `DescriptionOpener.yml`'s regex is anchored to `^This (skill|agent)\b`, so `gitea-workflow` ("This
is the human-facing entry point…") and `gitea-labels-milestones` ("This is a cross-cutting shared
skill…") both violate the rule and pass the linter.
"Reviews…"), while `:56` defers the same question to `Kyberforge.DescriptionOpener` and
`skill-author/SKILL.md:101` requires an imperative "Use when…" opener. The criterion is
unsatisfiable against the house's own skills, both of which open with "Use when".
- `DescriptionOpener.yml` is anchored to `^This (skill|agent)\b`, which misses a plain `This …`
opener; it is widened here to `^This\b`. The anchor itself stays. Composition prose that sits
*mid*-description — `gitea-workflow`'s "This is the human-facing entry point…" at character 377,
`gitea-labels-milestones`'s "This is a cross-cutting shared skill…" at character 300 — was never in
the opener rule's scope and correctly is not: under `scope: text.frontmatter.description` the `^`
anchors to the start of the whole folded value, and un-anchoring to reach mid-description text was
measured at 5 hits and 5 false positives and rejected (`LESSONS.md`, 2026-08-14). The real gap is
that no rule covered that text at all, which a new token-list rule, `Kyberforge.CompositionNote`,
closes: 10 alerts across four `gitea-*` skills, 0 false positives.
- `description-quality.md:45-50` has no FAIL condition for internal-mechanics content, which is why
`skill-author/SKILL.md:102` never bit.
@@ -159,24 +284,35 @@ retrofits kyberforge's own four author/audit skills, so the figures on landing a
With the gate hot and no baseline, a one-line
fix to `gitea-prs` cannot be committed until that skill meets the contract. This is deliberate — it
guarantees convergence and avoids a half-state — but it means the retrofit is lazy and *mandatory*
rather than deferred. The follow-up retrofit issue should be prioritised accordingly, and the risk it
rather than deferred. Issue #99 tracks it and should be prioritised accordingly, and the risk it
carries is the ordinary one for hot gates: a gate expensive enough to be inconvenient gets bypassed
with `SKIP=` and loses its authority.
**A second hot gate ships alongside it, and it is easy to miss.** `Kyberforge.CompositionNote` is
`level: error` like every other rule in that style, so `pre-commit run --all-files` is red on 10
alerts across `gitea-issues`, `gitea-labels-milestones`, `gitea-prs` and `gitea-workflow`
independently of anything `skill-size-check` reports. Someone scoping the #99 retrofit off the size
findings alone will fix those and still be blocked. The two gates want fixing together.
**A ceiling does not produce an average.** If every author writes to the 400-character FAIL, the
preload lands at 15,600 chars — a 34% cut, not the ~50% intended. The halving depends entirely on the
preload lands at 39 × 400 = 15,600 chars — a 33% cut off 23,427, not the ~50% intended. Writing to
the 250-character SUGGESTION instead lands at 9,750, a 58% cut. The halving depends entirely on the
250-character SUGGESTION tier being visible and respected. That tier works here in a way it does not
elsewhere in this repo: `skill-audit` already reports `PASS (N suggestions)` as a first-class
outcome. This is explicitly **not** the failure ADR-0013 records — Vale warnings are invisible
because vale's exit code keys on `error` alone, but these gates live in `validate.sh` and
`skill-audit`, where a SUGGESTION reaches the report. Realistic landing is 34-55% down, not a
guaranteed 50%.
`skill-audit`, where a SUGGESTION reaches the report. Realistic landing is somewhere in that 33-58%
band, not a guaranteed 50%.
**A word gate cannot detect the defect it is standing in for.** `git-commits` carries thirteen
Gotchas of which four restate steps in its own Workflow (`:32` ≡ step 9, `:33` ≡ step 9, `:36` ≡ step
2, `:31` ≡ the description) — 1,217 words that pass any plausible gate. The counts are a backstop to
the dispatch rule and the Gotchas constraint, not a substitute for them, and should not be read as
the mechanism.
**A word gate cannot detect the defect it is standing in for.** `git-commits` carries twelve Gotchas
of which four restate steps in its own Workflow (`:32` ≡ step 9, `:33` ≡ step 9, `:36` ≡ step 2,
`:31` ≡ the description). Its body is 1,102 words and its whole file 1,217, so it does fail the
900-word body FAIL — but for its length, not for the restatement. The four duplicated Gotchas are 114
words between them; delete every one and the file still fails, while a skill 250 words shorter with
the identical defect passes clean. The two properties are uncorrelated, which is why the counts are a
backstop to the dispatch rule and the Gotchas constraint — both of which are auditor judgment for the
semantic half, per the Enforcement table — and not a substitute for them. Reading the word gate as
the mechanism is the specific mistake this paragraph exists to prevent.
**Some skills legitimately need more description budget than others.** A tiered limit keyed to
sibling density was considered and rejected as too clever; the flat 250/400 pair means the `gitea-*`
@@ -184,22 +320,34 @@ and `git-*` families — where every sibling shares a keyword and boundary claus
— are the ones most likely to sit at the FAIL tier permanently. If the retrofit shows that family
routing degrades, the tier is the first thing to revisit.
**Four broken routing targets are live and are not fixed here.** `skill-audit` routes to
`/skill-improve` twice in its description plus `README.md:10`, and no such skill exists — the real
target is `skill-author`. `research` routes to `neuledge-context`, which exists only inside that
string. `agent-author` says "Do not use for read-only review — examine agent files manually",
routing away from `agent-audit`, the correct sibling. `gitea-issues` contains the literal string
`gitea-labels- milestones`, a stray space introduced by YAML folding, breaking the skill name in
preloaded text. The resolvable-target check added here will fail on all four the moment those files
are touched; fixing them is split into its own issue.
**Four broken routing targets were found; two are fixed here and two are live.** Tracked as issue
#100.
- `skill-audit` routed to `/skill-improve` twice in its description plus `README.md:10`, and no such
skill exists — the real target is `skill-author`. **Fixed here**, as a side effect of retrofitting
kyberforge's own skills.
- `agent-author` said "Do not use for read-only review — examine agent files manually", routing away
from `agent-audit`, the correct sibling. **Fixed here**, same way. Note this one was never
detectable by the resolvable-target check and never will be: "examine agent files manually" names
no target, and a check that resolves names cannot see a name that is absent. A misroute to nowhere
is a review finding, not a gate finding.
- `research` routes to `neuledge-context`, which exists only inside that string. **Live.**
- `gitea-issues` carries the literal string `gitea-labels- milestones` in its folded description, a
stray space introduced by YAML wrapping mid-token, breaking the skill name in preloaded text.
**Live** — the check reports it as a dangling `gitea-labels`.
So the check fires on 3 of the 4 against the base commit and on 2 at the tip of this change, and
`tests/test-skill-size-check.sh` probes exactly those three by name rather than asserting a count, so
it degrades to SKIP as #100 lands rather than going stale.
**Duplication between `skill-author` and `agent-author` survives un-gated.** The merge rule
deliberately excludes the author pair, so the commit-verification argument in four near-copies, the
root-cause grouping rule in four copies, and the wholesale clone of the "Improving an existing X"
flow all remain. Cache isolation makes them structurally unavoidable
(`skill-audit/SKILL.md:95` forbids cross-skill references; `LESSONS.md:107` records why), so the
options are a sync gate or continued drift. This is an input to the kyberforge-bodies follow-up
issue, not a solved problem.
options are a sync gate or continued drift. This is an input to issue #101, which carries both halves
of the kyberforge duplication problem — the deferred audit-pair merge and this — not a solved
problem.
**Provenance frontmatter is explicitly out of scope.** `LESSONS.md:63` asserts that non-routing
frontmatter (`source_keys`, `category`, `version`) is loaded at agent startup, which would make the
@@ -211,11 +359,14 @@ the ADR-0009 provenance machinery for no runtime gain. The metadata was added de
## Alternatives considered
Upstream citations below are relative to
`plugins/kyberforge/docs/research/examples/skill-write/`, as in Context above.
- **Keep pushiness, raise the budget to ~500 chars.** Undertriggering is the worse failure mode — a
skill that never fires is worth nothing regardless of cost — and `skill-creator/SKILL.md:67`
explicitly recommends being "pushy" against an observed undertriggering tendency. Rejected because
that claim is an unmeasured assertion about an older model, and because the correctness hazard in
`writing-skills:154-158` cuts the other way: a fat description is not merely expensive, it is a
`writing-skills/SKILL.md:154-158` cuts the other way: a fat description is not merely expensive, it is a
shortcut agents take instead of reading the body. Would have landed a 35% cut.
- **A trigger-eval loop to set lengths empirically.** `skill-creator/SKILL.md:337-404` specifies 20
queries per skill, 8-10 positive and 8-10 near-miss, with a 60/40 train/test split selecting on
@@ -226,8 +377,9 @@ the ADR-0009 provenance machinery for no runtime gain. The metadata was added de
skill's edit fail on account of another skill's growth, and because it is meaningless for an
external consumer installing a subset of the plugins.
- **500-word body FAIL, matching `writing-skills/SKILL.md:217-221`.** Best-grounded in upstream and
would align this repo with the tightest source. Rejected because it fails 30 of 39 skills, and a
blunt gate gets satisfied by deleting content rather than relocating it.
would align this repo with the tightest source. Rejected because it fails 28 of 39 skills body-only
(35 of 39 measured whole-file), and a blunt gate gets satisfied by deleting content rather than
relocating it.
- **A shrinking baseline file** recording each non-compliant skill's current numbers, failing only on
growth. Would have made the retrofit a visible burn-down instead of a wall. Rejected in favour of
hot gates.

View File

@@ -76,6 +76,9 @@ Include what the fresh context lacks:
- A direct role instruction opening the prompt: `You are a [role]. When invoked, [action].`
- One bounded job, stated so the agent knows what it must refuse.
- The dispatch, gates, inputs and outputs listed above.
- **Error handling** — what the agent does on malformed, missing or contradictory input: stop and
report, or degrade to a named fallback. Absent it, the agent invents a recovery, and a
subagent's invented recovery is invisible to its caller until the output is wrong.
- Non-obvious environment facts and project-specific conventions it cannot infer.
- One default per decision point with one escape hatch.
@@ -111,6 +114,8 @@ Flag as FAIL if:
Flag as SUGGESTION if:
- The body does not open with a direct role instruction
- The body specifies no error handling — nothing tells the agent what to do with malformed,
missing or contradictory input
- The job the agent describes is unbounded, or bounded only implicitly
- A rationale is missing from a rule the agent is expected to enforce — present but unexplained
- Comments are useful but verbose enough to bury the field they annotate

View File

@@ -35,6 +35,24 @@ You are a test agent. When invoked, do the thing.
EOF
}
# Helper: a description of EXACTLY <n> characters that carries a boundary
# clause and names no routing target. ADR-0020's missing-boundary-clause
# SUGGESTION fires on any description without one, so a fixture that omits it
# is never "otherwise clean" and a test refuting SUGGESTION would be asserting
# the boundary check's absence instead of the thing it names. The clause is
# paid for out of the measured budget rather than appended to it, because
# these tests measure the description LENGTH. "anything else" is not
# hyphenated, so no routing target comes with it.
desc_of_length() {
python3 - "$1" <<'PY'
import sys
n = int(sys.argv[1])
prefix = 'Use when doing the thing. Do not use for anything else. '
assert n >= len(prefix), 'requested description shorter than the boundary clause'
print(prefix + 'x' * (n - len(prefix)))
PY
}
# Helper: same shape as make_apm_agent, but the description is supplied
# verbatim — used by the ADR-0020 description-budget tests.
make_apm_agent_with_desc() {
@@ -637,7 +655,7 @@ EOF
@test "ADR-0020: agent description of exactly 250 chars raises no suggestion" {
local root="$TMPDIR/pkg"
make_apm_agent_with_desc "$root" "my-agent" "$(python3 -c "print('x' * 250)")"
make_apm_agent_with_desc "$root" "my-agent" "$(desc_of_length 250)"
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
assert_success
refute_output --partial "SUGGESTION"
@@ -645,7 +663,7 @@ EOF
@test "ADR-0020: agent description of 251 chars raises a SUGGESTION and still exits 0" {
local root="$TMPDIR/pkg"
make_apm_agent_with_desc "$root" "my-agent" "$(python3 -c "print('x' * 251)")"
make_apm_agent_with_desc "$root" "my-agent" "$(desc_of_length 251)"
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
assert_success
assert_output --partial "SUGGESTION"
@@ -654,7 +672,7 @@ EOF
@test "ADR-0020: agent description of exactly 400 chars is a SUGGESTION, not a FAIL" {
local root="$TMPDIR/pkg"
make_apm_agent_with_desc "$root" "my-agent" "$(python3 -c "print('x' * 400)")"
make_apm_agent_with_desc "$root" "my-agent" "$(desc_of_length 400)"
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
assert_success
assert_output --partial "SUGGESTION"
@@ -662,7 +680,7 @@ EOF
@test "ADR-0020: agent description of 401 chars FAILs and exits non-zero" {
local root="$TMPDIR/pkg"
make_apm_agent_with_desc "$root" "my-agent" "$(python3 -c "print('x' * 401)")"
make_apm_agent_with_desc "$root" "my-agent" "$(desc_of_length 401)"
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
assert_failure
assert_output --partial "description is 401 chars"
@@ -732,10 +750,18 @@ EOF
# body becomes the system prompt of a fresh context. ADR-0020 gates the
# former at 900 words and explicitly declines to gate the latter. If a body
# word gate is ever added here, it contradicts the ADR.
#
# The description carries a boundary clause so the ONLY thing this test can
# go red on is a body finding. Without one, the missing-boundary-clause
# SUGGESTION fires and the blanket `refute_output --partial "SUGGESTION"`
# below trips for a reason that has nothing to do with body length — which
# would look like the invariant breaking while proving nothing about it.
# AGENTS.md cites this test as the pin for that invariant, so it has to fail
# for one reason and one reason only.
{
echo "---"
echo "name: my-agent"
echo "description: A valid agent description."
echo "description: A valid agent description. Do not use for anything else."
echo "---"
echo ""
python3 -c "print(' '.join(['word'] * 1500))"
@@ -743,7 +769,13 @@ EOF
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
assert_success
refute_output --partial "FAIL"
# A 1,500-word body is 667% of the skill ceiling. Nothing may be said about
# it at any tier: not a FAIL, not a SUGGESTION, and not the word-count
# wording either tier would use if a gate were quietly added later.
refute_output --partial "SUGGESTION"
refute_output --partial "1500 words"
refute_output --partial "900-word"
refute_output --partial "body is"
}
@test "a bare plugin.json with no apm.yml is no longer plugin scope — falls through to project scope" {

View File

@@ -34,10 +34,13 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
boundary clause naming a real sibling skill or agent.
250 characters is the target, 400 the hard ceiling (ADR-0020).
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was
deleted. Add "Use proactively" only if the runtime should delegate here without
the user naming this agent.
deleted.
Never write "Use proactively" here. It steers the Claude Code runtime and does
nothing anywhere else, and this file compiles to a Copilot `.agent.md` too, where
agent-audit's KyberforgeCopilot.ProactivePhrase rule grades it a hard FAIL.
The phrase is CC-only; at this scope, a precise trigger clause does that job.
Example: "Use when a diff needs checking for injected credentials before it
merges. Not general code review -> code-reviewer." -->
merges. Not prose or style linting -> `lint-runner`." -->
<!-- model: sonnet
Optional. Aliases: sonnet, opus, haiku, fable. Or full model ID.
@@ -80,3 +83,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
## Output
FILL IN: What does the agent produce? Format, location, structure.
## Errors
FILL IN: What does the agent do on malformed, missing or contradictory input?
State whether it stops and reports, or degrades to a named fallback — and what it
tells the caller either way. An agent with no error handling invents a recovery,
and an invented recovery is invisible until the output is wrong.

View File

@@ -14,10 +14,14 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
boundary clause naming a real sibling skill or agent.
250 characters is the target, 400 the hard ceiling (ADR-0020).
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was
deleted. Add "Use proactively" only if the runtime should delegate here without
the user naming this agent.
deleted.
"Use proactively" is valid HERE and only here: it steers the Claude Code runtime
to offer this agent unprompted. Add it only if that is what you want. If you add
it, leave it OUT of the Copilot half of the pair — the phrase does nothing there
and agent-audit's KyberforgeCopilot.ProactivePhrase grades it a hard FAIL. The
pair must describe the same job; it does not have to be byte-identical.
Example: "Use when a diff needs checking for injected credentials before it
merges. Not general code review -> code-reviewer." -->
merges. Not prose or style linting -> `lint-runner`." -->
<!-- tools: Read, Bash, Grep
Optional. Allowlist of tool names: a comma-separated string or a YAML list.
@@ -99,3 +103,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
## Output
FILL IN: What does the agent produce? Format, location, structure.
## Errors
FILL IN: What does the agent do on malformed, missing or contradictory input?
State whether it stops and reports, or degrades to a named fallback — and what it
tells the caller either way. An agent with no error handling invents a recovery,
and an invented recovery is invisible until the output is wrong.

View File

@@ -19,9 +19,13 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
boundary clause naming a real sibling skill or agent.
250 characters is the target, 400 the hard ceiling (ADR-0020).
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was deleted.
Keep it identical in wording to the Claude Code half of the pair.
Never write "Use proactively" here. It steers the Claude Code runtime and does nothing
in Copilot, and agent-audit's KyberforgeCopilot.ProactivePhrase grades it a hard FAIL.
Otherwise keep the wording matched to the Claude Code half of the pair: agent-audit
checks that both halves describe the same job, not that they are byte-identical, so
dropping the CC-only phrase here is not a pair-consistency finding.
Example: "Use when a diff needs checking for injected credentials before it
merges. Not general code review -> code-reviewer." -->
merges. Not prose or style linting -> `lint-runner`." -->
<!-- tools: ["read", "search", "edit"]
Optional. Array of tool names. Omit = all available tools. [] = no tools.
@@ -68,3 +72,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
## Output
FILL IN: What does the agent produce? Format, location, structure.
## Errors
FILL IN: What does the agent do on malformed, missing or contradictory input?
State whether it stops and reports, or degrades to a named fallback — and what it
tells the caller either way. An agent with no error handling invents a recovery,
and an invented recovery is invisible until the output is wrong.

View File

@@ -44,15 +44,31 @@ Banned from a description; move it to the body or to `README.md`:
and ADR-0020 deleted it: the opener is `Use when`, matching every skill in this corpus, so one
router reads one shape.
**"Use proactively" is conditional.** Add it only where the runtime should delegate without the
user naming the agent — an agent invoked by name does not need it, and it costs activations
elsewhere when added by reflex. The same conditional governs indirect triggers ("even if the user
doesn't say X"): add one only where the user's natural phrasing genuinely omits the domain word.
**"Use proactively" is Claude Code-only, and conditional even there.** The phrase steers the
Claude Code runtime to offer an agent unprompted and does nothing anywhere else, so where it may
appear depends on the file:
**Boundary targets must resolve.** The name after the arrow is checked against real skills under
`plugins/*/.apm/skills/<name>/` and real agents under `plugins/*/.apm/agents/<name>.agent.md`. A
target that does not exist sends the router nowhere. Verify it before writing it — do not invent a
plausible sibling.
| File | Rule |
|---|---|
| Claude Code `.md` (project/user scope) | Allowed. Add it only where the runtime should delegate without the user naming the agent — an agent invoked by name does not need it, and it costs activations elsewhere when added by reflex. |
| Copilot `.agent.md` (project/user scope) | **Never.** Inert there, and `KyberforgeCopilot.ProactivePhrase` grades it a hard FAIL. |
| Vendor-neutral `.apm/agents/<name>.agent.md` (plugin/APM scope) | **Never.** Same Vale rule, same hard FAIL — the file matches the `**/*.agent.md` glob, and it compiles to a real Copilot agent downstream. |
A pair whose Claude Code half carries the phrase and whose Copilot half omits it is correct, not
inconsistent: `agent-audit` checks that both halves describe the same job, not that they match
word for word.
Indirect triggers ("even if the user doesn't say X") take a similar conditional at every scope:
add one only where the user's natural phrasing genuinely omits the domain word.
**Boundary targets must resolve.** Both forms are checked — the arrow and the prose form ("do not
use for X, use `y` instead") — so a typo dangles either way. Targets resolve against a universe
built by walking up **from the agent file itself**: the nearest ancestor holding
`plugins/*/.apm/{skills,agents}` (or, failing that, the nearest ancestor holding `.git`) contributes
every skill and agent under `<root>/plugins/*/`, plus the agent's own apm package and the packages
that package declares in `apm.yml` under `dependencies.apm`. A sibling plugin in the same monorepo
therefore resolves; a skill in an unrelated repo does not. A target outside that universe sends the
router nowhere. Verify it before writing it — do not invent a plausible sibling.
**Length.** 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only
with YAML folding resolved. Treat 250 as the target: the SUGGESTION tier is what moves the corpus
@@ -73,8 +89,18 @@ You are a <role>. When invoked, <primary action>.
## Output
<what it produces: format, location, structure>
## Errors
<what to do on malformed, missing or contradictory input: report and stop, or
which fallback to take — and what to say to the caller either way>
````
Four required elements: **inputs expected, process steps, output format, error handling.** The
last is the one that gets dropped, and dropping it is not neutral: an agent given a malformed
input and no instruction invents a recovery, and a subagent's invented recovery is invisible to
the caller until the output is wrong. Say explicitly whether the agent stops and reports, or
degrades to a named fallback.
One job per agent. An agent covering two jobs gets delegated to for the wrong one.
**Delegation discipline replaces the word gate.** A plugin/APM agent is a single file with no

View File

@@ -63,5 +63,7 @@ to installed skills instead of transcribed procedure.
- [ ] No `FILL IN:` placeholder and no `<!-- ... -->` template comment anywhere in the file
- [ ] System prompt body non-empty, and a read-only agent says so in prose as well as in
`disallowedTools`
- [ ] Body covers all four required elements: inputs expected, process steps, output format,
**error handling** — what the agent does on malformed, missing or contradictory input
Then return to the flow reference you came from.

View File

@@ -95,6 +95,8 @@ Both files:
- [ ] `name` present and kebab-case; `description` written to `references/contract.md`
- [ ] System prompt body present, non-empty and equivalent across the pair
- [ ] Body covers all four required elements: inputs expected, process steps, output format,
**error handling** — what the agent does on malformed, missing or contradictory input
- [ ] No `FILL IN:` placeholder and no `<!-- ... -->` template comment left
Copilot file only:

View File

@@ -285,6 +285,8 @@ else
if [[ "$SCOPE" == "plugin" ]]; then
echo " 1. Fill in $APM_FILE — replace every FILL IN: placeholder. Optional fields are" >&2
echo " scaffolded there as commented blocks; uncomment the ones that apply." >&2
echo " Description: 250 chars target / 400 ceiling (ADR-0020). The body has no" >&2
echo " word gate — delegate to a skill instead of restating what it does." >&2
echo " 2. Populate $SOURCES_DIR/sources.md with research sources, or delete it" >&2
echo " 3. Validate: $VALIDATE_HINT $APM_FILE" >&2
echo " It checks the frontmatter against the apm-agent-allowlist section of" >&2
@@ -292,6 +294,8 @@ else
else
echo " 1. Fill in $CC_FILE — replace every FILL IN: placeholder. Optional fields are" >&2
echo " scaffolded there as commented blocks; uncomment the ones that apply." >&2
echo " Description: 250 chars target / 400 ceiling (ADR-0020). The body has no" >&2
echo " word gate — delegate to a skill instead of restating what it does." >&2
echo " 2. Fill in $CP_FILE — same, and heed its closing comment: the Claude Code-only" >&2
echo " fields it names must not cross over from the file above." >&2
echo " 3. Validate: run $VALIDATE_HINT on each file" >&2

View File

@@ -6,10 +6,12 @@ Audit a skill directory against the agentskills.io specification and the house c
1. Runs `scripts/validate.sh` and `scripts/validate-provenance.sh` for structural and provenance checks, plus `scripts/vale-wrap.sh` — a Vale prefilter that deterministically flags non-imperative description openers, composition and architecture notes, vague wording, padding phrases, and "There is/are" sentence openers
2. Reads all files in the skill directory
3. Applies qualitative checks across six dimension groups, loading one rubric from `references/` per group
3. Applies qualitative checks across five dimension groups, loading one rubric from `references/` per group
4. Outputs a compact findings report — findings only, grouped by dimension, each with Why and Fix — and a result block with handoff to `skill-author`
`validate.sh` enforces two independent length families that must not be conflated: the agentskills.io spec conformance ceilings (500 lines, 2,770 words, both counting the whole file) and the ADR-0020 context budget (250/400 description characters, 600/900 body-only words, plus resolvable boundary targets).
`validate.sh` enforces two independent length families that must not be conflated: the agentskills.io spec conformance ceilings (500 lines, 2,770 words, both counting the whole file) and the ADR-0020 context budget (250/400 description characters, 600/900 body-only words).
Alongside those it runs four shape checks that are not length measurements at all. Two are FAILs: every routing target named in the description — in the compressed `Not <thing> -> <name>` arrow **and** in the prose form — must resolve to a real skill or agent, and every `references/<file>.md` the body names must exist on disk. Three are SUGGESTIONs: a missing boundary clause, a Gotchas section over five entries, and a Gotchas section over 25% of the body. The resolution universe for boundary targets is derived by walking up from the audited `SKILL.md` — the authoring root above it, its own apm package, and that package's declared `apm.yml` dependencies — so a fresh clone and a machine that has run `apm install` return the same verdict. When no universe can be determined the check prints `INFO ... DID NOT RUN` and does not silently pass.
## Usage
@@ -24,7 +26,7 @@ Provide the path to the skill directory to audit when invoking.
| File | Purpose |
|------|---------|
| `SKILL.md` | Skill instructions for agents |
| `scripts/validate.sh` | Structural validator — checks name format, name matches directory, description length, line count, placeholder detection, script executable bit, and interactive-prompt detection |
| `scripts/validate.sh` | Structural validator — checks name format, name matches directory, description presence and length, body-only word count, line and whole-file word ceilings, boundary-clause presence, boundary-target resolution, `references/` pointer existence, Gotchas entry count and body share, placeholder detection, script executable bit, and interactive-prompt detection |
| `scripts/validate-provenance.sh` | Provenance validator — checks sources.md completeness, source_keys/slug consistency, Contributing files existence, bidirectional linkage, Research doc: fields, and upstream research doc alignment |
| `scripts/vale-wrap.sh` | Vale prefilter wrapper — runs the bundled `Kyberforge` Vale styles against SKILL.md and reports alerts as deterministic FAILs ahead of Step 3's qualitative review |
| `assets/vale/.vale.ini` | Vale configuration — points Vale at the bundled `Kyberforge` style path, self-located relative to `vale-wrap.sh` |
@@ -38,6 +40,7 @@ Provide the path to the skill directory to audit when invoking.
| `references/patterns.md` | Rubric for the patterns dimension — which instruction construct fits which job, and how each is correctly formed |
| `references/file-structure.md` | Rubric for the file-structure and internal-consistency dimensions — permitted directories, cross-plugin path rules and their two structural exemptions, README drift |
| `references/formatting-and-scripts.md` | Rubric for the formatting and scripts dimensions — heading and fencing conventions, and the agentic-use criteria for bundled scripts |
| `references/validation-scripts.md` | Step 1 troubleshooting — the manual structural fallback when `validate.sh` cannot run, and the script exit codes that are easy to misread (loaded only on a script failure) |
| `references/sources.md` | Provenance record — agentskills.io sources that informed this skill and which files each contributed to |
| `tests/validate.bats` | (source-only) Bats test suite for validate.sh |
| `tests/validate-provenance.bats` | (source-only) Bats test suite for validate-provenance.sh |

View File

@@ -33,7 +33,9 @@ bash scripts/validate-provenance.sh <skill-dir>
scripts/vale-wrap.sh <skill-dir>/SKILL.md
```
`validate.sh` findings become the `### Structure` dimension — its FAILs and its SUGGESTIONs both. If it cannot run at all (no `python3`, Bash denied), report that as an INFO finding rather than guessing; what it measures is not reproducible by reading.
`validate.sh` findings become the `### Structure` dimension — its FAILs and its SUGGESTIONs both.
If any of the three fails, cannot run, or reports something needing interpretation, read `references/validation-scripts.md` — it carries the manual fallback and the misleading exit codes.
`validate-provenance.sh` prints nothing on success. Its FAIL and INFO findings become a separate `### Provenance` dimension, and it emits Why and Fix itself — surface those verbatim.

View File

@@ -69,8 +69,10 @@ table** plus the gates common to every branch, and each flow lives in its own se
`references/` file. Inlining all of them is a FAIL regardless of word count, because every
invocation then pays for every branch it did not take.
The reference shape in this repo is `apm-workflow`: a 554-word body dispatching to roughly 3,000
words of references across five mutually exclusive invocations.
The reference shape in this repo is `apm-workflow`: a **421-word body** dispatching to roughly
3,000 words of references across five mutually exclusive invocations. Its whole-file count is 554
words — cite 421 when calibrating a body, or the conflation this section warns against reappears
in the finding itself.
## Gotchas sections
@@ -86,17 +88,20 @@ sensibly.
Constraints:
- **Maximum five entries.** Past five, the section is a summary of the body rather than a set of
traps, and the agent stops reading it as a warning.
- **More than five entries is a SUGGESTION** — five is the guideline, not a ceiling. Past five, the
section is usually a summary of the body rather than a set of traps, and the agent stops reading
it as a warning. It stays advisory because whether a given gotcha earns its place is judgment;
`validate.sh` emits it through `suggest()` and the run still exits 0.
- **A Gotcha that paraphrases a step in the body below it is a FAIL.** It has no independent
content, and it teaches the agent that Gotchas can be skimmed because the real instruction is
coming.
coming. This one is the auditor's call — no script detects it.
- **A Gotchas section exceeding 25% of the body is a SUGGESTION** — the body has been inverted into
a preamble.
a preamble. Same tier and same reasoning as the entry count, and independent of it: either can
fire without the other.
- Place the section near the top. A gotcha read after the mistake is worthless, which is also why
Gotchas is the one construct exempt from moving to `references/`.
Worked negative example — `git-commits` carries thirteen entries, of which four restate content
Worked negative example — `git-commits` carries twelve entries, of which four restate content
that already appears below or in the description:
| Gotcha | Restates |
@@ -106,8 +111,10 @@ that already appears below or in the description:
| `:33` "Never skip hooks with `--no-verify`" | step 9 at `:52` |
| `:36` "Never commit secrets" | step 2 at `:45` |
All four are FAILs under this rule, and the section as a whole breaches the five-entry maximum. It
also passes every plausible word gate, which is the point of auditing the construct directly.
All four are FAILs under the paraphrase rule. The entry count and the section's share of the body
(387 of 1,102 words, 35%) are two further SUGGESTIONs on top — the script reports both, and neither
fails the run on its own. What makes this worth auditing directly is that the four paraphrase FAILs
pass every word gate there is; only reading the construct finds them.
## Calibrating control
@@ -144,7 +151,7 @@ Flag as FAIL if:
- A sentence answers "no" to the core test — it is padding
- The body exceeds 900 words counted body-only (`validate.sh` reports it)
- Two or more mutually exclusive flows are inlined instead of dispatched
- A Gotcha paraphrases a step in the body below it, or the section exceeds five entries
- A Gotcha paraphrases a step in the body below it
- A decision point presents a menu of options with no default
- An instruction repeats content already in the description
- A prescriptive sequence is used where flexibility is fine, or the reverse
@@ -152,6 +159,7 @@ Flag as FAIL if:
Flag as SUGGESTION if:
- The body exceeds 600 words counted body-only but stays at or under 900
- The Gotchas section carries more than five entries
- The Gotchas section exceeds 25% of the body
- A rationale is missing from an include/exclude rule — present but unexplained
- Gotchas are correct but placed late in the body rather than near the top

View File

@@ -28,6 +28,14 @@ resolving there. Flag any `../`, `../../`, or absolute repo path (`plugins/<plug
and its APM-native equivalent `.apm/skills/<other>/`) appearing in `SKILL.md`, `scripts/`,
`references/` or `assets/`.
**Referring to another skill's file.** There is one sanctioned spelling, and it is possessive:
`skill-audit's references/validation-scripts.md`. Write the skill by name and let the reader
resolve it — do not spell the repo path. The full path is the thing this section forbids, and
`references/validation-scripts.md` on its own is a hard ERROR from the ADR-0020 gate, which
requires an unqualified `references/` pointer to exist in the skill's OWN directory. The
possessive form is the only spelling both rules accept; the gate recognises it and skips the
on-disk check. Flag any other spelling of a cross-skill reference.
Two directories are exempt, and the exemptions are structural rather than discretionary:
- **`references/sources.md`.** Its `Research doc:` fields are development-time provenance pointers,

View File

@@ -15,7 +15,7 @@
- **URL:** https://agentskills.io/specification.md
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
- **Description:** Complete SKILL.md format specification — frontmatter fields, constraints, body content, optional directories, progressive disclosure levels, file references, validation
- **Contributing files:** SKILL.md, references/body-discipline.md, references/description-quality.md, references/patterns.md, references/file-structure.md, references/formatting-and-scripts.md
- **Contributing files:** SKILL.md, references/body-discipline.md, references/description-quality.md, references/patterns.md, references/file-structure.md, references/formatting-and-scripts.md, references/validation-scripts.md
- **Status:** `extracted`
## agentskills-best-practices
@@ -47,7 +47,7 @@
- **URL:** https://agentskills.io/skill-creation/using-scripts.md
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
- **Description:** Using scripts in skills — one-off commands, self-contained scripts with inline dependencies, designing scripts for agentic use (no interactive prompts, --help, structured output, idempotency)
- **Contributing files:** SKILL.md, references/formatting-and-scripts.md
- **Contributing files:** SKILL.md, references/formatting-and-scripts.md, references/validation-scripts.md
- **Status:** `extracted`
## agentskills-quickstart

View File

@@ -0,0 +1,118 @@
---
source_keys:
- agentskills-spec
- agentskills-using-scripts
---
# Validation Scripts Reference
Read this when a Step 1 script fails, cannot run, or reports something that needs interpreting.
Nothing here is needed on a clean run.
## Report the gap, do not guess
If a script cannot run at all — Bash denied, `python3` unavailable, PyYAML not importable, `vale`
not installed — say so as an **INFO** finding naming the script and the missing dependency, then
fall back to the manual checks below. An INFO never changes PASS/FAIL. Silently omitting the
dimension a script would have covered reports a clean audit that checked less than it claims to
have checked, and the Step 4 coverage line then names a dimension nothing actually examined.
## Manual structural fallback
`validate.sh` needs `python3` **and** PyYAML, and refuses to start without either — the description
value has to be measured after YAML folding is resolved, so skipping the ADR-0020 gates would be a
vacuous pass rather than a partial one. The two are checked separately, so the message already names
the right one — report it verbatim rather than diagnosing further:
```text
Error: python3 is required but was not found on PATH.
Error: PyYAML is required but is not importable by python3.
```
Without them — or with Bash denied, or on a permission error — work this list
by hand and file the results under `### Structure` exactly as the script's output would have been:
- **`name`** present, 1–64 characters, kebab-case (lowercase letters, digits and hyphens; no
leading, trailing or doubled hyphen), and **matching the skill's directory name** exactly.
- **`description`** present and non-empty; no unfilled `FILL IN:` placeholder in it. An absent or
empty description is a **FAIL**, never a silent skip — it is the one field preloaded into every
session, so a skill without one can never be routed to.
- **Description length**, measured on the folded YAML value with newlines collapsed to single
spaces — not on the raw block scalar, which counts indentation. 250 characters SUGGESTION, 400
FAIL (ADR-0020), 1,024 FAIL (agentskills.io spec).
- **Body length**, counting everything after the frontmatter's closing `---`. 600 words
SUGGESTION, 900 FAIL (ADR-0020).
- **Whole-file ceilings**, counting the file including frontmatter: 500 lines FAIL, 2,770 words
FAIL (agentskills.io spec). These are a different measurement from the two above — report them
as separate findings, never merged.
- **A boundary clause is present** — either the prose form (`do not` / `instead` / `rather than` /
`not for`) or ADR-0020's compressed `Not <thing> -> <name>` arrow. **SUGGESTION**, not FAIL:
the absence is deterministic, but whether this skill warrants one is the auditor's call.
- **Boundary targets resolve** — **FAIL** on a name that resolves to nothing. See the section
below; resolving these by hand is the one item on this list with a procedure of its own.
- **Every `references/<file>.md` named in the body exists on disk** — **FAIL**, not a suggestion.
A dispatch table or "read X" trigger naming a missing file sends the agent nowhere. Ignore
mentions inside fenced code blocks, and ignore a mention whose own line says the file is gone
(`removed`, `deleted`, `renamed`, `superseded`, `replaced`, `obsolete`, `deprecated`, `former`,
`gone`, `no longer`, `used to`) — that is a historical note, not a dispatch entry.
- **Gotchas discipline**, both **SUGGESTION**. Locate the section by a heading that *is* Gotchas
(`## Common Gotchas` counts; `## Gotcha handling` and `## Why gotchas matter` do not), running to
the next heading at the same level or shallower. More than five top-level entries is one
suggestion; a section over 25% of the body word count is a second, independent one. Count
entries at column 0 only — an indented child bullet is not an entry — and ignore fenced code
blocks for both.
- **No unfilled `FILL IN:` placeholder** anywhere in the body.
- **Every file in `scripts/`** carries the executable bit and contains no interactive prompt —
no bare `read`, no `select`, nothing that blocks on a TTY.
## Resolving boundary targets by hand
Targets are read from **both** boundary forms. The compressed `Not <thing> -> <name>` arrow and the
prose form are each parsed *and* target-checked, so a typo in prose phrasing fails exactly as an
arrow typo does — do not check only the names after an arrow.
Build the universe by walking up **from the `SKILL.md` under audit**, never from the validator's own
location. The nearest ancestor holding `plugins/*/.apm/skills/` or `plugins/*/.apm/agents/` is the
authoring root, falling back to the nearest ancestor holding `.git`. When one is found the universe
is every skill and agent under `<root>/plugins/*/`, plus the skill's own apm package, plus the
packages that package declares in its `apm.yml` under `dependencies.apm`. Deployed `.claude/` and
`.agents/` trees are consulted **only** when no authoring root exists — they are gitignored
`apm install` output, and reading them would make a fresh clone and a developer machine disagree.
Three ways to read the result wrong:
- **A hyphenated name used attributively is not a dangling target.** "Use pre-commit hooks instead
of ad-hoc scripts" reads as a route to `pre-commit` on wording alone. What separates a route from
prose is grammar: a route target is terminal — followed by punctuation, a conjunction, or a
boundary word — whereas a compound modifier is followed by the noun it modifies. A name followed
by an ordinary noun still *confirms* a route when it exists, but never raises a FAIL on its own.
- **A SUGGESTION-tier unresolved target is not a FAIL you may promote.** Terminal position alone is
not evidence of a route: "run `pre-commit` instead", "see `commit-msg`" and "use the clean-up
instead" are all terminal and all prose. A prose-form target earns a FAIL only when its own
sentence names another target that *does* resolve; otherwise the script reports it and moves on,
and so should you. Route notation — `/name` and `-> name` — is exempt and always FAILs, and it is
the fix to recommend when the author did mean a route.
- **`INFO boundary-target resolution DID NOT RUN` is not a pass.** The script prints it, and exits
0, when no universe could be determined for that path — the usual cause being a skill copy
audited outside its package. Report it as an INFO naming the unchecked targets and re-run against
the real directory; filing it as clean signs off targets nothing verified.
## Script-specific failures
- **`validate-provenance.sh` printed nothing.** That is a pass, not a skip. It also exits 0
silently when the skill has no `source_keys` and no `references/sources.md` — nothing to
validate is not a finding.
- **`vale` reports `0 files`.** Treat the pass as NOT RUN, not as clean, and fall back to full
Step 3 judgment for the dimensions it would have covered. The bundled `Kyberforge` style is
scoped by glob in `assets/vale/.vale.ini`; a file outside those globs is silently not linted.
- **`E100 Runtime error ... does not exist` (exit 2) from `vale-wrap.sh`.** An explicit relative
`--config` was passed. Pass none: the wrapper locates its own `assets/vale/.vale.ini` from its
own path, so a resolved script path plus an unresolved config path produces exactly this. Do not
read this exit code as vale being unavailable — that misreading sends the audit down the
fallback path while vale was installed and working the whole time.
- **The `vale` binary is genuinely absent** (`command not found`). Report one INFO naming it, then
fall back to full Step 3 judgment for the description, body-discipline and patterns dimensions —
the prefilter's whole coverage. Judge those by rubric rather than dropping them.
- **A path argument that does not exist is a hard error** in `vale-wrap.sh`, deliberately: bare
`vale` would fall back to reading stdin and print a clean-looking `0 errors ... in stdin`, which
the `0 files` guard above does not catch.

File diff suppressed because it is too large Load Diff

View File

@@ -8,7 +8,14 @@ setup() {
SCRIPT="$(cd "$BATS_TEST_DIRNAME/../scripts" && pwd)/validate.sh"
TMPDIR="$(mktemp -d)"
# Helper: create a minimal valid skill directory
# Helper: create a minimal valid skill directory.
#
# The description carries a boundary clause deliberately. ADR-0020's
# missing-boundary-clause SUGGESTION fires on any description without one, so
# a fixture that omits it is never "otherwise clean" — every test asserting
# SUGGESTION-freedom would be asserting the boundary check's absence instead
# of the thing it names. "anything else" is not hyphenated, so the clause adds
# a boundary marker without adding a routing target to resolve.
make_valid_skill() {
local dir="$1"
local name
@@ -17,7 +24,7 @@ setup() {
cat > "$dir/SKILL.md" <<EOF
---
name: $name
description: A valid skill description that is well within the limit.
description: A valid skill description that is well within the limit. Do not use for anything else.
---
## Step 1
@@ -26,6 +33,20 @@ Do the thing.
EOF
}
# Helper: a description of EXACTLY <n> characters that carries a boundary
# clause and names no routing target. The tests below measure the description
# LENGTH, so the clause has to be paid for out of the same budget rather than
# appended to it — hence the padding arithmetic instead of a fixed suffix.
desc_of_length() {
python3 - "$1" <<'PY'
import sys
n = int(sys.argv[1])
prefix = 'Use when doing the thing. Do not use for anything else. '
assert n >= len(prefix), 'requested description shorter than the boundary clause'
print(prefix + 'x' * (n - len(prefix)))
PY
}
# Helper: create a skill directory with an exact description length and an
# exact body word count. <desc> is used verbatim; <body_words> "word"
# tokens follow the frontmatter. Used by the ADR-0020 boundary tests.
@@ -300,7 +321,7 @@ EOF
@test "ADR-0020: description of exactly 250 chars raises no suggestion" {
local skill="$TMPDIR/my-skill"
make_sized_skill "$skill" "$(python3 -c "print('x' * 250)")" 10
make_sized_skill "$skill" "$(desc_of_length 250)" 10
run bash "$SCRIPT" "$skill"
assert_success
refute_output --partial "SUGGESTION"
@@ -308,7 +329,7 @@ EOF
@test "ADR-0020: description of 251 chars raises a SUGGESTION and still exits 0" {
local skill="$TMPDIR/my-skill"
make_sized_skill "$skill" "$(python3 -c "print('x' * 251)")" 10
make_sized_skill "$skill" "$(desc_of_length 251)" 10
run bash "$SCRIPT" "$skill"
assert_success
assert_output --partial "SUGGESTION"
@@ -318,7 +339,7 @@ EOF
@test "ADR-0020: description of exactly 400 chars is a SUGGESTION, not a FAIL" {
local skill="$TMPDIR/my-skill"
make_sized_skill "$skill" "$(python3 -c "print('x' * 400)")" 10
make_sized_skill "$skill" "$(desc_of_length 400)" 10
run bash "$SCRIPT" "$skill"
assert_success
assert_output --partial "SUGGESTION"
@@ -326,7 +347,7 @@ EOF
@test "ADR-0020: description of 401 chars FAILs and exits non-zero" {
local skill="$TMPDIR/my-skill"
make_sized_skill "$skill" "$(python3 -c "print('x' * 401)")" 10
make_sized_skill "$skill" "$(desc_of_length 401)" 10
run bash "$SCRIPT" "$skill"
assert_failure
assert_output --partial "description is 401 chars"
@@ -364,7 +385,7 @@ EOF
@test "ADR-0020: body of exactly 600 words raises no suggestion" {
local skill="$TMPDIR/my-skill"
make_sized_skill "$skill" "A short valid description." 600
make_sized_skill "$skill" "A short valid description. Do not use for anything else." 600
run bash "$SCRIPT" "$skill"
assert_success
refute_output --partial "SUGGESTION"
@@ -372,7 +393,7 @@ EOF
@test "ADR-0020: body of 601 words raises a SUGGESTION and still exits 0" {
local skill="$TMPDIR/my-skill"
make_sized_skill "$skill" "A short valid description." 601
make_sized_skill "$skill" "A short valid description. Do not use for anything else." 601
run bash "$SCRIPT" "$skill"
assert_success
assert_output --partial "body is 601 words"
@@ -381,7 +402,7 @@ EOF
@test "ADR-0020: body of exactly 900 words is a SUGGESTION, not a FAIL" {
local skill="$TMPDIR/my-skill"
make_sized_skill "$skill" "A short valid description." 900
make_sized_skill "$skill" "A short valid description. Do not use for anything else." 900
run bash "$SCRIPT" "$skill"
assert_success
assert_output --partial "body is 900 words"
@@ -389,7 +410,7 @@ EOF
@test "ADR-0020: body of 901 words FAILs and exits non-zero" {
local skill="$TMPDIR/my-skill"
make_sized_skill "$skill" "A short valid description." 901
make_sized_skill "$skill" "A short valid description. Do not use for anything else." 901
run bash "$SCRIPT" "$skill"
assert_failure
assert_output --partial "body is 901 words"
@@ -425,15 +446,33 @@ EOF
assert_output --partial "boundary target(s) resolve"
}
@test "ADR-0020: a boundary target naming a non-existent skill FAILs" {
@test "ADR-0020: a boundary target naming a non-existent skill FAILs when its sentence names one that resolves" {
local skill
skill="$(make_fixture_tree "$TMPDIR/tree" "my-skill")"
make_sized_skill "$skill" "Use when doing the thing. Do not use for the other thing — use fixture-missing-skill instead." 10
# `fixture-sibling-skill` is the corroborator: a prose-form target only earns
# a FAIL when its own sentence proves it is a routing sentence. See the
# shared resolver's CORROBORATION note, and the uncorroborated case below.
make_sized_skill "$skill" "Use when doing the thing. Do not use for the other thing — use fixture-sibling-skill or fixture-missing-skill instead." 10
run bash "$SCRIPT" "$skill"
assert_failure
assert_output --partial "routes to 'fixture-missing-skill'"
}
@test "ADR-0020: a LONE boundary target naming a non-existent skill is a SUGGESTION, not a FAIL" {
local skill
skill="$(make_fixture_tree "$TMPDIR/tree" "my-skill")"
# Same grammar as the case above and as "run \`pre-commit\` instead" — a
# route verb, a hyphenated name, terminal position. Nothing local separates a
# broken route from a tool name, so the target is named on every run but does
# not block: this gate ships with no baseline and no suppression mechanism.
make_sized_skill "$skill" "Use when doing the thing. Do not use for the other thing — use fixture-missing-skill instead." 10
run bash "$SCRIPT" "$skill"
assert_success
assert_output --partial "SUGGESTION"
assert_output --partial "routes to 'fixture-missing-skill'"
refute_output --partial "FAIL description routes to"
}
@test "ADR-0020: a boundary target naming an AGENT file resolves (agents are valid routing targets)" {
local skill
skill="$(make_fixture_tree "$TMPDIR/tree" "my-skill")"
@@ -452,15 +491,27 @@ EOF
assert_output --partial "routes to 'fixture-missing-improve'"
}
@test "ADR-0020: a backticked name that does not resolve FAILs" {
@test "ADR-0020: a backticked name that does not resolve FAILs when its sentence names one that resolves" {
local skill
skill="$(make_fixture_tree "$TMPDIR/tree" "my-skill")"
make_sized_skill "$skill" "Use when doing the thing. Composes \`fixture-missing-helper\` for the shared part." 10
make_sized_skill "$skill" "Use when doing the thing. Composes \`fixture-sibling-skill\` and \`fixture-missing-helper\` for the shared part." 10
run bash "$SCRIPT" "$skill"
assert_failure
assert_output --partial "routes to 'fixture-missing-helper'"
}
@test "ADR-0020: a /slash-command target is route NOTATION and FAILs on its own, uncorroborated" {
local skill
skill="$(make_fixture_tree "$TMPDIR/tree" "my-skill")"
# The escape hatch from the SUGGESTION tier: `/name` and `-> name` are never
# how prose cites a tool, so they are exempt from corroboration. An author
# who wants a route checked unconditionally writes one of those two forms.
make_sized_skill "$skill" "Use when doing the thing. Do not use for the other thing — use /fixture-missing-notation instead." 10
run bash "$SCRIPT" "$skill"
assert_failure
assert_output --partial "routes to 'fixture-missing-notation'"
}
@test "ADR-0020: a bare hyphenated word outside a boundary sentence is not read as a routing target" {
local skill
skill="$(make_fixture_tree "$TMPDIR/tree" "my-skill")"
@@ -502,9 +553,18 @@ EOF
}
@test "ADR-0020: the boundary check declines rather than false-FAILs when no authoring source is found" {
# Deliberately NOT built with make_fixture_tree: this skill sits in a bare
# temp directory with no plugins/*/.apm/ above it and no .git, so the resolver
# legitimately has no universe. That is a real path (a skill being drafted
# outside any repo), and the required behaviour is to DECLINE OUT LOUD rather
# than either false-FAIL or pass in silence — silence is what let a whole gate
# family go missing unnoticed. So the INFO text and the named unchecked target
# are both asserted, not just the absence of a failure.
local skill="$TMPDIR/orphan/my-skill"
make_sized_skill "$skill" "Use when doing the thing. Do not use for the other thing — use some-other-skill instead." 10
run bash "$SCRIPT" "$skill"
assert_success
refute_output --partial "routes to"
assert_output --partial "boundary-target resolution DID NOT RUN"
assert_output --partial "Unchecked target(s): some-other-skill"
}

View File

@@ -47,6 +47,7 @@ If the destination resolves inside an APM package, read `references/deployment-m
| `references/create.md` | The create flow end to end — prerequisites, package-intent gate, scaffold, frontmatter, scripts, references, sources (loaded on demand) |
| `references/improve.md` | The improve flow end to end — signal verification, root-cause grouping, announcement, edits (loaded on demand) |
| `references/contract.md` | The ADR-0020 description and body contract, the Gotchas constraint, the two size gates, body patterns, and org-policy embedding (loaded on demand) |
| `references/retrofit.md` | Bringing a pre-ADR-0020 skill into contract — ordered cut procedure, the mutually-exclusive-flows test, reference-file conventions, the collateral checklist, and a worked description retrofit (loaded from the improve flow when a budget is exceeded) |
| `references/deployment-modes.md` | APM package vs standalone differences and self-containment/cache-isolation rules (loaded on demand) |
| `references/scripts.md` | Package runners, inline dependency patterns, and full script contract (loaded on demand) |
| `references/sources.md` | Upstream research sources and which skill files each contributed to |

View File

@@ -19,10 +19,9 @@ metadata:
## Gotchas
- A skill's `name` and `description` are preloaded into every agent's context every session, invoked or not; the body loads only on invocation. The description is the scarce budget.
- The word gates are two different measurements, not one rule with two tiers. The 2,770-word / 500-line spec backstop counts the whole file including frontmatter; Step 3's gate counts the body alone. A file can sit well inside one and fail the other, so never unify them.
- Never spawn a subagent to audit or recheck your own work here. Run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent can have its worktree torn down by concurrent cleanup, destroying an uncommitted draft.
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires transcript analysis that is out of scope here — flag the opportunity as a suggestion instead.
- The word gates are two measurements, not two tiers of one rule: the 2,770-word / 500-line spec backstop counts the whole file, Step 3's gate the body alone. Never unify them.
- Never spawn a subagent to audit or recheck your own work — run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent's worktree can be torn down by concurrent cleanup, destroying an uncommitted draft.
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires out-of-scope transcript analysis — flag the opportunity as a suggestion instead.
## Step 1 — Dispatch
@@ -32,15 +31,15 @@ metadata:
| Directory exists, at least one improvement signal present | Improve | `references/improve.md` |
| Directory exists, no signals | Stop and ask | — |
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask: "No improvement signals found. Did you mean to create a new skill, or do you have feedback to apply?"
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask whether the user meant to create a new skill or has feedback to apply.
Read only the reference matching the resolved flow — each is self-contained. Capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
Read only the reference matching the resolved flow — each is self-contained. If the target sits inside a git worktree, capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
## Step 2 — Invocation axis
Decide before writing any description: model-invoked or hand-invoked?
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Worked example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`. Skip Step 3's description rules.
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Skip Step 3's description rules.
- **Model-invoked** — the default.
## Step 3 — Contract
@@ -51,12 +50,12 @@ Gates `/skill-audit` enforces in both flows:
- **Description** — a trigger clause, at most one capability clause, and a boundary clause shaped `Not <thing> -> <skill-name>` whose target resolves to a real skill or agent. 250 characters SUGGESTION, 400 FAIL, value only.
- **Body** — decision procedure only: ordered steps, branches, gates, and which reference to load when. 600 words SUGGESTION, 900 FAIL, body only. At two or more mutually exclusive flows a dispatch table is mandatory and each flow gets its own self-contained `references/` file.
- **Gotchas** — at most five, each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL.
- **Gotchas** — each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL; over five entries is a SUGGESTION only.
## Step 4 — Validate and close
Run `/skill-audit` on the resolved skill directory. It checks name-to-directory match, description presence, leftover `FILL IN:` placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those first. Resolve every FAIL before reporting done.
Run `/skill-audit` on the resolved skill directory; resolve every FAIL before reporting done. It checks name-to-directory match, placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those. Hand-check the one thing it misses: an empty body reports `PASS SKILL.md body word count 0 (ADR-0020 target: 600)`, so confirm at least one non-empty section exists.
With `metadata.version` present, bump the **minor** version on create (new skills start at `0.1.0`) and the **patch** version on improve.
**Commit verification.** Once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from the one captured at Step 1. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up first. Report done only once the hash has changed.
**Commit verification.** Inside a git worktree: once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from Step 1's. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up. Report done only once the hash has changed. Outside a worktree (a skill under `~/.claude/skills/`, say) nothing is committable — report done on a clean audit, naming that as the reason.

View File

@@ -10,14 +10,22 @@ name: SKILL_NAME
# Examples: my-tool, data-analyzer, pdf-processor
description: >
Use when FILL IN: trigger — when should an agent activate this skill?
Use when FILL IN: trigger.
FILL IN: at most ONE capability clause, stated specifically
(e.g. "parses and validates OpenAPI specs", not "helps with APIs").
Not FILL IN: near-miss case -> FILL IN: real sibling skill name.
Not FILL IN: near-miss case -> FILL IN: real sibling skill.
# Required. Preloaded into EVERY session whether or not the skill is invoked.
# Exactly three parts, in this order: trigger clause, at most one capability
# clause, boundary clause. Drop the boundary line if no near-miss skill exists.
# Budget: 250 characters target, 400 hard ceiling (counting this value only).
# Trigger clause: when should an agent activate this skill? Describe the user's
# intent, not the skill's internal mechanics.
# Budget: 250 characters target, 400 hard ceiling (counting this value only,
# with YAML folding resolved). This scaffold sits at 214 — keep the fill-in
# under the target rather than growing past it.
# Boundary clauses may be plural: write one per genuine near-miss, and none
# where no sibling could steal activations.
# Never let a hyphenated skill name wrap across two lines of this folded block
# — folding turns the break into a space and the routing target stops resolving.
# Banned here: capability lists, output-format detail, composition notes,
# implementation detail, and restating one trigger twice in two registers.
# The boundary target must resolve to a real skill or agent — it is checked.

View File

@@ -47,11 +47,15 @@ explicitly" only where the user's natural phrasing genuinely omits the domain wo
for `git-commits`, where the user says "commit". Adding one everywhere is what inflated this
corpus, and it was deleted as a blanket rule.
**Boundary targets must resolve.** The name after the arrow is checked against real skill
directories under `plugins/*/.apm/skills/<name>/` and real agents under
`plugins/*/.apm/agents/<name>.agent.md`. A boundary clause naming a target that does not exist
sends the router nowhere and fails the audit. Check the target exists before writing it — do not
invent a plausible sibling name.
**Boundary targets must resolve.** Both forms are checked — the arrow and the prose form ("do not
use for X, use `y` instead") — so a typo dangles either way. Targets resolve against a universe
built by walking up **from the SKILL.md itself**: the nearest ancestor holding
`plugins/*/.apm/{skills,agents}` (or, failing that, the nearest ancestor holding `.git`) contributes
every skill and agent under `<root>/plugins/*/`, plus the skill's own apm package and the packages
that package declares in `apm.yml` under `dependencies.apm`. A sibling plugin in the same monorepo
therefore resolves; a skill in an unrelated repo does not. A boundary clause naming a target
outside that universe sends the router nowhere and fails the audit. Check the target exists before
writing it — do not invent a plausible sibling name.
**Length.** 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only
with YAML folding resolved. The agentskills.io 1,024-character spec limit is unchanged and sits
@@ -61,7 +65,13 @@ as the outlier stop.
**Hand-invoked skills are exempt.** A skill carrying `disable-model-invocation: true` is absent
from the model-visible listing and is reached only by the user typing `/name`. It takes one plain
human-facing sentence — no trigger clause, no boundary clause, no indirect triggers. Worked
example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`.
example — the whole description of the `zoom-out` skill, which carries `disable-model-invocation`:
````markdown
Tell the agent to zoom out and give broader context or a higher-level perspective. Use when
you're unfamiliar with a section of code or need to understand how it fits into the bigger
picture.
````
## Body
@@ -90,6 +100,12 @@ blocks, rationale prose, and any content only one branch reaches. Each reference
self-contained for its concern, and every one is wired from the body with the literal conditional
form:
**The one exception, stated once so it is not re-litigated:** an output schema stays in the body
only when it applies to *every* flow and is short — roughly 50 words or less, which is the "Output
format template" pattern below. An output schema that is longer than that, or that only one flow
produces, moves to `references/` like any other schema. No third option exists, and the two rules
do not disagree.
````markdown
If <condition>, read `references/<file>.md`.
````
@@ -98,8 +114,9 @@ A generic pointer ("see references/ for details") is a Vale error — the agent
**Dispatch is mandatory at two or more mutually exclusive flows.** The body carries the dispatch
table and the gates common to every branch; each flow gets its own self-contained `references/`
file. Exemplar: `plugins/kyberforge/.apm/skills/apm-workflow/SKILL.md` — a 554-word body
dispatching to 3,006 words of references.
file. Exemplar: the `apm-workflow` skill — a **421-word body** dispatching to 3,006 words of
references. Calibrate against 421: that file's whole-file count is 554 words, and aiming at that
number instead overshoots the body budget by ~30%.
**Length.** 600 words SUGGESTION, 900 words FAIL, counting the **body only** — everything after
the frontmatter's closing `---`.
@@ -108,7 +125,7 @@ the frontmatter's closing `---`.
- Each entry must state a fact that **contradicts a reasonable default** — something the agent
gets wrong by acting sensibly. "Never commit secrets" is not one; the agent already knows.
- Maximum five entries.
- More than five entries is a SUGGESTION — five is the guideline, not a ceiling.
- A Gotcha that paraphrases a step in the body below it is a **FAIL**. If the rule is already a
step, it is not a gotcha.
- A Gotchas section exceeding 25% of the body is a SUGGESTION.
@@ -164,7 +181,9 @@ Do not modify flags.
| <condition> | <flow> | `references/<file>.md` |
````
**Output format template** (when the skill produces structured output):
**Output format template** (when the skill produces structured output on *every* flow, and the
schema is roughly 50 words or less — see the exception under Body above; anything longer or
flow-specific belongs in `references/`):
````markdown
Output format:
@@ -173,7 +192,8 @@ Output format:
```
````
For longer templates, place them in `assets/<name>.md` and reference conditionally.
For longer templates, place them in `references/<topic>.md` or `assets/<name>.md` and reference
conditionally.
## Embedding org-specific policy

View File

@@ -136,10 +136,16 @@ If no scripts are needed, delete `scripts/README.md` and the `scripts/` director
## Step 5 — Add references, assets, and tests (if needed)
**`references/`** — additional documentation loaded on demand. One topic per file. Reference
conditionally from SKILL.md with the literal form ``If <condition>, read `references/<file>.md` ``.
Keep reference chains one level deep — a reference file that references another reference file is
rarely loaded correctly.
**`references/`** — additional documentation loaded on demand. One topic per file, named in
kebab-case after the topic. Reference conditionally from SKILL.md with the literal form
``If <condition>, read `references/<file>.md` ``.
**Two hops from `SKILL.md`, never three.** A flow file may route on to a shared contract or
sub-topic file — that is the shipped pattern here (`SKILL.md` → `references/create.md` → this
file's own pointers to `contract.md`, `scripts.md` and `deployment-modes.md`). What does not work
is a third hop: a file reachable only through two intermediates is rarely loaded at the moment it
is needed. Every hop past the first also needs the same literal conditional form, so the agent
knows when to take it.
**`assets/`** — static resources: templates, schemas, lookup tables. Reference by relative path
from SKILL.md.

View File

@@ -76,6 +76,12 @@ contract first — the gates are hot and carry no baseline file, so a one-line f
non-compliant skill cannot be committed until the description and body meet
`references/contract.md`. Treat that retrofit as part of the same change, not a follow-up.
If the skill's description exceeds 250 characters, or its body-only word count exceeds 600, read
`references/retrofit.md` before editing. It carries the ordered cut procedure, the
mutually-exclusive-flows test, the reference-file conventions this flow needs, the collateral
checklist for `README.md` and `references/sources.md`, and a worked description retrofit. Do not
improvise the cuts — four dry runs invented six to ten different answers to the same questions.
If a signal points to a script or reference file, edit that file directly rather than adding a
workaround in SKILL.md.

View File

@@ -0,0 +1,152 @@
---
source_keys:
- agentskills-best-practices
- agentskills-optimizing-descriptions
---
# Retrofitting a skill to the ADR-0020 contract
Read this when `references/improve.md` Step 4 sends you here: the skill you are editing is over
the description or body budget and has to come into contract before any other change can be
committed. The gates are hot and carry no baseline file, so a one-line fix to a non-compliant
skill is blocked until this is done.
Measure first. Do not guess which gate fired: run `/skill-audit` on the directory and read its
`### Structure` dimension, which reports the description characters and the **body-only** word
count separately from the whole-file spec backstop. Retrofit against the number that actually
fired — a skill can sit a thousand words inside the whole-file backstop while failing the body
budget.
**Validate in place.** Audit the skill's real directory inside its package. Never audit a copy in a
scratch directory, and never move a skill out to work on it: the boundary-target universe is built
by walking up *from the file being checked*, so a copy with no authoring root above it resolves
against nothing and the check declines rather than running —
```text
INFO boundary-target resolution DID NOT RUN — no skill universe could be determined for
this path ... Unchecked target(s): totally-fake-target
```
The run still exits 0, so that line reads as a pass and is not one. Treat `DID NOT RUN` as **not
checked**, always. A retrofit signed off on a scratch copy carries an unverified boundary target
into the corpus, which is precisely the failure this gate exists to catch.
## Cut in this order
Work the list top down and stop as soon as the gate clears. The order is by ratio of tokens
removed to behaviour lost — inverting it is how a retrofit ends up deleting the one instruction
the skill existed to carry.
1. **Gotchas that paraphrase a step in the body below.** Zero information, and already a FAIL on
its own. Delete the Gotcha, keep the step.
2. **Spec restatements** — text that repeats a published specification, a tool's `--help`, or a
ceiling the validator already enforces. The agent gets this right without it. Delete, or move
the table to `references/` if a flow genuinely needs to look it up.
3. **Capability enumeration** — in a description, the feature list after the trigger clause; in a
body, the paragraph that recites what the skill can do. One capability clause survives in the
description; the rest belongs in `README.md`.
4. **Per-flow prose** — anything only one branch of the procedure ever reaches. This is the
largest single win in most bodies, and it is a *move*, not a delete: each flow gets its own
self-contained `references/` file, wired from a dispatch table.
If the body is still over after all four, the skill is doing two jobs. Split it, and say so
rather than compressing prose until it stops being readable.
## What "mutually exclusive flows" means
Two or more flows that a single invocation cannot both take. The three-way test, copied verbatim
from the body-discipline rubric `/skill-audit` judges against — nothing to load, it is quoted in
full here:
> separate subcommands, separate input types, separate lifecycle stages
Any one of the three is enough. Two flows that differ only in a parameter value are one flow.
At two or more mutually exclusive flows a dispatch table is **mandatory** regardless of word
count, because every invocation otherwise pays for every branch it did not take.
## Reference-file conventions
The create flow owns these rules, and this flow is forbidden from reading `references/create.md`,
so what a retrofit needs is restated here:
- **One topic per file.** A file mixing two concerns gets loaded for one of them and spends the
caller's context on the other.
- **Kebab-case filenames**, named after the topic rather than the flow that reads it —
`body-discipline.md`, not `step-3.md`.
- **Wire every file with the literal conditional form** ``If <condition>, read
`references/<file>.md` ``. A generic pointer ("see `references/` for details") is a Vale error.
- **Two hops from `SKILL.md`, never three.** A flow file may route on to a shared contract file;
a file reachable only through two intermediates is rarely loaded when it is needed.
- **`source_keys` frontmatter.** If the content you are moving drew on a research source, the new
file needs top-level `source_keys:` frontmatter listing those slugs, and every slug must already
exist as an `## <slug>` heading in `references/sources.md`. Moving sourced content out of
`SKILL.md` without carrying its slugs across breaks the provenance chain, and `/skill-audit`
reports the new file as an INFO with no `source_keys`.
## Collateral is mandatory, not optional
Moving content out of a `SKILL.md` leaves three files describing a structure that no longer
exists. `/skill-audit`'s provenance check exits clean on all three of these, so nothing catches
them for you. After every retrofit that adds, removes or renames a file:
- [ ] **`README.md` file table** — a row for every new `references/` file, and no row left for a
file that is gone. Say what triggers the load, not just what the file contains.
- [ ] **`references/README.md`**, where the skill has one — same update, same reason.
- [ ] **`references/sources.md` → `Contributing files`** — add the new file to every slug whose
content moved into it, and remove any file the retrofit deleted. This is the one that gets
missed: `sources.md` keeps citing sections of `SKILL.md` that no longer exist, the
provenance check still exits 0, and the stale claim survives review.
- [ ] Re-run `/skill-audit` and confirm its `### Provenance` dimension does not report the new
file as missing `source_keys`.
## Worked example — a description retrofit
`gitea-issues` before, 827 characters, the single most common shape in the corpus:
```text
Use when reading or writing Gitea issues: listing repo issues, getting a single issue's details/
comments/labels, creating an issue, updating its state, adding or editing comments, applying
labels via issue_write, or searching issues/PRs across repositories. Triggers on "create an
issue", "what issues are open", "get issue #N", "close issue #N", "comment on issue #N", "search
issues for X" — even when the user doesn't say "Gitea" explicitly. Composes gitea-labels-
milestones for all label inference/resolution and milestone lookup — do not use this skill to
manage label or milestone definitions themselves (create/edit/delete a label, create/close a
milestone), that's gitea-labels-milestones directly. Do not use for pull requests (use gitea-prs)
or for local git branch/commit work (use gitea-branches or git-branches).
```
After, 240 characters:
```text
Use when reading or writing Gitea issues — list, read, create, comment on, label, close, or
search — even when the user does not say "Gitea". Not pull requests -> `gitea-prs`. Not label or
milestone definitions -> `gitea-labels-milestones`.
```
What came out, and why:
| Removed | Why |
|---|---|
| The second trigger register — `Triggers on "create an issue", "what issues are open", …` | The same triggers restated as quoted user phrasings. Two registers of one trigger list is a FAIL, not a suggestion. |
| `applying labels via issue_write` | Implementation detail. The router does not choose a skill by which MCP call it makes. |
| `Composes gitea-labels-milestones for all label inference/resolution and milestone lookup` | A composition note. It changes no routing decision and belongs in `README.md`. |
| The parenthetical `(create/edit/delete a label, create/close a milestone)` | Capability enumeration inside a boundary clause. The boundary needs the target, not its feature list. |
| The `gitea-branches` / `git-branches` boundary | Dropped entirely. Neither was ever going to win an issue request, so the clause defended against nothing — an invented boundary costs characters and buys no routing accuracy. |
| `Do not use for pull requests (use gitea-prs)` prose form | Kept, but rewritten as `Not pull requests -> \`gitea-prs\`.` The rewrite buys characters and one uniform shape for the router — not safety. Both forms are parsed **and** target-checked, so a typo in the prose form dangles exactly as an arrow typo does. |
What stayed: one trigger clause, one capability clause, the indirect trigger (genuinely warranted
here — people say "create an issue", not "create a Gitea issue"), and the boundary clauses.
## Two rules the gates enforce but the prose does not spell out
**Boundary clauses may be plural.** Write one per genuine near-miss — the example above carries
two, because two different skills could each steal activations. "A boundary clause" in the
contract means *at least one*, not *exactly one*. What is banned is a boundary clause invented for
a skill that was never going to compete, not a second real one.
**Never let a hyphenated routing target wrap across lines in a folded `>` scalar.** YAML folding
replaces the newline with a space, so `gitea-labels-` at the end of one line and `milestones` at
the start of the next fold into `gitea-labels- milestones`. `validate.sh` then reads the target as
`gitea-labels`, finds no such skill, and reports a dangling boundary target — the live finding on
`gitea-issues` today. Reflow the line so the whole name sits on one of them. The same applies to
any backticked skill or agent name in a description.

View File

@@ -34,7 +34,7 @@ source_keys:
- **URL:** https://agentskills.io/skill-creation/best-practices.md
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
- **Description:** Best practices for skill creators — starting from real expertise, spending context wisely, calibrating control, instruction patterns (gotchas, templates, checklists, validation loops)
- **Contributing files:** SKILL.md, references/create.md, references/improve.md, references/contract.md
- **Contributing files:** SKILL.md, references/create.md, references/improve.md, references/contract.md, references/retrofit.md
- **Status:** `extracted`
## agentskills-optimizing-descriptions
@@ -42,7 +42,7 @@ source_keys:
- **URL:** https://agentskills.io/skill-creation/optimizing-descriptions.md
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
- **Description:** How to systematically test and improve skill descriptions for triggering accuracy — eval queries, trigger rate testing, train/validation splits, optimization loop
- **Contributing files:** SKILL.md, references/improve.md, references/contract.md
- **Contributing files:** SKILL.md, references/improve.md, references/contract.md, references/retrofit.md
- **Status:** `extracted`
## agentskills-evaluating-skills

View File

@@ -1,6 +1,6 @@
{
"name": "kyberforge",
"version": "1.5.0",
"version": "1.6.0",
"description": "Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.",
"author": {
"name": "Defame1297",

View File

@@ -1,6 +1,6 @@
{
"name": "kyberforge",
"version": "1.5.0",
"version": "1.6.0",
"description": "Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.",
"author": {
"name": "Defame1297",

View File

@@ -1,5 +1,5 @@
name: kyberforge
version: 1.5.0
version: 1.6.0
description: Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.
author:
name: Defame1297

View File

@@ -76,6 +76,9 @@ Include what the fresh context lacks:
- A direct role instruction opening the prompt: `You are a [role]. When invoked, [action].`
- One bounded job, stated so the agent knows what it must refuse.
- The dispatch, gates, inputs and outputs listed above.
- **Error handling** — what the agent does on malformed, missing or contradictory input: stop and
report, or degrade to a named fallback. Absent it, the agent invents a recovery, and a
subagent's invented recovery is invisible to its caller until the output is wrong.
- Non-obvious environment facts and project-specific conventions it cannot infer.
- One default per decision point with one escape hatch.
@@ -111,6 +114,8 @@ Flag as FAIL if:
Flag as SUGGESTION if:
- The body does not open with a direct role instruction
- The body specifies no error handling — nothing tells the agent what to do with malformed,
missing or contradictory input
- The job the agent describes is unbounded, or bounded only implicitly
- A rationale is missing from a rule the agent is expected to enforce — present but unexplained
- Comments are useful but verbose enough to bury the field they annotate

File diff suppressed because it is too large Load Diff

View File

@@ -34,10 +34,13 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
boundary clause naming a real sibling skill or agent.
250 characters is the target, 400 the hard ceiling (ADR-0020).
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was
deleted. Add "Use proactively" only if the runtime should delegate here without
the user naming this agent.
deleted.
Never write "Use proactively" here. It steers the Claude Code runtime and does
nothing anywhere else, and this file compiles to a Copilot `.agent.md` too, where
agent-audit's KyberforgeCopilot.ProactivePhrase rule grades it a hard FAIL.
The phrase is CC-only; at this scope, a precise trigger clause does that job.
Example: "Use when a diff needs checking for injected credentials before it
merges. Not general code review -> code-reviewer." -->
merges. Not prose or style linting -> `lint-runner`." -->
<!-- model: sonnet
Optional. Aliases: sonnet, opus, haiku, fable. Or full model ID.
@@ -80,3 +83,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
## Output
FILL IN: What does the agent produce? Format, location, structure.
## Errors
FILL IN: What does the agent do on malformed, missing or contradictory input?
State whether it stops and reports, or degrades to a named fallback — and what it
tells the caller either way. An agent with no error handling invents a recovery,
and an invented recovery is invisible until the output is wrong.

View File

@@ -14,10 +14,14 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
boundary clause naming a real sibling skill or agent.
250 characters is the target, 400 the hard ceiling (ADR-0020).
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was
deleted. Add "Use proactively" only if the runtime should delegate here without
the user naming this agent.
deleted.
"Use proactively" is valid HERE and only here: it steers the Claude Code runtime
to offer this agent unprompted. Add it only if that is what you want. If you add
it, leave it OUT of the Copilot half of the pair — the phrase does nothing there
and agent-audit's KyberforgeCopilot.ProactivePhrase grades it a hard FAIL. The
pair must describe the same job; it does not have to be byte-identical.
Example: "Use when a diff needs checking for injected credentials before it
merges. Not general code review -> code-reviewer." -->
merges. Not prose or style linting -> `lint-runner`." -->
<!-- tools: Read, Bash, Grep
Optional. Allowlist of tool names: a comma-separated string or a YAML list.
@@ -99,3 +103,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
## Output
FILL IN: What does the agent produce? Format, location, structure.
## Errors
FILL IN: What does the agent do on malformed, missing or contradictory input?
State whether it stops and reports, or degrades to a named fallback — and what it
tells the caller either way. An agent with no error handling invents a recovery,
and an invented recovery is invisible until the output is wrong.

View File

@@ -19,9 +19,13 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
boundary clause naming a real sibling skill or agent.
250 characters is the target, 400 the hard ceiling (ADR-0020).
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was deleted.
Keep it identical in wording to the Claude Code half of the pair.
Never write "Use proactively" here. It steers the Claude Code runtime and does nothing
in Copilot, and agent-audit's KyberforgeCopilot.ProactivePhrase grades it a hard FAIL.
Otherwise keep the wording matched to the Claude Code half of the pair: agent-audit
checks that both halves describe the same job, not that they are byte-identical, so
dropping the CC-only phrase here is not a pair-consistency finding.
Example: "Use when a diff needs checking for injected credentials before it
merges. Not general code review -> code-reviewer." -->
merges. Not prose or style linting -> `lint-runner`." -->
<!-- tools: ["read", "search", "edit"]
Optional. Array of tool names. Omit = all available tools. [] = no tools.
@@ -68,3 +72,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
## Output
FILL IN: What does the agent produce? Format, location, structure.
## Errors
FILL IN: What does the agent do on malformed, missing or contradictory input?
State whether it stops and reports, or degrades to a named fallback — and what it
tells the caller either way. An agent with no error handling invents a recovery,
and an invented recovery is invisible until the output is wrong.

View File

@@ -44,15 +44,31 @@ Banned from a description; move it to the body or to `README.md`:
and ADR-0020 deleted it: the opener is `Use when`, matching every skill in this corpus, so one
router reads one shape.
**"Use proactively" is conditional.** Add it only where the runtime should delegate without the
user naming the agent — an agent invoked by name does not need it, and it costs activations
elsewhere when added by reflex. The same conditional governs indirect triggers ("even if the user
doesn't say X"): add one only where the user's natural phrasing genuinely omits the domain word.
**"Use proactively" is Claude Code-only, and conditional even there.** The phrase steers the
Claude Code runtime to offer an agent unprompted and does nothing anywhere else, so where it may
appear depends on the file:
**Boundary targets must resolve.** The name after the arrow is checked against real skills under
`plugins/*/.apm/skills/<name>/` and real agents under `plugins/*/.apm/agents/<name>.agent.md`. A
target that does not exist sends the router nowhere. Verify it before writing it — do not invent a
plausible sibling.
| File | Rule |
|---|---|
| Claude Code `.md` (project/user scope) | Allowed. Add it only where the runtime should delegate without the user naming the agent — an agent invoked by name does not need it, and it costs activations elsewhere when added by reflex. |
| Copilot `.agent.md` (project/user scope) | **Never.** Inert there, and `KyberforgeCopilot.ProactivePhrase` grades it a hard FAIL. |
| Vendor-neutral `.apm/agents/<name>.agent.md` (plugin/APM scope) | **Never.** Same Vale rule, same hard FAIL — the file matches the `**/*.agent.md` glob, and it compiles to a real Copilot agent downstream. |
A pair whose Claude Code half carries the phrase and whose Copilot half omits it is correct, not
inconsistent: `agent-audit` checks that both halves describe the same job, not that they match
word for word.
Indirect triggers ("even if the user doesn't say X") take a similar conditional at every scope:
add one only where the user's natural phrasing genuinely omits the domain word.
**Boundary targets must resolve.** Both forms are checked — the arrow and the prose form ("do not
use for X, use `y` instead") — so a typo dangles either way. Targets resolve against a universe
built by walking up **from the agent file itself**: the nearest ancestor holding
`plugins/*/.apm/{skills,agents}` (or, failing that, the nearest ancestor holding `.git`) contributes
every skill and agent under `<root>/plugins/*/`, plus the agent's own apm package and the packages
that package declares in `apm.yml` under `dependencies.apm`. A sibling plugin in the same monorepo
therefore resolves; a skill in an unrelated repo does not. A target outside that universe sends the
router nowhere. Verify it before writing it — do not invent a plausible sibling.
**Length.** 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only
with YAML folding resolved. Treat 250 as the target: the SUGGESTION tier is what moves the corpus
@@ -73,8 +89,18 @@ You are a <role>. When invoked, <primary action>.
## Output
<what it produces: format, location, structure>
## Errors
<what to do on malformed, missing or contradictory input: report and stop, or
which fallback to take — and what to say to the caller either way>
````
Four required elements: **inputs expected, process steps, output format, error handling.** The
last is the one that gets dropped, and dropping it is not neutral: an agent given a malformed
input and no instruction invents a recovery, and a subagent's invented recovery is invisible to
the caller until the output is wrong. Say explicitly whether the agent stops and reports, or
degrades to a named fallback.
One job per agent. An agent covering two jobs gets delegated to for the wrong one.
**Delegation discipline replaces the word gate.** A plugin/APM agent is a single file with no

View File

@@ -63,5 +63,7 @@ to installed skills instead of transcribed procedure.
- [ ] No `FILL IN:` placeholder and no `<!-- ... -->` template comment anywhere in the file
- [ ] System prompt body non-empty, and a read-only agent says so in prose as well as in
`disallowedTools`
- [ ] Body covers all four required elements: inputs expected, process steps, output format,
**error handling** — what the agent does on malformed, missing or contradictory input
Then return to the flow reference you came from.

View File

@@ -95,6 +95,8 @@ Both files:
- [ ] `name` present and kebab-case; `description` written to `references/contract.md`
- [ ] System prompt body present, non-empty and equivalent across the pair
- [ ] Body covers all four required elements: inputs expected, process steps, output format,
**error handling** — what the agent does on malformed, missing or contradictory input
- [ ] No `FILL IN:` placeholder and no `<!-- ... -->` template comment left
Copilot file only:

View File

@@ -285,6 +285,8 @@ else
if [[ "$SCOPE" == "plugin" ]]; then
echo " 1. Fill in $APM_FILE — replace every FILL IN: placeholder. Optional fields are" >&2
echo " scaffolded there as commented blocks; uncomment the ones that apply." >&2
echo " Description: 250 chars target / 400 ceiling (ADR-0020). The body has no" >&2
echo " word gate — delegate to a skill instead of restating what it does." >&2
echo " 2. Populate $SOURCES_DIR/sources.md with research sources, or delete it" >&2
echo " 3. Validate: $VALIDATE_HINT $APM_FILE" >&2
echo " It checks the frontmatter against the apm-agent-allowlist section of" >&2
@@ -292,6 +294,8 @@ else
else
echo " 1. Fill in $CC_FILE — replace every FILL IN: placeholder. Optional fields are" >&2
echo " scaffolded there as commented blocks; uncomment the ones that apply." >&2
echo " Description: 250 chars target / 400 ceiling (ADR-0020). The body has no" >&2
echo " word gate — delegate to a skill instead of restating what it does." >&2
echo " 2. Fill in $CP_FILE — same, and heed its closing comment: the Claude Code-only" >&2
echo " fields it names must not cross over from the file above." >&2
echo " 3. Validate: run $VALIDATE_HINT on each file" >&2

View File

@@ -6,10 +6,12 @@ Audit a skill directory against the agentskills.io specification and the house c
1. Runs `scripts/validate.sh` and `scripts/validate-provenance.sh` for structural and provenance checks, plus `scripts/vale-wrap.sh` — a Vale prefilter that deterministically flags non-imperative description openers, composition and architecture notes, vague wording, padding phrases, and "There is/are" sentence openers
2. Reads all files in the skill directory
3. Applies qualitative checks across six dimension groups, loading one rubric from `references/` per group
3. Applies qualitative checks across five dimension groups, loading one rubric from `references/` per group
4. Outputs a compact findings report — findings only, grouped by dimension, each with Why and Fix — and a result block with handoff to `skill-author`
`validate.sh` enforces two independent length families that must not be conflated: the agentskills.io spec conformance ceilings (500 lines, 2,770 words, both counting the whole file) and the ADR-0020 context budget (250/400 description characters, 600/900 body-only words, plus resolvable boundary targets).
`validate.sh` enforces two independent length families that must not be conflated: the agentskills.io spec conformance ceilings (500 lines, 2,770 words, both counting the whole file) and the ADR-0020 context budget (250/400 description characters, 600/900 body-only words).
Alongside those it runs four shape checks that are not length measurements at all. Two are FAILs: every routing target named in the description — in the compressed `Not <thing> -> <name>` arrow **and** in the prose form — must resolve to a real skill or agent, and every `references/<file>.md` the body names must exist on disk. Three are SUGGESTIONs: a missing boundary clause, a Gotchas section over five entries, and a Gotchas section over 25% of the body. The resolution universe for boundary targets is derived by walking up from the audited `SKILL.md` — the authoring root above it, its own apm package, and that package's declared `apm.yml` dependencies — so a fresh clone and a machine that has run `apm install` return the same verdict. When no universe can be determined the check prints `INFO ... DID NOT RUN` and does not silently pass.
## Usage
@@ -24,7 +26,7 @@ Provide the path to the skill directory to audit when invoking.
| File | Purpose |
|------|---------|
| `SKILL.md` | Skill instructions for agents |
| `scripts/validate.sh` | Structural validator — checks name format, name matches directory, description length, line count, placeholder detection, script executable bit, and interactive-prompt detection |
| `scripts/validate.sh` | Structural validator — checks name format, name matches directory, description presence and length, body-only word count, line and whole-file word ceilings, boundary-clause presence, boundary-target resolution, `references/` pointer existence, Gotchas entry count and body share, placeholder detection, script executable bit, and interactive-prompt detection |
| `scripts/validate-provenance.sh` | Provenance validator — checks sources.md completeness, source_keys/slug consistency, Contributing files existence, bidirectional linkage, Research doc: fields, and upstream research doc alignment |
| `scripts/vale-wrap.sh` | Vale prefilter wrapper — runs the bundled `Kyberforge` Vale styles against SKILL.md and reports alerts as deterministic FAILs ahead of Step 3's qualitative review |
| `assets/vale/.vale.ini` | Vale configuration — points Vale at the bundled `Kyberforge` style path, self-located relative to `vale-wrap.sh` |
@@ -38,6 +40,7 @@ Provide the path to the skill directory to audit when invoking.
| `references/patterns.md` | Rubric for the patterns dimension — which instruction construct fits which job, and how each is correctly formed |
| `references/file-structure.md` | Rubric for the file-structure and internal-consistency dimensions — permitted directories, cross-plugin path rules and their two structural exemptions, README drift |
| `references/formatting-and-scripts.md` | Rubric for the formatting and scripts dimensions — heading and fencing conventions, and the agentic-use criteria for bundled scripts |
| `references/validation-scripts.md` | Step 1 troubleshooting — the manual structural fallback when `validate.sh` cannot run, and the script exit codes that are easy to misread (loaded only on a script failure) |
| `references/sources.md` | Provenance record — agentskills.io sources that informed this skill and which files each contributed to |
| `tests/validate.bats` | (source-only) Bats test suite for validate.sh |
| `tests/validate-provenance.bats` | (source-only) Bats test suite for validate-provenance.sh |

View File

@@ -33,7 +33,9 @@ bash scripts/validate-provenance.sh <skill-dir>
scripts/vale-wrap.sh <skill-dir>/SKILL.md
```
`validate.sh` findings become the `### Structure` dimension — its FAILs and its SUGGESTIONs both. If it cannot run at all (no `python3`, Bash denied), report that as an INFO finding rather than guessing; what it measures is not reproducible by reading.
`validate.sh` findings become the `### Structure` dimension — its FAILs and its SUGGESTIONs both.
If any of the three fails, cannot run, or reports something needing interpretation, read `references/validation-scripts.md` — it carries the manual fallback and the misleading exit codes.
`validate-provenance.sh` prints nothing on success. Its FAIL and INFO findings become a separate `### Provenance` dimension, and it emits Why and Fix itself — surface those verbatim.

View File

@@ -69,8 +69,10 @@ table** plus the gates common to every branch, and each flow lives in its own se
`references/` file. Inlining all of them is a FAIL regardless of word count, because every
invocation then pays for every branch it did not take.
The reference shape in this repo is `apm-workflow`: a 554-word body dispatching to roughly 3,000
words of references across five mutually exclusive invocations.
The reference shape in this repo is `apm-workflow`: a **421-word body** dispatching to roughly
3,000 words of references across five mutually exclusive invocations. Its whole-file count is 554
words — cite 421 when calibrating a body, or the conflation this section warns against reappears
in the finding itself.
## Gotchas sections
@@ -86,17 +88,20 @@ sensibly.
Constraints:
- **Maximum five entries.** Past five, the section is a summary of the body rather than a set of
traps, and the agent stops reading it as a warning.
- **More than five entries is a SUGGESTION** — five is the guideline, not a ceiling. Past five, the
section is usually a summary of the body rather than a set of traps, and the agent stops reading
it as a warning. It stays advisory because whether a given gotcha earns its place is judgment;
`validate.sh` emits it through `suggest()` and the run still exits 0.
- **A Gotcha that paraphrases a step in the body below it is a FAIL.** It has no independent
content, and it teaches the agent that Gotchas can be skimmed because the real instruction is
coming.
coming. This one is the auditor's call — no script detects it.
- **A Gotchas section exceeding 25% of the body is a SUGGESTION** — the body has been inverted into
a preamble.
a preamble. Same tier and same reasoning as the entry count, and independent of it: either can
fire without the other.
- Place the section near the top. A gotcha read after the mistake is worthless, which is also why
Gotchas is the one construct exempt from moving to `references/`.
Worked negative example — `git-commits` carries thirteen entries, of which four restate content
Worked negative example — `git-commits` carries twelve entries, of which four restate content
that already appears below or in the description:
| Gotcha | Restates |
@@ -106,8 +111,10 @@ that already appears below or in the description:
| `:33` "Never skip hooks with `--no-verify`" | step 9 at `:52` |
| `:36` "Never commit secrets" | step 2 at `:45` |
All four are FAILs under this rule, and the section as a whole breaches the five-entry maximum. It
also passes every plausible word gate, which is the point of auditing the construct directly.
All four are FAILs under the paraphrase rule. The entry count and the section's share of the body
(387 of 1,102 words, 35%) are two further SUGGESTIONs on top — the script reports both, and neither
fails the run on its own. What makes this worth auditing directly is that the four paraphrase FAILs
pass every word gate there is; only reading the construct finds them.
## Calibrating control
@@ -144,7 +151,7 @@ Flag as FAIL if:
- A sentence answers "no" to the core test — it is padding
- The body exceeds 900 words counted body-only (`validate.sh` reports it)
- Two or more mutually exclusive flows are inlined instead of dispatched
- A Gotcha paraphrases a step in the body below it, or the section exceeds five entries
- A Gotcha paraphrases a step in the body below it
- A decision point presents a menu of options with no default
- An instruction repeats content already in the description
- A prescriptive sequence is used where flexibility is fine, or the reverse
@@ -152,6 +159,7 @@ Flag as FAIL if:
Flag as SUGGESTION if:
- The body exceeds 600 words counted body-only but stays at or under 900
- The Gotchas section carries more than five entries
- The Gotchas section exceeds 25% of the body
- A rationale is missing from an include/exclude rule — present but unexplained
- Gotchas are correct but placed late in the body rather than near the top

View File

@@ -28,6 +28,14 @@ resolving there. Flag any `../`, `../../`, or absolute repo path (`plugins/<plug
and its APM-native equivalent `.apm/skills/<other>/`) appearing in `SKILL.md`, `scripts/`,
`references/` or `assets/`.
**Referring to another skill's file.** There is one sanctioned spelling, and it is possessive:
`skill-audit's references/validation-scripts.md`. Write the skill by name and let the reader
resolve it — do not spell the repo path. The full path is the thing this section forbids, and
`references/validation-scripts.md` on its own is a hard ERROR from the ADR-0020 gate, which
requires an unqualified `references/` pointer to exist in the skill's OWN directory. The
possessive form is the only spelling both rules accept; the gate recognises it and skips the
on-disk check. Flag any other spelling of a cross-skill reference.
Two directories are exempt, and the exemptions are structural rather than discretionary:
- **`references/sources.md`.** Its `Research doc:` fields are development-time provenance pointers,

View File

@@ -15,7 +15,7 @@
- **URL:** https://agentskills.io/specification.md
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
- **Description:** Complete SKILL.md format specification — frontmatter fields, constraints, body content, optional directories, progressive disclosure levels, file references, validation
- **Contributing files:** SKILL.md, references/body-discipline.md, references/description-quality.md, references/patterns.md, references/file-structure.md, references/formatting-and-scripts.md
- **Contributing files:** SKILL.md, references/body-discipline.md, references/description-quality.md, references/patterns.md, references/file-structure.md, references/formatting-and-scripts.md, references/validation-scripts.md
- **Status:** `extracted`
## agentskills-best-practices
@@ -47,7 +47,7 @@
- **URL:** https://agentskills.io/skill-creation/using-scripts.md
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
- **Description:** Using scripts in skills — one-off commands, self-contained scripts with inline dependencies, designing scripts for agentic use (no interactive prompts, --help, structured output, idempotency)
- **Contributing files:** SKILL.md, references/formatting-and-scripts.md
- **Contributing files:** SKILL.md, references/formatting-and-scripts.md, references/validation-scripts.md
- **Status:** `extracted`
## agentskills-quickstart

View File

@@ -0,0 +1,118 @@
---
source_keys:
- agentskills-spec
- agentskills-using-scripts
---
# Validation Scripts Reference
Read this when a Step 1 script fails, cannot run, or reports something that needs interpreting.
Nothing here is needed on a clean run.
## Report the gap, do not guess
If a script cannot run at all — Bash denied, `python3` unavailable, PyYAML not importable, `vale`
not installed — say so as an **INFO** finding naming the script and the missing dependency, then
fall back to the manual checks below. An INFO never changes PASS/FAIL. Silently omitting the
dimension a script would have covered reports a clean audit that checked less than it claims to
have checked, and the Step 4 coverage line then names a dimension nothing actually examined.
## Manual structural fallback
`validate.sh` needs `python3` **and** PyYAML, and refuses to start without either — the description
value has to be measured after YAML folding is resolved, so skipping the ADR-0020 gates would be a
vacuous pass rather than a partial one. The two are checked separately, so the message already names
the right one — report it verbatim rather than diagnosing further:
```text
Error: python3 is required but was not found on PATH.
Error: PyYAML is required but is not importable by python3.
```
Without them — or with Bash denied, or on a permission error — work this list
by hand and file the results under `### Structure` exactly as the script's output would have been:
- **`name`** present, 1–64 characters, kebab-case (lowercase letters, digits and hyphens; no
leading, trailing or doubled hyphen), and **matching the skill's directory name** exactly.
- **`description`** present and non-empty; no unfilled `FILL IN:` placeholder in it. An absent or
empty description is a **FAIL**, never a silent skip — it is the one field preloaded into every
session, so a skill without one can never be routed to.
- **Description length**, measured on the folded YAML value with newlines collapsed to single
spaces — not on the raw block scalar, which counts indentation. 250 characters SUGGESTION, 400
FAIL (ADR-0020), 1,024 FAIL (agentskills.io spec).
- **Body length**, counting everything after the frontmatter's closing `---`. 600 words
SUGGESTION, 900 FAIL (ADR-0020).
- **Whole-file ceilings**, counting the file including frontmatter: 500 lines FAIL, 2,770 words
FAIL (agentskills.io spec). These are a different measurement from the two above — report them
as separate findings, never merged.
- **A boundary clause is present** — either the prose form (`do not` / `instead` / `rather than` /
`not for`) or ADR-0020's compressed `Not <thing> -> <name>` arrow. **SUGGESTION**, not FAIL:
the absence is deterministic, but whether this skill warrants one is the auditor's call.
- **Boundary targets resolve** — **FAIL** on a name that resolves to nothing. See the section
below; resolving these by hand is the one item on this list with a procedure of its own.
- **Every `references/<file>.md` named in the body exists on disk** — **FAIL**, not a suggestion.
A dispatch table or "read X" trigger naming a missing file sends the agent nowhere. Ignore
mentions inside fenced code blocks, and ignore a mention whose own line says the file is gone
(`removed`, `deleted`, `renamed`, `superseded`, `replaced`, `obsolete`, `deprecated`, `former`,
`gone`, `no longer`, `used to`) — that is a historical note, not a dispatch entry.
- **Gotchas discipline**, both **SUGGESTION**. Locate the section by a heading that *is* Gotchas
(`## Common Gotchas` counts; `## Gotcha handling` and `## Why gotchas matter` do not), running to
the next heading at the same level or shallower. More than five top-level entries is one
suggestion; a section over 25% of the body word count is a second, independent one. Count
entries at column 0 only — an indented child bullet is not an entry — and ignore fenced code
blocks for both.
- **No unfilled `FILL IN:` placeholder** anywhere in the body.
- **Every file in `scripts/`** carries the executable bit and contains no interactive prompt —
no bare `read`, no `select`, nothing that blocks on a TTY.
## Resolving boundary targets by hand
Targets are read from **both** boundary forms. The compressed `Not <thing> -> <name>` arrow and the
prose form are each parsed *and* target-checked, so a typo in prose phrasing fails exactly as an
arrow typo does — do not check only the names after an arrow.
Build the universe by walking up **from the `SKILL.md` under audit**, never from the validator's own
location. The nearest ancestor holding `plugins/*/.apm/skills/` or `plugins/*/.apm/agents/` is the
authoring root, falling back to the nearest ancestor holding `.git`. When one is found the universe
is every skill and agent under `<root>/plugins/*/`, plus the skill's own apm package, plus the
packages that package declares in its `apm.yml` under `dependencies.apm`. Deployed `.claude/` and
`.agents/` trees are consulted **only** when no authoring root exists — they are gitignored
`apm install` output, and reading them would make a fresh clone and a developer machine disagree.
Three ways to read the result wrong:
- **A hyphenated name used attributively is not a dangling target.** "Use pre-commit hooks instead
of ad-hoc scripts" reads as a route to `pre-commit` on wording alone. What separates a route from
prose is grammar: a route target is terminal — followed by punctuation, a conjunction, or a
boundary word — whereas a compound modifier is followed by the noun it modifies. A name followed
by an ordinary noun still *confirms* a route when it exists, but never raises a FAIL on its own.
- **A SUGGESTION-tier unresolved target is not a FAIL you may promote.** Terminal position alone is
not evidence of a route: "run `pre-commit` instead", "see `commit-msg`" and "use the clean-up
instead" are all terminal and all prose. A prose-form target earns a FAIL only when its own
sentence names another target that *does* resolve; otherwise the script reports it and moves on,
and so should you. Route notation — `/name` and `-> name` — is exempt and always FAILs, and it is
the fix to recommend when the author did mean a route.
- **`INFO boundary-target resolution DID NOT RUN` is not a pass.** The script prints it, and exits
0, when no universe could be determined for that path — the usual cause being a skill copy
audited outside its package. Report it as an INFO naming the unchecked targets and re-run against
the real directory; filing it as clean signs off targets nothing verified.
## Script-specific failures
- **`validate-provenance.sh` printed nothing.** That is a pass, not a skip. It also exits 0
silently when the skill has no `source_keys` and no `references/sources.md` — nothing to
validate is not a finding.
- **`vale` reports `0 files`.** Treat the pass as NOT RUN, not as clean, and fall back to full
Step 3 judgment for the dimensions it would have covered. The bundled `Kyberforge` style is
scoped by glob in `assets/vale/.vale.ini`; a file outside those globs is silently not linted.
- **`E100 Runtime error ... does not exist` (exit 2) from `vale-wrap.sh`.** An explicit relative
`--config` was passed. Pass none: the wrapper locates its own `assets/vale/.vale.ini` from its
own path, so a resolved script path plus an unresolved config path produces exactly this. Do not
read this exit code as vale being unavailable — that misreading sends the audit down the
fallback path while vale was installed and working the whole time.
- **The `vale` binary is genuinely absent** (`command not found`). Report one INFO naming it, then
fall back to full Step 3 judgment for the description, body-discipline and patterns dimensions —
the prefilter's whole coverage. Judge those by rubric rather than dropping them.
- **A path argument that does not exist is a hard error** in `vale-wrap.sh`, deliberately: bare
`vale` would fall back to reading stdin and print a clean-looking `0 errors ... in stdin`, which
the `0 files` guard above does not catch.

File diff suppressed because it is too large Load Diff

View File

@@ -47,6 +47,7 @@ If the destination resolves inside an APM package, read `references/deployment-m
| `references/create.md` | The create flow end to end — prerequisites, package-intent gate, scaffold, frontmatter, scripts, references, sources (loaded on demand) |
| `references/improve.md` | The improve flow end to end — signal verification, root-cause grouping, announcement, edits (loaded on demand) |
| `references/contract.md` | The ADR-0020 description and body contract, the Gotchas constraint, the two size gates, body patterns, and org-policy embedding (loaded on demand) |
| `references/retrofit.md` | Bringing a pre-ADR-0020 skill into contract — ordered cut procedure, the mutually-exclusive-flows test, reference-file conventions, the collateral checklist, and a worked description retrofit (loaded from the improve flow when a budget is exceeded) |
| `references/deployment-modes.md` | APM package vs standalone differences and self-containment/cache-isolation rules (loaded on demand) |
| `references/scripts.md` | Package runners, inline dependency patterns, and full script contract (loaded on demand) |
| `references/sources.md` | Upstream research sources and which skill files each contributed to |

View File

@@ -19,10 +19,9 @@ metadata:
## Gotchas
- A skill's `name` and `description` are preloaded into every agent's context every session, invoked or not; the body loads only on invocation. The description is the scarce budget.
- The word gates are two different measurements, not one rule with two tiers. The 2,770-word / 500-line spec backstop counts the whole file including frontmatter; Step 3's gate counts the body alone. A file can sit well inside one and fail the other, so never unify them.
- Never spawn a subagent to audit or recheck your own work here. Run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent can have its worktree torn down by concurrent cleanup, destroying an uncommitted draft.
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires transcript analysis that is out of scope here — flag the opportunity as a suggestion instead.
- The word gates are two measurements, not two tiers of one rule: the 2,770-word / 500-line spec backstop counts the whole file, Step 3's gate the body alone. Never unify them.
- Never spawn a subagent to audit or recheck your own work — run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent's worktree can be torn down by concurrent cleanup, destroying an uncommitted draft.
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires out-of-scope transcript analysis — flag the opportunity as a suggestion instead.
## Step 1 — Dispatch
@@ -32,15 +31,15 @@ metadata:
| Directory exists, at least one improvement signal present | Improve | `references/improve.md` |
| Directory exists, no signals | Stop and ask | — |
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask: "No improvement signals found. Did you mean to create a new skill, or do you have feedback to apply?"
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask whether the user meant to create a new skill or has feedback to apply.
Read only the reference matching the resolved flow — each is self-contained. Capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
Read only the reference matching the resolved flow — each is self-contained. If the target sits inside a git worktree, capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
## Step 2 — Invocation axis
Decide before writing any description: model-invoked or hand-invoked?
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Worked example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`. Skip Step 3's description rules.
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Skip Step 3's description rules.
- **Model-invoked** — the default.
## Step 3 — Contract
@@ -51,12 +50,12 @@ Gates `/skill-audit` enforces in both flows:
- **Description** — a trigger clause, at most one capability clause, and a boundary clause shaped `Not <thing> -> <skill-name>` whose target resolves to a real skill or agent. 250 characters SUGGESTION, 400 FAIL, value only.
- **Body** — decision procedure only: ordered steps, branches, gates, and which reference to load when. 600 words SUGGESTION, 900 FAIL, body only. At two or more mutually exclusive flows a dispatch table is mandatory and each flow gets its own self-contained `references/` file.
- **Gotchas** — at most five, each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL.
- **Gotchas** — each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL; over five entries is a SUGGESTION only.
## Step 4 — Validate and close
Run `/skill-audit` on the resolved skill directory. It checks name-to-directory match, description presence, leftover `FILL IN:` placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those first. Resolve every FAIL before reporting done.
Run `/skill-audit` on the resolved skill directory; resolve every FAIL before reporting done. It checks name-to-directory match, placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those. Hand-check the one thing it misses: an empty body reports `PASS SKILL.md body word count 0 (ADR-0020 target: 600)`, so confirm at least one non-empty section exists.
With `metadata.version` present, bump the **minor** version on create (new skills start at `0.1.0`) and the **patch** version on improve.
**Commit verification.** Once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from the one captured at Step 1. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up first. Report done only once the hash has changed.
**Commit verification.** Inside a git worktree: once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from Step 1's. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up. Report done only once the hash has changed. Outside a worktree (a skill under `~/.claude/skills/`, say) nothing is committable — report done on a clean audit, naming that as the reason.

View File

@@ -10,14 +10,22 @@ name: SKILL_NAME
# Examples: my-tool, data-analyzer, pdf-processor
description: >
Use when FILL IN: trigger — when should an agent activate this skill?
Use when FILL IN: trigger.
FILL IN: at most ONE capability clause, stated specifically
(e.g. "parses and validates OpenAPI specs", not "helps with APIs").
Not FILL IN: near-miss case -> FILL IN: real sibling skill name.
Not FILL IN: near-miss case -> FILL IN: real sibling skill.
# Required. Preloaded into EVERY session whether or not the skill is invoked.
# Exactly three parts, in this order: trigger clause, at most one capability
# clause, boundary clause. Drop the boundary line if no near-miss skill exists.
# Budget: 250 characters target, 400 hard ceiling (counting this value only).
# Trigger clause: when should an agent activate this skill? Describe the user's
# intent, not the skill's internal mechanics.
# Budget: 250 characters target, 400 hard ceiling (counting this value only,
# with YAML folding resolved). This scaffold sits at 214 — keep the fill-in
# under the target rather than growing past it.
# Boundary clauses may be plural: write one per genuine near-miss, and none
# where no sibling could steal activations.
# Never let a hyphenated skill name wrap across two lines of this folded block
# — folding turns the break into a space and the routing target stops resolving.
# Banned here: capability lists, output-format detail, composition notes,
# implementation detail, and restating one trigger twice in two registers.
# The boundary target must resolve to a real skill or agent — it is checked.

View File

@@ -47,11 +47,15 @@ explicitly" only where the user's natural phrasing genuinely omits the domain wo
for `git-commits`, where the user says "commit". Adding one everywhere is what inflated this
corpus, and it was deleted as a blanket rule.
**Boundary targets must resolve.** The name after the arrow is checked against real skill
directories under `plugins/*/.apm/skills/<name>/` and real agents under
`plugins/*/.apm/agents/<name>.agent.md`. A boundary clause naming a target that does not exist
sends the router nowhere and fails the audit. Check the target exists before writing it — do not
invent a plausible sibling name.
**Boundary targets must resolve.** Both forms are checked — the arrow and the prose form ("do not
use for X, use `y` instead") — so a typo dangles either way. Targets resolve against a universe
built by walking up **from the SKILL.md itself**: the nearest ancestor holding
`plugins/*/.apm/{skills,agents}` (or, failing that, the nearest ancestor holding `.git`) contributes
every skill and agent under `<root>/plugins/*/`, plus the skill's own apm package and the packages
that package declares in `apm.yml` under `dependencies.apm`. A sibling plugin in the same monorepo
therefore resolves; a skill in an unrelated repo does not. A boundary clause naming a target
outside that universe sends the router nowhere and fails the audit. Check the target exists before
writing it — do not invent a plausible sibling name.
**Length.** 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only
with YAML folding resolved. The agentskills.io 1,024-character spec limit is unchanged and sits
@@ -61,7 +65,13 @@ as the outlier stop.
**Hand-invoked skills are exempt.** A skill carrying `disable-model-invocation: true` is absent
from the model-visible listing and is reached only by the user typing `/name`. It takes one plain
human-facing sentence — no trigger clause, no boundary clause, no indirect triggers. Worked
example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`.
example — the whole description of the `zoom-out` skill, which carries `disable-model-invocation`:
````markdown
Tell the agent to zoom out and give broader context or a higher-level perspective. Use when
you're unfamiliar with a section of code or need to understand how it fits into the bigger
picture.
````
## Body
@@ -90,6 +100,12 @@ blocks, rationale prose, and any content only one branch reaches. Each reference
self-contained for its concern, and every one is wired from the body with the literal conditional
form:
**The one exception, stated once so it is not re-litigated:** an output schema stays in the body
only when it applies to *every* flow and is short — roughly 50 words or less, which is the "Output
format template" pattern below. An output schema that is longer than that, or that only one flow
produces, moves to `references/` like any other schema. No third option exists, and the two rules
do not disagree.
````markdown
If <condition>, read `references/<file>.md`.
````
@@ -98,8 +114,9 @@ A generic pointer ("see references/ for details") is a Vale error — the agent
**Dispatch is mandatory at two or more mutually exclusive flows.** The body carries the dispatch
table and the gates common to every branch; each flow gets its own self-contained `references/`
file. Exemplar: `plugins/kyberforge/.apm/skills/apm-workflow/SKILL.md` — a 554-word body
dispatching to 3,006 words of references.
file. Exemplar: the `apm-workflow` skill — a **421-word body** dispatching to 3,006 words of
references. Calibrate against 421: that file's whole-file count is 554 words, and aiming at that
number instead overshoots the body budget by ~30%.
**Length.** 600 words SUGGESTION, 900 words FAIL, counting the **body only** — everything after
the frontmatter's closing `---`.
@@ -108,7 +125,7 @@ the frontmatter's closing `---`.
- Each entry must state a fact that **contradicts a reasonable default** — something the agent
gets wrong by acting sensibly. "Never commit secrets" is not one; the agent already knows.
- Maximum five entries.
- More than five entries is a SUGGESTION — five is the guideline, not a ceiling.
- A Gotcha that paraphrases a step in the body below it is a **FAIL**. If the rule is already a
step, it is not a gotcha.
- A Gotchas section exceeding 25% of the body is a SUGGESTION.
@@ -164,7 +181,9 @@ Do not modify flags.
| <condition> | <flow> | `references/<file>.md` |
````
**Output format template** (when the skill produces structured output):
**Output format template** (when the skill produces structured output on *every* flow, and the
schema is roughly 50 words or less — see the exception under Body above; anything longer or
flow-specific belongs in `references/`):
````markdown
Output format:
@@ -173,7 +192,8 @@ Output format:
```
````
For longer templates, place them in `assets/<name>.md` and reference conditionally.
For longer templates, place them in `references/<topic>.md` or `assets/<name>.md` and reference
conditionally.
## Embedding org-specific policy

View File

@@ -136,10 +136,16 @@ If no scripts are needed, delete `scripts/README.md` and the `scripts/` director
## Step 5 — Add references, assets, and tests (if needed)
**`references/`** — additional documentation loaded on demand. One topic per file. Reference
conditionally from SKILL.md with the literal form ``If <condition>, read `references/<file>.md` ``.
Keep reference chains one level deep — a reference file that references another reference file is
rarely loaded correctly.
**`references/`** — additional documentation loaded on demand. One topic per file, named in
kebab-case after the topic. Reference conditionally from SKILL.md with the literal form
``If <condition>, read `references/<file>.md` ``.
**Two hops from `SKILL.md`, never three.** A flow file may route on to a shared contract or
sub-topic file — that is the shipped pattern here (`SKILL.md` → `references/create.md` → this
file's own pointers to `contract.md`, `scripts.md` and `deployment-modes.md`). What does not work
is a third hop: a file reachable only through two intermediates is rarely loaded at the moment it
is needed. Every hop past the first also needs the same literal conditional form, so the agent
knows when to take it.
**`assets/`** — static resources: templates, schemas, lookup tables. Reference by relative path
from SKILL.md.

View File

@@ -76,6 +76,12 @@ contract first — the gates are hot and carry no baseline file, so a one-line f
non-compliant skill cannot be committed until the description and body meet
`references/contract.md`. Treat that retrofit as part of the same change, not a follow-up.
If the skill's description exceeds 250 characters, or its body-only word count exceeds 600, read
`references/retrofit.md` before editing. It carries the ordered cut procedure, the
mutually-exclusive-flows test, the reference-file conventions this flow needs, the collateral
checklist for `README.md` and `references/sources.md`, and a worked description retrofit. Do not
improvise the cuts — four dry runs invented six to ten different answers to the same questions.
If a signal points to a script or reference file, edit that file directly rather than adding a
workaround in SKILL.md.

View File

@@ -0,0 +1,152 @@
---
source_keys:
- agentskills-best-practices
- agentskills-optimizing-descriptions
---
# Retrofitting a skill to the ADR-0020 contract
Read this when `references/improve.md` Step 4 sends you here: the skill you are editing is over
the description or body budget and has to come into contract before any other change can be
committed. The gates are hot and carry no baseline file, so a one-line fix to a non-compliant
skill is blocked until this is done.
Measure first. Do not guess which gate fired: run `/skill-audit` on the directory and read its
`### Structure` dimension, which reports the description characters and the **body-only** word
count separately from the whole-file spec backstop. Retrofit against the number that actually
fired — a skill can sit a thousand words inside the whole-file backstop while failing the body
budget.
**Validate in place.** Audit the skill's real directory inside its package. Never audit a copy in a
scratch directory, and never move a skill out to work on it: the boundary-target universe is built
by walking up *from the file being checked*, so a copy with no authoring root above it resolves
against nothing and the check declines rather than running —
```text
INFO boundary-target resolution DID NOT RUN — no skill universe could be determined for
this path ... Unchecked target(s): totally-fake-target
```
The run still exits 0, so that line reads as a pass and is not one. Treat `DID NOT RUN` as **not
checked**, always. A retrofit signed off on a scratch copy carries an unverified boundary target
into the corpus, which is precisely the failure this gate exists to catch.
## Cut in this order
Work the list top down and stop as soon as the gate clears. The order is by ratio of tokens
removed to behaviour lost — inverting it is how a retrofit ends up deleting the one instruction
the skill existed to carry.
1. **Gotchas that paraphrase a step in the body below.** Zero information, and already a FAIL on
its own. Delete the Gotcha, keep the step.
2. **Spec restatements** — text that repeats a published specification, a tool's `--help`, or a
ceiling the validator already enforces. The agent gets this right without it. Delete, or move
the table to `references/` if a flow genuinely needs to look it up.
3. **Capability enumeration** — in a description, the feature list after the trigger clause; in a
body, the paragraph that recites what the skill can do. One capability clause survives in the
description; the rest belongs in `README.md`.
4. **Per-flow prose** — anything only one branch of the procedure ever reaches. This is the
largest single win in most bodies, and it is a *move*, not a delete: each flow gets its own
self-contained `references/` file, wired from a dispatch table.
If the body is still over after all four, the skill is doing two jobs. Split it, and say so
rather than compressing prose until it stops being readable.
## What "mutually exclusive flows" means
Two or more flows that a single invocation cannot both take. The three-way test, copied verbatim
from the body-discipline rubric `/skill-audit` judges against — nothing to load, it is quoted in
full here:
> separate subcommands, separate input types, separate lifecycle stages
Any one of the three is enough. Two flows that differ only in a parameter value are one flow.
At two or more mutually exclusive flows a dispatch table is **mandatory** regardless of word
count, because every invocation otherwise pays for every branch it did not take.
## Reference-file conventions
The create flow owns these rules, and this flow is forbidden from reading `references/create.md`,
so what a retrofit needs is restated here:
- **One topic per file.** A file mixing two concerns gets loaded for one of them and spends the
caller's context on the other.
- **Kebab-case filenames**, named after the topic rather than the flow that reads it —
`body-discipline.md`, not `step-3.md`.
- **Wire every file with the literal conditional form** ``If <condition>, read
`references/<file>.md` ``. A generic pointer ("see `references/` for details") is a Vale error.
- **Two hops from `SKILL.md`, never three.** A flow file may route on to a shared contract file;
a file reachable only through two intermediates is rarely loaded when it is needed.
- **`source_keys` frontmatter.** If the content you are moving drew on a research source, the new
file needs top-level `source_keys:` frontmatter listing those slugs, and every slug must already
exist as an `## <slug>` heading in `references/sources.md`. Moving sourced content out of
`SKILL.md` without carrying its slugs across breaks the provenance chain, and `/skill-audit`
reports the new file as an INFO with no `source_keys`.
## Collateral is mandatory, not optional
Moving content out of a `SKILL.md` leaves three files describing a structure that no longer
exists. `/skill-audit`'s provenance check exits clean on all three of these, so nothing catches
them for you. After every retrofit that adds, removes or renames a file:
- [ ] **`README.md` file table** — a row for every new `references/` file, and no row left for a
file that is gone. Say what triggers the load, not just what the file contains.
- [ ] **`references/README.md`**, where the skill has one — same update, same reason.
- [ ] **`references/sources.md` → `Contributing files`** — add the new file to every slug whose
content moved into it, and remove any file the retrofit deleted. This is the one that gets
missed: `sources.md` keeps citing sections of `SKILL.md` that no longer exist, the
provenance check still exits 0, and the stale claim survives review.
- [ ] Re-run `/skill-audit` and confirm its `### Provenance` dimension does not report the new
file as missing `source_keys`.
## Worked example — a description retrofit
`gitea-issues` before, 827 characters, the single most common shape in the corpus:
```text
Use when reading or writing Gitea issues: listing repo issues, getting a single issue's details/
comments/labels, creating an issue, updating its state, adding or editing comments, applying
labels via issue_write, or searching issues/PRs across repositories. Triggers on "create an
issue", "what issues are open", "get issue #N", "close issue #N", "comment on issue #N", "search
issues for X" — even when the user doesn't say "Gitea" explicitly. Composes gitea-labels-
milestones for all label inference/resolution and milestone lookup — do not use this skill to
manage label or milestone definitions themselves (create/edit/delete a label, create/close a
milestone), that's gitea-labels-milestones directly. Do not use for pull requests (use gitea-prs)
or for local git branch/commit work (use gitea-branches or git-branches).
```
After, 240 characters:
```text
Use when reading or writing Gitea issues — list, read, create, comment on, label, close, or
search — even when the user does not say "Gitea". Not pull requests -> `gitea-prs`. Not label or
milestone definitions -> `gitea-labels-milestones`.
```
What came out, and why:
| Removed | Why |
|---|---|
| The second trigger register — `Triggers on "create an issue", "what issues are open", …` | The same triggers restated as quoted user phrasings. Two registers of one trigger list is a FAIL, not a suggestion. |
| `applying labels via issue_write` | Implementation detail. The router does not choose a skill by which MCP call it makes. |
| `Composes gitea-labels-milestones for all label inference/resolution and milestone lookup` | A composition note. It changes no routing decision and belongs in `README.md`. |
| The parenthetical `(create/edit/delete a label, create/close a milestone)` | Capability enumeration inside a boundary clause. The boundary needs the target, not its feature list. |
| The `gitea-branches` / `git-branches` boundary | Dropped entirely. Neither was ever going to win an issue request, so the clause defended against nothing — an invented boundary costs characters and buys no routing accuracy. |
| `Do not use for pull requests (use gitea-prs)` prose form | Kept, but rewritten as `Not pull requests -> \`gitea-prs\`.` The rewrite buys characters and one uniform shape for the router — not safety. Both forms are parsed **and** target-checked, so a typo in the prose form dangles exactly as an arrow typo does. |
What stayed: one trigger clause, one capability clause, the indirect trigger (genuinely warranted
here — people say "create an issue", not "create a Gitea issue"), and the boundary clauses.
## Two rules the gates enforce but the prose does not spell out
**Boundary clauses may be plural.** Write one per genuine near-miss — the example above carries
two, because two different skills could each steal activations. "A boundary clause" in the
contract means *at least one*, not *exactly one*. What is banned is a boundary clause invented for
a skill that was never going to compete, not a second real one.
**Never let a hyphenated routing target wrap across lines in a folded `>` scalar.** YAML folding
replaces the newline with a space, so `gitea-labels-` at the end of one line and `milestones` at
the start of the next fold into `gitea-labels- milestones`. `validate.sh` then reads the target as
`gitea-labels`, finds no such skill, and reports a dangling boundary target — the live finding on
`gitea-issues` today. Reflow the line so the whole name sits on one of them. The same applies to
any backticked skill or agent name in a description.

View File

@@ -34,7 +34,7 @@ source_keys:
- **URL:** https://agentskills.io/skill-creation/best-practices.md
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
- **Description:** Best practices for skill creators — starting from real expertise, spending context wisely, calibrating control, instruction patterns (gotchas, templates, checklists, validation loops)
- **Contributing files:** SKILL.md, references/create.md, references/improve.md, references/contract.md
- **Contributing files:** SKILL.md, references/create.md, references/improve.md, references/contract.md, references/retrofit.md
- **Status:** `extracted`
## agentskills-optimizing-descriptions
@@ -42,7 +42,7 @@ source_keys:
- **URL:** https://agentskills.io/skill-creation/optimizing-descriptions.md
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
- **Description:** How to systematically test and improve skill descriptions for triggering accuracy — eval queries, trigger rate testing, train/validation splits, optimization loop
- **Contributing files:** SKILL.md, references/improve.md, references/contract.md
- **Contributing files:** SKILL.md, references/improve.md, references/contract.md, references/retrofit.md
- **Status:** `extracted`
## agentskills-evaluating-skills

File diff suppressed because it is too large Load Diff

View File

@@ -36,6 +36,22 @@ STRICT=false
if [[ "${RUN_TESTS_STRICT:-}" == "1" ]]; then
STRICT=true
fi
# Latched, then REMOVED from the environment. The value has done its only job by
# this line -- it is now held in the STRICT shell local -- and leaving it exported
# makes strictness leak down the whole process tree: every test-*.sh dispatched
# through batch_run below inherits it, and any of them that itself invokes
# run-tests.sh (tests/test-run-tests.sh drives a copy of this script over fixture
# trees) silently turns a deliberately non-strict fixture strict.
#
# That is not symmetric with `--strict`, which never leaked: the flag only ever
# sets the shell local above, so `bash tests/run-tests.sh --strict` (the spelling
# the run-tests pre-push hook uses) always gave children a clean environment. Only
# the env-var spelling leaked, and it broke exactly two assertions in
# tests/test-run-tests.sh -- its cases 10c and 10g. Unsetting here makes the two
# documented invocations equivalent in what a CHILD sees, not just in the parent's
# verdict, so no future suite has to defend itself the way test-run-tests.sh's
# run_fake() does with `env -u`.
unset RUN_TESTS_STRICT
# A loop rather than the `[[ "${1:-}" == --bats-only ]]` test this used to be, so
# the two flags compose and an unknown flag is rejected instead of ignored. A
# silently-ignored `--strict` is the one typo that would turn the gate back off.

406
tests/test-adr0020-body-checks.sh Executable file
View File

@@ -0,0 +1,406 @@
#!/usr/bin/env bash
# Regression test for the ADR-0020 body-shape checks and, just as importantly,
# for the false-positive fixes each of them needed. Every check here was
# completely untested.
#
# * Gotchas section over 5 entries — SUGGESTION
# * Gotchas section over 25% of the body — SUGGESTION
# * a references/<file>.md named but absent — ERROR (a broken pointer is not a
# style opinion)
# * description with no boundary clause — SUGGESTION
#
# The false-positive half is not optional extra coverage. Each of these checks
# scans prose, and the first naive version of each one fired on ordinary writing:
# a ```-fenced EXAMPLE of a Gotchas section became the section itself, indented
# child bullets were counted as top-level entries, `## Gotcha handling` was read
# as the Gotchas section, and a documented-then-removed references/ file became a
# hard ERROR. The skills most likely to carry such an example are skill-author and
# skill-audit — the two that DOCUMENT these conventions — so a gate that fires on
# them is a gate nobody can turn on.
#
# Every case is a matched pair: the check fires just over its boundary, and stays
# silent just under it (or on the shape it must not match). A test asserting only
# that a bad file fails proves nothing about a check that fires on everything.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
HOOK="$REPO_ROOT/scripts/skill-size-check.sh"
PASS=0
FAIL=0
pass() { echo " PASS: $1"; PASS=$((PASS + 1)); }
fail() { echo " FAIL: $1"; FAIL=$((FAIL + 1)); }
TMPDIR_T="$(mktemp -d)"
trap 'rm -rf "$TMPDIR_T"' EXIT
# A description with a boundary clause and no routing target — the neutral
# default, so a fixture about Gotchas or references does not also trip the
# missing-boundary-clause SUGGESTION and stop isolating what it names.
CLEAN_DESC="Use when doing the thing. Do not use for anything else."
# make_skill <name> <desc> — SKILL.md with the body read from stdin. Echoes the
# path. Each skill gets its own directory so references/ fixtures are isolated.
make_skill() {
local name="$1" desc="$2" dir
dir="$TMPDIR_T/$name"
mkdir -p "$dir"
{
echo "---"
echo "name: $name"
echo "description: $desc"
echo "---"
cat
} > "$dir/SKILL.md"
echo "$dir/SKILL.md"
}
# expect <label> <file> <expect: silent|suggests|errors> [needle] [absent-needle]
#
# `silent` is the strict one: exit 0 AND completely empty output. It is what
# makes every "must not fire" case below real — a check that fired with some
# other wording would still be caught.
expect() {
local label="$1" file="$2" mode="$3" needle="${4:-}" absent="${5:-}" out status=0
set +e
out="$(bash "$HOOK" "$file" 2>&1)"
status=$?
set -e
case "$mode" in
silent)
if [[ $status -eq 0 && -z "$out" ]]; then
pass "$label"
else
fail "$label (exit $status, output: ${out:-<empty>})"
fi
;;
suggests)
if [[ $status -ne 0 ]]; then
fail "$label — a SUGGESTION must never change the exit code (exit $status, output: $out)"
elif [[ "$out" != *"SUGGESTION"* || "$out" != *"$needle"* ]]; then
fail "$label (exit $status, output: ${out:-<empty>})"
elif [[ -n "$absent" && "$out" == *"$absent"* ]]; then
fail "$label — output also contained '$absent', which must not fire here: $out"
else
pass "$label"
fi
;;
errors)
if [[ $status -eq 0 ]]; then
fail "$label — expected a hard ERROR, got exit 0 (output: ${out:-<empty>})"
elif [[ "$out" != *"$needle"* ]]; then
fail "$label (exit $status, output: ${out:-<empty>})"
else
pass "$label"
fi
;;
quiet-about)
# Exit 0 and the named text absent, but other output permitted. Used where
# a second, unrelated finding legitimately fires.
if [[ $status -ne 0 ]]; then
fail "$label (exit $status, output: $out)"
elif [[ "$out" == *"$needle"* ]]; then
fail "$label — '$needle' fired when it must not: $out"
else
pass "$label"
fi
;;
esac
}
filler() { python3 -c "print(' '.join(['word'] * $1))"; }
# ---------------------------------------------------------------------------
# Gotchas: entry count (guideline 5)
# ---------------------------------------------------------------------------
# The filler after the section keeps the 25% fraction check well clear, so these
# two cases isolate the ENTRY count. Without it a six-entry section in a short
# body would fire both and the pair would not distinguish them.
echo ""
echo "--- Gotchas entry count: 5 is fine, 6 is a SUGGESTION ---"
F_FIVE="$(make_skill gotchas-five "$CLEAN_DESC" <<EOF
## Gotchas
- first trap here
- second trap here
- third trap here
- fourth trap here
- fifth trap here
## Notes
$(filler 200)
EOF
)"
expect "a Gotchas section with exactly 5 entries is silent" "$F_FIVE" silent
F_SIX="$(make_skill gotchas-six "$CLEAN_DESC" <<EOF
## Gotchas
- first trap here
- second trap here
- third trap here
- fourth trap here
- fifth trap here
- sixth trap here
## Notes
$(filler 200)
EOF
)"
expect "a Gotchas section with 6 entries raises a SUGGESTION and still exits 0" \
"$F_SIX" suggests "Gotchas section has 6 entries" "over the 25% guideline"
# ---------------------------------------------------------------------------
# Gotchas: share of the body (guideline 25%)
# ---------------------------------------------------------------------------
# Exact boundary arithmetic, not an approximation. The section is prose (no list
# items) so the entry check cannot fire and confuse the result; body words are
# then section + filler + the two two-word headings. At a body of 100 words a
# 25-word section is exactly the guideline (inclusive — `>` is the comparison, so
# it passes) and a 26-word section is one word past it.
echo ""
echo "--- Gotchas share of body: exactly 25% is fine, 26% is a SUGGESTION ---"
F_AT="$(make_skill gotchas-at-fraction "$CLEAN_DESC" <<EOF
## Gotchas
$(filler 25)
## Notes
$(filler 71)
EOF
)"
expect "a Gotchas section at exactly 25% of the body is silent" "$F_AT" silent
F_OVER="$(make_skill gotchas-over-fraction "$CLEAN_DESC" <<EOF
## Gotchas
$(filler 26)
## Notes
$(filler 70)
EOF
)"
expect "a Gotchas section at 26% of the body raises a SUGGESTION" \
"$F_OVER" suggests "Gotchas section is 26 of 100 body words (26%)" "entries"
# ---------------------------------------------------------------------------
# Gotchas: false positives
# ---------------------------------------------------------------------------
echo ""
echo "--- a ## Gotchas heading inside a fenced block is not the Gotchas section ---"
# skill-author and skill-audit both document this convention by showing it. If a
# fenced example counted, the two skills that define the rule would be the two
# most likely to fail it.
F_FENCED_HEADING="$(make_skill gotchas-fenced-heading "$CLEAN_DESC" <<EOF
## How to write one
\`\`\`markdown
## Gotchas
- example one
- example two
- example three
- example four
- example five
- example six
- example seven
\`\`\`
$(filler 200)
EOF
)"
expect "a fenced ## Gotchas heading is not read as the section" \
"$F_FENCED_HEADING" quiet-about "Gotchas section"
echo ""
echo "--- list items inside a fenced block do not count as Gotchas entries ---"
F_FENCED_ENTRIES="$(make_skill gotchas-fenced-entries "$CLEAN_DESC" <<EOF
## Gotchas
- a real trap
- another real trap
Shown as an example of what NOT to write:
\`\`\`markdown
- fake one
- fake two
- fake three
- fake four
- fake five
- fake six
- fake seven
- fake eight
\`\`\`
## Notes
$(filler 300)
EOF
)"
expect "eight fenced bullets plus two real ones counts as two entries, not ten" \
"$F_FENCED_ENTRIES" quiet-about "Gotchas section has"
echo ""
echo "--- the heading must END in gotcha(s): '## Gotcha handling' is not the section ---"
# `## Gotcha handling` and `## Why gotchas matter` are prose sections. Treating
# one as the Gotchas section measures a span that was never a gotcha list.
F_HANDLING="$(make_skill gotcha-handling "$CLEAN_DESC" <<EOF
## Gotcha handling
- item one
- item two
- item three
- item four
- item five
- item six
- item seven
## Notes
$(filler 200)
EOF
)"
expect "'## Gotcha handling' is not matched as the Gotchas section" \
"$F_HANDLING" quiet-about "Gotchas section"
# The control for the three cases above. Without it, "no Gotchas finding" could
# equally mean the whole check is dead, and all three would still be green.
F_CONTROL="$(make_skill gotchas-control "$CLEAN_DESC" <<EOF
## Common gotchas
- item one
- item two
- item three
- item four
- item five
- item six
- item seven
## Notes
$(filler 200)
EOF
)"
expect "control: a real '## Common gotchas' heading with 7 entries IS matched" \
"$F_CONTROL" suggests "Gotchas section has 7 entries"
# ---------------------------------------------------------------------------
# references/<file>.md pointers
# ---------------------------------------------------------------------------
echo ""
echo "--- a references/ pointer that is not on disk is a hard ERROR ---"
F_REF_MISSING="$(make_skill ref-missing "$CLEAN_DESC" <<EOF
If the caller needs the long form, read references/nowhere.md first.
EOF
)"
expect "a body pointing at an absent references/nowhere.md ERRORs" \
"$F_REF_MISSING" errors "points at references/nowhere.md"
F_REF_PRESENT="$(make_skill ref-present "$CLEAN_DESC" <<EOF
If the caller needs the long form, read references/here.md first.
EOF
)"
mkdir -p "$TMPDIR_T/ref-present/references"
echo "content" > "$TMPDIR_T/ref-present/references/here.md"
expect "the same pointer is silent once the file exists" "$F_REF_PRESENT" silent
echo ""
echo "--- a references/ pointer inside a fenced block is not a dispatch entry ---"
F_REF_FENCED="$(make_skill ref-fenced "$CLEAN_DESC" <<EOF
Dispatch tables look like this:
\`\`\`markdown
If X, read references/example-file.md.
\`\`\`
EOF
)"
expect "a fenced references/example-file.md does not ERROR" "$F_REF_FENCED" silent
echo ""
echo "--- a references/ pointer in a same-line removal context is history, not dispatch ---"
# Narrow on purpose: a live dispatch table never describes its own target as
# removed, so the exemption costs no recall. Each phrasing is checked separately
# because they are separate alternatives in one regex, and a typo in any one of
# them turns ordinary prose back into a hard ERROR.
#
# The file names are deliberately NEUTRAL (detail-a.md, not gone-a.md). An
# earlier draft of this block named them gone-*.md and every case passed for the
# wrong reason: "gone" is itself one of the removal words, so the exemption fired
# off the FILENAME and the phrase under test was never exercised. The control
# below is what surfaced that — it is the assertion that keeps these five honest.
i=0
for phrase in \
"The old references/detail-a.md was removed in v2." \
"references/detail-b.md is no longer part of this skill." \
"references/detail-c.md is deprecated and should not be read." \
"references/detail-d.md was renamed, so nothing points at it now." \
"references/detail-e.md has been superseded by the body itself."; do
i=$((i + 1))
F_REF_PAST="$(make_skill "ref-past-$i" "$CLEAN_DESC" <<EOF
$phrase
EOF
)"
expect "removal-context pointer is not an ERROR: \"$phrase\"" "$F_REF_PAST" silent
done
# The control: the SAME sentence shape without a removal word must still ERROR,
# or the exemption above has swallowed the check rather than narrowed it.
F_REF_LIVE="$(make_skill ref-live "$CLEAN_DESC" <<EOF
The details live in references/detail-a.md, which the agent should read first.
EOF
)"
expect "control: the same pointer with no removal word still ERRORs" \
"$F_REF_LIVE" errors "points at references/detail-a.md"
# ---------------------------------------------------------------------------
# Missing boundary clause
# ---------------------------------------------------------------------------
echo ""
echo "--- a description with no boundary clause raises a SUGGESTION ---"
F_NO_BOUNDARY="$(make_skill no-boundary "Use when the user wants the thing done." <<EOF
Do the thing.
EOF
)"
expect "a description with no boundary clause raises a SUGGESTION and still exits 0" \
"$F_NO_BOUNDARY" suggests "description has no boundary clause"
# Both accepted shapes, asserted separately: the prose markers and ADR-0020's
# compressed arrow form. Dropping either from the detector would leave the other
# green.
F_PROSE_BOUNDARY="$(make_skill prose-boundary "Use when the user wants the thing done. Do not use for anything else." <<EOF
Do the thing.
EOF
)"
expect "the prose boundary form satisfies the check" "$F_PROSE_BOUNDARY" silent
F_ARROW_BOUNDARY="$(make_skill arrow-boundary "Use when the user wants the thing done. Not the other thing -> sibling-skill." <<EOF
Do the thing.
EOF
)"
expect "ADR-0020's compressed 'Not X -> y' form satisfies the check" \
"$F_ARROW_BOUNDARY" quiet-about "no boundary clause"
echo ""
echo "Results: $PASS passed, $FAIL failed"
[[ $FAIL -eq 0 ]]

331
tests/test-adr0020-contract.sh Executable file
View File

@@ -0,0 +1,331 @@
#!/usr/bin/env bash
# Regression test for the three STRUCTURAL claims the ADR-0020 gate family makes
# about itself. None of them was pinned anywhere before this file, and each one
# fails silently — which is the whole reason they need a test rather than a
# comment:
#
# 1. "ONE resolver, embedded VERBATIM in three scripts." The block between the
# BEGIN/END markers is copied, not imported, because a cache-installed
# plugin's scripts cannot read files outside their own plugin directory.
# Nothing but this file asserts the three copies are still identical, and a
# one-line edit to a single copy is invisible: every constant-agreement
# assertion in tests/test-skill-size-check.sh still passes, because the
# CONSTANTS are not what drifted.
# 2. Both interpreter preflights, in all three scripts. python3 and PyYAML are
# declared HARD dependencies precisely so a missing one cannot turn into a
# vacuous pass, and the two are checked separately so the message names the
# thing to install rather than the wrong one.
# 3. `verbose: true` on the skill-size-check hook. It is the ENTIRE delivery
# mechanism for the SUGGESTION tier: pre-commit prints nothing at all for a
# passing hook, and a SUGGESTION deliberately does not fail, so dropping
# one word from the config silences the tier ADR-0020 depends on while
# every test and every hook still reports green.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
HOOK="$REPO_ROOT/scripts/skill-size-check.sh"
SKILL_VALIDATE="$REPO_ROOT/plugins/kyberforge/.apm/skills/skill-audit/scripts/validate.sh"
AGENT_VALIDATE="$REPO_ROOT/plugins/kyberforge/.apm/skills/agent-audit/scripts/validate.sh"
PASS=0
FAIL=0
pass() { echo " PASS: $1"; PASS=$((PASS + 1)); }
fail() { echo " FAIL: $1"; FAIL=$((FAIL + 1)); }
TMPDIR_T="$(mktemp -d)"
trap 'rm -rf "$TMPDIR_T"' EXIT
BEGIN_MARKER='# ===== BEGIN ADR-0020 SHARED BOUNDARY RESOLVER ====='
END_MARKER='# ===== END ADR-0020 SHARED BOUNDARY RESOLVER ====='
# ---------------------------------------------------------------------------
# 1. The shared resolver block is byte-identical in all three scripts
# ---------------------------------------------------------------------------
echo ""
echo "--- the ADR-0020 shared resolver block is byte-identical in all three scripts ---"
# Marker discipline first. An unbalanced or duplicated marker pair makes the
# extraction below silently measure the wrong span — a sed range that never
# closes swallows the rest of the file, and one that opens twice concatenates
# two spans. Both would still compare "equal" if all three were mangled the
# same way, so the shape is asserted before the contents.
MARKERS_OK=true
for f in "$HOOK" "$SKILL_VALIDATE" "$AGENT_VALIDATE"; do
if [[ ! -f "$f" ]]; then
fail "script not found: $f"
MARKERS_OK=false
continue
fi
b="$(grep -cFx "$BEGIN_MARKER" "$f" || true)"
e="$(grep -cFx "$END_MARKER" "$f" || true)"
if [[ "$b" == "1" && "$e" == "1" ]]; then
pass "${f#"$REPO_ROOT/"} carries exactly one BEGIN and one END marker"
else
fail "${f#"$REPO_ROOT/"} has $b BEGIN and $e END markers, expected 1 and 1"
MARKERS_OK=false
fi
done
if ! $MARKERS_OK; then
fail "skipping the byte-identity comparison — the marker pairs are not well-formed, so any extraction would measure the wrong span"
else
HASHES=()
LINECOUNTS=()
for f in "$HOOK" "$SKILL_VALIDATE" "$AGENT_VALIDATE"; do
out="$TMPDIR_T/block-$(echo "$f" | md5sum | cut -c1-8).txt"
sed -n "/^${BEGIN_MARKER}\$/,/^${END_MARKER}\$/p" "$f" > "$out"
HASHES+=("$(md5sum < "$out" | cut -d' ' -f1)")
LINECOUNTS+=("$(wc -l < "$out" | tr -d ' ')")
done
if [[ "${HASHES[0]}" == "${HASHES[1]}" && "${HASHES[1]}" == "${HASHES[2]}" ]]; then
pass "all three copies hash to ${HASHES[0]} (${LINECOUNTS[0]} lines) — agreement by construction, not by coincidence"
else
fail "the shared resolver has DRIFTED: skill-size-check=${HASHES[0]} (${LINECOUNTS[0]} lines), skill-audit=${HASHES[1]} (${LINECOUNTS[1]} lines), agent-audit=${HASHES[2]} (${LINECOUNTS[2]} lines). Edit one copy, then paste it over the other two."
fi
# A block that has been emptied out would hash equal in all three and pass the
# comparison above while enforcing nothing. The resolver is ~570 lines; 100 is
# a floor low enough never to need maintenance and high enough that a gutted
# block cannot sneak past.
if [[ "${LINECOUNTS[0]}" -gt 100 ]]; then
pass "the extracted block is ${LINECOUNTS[0]} lines — the comparison is over real content, not an empty span"
else
fail "the extracted shared block is only ${LINECOUNTS[0]} lines — three identical empty spans would compare equal and assert nothing"
fi
fi
# ---------------------------------------------------------------------------
# 2. Both interpreter preflights, in all three scripts
# ---------------------------------------------------------------------------
# The two are checked separately on purpose: `python3 -c 'import yaml'` fails
# identically whether python3 is missing or PyYAML is, and naming the wrong one
# sends the reader to install the wrong thing.
REAL_PYTHON="$(command -v python3)"
# Absolute path, deliberately. The no-python3 fixture below replaces PATH
# wholesale, so a bare `bash` (or `/usr/bin/env bash`) would be resolved against
# that stripped PATH and die with "No such file or directory" before the script
# under test ever starts -- a 127 that looks like the preflight firing.
BASH_BIN="$(command -v bash)"
# A PATH that genuinely has no python3 on it. Built by symlinking the handful of
# binaries the three scripts touch before their own preflight rather than by
# hiding python3 from a full PATH, because there is no portable way to subtract
# one entry from a directory. `bash` is invoked by absolute path below so the
# interpreter itself does not have to be on this PATH.
NOPY_BIN="$TMPDIR_T/nopython-bin"
mkdir -p "$NOPY_BIN"
for b in awk cat cut dirname basename grep sed pwd rm mkdir tr; do
src="$(command -v "$b" 2>/dev/null || true)"
[[ -n "$src" ]] && ln -sf "$src" "$NOPY_BIN/$b"
done
# A python3 that runs but cannot import yaml. A shim on PATH re-execs the real
# interpreter with a PYTHONPATH entry holding a `yaml` module that raises on
# import; PYTHONPATH precedes site-packages on sys.path, so it shadows a real
# PyYAML install without touching it.
SHADOW="$TMPDIR_T/shadow"
mkdir -p "$SHADOW"
printf 'raise ImportError("PyYAML deliberately unavailable in this fixture")\n' \
> "$SHADOW/yaml.py"
NOYAML_BIN="$TMPDIR_T/noyaml-bin"
mkdir -p "$NOYAML_BIN"
cat > "$NOYAML_BIN/python3" <<EOF
#!/bin/sh
PYTHONPATH="$SHADOW\${PYTHONPATH:+:\$PYTHONPATH}" exec "$REAL_PYTHON" "\$@"
EOF
chmod +x "$NOYAML_BIN/python3"
# Sanity-check the two fixtures themselves before trusting any verdict they
# produce. A shim that silently still imports yaml would make every PyYAML
# assertion below pass for the wrong reason.
if PATH="$NOYAML_BIN:$PATH" python3 -c 'import yaml' 2>/dev/null; then
fail "the no-PyYAML shim does not actually shadow PyYAML — every PyYAML assertion below would be vacuous"
else
pass "fixture check: the no-PyYAML shim makes 'import yaml' fail while python3 still runs"
fi
if PATH="$NOPY_BIN" command -v python3 > /dev/null 2>&1; then
fail "the no-python3 PATH still resolves python3 — every python3 assertion below would be vacuous"
else
pass "fixture check: the no-python3 PATH resolves no python3"
fi
# A minimal, entirely clean subject for each script. The preflight must fire
# before any measurement, so the subject's own content is irrelevant — which is
# exactly what makes a clean one the right choice: nothing else can produce the
# non-zero exit these cases assert.
SUBJECT_SKILL_DIR="$TMPDIR_T/subject/my-skill"
mkdir -p "$SUBJECT_SKILL_DIR"
cat > "$SUBJECT_SKILL_DIR/SKILL.md" <<'EOF'
---
name: my-skill
description: A short valid description. Do not use for anything else.
---
Do the thing.
EOF
SUBJECT_AGENT_ROOT="$TMPDIR_T/subject-agent"
mkdir -p "$SUBJECT_AGENT_ROOT/.apm/agents"
cat > "$SUBJECT_AGENT_ROOT/apm.yml" <<'EOF'
name: test-package
version: 0.1.0
type: skill
EOF
cat > "$SUBJECT_AGENT_ROOT/.apm/agents/my-agent.agent.md" <<'EOF'
---
name: my-agent
description: A short valid description. Do not use for anything else.
---
You are a test agent. When invoked, do the thing.
EOF
# probe_preflight <label> <env-kind: nopython|noyaml> <expect-needle> <cmd...>
probe_preflight() {
local label="$1" kind="$2" needle="$3"
shift 3
local out status=0
set +e
if [[ "$kind" == nopython ]]; then
out="$(env -i PATH="$NOPY_BIN" HOME="$HOME" "$BASH_BIN" "$@" 2>&1)"
else
out="$(env PATH="$NOYAML_BIN:$PATH" "$BASH_BIN" "$@" 2>&1)"
fi
status=$?
set -e
if [[ $status -eq 0 ]]; then
fail "$label exited 0 — a missing hard dependency became a vacuous pass (output: ${out:-<empty>})"
elif [[ "$out" != *"$needle"* ]]; then
# The needle is the DIAGNOSTIC ("python3 is required"), not the bare word.
# Deleting the preflight entirely would still produce a non-zero exit and a
# message mentioning python3 -- bash's own "python3: command not found" --
# so a bare-word needle would go green on a script with no preflight at all.
fail "$label exited $status but never produced the '$needle' diagnostic (output: ${out:-<empty>})"
elif [[ "$kind" == nopython && "$out" == *PyYAML* ]]; then
fail "$label reported PyYAML when python3 itself is missing — that sends the reader to install the wrong thing (output: $out)"
else
pass "$label"
fi
}
echo ""
echo "--- a PATH with no python3 is a hard failure in all three scripts, naming python3 ---"
probe_preflight "scripts/skill-size-check.sh reports missing python3" \
nopython "python3 is required" \
"$HOOK" "$SUBJECT_SKILL_DIR/SKILL.md"
probe_preflight "skill-audit/scripts/validate.sh reports missing python3" \
nopython "python3 is required" \
"$SKILL_VALIDATE" "$SUBJECT_SKILL_DIR"
probe_preflight "agent-audit/scripts/validate.sh reports missing python3" \
nopython "python3 is required" \
"$AGENT_VALIDATE" "$SUBJECT_AGENT_ROOT/.apm/agents/my-agent.agent.md"
echo ""
echo "--- a python3 that cannot import yaml is a hard failure in all three scripts, naming PyYAML ---"
probe_preflight "scripts/skill-size-check.sh reports missing PyYAML" \
noyaml "PyYAML is required" \
"$HOOK" "$SUBJECT_SKILL_DIR/SKILL.md"
probe_preflight "skill-audit/scripts/validate.sh reports missing PyYAML" \
noyaml "PyYAML is required" \
"$SKILL_VALIDATE" "$SUBJECT_SKILL_DIR"
probe_preflight "agent-audit/scripts/validate.sh reports missing PyYAML" \
noyaml "PyYAML is required" \
"$AGENT_VALIDATE" "$SUBJECT_AGENT_ROOT/.apm/agents/my-agent.agent.md"
# The control. Without it, "fails when the dependency is missing" is satisfied by
# a script that fails unconditionally, and the two cases above would be green on
# a gate that never runs at all.
echo ""
echo "--- control: with both dependencies present the same subjects pass ---"
for probe in "$HOOK:$SUBJECT_SKILL_DIR/SKILL.md" \
"$SKILL_VALIDATE:$SUBJECT_SKILL_DIR" \
"$AGENT_VALIDATE:$SUBJECT_AGENT_ROOT/.apm/agents/my-agent.agent.md"; do
script="${probe%%:*}"
arg="${probe#*:}"
set +e
ctl_out="$(bash "$script" "$arg" 2>&1)"
ctl_rc=$?
set -e
if [[ $ctl_rc -eq 0 ]]; then
pass "${script#"$REPO_ROOT/"} exits 0 on a clean subject with python3 and PyYAML available"
else
fail "${script#"$REPO_ROOT/"} failed a clean subject (exit $ctl_rc): $ctl_out"
fi
done
# ---------------------------------------------------------------------------
# 3. verbose: true on the skill-size-check hook, in BOTH manifests
# ---------------------------------------------------------------------------
# .pre-commit-config.yaml governs this repo; .pre-commit-hooks.yaml is what a
# CONSUMER repo gets when it points at this one. Dropping the flag from either
# silences the SUGGESTION tier for that audience alone, which is the hardest
# version of the defect to notice.
echo ""
echo "--- the skill-size-check hook declares verbose: true in both manifests ---"
VERBOSE_REPORT="$(python3 - "$REPO_ROOT" <<'PY'
import os
import sys
import yaml
root = sys.argv[1]
def emit(status, msg):
print("%s\t%s" % (status, msg))
# Repo config: nested repos[].hooks[].
path = os.path.join(root, '.pre-commit-config.yaml')
try:
with open(path, encoding='utf-8') as fh:
cfg = yaml.safe_load(fh) or {}
except Exception as exc:
emit('FAIL', '.pre-commit-config.yaml did not parse: %s' % exc)
cfg = {}
found = None
for repo in cfg.get('repos') or []:
for hook in (repo.get('hooks') or []):
if hook.get('id') == 'skill-size-check':
found = hook
if found is None:
emit('FAIL', '.pre-commit-config.yaml declares no hook with id skill-size-check')
elif found.get('verbose') is True:
emit('PASS', '.pre-commit-config.yaml: skill-size-check is verbose: true')
else:
emit('FAIL', '.pre-commit-config.yaml: skill-size-check has verbose=%r — '
'pre-commit prints nothing for a passing hook, so every '
'ADR-0020 SUGGESTION is swallowed' % (found.get('verbose'),))
# Consumer manifest: a flat list of hooks.
path = os.path.join(root, '.pre-commit-hooks.yaml')
try:
with open(path, encoding='utf-8') as fh:
hooks = yaml.safe_load(fh) or []
except Exception as exc:
emit('FAIL', '.pre-commit-hooks.yaml did not parse: %s' % exc)
hooks = []
found = None
for hook in hooks:
if isinstance(hook, dict) and hook.get('id') == 'kyberforge-skill-size-check':
found = hook
if found is None:
emit('FAIL', '.pre-commit-hooks.yaml declares no hook with id kyberforge-skill-size-check')
elif found.get('verbose') is True:
emit('PASS', '.pre-commit-hooks.yaml: kyberforge-skill-size-check is verbose: true')
else:
emit('FAIL', '.pre-commit-hooks.yaml: kyberforge-skill-size-check has verbose=%r — '
'a consumer repo would never see the SUGGESTION tier'
% (found.get('verbose'),))
PY
)"
while IFS=$'\t' read -r status msg; do
[[ -n "$status" ]] || continue
if [[ "$status" == PASS ]]; then
pass "$msg"
else
fail "$msg"
fi
done <<< "$VERBOSE_REPORT"
echo ""
echo "Results: $PASS passed, $FAIL failed"
[[ $FAIL -eq 0 ]]

View File

@@ -0,0 +1,365 @@
#!/usr/bin/env bash
# Differential test: scripts/skill-size-check.sh (the pre-commit hook) and
# skill-audit/scripts/validate.sh (the in-skill auditor) must reach the SAME
# ADR-0020 verdict on the same file.
#
# Why this exists as a separate suite. tests/test-skill-size-check.sh already
# asserts the two agree on their CONSTANTS, and that assertion is necessary but
# demonstrably not sufficient: a previous review found the two scripts disagreeing
# on real files while every constant matched perfectly. Constants are one of the
# ways two hand-duplicated implementations diverge; comparison operators, message
# wording, which value gets measured, and which branch runs first are the others,
# and none of them is visible to a constant check.
#
# The consequence of divergence is specific and bad: skill-audit reports a skill
# ready to ship and the commit hook then rejects it, or worse, the reverse. So the
# comparison here is over VERDICTS on files, not over source text.
#
# Scope: the ADR-0020 axes the two scripts share — description length and tier,
# body word count and tier, dangling routing targets, missing references/
# pointers, the two Gotchas suggestions, the missing-boundary-clause suggestion,
# a declined resolution, and an empty description. The two scripts legitimately
# differ elsewhere (validate.sh also checks name/directory agreement, script
# executability and the 1024-char spec backstop; the hook checks whole-file lines
# and words), and those lines are ignored rather than being forced into a shared
# shape they were never meant to have.
#
# Run over the real 39-skill corpus AND over purpose-built fixtures that sit ON
# each boundary. The corpus alone is not enough — it happens not to contain a
# file at exactly 900 body words, which is precisely where an inclusive/exclusive
# comparison mismatch would hide.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
HOOK="$REPO_ROOT/scripts/skill-size-check.sh"
SKILL_VALIDATE="$REPO_ROOT/plugins/kyberforge/.apm/skills/skill-audit/scripts/validate.sh"
TMPDIR_T="$(mktemp -d)"
trap 'rm -rf "$TMPDIR_T"' EXIT
# ---------------------------------------------------------------------------
# Fixtures: one per ADR-0020 axis, placed ON the boundary wherever there is one.
# ---------------------------------------------------------------------------
# Built inside a synthetic plugin monorepo so boundary-target resolution actually
# runs for both scripts (in a bare temp dir both would decline, and "both
# declined" is agreement about nothing).
FIXTURE_ROOT="$TMPDIR_T/fixtures"
FX="$FIXTURE_ROOT/plugins/fixture-plugin/.apm/skills"
mkdir -p "$FX/sibling-skill" "$FIXTURE_ROOT/plugins/fixture-plugin/.apm/agents"
make_fx() {
local name="$1" desc="$2" body_words="$3"
mkdir -p "$FX/$name"
{
echo "---"
echo "name: $name"
echo "description: $desc"
echo "---"
echo ""
python3 -c "print(' '.join(['word'] * $body_words))"
} > "$FX/$name/SKILL.md"
}
desc_of_length() {
python3 - "$1" <<'PY'
import sys
n = int(sys.argv[1])
prefix = 'Use when doing the thing. Do not use for anything else. '
print(prefix + 'x' * (n - len(prefix)))
PY
}
CLEAN="Use when doing the thing. Do not use for anything else."
# Description tier boundaries, both sides of both thresholds.
make_fx desc-249 "$(desc_of_length 249)" 10
make_fx desc-250 "$(desc_of_length 250)" 10
make_fx desc-251 "$(desc_of_length 251)" 10
make_fx desc-400 "$(desc_of_length 400)" 10
make_fx desc-401 "$(desc_of_length 401)" 10
# Body tier boundaries, both sides of both thresholds.
make_fx body-599 "$CLEAN" 599
make_fx body-600 "$CLEAN" 600
make_fx body-601 "$CLEAN" 601
make_fx body-900 "$CLEAN" 900
make_fx body-901 "$CLEAN" 901
# Folding: the value has to be measured after YAML folding in both scripts.
mkdir -p "$FX/folded-desc"
{
echo "---"
echo "name: folded-desc"
echo "description: >"
python3 -c "print('\n'.join([' ' + 'x' * 40] * 11))"
echo "---"
echo ""
echo "Do the thing."
} > "$FX/folded-desc/SKILL.md"
# Routing targets, one per tier the resolver can produce: resolves, dangles with
# in-sentence corroboration (ERROR/FAIL), dangles alone (SUGGESTION on both
# sides), route notation (ERROR/FAIL without corroboration), attributive
# (silent). Each tier is here because the two scripts have to agree on the TIER,
# not merely on the finding — a copy that promoted or demoted one of them would
# otherwise pass this comparison.
make_fx target-resolves "Use when doing the thing. Do not use for the other thing — use sibling-skill instead." 10
make_fx target-dangles "Use when doing the thing. Do not use for the other thing — use sibling-skill or no-such-skill instead." 10
make_fx target-dangles-lone "Use when doing the thing. Do not use for the other thing — use no-such-lone-skill instead." 10
make_fx target-dangles-notation "Use when doing the thing. Do not use for the other thing — use /no-such-notation-skill instead." 10
make_fx target-attributive "Use when doing the thing. Use pre-commit hooks instead of ad-hoc scripts." 10
# No boundary clause at all.
make_fx no-boundary "Use when the user wants the thing done." 10
# A missing references/ pointer, and a present one.
make_fx ref-missing "$CLEAN" 10
printf '\nIf the caller needs detail, read references/absent.md first.\n' >> "$FX/ref-missing/SKILL.md"
make_fx ref-present "$CLEAN" 10
printf '\nIf the caller needs detail, read references/there.md first.\n' >> "$FX/ref-present/SKILL.md"
mkdir -p "$FX/ref-present/references"
echo "detail" > "$FX/ref-present/references/there.md"
# Gotchas, over each guideline.
make_fx gotchas-many "$CLEAN" 0
cat >> "$FX/gotchas-many/SKILL.md" <<'EOF'
## Gotchas
- one trap here
- two trap here
- three trap here
- four trap here
- five trap here
- six trap here
- seven trap here
## Notes
EOF
python3 -c "print(' '.join(['word'] * 200))" >> "$FX/gotchas-many/SKILL.md"
# Gotchas over the 25% body-fraction guideline. Prose, not list items, so the
# entry guideline cannot fire and the two suggestions stay separable: 26 section
# words in a 100-word body is one word past the threshold.
make_fx gotchas-fraction "$CLEAN" 0
{
echo ""
echo "## Gotchas"
echo ""
python3 -c "print(' '.join(['word'] * 26))"
echo ""
echo "## Notes"
echo ""
python3 -c "print(' '.join(['word'] * 70))"
} >> "$FX/gotchas-fraction/SKILL.md"
# Empty description — the shape that used to exit 0 in silence.
mkdir -p "$FX/empty-desc"
printf -- '---\nname: empty-desc\ndescription:\nmodel: sonnet\n---\n\nDo the thing.\n' \
> "$FX/empty-desc/SKILL.md"
# Every boundary shape at once, so a divergence that only appears when several
# findings fire together is not missed.
make_fx combined "$(desc_of_length 401)" 901
printf '\nIf the caller needs detail, read references/absent.md first.\n' >> "$FX/combined/SKILL.md"
# A skill with routing targets and NO authoring root above it — deliberately
# OUTSIDE the fixture plugin tree. Both scripts must decline out loud, and both
# must decline identically; "both declined" is only meaningful as agreement if
# the declining path is exercised on purpose somewhere.
ORPHAN_ROOT="$TMPDIR_T/orphan"
mkdir -p "$ORPHAN_ROOT/no-universe"
{
echo "---"
echo "name: no-universe"
echo "description: Use when doing the thing. Do not use for the other thing — use some-other-skill instead."
echo "---"
echo ""
echo "Do the thing."
} > "$ORPHAN_ROOT/no-universe/SKILL.md"
# ---------------------------------------------------------------------------
# The comparison
# ---------------------------------------------------------------------------
python3 - "$HOOK" "$SKILL_VALIDATE" "$REPO_ROOT" "$FX" "$ORPHAN_ROOT" <<'PYTHON'
import glob
import os
import re
import subprocess
import sys
hook, validate, repo_root, fixture_dir, orphan_dir = sys.argv[1:6]
passes = 0
failures = 0
def ok(msg):
global passes
passes += 1
print(" PASS: %s" % msg)
def bad(msg):
global failures
failures += 1
print(" FAIL: %s" % msg)
# Tier prefixes. The hook writes `ERROR: ` / `SUGGESTION: ` / `INFO: `; the
# auditor writes `FAIL ` / `SUGGESTION ` / `INFO ` and additionally `PASS `
# lines, which carry no finding and are dropped.
TIERS = (
('ERROR', ('ERROR:', 'FAIL ')),
('SUGGESTION', ('SUGGESTION:', 'SUGGESTION ')),
('INFO', ('INFO:', 'INFO ')),
)
# Each rule turns a finding line into a canonical token. Wording differs between
# the two scripts by design (one addresses a committer, the other an auditor), so
# the tokens deliberately capture the MEASUREMENT and not the sentence.
RULES = (
('DESC_CHARS', re.compile(r'description is (\d+) char')),
('BODY_WORDS', re.compile(r'body is (\d+) words')),
('ROUTE', re.compile(r"routes to '([^']+)'")),
('MISSING_REF', re.compile(r'points at (references/[^\s,]+)')),
('GOTCHA_ENTRIES', re.compile(r'Gotchas section has (\d+) entries')),
('GOTCHA_FRACTION', re.compile(r'Gotchas section is (\d+) of (\d+) body words')),
('NO_BOUNDARY_CLAUSE', re.compile(r'(description has no boundary clause)')),
('RESOLUTION_DECLINED', re.compile(r'(boundary-target resolution DID NOT RUN)')),
('DESC_EMPTY', re.compile(r'(description field is missing or empty)')),
)
def verdict(output):
"""The set of ADR-0020 findings in a script's output, tier included.
Lines that match no rule are dropped rather than compared: the two scripts
legitimately check different things outside ADR-0020 (name/directory
agreement, script executability, the 1024-char spec backstop, whole-file
line and word ceilings), and forcing those into the comparison would report
a difference that is not a disagreement.
"""
found = set()
for raw in output.splitlines():
line = raw.strip()
tier = None
for name, prefixes in TIERS:
if any(line.startswith(p) for p in prefixes):
tier = name
break
if tier is None:
continue
for token, pattern in RULES:
match = pattern.search(line)
if match:
found.add((tier, token) + tuple(match.groups()))
return found
def run(cmd):
proc = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT)
return proc.returncode, proc.stdout.decode('utf-8', 'replace')
def compare(label, skill_dir):
skill_md = os.path.join(skill_dir, 'SKILL.md')
hook_rc, hook_out = run(['bash', hook, skill_md])
audit_rc, audit_out = run(['bash', validate, skill_dir])
hook_v = verdict(hook_out)
audit_v = verdict(audit_out)
problems = []
only_hook = sorted(hook_v - audit_v)
only_audit = sorted(audit_v - hook_v)
if only_hook:
problems.append('only the hook reported %s' % (only_hook,))
if only_audit:
problems.append('only skill-audit reported %s' % (only_audit,))
# Exit codes are compared on the ADR-0020 axis only: an ERROR-tier ADR-0020
# finding must make BOTH scripts non-zero, and neither may be turned
# non-zero by a SUGGESTION. The raw codes cannot be compared directly —
# validate.sh also fails on checks the hook does not run at all.
hook_err = any(t == 'ERROR' for t, *_ in hook_v)
audit_err = any(t == 'ERROR' for t, *_ in audit_v)
if hook_err and hook_rc == 0:
problems.append('the hook reported an ADR-0020 ERROR but exited 0')
if audit_err and audit_rc == 0:
problems.append('skill-audit reported an ADR-0020 FAIL but exited 0')
if not hook_err and hook_rc != 0 and not _non_adr_hook_error(hook_out):
problems.append('the hook exited %d with no ADR-0020 ERROR and no spec-ceiling ERROR'
% hook_rc)
if problems:
bad('%s: %s' % (label, '; '.join(problems)))
else:
return True
return False
def _non_adr_hook_error(output):
"""True if the hook failed on a spec ceiling rather than an ADR-0020 gate.
MAX_LINES / MAX_WORDS are the hook's other ERROR sources and are outside
this comparison, so a non-zero exit explained by one of them is not a
disagreement.
"""
return bool(re.search(r'ERROR: .*(-line ceiling|-word ceiling \(~5,000 tokens)', output))
# --- The real corpus -------------------------------------------------------
corpus = sorted(glob.glob(os.path.join(repo_root, 'plugins', '*', '.apm', 'skills', '*')))
corpus = [d for d in corpus if os.path.isfile(os.path.join(d, 'SKILL.md'))]
print("")
print("--- the two scripts agree on every skill in the live corpus (%d files) ---" % len(corpus))
if len(corpus) < 30:
bad('only %d corpus skills were discovered — the glob is wrong, so this leg '
'proves nothing' % len(corpus))
else:
ok('discovered %d corpus skills to compare' % len(corpus))
agreed = 0
for skill_dir in corpus:
rel = os.path.relpath(skill_dir, repo_root)
if compare(rel, skill_dir):
agreed += 1
if agreed == len(corpus):
ok('all %d corpus skills produce identical ADR-0020 verdicts from both scripts' % agreed)
# --- Boundary fixtures -----------------------------------------------------
fixtures = sorted(d for d in (glob.glob(os.path.join(fixture_dir, '*'))
+ glob.glob(os.path.join(orphan_dir, '*')))
if os.path.isfile(os.path.join(d, 'SKILL.md')))
print("")
print("--- the two scripts agree on every boundary fixture (%d files) ---" % len(fixtures))
if len(fixtures) < 15:
bad('only %d fixtures were built — the fixture set is incomplete, so the '
'boundaries the corpus does not cover are untested' % len(fixtures))
else:
ok('built %d boundary fixtures to compare' % len(fixtures))
fx_agreed = 0
for skill_dir in fixtures:
if compare(os.path.basename(skill_dir), skill_dir):
fx_agreed += 1
if fx_agreed == len(fixtures):
ok('all %d boundary fixtures produce identical ADR-0020 verdicts from both scripts' % fx_agreed)
# --- The comparison must not be vacuous ------------------------------------
# Everything above would also pass if verdict() extracted nothing at all. So the
# fixtures are required to have produced findings across every axis this suite
# claims to compare — if a rule stops matching (a reworded message, say), that is
# a silent loss of coverage and it fails here instead.
print("")
print("--- the comparison actually extracted findings on every axis it claims to cover ---")
seen_tokens = set()
for skill_dir in fixtures:
_, out = run(['bash', hook, os.path.join(skill_dir, 'SKILL.md')])
for entry in verdict(out):
seen_tokens.add(entry[1])
_, out = run(['bash', validate, skill_dir])
for entry in verdict(out):
seen_tokens.add(entry[1])
expected_tokens = {t for t, _ in RULES}
missing = sorted(expected_tokens - seen_tokens)
if missing:
bad('no fixture produced a finding for %s — verdict() may no longer match '
'those messages, and any disagreement on them would go unseen' % missing)
else:
ok('every one of the %d compared axes was exercised by at least one fixture'
% len(expected_tokens))
print("")
print("Results: %d passed, %d failed" % (passes, failures))
sys.exit(1 if failures else 0)
PYTHON

286
tests/test-adr0020-frontmatter.sh Executable file
View File

@@ -0,0 +1,286 @@
#!/usr/bin/env bash
# Regression test for the two ways an ADR-0020 gate can be made to check NOTHING
# while still exiting 0. Both were live defects, both were silent, and both sit
# in the shared resolver block that all three scripts embed verbatim — so every
# case below runs against all three.
#
# 1. THE FRONTMATTER BLOCKER. The frontmatter matcher used to be `^---\n`. A
# UTF-8 BOM, a leading blank line, a trailing space after either marker, or
# CRLF line endings all defeated it, and the miss was not reported: every
# ADR-0020 check was skipped and the file passed. Measured at the time: a
# 550-character description with a 1,000-word body exited 0 behind a BOM.
# So this file asserts two complementary things — that each of those four
# shapes is now TOLERATED (the findings actually fire), and that
# frontmatter which genuinely cannot be parsed is a hard ERROR rather than
# a quiet skip. A file that cannot be measured must never report green.
#
# 2. THE VALUELESS DESCRIPTION. `description:` with no value, followed by
# another key, let a line regex's `\s*` cross the newline and capture the
# NEXT key. The value then looked present (so "missing or empty" never
# fired) and was empty once folded (so every ADR-0020 gate early-returned).
# An agent file with one exited 0 with zero output through a BLOCKING
# pre-push gate. All five spellings of "no value" are pinned here.
#
# Both fixtures carry an over-ceiling description AND an over-ceiling body on
# purpose: asserting a non-zero exit alone would be satisfied by the "cannot
# parse" error itself, so the tolerated shapes are asserted on the CONTENT of
# the findings, not on the exit code.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
HOOK="$REPO_ROOT/scripts/skill-size-check.sh"
SKILL_VALIDATE="$REPO_ROOT/plugins/kyberforge/.apm/skills/skill-audit/scripts/validate.sh"
AGENT_VALIDATE="$REPO_ROOT/plugins/kyberforge/.apm/skills/agent-audit/scripts/validate.sh"
PASS=0
FAIL=0
pass() { echo " PASS: $1"; PASS=$((PASS + 1)); }
fail() { echo " FAIL: $1"; FAIL=$((FAIL + 1)); }
TMPDIR_T="$(mktemp -d)"
trap 'rm -rf "$TMPDIR_T"' EXIT
DESC_CHARS=450
BODY_WORDS=1000
# write_fixture <kind> <path> <name> — one generator for both file shapes.
#
# Byte-level control is the point: BOM placement, line endings and trailing
# whitespace are exactly what is under test, so the file is emitted in binary
# mode rather than through a shell heredoc that would normalise them.
write_fixture() {
python3 - "$1" "$2" "$3" "$DESC_CHARS" "$BODY_WORDS" <<'PY'
import sys
kind, path, name, desc_chars, body_words = sys.argv[1:6]
desc = 'x' * int(desc_chars)
body = ' '.join(['word'] * int(body_words))
# The default, well-formed shape. Variants below mutate it.
open_marker = '---'
close_marker = '---'
prefix = ''
newline = '\n'
fm_lines = ['name: ' + name, 'description: ' + desc]
if kind == 'plain':
pass
elif kind == 'bom':
prefix = ''
elif kind == 'leading-blanks':
prefix = '\n\n \n'
elif kind == 'trailing-ws':
open_marker = '--- '
close_marker = '---\t '
elif kind == 'crlf':
newline = '\r\n'
elif kind == 'no-close':
close_marker = None
elif kind == 'yaml-list':
fm_lines = ['- one', '- two']
elif kind == 'yaml-string':
fm_lines = ['just a bare scalar, not a mapping']
elif kind == 'yaml-none':
fm_lines = []
elif kind == 'yaml-malformed':
fm_lines = ['name: ' + name, 'description: "unterminated', 'tabs:\t- a']
elif kind == 'desc-no-value':
# The exact shape that exited 0 with zero output: a line regex's `\s*`
# crosses the newline and captures `model: sonnet` as the description.
fm_lines = ['name: ' + name, 'description:', 'model: sonnet']
elif kind == 'desc-null':
fm_lines = ['name: ' + name, 'description: null']
elif kind == 'desc-single-quoted-empty':
fm_lines = ['name: ' + name, "description: ''"]
elif kind == 'desc-double-quoted-empty':
fm_lines = ['name: ' + name, 'description: ""']
elif kind == 'desc-empty-fold':
fm_lines = ['name: ' + name, 'description: >']
else:
raise SystemExit('unknown fixture kind: %s' % kind)
parts = [prefix, open_marker, newline]
for line in fm_lines:
parts.append(line)
parts.append(newline)
if close_marker is not None:
parts.append(close_marker)
parts.append(newline)
parts.append(newline)
parts.append(body)
parts.append(newline)
with open(path, 'wb') as fh:
fh.write(''.join(parts).encode('utf-8'))
PY
}
# Builds all three subjects for one fixture kind and echoes nothing; the paths
# are fixed by convention so the probes below can find them.
#
# skill-audit takes a DIRECTORY (SKILL.md inside it, name matching the dir);
# agent-audit takes a FILE inside an apm package. The hook takes the SKILL.md
# directly, so it and skill-audit share one file.
build_subjects() {
local kind="$1" base="$TMPDIR_T/$1"
rm -rf "$base"
mkdir -p "$base/skill/my-skill" "$base/agent/.apm/agents"
cat > "$base/agent/apm.yml" <<'EOF'
name: test-package
version: 0.1.0
type: skill
EOF
write_fixture "$kind" "$base/skill/my-skill/SKILL.md" my-skill
write_fixture "$kind" "$base/agent/.apm/agents/my-agent.agent.md" my-agent
}
# probe_all <label> <kind> <needle>... — runs all three scripts over the fixture
# and requires every one of them to exit non-zero AND report every needle. One
# assertion per script would let two of them drift apart while the suite stayed
# green; the whole point of the shared resolver block is that they cannot.
#
# A needle written `@skills:<text>` is asserted for the hook and skill-audit but
# NOT for agent-audit. There is exactly one such needle in this file — the body
# word ceiling — and the exemption is the ADR, not a workaround: ADR-0020 gives
# agents the description gates and deliberately NO body word gate, because an
# agent body becomes the system prompt of a fresh context rather than competing
# with the caller's live conversation. Demanding a body finding from agent-audit
# would be demanding the ADR be contradicted.
probe_all() {
local label="$1" kind="$2"
shift 2
local base="$TMPDIR_T/$kind"
local -a targets=(
"hook|$HOOK|$base/skill/my-skill/SKILL.md"
"skill-audit|$SKILL_VALIDATE|$base/skill/my-skill"
"agent-audit|$AGENT_VALIDATE|$base/agent/.apm/agents/my-agent.agent.md"
)
local problems=""
for target in "${targets[@]}"; do
local who="${target%%|*}" rest="${target#*|}"
local script="${rest%%|*}" arg="${rest#*|}"
local out status=0
set +e
out="$(bash "$script" "$arg" 2>&1)"
status=$?
set -e
if [[ $status -eq 0 ]]; then
problems="$problems [$who exited 0: ${out:-<no output>}]"
continue
fi
for needle in "$@"; do
if [[ "$needle" == @skills:* ]]; then
# Spelled as a full `if`, not `[[ ... ]] && continue`. Under `set -e` the
# short form's exit status is the test's when it is false, and relying on
# the &&-list exemption to keep that from aborting the run is a footgun
# one edit away from biting.
if [[ "$who" == agent-audit ]]; then
continue
fi
needle="${needle#@skills:}"
fi
if [[ "$out" != *"$needle"* ]]; then
problems="$problems [$who never said '$needle': $out]"
fi
done
done
if [[ -z "$problems" ]]; then
pass "$label"
else
fail "$label —$problems"
fi
}
# ---------------------------------------------------------------------------
# 1a. Tolerated frontmatter shapes — the gates must RUN, not merely not-pass
# ---------------------------------------------------------------------------
# The needles are the FINDINGS, not the exit code. A script that rejected the BOM
# outright would exit non-zero too, and would still be skipping every ADR-0020
# measurement — which is the defect, one error message later.
echo ""
echo "--- a BOM, leading blanks, trailing marker whitespace and CRLF are all tolerated, and the gates still fire ---"
for kind in plain bom leading-blanks trailing-ws crlf; do
build_subjects "$kind"
done
probe_all "control: a well-formed over-ceiling file fails on BOTH the description and the body" \
plain "description is $DESC_CHARS char" "@skills:body is $BODY_WORDS words"
probe_all "a UTF-8 BOM does not hide an over-ceiling description or body" \
bom "description is $DESC_CHARS char" "@skills:body is $BODY_WORDS words"
probe_all "leading blank lines before the opening --- do not hide the findings" \
leading-blanks "description is $DESC_CHARS char" "@skills:body is $BODY_WORDS words"
probe_all "trailing whitespace after either --- marker does not hide the findings" \
trailing-ws "description is $DESC_CHARS char" "@skills:body is $BODY_WORDS words"
# Note on what this last one can and cannot detect. read_text() opens the file in
# TEXT mode, so Python's universal-newline translation turns \r\n into \n before
# the frontmatter matcher ever sees it — verified by mutation: reverting
# FRONTMATTER_RE to the old `^---\n(.*?)\n---` breaks the leading-blanks and
# trailing-whitespace cases above but NOT this one. So this case pins the
# end-to-end behaviour (a CRLF file is measured, not skipped) rather than the
# `\r?\n` alternations in the regex, and it would catch a future switch to binary
# reads or to a newline='' open. Kept for that reason, and labelled so nobody
# reads it as covering more than it does.
probe_all "CRLF line endings do not hide the findings" \
crlf "description is $DESC_CHARS char" "@skills:body is $BODY_WORDS words"
# ---------------------------------------------------------------------------
# 1b. Unparseable frontmatter is a hard ERROR, never a quiet skip
# ---------------------------------------------------------------------------
echo ""
echo "--- genuinely unparseable frontmatter exits non-zero with a message, rather than passing quietly ---"
for kind in no-close yaml-list yaml-string yaml-none yaml-malformed; do
build_subjects "$kind"
done
probe_all "frontmatter with no closing --- is reported, not skipped" \
no-close "frontmatter"
probe_all "frontmatter that parses to a LIST is reported, not skipped" \
yaml-list "frontmatter"
probe_all "frontmatter that parses to a STRING is reported, not skipped" \
yaml-string "frontmatter"
probe_all "frontmatter that parses to None (empty block) is reported, not skipped" \
yaml-none "frontmatter"
probe_all "malformed YAML in the frontmatter is reported, not skipped" \
yaml-malformed "frontmatter"
# ---------------------------------------------------------------------------
# 2. A valueless description is a hard FAIL in all three scripts
# ---------------------------------------------------------------------------
# All five spellings mean the same thing to a YAML parser — an empty value — and
# all five have to be decided on the FOLDED value rather than on a line regex.
# `description:` followed by `model: sonnet` is the one that shipped: it made the
# value look present, skipped the "missing or empty" failure, and then
# early-returned out of every ADR-0020 gate on the genuinely empty folded value.
echo ""
echo "--- every spelling of a valueless description hard-FAILs in all three scripts ---"
for kind in desc-no-value desc-null desc-single-quoted-empty desc-double-quoted-empty desc-empty-fold; do
build_subjects "$kind"
done
probe_all "'description:' with no value (next key not captured as the value) FAILs" \
desc-no-value "description field is missing or empty"
probe_all "'description: null' FAILs" \
desc-null "description field is missing or empty"
probe_all "\"description: ''\" FAILs" \
desc-single-quoted-empty "description field is missing or empty"
probe_all "'description: \"\"' FAILs" \
desc-double-quoted-empty "description field is missing or empty"
probe_all "'description: >' with nothing folded under it FAILs" \
desc-empty-fold "description field is missing or empty"
# The specific regression, spelled out: the valueless-description agent file must
# not merely fail — it must not be SILENT. Zero output on a blocking gate is what
# made this un-diagnosable, so the output is asserted non-empty independently.
echo ""
echo "--- the valueless-description agent file produces output, not silence ---"
build_subjects desc-no-value
set +e
SILENT_OUT="$(bash "$AGENT_VALIDATE" "$TMPDIR_T/desc-no-value/agent/.apm/agents/my-agent.agent.md" 2>&1)"
SILENT_RC=$?
set -e
if [[ $SILENT_RC -ne 0 && -n "$SILENT_OUT" ]]; then
pass "agent-audit reports a valueless description rather than exiting 0 with zero output"
else
fail "agent-audit exited $SILENT_RC with output '${SILENT_OUT:-<empty>}' — the original defect was exit 0 and total silence on a blocking pre-push gate"
fi
echo ""
echo "Results: $PASS passed, $FAIL failed"
[[ $FAIL -eq 0 ]]

409
tests/test-adr0020-targets.sh Executable file
View File

@@ -0,0 +1,409 @@
#!/usr/bin/env bash
# Regression test for the two properties of ADR-0020 boundary-target resolution
# that decide whether the gate can be trusted at all.
#
# 1. MACHINE INDEPENDENCE. The resolution universe is derived by walking up
# FROM THE TARGET FILE to an authoring root, and when one is found the
# deployed .claude/ and .agents/ trees are deliberately NOT consulted. Those
# trees are `apm install` output — gitignored, and present only on a machine
# that has run it. Four cross-plugin targets in this repo resolved through
# .claude/skills/ alone, so the same commit measured 2 dangling targets on a
# developer machine and 6 on a fresh clone. A gate shipping hot with no
# baseline cannot give two answers, so this file asserts the verdict is
# identical with and without a deployed tree — on a synthetic fixture AND on
# the real 39-skill corpus.
#
# 2. THE BARE-TARGET GRAMMAR RULE. A hyphenated token used as a compound
# MODIFIER ("pre-commit hooks", "pull-request template") is prose, not a
# route; a terminal one is a real target. Getting this wrong in either
# direction is fatal: firing on prose makes the gate untrustworthy and it
# gets turned off, while suppressing too much deletes the only two true
# positives the corpus has. Both live true positives are BARE, which is why
# the rule keys on the FOLLOWER TOKEN rather than on marking, and why they
# are pinned by name below — a future false-positive fix must not be able to
# quietly take them with it.
#
# 3. IN-SENTENCE CORROBORATION. Terminal position alone is not evidence of a
# route: "run `pre-commit` instead", "see `commit-msg`", "use the clean-up
# instead" and "run unit-tests" are all terminal, all prose, and all were
# hard FAILs with no suppression mechanism anywhere in the gate. A
# prose-form target therefore blocks only when its own sentence names
# another target that RESOLVES; otherwise it is reported at SUGGESTION tier
# and the commit proceeds. Route NOTATION (`/name`, `-> name`) is exempt
# and always blocks. Both halves are asserted below: the prose class must
# report-not-block, and the notation and corroborated forms must still
# ERROR, or the fix would have eaten the gate rather than narrowed it.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
HOOK="$REPO_ROOT/scripts/skill-size-check.sh"
PASS=0
FAIL=0
pass() { echo " PASS: $1"; PASS=$((PASS + 1)); }
fail() { echo " FAIL: $1"; FAIL=$((FAIL + 1)); }
TMPDIR_T="$(mktemp -d)"
trap 'rm -rf "$TMPDIR_T"' EXIT
# write_skill <skill-dir> <name> <desc>
write_skill() {
mkdir -p "$1"
{
echo "---"
echo "name: $2"
echo "description: $3"
echo "---"
echo ""
echo "Do the thing."
} > "$1/SKILL.md"
}
# ---------------------------------------------------------------------------
# 1. Machine independence — synthetic fixture
# ---------------------------------------------------------------------------
# Two trees, identical except that one also carries a deployed .claude/ tree
# holding a skill and an agent that exist NOWHERE in plugins/. The subject routes
# to one name that lives in a sibling plugin (must resolve in both) and one that
# lives only in .claude/ (must DANGLE in both — an authoring root exists, so the
# deployed tree is not part of the universe).
#
# If the deployed tree were consulted, the second target would resolve on the
# machine that has run `apm install` and dangle on a fresh clone. That is the
# 2-vs-6 defect exactly, at fixture scale.
echo ""
echo "--- the same file gets the same verdict with and without a deployed .claude/ tree ---"
build_tree() {
local root="$1"
write_skill "$root/plugins/other-plugin/.apm/skills/cross-plugin-skill" cross-plugin-skill \
"Use when doing the other thing. Do not use for anything else."
write_skill "$root/plugins/subject-plugin/.apm/skills/my-skill" my-skill \
"Use when doing the thing. Do not use for the other thing — use cross-plugin-skill or deployed-only-skill instead."
}
build_tree "$TMPDIR_T/no-claude"
build_tree "$TMPDIR_T/with-claude"
# The deployed tree, present only in the second root. Both a skill and an agent,
# because both are valid routing targets and both would leak.
mkdir -p "$TMPDIR_T/with-claude/.claude/skills/deployed-only-skill" \
"$TMPDIR_T/with-claude/.claude/agents"
: > "$TMPDIR_T/with-claude/.claude/agents/deployed-only-agent.md"
run_subject() {
local root="$1" out
set +e
out="$(bash "$HOOK" "$root/plugins/subject-plugin/.apm/skills/my-skill/SKILL.md" 2>&1)"
set -e
# Normalise the tree root out of the paths so the two runs are comparable.
printf '%s\n' "$out" | sed "s#$root#<ROOT>#g"
}
NO_CLAUDE_OUT="$(run_subject "$TMPDIR_T/no-claude")"
WITH_CLAUDE_OUT="$(run_subject "$TMPDIR_T/with-claude")"
if [[ "$NO_CLAUDE_OUT" == "$WITH_CLAUDE_OUT" ]]; then
pass "identical output with and without a deployed .claude/ tree"
else
fail "the deployed .claude/ tree changed the verdict — without: [$NO_CLAUDE_OUT] with: [$WITH_CLAUDE_OUT]"
fi
# Identical-but-wrong is still possible (both could resolve everything, or
# neither could resolve anything), so the CONTENT is asserted too: the
# sibling-plugin name must resolve and the deployed-only name must not.
if [[ "$WITH_CLAUDE_OUT" == *"routes to 'deployed-only-skill'"* ]]; then
pass "a name that exists only in .claude/ still dangles when an authoring root is present"
else
fail "the deployed-only target did not dangle — the deployed tree is being read into the universe: $WITH_CLAUDE_OUT"
fi
if [[ "$WITH_CLAUDE_OUT" != *"routes to 'cross-plugin-skill'"* ]]; then
pass "a name in a SIBLING PLUGIN resolves, so the comparison above is not 'nothing resolves'"
else
fail "the sibling-plugin target dangled — the monorepo universe is not being built: $WITH_CLAUDE_OUT"
fi
# The other half of the rule: with NO authoring root, deployed trees ARE the
# universe. That is the consumer case, and without this the rule above could be
# implemented as "never read .claude/", which would leave consumers with no
# resolution at all.
echo ""
echo "--- with no authoring root, a deployed .claude/ tree IS the universe ---"
CONSUMER="$TMPDIR_T/consumer"
mkdir -p "$CONSUMER/.claude/skills/deployed-only-skill"
write_skill "$CONSUMER/.claude/skills/my-skill" my-skill \
"Use when doing the thing. Do not use for the other thing — use deployed-only-skill instead."
set +e
CONSUMER_OUT="$(bash "$HOOK" "$CONSUMER/.claude/skills/my-skill/SKILL.md" 2>&1)"
CONSUMER_RC=$?
set -e
if [[ $CONSUMER_RC -eq 0 && "$CONSUMER_OUT" != *"routes to"* && "$CONSUMER_OUT" != *"DID NOT RUN"* ]]; then
pass "a sibling in a deployed .claude/skills/ tree resolves when there is no authoring root"
else
fail "the consumer path did not resolve through the deployed tree (exit $CONSUMER_RC): ${CONSUMER_OUT:-<empty>}"
fi
# ---------------------------------------------------------------------------
# 1b. Machine independence — the real corpus
# ---------------------------------------------------------------------------
# The fixture above proves the rule; this proves it at the scale where it broke.
#
# The A/B is built rather than borrowed. plugins/ is copied TWICE: once bare (a
# fresh clone), and once with a synthetic .claude/skills/ tree deployed beside it
# holding a directory for every name the corpus currently reports as dangling. If
# deployed trees leaked back into the universe, the second copy would resolve
# those names and report an empty dangling set while the first reported two —
# 2-vs-6, reproduced deterministically.
#
# Deliberately NOT keyed on whether THIS machine has run `apm install`. Doing that
# would make the suite fail on a fresh clone (where there is no .claude/ to
# contrast against) — a test of machine independence that is itself
# machine-dependent. The live tree is still compared, as a third data point, but
# nothing here requires it to be in either state.
echo ""
echo "--- the real corpus reports the same dangling targets with and without a deployed tree ---"
dangling_set() {
local -a files=()
local f
# Collected with a `while read` loop, not `mapfile`: macOS ships bash 3.2,
# which has no `mapfile`, and tests/test-vale-wrap.sh scans tests/*.sh for
# exactly that hazard. `find` rather than a glob so both roots walk identically.
while IFS= read -r f; do
files+=("$f")
done < <(find "$1" -path '*/.apm/skills/*/SKILL.md' | sort)
if [[ ${#files[@]} -eq 0 ]]; then
echo "NO-FILES-FOUND"
return
fi
set +e
bash "$HOOK" ${files[@]+"${files[@]}"} 2>&1 \
| grep -oE "routes to '[^']+'" \
| sed "s/routes to '//; s/'//" \
| sort -u
set -e
}
LIVE_DANGLING="$(dangling_set "$REPO_ROOT/plugins")"
FRESH_ROOT="$TMPDIR_T/fresh-clone"
mkdir -p "$FRESH_ROOT"
cp -R "$REPO_ROOT/plugins" "$FRESH_ROOT/plugins"
[[ -f "$REPO_ROOT/apm.yml" ]] && cp "$REPO_ROOT/apm.yml" "$FRESH_ROOT/apm.yml"
FRESH_DANGLING="$(dangling_set "$FRESH_ROOT/plugins")"
DEPLOYED_ROOT="$TMPDIR_T/deployed-clone"
mkdir -p "$DEPLOYED_ROOT/.claude/skills" "$DEPLOYED_ROOT/.claude/agents"
cp -R "$REPO_ROOT/plugins" "$DEPLOYED_ROOT/plugins"
[[ -f "$REPO_ROOT/apm.yml" ]] && cp "$REPO_ROOT/apm.yml" "$DEPLOYED_ROOT/apm.yml"
# Deploy exactly the names that currently dangle. That is the strongest possible
# bait: if the deployed tree were consulted, every one of them would resolve and
# the dangling set would collapse to empty.
DEPLOY_COUNT=0
while IFS= read -r name; do
[[ -n "$name" ]] || continue
mkdir -p "$DEPLOYED_ROOT/.claude/skills/$name"
DEPLOY_COUNT=$((DEPLOY_COUNT + 1))
done <<< "$FRESH_DANGLING"
DEPLOYED_DANGLING="$(dangling_set "$DEPLOYED_ROOT/plugins")"
if [[ "$DEPLOY_COUNT" -gt 0 ]]; then
pass "precondition: $DEPLOY_COUNT dangling name(s) deployed into the contrast tree's .claude/skills/, so the A/B has something to distinguish"
else
fail "no dangling names to deploy — the corpus reports none, so this A/B distinguishes nothing. Deploy a known-absent name explicitly instead of deriving one."
fi
if [[ ! -d "$FRESH_ROOT/.claude" && ! -d "$FRESH_ROOT/.agents" ]]; then
pass "precondition: the fresh-clone copy has no deployed tree of its own"
else
fail "the fresh-clone copy picked up a deployed tree — it is not a fresh-clone fixture"
fi
if [[ "$FRESH_DANGLING" == "$DEPLOYED_DANGLING" ]]; then
pass "deploying every dangling name into .claude/skills/ changes nothing: $(echo "$FRESH_DANGLING" | tr '\n' ' ')"
else
fail "the corpus verdict depends on whether apm install has been run — fresh clone: [$(echo "$FRESH_DANGLING" | tr '\n' ' ')] with a deployed tree: [$(echo "$DEPLOYED_DANGLING" | tr '\n' ' ')]"
fi
# Third data point: whatever state THIS machine happens to be in, the live tree
# must agree with a bare copy of the same plugins/. No precondition on that state
# — see the section header.
if [[ "$LIVE_DANGLING" == "$FRESH_DANGLING" ]]; then
pass "the live tree agrees with a bare copy (this machine $( [[ -d "$REPO_ROOT/.claude/skills" ]] && echo "HAS" || echo "has no" ) deployed .claude/skills/ tree)"
else
fail "the live tree disagrees with a bare copy of the same plugins/ — live: [$(echo "$LIVE_DANGLING" | tr '\n' ' ')] fresh clone: [$(echo "$FRESH_DANGLING" | tr '\n' ' ')]"
fi
# ---------------------------------------------------------------------------
# 1c. The two live true positives, pinned by name
# ---------------------------------------------------------------------------
# ADR-0020 records these as real broken routing targets and splits fixing them
# into Gitea issue #100. Until that lands they are the ONLY evidence the dangling
# check finds anything at all in real prose, so they are asserted as an exact set
# rather than a "contains" — a false-positive fix that suppressed one of them
# would otherwise land green.
#
# `gitea-labels` is the subtler of the two and is worth keeping: it is not
# written anywhere as `gitea-labels`. gitea-issues' description says "Composes
# `gitea-labels-\n milestones`" in a `>`-folded scalar, and the fold joins the
# lines into "gitea-labels- milestones" — the trailing hyphen is what keeps the
# token terminal and therefore danglable.
#
# WHEN ISSUE #100 IS FIXED: update EXPECTED_DANGLING to match. Do not delete the
# assertion — an empty expected set is fine and still pins that no NEW dangling
# target appeared.
echo ""
echo "--- the two live dangling targets in the corpus are exactly the two ADR-0020 records ---"
EXPECTED_DANGLING="$(printf '%s\n' gitea-labels neuledge-context)"
if [[ "$LIVE_DANGLING" == "$EXPECTED_DANGLING" ]]; then
pass "the corpus dangling set is exactly {gitea-labels, neuledge-context}"
else
fail "the corpus dangling set changed — expected [$(echo "$EXPECTED_DANGLING" | tr '\n' ' ')], got [$(echo "$LIVE_DANGLING" | tr '\n' ' ')]. If a retrofit fixed one, update EXPECTED_DANGLING; if a false-positive fix silently deleted one, that is the regression this asserts."
fi
for probe in \
"plugins/bin/.apm/skills/research/SKILL.md:neuledge-context" \
"plugins/gitea/.apm/skills/gitea-issues/SKILL.md:gitea-labels"; do
probe_file="$REPO_ROOT/${probe%%:*}"
probe_name="${probe##*:}"
if [[ ! -f "$probe_file" ]]; then
fail "the true-positive fixture ${probe%%:*} no longer exists — this pin has become vacuous"
continue
fi
set +e
probe_out="$(bash "$HOOK" "$probe_file" 2>&1)"
set -e
if [[ "$probe_out" == *"routes to '$probe_name'"* ]]; then
pass "detects the dangling '$probe_name' target in ${probe%%:*}"
else
fail "did not detect the dangling '$probe_name' target in ${probe%%:*} — a false-positive fix has taken a true positive with it: $probe_out"
fi
done
# ---------------------------------------------------------------------------
# 2. The bare-target grammar rule
# ---------------------------------------------------------------------------
# Every fixture is built inside a real plugin tree. In a bare temp directory the
# resolver would decline ("DID NOT RUN") and every must-not-error case would pass
# vacuously, proving nothing about extraction.
echo ""
echo "--- attributive compound modifiers are prose, not routing targets ---"
GRAMMAR_ROOT="$TMPDIR_T/grammar"
write_skill "$GRAMMAR_ROOT/plugins/p/.apm/skills/sibling-skill" sibling-skill \
"Use when doing the other thing. Do not use for anything else."
# grammar_case <slug> <expect: silent|errors> <needle> <description>
grammar_case() {
local slug="$1" mode="$2" needle="$3" desc="$4" out status=0
write_skill "$GRAMMAR_ROOT/plugins/p/.apm/skills/$slug" "$slug" "$desc"
set +e
out="$(bash "$HOOK" "$GRAMMAR_ROOT/plugins/p/.apm/skills/$slug/SKILL.md" 2>&1)"
status=$?
set -e
if [[ "$out" == *"DID NOT RUN"* ]]; then
fail "\"$desc\" — the resolver declined, so this case asserts nothing about extraction: $out"
return
fi
case "$mode" in
silent)
if [[ $status -eq 0 && -z "$out" ]]; then
pass "not a dangling target: \"$desc\""
else
fail "\"$desc\" (exit $status, output: ${out:-<empty>})"
fi
;;
errors)
if [[ $status -ne 0 && "$out" == *"$needle"* ]]; then
pass "still a dangling target: \"$desc\""
else
fail "\"$desc\" should have ERRORed with $needle (exit $status, output: ${out:-<empty>})"
fi
;;
suggests)
# Reported, not blocking. Both halves matter: an ERROR here would be the
# unsuppressable false positive this tier exists to remove, and silence
# would mean the gate stopped noticing the target at all.
if [[ $status -eq 0 && "$out" == *"SUGGESTION"*"$needle"* && "$out" != *"ERROR"* ]]; then
pass "reported but not blocking: \"$desc\""
else
fail "\"$desc\" should have exited 0 with a SUGGESTION naming $needle (exit $status, output: ${out:-<empty>})"
fi
;;
esac
}
# The four phrasings that were hard dangling FAILs with no suppression. All four
# are lifted from real descriptions in this corpus.
grammar_case fp-precommit-hooks silent "" \
"Use when running the linter. Use pre-commit hooks instead of ad-hoc scripts."
grammar_case fp-pull-request silent "" \
"Use when opening changes. Invoke the pull-request template instead of writing one by hand."
grammar_case fp-conventional silent "" \
"Use when writing history. Use conventional-commits formatting rather than free-form messages."
grammar_case fp-prepush-backticked silent "" \
"Use when checking a branch. Do not use for local edits — run the \`pre-push\` hooks instead."
echo ""
echo "--- a lone unresolvable token in terminal position is REPORTED, not blocking ---"
# The class this tier was added for. Every one of these is grammatically
# identical to a real broken route — "route verb + name + terminal" is also how
# prose cites a hook, a linter, a file format or an English compound — and every
# one of them was a hard FAIL with no suppression mechanism anywhere in the gate.
# The skills most exposed are exactly the ones the ADR-0020 retrofit sends
# authors back to rewrite first: pc-run, pc-author, vale-run, vale-config and the
# apm-* family are all ABOUT hyphenated tools.
grammar_case fp-precommit-terminal suggests "routes to 'pre-commit'" \
"Use when running the linter. Do not use for running hooks — run \`pre-commit\` instead."
grammar_case fp-commit-msg suggests "routes to 'commit-msg'" \
"Use when writing history. Do not use for the commit message — see \`commit-msg\`."
grammar_case fp-type-check suggests "routes to 'type-check'" \
"Use when compiling. Do not use for type errors — run \`type-check\` first."
grammar_case fp-semantic-release suggests "routes to 'semantic-release'" \
"Use when tagging a version. Instead, use \`semantic-release\`."
# Single-word tool names are the same defect: `eslint` in terminal position hit
# the marked-target path and hard-FAILed just as `pre-commit` did.
grammar_case fp-eslint suggests "routes to 'eslint'" \
"Use when linting JS. Do not use for style — run \`eslint\` instead."
# Bare English compounds, which the backtick path never sees at all.
grammar_case fp-clean-up suggests "routes to 'clean-up'" \
"Use when doing the thing. Do not use for the old flow — use the clean-up instead."
grammar_case fp-built-in suggests "routes to 'built-in'" \
"Use when doing the thing. Do not use for the custom path — Instead, prefer the built-in."
grammar_case fp-write-up suggests "routes to 'write-up'" \
"Use when doing the thing. Do not use for the summary — see the write-up."
grammar_case fp-front-end suggests "routes to 'front-end'" \
"Use when doing the thing. Do not use for the API layer — use the front-end."
grammar_case fp-unit-tests suggests "routes to 'unit-tests'" \
"Use when testing. Do not run end-to-end, run unit-tests."
echo ""
echo "--- terminal targets that resolve to nothing still ERROR ---"
# The controls. Without them the cases above are satisfied by a check that never
# fires, and the narrowing would have eaten the gate rather than sharpened it.
# Three forms, three code paths:
# * route NOTATION — `/name` and `-> name` — is exempt from corroboration and
# blocks on its own. Nobody writes `/pre-commit` or `-> pre-commit` to mean
# the hook, so there is no ambiguity to resolve, and an author who wants a
# route checked unconditionally has two ways to say so.
# * a PROSE-form target — backticked or bare — blocks when its own sentence
# names another target that resolves. `sibling-skill` is that corroborator
# here; it is the same shape as both live true positives, which sit beside
# `write-docs` and `gitea-labels-milestones` respectively.
grammar_case tp-arrow errors "routes to 'no-such-arrow-target'" \
"Use when doing the thing. Not the other thing → no-such-arrow-target."
grammar_case tp-slash errors "routes to 'no-such-slash-skill'" \
"Use when doing the thing. Do not use for improvements — use /no-such-slash-skill instead."
grammar_case tp-backticked errors "routes to 'no-such-backticked-skill'" \
"Use when doing the thing. Do not use for improvements — use \`sibling-skill\` or \`no-such-backticked-skill\` instead."
grammar_case tp-bare-terminal errors "routes to 'no-such-bare-skill'" \
"Use when doing the thing. Do not use for improvements — use sibling-skill or no-such-bare-skill instead."
# And the confirming half of the grammar rule: a compound-modifier target is
# CONFIRM-ONLY, not ignored. When the name does exist it still counts as a route
# — the rule suppresses the ERROR, it does not delete the target.
echo ""
echo "--- an attributive target that DOES resolve is still a route, not a discarded token ---"
write_skill "$GRAMMAR_ROOT/plugins/p/.apm/skills/attributive-subject" attributive-subject \
"Use when doing the thing. Do not use for the other thing — use the sibling-skill helper instead."
set +e
ATTR_OUT="$(bash "$HOOK" "$GRAMMAR_ROOT/plugins/p/.apm/skills/attributive-subject/SKILL.md" 2>&1)"
ATTR_RC=$?
set -e
if [[ $ATTR_RC -eq 0 && -z "$ATTR_OUT" ]]; then
pass "a resolving attributive target neither errors nor is reported"
else
fail "an attributive target naming a REAL skill produced output (exit $ATTR_RC): $ATTR_OUT"
fi
echo ""
echo "Results: $PASS passed, $FAIL failed"
[[ $FAIL -eq 0 ]]

View File

@@ -538,13 +538,24 @@ else
fi
# --- 10g. An ambient RUN_TESTS_STRICT=1 must not reach a fixture that did not ask
# for it. This suite is discovered and run by run-tests.sh itself, so under the
# gate the variable is exported into every child here. That is not a hypothetical:
# `bash tests/run-tests.sh` reported 19 passed while
# `RUN_TESTS_STRICT=1 bash tests/run-tests.sh` reported this file as the one
# failure, because case 10c's deliberately-non-strict run inherited strict and
# went red. The gate was therefore RED for everyone, and the only way to see it
# was to run the gate.
# for it. This suite is discovered and run by run-tests.sh itself, so when the
# outer run is launched with the env-var spelling the variable used to be exported
# into every child here. That was not a hypothetical: `bash tests/run-tests.sh`
# was green while `RUN_TESTS_STRICT=1 bash tests/run-tests.sh` reported this file
# as the one failure, because case 10c's deliberately-non-strict run inherited
# strict and went red.
#
# Two corrections to the record, because both were overstated before:
# * The blast radius was TWO assertions, not six -- cases 10c and 10g here, and
# nothing else in the repo reads RUN_TESTS_STRICT.
# * The pre-push GATE was never red. It runs `bash tests/run-tests.sh --strict`
# (see .pre-commit-config.yaml), and the flag sets a shell local that is never
# exported, so the flag spelling never leaked. Only the env-var spelling did.
#
# run-tests.sh now `unset`s the variable immediately after latching it, so the
# leak is closed at its source and the two spellings hand children an identical
# environment (case 10i pins that directly). The `env -u` in run_fake() is kept as
# this suite's own defence-in-depth rather than as the fix.
#
# The variable is exported here rather than passed as a prefix on purpose: a
# prefix (`RUN_TESTS_STRICT=1 run_fake ...`) applies to the function call, and
@@ -604,6 +615,60 @@ else
pass "--strict reports skipped suites on stderr and suppresses the duplicate stdout list"
fi
# --- 10i. The two documented invocations are equivalent in what a CHILD sees ---
# `bash tests/run-tests.sh --strict` and `RUN_TESTS_STRICT=1 bash
# tests/run-tests.sh` are documented as the same switch, and case 10b already
# asserts they produce the same PARENT verdict. That is the weaker half: the two
# differed in the ENVIRONMENT they handed every dispatched test-*.sh, because the
# flag sets a shell local while the env var stayed exported down the whole process
# tree. A dispatched suite could therefore behave differently depending on which
# spelling launched the run above it -- which is how case 10c went red under one
# invocation and green under the other.
#
# So this asks the children directly rather than reading the parent's summary. The
# case script reports whether RUN_TESTS_STRICT is present in its own environment
# AT ALL (`${VAR+set}`, not `${VAR:-}` -- an exported empty value is still a leak),
# and both spellings must report it absent. Equivalence is asserted between the two
# observations, not just against a hardcoded expectation, so the two cannot drift
# apart in some future direction neither case anticipated.
echo ""
echo "--- --strict and RUN_TESTS_STRICT=1 hand children the same environment ---"
DIR10I="$(make_fake_repo)"
FIXTURES+=("$DIR10I")
install_healthy_bats_runner "$DIR10I"
add_case "$DIR10I" test-reports-its-env.sh <<'EOF'
#!/usr/bin/env bash
if [[ -n "${RUN_TESTS_STRICT+set}" ]]; then
echo "CHILD-SAW-STRICT=[${RUN_TESTS_STRICT}]"
else
echo "CHILD-SAW-STRICT=<absent>"
fi
EOF
# Flag spelling: run_fake scrubs the ambient variable first, so what the child
# sees here is purely a function of what run-tests.sh itself exports.
run_fake "$DIR10I" --strict
FLAG_CHILD="$(echo "$FAKE_OUT" | grep -o 'CHILD-SAW-STRICT=.*' | head -n 1 || true)"
FLAG_RC=$FAKE_RC
# Env spelling: invoked directly, NOT through run_fake, because run_fake's `env -u`
# would strip the very variable under test.
I_PRIV="$(mktemp -d)"
FIXTURES+=("$I_PRIV")
ENV_RC=0
ENV_OUT="$(TMPDIR="$I_PRIV" TEST_DIR="$DIR10I/cases" RUN_TESTS_STRICT=1 \
bash "$DIR10I/tests/run-tests.sh" 2>&1)" || ENV_RC=$?
ENV_CHILD="$(echo "$ENV_OUT" | grep -o 'CHILD-SAW-STRICT=.*' | head -n 1 || true)"
if [[ -z "$FLAG_CHILD" || -z "$ENV_CHILD" ]]; then
fail "the reporting case script never ran under one of the two invocations (flag: '${FLAG_CHILD:-<none>}', env: '${ENV_CHILD:-<none>}')"
elif [[ "$FLAG_CHILD" != "$ENV_CHILD" ]]; then
fail "the two documented invocations hand children different environments — flag: $FLAG_CHILD, env: $ENV_CHILD"
elif [[ "$ENV_CHILD" != "CHILD-SAW-STRICT=<absent>" ]]; then
fail "RUN_TESTS_STRICT is still exported to dispatched suites ($ENV_CHILD) — a suite that itself runs run-tests.sh inherits strictness it never asked for"
elif [[ $FLAG_RC -ne 0 || $ENV_RC -ne 0 ]]; then
fail "a clean fixture failed under one of the two invocations (flag rc=$FLAG_RC, env rc=$ENV_RC): $FAKE_OUT / $ENV_OUT"
else
pass "--strict and RUN_TESTS_STRICT=1 both dispatch children with RUN_TESTS_STRICT absent"
fi
echo ""
echo "Results: $PASS passed, $FAIL failed"
[[ $FAIL -eq 0 ]]

View File

@@ -294,6 +294,61 @@ make_budget_fixture() {
echo "$file"
}
# desc_of_length <n> — a description of EXACTLY n characters that carries a
# boundary clause and names no routing target.
#
# ADR-0020's missing-boundary-clause SUGGESTION fires on every description
# without one, so a fixture that omits it is never "otherwise clean": a test
# asserting silence would be asserting the boundary check's ABSENCE rather than
# the length boundary it names. The clause is paid for out of the same budget
# being measured (padding arithmetic, not a fixed suffix) so the character count
# stays exact. "anything else" is unhyphenated, so no routing target rides along.
desc_of_length() {
python3 - "$1" <<'PY'
import sys
n = int(sys.argv[1])
prefix = 'Use when doing the thing. Do not use for anything else. '
assert n >= len(prefix), 'requested description shorter than the boundary clause'
print(prefix + 'x' * (n - len(prefix)))
PY
}
# make_tree_fixture <label> <desc> <body_words> — a SKILL.md inside a synthetic
# apm plugin monorepo, so the boundary-target resolver has a universe.
#
# Resolution walks up FROM THE TARGET FILE to an authoring root (the nearest
# ancestor holding plugins/*/.apm/{skills,agents}, falling back to .git); it is
# never derived from the checker's own location, because deriving it from
# ${BASH_SOURCE} leaked this repo's 39-skill universe into every consumer repo
# running the hook. A fixture in a bare mktemp -d therefore has NO universe and
# correctly reports "DID NOT RUN" — that is not a bug to paper over with a
# looser assertion, it is why the fixture has to be a real tree:
#
# <root>/plugins/subject-plugin/.apm/skills/<label>/SKILL.md <- the subject
# <root>/plugins/subject-plugin/.apm/skills/sibling-skill/ <- same package
# <root>/plugins/subject-plugin/.apm/agents/sibling-agent.agent.md
# <root>/plugins/other-plugin/.apm/skills/cross-plugin-skill/ <- sibling plugin
#
# The sibling plugin is what makes "every plugin in the monorepo contributes its
# names" testable; without it a cross-plugin target and a typo are the same.
make_tree_fixture() {
local label="$1" desc="$2" body_words="$3" root apm
root="$TMPDIR/tree-$label"
apm="$root/plugins/subject-plugin/.apm"
mkdir -p "$apm/skills/$label" "$apm/skills/sibling-skill" "$apm/agents" \
"$root/plugins/other-plugin/.apm/skills/cross-plugin-skill"
: > "$apm/agents/sibling-agent.agent.md"
{
echo "---"
echo "name: $label"
echo "description: $desc"
echo "---"
echo ""
python3 -c "print(' '.join(['word'] * $body_words))"
} > "$apm/skills/$label/SKILL.md"
echo "$apm/skills/$label/SKILL.md"
}
# expect_gate <label> <expected: pass|suggest|fail> <file> [needle]
expect_gate() {
local label="$1" expected="$2" file="$3" needle="${4:-}" out status
@@ -323,15 +378,26 @@ expect_gate() {
fail "$label (exit $status, output: ${out:-<empty>})"
fi
;;
# A check that DECLINED to run must say so and must not fail the file. The
# ERROR guard is the point: a declined check that also errored would satisfy
# a bare "output contains INFO" assertion.
info)
if [[ $status -eq 0 && "$out" == *"INFO"* && "$out" == *"$needle"* \
&& "$out" != *"ERROR"* ]]; then
pass "$label"
else
fail "$label (exit $status, output: ${out:-<empty>})"
fi
;;
esac
}
echo ""
echo "--- description budget: $DESC_SUGGEST_CHARS SUGGESTION / $DESC_MAX_CHARS FAIL, both inclusive ---"
D_AT_SUGGEST="$(python3 -c "print('x' * $DESC_SUGGEST_CHARS)")"
D_OVER_SUGGEST="$(python3 -c "print('x' * $((DESC_SUGGEST_CHARS + 1)))")"
D_AT_MAX="$(python3 -c "print('x' * $DESC_MAX_CHARS)")"
D_OVER_MAX="$(python3 -c "print('x' * $((DESC_MAX_CHARS + 1)))")"
D_AT_SUGGEST="$(desc_of_length "$DESC_SUGGEST_CHARS")"
D_OVER_SUGGEST="$(desc_of_length "$((DESC_SUGGEST_CHARS + 1))")"
D_AT_MAX="$(desc_of_length "$DESC_MAX_CHARS")"
D_OVER_MAX="$(desc_of_length "$((DESC_MAX_CHARS + 1))")"
expect_gate "description at exactly $DESC_SUGGEST_CHARS chars is silent" \
pass "$(make_budget_fixture desc-at-suggest "$D_AT_SUGGEST" 10)"
expect_gate "description at $((DESC_SUGGEST_CHARS + 1)) chars suggests and exits 0" \
@@ -359,18 +425,23 @@ FOLDED="$TMPDIR/folded.md"
expect_gate "a >-folded 450-char description fails (raw first line would read as 1 char)" \
fail "$FOLDED" "description is 450 characters"
# Every body fixture below carries a boundary clause for the same reason
# desc_of_length() does: without one the missing-boundary-clause SUGGESTION
# fires and a body-budget test that asserts silence stops isolating the body
# budget. It is short, so the description gate stays quiet too.
CLEAN_DESC="Short valid description. Do not use for anything else."
echo ""
echo "--- body budget: $BODY_SUGGEST_WORDS SUGGESTION / $BODY_MAX_WORDS FAIL, body only, both inclusive ---"
expect_gate "body at exactly $BODY_SUGGEST_WORDS words is silent" \
pass "$(make_budget_fixture body-at-suggest "Short valid description." "$BODY_SUGGEST_WORDS")"
pass "$(make_budget_fixture body-at-suggest "$CLEAN_DESC" "$BODY_SUGGEST_WORDS")"
expect_gate "body at $((BODY_SUGGEST_WORDS + 1)) words suggests and exits 0" \
suggest "$(make_budget_fixture body-over-suggest "Short valid description." "$((BODY_SUGGEST_WORDS + 1))")" \
suggest "$(make_budget_fixture body-over-suggest "$CLEAN_DESC" "$((BODY_SUGGEST_WORDS + 1))")" \
"body is $((BODY_SUGGEST_WORDS + 1)) words"
expect_gate "body at exactly $BODY_MAX_WORDS words suggests, does not fail" \
suggest "$(make_budget_fixture body-at-max "Short valid description." "$BODY_MAX_WORDS")" \
suggest "$(make_budget_fixture body-at-max "$CLEAN_DESC" "$BODY_MAX_WORDS")" \
"body is $BODY_MAX_WORDS words"
expect_gate "body at $((BODY_MAX_WORDS + 1)) words fails" \
fail "$(make_budget_fixture body-over-max "Short valid description." "$((BODY_MAX_WORDS + 1))")" \
fail "$(make_budget_fixture body-over-max "$CLEAN_DESC" "$((BODY_MAX_WORDS + 1))")" \
"$BODY_MAX_WORDS-word ceiling"
# The two word gates measure different things and must stay separable: a file
@@ -382,7 +453,7 @@ BODY_ONLY_DESC="$(python3 -c "print(' '.join(['w'] * 100))")"
expect_gate "frontmatter words do not count toward the $BODY_MAX_WORDS-word body ceiling" \
suggest "$(make_budget_fixture body-independent "$BODY_ONLY_DESC" "$((BODY_MAX_WORDS - 5))")" \
"words"
BIG_BODY="$(make_budget_fixture body-over-not-whole-file "Short valid description." "$((BODY_MAX_WORDS + 1))")"
BIG_BODY="$(make_budget_fixture body-over-not-whole-file "$CLEAN_DESC" "$((BODY_MAX_WORDS + 1))")"
BIG_BODY_WORDS="$(wc -w < "$BIG_BODY")"
if [[ "$BIG_BODY_WORDS" -le "$MAX_WORDS" ]]; then
pass "the body-gate fixture is $BIG_BODY_WORDS whole-file words, well under MAX_WORDS=$MAX_WORDS — it fails on the body gate alone"
@@ -393,21 +464,37 @@ fi
echo ""
echo "--- resolvable boundary targets ---"
# Resolution is against the AUTHORING SOURCE (plugins/*/.apm/skills/ and
# plugins/*/.apm/agents/), found here via the script's own repo root — these
# fixtures live in a temp dir with no plugin tree of their own, so a resolving
# target proves the repo-root path works.
expect_gate "a boundary target naming a real skill resolves" \
pass "$(make_budget_fixture target-ok \
"Use when doing the thing. Do not use for commits — use git-commits instead." 10)"
expect_gate "a boundary target naming a real AGENT resolves (agents are valid targets)" \
pass "$(make_budget_fixture target-agent-ok \
"Use when doing the thing. Do not use when the caller is an agent — invoke git-orchestrate instead." 10)"
expect_gate "a boundary target that resolves to nothing fails" \
fail "$(make_budget_fixture target-missing \
"Use when doing the thing. Do not use for improvements — use no-such-skill-anywhere instead." 10)" \
# plugins/*/.apm/agents/), reached by walking up FROM THE SKILL FILE. These
# fixtures therefore build their own synthetic monorepo (make_tree_fixture) and
# name only fixture-local targets: they must not depend on this repo's live
# skills, or renaming git-commits would break a test about extraction grammar.
expect_gate "a boundary target naming a sibling skill in the same package resolves" \
pass "$(make_tree_fixture target-ok \
"Use when doing the thing. Do not use for commits — use sibling-skill instead." 10)"
expect_gate "a boundary target naming a skill in a SIBLING PLUGIN resolves (that is what a monorepo means)" \
pass "$(make_tree_fixture target-cross-plugin \
"Use when doing the thing. Do not use for the other thing — use cross-plugin-skill instead." 10)"
expect_gate "a boundary target naming an AGENT resolves (agents are valid targets)" \
pass "$(make_tree_fixture target-agent-ok \
"Use when doing the thing. Do not use when the caller is an agent — invoke sibling-agent instead." 10)"
# CORROBORATED: `sibling-skill` resolves in the same sentence, which is what
# promotes a prose-form target from "reported" to "blocking". A lone prose-form
# target is deliberately not fatal — see the case below and the shared resolver's
# CORROBORATION note.
expect_gate "a boundary target that resolves to nothing fails when its sentence names one that does" \
fail "$(make_tree_fixture target-missing \
"Use when doing the thing. Do not use for improvements — use sibling-skill or no-such-skill-anywhere instead." 10)" \
"routes to 'no-such-skill-anywhere'"
# UNCORROBORATED: identical grammar to the case above, and identical grammar to
# "run `pre-commit` instead". Reported at SUGGESTION tier, exit 0 — a gate that
# ships hot with no baseline and no suppression mechanism must not block a commit
# on a token it cannot tell from a tool name.
expect_gate "a lone boundary target that resolves to nothing is reported, not fatal" \
suggest "$(make_tree_fixture target-missing-lone \
"Use when doing the thing. Do not use for improvements — use no-such-lone-skill instead." 10)" \
"routes to 'no-such-lone-skill'"
expect_gate "a /slash-command boundary target that resolves to nothing fails" \
fail "$(make_budget_fixture target-missing-slash \
fail "$(make_tree_fixture target-missing-slash \
"Use when doing the thing. Do not use for improvements — use /no-such-slash-skill instead." 10)" \
"routes to 'no-such-slash-skill'"
# False-positive guards. These phrasings are lifted from real descriptions:
@@ -415,16 +502,38 @@ expect_gate "a /slash-command boundary target that resolves to nothing fails" \
# gitea-files says "(use Read/Write/Edit)", gitea-labels-milestones says
# "through `issue_write`/`pull_request_write`". None of them is a routing
# target, and reading any of them as one makes the gate untrustworthy.
#
# Each carries a boundary clause in a SEPARATE sentence. That is not decoration:
# target extraction is decided per sentence, so the clause satisfies the
# missing-boundary-clause SUGGESTION (keeping the expected output empty) while
# leaving the sentence under test outside a boundary context, which is the exact
# condition each of these is about. They are built as trees so a universe exists
# — in a bare temp dir the resolver would decline and the guard would pass
# vacuously, proving nothing about extraction.
expect_gate "'run pre-commit hooks' outside a boundary sentence is not a routing target" \
pass "$(make_budget_fixture fp-precommit \
"Use when the user wants to run pre-commit hooks or install git hooks." 10)"
pass "$(make_tree_fixture fp-precommit \
"Use when the user wants to run pre-commit hooks or install git hooks. Do not use for anything else." 10)"
expect_gate "an arrow chain outside a boundary clause is not a routing target" \
pass "$(make_budget_fixture fp-arrow \
"Reproduce → minimise → instrument → fix → regression-test. Use when a bug is reported." 10)"
pass "$(make_tree_fixture fp-arrow \
"Reproduce → minimise → instrument → fix → regression-test. Use when a bug is reported. Do not use for anything else." 10)"
expect_gate "tool names and MCP tool names are not routing targets" \
pass "$(make_budget_fixture fp-tools \
pass "$(make_tree_fixture fp-tools \
"Use when writing issues. Do not use for local files (use Read/Write/Edit) — that write goes through \`issue_write\`/\`pull_request_write\` instead." 10)"
echo ""
echo "--- with NO authoring root the resolver declines OUT LOUD and does not fail the file ---"
# The consumer/draft case, and a real one: a SKILL.md in a bare directory with no
# plugins/*/.apm/ above it and no .git has no universe to resolve against. The
# required behaviour is neither a false FAIL nor silence — silence is how a whole
# gate family goes missing unnoticed — so the INFO and the named unchecked target
# are both asserted, along with exit 0. This is the same path make_tree_fixture
# exists to escape, kept pinned so a future "just use the repo root" shortcut
# (the ${BASH_SOURCE} universe leak ADR-0020 removed) fails here.
expect_gate "a fixture with no authoring root reports DID NOT RUN and exits 0" \
info "$(make_budget_fixture no-universe \
"Use when doing the thing. Do not use for improvements — use some-other-skill instead." 10)" \
"Unchecked target(s): some-other-skill"
echo ""
echo "--- the three live dangling routing targets are caught (issue #100) ---"
# ADR-0020 records four broken routing targets and splits fixing them into its