fix(kyberforge): correct the routing-tier contract and make agent-audit rubrics conditional

Five documents told authors that a prose-form dangling routing target blocks. The
gate reports it as a SUGGESTION and exits 0. Verified on fixtures: `-> name` and
`/name` are blocking ERRORs, the prose form is SUGGESTION-tier unless a second
resolving target in the same sentence corroborates it. ADR-0020 and gates.md were
right; contract.md, retrofit.md, description-quality.md, finding-criteria.md and
agent-author's contract.md were wrong — and they are what an author and an auditor
actually read. The whole 39-skill corpus was retrofitted against them.

skill-audit was also self-contradictory: it imports validate.sh's SUGGESTIONs into
the Structure dimension verbatim while its own rubric grades the same target a FAIL,
so one target got reported twice at two tiers. The script owns the grade; the rubric
now says so.

The YAML-fold trap that broke gitea-labels-milestones (#100) was warned about only
in retrofit.md, reachable only from the improve flow when a budget is exceeded. It
is now in both contract.md files, which SKILL.md mandates on the create flow too.

agent-audit loaded both rubrics unconditionally on every run — 3,323 words for a
clean audit against skill-audit's 1,636. dac9cad fixed exactly this in skill-audit
and edited agent-audit in the same commit without applying it. Same treatment: the
criteria move to a new finding-criteria.md and load per dimension. Clean run now
2,083 words, a 37% cut.

Routing: apm-workflow's description shed dependency installation while still owning
the flow, and apm-install's boundary did not exclude it, so "install my apm
dependencies" matched the CLI-binary skill with no route back. Fixed on both sides.
forge regains two of the three phrasings the retrofit deleted.

forge Step 1 called grill-with-docs unconditionally — a skill in plugins/bin, which
kyberforge does not declare as a dependency. It resolves here only because the
walk-up sweeps sibling plugins; a standalone install dead-ends. Step 1 now names
the cross-plugin dependency and gives an inline fallback. Declaring it properly in
apm.yml remains the better fix.

Also: both audit SKILL.md files now grade exit 2 as "did not run, dimension
unverified" rather than as findings; skill-audit's README row described content that
moved, which its own finding-criteria.md grades a FAIL; and body-discipline.md's
`git show <sha>:plugins/...` command is fenced, since an installed plugin cache has
no repo and file-structure.md makes a bare repo path a FAIL.

Refs: #100, #101, #125
ADR: 0020
This commit is contained in:
2026-09-01 12:38:22 +00:00
parent 59f27dbd94
commit fc305ba7d9
36 changed files with 424 additions and 228 deletions

View File

@@ -36,7 +36,7 @@ Provide the path to the skill directory to audit when invoking.
| `assets/vale/styles/Kyberforge/SentenceOpenerThereIs.yml` | Vale rule — flags body sentences starting with "There is"/"There are" |
| `assets/vale/styles/Kyberforge/VagueWording.yml` | Vale rule — flags known filler wording (e.g. "helps with", "utilize") |
| `references/finding-criteria.md` | Every dimension's FAIL and SUGGESTION criteria — the one Step 3 file loaded on every run; it decides which rubrics below are worth loading |
| `references/description-quality.md` | Rubric for the description dimension — three-part shape, the 250/400-character budget, the hand-invoked (`disable-model-invocation`) contract, and the internal-mechanics FAIL |
| `references/description-quality.md` | Rubric for the description dimension — why the description is the expensive part, the hand-invoked (`disable-model-invocation`) contract, the three-part shape, when an indirect trigger is warranted, near-miss exclusions, and a before/after pair |
| `references/body-discipline.md` | Rubric for the body-discipline dimension — the core test, the 600/900 body-only budget against the 2,770-word whole-file backstop, the mandatory-dispatch rule, and the Gotchas constraints |
| `references/patterns.md` | Rubric for the patterns dimension — which instruction construct fits which job, and how each is correctly formed |
| `references/file-structure.md` | Rubric for the file-structure and internal-consistency dimensions — permitted directories, cross-plugin path rules and their two structural exemptions, README drift |

View File

@@ -33,11 +33,11 @@ bash scripts/validate-provenance.sh <skill-dir>
bash scripts/vale-wrap.sh <skill-dir>/SKILL.md
```
`validate.sh` findings become the `### Structure` dimension — its FAILs and its SUGGESTIONs both.
`validate.sh` findings become the `### Structure` dimension — its FAILs and its SUGGESTIONs both, at the tier the script assigned. Report each once; never re-grade one under another dimension. Unresolved boundary targets are where this bites, because their tier turns on notation.
If any of the three cannot run, or exits non-zero for a reason other than findings, read `references/validation-scripts.md` — it carries the manual fallback and the misleading exit codes. Ordinary content FAILs are the expected outcome here and need no fallback.
`validate-provenance.sh` prints nothing on success. Its FAIL and INFO findings become a separate `### Provenance` dimension, and it emits Why and Fix itself — surface those verbatim.
`validate-provenance.sh` prints nothing on success, so read its exit code before you read its silence. **0** is a genuine pass. **1** means real findings: its FAILs and INFOs become a separate `### Provenance` dimension, and it emits Why and Fix itself — surface those verbatim. **2** means the check never ran — a usage or environment error, reason on stderr, no findings and often no stdout at all. On a 2, report `### Provenance` as unverified and quote the stderr reason. Never grade an exit 2 as a clean pass: empty stdout there means nothing was checked, not that nothing was wrong.
`vale-wrap.sh` applies the bundled `Kyberforge` style as a prefilter. Pass no `--config`; the wrapper locates its own. Every rule is graded `error`, so every alert is a FAIL. Report each one citing its rule ID, filed under the dimension it belongs to, and do not re-derive it by judgment:

View File

@@ -135,9 +135,16 @@ Constraints:
Worked negative example — **`git-commits` v0.1.2 at commit `5e23250`, a fixed pre-retrofit
snapshot, not the current file.** The live skill is v0.1.3 and matches none of the citations below;
they are quoted as they stood before the ADR-0020 retrofit, and are not to be refreshed against
`HEAD`. Read the snapshot with
`git show 5e23250:plugins/git/.apm/skills/git-commits/SKILL.md`. That body carried
twelve Gotchas, four of which restated content already below them or already in the description:
`HEAD`. The snapshot is reachable only from a checkout of the authoring repo — an installed plugin
cache holds no git history and no such path — so read the citations below as quoted rather than
going to look for the file. From a checkout:
```text
git show 5e23250:<the git plugin>/.apm/skills/git-commits/SKILL.md
```
That body carried twelve Gotchas, four of which restated content already below them or already in
the description:
| Gotcha | Restates |
|---|---|

View File

@@ -40,8 +40,11 @@ A model-invoked description carries exactly three things:
the agent is deciding whether to act, not reading a catalogue entry.
2. **At most one capability clause.** What it does, in one clause. Never an enumeration.
3. **Boundary clause.** Compressed form: `Not <thing> -> <skill-name>.` The target must resolve to
a real skill directory or agent file in the authoring source; `validate.sh` checks that
deterministically and a dangling target already surfaces as a Structure FAIL.
a real skill directory or agent file in the authoring source. `validate.sh` checks that
deterministically and grades it by notation: an unresolved `/name` or arrow target is an ERROR
and reaches the report as a Structure FAIL, while an unresolved prose-form target ("use `y`
instead") is only a SUGGESTION unless a second target in the same sentence resolves. Take the
script's tier as given and report it once, under Structure.
Everything else belongs in the body or in `README.md`.

View File

@@ -42,8 +42,6 @@ Flag as FAIL if:
- **Vague capabilities** ("helps with APIs" where "parses and validates OpenAPI specs" was
available). `Kyberforge.VagueWording` catches the known filler; imprecision outside that list is
judgment.
- **A boundary clause naming a target that does not resolve** to a real skill directory or agent
file in the authoring source. `validate.sh` reports the unresolved name.
- **Trigger-list, boundary or indirect-trigger content on a hand-invoked skill** — see Step 0 of
`references/description-quality.md`.
- **Over 1024 characters** — the agentskills.io specification ceiling, unchanged and independent
@@ -56,6 +54,14 @@ Flag as SUGGESTION if:
- A near-miss exclusion is present but targets a weak near-miss.
- An indirect trigger is present and warranted but could name the omitted phrasing more precisely.
**An unresolved boundary target is not graded here.** `validate.sh` owns that call and tiers it by
notation — `/name` or an arrow form is an ERROR, the bare prose form a SUGGESTION unless a second
target in the same sentence resolves — and Step 1 has already filed it under `### Structure` at that
tier. Re-grading it as a description FAIL puts one target in the report twice at two tiers. What is
left to judgment here is semantic and the script cannot reach it: whether a target that *does*
resolve is the right sibling to exclude, and whether a clause naming no target at all ("examine the
files manually") should have named one.
## body-discipline — `references/body-discipline.md`
Flag as FAIL if: