fix(kyberforge): correct the routing-tier contract and make agent-audit rubrics conditional

Five documents told authors that a prose-form dangling routing target blocks. The
gate reports it as a SUGGESTION and exits 0. Verified on fixtures: `-> name` and
`/name` are blocking ERRORs, the prose form is SUGGESTION-tier unless a second
resolving target in the same sentence corroborates it. ADR-0020 and gates.md were
right; contract.md, retrofit.md, description-quality.md, finding-criteria.md and
agent-author's contract.md were wrong — and they are what an author and an auditor
actually read. The whole 39-skill corpus was retrofitted against them.

skill-audit was also self-contradictory: it imports validate.sh's SUGGESTIONs into
the Structure dimension verbatim while its own rubric grades the same target a FAIL,
so one target got reported twice at two tiers. The script owns the grade; the rubric
now says so.

The YAML-fold trap that broke gitea-labels-milestones (#100) was warned about only
in retrofit.md, reachable only from the improve flow when a budget is exceeded. It
is now in both contract.md files, which SKILL.md mandates on the create flow too.

agent-audit loaded both rubrics unconditionally on every run — 3,323 words for a
clean audit against skill-audit's 1,636. dac9cad fixed exactly this in skill-audit
and edited agent-audit in the same commit without applying it. Same treatment: the
criteria move to a new finding-criteria.md and load per dimension. Clean run now
2,083 words, a 37% cut.

Routing: apm-workflow's description shed dependency installation while still owning
the flow, and apm-install's boundary did not exclude it, so "install my apm
dependencies" matched the CLI-binary skill with no route back. Fixed on both sides.
forge regains two of the three phrasings the retrofit deleted.

forge Step 1 called grill-with-docs unconditionally — a skill in plugins/bin, which
kyberforge does not declare as a dependency. It resolves here only because the
walk-up sweeps sibling plugins; a standalone install dead-ends. Step 1 now names
the cross-plugin dependency and gives an inline fallback. Declaring it properly in
apm.yml remains the better fix.

Also: both audit SKILL.md files now grade exit 2 as "did not run, dimension
unverified" rather than as findings; skill-audit's README row described content that
moved, which its own finding-criteria.md grades a FAIL; and body-discipline.md's
`git show <sha>:plugins/...` command is fenced, since an installed plugin cache has
no repo and file-structure.md makes a bare repo path a FAIL.

Refs: #100, #101, #125
ADR: 0020
This commit is contained in:
2026-09-01 12:38:22 +00:00
parent 59f27dbd94
commit fc305ba7d9
36 changed files with 424 additions and 228 deletions

View File

@@ -10,8 +10,9 @@ Additional documentation agents load on demand.
| File | Purpose |
|------|---------|
| `description-quality.md` | Rubric for the description dimension — three-part shape, the 250/400-character budget, the hand-invoked contract, and the internal-mechanics FAIL. |
| `body-and-delegation.md` | Rubric for the body, delegation and comment-discipline dimensions — the delegation FAIL, why agents take no body word gate, and what an agent body is for. |
| `finding-criteria.md` | Every dimension's FAIL and SUGGESTION criteria — the one Step 3 file read on every run; it decides which rubrics below are worth loading. |
| `description-quality.md` | Rubric for the description dimension — why the description is the expensive part, the hand-invoked contract, the three-part shape, indirect triggers, and near-miss exclusions. |
| `body-and-delegation.md` | Rubric for the body, delegation and comment-discipline dimensions — the core test, the delegation FAIL, why agents take no body word gate, and what an agent body is for. |
| `scope-plugin-apm.md` | Contract for a single vendor-neutral `.apm/agents/<name>.agent.md` file — allowlist, dimension routing, and the dimensions that do not apply. |
| `scope-project-user.md` | Contract for a Claude Code / Copilot file pair — counterpart derivation, provider field rules, and pair consistency. |
| `validation-scripts.md` | Loaded only when a Step 1 script fails or cannot run — scope-detection walk-up, manual fallback checks, and known script failures. |

View File

@@ -98,24 +98,8 @@ to a shipped file. At plugin/APM scope the stakes are higher than tidiness: `apm
frontmatter verbatim to every target, `<!-- ... -->` is not valid YAML, and `validate.sh` FAILs a
frontmatter block that still contains one.
## Auditing guidance
## Where the criteria live
Flag as FAIL if:
- The body restates a procedure owned by a skill the agent can invoke — Fix: invoke `<skill>`
instead
- A sentence answers "no" to the core test — it is padding
- A decision point presents a menu of options with no default
- An instruction repeats content already in the description
- Frontmatter comments are template scaffolding rather than instruction, or are HTML comments at
plugin/APM scope
- A prescriptive sequence is used where flexibility is fine, or the reverse
Flag as SUGGESTION if:
- The body does not open with a direct role instruction
- The body specifies no error handling — nothing tells the agent what to do with malformed,
missing or contradictory input
- The job the agent describes is unbounded, or bounded only implicitly
- A rationale is missing from a rule the agent is expected to enforce — present but unexplained
- Comments are useful but verbose enough to bury the field they annotate
Every FAIL and SUGGESTION criterion for these dimensions is in `references/finding-criteria.md`,
which Step 3 reads on every run. This file is the reasoning behind them, loaded only when that file
puts the body, delegation or comment-discipline dimension in play.

View File

@@ -82,49 +82,8 @@ description: >
dispatched and safety-gated. Not conversational git help -> git-workflow.
```
## Auditing guidance
## Where the criteria live
Flag as FAIL if:
- **Over 400 characters.** Measured on the folded YAML value, not the raw source lines.
`validate.sh` reports the number; do not re-derive it, but do point the Fix at what to cut. Agent
descriptions have no platform-documented ceiling of their own — unlike a skill's 1,024-character
spec limit, the 400-character house ceiling is the only hard limit there is, so do not go looking
for a backstop behind it.
- **Internal mechanics appear in the description.** Any of:
- capability enumeration or a feature list;
- output-format detail ("Produces a compact findings report with Why and Fix per finding");
- composition or architecture notes ("composes X rather than duplicating Y", "a cross-cutting
shared agent", "the human-facing entry point", "replaces the old flat invocation");
- implementation detail ("self-validates via a bundled deterministic script").
None of it can change a routing decision and all of it is preloaded.
`Kyberforge.CompositionNote` catches the common phrasings deterministically; the rest is
judgment. This is the rule that deflates a description, so apply it before reaching for length.
- **The same trigger stated twice in two registers** — a verb list, then the same verbs re-quoted
as user phrasings, usually in the same order. One register, whichever routes better.
- **Descriptive rather than imperative phrasing** (`This agent ...`, `This is the ...`).
`Kyberforge.DescriptionOpener` catches any opener matching `^This`. There is no action-verb rule
here and never was a defensible one: an `Orchestrates ...` or `Audits ...` opener is a catalogue
entry, not a trigger.
- **Vague capabilities** ("helps with agents" where "audits an agent definition pair" was
available). `Kyberforge.VagueWording` catches the known filler; imprecision outside that list is
judgment.
- **A boundary clause naming a target that does not resolve** to a real skill directory or agent
file in the authoring source. `validate.sh` resolves this for agent files at both scopes and
reports each unresolved target itself — take its verdict rather than re-resolving the name by
hand, because a hand-walk over a different universe can contradict it. What is left to you is
semantic and the script cannot reach it: whether a target that *does* resolve is the right
sibling to exclude, and whether a clause naming no target at all ("examine the files manually")
should have named one.
- **`Use proactively` in a Copilot or vendor-neutral description.**
`KyberforgeCopilot.ProactivePhrase` catches it. The phrase steers the Claude Code runtime and
does nothing anywhere else, so in a `.agent.md` it is preloaded text that buys no behaviour.
- **Trigger-list, boundary or indirect-trigger content on a hand-invoked agent** — see Step 0.
Flag as SUGGESTION if:
- **Over 250 characters** but at or under 400. This tier is what moves the corpus average; the FAIL
tier only stops outliers. Report it rather than treating a 399-character description as clean.
- A near-miss exclusion is present but targets a weak near-miss.
- An indirect trigger is present and warranted but could name the omitted phrasing more precisely.
Every FAIL and SUGGESTION criterion for this dimension is in `references/finding-criteria.md`,
which Step 3 reads on every run. This file is the reasoning behind them, loaded only when that file
puts the description dimension in play.

View File

@@ -0,0 +1,98 @@
---
source_keys:
- context7-websites-code-claude
- claude-code-plugins-docs
- claude-code-subagents-docs
- context7-github-en-copilot
- github-custom-agents-configuration
---
# Finding Criteria
Every FAIL and SUGGESTION criterion, for every qualitative dimension, and nothing else. The
reasoning each criterion stands on, its worked examples and its house rules stay in that
dimension's rubric, which Step 3 loads only for a dimension this file puts in play.
Two rules on using it:
- A criterion that plainly applies is a finding. Write it up citing file and line.
- A criterion that might apply, or whose call the wording here does not settle, is a reason to load
that dimension's rubric — never a reason to drop the candidate. This file decides which rubrics
to read; it does not settle a close call on its own.
## description — `references/description-quality.md`
Flag as FAIL if:
- **Over 400 characters.** Measured on the folded YAML value, not the raw source lines.
`validate.sh` reports the number; do not re-derive it, but do point the Fix at what to cut. Agent
descriptions have no platform-documented ceiling of their own, so 400 is the only hard limit
there is — do not go looking for a backstop behind it.
- **Internal mechanics appear in the description.** Any of:
- capability enumeration or a feature list;
- output-format detail ("Produces a compact findings report with Why and Fix per finding");
- composition or architecture notes ("composes X rather than duplicating Y", "a cross-cutting
shared agent", "the human-facing entry point", "replaces the old flat invocation");
- implementation detail ("self-validates via a bundled deterministic script").
None of it can change a routing decision and all of it is preloaded.
`Kyberforge.CompositionNote` catches the common phrasings deterministically; the rest is
judgment. This is the rule that deflates a description, so apply it before reaching for length.
- **The same trigger stated twice in two registers** — a verb list, then the same verbs re-quoted
as user phrasings, usually in the same order. One register, whichever routes better.
- **Descriptive rather than imperative phrasing** (`This agent ...`, `This is the ...`).
`Kyberforge.DescriptionOpener` catches any opener matching `^This`.
- **Vague capabilities** ("helps with agents" where "audits an agent definition pair" was
available). `Kyberforge.VagueWording` catches the known filler; imprecision outside that list is
judgment.
- **`Use proactively` in a Copilot or vendor-neutral description.**
`KyberforgeCopilot.ProactivePhrase` catches it. The phrase steers the Claude Code runtime and
does nothing anywhere else, so in a `.agent.md` it is preloaded text that buys no behaviour.
- **Trigger-list, boundary or indirect-trigger content on a hand-invoked agent** — see Step 0 of
`references/description-quality.md`.
Flag as SUGGESTION if:
- **Over 250 characters** but at or under 400. This tier is what moves the corpus average; the FAIL
tier only stops outliers. Report it rather than treating a 399-character description as clean.
- A near-miss exclusion is present but targets a weak near-miss.
- An indirect trigger is present and warranted but could name the omitted phrasing more precisely.
**An unresolved boundary target is not graded here.** `validate.sh` resolves boundary targets for
agent files at both scopes and tiers the verdict itself — route notation (`/name`, an arrow form)
is an ERROR, the bare prose form a SUGGESTION unless a second target in the same sentence resolves.
Step 1 has already filed it under `### Structure` at that tier. Take the script's verdict rather
than re-resolving the name by hand, and do not re-grade it under description: a hand-walk over a
different universe can contradict the script, and re-grading puts one target in the report twice.
What is left to judgment is semantic and the script cannot reach it: whether a target that *does*
resolve is the right sibling to exclude, and whether a clause naming no target at all ("examine the
files manually") should have named one.
## body, delegation and comment-discipline — `references/body-and-delegation.md`
Flag as FAIL if:
- The body restates a procedure owned by a skill the agent can invoke — Fix: invoke `<skill>`
instead
- A sentence answers "no" to the core test — it is padding
- A decision point presents a menu of options with no default
- An instruction repeats content already in the description
- Frontmatter comments are template scaffolding rather than instruction, or are HTML comments at
plugin/APM scope
- A prescriptive sequence is used where flexibility is fine, or the reverse
Flag as SUGGESTION if:
- The body does not open with a direct role instruction
- The body specifies no error handling — nothing tells the agent what to do with malformed,
missing or contradictory input
- The job the agent describes is unbounded, or bounded only implicitly
- A rationale is missing from a rule the agent is expected to enforce — present but unexplained
- Comments are useful but verbose enough to bury the field they annotate
**Never report an agent body as too long on a word count.** ADR-0020 gates a skill body at
600/900 words and deliberately gates an agent body at nothing, because an agent body *becomes* the
system prompt of a fresh context rather than competing with a live conversation. There is no number
to cite. The one length signal that applies is the Copilot runtime's 30,000-character body limit,
which `validate.sh` already reports as a SUGGESTION. Length is judged through the delegation FAIL
above instead.

View File

@@ -14,7 +14,7 @@ source_keys:
- **URL:** context7:/websites/code_claude
- **Research doc:** plugins/kyberforge/docs/research/docs/claude-code-plugins/sources.md
- **Description:** Official Claude Code documentation site indexed by Context7 — plugin manifest schema, subagent definition types, marketplace JSON format, agent markdown file format
- **Contributing files:** SKILL.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-project-user.md
- **Contributing files:** SKILL.md, references/finding-criteria.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-project-user.md
- **Status:** `extracted`
## claude-code-plugins-docs
@@ -22,7 +22,7 @@ source_keys:
- **URL:** https://code.claude.com/docs/en/plugins
- **Research doc:** plugins/kyberforge/docs/research/docs/claude-code-plugins/sources.md
- **Description:** Official Claude Code plugin authoring guide — plugin structure, manifest fields, loading methods, skill namespacing, agent activation, marketplace submission
- **Contributing files:** SKILL.md, references/field-inventory.md, references/body-and-delegation.md, references/scope-plugin-apm.md, references/validation-scripts.md
- **Contributing files:** SKILL.md, references/finding-criteria.md, references/field-inventory.md, references/body-and-delegation.md, references/scope-plugin-apm.md, references/validation-scripts.md
- **Status:** `extracted`
## claude-code-subagents-docs
@@ -30,7 +30,7 @@ source_keys:
- **URL:** https://code.claude.com/docs/en/sub-agents
- **Research doc:** plugins/kyberforge/docs/research/docs/claude-code-plugins/sources.md
- **Description:** Official Claude Code subagent reference — definition format, all frontmatter fields, scope priority, built-in agents, CLI flags, environment variables, known limitations
- **Contributing files:** SKILL.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-plugin-apm.md, references/scope-project-user.md, references/validation-scripts.md
- **Contributing files:** SKILL.md, references/finding-criteria.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-plugin-apm.md, references/scope-project-user.md, references/validation-scripts.md
- **Status:** `extracted`
## context7-github-en-copilot
@@ -38,7 +38,7 @@ source_keys:
- **URL:** context7:/websites/github_en_copilot
- **Research doc:** plugins/kyberforge/docs/research/docs/github-copilot-plugins/sources.md
- **Description:** Official GitHub Copilot documentation indexed by Context7; covers CLI plugins, custom agents, SDK, and marketplace
- **Contributing files:** SKILL.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-project-user.md
- **Contributing files:** SKILL.md, references/finding-criteria.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-project-user.md
- **Status:** `extracted`
## github-custom-agents-configuration
@@ -46,7 +46,7 @@ source_keys:
- **URL:** https://docs.github.com/en/copilot/reference/custom-agents-configuration
- **Research doc:** plugins/kyberforge/docs/research/docs/github-copilot-plugins/sources.md
- **Description:** Reference for cloud and IDE custom agent definition format — frontmatter fields, tool aliases, MCP server config, secrets interpolation, scoping hierarchy
- **Contributing files:** SKILL.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-plugin-apm.md, references/scope-project-user.md, references/validation-scripts.md
- **Contributing files:** SKILL.md, references/finding-criteria.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-plugin-apm.md, references/scope-project-user.md, references/validation-scripts.md
- **Status:** `extracted`
## github-cli-plugin-reference