Files
holocron/plugins/kyberforge/.apm/skills/skill-audit/references/description-quality.md
Defame1297 fc305ba7d9 fix(kyberforge): correct the routing-tier contract and make agent-audit rubrics conditional
Five documents told authors that a prose-form dangling routing target blocks. The
gate reports it as a SUGGESTION and exits 0. Verified on fixtures: `-> name` and
`/name` are blocking ERRORs, the prose form is SUGGESTION-tier unless a second
resolving target in the same sentence corroborates it. ADR-0020 and gates.md were
right; contract.md, retrofit.md, description-quality.md, finding-criteria.md and
agent-author's contract.md were wrong — and they are what an author and an auditor
actually read. The whole 39-skill corpus was retrofitted against them.

skill-audit was also self-contradictory: it imports validate.sh's SUGGESTIONs into
the Structure dimension verbatim while its own rubric grades the same target a FAIL,
so one target got reported twice at two tiers. The script owns the grade; the rubric
now says so.

The YAML-fold trap that broke gitea-labels-milestones (#100) was warned about only
in retrofit.md, reachable only from the improve flow when a budget is exceeded. It
is now in both contract.md files, which SKILL.md mandates on the create flow too.

agent-audit loaded both rubrics unconditionally on every run — 3,323 words for a
clean audit against skill-audit's 1,636. dac9cad fixed exactly this in skill-audit
and edited agent-audit in the same commit without applying it. Same treatment: the
criteria move to a new finding-criteria.md and load per dimension. Clean run now
2,083 words, a 37% cut.

Routing: apm-workflow's description shed dependency installation while still owning
the flow, and apm-install's boundary did not exclude it, so "install my apm
dependencies" matched the CLI-binary skill with no route back. Fixed on both sides.
forge regains two of the three phrasings the retrofit deleted.

forge Step 1 called grill-with-docs unconditionally — a skill in plugins/bin, which
kyberforge does not declare as a dependency. It resolves here only because the
walk-up sweeps sibling plugins; a standalone install dead-ends. Step 1 now names
the cross-plugin dependency and gives an inline fallback. Declaring it properly in
apm.yml remains the better fix.

Also: both audit SKILL.md files now grade exit 2 as "did not run, dimension
unverified" rather than as findings; skill-audit's README row described content that
moved, which its own finding-criteria.md grades a FAIL; and body-discipline.md's
`git show <sha>:plugins/...` command is fenced, since an installed plugin cache has
no repo and file-structure.md makes a bare repo path a FAIL.

Refs: #100, #101, #125
ADR: 0020
2026-09-01 12:38:22 +00:00

4.3 KiB

source_keys
source_keys
agentskills-spec
agentskills-optimizing-descriptions

Description Quality Reference

Upstream source: agentskills.io — optimizing-descriptions, specification. House contract: ADR-0020, the context budget. The house contract is narrower than the spec rather than a reinterpretation of it: where both speak, both must be satisfied.

Why the description is the expensive part

At startup an agent loads only the name and description of every installed skill. The body is never seen until the skill triggers. The description therefore carries the entire triggering burden and is paid for in every session, whether the skill fires or not.

A second cost is less obvious and is a correctness hazard rather than a token cost: a description that summarises the workflow is a shortcut the agent takes instead of reading the body. A measured failure upstream — a description saying "code review between tasks" — produced one review where the body's flowchart specified two.

Step 0 — establish which contract applies

Read the frontmatter before judging a single word.

  • disable-model-invocation: true — the skill is hand-invoked. Its description is never matched against user intent, so it is not a routing string. It carries one plain human-facing sentence stating what the skill does. Audit it for that and nothing else. Reporting a missing trigger clause, a missing boundary clause or absent indirect triggers on a hand-invoked skill is a wrong finding, not a strict one.
  • No such flag — the skill is model-invoked and the rest of this file applies.

The three-part shape

A model-invoked description carries exactly three things:

  1. Trigger clause. When to invoke, phrased imperatively: Use when .... Not This skill ... — the agent is deciding whether to act, not reading a catalogue entry.
  2. At most one capability clause. What it does, in one clause. Never an enumeration.
  3. Boundary clause. Compressed form: Not <thing> -> <skill-name>. The target must resolve to a real skill directory or agent file in the authoring source. validate.sh checks that deterministically and grades it by notation: an unresolved /name or arrow target is an ERROR and reaches the report as a Structure FAIL, while an unresolved prose-form target ("use y instead") is only a SUGGESTION unless a second target in the same sentence resolves. Take the script's tier as given and report it once, under Structure.

Everything else belongs in the body or in README.md.

Indirect triggers — conditional, never blanket

Add "even if the user doesn't say X" only where the user's natural phrasing genuinely omits the domain word. True for the gitea-* family: people say "create an issue", not "create a Gitea issue". False for git-commits: nobody asks for a commit without saying commit. A blanket indirect-trigger clause on a skill whose domain word is unavoidable is padding charged to every session.

Near-miss exclusions

Add a boundary clause only where a sibling skill could plausibly steal the activation. Use strong near-misses — queries that share keywords but need something different — not weak ones ("write a fibonacci function"). One boundary clause per genuine near-miss; a list of four is enumeration wearing a boundary's clothes.

Before / after

# FAIL — enumeration first, mechanics as the opener, a blanket indirect trigger,
# and 300+ characters of it preloaded into every session forever.
description: >
  Analyze CSV and tabular data files — compute summary statistics, add derived
  columns, generate charts, and clean messy data. Use when the user has a CSV,
  TSV, or Excel file and wants to explore, transform, or visualize the data,
  even if they don't explicitly mention "CSV" or "analysis."

# PASS — trigger, one capability clause, boundary. The four verbs the FAIL
# version enumerates are the body's job; the router cannot act on them.
description: >
  Use when the user has a CSV, TSV, or Excel file and wants it explored,
  transformed, or charted. Not schema design -> data-model.

(data-model is illustrative. In a real description the target has to resolve.)

The FAIL and SUGGESTION criteria for this dimension live in references/finding-criteria.md, which Step 3 loads on every run.