Diffing each retrofitted SKILL.md against its replacement references/ files found rules that existed on main and now existed nowhere — relocated in intent, deleted in fact. A trim that loses a rule is not progressive disclosure, it is data loss with a smaller word count. Three had no survivor. The least-privilege guidance for `tools` kept its mechanics and lost the "restrict to what the agent needs" half, so the remaining text read as encouragement to omit the field. The improve flow lost its regression check, so nothing compared the closing audit against the pre-edit state and a PASS quietly becoming a SUGGESTION went unnoticed — restored on both halves of the author pair, since agent-author had dropped its equivalent too. And agent bodies lost "would the agent get this wrong without it?", which mattered more than it looks: ADR-0020 deliberately sets no body word gate for agents, three of the four already sit between 933 and 1,199 words, and the delegation check only fires on procedure a skill already owns. That heuristic was the only brake left. Two more were reachable only from the wrong scope. agent-author tells the reader to load only the file for the resolved scope, but the mcp__ glob syntax for disallowedTools and the five tools no subagent ever receives had both landed in project-user-scope.md. disallowedTools is the ONLY permitted fence at plugin/APM scope, so the scope that needs the syntax most could not reach it, and a plugin-scope run could write a body telling the agent to ask the user a question. Two documents were actively wrong rather than merely thin. agent-audit told auditors that validate.sh resolves boundary targets for skills only; it runs at both scopes, so the auditor was hand-resolving what the script had already decided and could contradict it. And skill-audit routed to its script-troubleshooting reference whenever validate.sh "fails" — but it exits 1 on ordinary content FAILs, the normal outcome for the whole #99 population, so 1,302 words loaded on nearly every audit. A context-budget regression inside the skill that enforces the context budget. Finally, two illustrations taught the shape the gate ERRORs on, unfenced, while an adjacent rubric called it a hard ERROR. LESSONS.md records the reference-chain depth rule flipping from "one level deep" to "two hops, never three". ADR-0020 is silent on it and the reversal rode entirely on the diff; the looser rule is what mandatory dispatch requires. Refs: #99 ADR: 0020 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015W3iwF9ncfRZddGBxsMCYi
skill-audit
Audit a skill directory against the agentskills.io specification and the house context-budget contract (ADR-0020). Runs structural validation then a qualitative review across description quality, body discipline, patterns, formatting, file structure, scripts, and internal consistency, plus a provenance chain check.
What it does
- Runs
scripts/validate.shandscripts/validate-provenance.shfor structural and provenance checks, plusscripts/vale-wrap.sh— a Vale prefilter that deterministically flags non-imperative description openers, composition and architecture notes, vague wording, padding phrases, and "There is/are" sentence openers - Reads all files in the skill directory
- Applies qualitative checks across five dimension groups, loading one rubric from
references/per group - Outputs a compact findings report — findings only, grouped by dimension, each with Why and Fix — and a result block with handoff to
skill-author
validate.sh enforces two independent length families that must not be conflated: the agentskills.io spec conformance ceilings (500 lines, 2,770 words, both counting the whole file) and the ADR-0020 context budget (250/400 description characters, 600/900 body-only words).
Alongside those it runs four shape checks that are not length measurements at all. Two are FAILs: every routing target named in the description — in the compressed Not <thing> -> <name> arrow and in the prose form — must resolve to a real skill or agent, and every references/<file>.md the body names must exist on disk. Three are SUGGESTIONs: a missing boundary clause, a Gotchas section over five entries, and a Gotchas section over 25% of the body. The resolution universe for boundary targets is derived by walking up from the audited SKILL.md — the authoring root above it, its own apm package, and that package's declared apm.yml dependencies — so a fresh clone and a machine that has run apm install return the same verdict. When no universe can be determined the check prints INFO ... DID NOT RUN and does not silently pass.
Usage
/skill-audit
Provide the path to the skill directory to audit when invoking.
Files
| File | Purpose |
|---|---|
SKILL.md |
Skill instructions for agents |
scripts/validate.sh |
Structural validator — checks name format, name matches directory, description presence and length, body-only word count, line and whole-file word ceilings, boundary-clause presence, boundary-target resolution, references/ pointer existence, Gotchas entry count and body share, placeholder detection, script executable bit, and interactive-prompt detection |
scripts/validate-provenance.sh |
Provenance validator — checks sources.md completeness, source_keys/slug consistency, Contributing files existence, bidirectional linkage, Research doc: fields, and upstream research doc alignment |
scripts/vale-wrap.sh |
Vale prefilter wrapper — runs the bundled Kyberforge Vale styles against SKILL.md and reports alerts as deterministic FAILs ahead of Step 3's qualitative review |
assets/vale/.vale.ini |
Vale configuration — points Vale at the bundled Kyberforge style path, self-located relative to vale-wrap.sh |
assets/vale/styles/Kyberforge/CompositionNote.yml |
Vale rule — flags composition and architecture notes in a description (e.g. "cross-cutting", "entry point", "rather than duplicating") |
assets/vale/styles/Kyberforge/DescriptionOpener.yml |
Vale rule — flags non-imperative "This..." description openers |
assets/vale/styles/Kyberforge/PaddingPhrase.yml |
Vale rule — flags generic "see references/" padding phrasing in conditional references |
assets/vale/styles/Kyberforge/SentenceOpenerThereIs.yml |
Vale rule — flags body sentences starting with "There is"/"There are" |
assets/vale/styles/Kyberforge/VagueWording.yml |
Vale rule — flags known filler wording (e.g. "helps with", "utilize") |
references/description-quality.md |
Rubric for the description dimension — three-part shape, the 250/400-character budget, the hand-invoked (disable-model-invocation) contract, and the internal-mechanics FAIL |
references/body-discipline.md |
Rubric for the body-discipline dimension — the core test, the 600/900 body-only budget against the 2,770-word whole-file backstop, the mandatory-dispatch rule, and the Gotchas constraints |
references/patterns.md |
Rubric for the patterns dimension — which instruction construct fits which job, and how each is correctly formed |
references/file-structure.md |
Rubric for the file-structure and internal-consistency dimensions — permitted directories, cross-plugin path rules and their two structural exemptions, README drift |
references/formatting-and-scripts.md |
Rubric for the formatting and scripts dimensions — heading and fencing conventions, and the agentic-use criteria for bundled scripts |
references/validation-scripts.md |
Step 1 troubleshooting — the manual structural fallback when validate.sh cannot run, and the script exit codes that are easy to misread (loaded only on a script failure) |
references/sources.md |
Provenance record — agentskills.io sources that informed this skill and which files each contributed to |
tests/validate.bats |
(source-only) Bats test suite for validate.sh |
tests/validate-provenance.bats |
(source-only) Bats test suite for validate-provenance.sh |
tests/README.md |
(source-only) Setup instructions for bats-support and bats-assert test dependencies |
Rows marked (source-only) exist in the authoring source (.apm/skills/skill-audit/) but are
not present in an installed plugin: scripts/sync-plugin-content.sh strips
<category>/<name>/tests when it generates the flat mirror, because these are dev-time fixtures no
plugin host needs to discover (ADR-0017). Run them from a repo checkout, not from an install.