Three related half-applied changes from #130, each leaving the corpus in a state its own documentation contradicts. Why: - `assets/templates/SKILL.md` shipped `metadata:` fully commented out, and `new-skill.sh` only substitutes SKILL_NAME. Every scaffolded skill therefore lacked the `metadata.version` ADR-0022 made mandatory and was blocked at first commit by the very hook this PR added. The commented example also read `"1.0"` — neither the `0.1.0` new-skill seed nor valid semver. - `agent-audit/references/scope-project-user.md` still joined `disable-model-invocation` and `user-invocable` with a slash — #125's defect verbatim — while pointing the reader at the file this PR had just corrected to say the opposite. - ADR-0022 required the "when present" bump conditional dropped and `metadata.version` moved into create.md's required list. It was dropped from SKILL.md but left in README.md, and the field was edited in place under a heading that still authorises removing it entirely. Implementation notes: - The template emits `metadata: version: "0.1.0"` live, captioned as required, with the optional keys left commented. `new-skill.bats` gains a case asserting a live key and three-part semver, so this cannot regress. - `description-quality.md` now asserts only what the vendored Copilot research supports: two fields with opposite defaults, and the retired `infer` replaced by the pair rather than by either alone. The unsupported negative it previously stated as fact is gone. - The `1.0.0` retrofit seed is stated in improve.md and retrofit.md, which the retrofit flow actually reads — create.md, where it lived, is unreachable from that path. The compression item moved out of the file-churn checklist, whose preamble excluded the wording-only change it covers. - Executable git commands in these three skills now carry the ADR-0023 rtk prefix. Refs: #125, #127 ADR: 0022, 0023 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EeH8SCbcrCAQrtymkNuhKP
5.8 KiB
name, description, allowed-tools, metadata
| name | description | allowed-tools | metadata | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| agent-audit | Use when the user wants an agent definition audited — "audit this agent", "review my agent file", "is this ready to ship" — or after hand-editing an agent outside agent-author. Not applying fixes -> agent-author. Not a skill directory -> skill-audit. | Bash Read |
|
Gotchas
- Do not narrate PASS/FAIL per check while auditing. Gather findings internally and surface them only in the Step 4 report. Narrating each check as you go is the default failure mode here.
- Agents take the same 250/400-character description gates as skills and no body word gate at all — an agent body becomes the system prompt of a fresh context, so the 900-word skill ceiling does not transfer. Judge an over-long agent body through the delegation check, never by word count.
- At plugin/APM scope the agent is a single vendor-neutral file by design, so provider safety stops meaning Claude-Code-versus-Copilot field leakage there.
- Vale reporting
0 filesscanned means NOT RUN, not clean. Fall back to full Step 3 judgment for every dimension it would have covered.
Step 1 — Deterministic checks
Resolve all three paths against this skill's own directory so they work from a repo checkout and an installed plugin cache alike. Run exactly:
bash scripts/validate.sh <agent-file>
bash scripts/validate-provenance.sh <agent-file>
bash scripts/vale-wrap.sh <agent-file> [<counterpart-file>]
validate.sh takes either half of a project/user-scope pair or the single plugin/APM-scope file, detects the provider from the extension and the scope by walking up, then checks required fields, kebab-case name, FILL IN: placeholders, template HTML comments left in frontmatter, the ADR-0020 description budget (250 chars SUGGESTION, 400 FAIL, measured on the folded YAML value) and the fields that scope permits. Its findings become the ### Structure dimension — its FAILs and its SUGGESTIONs both — except the ones the Step 2 scope contract re-routes.
If a validation script fails or cannot run — Bash denied, python3 or vale absent, references/field-inventory.md missing — read references/validation-scripts.md; what these scripts measure is not reproducible by reading.
validate-provenance.sh prints nothing on success, so read its exit code before you read its silence. 0 is a genuine pass, including the silent exit 0 at project or user scope, where plugin-scope provenance does not apply. 1 means real findings: its FAILs and INFOs become a separate ### Provenance dimension, and it emits Why and Fix itself — surface those verbatim. 2 means the check never ran — a bad argument or a missing dependency, reason on stderr, no findings and often no stdout at all. On a 2, report ### Provenance as unverified and quote the stderr reason; never grade it as a clean pass. validate.sh uses the same 2 tier.
vale-wrap.sh applies the bundled Kyberforge style as a prefilter. Pass no --config; the wrapper locates its own. At project/user scope pass both files of the pair, not only the one you were handed. Every rule is graded error, so every alert is a FAIL. Report each one citing its rule ID, filed under the dimension it belongs to, and do not re-derive it by judgment:
| Rule | Dimension |
|---|---|
Kyberforge.DescriptionOpener, Kyberforge.CompositionNote, Kyberforge.VagueWording, KyberforgeCopilot.ProactivePhrase |
description |
Kyberforge.SentenceOpenerThereIs, Kyberforge.PaddingPhrase |
body |
Step 2 — Read the agent and load its scope contract
Read the agent file end to end, and at project/user scope its counterpart too. A path containing .apm/agents/ is plugin/APM scope; anything else is project or user scope. Each contract names the dimensions that apply there and where validate.sh findings other than Structure belong:
| Scope | Read |
|---|---|
| plugin/APM | references/scope-plugin-apm.md |
| project, user | references/scope-project-user.md |
Step 3 — Qualitative audit
Read references/finding-criteria.md first — every dimension's FAIL and SUGGESTION criteria. Load the rubric below only for a dimension the criteria put in play: one carrying a candidate finding, or one where the criterion alone does not settle the call.
| Dimension | Rubric |
|---|---|
| description | references/description-quality.md |
| body, delegation, comment-discipline | references/body-and-delegation.md |
Each rubric is the reasoning behind its criteria, not a second copy of them. Cite file and line number for every finding.
Step 4 — Report
Open with a coverage line naming every dimension checked. At project/user scope:
Checked: structure · provider-safety · description · body · delegation · comment-discipline · pair-consistency · provenance
At plugin/APM scope, drop pair-consistency — there is no pair to check.
Then output only the dimensions that have findings, grouped under H3 headings, FAILs before SUGGESTIONs within each. Omit clean dimensions — their absence is what confirms they passed.
Each finding:
FAIL/SUGGESTION <finding> — file:line
Why: <why this is a problem>
Fix: <exact change — quote before/after where applicable>
Close with a ## Result block holding one line: PASS, PASS (N suggestions), or FAIL (N fails · M suggestions), each optionally followed by · P info. INFO findings are observational and never change PASS/FAIL; omit · P info when there are none. Add a second line, Run agent-author to address findings., whenever there is at least one finding. Do not apply fixes — report and propose only.