Diffing each retrofitted SKILL.md against its replacement references/ files found rules that existed on main and now existed nowhere — relocated in intent, deleted in fact. A trim that loses a rule is not progressive disclosure, it is data loss with a smaller word count. Three had no survivor. The least-privilege guidance for `tools` kept its mechanics and lost the "restrict to what the agent needs" half, so the remaining text read as encouragement to omit the field. The improve flow lost its regression check, so nothing compared the closing audit against the pre-edit state and a PASS quietly becoming a SUGGESTION went unnoticed — restored on both halves of the author pair, since agent-author had dropped its equivalent too. And agent bodies lost "would the agent get this wrong without it?", which mattered more than it looks: ADR-0020 deliberately sets no body word gate for agents, three of the four already sit between 933 and 1,199 words, and the delegation check only fires on procedure a skill already owns. That heuristic was the only brake left. Two more were reachable only from the wrong scope. agent-author tells the reader to load only the file for the resolved scope, but the mcp__ glob syntax for disallowedTools and the five tools no subagent ever receives had both landed in project-user-scope.md. disallowedTools is the ONLY permitted fence at plugin/APM scope, so the scope that needs the syntax most could not reach it, and a plugin-scope run could write a body telling the agent to ask the user a question. Two documents were actively wrong rather than merely thin. agent-audit told auditors that validate.sh resolves boundary targets for skills only; it runs at both scopes, so the auditor was hand-resolving what the script had already decided and could contradict it. And skill-audit routed to its script-troubleshooting reference whenever validate.sh "fails" — but it exits 1 on ordinary content FAILs, the normal outcome for the whole #99 population, so 1,302 words loaded on nearly every audit. A context-budget regression inside the skill that enforces the context budget. Finally, two illustrations taught the shape the gate ERRORs on, unfenced, while an adjacent rubric called it a hard ERROR. LESSONS.md records the reference-chain depth rule flipping from "one level deep" to "two hops, never three". ADR-0020 is silent on it and the reversal rode entirely on the diff; the looser rule is what mandatory dispatch requires. Refs: #99 ADR: 0020 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015W3iwF9ncfRZddGBxsMCYi
169 lines
7.4 KiB
Markdown
169 lines
7.4 KiB
Markdown
---
|
|
source_keys:
|
|
- agentskills-spec
|
|
- agentskills-best-practices
|
|
---
|
|
|
|
# Body Discipline Reference
|
|
|
|
Upstream source: agentskills.io — skill-authoring, best-practices.
|
|
House contract: ADR-0020, the context budget.
|
|
|
|
## The core test
|
|
|
|
For every sentence in the body, ask: **"Would the agent get this wrong without this instruction?"**
|
|
|
|
If no — cut it. The agent already knows it from general training. Adding it wastes tokens and
|
|
dilutes the signal of what matters.
|
|
|
|
## What the body is for
|
|
|
|
The body carries the **decision procedure only**: ordered steps, decision branches, gates, and
|
|
which reference to load when.
|
|
|
|
Include content the agent lacks:
|
|
|
|
- Project-specific conventions and domain procedures it cannot infer
|
|
- Non-obvious edge cases and environment-specific gotchas
|
|
- The specific tools or sequences to use — not the full range of options
|
|
- One default per decision point with one escape hatch
|
|
|
|
Move to `references/`, behind an explicit "If X, read `references/<file>.md`" trigger — the literal
|
|
conditional form, never a generic pointer. Write the real filename in the skill under audit; the
|
|
angle brackets are a placeholder here, and a literal `references/file.md` in a body is an ERROR
|
|
from the ADR-0020 gate because no such file exists on disk. Move:
|
|
|
|
- Lookup tables and spec restatements
|
|
- Output schemas, templates and example blocks
|
|
- Rationale and justification prose
|
|
- Anything only one branch of the procedure ever reaches
|
|
|
|
Do not include at all:
|
|
|
|
- Concepts the agent already knows (what JSON is, how HTTP works, what a CSV is)
|
|
- Exhaustive option lists — pick a default; the agent does not benefit from choosing
|
|
- Steps the agent handles independently — over-specifying leads to unproductive paths
|
|
- Restatements of the description, which is already in context
|
|
|
|
## Two length families, measured differently
|
|
|
|
Do not conflate these, and do not report them as one finding.
|
|
|
|
| Gate | SUGGESTION | FAIL | Counts |
|
|
|---|---|---|---|
|
|
| Body budget (house, ADR-0020) | 600 words | 900 words | the **body only** — everything after the frontmatter's closing `---` |
|
|
| Spec conformance (agentskills.io) | — | 2,770 words / 500 lines | the **whole file**, frontmatter included |
|
|
|
|
The 2,770-word ceiling is a token-conformance backstop calibrated to the densest prose in the
|
|
corpus; it says nothing about quality and a file can sit a thousand words inside it while failing
|
|
the body budget. The 900-word ceiling is the quality gate: a body is loaded into the caller's live
|
|
context and competes with the conversation already there. `validate.sh` reports both. Cite whichever
|
|
one actually fired.
|
|
|
|
A word count cannot detect the defect it stands in for. Treat both numbers as backstops to the
|
|
dispatch rule and the Gotchas constraint below, never as a substitute for them.
|
|
|
|
## Dispatch is mandatory at two or more mutually exclusive flows
|
|
|
|
If a skill handles two or more flows that a single invocation cannot both take — separate
|
|
subcommands, separate input types, separate lifecycle stages — the body carries a **dispatch
|
|
table** plus the gates common to every branch, and each flow lives in its own self-contained
|
|
`references/` file. Inlining all of them is a FAIL regardless of word count, because every
|
|
invocation then pays for every branch it did not take.
|
|
|
|
The reference shape in this repo is `apm-workflow`: a **421-word body** dispatching to roughly
|
|
3,000 words of references across five mutually exclusive invocations. Its whole-file count is 554
|
|
words — cite 421 when calibrating a body, or the conflation this section warns against reappears
|
|
in the finding itself.
|
|
|
|
## Gotchas sections
|
|
|
|
The highest-value construct in a body, and the easiest to fill with noise. A Gotcha must state a
|
|
fact that **contradicts a reasonable default** — something the agent gets wrong precisely by acting
|
|
sensibly.
|
|
|
|
```markdown
|
|
## Gotchas
|
|
- The `users` table uses soft deletes. Always include `WHERE deleted_at IS NULL`.
|
|
- User ID is `user_id` in the database, `uid` in auth, `accountId` in billing. Same value.
|
|
```
|
|
|
|
Constraints:
|
|
|
|
- **More than five entries is a SUGGESTION** — five is the guideline, not a ceiling. Past five, the
|
|
section is usually a summary of the body rather than a set of traps, and the agent stops reading
|
|
it as a warning. It stays advisory because whether a given gotcha earns its place is judgment;
|
|
`validate.sh` emits it through `suggest()` and the run still exits 0.
|
|
- **A Gotcha that paraphrases a step in the body below it is a FAIL.** It has no independent
|
|
content, and it teaches the agent that Gotchas can be skimmed because the real instruction is
|
|
coming. This one is the auditor's call — no script detects it.
|
|
- **A Gotchas section exceeding 25% of the body is a SUGGESTION** — the body has been inverted into
|
|
a preamble. Same tier and same reasoning as the entry count, and independent of it: either can
|
|
fire without the other.
|
|
- Place the section near the top. A gotcha read after the mistake is worthless, which is also why
|
|
Gotchas is the one construct exempt from moving to `references/`.
|
|
|
|
Worked negative example — `git-commits` carries twelve entries, of which four restate content
|
|
that already appears below or in the description:
|
|
|
|
| Gotcha | Restates |
|
|
|---|---|
|
|
| `:31` "Communicates SemVer impact" | the description |
|
|
| `:32` "Confirmation gates are mandatory for destructive operations" | step 9 at `:52` |
|
|
| `:33` "Never skip hooks with `--no-verify`" | step 9 at `:52` |
|
|
| `:36` "Never commit secrets" | step 2 at `:45` |
|
|
|
|
All four are FAILs under the paraphrase rule. The entry count and the section's share of the body
|
|
(387 of 1,102 words, 35%) are two further SUGGESTIONs on top — the script reports both, and neither
|
|
fails the run on its own. What makes this worth auditing directly is that the four paraphrase FAILs
|
|
pass every word gate there is; only reading the construct finds them.
|
|
|
|
## Calibrating control
|
|
|
|
**Be prescriptive** when operations are fragile, consistency matters, or a specific sequence must be
|
|
followed:
|
|
|
|
```markdown
|
|
Run exactly:
|
|
\`\`\`bash
|
|
python scripts/migrate.py --verify --backup
|
|
\`\`\`
|
|
Do not modify the command or add additional flags.
|
|
```
|
|
|
|
**Give freedom** when multiple approaches are valid. Explaining *why* outperforms rigid directives —
|
|
agents make better decisions when they understand the purpose.
|
|
|
|
## Defaults not menus
|
|
|
|
Never present a list of equivalent options — pick one and mention the alternative briefly:
|
|
|
|
```markdown
|
|
# Too many options
|
|
Use pypdf, pdfplumber, PyMuPDF, or pdf2image...
|
|
|
|
# Default with escape hatch
|
|
Use pdfplumber for text extraction. For scanned PDFs requiring OCR, use pdf2image instead.
|
|
```
|
|
|
|
## Auditing guidance
|
|
|
|
Flag as FAIL if:
|
|
|
|
- A sentence answers "no" to the core test — it is padding
|
|
- The body exceeds 900 words counted body-only (`validate.sh` reports it)
|
|
- Two or more mutually exclusive flows are inlined instead of dispatched
|
|
- A Gotcha paraphrases a step in the body below it
|
|
- A decision point presents a menu of options with no default
|
|
- An instruction repeats content already in the description
|
|
- A prescriptive sequence is used where flexibility is fine, or the reverse
|
|
|
|
Flag as SUGGESTION if:
|
|
|
|
- The body exceeds 600 words counted body-only but stays at or under 900
|
|
- The Gotchas section carries more than five entries
|
|
- The Gotchas section exceeds 25% of the body
|
|
- A rationale is missing from an include/exclude rule — present but unexplained
|
|
- Gotchas are correct but placed late in the body rather than near the top
|
|
- Content that only one branch reaches is inlined where a `references/` file would serve
|