Six defects, each one a place where two files that an author reads in the same sitting told them different things — or where the trim dropped a rule and nothing noticed because no gate covers prose. **"Use proactively" contradicted itself across the pair.** All three agent templates said to add it where the runtime should delegate unprompted, while `agent-audit`'s `KyberforgeCopilot.ProactivePhrase` rule grades it a hard FAIL in any `*.agent.md` — which is the Copilot half of every project/user pair *and* the vendor-neutral plugin-scope file, since that compiles to a real Copilot agent downstream. Following the template produced a file the repo's own gate rejects. The phrase is now permitted in exactly one place, the Claude Code `.md`, and `references/contract.md` carries the per-file table plus the consequence authors ask about next: a pair whose CC half has it and whose Copilot half does not is correct, because `agent-audit` checks that both halves describe the same job, not that they match word for word. **The output-schema rule contradicted itself inside one file.** `contract.md` said any content only one branch reaches moves to `references/`, and then offered an "Output format template" body pattern with no qualification. Stated once now, so it is not re-litigated: an output schema stays in the body only when every flow produces it and it is roughly 50 words or less. No third option. **Gotchas tiers disagreed with the script.** `validate.sh` emits the entry count through `suggest()` and exits 0, while `skill-author` and `skill-audit` both called more than five entries a FAIL. Whether a given gotcha earns its place is judgment, so the prose moves to the script's tier rather than the reverse. The paraphrase rule stays a FAIL and is explicitly marked as the auditor's call — no script detects it. **The dispatch exemplar was cited at the wrong number.** `apm-workflow`'s body is 421 words; 554 is its whole-file count. Both `contract.md` and `body-discipline.md` cited 554 while describing a body budget, so an author calibrating against the exemplar overshot by ~30% — the exact whole-file/body-only conflation those two sections exist to warn against, reproduced inside the warning. **"Error handling" came back as a required body element.** It was one of four and is the one that gets dropped, and dropping it is not neutral: an agent handed malformed input with no instruction invents a recovery, and a subagent's invented recovery is invisible to its caller until the output is wrong. Restored in `agent-audit`'s rubric as a SUGGESTION, in `agent-author`'s contract and both scope checklists as a required element, and as an `## Errors` section in all three templates. **`skill-author` Step 4 gains the one check the audit misses.** An empty body reports `PASS SKILL.md body word count 0` — a word gate cannot tell "concise" from "absent". Step 4 now hand-checks for a non-empty section, and its commit verification is conditioned on actually being inside a git worktree, which a skill under `~/.claude/skills/` is not. Also here: absolute repo paths removed from `skill-author`'s SKILL.md and contract.md in favour of naming the skill (`zoom-out`'s description is quoted inline instead of pointed at), the boundary-target universe documented to match the resolver, a two-hops-from-SKILL.md limit on reference chains, and `new-agent.sh`'s next-steps output naming the description budget and the deliberate absence of an agent body gate. Refs: ADR-0020
167 lines
7.2 KiB
Markdown
167 lines
7.2 KiB
Markdown
---
|
|
source_keys:
|
|
- agentskills-spec
|
|
- agentskills-best-practices
|
|
---
|
|
|
|
# Body Discipline Reference
|
|
|
|
Upstream source: agentskills.io — skill-authoring, best-practices.
|
|
House contract: ADR-0020, the context budget.
|
|
|
|
## The core test
|
|
|
|
For every sentence in the body, ask: **"Would the agent get this wrong without this instruction?"**
|
|
|
|
If no — cut it. The agent already knows it from general training. Adding it wastes tokens and
|
|
dilutes the signal of what matters.
|
|
|
|
## What the body is for
|
|
|
|
The body carries the **decision procedure only**: ordered steps, decision branches, gates, and
|
|
which reference to load when.
|
|
|
|
Include content the agent lacks:
|
|
|
|
- Project-specific conventions and domain procedures it cannot infer
|
|
- Non-obvious edge cases and environment-specific gotchas
|
|
- The specific tools or sequences to use — not the full range of options
|
|
- One default per decision point with one escape hatch
|
|
|
|
Move to `references/`, behind an explicit "If X, read `references/file.md`" trigger — the literal
|
|
conditional form, never a generic pointer:
|
|
|
|
- Lookup tables and spec restatements
|
|
- Output schemas, templates and example blocks
|
|
- Rationale and justification prose
|
|
- Anything only one branch of the procedure ever reaches
|
|
|
|
Do not include at all:
|
|
|
|
- Concepts the agent already knows (what JSON is, how HTTP works, what a CSV is)
|
|
- Exhaustive option lists — pick a default; the agent does not benefit from choosing
|
|
- Steps the agent handles independently — over-specifying leads to unproductive paths
|
|
- Restatements of the description, which is already in context
|
|
|
|
## Two length families, measured differently
|
|
|
|
Do not conflate these, and do not report them as one finding.
|
|
|
|
| Gate | SUGGESTION | FAIL | Counts |
|
|
|---|---|---|---|
|
|
| Body budget (house, ADR-0020) | 600 words | 900 words | the **body only** — everything after the frontmatter's closing `---` |
|
|
| Spec conformance (agentskills.io) | — | 2,770 words / 500 lines | the **whole file**, frontmatter included |
|
|
|
|
The 2,770-word ceiling is a token-conformance backstop calibrated to the densest prose in the
|
|
corpus; it says nothing about quality and a file can sit a thousand words inside it while failing
|
|
the body budget. The 900-word ceiling is the quality gate: a body is loaded into the caller's live
|
|
context and competes with the conversation already there. `validate.sh` reports both. Cite whichever
|
|
one actually fired.
|
|
|
|
A word count cannot detect the defect it stands in for. Treat both numbers as backstops to the
|
|
dispatch rule and the Gotchas constraint below, never as a substitute for them.
|
|
|
|
## Dispatch is mandatory at two or more mutually exclusive flows
|
|
|
|
If a skill handles two or more flows that a single invocation cannot both take — separate
|
|
subcommands, separate input types, separate lifecycle stages — the body carries a **dispatch
|
|
table** plus the gates common to every branch, and each flow lives in its own self-contained
|
|
`references/` file. Inlining all of them is a FAIL regardless of word count, because every
|
|
invocation then pays for every branch it did not take.
|
|
|
|
The reference shape in this repo is `apm-workflow`: a **421-word body** dispatching to roughly
|
|
3,000 words of references across five mutually exclusive invocations. Its whole-file count is 554
|
|
words — cite 421 when calibrating a body, or the conflation this section warns against reappears
|
|
in the finding itself.
|
|
|
|
## Gotchas sections
|
|
|
|
The highest-value construct in a body, and the easiest to fill with noise. A Gotcha must state a
|
|
fact that **contradicts a reasonable default** — something the agent gets wrong precisely by acting
|
|
sensibly.
|
|
|
|
```markdown
|
|
## Gotchas
|
|
- The `users` table uses soft deletes. Always include `WHERE deleted_at IS NULL`.
|
|
- User ID is `user_id` in the database, `uid` in auth, `accountId` in billing. Same value.
|
|
```
|
|
|
|
Constraints:
|
|
|
|
- **More than five entries is a SUGGESTION** — five is the guideline, not a ceiling. Past five, the
|
|
section is usually a summary of the body rather than a set of traps, and the agent stops reading
|
|
it as a warning. It stays advisory because whether a given gotcha earns its place is judgment;
|
|
`validate.sh` emits it through `suggest()` and the run still exits 0.
|
|
- **A Gotcha that paraphrases a step in the body below it is a FAIL.** It has no independent
|
|
content, and it teaches the agent that Gotchas can be skimmed because the real instruction is
|
|
coming. This one is the auditor's call — no script detects it.
|
|
- **A Gotchas section exceeding 25% of the body is a SUGGESTION** — the body has been inverted into
|
|
a preamble. Same tier and same reasoning as the entry count, and independent of it: either can
|
|
fire without the other.
|
|
- Place the section near the top. A gotcha read after the mistake is worthless, which is also why
|
|
Gotchas is the one construct exempt from moving to `references/`.
|
|
|
|
Worked negative example — `git-commits` carries twelve entries, of which four restate content
|
|
that already appears below or in the description:
|
|
|
|
| Gotcha | Restates |
|
|
|---|---|
|
|
| `:31` "Communicates SemVer impact" | the description |
|
|
| `:32` "Confirmation gates are mandatory for destructive operations" | step 9 at `:52` |
|
|
| `:33` "Never skip hooks with `--no-verify`" | step 9 at `:52` |
|
|
| `:36` "Never commit secrets" | step 2 at `:45` |
|
|
|
|
All four are FAILs under the paraphrase rule. The entry count and the section's share of the body
|
|
(387 of 1,102 words, 35%) are two further SUGGESTIONs on top — the script reports both, and neither
|
|
fails the run on its own. What makes this worth auditing directly is that the four paraphrase FAILs
|
|
pass every word gate there is; only reading the construct finds them.
|
|
|
|
## Calibrating control
|
|
|
|
**Be prescriptive** when operations are fragile, consistency matters, or a specific sequence must be
|
|
followed:
|
|
|
|
```markdown
|
|
Run exactly:
|
|
\`\`\`bash
|
|
python scripts/migrate.py --verify --backup
|
|
\`\`\`
|
|
Do not modify the command or add additional flags.
|
|
```
|
|
|
|
**Give freedom** when multiple approaches are valid. Explaining *why* outperforms rigid directives —
|
|
agents make better decisions when they understand the purpose.
|
|
|
|
## Defaults not menus
|
|
|
|
Never present a list of equivalent options — pick one and mention the alternative briefly:
|
|
|
|
```markdown
|
|
# Too many options
|
|
Use pypdf, pdfplumber, PyMuPDF, or pdf2image...
|
|
|
|
# Default with escape hatch
|
|
Use pdfplumber for text extraction. For scanned PDFs requiring OCR, use pdf2image instead.
|
|
```
|
|
|
|
## Auditing guidance
|
|
|
|
Flag as FAIL if:
|
|
|
|
- A sentence answers "no" to the core test — it is padding
|
|
- The body exceeds 900 words counted body-only (`validate.sh` reports it)
|
|
- Two or more mutually exclusive flows are inlined instead of dispatched
|
|
- A Gotcha paraphrases a step in the body below it
|
|
- A decision point presents a menu of options with no default
|
|
- An instruction repeats content already in the description
|
|
- A prescriptive sequence is used where flexibility is fine, or the reverse
|
|
|
|
Flag as SUGGESTION if:
|
|
|
|
- The body exceeds 600 words counted body-only but stays at or under 900
|
|
- The Gotchas section carries more than five entries
|
|
- The Gotchas section exceeds 25% of the body
|
|
- A rationale is missing from an include/exclude rule — present but unexplained
|
|
- Gotchas are correct but placed late in the body rather than near the top
|
|
- Content that only one branch reaches is inlined where a `references/` file would serve
|