Skill name+description pairs are preloaded into every session, costing ~6,200 tokens across 39 skills before any skill is invoked. The authoring rules mandated that growth: skill-author:104 and description-quality.md:21 both required padding, while skill-author:102 (the deflating rule) had no FAIL condition behind it. Gates (blocking, no baseline file): - description 250 chars SUGGESTION / 400 FAIL, measured on the folded YAML value - body-only 600 words SUGGESTION / 900 FAIL, independent of the unchanged whole-file 2770-word / 500-line spec backstop - every boundary-clause routing target must resolve to a real skill or agent; catches skill-improve, neuledge-context and gitea-labels - agents take the description gates but deliberately no body gate; a test pins that absence Vale: DescriptionOpener widened to ^This\b, new CompositionNote rule banning architecture notes from descriptions. 10 hits, 0 false positives. Kyberforge's own four skills retrofitted: descriptions 3,364 -> 938 chars (-72%), bodies 8,306 -> 2,487 words (-70%), all via the apm-workflow dispatch pattern. Fixes the skill-improve dangling route and the agent-author misroute to manual review. Also fixes a pre-existing false positive where any line-initial 'read ' was flagged as interactive input, which had already caused two scripts to be rewritten around it. Refs: ADR-0020
6.5 KiB
source_keys
| source_keys | ||
|---|---|---|
|
Body Discipline Reference
Upstream source: agentskills.io — skill-authoring, best-practices. House contract: ADR-0020, the context budget.
The core test
For every sentence in the body, ask: "Would the agent get this wrong without this instruction?"
If no — cut it. The agent already knows it from general training. Adding it wastes tokens and dilutes the signal of what matters.
What the body is for
The body carries the decision procedure only: ordered steps, decision branches, gates, and which reference to load when.
Include content the agent lacks:
- Project-specific conventions and domain procedures it cannot infer
- Non-obvious edge cases and environment-specific gotchas
- The specific tools or sequences to use — not the full range of options
- One default per decision point with one escape hatch
Move to references/, behind an explicit "If X, read references/file.md" trigger — the literal
conditional form, never a generic pointer:
- Lookup tables and spec restatements
- Output schemas, templates and example blocks
- Rationale and justification prose
- Anything only one branch of the procedure ever reaches
Do not include at all:
- Concepts the agent already knows (what JSON is, how HTTP works, what a CSV is)
- Exhaustive option lists — pick a default; the agent does not benefit from choosing
- Steps the agent handles independently — over-specifying leads to unproductive paths
- Restatements of the description, which is already in context
Two length families, measured differently
Do not conflate these, and do not report them as one finding.
| Gate | SUGGESTION | FAIL | Counts |
|---|---|---|---|
| Body budget (house, ADR-0020) | 600 words | 900 words | the body only — everything after the frontmatter's closing --- |
| Spec conformance (agentskills.io) | — | 2,770 words / 500 lines | the whole file, frontmatter included |
The 2,770-word ceiling is a token-conformance backstop calibrated to the densest prose in the
corpus; it says nothing about quality and a file can sit a thousand words inside it while failing
the body budget. The 900-word ceiling is the quality gate: a body is loaded into the caller's live
context and competes with the conversation already there. validate.sh reports both. Cite whichever
one actually fired.
A word count cannot detect the defect it stands in for. Treat both numbers as backstops to the dispatch rule and the Gotchas constraint below, never as a substitute for them.
Dispatch is mandatory at two or more mutually exclusive flows
If a skill handles two or more flows that a single invocation cannot both take — separate
subcommands, separate input types, separate lifecycle stages — the body carries a dispatch
table plus the gates common to every branch, and each flow lives in its own self-contained
references/ file. Inlining all of them is a FAIL regardless of word count, because every
invocation then pays for every branch it did not take.
The reference shape in this repo is apm-workflow: a 554-word body dispatching to roughly 3,000
words of references across five mutually exclusive invocations.
Gotchas sections
The highest-value construct in a body, and the easiest to fill with noise. A Gotcha must state a fact that contradicts a reasonable default — something the agent gets wrong precisely by acting sensibly.
## Gotchas
- The `users` table uses soft deletes. Always include `WHERE deleted_at IS NULL`.
- User ID is `user_id` in the database, `uid` in auth, `accountId` in billing. Same value.
Constraints:
- Maximum five entries. Past five, the section is a summary of the body rather than a set of traps, and the agent stops reading it as a warning.
- A Gotcha that paraphrases a step in the body below it is a FAIL. It has no independent content, and it teaches the agent that Gotchas can be skimmed because the real instruction is coming.
- A Gotchas section exceeding 25% of the body is a SUGGESTION — the body has been inverted into a preamble.
- Place the section near the top. A gotcha read after the mistake is worthless, which is also why
Gotchas is the one construct exempt from moving to
references/.
Worked negative example — git-commits carries thirteen entries, of which four restate content
that already appears below or in the description:
| Gotcha | Restates |
|---|---|
:31 "Communicates SemVer impact" |
the description |
:32 "Confirmation gates are mandatory for destructive operations" |
step 9 at :52 |
:33 "Never skip hooks with --no-verify" |
step 9 at :52 |
:36 "Never commit secrets" |
step 2 at :45 |
All four are FAILs under this rule, and the section as a whole breaches the five-entry maximum. It also passes every plausible word gate, which is the point of auditing the construct directly.
Calibrating control
Be prescriptive when operations are fragile, consistency matters, or a specific sequence must be followed:
Run exactly:
\`\`\`bash
python scripts/migrate.py --verify --backup
\`\`\`
Do not modify the command or add additional flags.
Give freedom when multiple approaches are valid. Explaining why outperforms rigid directives — agents make better decisions when they understand the purpose.
Defaults not menus
Never present a list of equivalent options — pick one and mention the alternative briefly:
# Too many options
Use pypdf, pdfplumber, PyMuPDF, or pdf2image...
# Default with escape hatch
Use pdfplumber for text extraction. For scanned PDFs requiring OCR, use pdf2image instead.
Auditing guidance
Flag as FAIL if:
- A sentence answers "no" to the core test — it is padding
- The body exceeds 900 words counted body-only (
validate.shreports it) - Two or more mutually exclusive flows are inlined instead of dispatched
- A Gotcha paraphrases a step in the body below it, or the section exceeds five entries
- A decision point presents a menu of options with no default
- An instruction repeats content already in the description
- A prescriptive sequence is used where flexibility is fine, or the reverse
Flag as SUGGESTION if:
- The body exceeds 600 words counted body-only but stays at or under 900
- The Gotchas section exceeds 25% of the body
- A rationale is missing from an include/exclude rule — present but unexplained
- Gotchas are correct but placed late in the body rather than near the top
- Content that only one branch reaches is inlined where a
references/file would serve