Files
holocron/plugins/kyberforge/skills/skill-audit/references/body-discipline.md
Defame1297 311e7cd22c fix(kyberforge): reconcile the authoring rules the ADR-0020 trim left disagreeing
Six defects, each one a place where two files that an author reads in the same
sitting told them different things — or where the trim dropped a rule and nothing
noticed because no gate covers prose.

**"Use proactively" contradicted itself across the pair.** All three agent
templates said to add it where the runtime should delegate unprompted, while
`agent-audit`'s `KyberforgeCopilot.ProactivePhrase` rule grades it a hard FAIL in
any `*.agent.md` — which is the Copilot half of every project/user pair *and* the
vendor-neutral plugin-scope file, since that compiles to a real Copilot agent
downstream. Following the template produced a file the repo's own gate rejects.
The phrase is now permitted in exactly one place, the Claude Code `.md`, and
`references/contract.md` carries the per-file table plus the consequence authors
ask about next: a pair whose CC half has it and whose Copilot half does not is
correct, because `agent-audit` checks that both halves describe the same job, not
that they match word for word.

**The output-schema rule contradicted itself inside one file.** `contract.md`
said any content only one branch reaches moves to `references/`, and then offered
an "Output format template" body pattern with no qualification. Stated once now,
so it is not re-litigated: an output schema stays in the body only when every flow
produces it and it is roughly 50 words or less. No third option.

**Gotchas tiers disagreed with the script.** `validate.sh` emits the entry count
through `suggest()` and exits 0, while `skill-author` and `skill-audit` both
called more than five entries a FAIL. Whether a given gotcha earns its place is
judgment, so the prose moves to the script's tier rather than the reverse. The
paraphrase rule stays a FAIL and is explicitly marked as the auditor's call — no
script detects it.

**The dispatch exemplar was cited at the wrong number.** `apm-workflow`'s body is
421 words; 554 is its whole-file count. Both `contract.md` and `body-discipline.md`
cited 554 while describing a body budget, so an author calibrating against the
exemplar overshot by ~30% — the exact whole-file/body-only conflation those two
sections exist to warn against, reproduced inside the warning.

**"Error handling" came back as a required body element.** It was one of four and
is the one that gets dropped, and dropping it is not neutral: an agent handed
malformed input with no instruction invents a recovery, and a subagent's invented
recovery is invisible to its caller until the output is wrong. Restored in
`agent-audit`'s rubric as a SUGGESTION, in `agent-author`'s contract and both
scope checklists as a required element, and as an `## Errors` section in all three
templates.

**`skill-author` Step 4 gains the one check the audit misses.** An empty body
reports `PASS SKILL.md body word count 0` — a word gate cannot tell "concise"
from "absent". Step 4 now hand-checks for a non-empty section, and its commit
verification is conditioned on actually being inside a git worktree, which a skill
under `~/.claude/skills/` is not.

Also here: absolute repo paths removed from `skill-author`'s SKILL.md and
contract.md in favour of naming the skill (`zoom-out`'s description is quoted
inline instead of pointed at), the boundary-target universe documented to match
the resolver, a two-hops-from-SKILL.md limit on reference chains, and
`new-agent.sh`'s next-steps output naming the description budget and the
deliberate absence of an agent body gate.

Refs: ADR-0020
2026-08-16 16:40:51 +00:00

7.2 KiB

source_keys
source_keys
agentskills-spec
agentskills-best-practices

Body Discipline Reference

Upstream source: agentskills.io — skill-authoring, best-practices. House contract: ADR-0020, the context budget.

The core test

For every sentence in the body, ask: "Would the agent get this wrong without this instruction?"

If no — cut it. The agent already knows it from general training. Adding it wastes tokens and dilutes the signal of what matters.

What the body is for

The body carries the decision procedure only: ordered steps, decision branches, gates, and which reference to load when.

Include content the agent lacks:

  • Project-specific conventions and domain procedures it cannot infer
  • Non-obvious edge cases and environment-specific gotchas
  • The specific tools or sequences to use — not the full range of options
  • One default per decision point with one escape hatch

Move to references/, behind an explicit "If X, read references/file.md" trigger — the literal conditional form, never a generic pointer:

  • Lookup tables and spec restatements
  • Output schemas, templates and example blocks
  • Rationale and justification prose
  • Anything only one branch of the procedure ever reaches

Do not include at all:

  • Concepts the agent already knows (what JSON is, how HTTP works, what a CSV is)
  • Exhaustive option lists — pick a default; the agent does not benefit from choosing
  • Steps the agent handles independently — over-specifying leads to unproductive paths
  • Restatements of the description, which is already in context

Two length families, measured differently

Do not conflate these, and do not report them as one finding.

Gate SUGGESTION FAIL Counts
Body budget (house, ADR-0020) 600 words 900 words the body only — everything after the frontmatter's closing ---
Spec conformance (agentskills.io) — 2,770 words / 500 lines the whole file, frontmatter included

The 2,770-word ceiling is a token-conformance backstop calibrated to the densest prose in the corpus; it says nothing about quality and a file can sit a thousand words inside it while failing the body budget. The 900-word ceiling is the quality gate: a body is loaded into the caller's live context and competes with the conversation already there. validate.sh reports both. Cite whichever one actually fired.

A word count cannot detect the defect it stands in for. Treat both numbers as backstops to the dispatch rule and the Gotchas constraint below, never as a substitute for them.

Dispatch is mandatory at two or more mutually exclusive flows

If a skill handles two or more flows that a single invocation cannot both take — separate subcommands, separate input types, separate lifecycle stages — the body carries a dispatch table plus the gates common to every branch, and each flow lives in its own self-contained references/ file. Inlining all of them is a FAIL regardless of word count, because every invocation then pays for every branch it did not take.

The reference shape in this repo is apm-workflow: a 421-word body dispatching to roughly 3,000 words of references across five mutually exclusive invocations. Its whole-file count is 554 words — cite 421 when calibrating a body, or the conflation this section warns against reappears in the finding itself.

Gotchas sections

The highest-value construct in a body, and the easiest to fill with noise. A Gotcha must state a fact that contradicts a reasonable default — something the agent gets wrong precisely by acting sensibly.

## Gotchas
- The `users` table uses soft deletes. Always include `WHERE deleted_at IS NULL`.
- User ID is `user_id` in the database, `uid` in auth, `accountId` in billing. Same value.

Constraints:

  • More than five entries is a SUGGESTION — five is the guideline, not a ceiling. Past five, the section is usually a summary of the body rather than a set of traps, and the agent stops reading it as a warning. It stays advisory because whether a given gotcha earns its place is judgment; validate.sh emits it through suggest() and the run still exits 0.
  • A Gotcha that paraphrases a step in the body below it is a FAIL. It has no independent content, and it teaches the agent that Gotchas can be skimmed because the real instruction is coming. This one is the auditor's call — no script detects it.
  • A Gotchas section exceeding 25% of the body is a SUGGESTION — the body has been inverted into a preamble. Same tier and same reasoning as the entry count, and independent of it: either can fire without the other.
  • Place the section near the top. A gotcha read after the mistake is worthless, which is also why Gotchas is the one construct exempt from moving to references/.

Worked negative example — git-commits carries twelve entries, of which four restate content that already appears below or in the description:

Gotcha Restates
:31 "Communicates SemVer impact" the description
:32 "Confirmation gates are mandatory for destructive operations" step 9 at :52
:33 "Never skip hooks with --no-verify" step 9 at :52
:36 "Never commit secrets" step 2 at :45

All four are FAILs under the paraphrase rule. The entry count and the section's share of the body (387 of 1,102 words, 35%) are two further SUGGESTIONs on top — the script reports both, and neither fails the run on its own. What makes this worth auditing directly is that the four paraphrase FAILs pass every word gate there is; only reading the construct finds them.

Calibrating control

Be prescriptive when operations are fragile, consistency matters, or a specific sequence must be followed:

Run exactly:
\`\`\`bash
python scripts/migrate.py --verify --backup
\`\`\`
Do not modify the command or add additional flags.

Give freedom when multiple approaches are valid. Explaining why outperforms rigid directives — agents make better decisions when they understand the purpose.

Defaults not menus

Never present a list of equivalent options — pick one and mention the alternative briefly:

# Too many options
Use pypdf, pdfplumber, PyMuPDF, or pdf2image...

# Default with escape hatch
Use pdfplumber for text extraction. For scanned PDFs requiring OCR, use pdf2image instead.

Auditing guidance

Flag as FAIL if:

  • A sentence answers "no" to the core test — it is padding
  • The body exceeds 900 words counted body-only (validate.sh reports it)
  • Two or more mutually exclusive flows are inlined instead of dispatched
  • A Gotcha paraphrases a step in the body below it
  • A decision point presents a menu of options with no default
  • An instruction repeats content already in the description
  • A prescriptive sequence is used where flexibility is fine, or the reverse

Flag as SUGGESTION if:

  • The body exceeds 600 words counted body-only but stays at or under 900
  • The Gotchas section carries more than five entries
  • The Gotchas section exceeds 25% of the body
  • A rationale is missing from an include/exclude rule — present but unexplained
  • Gotchas are correct but placed late in the body rather than near the top
  • Content that only one branch reaches is inlined where a references/ file would serve