Six defects, each one a place where two files that an author reads in the same sitting told them different things — or where the trim dropped a rule and nothing noticed because no gate covers prose. **"Use proactively" contradicted itself across the pair.** All three agent templates said to add it where the runtime should delegate unprompted, while `agent-audit`'s `KyberforgeCopilot.ProactivePhrase` rule grades it a hard FAIL in any `*.agent.md` — which is the Copilot half of every project/user pair *and* the vendor-neutral plugin-scope file, since that compiles to a real Copilot agent downstream. Following the template produced a file the repo's own gate rejects. The phrase is now permitted in exactly one place, the Claude Code `.md`, and `references/contract.md` carries the per-file table plus the consequence authors ask about next: a pair whose CC half has it and whose Copilot half does not is correct, because `agent-audit` checks that both halves describe the same job, not that they match word for word. **The output-schema rule contradicted itself inside one file.** `contract.md` said any content only one branch reaches moves to `references/`, and then offered an "Output format template" body pattern with no qualification. Stated once now, so it is not re-litigated: an output schema stays in the body only when every flow produces it and it is roughly 50 words or less. No third option. **Gotchas tiers disagreed with the script.** `validate.sh` emits the entry count through `suggest()` and exits 0, while `skill-author` and `skill-audit` both called more than five entries a FAIL. Whether a given gotcha earns its place is judgment, so the prose moves to the script's tier rather than the reverse. The paraphrase rule stays a FAIL and is explicitly marked as the auditor's call — no script detects it. **The dispatch exemplar was cited at the wrong number.** `apm-workflow`'s body is 421 words; 554 is its whole-file count. Both `contract.md` and `body-discipline.md` cited 554 while describing a body budget, so an author calibrating against the exemplar overshot by ~30% — the exact whole-file/body-only conflation those two sections exist to warn against, reproduced inside the warning. **"Error handling" came back as a required body element.** It was one of four and is the one that gets dropped, and dropping it is not neutral: an agent handed malformed input with no instruction invents a recovery, and a subagent's invented recovery is invisible to its caller until the output is wrong. Restored in `agent-audit`'s rubric as a SUGGESTION, in `agent-author`'s contract and both scope checklists as a required element, and as an `## Errors` section in all three templates. **`skill-author` Step 4 gains the one check the audit misses.** An empty body reports `PASS SKILL.md body word count 0` — a word gate cannot tell "concise" from "absent". Step 4 now hand-checks for a non-empty section, and its commit verification is conditioned on actually being inside a git worktree, which a skill under `~/.claude/skills/` is not. Also here: absolute repo paths removed from `skill-author`'s SKILL.md and contract.md in favour of naming the skill (`zoom-out`'s description is quoted inline instead of pointed at), the boundary-target universe documented to match the resolver, a two-hops-from-SKILL.md limit on reference chains, and `new-agent.sh`'s next-steps output naming the description budget and the deliberate absence of an agent body gate. Refs: ADR-0020
7.5 KiB
source_keys
| source_keys | ||
|---|---|---|
|
The agent description and body contract
House contract, set by ADR-0020. The counts and the boundary targets are enforced by
agent-audit's scripts/validate.sh; the prose patterns by the Vale styles it bundles; the
judgment calls by its reference files.
Why the budget exists
An agent's name and description is loaded into every session's context at startup, whether or
not the agent is ever delegated to — the same cost a skill's description carries, so agents take
the same numbers. The body is different: it is not loaded into the caller's conversation at all,
it becomes the system prompt of a fresh context when the agent runs. That is why the body has no
word gate here and a skill body has one.
Description
A description carries exactly three things:
- Trigger clause — when to delegate, imperative: "Use when …", never "This agent …". Describe the user's intent and the triggering condition, not the agent's internal mechanics.
- At most one capability clause — what it does, one clause, no enumeration. Be specific ("reviews a diff for injected credentials", not "helps with security").
- Boundary clause — form:
Not <thing> -> <name>.Add one only where a near-miss agent or skill could steal delegations.
Banned from a description; move it to the body or to README.md:
- Capability enumeration or feature lists
- Per-scope emission mechanics — which files the author skill writes at which scope changes no delegation decision
- Output-format detail ("Produces a compact findings report with Why and Fix per finding")
- Composition or architecture notes ("composes X rather than duplicating Y", "cross-cutting")
- Implementation detail ("Self-validates via a bundled deterministic script")
- Restating the same trigger twice in two registers — a verb list, then the same verbs re-quoted as user phrasings. This is a FAIL, not a suggestion.
Do not open with an action verb. "Reviews…", "Analyzes…", "Generates…" was the old house rule
and ADR-0020 deleted it: the opener is Use when, matching every skill in this corpus, so one
router reads one shape.
"Use proactively" is Claude Code-only, and conditional even there. The phrase steers the Claude Code runtime to offer an agent unprompted and does nothing anywhere else, so where it may appear depends on the file:
| File | Rule |
|---|---|
Claude Code .md (project/user scope) |
Allowed. Add it only where the runtime should delegate without the user naming the agent — an agent invoked by name does not need it, and it costs activations elsewhere when added by reflex. |
Copilot .agent.md (project/user scope) |
Never. Inert there, and KyberforgeCopilot.ProactivePhrase grades it a hard FAIL. |
Vendor-neutral .apm/agents/<name>.agent.md (plugin/APM scope) |
Never. Same Vale rule, same hard FAIL — the file matches the **/*.agent.md glob, and it compiles to a real Copilot agent downstream. |
A pair whose Claude Code half carries the phrase and whose Copilot half omits it is correct, not
inconsistent: agent-audit checks that both halves describe the same job, not that they match
word for word.
Indirect triggers ("even if the user doesn't say X") take a similar conditional at every scope: add one only where the user's natural phrasing genuinely omits the domain word.
Boundary targets must resolve. Both forms are checked — the arrow and the prose form ("do not
use for X, use y instead") — so a typo dangles either way. Targets resolve against a universe
built by walking up from the agent file itself: the nearest ancestor holding
plugins/*/.apm/{skills,agents} (or, failing that, the nearest ancestor holding .git) contributes
every skill and agent under <root>/plugins/*/, plus the agent's own apm package and the packages
that package declares in apm.yml under dependencies.apm. A sibling plugin in the same monorepo
therefore resolves; a skill in an unrelated repo does not. A target outside that universe sends the
router nowhere. Verify it before writing it — do not invent a plausible sibling.
Length. 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only with YAML folding resolved. Treat 250 as the target: the SUGGESTION tier is what moves the corpus average, the FAIL tier only stops outliers.
Body
Write the body as a direct role instruction, addressed to the agent:
You are a <role>. When invoked, <primary action>.
## Inputs
<what the agent is given: files, context, parameters>
## Process
<ordered steps; be explicit where ordering matters>
## Output
<what it produces: format, location, structure>
## Errors
<what to do on malformed, missing or contradictory input: report and stop, or
which fallback to take — and what to say to the caller either way>
Four required elements: inputs expected, process steps, output format, error handling. The last is the one that gets dropped, and dropping it is not neutral: an agent given a malformed input and no instruction invents a recovery, and a subagent's invented recovery is invisible to the caller until the output is wrong. Say explicitly whether the agent stops and reports, or degrades to a named fallback.
One job per agent. An agent covering two jobs gets delegated to for the wrong one.
Delegation discipline replaces the word gate. A plugin/APM agent is a single file with no
sibling references/ directory: it cannot disclose progressively to itself, so its only way to
stay short is to invoke rather than restate. A body that transcribes a procedure a skill it
can invoke already owns is an agent-audit FAIL, and the fix is one line — "invoke <skill>".
- Restating: "To commit, check the message against Conventional Commits: type, scope, description; header under 100 chars; …"
- Delegating: "Author commits with
git-commits."
The same holds for a procedure another agent owns. What belongs in the body is what no invocable skill covers: the agent's role, its boundaries, the order it works in, and the format it returns.
State a read-only boundary in prose, not only in frontmatter. disallowedTools denies the
tools it names and nothing else — never Bash, which an agent with no tools field inherits — so
an agent fenced only in frontmatter can still write through a shell redirect.
Invocation axis
Decide before writing the description whether the agent is model-delegated (the runtime picks it)
or reached only by name (@agent-<name>).
Only Copilot's cloud/IDE format expresses that in frontmatter: disable-model-invocation: true
requires explicit invocation, and user-invocable: false hides an agent from manual invocation.
Both live in .github/copilot/agents/<name>.md and are inert in the CLI format. Claude Code has
no equivalent field, and neither does the vendor-neutral plugin/APM file, so at those scopes a
name-invoked agent still needs a description precise enough not to steal delegations — the
boundary clause is doing that work.
One gate, two measurements
| Gate | SUGGESTION | FAIL | Counts |
|---|---|---|---|
| description | 250 chars | 400 chars | the description: value only |
| body (Copilot limit) | 30,000 chars | — | the body only; content past it is truncated silently |
The 30,000-character Copilot ceiling is a runtime truncation limit, not a quality target, and it applies to a plugin/APM file too — that file compiles into a real Copilot agent downstream. An agent body long enough to approach it has a delegation defect, not a length problem.