Retrofits all 39 skills to ADR-0020's description/body context contract, then fixes what six rounds of independent review found in that retrofit — including four ways the hot gate itself failed open. Closes #99, #107, #108, #110, #111, #114, #115, #120. ## The retrofit (waves 1-5) | | Start | Now | |---|---|---| | Description FAILs (>400 chars) | 26 | **0** | | Body FAILs (>900 words, body-only) | 9 | **0** | | Dangling routing targets | 2 | **0** | | `Kyberforge.CompositionNote` | 10 | **0** | | Preload tax | 21,005 chars | **~10,500** | Under the 12,000-char success criterion. Per-wave detail is on #99. ## The review fixes **The gate failed open four ways, three of them found after the retrofit shipped.** An unrecognised follower token made a dangling target vanish. A skill directory with no `SKILL.md` resolved as a valid target, so a commit could be green locally and red in a fresh clone — three existing fixtures were relying on that, one of which made the install-leak A/B pass vacuously. Then the free-standing `/name` sweep turned out to be gated on the sentence carrying a boundary marker, so route notation in any other sentence was invisible — not an ERROR, not a SUGGESTION, not an INFO — which left the documented "`/name` always blocks" promise false from a second direction. All four fixed and pinned. **Two checks were silently not running.** `validate-provenance.sh` checks 7-8 were dead across nine skills. Waking them exposed a deeper problem: they assume `Research doc:` names a source index, but 30 of 121 entries point at topic content documents, so every new check-7 INFO was a false positive and check 8 was saved from a false-FAIL flood only by an *unannounced* skip. Checks 7/8 are now scoped to source indexes and every skip announces itself (#121). **The retrofit's own anti-goal, four times.** ADR-0020 warns that a blunt gate gets satisfied by deleting content rather than relocating it. `diagnose` and `skill-audit` relocated prose and then read it unconditionally; `prototype` and `vale-config` deleted rules outright that survived nowhere. All four addressed. ## Verification - `bash tests/run-tests.sh --strict` — 24 suites, 0 skipped, 0 failed - `bash tests/run-bats.sh` — 325 tests, 0 failures - `pre-commit run --all-files` — 17/17 - `pre-commit run --hook-stage pre-push --all-files` — 16/16, with `apm marketplace check` and `apm pack --check-clean` run against the remote, not skipped - `scripts/skill-size-check.sh` over all 39 skills — rc 0, 0 ERROR/FAIL, SUGGESTION-only - Preload tax measured at **10,498 chars**, max description 390 — both inside budget - Every new test proven non-vacuous by a deliberate mutation of the behaviour it covers **Per-commit sync, stated accurately:** the ten commits from the latest review round each pass `check-plugin-content-sync` in isolation, verified by checking each out in a detached worktree with a clean between. The earlier gitea window (`dfacf05..bedbd1d`, nine commits) does **not** — its mirror was regenerated in one batch at `bbc7300`. An earlier revision of this description claimed the property held for every commit; it does not, and a bisect through that window lands on a red commit. **Squash-merge** to collapse it, or accept that this range is not bisectable. ## Version bump Six plugins and the catalog take a **patch**, not a minor. The branch is **89 commits — 40 `fix` / 30 `refactor` / 12 `docs` / 5 `chore` / 2 `test` — zero `feat`, zero `!`, zero `BREAKING CHANGE`** — and adds no skill, agent, command or hook. (Two earlier revisions of this section cited a stale histogram, most recently 78 commits; the figures above are measured at HEAD.) Both rules this repo ships (`forge/references/version-bump.md`, landing in this PR, and `git-commits/references/conventional-commits-spec.md`) make that a patch, and the catalog set is unchanged at 7 entries. Not settled by that: four published files were removed from the installed tree, three moved, and `caveman` gained `disable-model-invocation`, retiring its old triggers. Under a strict reading those are major-class and currently ship under `refactor:` with no marker. Whether the deployed skill surface is a public contract is written down nowhere — worth deciding, but it outlives this PR. ## Deliberately not in scope #112 (cherry-pick ownership, now resolved in favour of `git-commits`), #113 (`rtk git` normalisation), #116 (research fan-out), #101 (audit-skill merge), #122 (non-spec skill-root files), #123 (no PRD producer) stay open. #117 is the one worth reading: the contract's remedy is to move prose into `references/`, which is exactly where neither the size gate nor Vale looks — and the blind spot is wider than #117 currently records, since there is no root `.vale.ini` at all, so every ADR, `CONTEXT.md` and `README.md` is unlinted too. That blind spot let this branch carry two `level: error` `Kyberforge.SentenceOpenerThereIs` violations into `references/` files it created — `provider-adapter-author/references/provider-matrix.md:31` and `agent-audit/references/finding-criteria.md:95`. Both are reworded in `afadaae`, confirmed by routing each file through the audit's own `vale-wrap.sh` (1 error each before, 0 after). Five further occurrences sit in `references/` files already on `main`; those are the pre-existing corpus and stay with #117, which is the real fix. Also unfixed and not this PR's: `apm install` appends a duplicate `SessionStart` entry to `.claude/settings.json`, so a fresh clone cannot get pre-push green without an edit AGENTS.md warns against. Reproduces identically on `main`. Co-authored-by: Defame1297 <gitea@rkdr.net> Reviewed-on: https://git.dev.rkdr.net/Defame1297/holocron/pulls/129 Co-authored-by: Claude Code AI - Gitea MCP <claude@noreply.git.dev.rkdr.net> Co-committed-by: Claude Code AI - Gitea MCP <claude@noreply.git.dev.rkdr.net>
9.2 KiB
source_keys
| source_keys | ||
|---|---|---|
|
The agent description and body contract
House contract, set by ADR-0020. The counts and the boundary targets are enforced by
agent-audit's scripts/validate.sh; the prose patterns by the Vale styles it bundles; the
judgment calls by its reference files.
Why the budget exists
An agent's name and description is loaded into every session's context at startup, whether or
not the agent is ever delegated to — the same cost a skill's description carries, so agents take
the same numbers. The body is different: it is not loaded into the caller's conversation at all,
it becomes the system prompt of a fresh context when the agent runs. That is why the body has no
word gate here and a skill body has one.
Description
A description carries exactly three things:
- Trigger clause — when to delegate, imperative: "Use when …", never "This agent …". Describe the user's intent and the triggering condition, not the agent's internal mechanics.
- At most one capability clause — what it does, one clause, no enumeration. Be specific ("reviews a diff for injected credentials", not "helps with security").
- Boundary clause — form:
Not <thing> -> <name>.Add one only where a near-miss agent or skill could steal delegations.
Banned from a description; move it to the body or to README.md:
- Capability enumeration or feature lists
- Per-scope emission mechanics — which files the author skill writes at which scope changes no delegation decision
- Output-format detail ("Produces a compact findings report with Why and Fix per finding")
- Composition or architecture notes ("composes X rather than duplicating Y", "cross-cutting")
- Implementation detail ("Self-validates via a bundled deterministic script")
- Restating the same trigger twice in two registers — a verb list, then the same verbs re-quoted as user phrasings. This is a FAIL, not a suggestion.
Do not open with an action verb. "Reviews…", "Analyzes…", "Generates…" was the old house rule
and ADR-0020 deleted it: the opener is Use when, matching every skill in this corpus, so one
router reads one shape.
"Use proactively" is Claude Code-only, and conditional even there. The phrase steers the Claude Code runtime to offer an agent unprompted and does nothing anywhere else, so where it may appear depends on the file:
| File | Rule |
|---|---|
Claude Code .md (project/user scope) |
Allowed. Add it only where the runtime should delegate without the user naming the agent — an agent invoked by name does not need it, and it costs activations elsewhere when added by reflex. |
Copilot .agent.md (project/user scope) |
Never. Inert there, and KyberforgeCopilot.ProactivePhrase grades it a hard FAIL. |
Vendor-neutral .apm/agents/<name>.agent.md (plugin/APM scope) |
Never. Same Vale rule, same hard FAIL — the file matches the **/*.agent.md glob, and it compiles to a real Copilot agent downstream. |
A pair whose Claude Code half carries the phrase and whose Copilot half omits it is correct, not
inconsistent: agent-audit checks that both halves describe the same job, not that they match
word for word.
Indirect triggers ("even if the user doesn't say X") take a similar conditional at every scope: add one only where the user's natural phrasing genuinely omits the domain word.
Boundary targets must resolve, and the notation decides how hard the gate bites. Route
notation — /name, or any arrow form (-> name, -> `name`) — is checked
unconditionally: an unresolved target there is a blocking ERROR. The prose form ("do not use
for X, use y instead") is only a SUGGESTION by default, because a bare hyphenated word in a
boundary clause is as likely to be a tool, a file format or an English compound as a route. It
is promoted to a blocking ERROR only when a second target in the same sentence does resolve,
which corroborates that the name was meant as a route. So a typo does not dangle equally
either way — write the arrow when you want the target checked. Targets resolve against a universe
built by walking up from the agent file itself: the nearest ancestor holding
plugins/*/.apm/{skills,agents} (or, failing that, the nearest ancestor holding .git) contributes
every skill and agent under <root>/plugins/*/, plus the agent's own apm package and the packages
that package declares in apm.yml under dependencies.apm. A sibling plugin in the same monorepo
therefore resolves; a skill in an unrelated repo does not. A target outside that universe sends the
router nowhere — a blocking failure in arrow or /name form, and in prose form only a SUGGESTION
nobody is forced to act on, which is the worse outcome because it ships. Verify it before writing
it — do not invent a plausible sibling.
Never let a hyphenated routing target wrap across lines in a folded > scalar. YAML folding
replaces the newline with a space, so gitea-labels- at the end of one line and milestones at
the start of the next fold into gitea-labels- milestones. The gate then reads the target as
gitea-labels, finds no such skill, and reports it dangling — nothing in the source lines looks
wrong. Reflow so the whole name sits on one line. The same applies to any backticked skill or
agent name anywhere in a description.
That universe is the apm marketplace and stops there. A host built-in is not a routing target:
/compact, /clear and /init are Claude Code slash commands with no counterpart in Copilot CLI
or Codex, and .apm/ source compiles for all three, so routing to one is a portability defect. The
gate is right to fail it and there is no allowlist. If a built-in genuinely needs mentioning, write
it un-slashed — the `compact` built-in — which makes no routing claim and is not checked.
Length. 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only with YAML folding resolved. Treat 250 as the target: the SUGGESTION tier is what moves the corpus average, the FAIL tier only stops outliers.
Body
Write the body as a direct role instruction, addressed to the agent:
You are a <role>. When invoked, <primary action>.
## Inputs
<what the agent is given: files, context, parameters>
## Process
<ordered steps; be explicit where ordering matters>
## Output
<what it produces: format, location, structure>
## Errors
<what to do on malformed, missing or contradictory input: report and stop, or
which fallback to take — and what to say to the caller either way>
Four required elements: inputs expected, process steps, output format, error handling. The last is the one that gets dropped, and dropping it is not neutral: an agent given a malformed input and no instruction invents a recovery, and a subagent's invented recovery is invisible to the caller until the output is wrong. Say explicitly whether the agent stops and reports, or degrades to a named fallback.
One job per agent. An agent covering two jobs gets delegated to for the wrong one.
Delegation discipline replaces the word gate. A plugin/APM agent is a single file with no
sibling references/ directory: it cannot disclose progressively to itself, so its only way to
stay short is to invoke rather than restate. A body that transcribes a procedure a skill it
can invoke already owns is an agent-audit FAIL, and the fix is one line — "invoke <skill>".
- Restating: "To commit, check the message against Conventional Commits: type, scope, description; header under 100 chars; …"
- Delegating: "Author commits with
git-commits."
The same holds for a procedure another agent owns. What belongs in the body is what no invocable skill covers: the agent's role, its boundaries, the order it works in, and the format it returns.
State a read-only boundary in prose, not only in frontmatter. disallowedTools denies the
tools it names and nothing else — never Bash, which an agent with no tools field inherits — so
an agent fenced only in frontmatter can still write through a shell redirect.
Invocation axis
Decide before writing the description whether the agent is model-delegated (the runtime picks it)
or reached only by name (@agent-<name>).
Only Copilot's cloud/IDE format expresses that in frontmatter: disable-model-invocation: true
requires explicit invocation, and user-invocable: false hides an agent from manual invocation.
Both live in .github/copilot/agents/<name>.md and are inert in the CLI format. Claude Code has
no equivalent field, and neither does the vendor-neutral plugin/APM file, so at those scopes a
name-invoked agent still needs a description precise enough not to steal delegations — the
boundary clause is doing that work.
One gate, two measurements
| Gate | SUGGESTION | FAIL | Counts |
|---|---|---|---|
| description | 250 chars | 400 chars | the description: value only |
| body (Copilot limit) | 30,000 chars | — | the body only; content past it is truncated silently |
The 30,000-character Copilot ceiling is a runtime truncation limit, not a quality target, and it applies to a plugin/APM file too — that file compiles into a real Copilot agent downstream. An agent body long enough to approach it has a delegation defect, not a length problem.