Files
holocron/plugins/kyberforge/.apm/skills/agent-author/references/contract.md
Defame1297 620f20b0fd refactor(kyberforge)!: merge skill-audit and agent-audit into factory-audit
Why

The two audit skills carried 1,724 lines of byte-identical duplication: the ADR-0020 boundary
resolver (1,061), vale-wrap.sh (526), the Vale style rules (44) and the Contributing-files parser
(93). Nothing shared them — they were held in sync by a 413-line pre-push gate and its 797-line
test suite. Sync-by-gate had already failed once: at 484357a the two parser copies drifted into
different spellings of the bullet loop while a docstring asserted they were identical. That drift
was behaviour-neutral and was re-unified by hand at 598a7c3, so the copies were identical at merge
time — but nothing had caught it, and the next drift need not be neutral.

Implementation Notes

Self-containment binds BETWEEN skills, not within one. The agentskills.io spec forbids reaching
across skill directories, which is why two separate skills needed embedded copies; two files inside
ONE skill may source a third. That is the whole reason the merge removes duplication rather than
relocating it.

The union of both bodies measured 1,532 words against BODY_MAX_WORDS=900, and only 211 of those
words were shared, so SKILL.md is a dispatch body. Step 0 resolves the flow from the target path
before any validation, and its table mirrors validate.sh's detection exactly: a directory holding
SKILL.md or a SKILL.md file (skill); a *.agent.md, or a .md directly under an agents/ directory
(agent); anything else stops without running a validator. Steps 1-3 live in
references/skill-flow.md and references/agent-flow.md, and gotchas that apply to one flow live in
that flow's file, since it is loaded on every invocation anyway. If validate.sh reports on the
other artifact type, the body restarts at Step 0.

Named factory-audit rather than forge-audit because forge is a live skill, and a family prefix that
matches a live sibling reads as ownership rather than membership.

The description carries one arrow per boundary target, because ADR-0020 resolves only the first
target after an arrow. It drops the quoted "audit this skill"-style phrases, which restated
"audited" in a second register (ADR-0020's duplicate-register rule). 241 characters, Gotchas 16%
of the body: no size SUGGESTIONs.

The boundary resolver stays embedded in two files rather than imported: a cache-installed plugin
cannot read outside its own directory, and the repo-root hook resolves via .pre-commit-hooks.yaml
where entry[0] is the only token pre-commit rewrites, so no single file is reachable by both.
tests/test-adr0020-contract.sh hashes both copies for byte-identity, and asserts validate.sh sources
the resolver and that no third copy exists.

The entry scripts classify the target from its resolved parent directory, so a bare agent filename
typed inside agents/ works; resolve SCRIPT_DIR CDPATH-safely; and exit 2 when a lib-*.sh is
missing, rather than dying with exit 1, the tier the flows relay as real findings.

The provenance run functions stash their findings code in KYBERFORGE_PROV_RC and
return 0, so validate-provenance.sh calls them UNTESTED. Testing a function's
status (`f || RC=$?`) disables errexit for its entire body, and no subshell or
`set -e` inside can re-arm it once the call sits in a condition context
(measured, both spellings). Their error paths use `exit`, which is unaffected
either way; this keeps errexit armed for anything added later.

Case 0's readability guard reads the file instead of asking `[[ -r ]]`. `-r` is
access(2), which answers yes for uid 0 even on a mode-000 file, and this repo's
dev environment is root -- so the guard could never fire where it exists to fire.
A read attempt is also the stricter question, catching EIO. This is the reasoning
scripts/check-vale-style-sync.sh carried before this commit deleted it; the
hazard did not go with it.

All three entry scripts are CDPATH-safe, vale-wrap.sh included: both of its cd sites are cleared,
the --config resolution and the directory-mirror walk, where an exported CDPATH would otherwise
print a decoy path into the -print0 stream and build the mirror from the decoy's files. The two
remaining bare cd calls take absolute paths, which CDPATH is never consulted for.

Impact

BREAKING: skill-audit and agent-audit no longer exist as invocable skills. kyberforge goes to
2.0.0 (catalog 0.4.7).

Check logic is unchanged: differential runs of the old and new validators across every skill and
agent produced byte-identical stdout, stderr and exit codes, and the reconstructed Python payloads
differ only in comments and the references/field-inventory.md -> agent-field-inventory.md rename.
One doctrine governs the tiers: exit 0 is audited and clean, exit 1 is audited with findings OR a
target present but unreadable, exit 2 is that nothing was audited at all. Edge paths DID change,
deliberately (full table in ADR-0025):
- a missing target exits 2 (never ran), not 1, under its own "does not exist" message; detection is
  by path shape, so a shape-matching path that is simply absent used to reach the validator and come
  back as a FAIL against a file that never existed;
- an unshaped target exits 2 under the generic "matches neither" message, and a directory with no
  SKILL.md under a third, distinct one -- three exit-2 messages, not one;
- a dangling symlink or a symlink loop stays exit 1: it is present but broken, which is a finding
  about the artifact rather than a usage error;
- a SKILL.md file path is audited as its skill directory instead of refused;
- a .md agent outside an agents/ directory is refused rather than audited;
- a missing script library, a missing python3, a missing PyYAML, and no argument at all each exit 2.
  validate-provenance.sh already exited 2 for the last two; validate.sh now matches it.

.pre-commit-hooks.yaml is a published contract consumed by external repos. Both hook IDs and both
files: regexes are unchanged; only entry: and description: moved.

scripts/check-vale-style-sync.sh (413), scripts/sync-vale-styles.sh (21),
tests/test-check-vale-style-sync.sh (797) and agent-audit/scripts/README.md (47) are deleted. The
checker made 17 assertions: 6 compared the two Vale copies and are moot; 10 are rehomed into
tests/test-vale-wrap.sh (case 0, cases 28-31, and the suite's Vale-absent skip); and the
cross-manifest files: agreement check, which selected hooks by entry: and so could not survive both
hooks sharing one, is ported as case 33 pairing hooks by id:. Cases 28, 30 and 33 carry mutation
self-tests; narrowing the local skill prefilter to 6 of 38 SKILL.md files now fails the suite.

Skills go 39 to 38. Pre-push goes 9 repo-authored hooks to 8.

ADR: 0025
BREAKING-CHANGE: the skill-audit and agent-audit skills are removed. Both flows are served by
  factory-audit, which auto-detects whether it was handed a skill directory or an agent file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YR2CjVumUbEGWcMikcoXBD
2026-09-16 09:13:57 +00:00

9.1 KiB

source_keys
source_keys
claude-code-subagents-docs
github-custom-agents-configuration

The agent description and body contract

House contract. The counts and the boundary targets are enforced by factory-audit's scripts/validate.sh; the prose patterns by the Vale styles it bundles; the judgment calls by its reference files.

Why the budget exists

An agent's name and description is loaded into every session's context at startup, whether or not the agent is ever delegated to — the same cost a skill's description carries, so agents take the same numbers. The body is different: it is not loaded into the caller's conversation at all, it becomes the system prompt of a fresh context when the agent runs. That is why the body has no word gate here and a skill body has one.

Description

A description carries exactly three things:

  1. Trigger clause — when to delegate, imperative: "Use when …", never "This agent …". Describe the user's intent and the triggering condition, not the agent's internal mechanics.
  2. At most one capability clause — what it does, one clause, no enumeration. Be specific ("reviews a diff for injected credentials", not "helps with security").
  3. Boundary clause — form: Not <thing> -> <name>. Add one only where a near-miss agent or skill could steal delegations.

Banned from a description; move it to the body or to README.md:

  • Capability enumeration or feature lists
  • Per-scope emission mechanics — which files the author skill writes at which scope changes no delegation decision
  • Output-format detail ("Produces a compact findings report with Why and Fix per finding")
  • Composition or architecture notes ("composes X rather than duplicating Y", "cross-cutting")
  • Implementation detail ("Self-validates via a bundled deterministic script")
  • Restating the same trigger twice in two registers — a verb list, then the same verbs re-quoted as user phrasings. This is a FAIL, not a suggestion.

Do not open with an action verb. The opener is Use when, matching every skill in this corpus, so one router reads one shape.

"Use proactively" is Claude Code-only, and conditional even there. The phrase steers the Claude Code runtime to offer an agent unprompted and does nothing anywhere else, so where it may appear depends on the file:

File Rule
Claude Code .md (project/user scope) Allowed. Add it only where the runtime should delegate without the user naming the agent — an agent invoked by name does not need it, and it costs activations elsewhere when added by reflex.
Copilot .agent.md (project/user scope) Never. Inert there, and KyberforgeCopilot.ProactivePhrase grades it a hard FAIL.
Vendor-neutral .apm/agents/<name>.agent.md (plugin/APM scope) Never. Same Vale rule, same hard FAIL — the file matches the **/*.agent.md glob, and it compiles to a real Copilot agent downstream.

A pair whose Claude Code half carries the phrase and whose Copilot half omits it is correct, not inconsistent: factory-audit checks that both halves describe the same job, not that they match word for word.

Indirect triggers ("even if the user doesn't say X") take a similar conditional at every scope: add one only where the user's natural phrasing genuinely omits the domain word.

Boundary targets must resolve, and the notation decides how hard the gate bites. Route notation — /name, or any arrow form (-> name, -> `name`) — is checked unconditionally: an unresolved target there is a blocking ERROR. The prose form ("do not use for X, use y instead") is only a SUGGESTION by default, because a bare hyphenated word in a boundary clause is as likely to be a tool, a file format or an English compound as a route. It is promoted to a blocking ERROR only when a second target in the same sentence does resolve, which corroborates that the name was meant as a route. So a typo does not dangle equally either way — write the arrow when you want the target checked. Targets resolve against a universe built by walking up from the agent file itself: the nearest ancestor holding plugins/*/.apm/{skills,agents} (or, failing that, the nearest ancestor holding .git) contributes every skill and agent under <root>/plugins/*/, plus the agent's own apm package and the packages that package declares in apm.yml under dependencies.apm. A sibling plugin in the same monorepo therefore resolves; a skill in an unrelated repo does not. A target outside that universe sends the router nowhere — a blocking failure in arrow or /name form, and in prose form only a SUGGESTION nobody is forced to act on, which is the worse outcome because it ships. Verify it before writing it — do not invent a plausible sibling.

Never let a hyphenated routing target wrap across lines in a folded > scalar. YAML folding replaces the newline with a space, so gitea-labels- at the end of one line and milestones at the start of the next fold into gitea-labels- milestones. The gate then reads the target as gitea-labels, finds no such skill, and reports it dangling — nothing in the source lines looks wrong. Reflow so the whole name sits on one line. The same applies to any backticked skill or agent name anywhere in a description.

That universe is the apm marketplace and stops there. A host built-in is not a routing target: /compact, /clear and /init are Claude Code slash commands with no counterpart in Copilot CLI or Codex, and .apm/ source compiles for all three, so routing to one is a portability defect. The gate is right to fail it and there is no allowlist. If a built-in genuinely needs mentioning, write it un-slashed — the `compact` built-in — which makes no routing claim and is not checked.

Length. 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only with YAML folding resolved. Treat 250 as the target: the SUGGESTION tier is what moves the corpus average, the FAIL tier only stops outliers.

Body

Write the body as a direct role instruction, addressed to the agent:

You are a <role>. When invoked, <primary action>.

## Inputs
<what the agent is given: files, context, parameters>

## Process
<ordered steps; be explicit where ordering matters>

## Output
<what it produces: format, location, structure>

## Errors
<what to do on malformed, missing or contradictory input: report and stop, or
 which fallback to take — and what to say to the caller either way>

Four required elements: inputs expected, process steps, output format, error handling. The last is the one that gets dropped, and dropping it is not neutral: an agent given a malformed input and no instruction invents a recovery, and a subagent's invented recovery is invisible to the caller until the output is wrong. Say explicitly whether the agent stops and reports, or degrades to a named fallback.

One job per agent. An agent covering two jobs gets delegated to for the wrong one.

Delegation discipline replaces the word gate. A plugin/APM agent is a single file with no sibling references/ directory: it cannot disclose progressively to itself, so its only way to stay short is to invoke rather than restate. A body that transcribes a procedure a skill it can invoke already owns is a factory-audit FAIL, and the fix is one line — "invoke <skill>".

  • Restating: "To commit, check the message against Conventional Commits: type, scope, description; header under 100 chars; …"
  • Delegating: "Author commits with git-commits."

The same holds for a procedure another agent owns. What belongs in the body is what no invocable skill covers: the agent's role, its boundaries, the order it works in, and the format it returns.

State a read-only boundary in prose, not only in frontmatter. disallowedTools denies the tools it names and nothing else — never Bash, which an agent with no tools field inherits — so an agent fenced only in frontmatter can still write through a shell redirect.

Invocation axis

Decide before writing the description whether the agent is model-delegated (the runtime picks it) or reached only by name (@agent-<name>).

Only Copilot's cloud/IDE format expresses that in frontmatter: disable-model-invocation: true requires explicit invocation, and user-invocable: false hides an agent from manual invocation. Both live in .github/copilot/agents/<name>.md and are inert in the CLI format. Claude Code has no equivalent field, and neither does the vendor-neutral plugin/APM file, so at those scopes a name-invoked agent still needs a description precise enough not to steal delegations — the boundary clause is doing that work.

One gate, two measurements

Gate SUGGESTION FAIL Counts
description 250 chars 400 chars the description: value only
body (Copilot limit) 30,000 chars — the body only; content past it is truncated silently

The 30,000-character Copilot ceiling is a runtime truncation limit, not a quality target, and it applies to a plugin/APM file too — that file compiles into a real Copilot agent downstream. An agent body long enough to approach it has a delegation defect, not a length problem.