feat(kyberforge): ADR-0020 context contract for skills and agents #103

Merged
Defame1297 merged 18 commits from refactor/trim-skills-agents-context into main 2026-08-16 21:20:02 +00:00
26 changed files with 332 additions and 108 deletions
Showing only changes of commit 311e7cd22c - Show all commits

View File

@@ -76,6 +76,9 @@ Include what the fresh context lacks:
- A direct role instruction opening the prompt: `You are a [role]. When invoked, [action].`
- One bounded job, stated so the agent knows what it must refuse.
- The dispatch, gates, inputs and outputs listed above.
- **Error handling** — what the agent does on malformed, missing or contradictory input: stop and
report, or degrade to a named fallback. Absent it, the agent invents a recovery, and a
subagent's invented recovery is invisible to its caller until the output is wrong.
- Non-obvious environment facts and project-specific conventions it cannot infer.
- One default per decision point with one escape hatch.
@@ -111,6 +114,8 @@ Flag as FAIL if:
Flag as SUGGESTION if:
- The body does not open with a direct role instruction
- The body specifies no error handling — nothing tells the agent what to do with malformed,
missing or contradictory input
- The job the agent describes is unbounded, or bounded only implicitly
- A rationale is missing from a rule the agent is expected to enforce — present but unexplained
- Comments are useful but verbose enough to bury the field they annotate

View File

@@ -34,10 +34,13 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
boundary clause naming a real sibling skill or agent.
250 characters is the target, 400 the hard ceiling (ADR-0020).
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was
deleted. Add "Use proactively" only if the runtime should delegate here without
the user naming this agent.
deleted.
Never write "Use proactively" here. It steers the Claude Code runtime and does
nothing anywhere else, and this file compiles to a Copilot `.agent.md` too, where
agent-audit's KyberforgeCopilot.ProactivePhrase rule grades it a hard FAIL.
The phrase is CC-only; at this scope, a precise trigger clause does that job.
Example: "Use when a diff needs checking for injected credentials before it
merges. Not general code review -> code-reviewer." -->
merges. Not prose or style linting -> `lint-runner`." -->
<!-- model: sonnet
Optional. Aliases: sonnet, opus, haiku, fable. Or full model ID.
@@ -80,3 +83,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
## Output
FILL IN: What does the agent produce? Format, location, structure.
## Errors
FILL IN: What does the agent do on malformed, missing or contradictory input?
State whether it stops and reports, or degrades to a named fallback — and what it
tells the caller either way. An agent with no error handling invents a recovery,
and an invented recovery is invisible until the output is wrong.

View File

@@ -14,10 +14,14 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
boundary clause naming a real sibling skill or agent.
250 characters is the target, 400 the hard ceiling (ADR-0020).
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was
deleted. Add "Use proactively" only if the runtime should delegate here without
the user naming this agent.
deleted.
"Use proactively" is valid HERE and only here: it steers the Claude Code runtime
to offer this agent unprompted. Add it only if that is what you want. If you add
it, leave it OUT of the Copilot half of the pair — the phrase does nothing there
and agent-audit's KyberforgeCopilot.ProactivePhrase grades it a hard FAIL. The
pair must describe the same job; it does not have to be byte-identical.
Example: "Use when a diff needs checking for injected credentials before it
merges. Not general code review -> code-reviewer." -->
merges. Not prose or style linting -> `lint-runner`." -->
<!-- tools: Read, Bash, Grep
Optional. Allowlist of tool names: a comma-separated string or a YAML list.
@@ -99,3 +103,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
## Output
FILL IN: What does the agent produce? Format, location, structure.
## Errors
FILL IN: What does the agent do on malformed, missing or contradictory input?
State whether it stops and reports, or degrades to a named fallback — and what it
tells the caller either way. An agent with no error handling invents a recovery,
and an invented recovery is invisible until the output is wrong.

View File

@@ -19,9 +19,13 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
boundary clause naming a real sibling skill or agent.
250 characters is the target, 400 the hard ceiling (ADR-0020).
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was deleted.
Keep it identical in wording to the Claude Code half of the pair.
Never write "Use proactively" here. It steers the Claude Code runtime and does nothing
in Copilot, and agent-audit's KyberforgeCopilot.ProactivePhrase grades it a hard FAIL.
Otherwise keep the wording matched to the Claude Code half of the pair: agent-audit
checks that both halves describe the same job, not that they are byte-identical, so
dropping the CC-only phrase here is not a pair-consistency finding.
Example: "Use when a diff needs checking for injected credentials before it
merges. Not general code review -> code-reviewer." -->
merges. Not prose or style linting -> `lint-runner`." -->
<!-- tools: ["read", "search", "edit"]
Optional. Array of tool names. Omit = all available tools. [] = no tools.
@@ -68,3 +72,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
## Output
FILL IN: What does the agent produce? Format, location, structure.
## Errors
FILL IN: What does the agent do on malformed, missing or contradictory input?
State whether it stops and reports, or degrades to a named fallback — and what it
tells the caller either way. An agent with no error handling invents a recovery,
and an invented recovery is invisible until the output is wrong.

View File

@@ -44,15 +44,31 @@ Banned from a description; move it to the body or to `README.md`:
and ADR-0020 deleted it: the opener is `Use when`, matching every skill in this corpus, so one
router reads one shape.
**"Use proactively" is conditional.** Add it only where the runtime should delegate without the
user naming the agent — an agent invoked by name does not need it, and it costs activations
elsewhere when added by reflex. The same conditional governs indirect triggers ("even if the user
doesn't say X"): add one only where the user's natural phrasing genuinely omits the domain word.
**"Use proactively" is Claude Code-only, and conditional even there.** The phrase steers the
Claude Code runtime to offer an agent unprompted and does nothing anywhere else, so where it may
appear depends on the file:
**Boundary targets must resolve.** The name after the arrow is checked against real skills under
`plugins/*/.apm/skills/<name>/` and real agents under `plugins/*/.apm/agents/<name>.agent.md`. A
target that does not exist sends the router nowhere. Verify it before writing it — do not invent a
plausible sibling.
| File | Rule |
|---|---|
| Claude Code `.md` (project/user scope) | Allowed. Add it only where the runtime should delegate without the user naming the agent — an agent invoked by name does not need it, and it costs activations elsewhere when added by reflex. |
| Copilot `.agent.md` (project/user scope) | **Never.** Inert there, and `KyberforgeCopilot.ProactivePhrase` grades it a hard FAIL. |
| Vendor-neutral `.apm/agents/<name>.agent.md` (plugin/APM scope) | **Never.** Same Vale rule, same hard FAIL — the file matches the `**/*.agent.md` glob, and it compiles to a real Copilot agent downstream. |
A pair whose Claude Code half carries the phrase and whose Copilot half omits it is correct, not
inconsistent: `agent-audit` checks that both halves describe the same job, not that they match
word for word.
Indirect triggers ("even if the user doesn't say X") take a similar conditional at every scope:
add one only where the user's natural phrasing genuinely omits the domain word.
**Boundary targets must resolve.** Both forms are checked — the arrow and the prose form ("do not
use for X, use `y` instead") — so a typo dangles either way. Targets resolve against a universe
built by walking up **from the agent file itself**: the nearest ancestor holding
`plugins/*/.apm/{skills,agents}` (or, failing that, the nearest ancestor holding `.git`) contributes
every skill and agent under `<root>/plugins/*/`, plus the agent's own apm package and the packages
that package declares in `apm.yml` under `dependencies.apm`. A sibling plugin in the same monorepo
therefore resolves; a skill in an unrelated repo does not. A target outside that universe sends the
router nowhere. Verify it before writing it — do not invent a plausible sibling.
**Length.** 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only
with YAML folding resolved. Treat 250 as the target: the SUGGESTION tier is what moves the corpus
@@ -73,8 +89,18 @@ You are a <role>. When invoked, <primary action>.
## Output
<what it produces: format, location, structure>
## Errors
<what to do on malformed, missing or contradictory input: report and stop, or
which fallback to take — and what to say to the caller either way>
````
Four required elements: **inputs expected, process steps, output format, error handling.** The
last is the one that gets dropped, and dropping it is not neutral: an agent given a malformed
input and no instruction invents a recovery, and a subagent's invented recovery is invisible to
the caller until the output is wrong. Say explicitly whether the agent stops and reports, or
degrades to a named fallback.
One job per agent. An agent covering two jobs gets delegated to for the wrong one.
**Delegation discipline replaces the word gate.** A plugin/APM agent is a single file with no

View File

@@ -63,5 +63,7 @@ to installed skills instead of transcribed procedure.
- [ ] No `FILL IN:` placeholder and no `<!-- ... -->` template comment anywhere in the file
- [ ] System prompt body non-empty, and a read-only agent says so in prose as well as in
`disallowedTools`
- [ ] Body covers all four required elements: inputs expected, process steps, output format,
**error handling** — what the agent does on malformed, missing or contradictory input
Then return to the flow reference you came from.

View File

@@ -95,6 +95,8 @@ Both files:
- [ ] `name` present and kebab-case; `description` written to `references/contract.md`
- [ ] System prompt body present, non-empty and equivalent across the pair
- [ ] Body covers all four required elements: inputs expected, process steps, output format,
**error handling** — what the agent does on malformed, missing or contradictory input
- [ ] No `FILL IN:` placeholder and no `<!-- ... -->` template comment left
Copilot file only:

View File

@@ -285,6 +285,8 @@ else
if [[ "$SCOPE" == "plugin" ]]; then
echo " 1. Fill in $APM_FILE — replace every FILL IN: placeholder. Optional fields are" >&2
echo " scaffolded there as commented blocks; uncomment the ones that apply." >&2
echo " Description: 250 chars target / 400 ceiling (ADR-0020). The body has no" >&2
echo " word gate — delegate to a skill instead of restating what it does." >&2
echo " 2. Populate $SOURCES_DIR/sources.md with research sources, or delete it" >&2
echo " 3. Validate: $VALIDATE_HINT $APM_FILE" >&2
echo " It checks the frontmatter against the apm-agent-allowlist section of" >&2
@@ -292,6 +294,8 @@ else
else
echo " 1. Fill in $CC_FILE — replace every FILL IN: placeholder. Optional fields are" >&2
echo " scaffolded there as commented blocks; uncomment the ones that apply." >&2
echo " Description: 250 chars target / 400 ceiling (ADR-0020). The body has no" >&2
echo " word gate — delegate to a skill instead of restating what it does." >&2
echo " 2. Fill in $CP_FILE — same, and heed its closing comment: the Claude Code-only" >&2
echo " fields it names must not cross over from the file above." >&2
echo " 3. Validate: run $VALIDATE_HINT on each file" >&2

View File

@@ -69,8 +69,10 @@ table** plus the gates common to every branch, and each flow lives in its own se
`references/` file. Inlining all of them is a FAIL regardless of word count, because every
invocation then pays for every branch it did not take.
The reference shape in this repo is `apm-workflow`: a 554-word body dispatching to roughly 3,000
words of references across five mutually exclusive invocations.
The reference shape in this repo is `apm-workflow`: a **421-word body** dispatching to roughly
3,000 words of references across five mutually exclusive invocations. Its whole-file count is 554
words — cite 421 when calibrating a body, or the conflation this section warns against reappears
in the finding itself.
## Gotchas sections
@@ -86,17 +88,20 @@ sensibly.
Constraints:
- **Maximum five entries.** Past five, the section is a summary of the body rather than a set of
traps, and the agent stops reading it as a warning.
- **More than five entries is a SUGGESTION** — five is the guideline, not a ceiling. Past five, the
section is usually a summary of the body rather than a set of traps, and the agent stops reading
it as a warning. It stays advisory because whether a given gotcha earns its place is judgment;
`validate.sh` emits it through `suggest()` and the run still exits 0.
- **A Gotcha that paraphrases a step in the body below it is a FAIL.** It has no independent
content, and it teaches the agent that Gotchas can be skimmed because the real instruction is
coming.
coming. This one is the auditor's call — no script detects it.
- **A Gotchas section exceeding 25% of the body is a SUGGESTION** — the body has been inverted into
a preamble.
a preamble. Same tier and same reasoning as the entry count, and independent of it: either can
fire without the other.
- Place the section near the top. A gotcha read after the mistake is worthless, which is also why
Gotchas is the one construct exempt from moving to `references/`.
Worked negative example — `git-commits` carries thirteen entries, of which four restate content
Worked negative example — `git-commits` carries twelve entries, of which four restate content
that already appears below or in the description:
| Gotcha | Restates |
@@ -106,8 +111,10 @@ that already appears below or in the description:
| `:33` "Never skip hooks with `--no-verify`" | step 9 at `:52` |
| `:36` "Never commit secrets" | step 2 at `:45` |
All four are FAILs under this rule, and the section as a whole breaches the five-entry maximum. It
also passes every plausible word gate, which is the point of auditing the construct directly.
All four are FAILs under the paraphrase rule. The entry count and the section's share of the body
(387 of 1,102 words, 35%) are two further SUGGESTIONs on top — the script reports both, and neither
fails the run on its own. What makes this worth auditing directly is that the four paraphrase FAILs
pass every word gate there is; only reading the construct finds them.
## Calibrating control
@@ -144,7 +151,7 @@ Flag as FAIL if:
- A sentence answers "no" to the core test — it is padding
- The body exceeds 900 words counted body-only (`validate.sh` reports it)
- Two or more mutually exclusive flows are inlined instead of dispatched
- A Gotcha paraphrases a step in the body below it, or the section exceeds five entries
- A Gotcha paraphrases a step in the body below it
- A decision point presents a menu of options with no default
- An instruction repeats content already in the description
- A prescriptive sequence is used where flexibility is fine, or the reverse
@@ -152,6 +159,7 @@ Flag as FAIL if:
Flag as SUGGESTION if:
- The body exceeds 600 words counted body-only but stays at or under 900
- The Gotchas section carries more than five entries
- The Gotchas section exceeds 25% of the body
- A rationale is missing from an include/exclude rule — present but unexplained
- Gotchas are correct but placed late in the body rather than near the top

View File

@@ -19,10 +19,9 @@ metadata:
## Gotchas
- A skill's `name` and `description` are preloaded into every agent's context every session, invoked or not; the body loads only on invocation. The description is the scarce budget.
- The word gates are two different measurements, not one rule with two tiers. The 2,770-word / 500-line spec backstop counts the whole file including frontmatter; Step 3's gate counts the body alone. A file can sit well inside one and fail the other, so never unify them.
- Never spawn a subagent to audit or recheck your own work here. Run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent can have its worktree torn down by concurrent cleanup, destroying an uncommitted draft.
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires transcript analysis that is out of scope here — flag the opportunity as a suggestion instead.
- The word gates are two measurements, not two tiers of one rule: the 2,770-word / 500-line spec backstop counts the whole file, Step 3's gate the body alone. Never unify them.
- Never spawn a subagent to audit or recheck your own work — run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent's worktree can be torn down by concurrent cleanup, destroying an uncommitted draft.
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires out-of-scope transcript analysis — flag the opportunity as a suggestion instead.
## Step 1 — Dispatch
@@ -32,15 +31,15 @@ metadata:
| Directory exists, at least one improvement signal present | Improve | `references/improve.md` |
| Directory exists, no signals | Stop and ask | — |
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask: "No improvement signals found. Did you mean to create a new skill, or do you have feedback to apply?"
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask whether the user meant to create a new skill or has feedback to apply.
Read only the reference matching the resolved flow — each is self-contained. Capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
Read only the reference matching the resolved flow — each is self-contained. If the target sits inside a git worktree, capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
## Step 2 — Invocation axis
Decide before writing any description: model-invoked or hand-invoked?
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Worked example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`. Skip Step 3's description rules.
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Skip Step 3's description rules.
- **Model-invoked** — the default.
## Step 3 — Contract
@@ -51,12 +50,12 @@ Gates `/skill-audit` enforces in both flows:
- **Description** — a trigger clause, at most one capability clause, and a boundary clause shaped `Not <thing> -> <skill-name>` whose target resolves to a real skill or agent. 250 characters SUGGESTION, 400 FAIL, value only.
- **Body** — decision procedure only: ordered steps, branches, gates, and which reference to load when. 600 words SUGGESTION, 900 FAIL, body only. At two or more mutually exclusive flows a dispatch table is mandatory and each flow gets its own self-contained `references/` file.
- **Gotchas** — at most five, each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL.
- **Gotchas** — each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL; over five entries is a SUGGESTION only.
## Step 4 — Validate and close
Run `/skill-audit` on the resolved skill directory. It checks name-to-directory match, description presence, leftover `FILL IN:` placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those first. Resolve every FAIL before reporting done.
Run `/skill-audit` on the resolved skill directory; resolve every FAIL before reporting done. It checks name-to-directory match, placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those. Hand-check the one thing it misses: an empty body reports `PASS SKILL.md body word count 0 (ADR-0020 target: 600)`, so confirm at least one non-empty section exists.
With `metadata.version` present, bump the **minor** version on create (new skills start at `0.1.0`) and the **patch** version on improve.
**Commit verification.** Once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from the one captured at Step 1. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up first. Report done only once the hash has changed.
**Commit verification.** Inside a git worktree: once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from Step 1's. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up. Report done only once the hash has changed. Outside a worktree (a skill under `~/.claude/skills/`, say) nothing is committable — report done on a clean audit, naming that as the reason.

View File

@@ -10,14 +10,22 @@ name: SKILL_NAME
# Examples: my-tool, data-analyzer, pdf-processor
description: >
Use when FILL IN: trigger — when should an agent activate this skill?
Use when FILL IN: trigger.
FILL IN: at most ONE capability clause, stated specifically
(e.g. "parses and validates OpenAPI specs", not "helps with APIs").
Not FILL IN: near-miss case -> FILL IN: real sibling skill name.
Not FILL IN: near-miss case -> FILL IN: real sibling skill.
# Required. Preloaded into EVERY session whether or not the skill is invoked.
# Exactly three parts, in this order: trigger clause, at most one capability
# clause, boundary clause. Drop the boundary line if no near-miss skill exists.
# Budget: 250 characters target, 400 hard ceiling (counting this value only).
# Trigger clause: when should an agent activate this skill? Describe the user's
# intent, not the skill's internal mechanics.
# Budget: 250 characters target, 400 hard ceiling (counting this value only,
# with YAML folding resolved). This scaffold sits at 214 — keep the fill-in
# under the target rather than growing past it.
# Boundary clauses may be plural: write one per genuine near-miss, and none
# where no sibling could steal activations.
# Never let a hyphenated skill name wrap across two lines of this folded block
# — folding turns the break into a space and the routing target stops resolving.
# Banned here: capability lists, output-format detail, composition notes,
# implementation detail, and restating one trigger twice in two registers.
# The boundary target must resolve to a real skill or agent — it is checked.

View File

@@ -47,11 +47,15 @@ explicitly" only where the user's natural phrasing genuinely omits the domain wo
for `git-commits`, where the user says "commit". Adding one everywhere is what inflated this
corpus, and it was deleted as a blanket rule.
**Boundary targets must resolve.** The name after the arrow is checked against real skill
directories under `plugins/*/.apm/skills/<name>/` and real agents under
`plugins/*/.apm/agents/<name>.agent.md`. A boundary clause naming a target that does not exist
sends the router nowhere and fails the audit. Check the target exists before writing it — do not
invent a plausible sibling name.
**Boundary targets must resolve.** Both forms are checked — the arrow and the prose form ("do not
use for X, use `y` instead") — so a typo dangles either way. Targets resolve against a universe
built by walking up **from the SKILL.md itself**: the nearest ancestor holding
`plugins/*/.apm/{skills,agents}` (or, failing that, the nearest ancestor holding `.git`) contributes
every skill and agent under `<root>/plugins/*/`, plus the skill's own apm package and the packages
that package declares in `apm.yml` under `dependencies.apm`. A sibling plugin in the same monorepo
therefore resolves; a skill in an unrelated repo does not. A boundary clause naming a target
outside that universe sends the router nowhere and fails the audit. Check the target exists before
writing it — do not invent a plausible sibling name.
**Length.** 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only
with YAML folding resolved. The agentskills.io 1,024-character spec limit is unchanged and sits
@@ -61,7 +65,13 @@ as the outlier stop.
**Hand-invoked skills are exempt.** A skill carrying `disable-model-invocation: true` is absent
from the model-visible listing and is reached only by the user typing `/name`. It takes one plain
human-facing sentence — no trigger clause, no boundary clause, no indirect triggers. Worked
example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`.
example — the whole description of the `zoom-out` skill, which carries `disable-model-invocation`:
````markdown
Tell the agent to zoom out and give broader context or a higher-level perspective. Use when
you're unfamiliar with a section of code or need to understand how it fits into the bigger
picture.
````
## Body
@@ -90,6 +100,12 @@ blocks, rationale prose, and any content only one branch reaches. Each reference
self-contained for its concern, and every one is wired from the body with the literal conditional
form:
**The one exception, stated once so it is not re-litigated:** an output schema stays in the body
only when it applies to *every* flow and is short — roughly 50 words or less, which is the "Output
format template" pattern below. An output schema that is longer than that, or that only one flow
produces, moves to `references/` like any other schema. No third option exists, and the two rules
do not disagree.
````markdown
If <condition>, read `references/<file>.md`.
````
@@ -98,8 +114,9 @@ A generic pointer ("see references/ for details") is a Vale error — the agent
**Dispatch is mandatory at two or more mutually exclusive flows.** The body carries the dispatch
table and the gates common to every branch; each flow gets its own self-contained `references/`
file. Exemplar: `plugins/kyberforge/.apm/skills/apm-workflow/SKILL.md` — a 554-word body
dispatching to 3,006 words of references.
file. Exemplar: the `apm-workflow` skill — a **421-word body** dispatching to 3,006 words of
references. Calibrate against 421: that file's whole-file count is 554 words, and aiming at that
number instead overshoots the body budget by ~30%.
**Length.** 600 words SUGGESTION, 900 words FAIL, counting the **body only** — everything after
the frontmatter's closing `---`.
@@ -108,7 +125,7 @@ the frontmatter's closing `---`.
- Each entry must state a fact that **contradicts a reasonable default** — something the agent
gets wrong by acting sensibly. "Never commit secrets" is not one; the agent already knows.
- Maximum five entries.
- More than five entries is a SUGGESTION — five is the guideline, not a ceiling.
- A Gotcha that paraphrases a step in the body below it is a **FAIL**. If the rule is already a
step, it is not a gotcha.
- A Gotchas section exceeding 25% of the body is a SUGGESTION.
@@ -164,7 +181,9 @@ Do not modify flags.
| <condition> | <flow> | `references/<file>.md` |
````
**Output format template** (when the skill produces structured output):
**Output format template** (when the skill produces structured output on *every* flow, and the
schema is roughly 50 words or less — see the exception under Body above; anything longer or
flow-specific belongs in `references/`):
````markdown
Output format:
@@ -173,7 +192,8 @@ Output format:
```
````
For longer templates, place them in `assets/<name>.md` and reference conditionally.
For longer templates, place them in `references/<topic>.md` or `assets/<name>.md` and reference
conditionally.
## Embedding org-specific policy

View File

@@ -136,10 +136,16 @@ If no scripts are needed, delete `scripts/README.md` and the `scripts/` director
## Step 5 — Add references, assets, and tests (if needed)
**`references/`** — additional documentation loaded on demand. One topic per file. Reference
conditionally from SKILL.md with the literal form ``If <condition>, read `references/<file>.md` ``.
Keep reference chains one level deep — a reference file that references another reference file is
rarely loaded correctly.
**`references/`** — additional documentation loaded on demand. One topic per file, named in
kebab-case after the topic. Reference conditionally from SKILL.md with the literal form
``If <condition>, read `references/<file>.md` ``.
**Two hops from `SKILL.md`, never three.** A flow file may route on to a shared contract or
sub-topic file — that is the shipped pattern here (`SKILL.md` → `references/create.md` → this
file's own pointers to `contract.md`, `scripts.md` and `deployment-modes.md`). What does not work
is a third hop: a file reachable only through two intermediates is rarely loaded at the moment it
is needed. Every hop past the first also needs the same literal conditional form, so the agent
knows when to take it.
**`assets/`** — static resources: templates, schemas, lookup tables. Reference by relative path
from SKILL.md.

View File

@@ -76,6 +76,9 @@ Include what the fresh context lacks:
- A direct role instruction opening the prompt: `You are a [role]. When invoked, [action].`
- One bounded job, stated so the agent knows what it must refuse.
- The dispatch, gates, inputs and outputs listed above.
- **Error handling** — what the agent does on malformed, missing or contradictory input: stop and
report, or degrade to a named fallback. Absent it, the agent invents a recovery, and a
subagent's invented recovery is invisible to its caller until the output is wrong.
- Non-obvious environment facts and project-specific conventions it cannot infer.
- One default per decision point with one escape hatch.
@@ -111,6 +114,8 @@ Flag as FAIL if:
Flag as SUGGESTION if:
- The body does not open with a direct role instruction
- The body specifies no error handling — nothing tells the agent what to do with malformed,
missing or contradictory input
- The job the agent describes is unbounded, or bounded only implicitly
- A rationale is missing from a rule the agent is expected to enforce — present but unexplained
- Comments are useful but verbose enough to bury the field they annotate

View File

@@ -34,10 +34,13 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
boundary clause naming a real sibling skill or agent.
250 characters is the target, 400 the hard ceiling (ADR-0020).
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was
deleted. Add "Use proactively" only if the runtime should delegate here without
the user naming this agent.
deleted.
Never write "Use proactively" here. It steers the Claude Code runtime and does
nothing anywhere else, and this file compiles to a Copilot `.agent.md` too, where
agent-audit's KyberforgeCopilot.ProactivePhrase rule grades it a hard FAIL.
The phrase is CC-only; at this scope, a precise trigger clause does that job.
Example: "Use when a diff needs checking for injected credentials before it
merges. Not general code review -> code-reviewer." -->
merges. Not prose or style linting -> `lint-runner`." -->
<!-- model: sonnet
Optional. Aliases: sonnet, opus, haiku, fable. Or full model ID.
@@ -80,3 +83,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
## Output
FILL IN: What does the agent produce? Format, location, structure.
## Errors
FILL IN: What does the agent do on malformed, missing or contradictory input?
State whether it stops and reports, or degrades to a named fallback — and what it
tells the caller either way. An agent with no error handling invents a recovery,
and an invented recovery is invisible until the output is wrong.

View File

@@ -14,10 +14,14 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
boundary clause naming a real sibling skill or agent.
250 characters is the target, 400 the hard ceiling (ADR-0020).
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was
deleted. Add "Use proactively" only if the runtime should delegate here without
the user naming this agent.
deleted.
"Use proactively" is valid HERE and only here: it steers the Claude Code runtime
to offer this agent unprompted. Add it only if that is what you want. If you add
it, leave it OUT of the Copilot half of the pair — the phrase does nothing there
and agent-audit's KyberforgeCopilot.ProactivePhrase grades it a hard FAIL. The
pair must describe the same job; it does not have to be byte-identical.
Example: "Use when a diff needs checking for injected credentials before it
merges. Not general code review -> code-reviewer." -->
merges. Not prose or style linting -> `lint-runner`." -->
<!-- tools: Read, Bash, Grep
Optional. Allowlist of tool names: a comma-separated string or a YAML list.
@@ -99,3 +103,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
## Output
FILL IN: What does the agent produce? Format, location, structure.
## Errors
FILL IN: What does the agent do on malformed, missing or contradictory input?
State whether it stops and reports, or degrades to a named fallback — and what it
tells the caller either way. An agent with no error handling invents a recovery,
and an invented recovery is invisible until the output is wrong.

View File

@@ -19,9 +19,13 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
boundary clause naming a real sibling skill or agent.
250 characters is the target, 400 the hard ceiling (ADR-0020).
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was deleted.
Keep it identical in wording to the Claude Code half of the pair.
Never write "Use proactively" here. It steers the Claude Code runtime and does nothing
in Copilot, and agent-audit's KyberforgeCopilot.ProactivePhrase grades it a hard FAIL.
Otherwise keep the wording matched to the Claude Code half of the pair: agent-audit
checks that both halves describe the same job, not that they are byte-identical, so
dropping the CC-only phrase here is not a pair-consistency finding.
Example: "Use when a diff needs checking for injected credentials before it
merges. Not general code review -> code-reviewer." -->
merges. Not prose or style linting -> `lint-runner`." -->
<!-- tools: ["read", "search", "edit"]
Optional. Array of tool names. Omit = all available tools. [] = no tools.
@@ -68,3 +72,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
## Output
FILL IN: What does the agent produce? Format, location, structure.
## Errors
FILL IN: What does the agent do on malformed, missing or contradictory input?
State whether it stops and reports, or degrades to a named fallback — and what it
tells the caller either way. An agent with no error handling invents a recovery,
and an invented recovery is invisible until the output is wrong.

View File

@@ -44,15 +44,31 @@ Banned from a description; move it to the body or to `README.md`:
and ADR-0020 deleted it: the opener is `Use when`, matching every skill in this corpus, so one
router reads one shape.
**"Use proactively" is conditional.** Add it only where the runtime should delegate without the
user naming the agent — an agent invoked by name does not need it, and it costs activations
elsewhere when added by reflex. The same conditional governs indirect triggers ("even if the user
doesn't say X"): add one only where the user's natural phrasing genuinely omits the domain word.
**"Use proactively" is Claude Code-only, and conditional even there.** The phrase steers the
Claude Code runtime to offer an agent unprompted and does nothing anywhere else, so where it may
appear depends on the file:
**Boundary targets must resolve.** The name after the arrow is checked against real skills under
`plugins/*/.apm/skills/<name>/` and real agents under `plugins/*/.apm/agents/<name>.agent.md`. A
target that does not exist sends the router nowhere. Verify it before writing it — do not invent a
plausible sibling.
| File | Rule |
|---|---|
| Claude Code `.md` (project/user scope) | Allowed. Add it only where the runtime should delegate without the user naming the agent — an agent invoked by name does not need it, and it costs activations elsewhere when added by reflex. |
| Copilot `.agent.md` (project/user scope) | **Never.** Inert there, and `KyberforgeCopilot.ProactivePhrase` grades it a hard FAIL. |
| Vendor-neutral `.apm/agents/<name>.agent.md` (plugin/APM scope) | **Never.** Same Vale rule, same hard FAIL — the file matches the `**/*.agent.md` glob, and it compiles to a real Copilot agent downstream. |
A pair whose Claude Code half carries the phrase and whose Copilot half omits it is correct, not
inconsistent: `agent-audit` checks that both halves describe the same job, not that they match
word for word.
Indirect triggers ("even if the user doesn't say X") take a similar conditional at every scope:
add one only where the user's natural phrasing genuinely omits the domain word.
**Boundary targets must resolve.** Both forms are checked — the arrow and the prose form ("do not
use for X, use `y` instead") — so a typo dangles either way. Targets resolve against a universe
built by walking up **from the agent file itself**: the nearest ancestor holding
`plugins/*/.apm/{skills,agents}` (or, failing that, the nearest ancestor holding `.git`) contributes
every skill and agent under `<root>/plugins/*/`, plus the agent's own apm package and the packages
that package declares in `apm.yml` under `dependencies.apm`. A sibling plugin in the same monorepo
therefore resolves; a skill in an unrelated repo does not. A target outside that universe sends the
router nowhere. Verify it before writing it — do not invent a plausible sibling.
**Length.** 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only
with YAML folding resolved. Treat 250 as the target: the SUGGESTION tier is what moves the corpus
@@ -73,8 +89,18 @@ You are a <role>. When invoked, <primary action>.
## Output
<what it produces: format, location, structure>
## Errors
<what to do on malformed, missing or contradictory input: report and stop, or
which fallback to take — and what to say to the caller either way>
````
Four required elements: **inputs expected, process steps, output format, error handling.** The
last is the one that gets dropped, and dropping it is not neutral: an agent given a malformed
input and no instruction invents a recovery, and a subagent's invented recovery is invisible to
the caller until the output is wrong. Say explicitly whether the agent stops and reports, or
degrades to a named fallback.
One job per agent. An agent covering two jobs gets delegated to for the wrong one.
**Delegation discipline replaces the word gate.** A plugin/APM agent is a single file with no

View File

@@ -63,5 +63,7 @@ to installed skills instead of transcribed procedure.
- [ ] No `FILL IN:` placeholder and no `<!-- ... -->` template comment anywhere in the file
- [ ] System prompt body non-empty, and a read-only agent says so in prose as well as in
`disallowedTools`
- [ ] Body covers all four required elements: inputs expected, process steps, output format,
**error handling** — what the agent does on malformed, missing or contradictory input
Then return to the flow reference you came from.

View File

@@ -95,6 +95,8 @@ Both files:
- [ ] `name` present and kebab-case; `description` written to `references/contract.md`
- [ ] System prompt body present, non-empty and equivalent across the pair
- [ ] Body covers all four required elements: inputs expected, process steps, output format,
**error handling** — what the agent does on malformed, missing or contradictory input
- [ ] No `FILL IN:` placeholder and no `<!-- ... -->` template comment left
Copilot file only:

View File

@@ -285,6 +285,8 @@ else
if [[ "$SCOPE" == "plugin" ]]; then
echo " 1. Fill in $APM_FILE — replace every FILL IN: placeholder. Optional fields are" >&2
echo " scaffolded there as commented blocks; uncomment the ones that apply." >&2
echo " Description: 250 chars target / 400 ceiling (ADR-0020). The body has no" >&2
echo " word gate — delegate to a skill instead of restating what it does." >&2
echo " 2. Populate $SOURCES_DIR/sources.md with research sources, or delete it" >&2
echo " 3. Validate: $VALIDATE_HINT $APM_FILE" >&2
echo " It checks the frontmatter against the apm-agent-allowlist section of" >&2
@@ -292,6 +294,8 @@ else
else
echo " 1. Fill in $CC_FILE — replace every FILL IN: placeholder. Optional fields are" >&2
echo " scaffolded there as commented blocks; uncomment the ones that apply." >&2
echo " Description: 250 chars target / 400 ceiling (ADR-0020). The body has no" >&2
echo " word gate — delegate to a skill instead of restating what it does." >&2
echo " 2. Fill in $CP_FILE — same, and heed its closing comment: the Claude Code-only" >&2
echo " fields it names must not cross over from the file above." >&2
echo " 3. Validate: run $VALIDATE_HINT on each file" >&2

View File

@@ -69,8 +69,10 @@ table** plus the gates common to every branch, and each flow lives in its own se
`references/` file. Inlining all of them is a FAIL regardless of word count, because every
invocation then pays for every branch it did not take.
The reference shape in this repo is `apm-workflow`: a 554-word body dispatching to roughly 3,000
words of references across five mutually exclusive invocations.
The reference shape in this repo is `apm-workflow`: a **421-word body** dispatching to roughly
3,000 words of references across five mutually exclusive invocations. Its whole-file count is 554
words — cite 421 when calibrating a body, or the conflation this section warns against reappears
in the finding itself.
## Gotchas sections
@@ -86,17 +88,20 @@ sensibly.
Constraints:
- **Maximum five entries.** Past five, the section is a summary of the body rather than a set of
traps, and the agent stops reading it as a warning.
- **More than five entries is a SUGGESTION** — five is the guideline, not a ceiling. Past five, the
section is usually a summary of the body rather than a set of traps, and the agent stops reading
it as a warning. It stays advisory because whether a given gotcha earns its place is judgment;
`validate.sh` emits it through `suggest()` and the run still exits 0.
- **A Gotcha that paraphrases a step in the body below it is a FAIL.** It has no independent
content, and it teaches the agent that Gotchas can be skimmed because the real instruction is
coming.
coming. This one is the auditor's call — no script detects it.
- **A Gotchas section exceeding 25% of the body is a SUGGESTION** — the body has been inverted into
a preamble.
a preamble. Same tier and same reasoning as the entry count, and independent of it: either can
fire without the other.
- Place the section near the top. A gotcha read after the mistake is worthless, which is also why
Gotchas is the one construct exempt from moving to `references/`.
Worked negative example — `git-commits` carries thirteen entries, of which four restate content
Worked negative example — `git-commits` carries twelve entries, of which four restate content
that already appears below or in the description:
| Gotcha | Restates |
@@ -106,8 +111,10 @@ that already appears below or in the description:
| `:33` "Never skip hooks with `--no-verify`" | step 9 at `:52` |
| `:36` "Never commit secrets" | step 2 at `:45` |
All four are FAILs under this rule, and the section as a whole breaches the five-entry maximum. It
also passes every plausible word gate, which is the point of auditing the construct directly.
All four are FAILs under the paraphrase rule. The entry count and the section's share of the body
(387 of 1,102 words, 35%) are two further SUGGESTIONs on top — the script reports both, and neither
fails the run on its own. What makes this worth auditing directly is that the four paraphrase FAILs
pass every word gate there is; only reading the construct finds them.
## Calibrating control
@@ -144,7 +151,7 @@ Flag as FAIL if:
- A sentence answers "no" to the core test — it is padding
- The body exceeds 900 words counted body-only (`validate.sh` reports it)
- Two or more mutually exclusive flows are inlined instead of dispatched
- A Gotcha paraphrases a step in the body below it, or the section exceeds five entries
- A Gotcha paraphrases a step in the body below it
- A decision point presents a menu of options with no default
- An instruction repeats content already in the description
- A prescriptive sequence is used where flexibility is fine, or the reverse
@@ -152,6 +159,7 @@ Flag as FAIL if:
Flag as SUGGESTION if:
- The body exceeds 600 words counted body-only but stays at or under 900
- The Gotchas section carries more than five entries
- The Gotchas section exceeds 25% of the body
- A rationale is missing from an include/exclude rule — present but unexplained
- Gotchas are correct but placed late in the body rather than near the top

View File

@@ -19,10 +19,9 @@ metadata:
## Gotchas
- A skill's `name` and `description` are preloaded into every agent's context every session, invoked or not; the body loads only on invocation. The description is the scarce budget.
- The word gates are two different measurements, not one rule with two tiers. The 2,770-word / 500-line spec backstop counts the whole file including frontmatter; Step 3's gate counts the body alone. A file can sit well inside one and fail the other, so never unify them.
- Never spawn a subagent to audit or recheck your own work here. Run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent can have its worktree torn down by concurrent cleanup, destroying an uncommitted draft.
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires transcript analysis that is out of scope here — flag the opportunity as a suggestion instead.
- The word gates are two measurements, not two tiers of one rule: the 2,770-word / 500-line spec backstop counts the whole file, Step 3's gate the body alone. Never unify them.
- Never spawn a subagent to audit or recheck your own work — run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent's worktree can be torn down by concurrent cleanup, destroying an uncommitted draft.
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires out-of-scope transcript analysis — flag the opportunity as a suggestion instead.
## Step 1 — Dispatch
@@ -32,15 +31,15 @@ metadata:
| Directory exists, at least one improvement signal present | Improve | `references/improve.md` |
| Directory exists, no signals | Stop and ask | — |
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask: "No improvement signals found. Did you mean to create a new skill, or do you have feedback to apply?"
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask whether the user meant to create a new skill or has feedback to apply.
Read only the reference matching the resolved flow — each is self-contained. Capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
Read only the reference matching the resolved flow — each is self-contained. If the target sits inside a git worktree, capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
## Step 2 — Invocation axis
Decide before writing any description: model-invoked or hand-invoked?
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Worked example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`. Skip Step 3's description rules.
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Skip Step 3's description rules.
- **Model-invoked** — the default.
## Step 3 — Contract
@@ -51,12 +50,12 @@ Gates `/skill-audit` enforces in both flows:
- **Description** — a trigger clause, at most one capability clause, and a boundary clause shaped `Not <thing> -> <skill-name>` whose target resolves to a real skill or agent. 250 characters SUGGESTION, 400 FAIL, value only.
- **Body** — decision procedure only: ordered steps, branches, gates, and which reference to load when. 600 words SUGGESTION, 900 FAIL, body only. At two or more mutually exclusive flows a dispatch table is mandatory and each flow gets its own self-contained `references/` file.
- **Gotchas** — at most five, each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL.
- **Gotchas** — each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL; over five entries is a SUGGESTION only.
## Step 4 — Validate and close
Run `/skill-audit` on the resolved skill directory. It checks name-to-directory match, description presence, leftover `FILL IN:` placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those first. Resolve every FAIL before reporting done.
Run `/skill-audit` on the resolved skill directory; resolve every FAIL before reporting done. It checks name-to-directory match, placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those. Hand-check the one thing it misses: an empty body reports `PASS SKILL.md body word count 0 (ADR-0020 target: 600)`, so confirm at least one non-empty section exists.
With `metadata.version` present, bump the **minor** version on create (new skills start at `0.1.0`) and the **patch** version on improve.
**Commit verification.** Once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from the one captured at Step 1. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up first. Report done only once the hash has changed.
**Commit verification.** Inside a git worktree: once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from Step 1's. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up. Report done only once the hash has changed. Outside a worktree (a skill under `~/.claude/skills/`, say) nothing is committable — report done on a clean audit, naming that as the reason.

View File

@@ -10,14 +10,22 @@ name: SKILL_NAME
# Examples: my-tool, data-analyzer, pdf-processor
description: >
Use when FILL IN: trigger — when should an agent activate this skill?
Use when FILL IN: trigger.
FILL IN: at most ONE capability clause, stated specifically
(e.g. "parses and validates OpenAPI specs", not "helps with APIs").
Not FILL IN: near-miss case -> FILL IN: real sibling skill name.
Not FILL IN: near-miss case -> FILL IN: real sibling skill.
# Required. Preloaded into EVERY session whether or not the skill is invoked.
# Exactly three parts, in this order: trigger clause, at most one capability
# clause, boundary clause. Drop the boundary line if no near-miss skill exists.
# Budget: 250 characters target, 400 hard ceiling (counting this value only).
# Trigger clause: when should an agent activate this skill? Describe the user's
# intent, not the skill's internal mechanics.
# Budget: 250 characters target, 400 hard ceiling (counting this value only,
# with YAML folding resolved). This scaffold sits at 214 — keep the fill-in
# under the target rather than growing past it.
# Boundary clauses may be plural: write one per genuine near-miss, and none
# where no sibling could steal activations.
# Never let a hyphenated skill name wrap across two lines of this folded block
# — folding turns the break into a space and the routing target stops resolving.
# Banned here: capability lists, output-format detail, composition notes,
# implementation detail, and restating one trigger twice in two registers.
# The boundary target must resolve to a real skill or agent — it is checked.

View File

@@ -47,11 +47,15 @@ explicitly" only where the user's natural phrasing genuinely omits the domain wo
for `git-commits`, where the user says "commit". Adding one everywhere is what inflated this
corpus, and it was deleted as a blanket rule.
**Boundary targets must resolve.** The name after the arrow is checked against real skill
directories under `plugins/*/.apm/skills/<name>/` and real agents under
`plugins/*/.apm/agents/<name>.agent.md`. A boundary clause naming a target that does not exist
sends the router nowhere and fails the audit. Check the target exists before writing it — do not
invent a plausible sibling name.
**Boundary targets must resolve.** Both forms are checked — the arrow and the prose form ("do not
use for X, use `y` instead") — so a typo dangles either way. Targets resolve against a universe
built by walking up **from the SKILL.md itself**: the nearest ancestor holding
`plugins/*/.apm/{skills,agents}` (or, failing that, the nearest ancestor holding `.git`) contributes
every skill and agent under `<root>/plugins/*/`, plus the skill's own apm package and the packages
that package declares in `apm.yml` under `dependencies.apm`. A sibling plugin in the same monorepo
therefore resolves; a skill in an unrelated repo does not. A boundary clause naming a target
outside that universe sends the router nowhere and fails the audit. Check the target exists before
writing it — do not invent a plausible sibling name.
**Length.** 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only
with YAML folding resolved. The agentskills.io 1,024-character spec limit is unchanged and sits
@@ -61,7 +65,13 @@ as the outlier stop.
**Hand-invoked skills are exempt.** A skill carrying `disable-model-invocation: true` is absent
from the model-visible listing and is reached only by the user typing `/name`. It takes one plain
human-facing sentence — no trigger clause, no boundary clause, no indirect triggers. Worked
example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`.
example — the whole description of the `zoom-out` skill, which carries `disable-model-invocation`:
````markdown
Tell the agent to zoom out and give broader context or a higher-level perspective. Use when
you're unfamiliar with a section of code or need to understand how it fits into the bigger
picture.
````
## Body
@@ -90,6 +100,12 @@ blocks, rationale prose, and any content only one branch reaches. Each reference
self-contained for its concern, and every one is wired from the body with the literal conditional
form:
**The one exception, stated once so it is not re-litigated:** an output schema stays in the body
only when it applies to *every* flow and is short — roughly 50 words or less, which is the "Output
format template" pattern below. An output schema that is longer than that, or that only one flow
produces, moves to `references/` like any other schema. No third option exists, and the two rules
do not disagree.
````markdown
If <condition>, read `references/<file>.md`.
````
@@ -98,8 +114,9 @@ A generic pointer ("see references/ for details") is a Vale error — the agent
**Dispatch is mandatory at two or more mutually exclusive flows.** The body carries the dispatch
table and the gates common to every branch; each flow gets its own self-contained `references/`
file. Exemplar: `plugins/kyberforge/.apm/skills/apm-workflow/SKILL.md` — a 554-word body
dispatching to 3,006 words of references.
file. Exemplar: the `apm-workflow` skill — a **421-word body** dispatching to 3,006 words of
references. Calibrate against 421: that file's whole-file count is 554 words, and aiming at that
number instead overshoots the body budget by ~30%.
**Length.** 600 words SUGGESTION, 900 words FAIL, counting the **body only** — everything after
the frontmatter's closing `---`.
@@ -108,7 +125,7 @@ the frontmatter's closing `---`.
- Each entry must state a fact that **contradicts a reasonable default** — something the agent
gets wrong by acting sensibly. "Never commit secrets" is not one; the agent already knows.
- Maximum five entries.
- More than five entries is a SUGGESTION — five is the guideline, not a ceiling.
- A Gotcha that paraphrases a step in the body below it is a **FAIL**. If the rule is already a
step, it is not a gotcha.
- A Gotchas section exceeding 25% of the body is a SUGGESTION.
@@ -164,7 +181,9 @@ Do not modify flags.
| <condition> | <flow> | `references/<file>.md` |
````
**Output format template** (when the skill produces structured output):
**Output format template** (when the skill produces structured output on *every* flow, and the
schema is roughly 50 words or less — see the exception under Body above; anything longer or
flow-specific belongs in `references/`):
````markdown
Output format:
@@ -173,7 +192,8 @@ Output format:
```
````
For longer templates, place them in `assets/<name>.md` and reference conditionally.
For longer templates, place them in `references/<topic>.md` or `assets/<name>.md` and reference
conditionally.
## Embedding org-specific policy

View File

@@ -136,10 +136,16 @@ If no scripts are needed, delete `scripts/README.md` and the `scripts/` director
## Step 5 — Add references, assets, and tests (if needed)
**`references/`** — additional documentation loaded on demand. One topic per file. Reference
conditionally from SKILL.md with the literal form ``If <condition>, read `references/<file>.md` ``.
Keep reference chains one level deep — a reference file that references another reference file is
rarely loaded correctly.
**`references/`** — additional documentation loaded on demand. One topic per file, named in
kebab-case after the topic. Reference conditionally from SKILL.md with the literal form
``If <condition>, read `references/<file>.md` ``.
**Two hops from `SKILL.md`, never three.** A flow file may route on to a shared contract or
sub-topic file — that is the shipped pattern here (`SKILL.md` → `references/create.md` → this
file's own pointers to `contract.md`, `scripts.md` and `deployment-modes.md`). What does not work
is a third hop: a file reachable only through two intermediates is rarely loaded at the moment it
is needed. Every hop past the first also needs the same literal conditional form, so the agent
knows when to take it.
**`assets/`** — static resources: templates, schemas, lookup tables. Reference by relative path
from SKILL.md.