fix(kyberforge): reconcile the authoring rules the ADR-0020 trim left disagreeing

Six defects, each one a place where two files that an author reads in the same
sitting told them different things — or where the trim dropped a rule and nothing
noticed because no gate covers prose.

**"Use proactively" contradicted itself across the pair.** All three agent
templates said to add it where the runtime should delegate unprompted, while
`agent-audit`'s `KyberforgeCopilot.ProactivePhrase` rule grades it a hard FAIL in
any `*.agent.md` — which is the Copilot half of every project/user pair *and* the
vendor-neutral plugin-scope file, since that compiles to a real Copilot agent
downstream. Following the template produced a file the repo's own gate rejects.
The phrase is now permitted in exactly one place, the Claude Code `.md`, and
`references/contract.md` carries the per-file table plus the consequence authors
ask about next: a pair whose CC half has it and whose Copilot half does not is
correct, because `agent-audit` checks that both halves describe the same job, not
that they match word for word.

**The output-schema rule contradicted itself inside one file.** `contract.md`
said any content only one branch reaches moves to `references/`, and then offered
an "Output format template" body pattern with no qualification. Stated once now,
so it is not re-litigated: an output schema stays in the body only when every flow
produces it and it is roughly 50 words or less. No third option.

**Gotchas tiers disagreed with the script.** `validate.sh` emits the entry count
through `suggest()` and exits 0, while `skill-author` and `skill-audit` both
called more than five entries a FAIL. Whether a given gotcha earns its place is
judgment, so the prose moves to the script's tier rather than the reverse. The
paraphrase rule stays a FAIL and is explicitly marked as the auditor's call — no
script detects it.

**The dispatch exemplar was cited at the wrong number.** `apm-workflow`'s body is
421 words; 554 is its whole-file count. Both `contract.md` and `body-discipline.md`
cited 554 while describing a body budget, so an author calibrating against the
exemplar overshot by ~30% — the exact whole-file/body-only conflation those two
sections exist to warn against, reproduced inside the warning.

**"Error handling" came back as a required body element.** It was one of four and
is the one that gets dropped, and dropping it is not neutral: an agent handed
malformed input with no instruction invents a recovery, and a subagent's invented
recovery is invisible to its caller until the output is wrong. Restored in
`agent-audit`'s rubric as a SUGGESTION, in `agent-author`'s contract and both
scope checklists as a required element, and as an `## Errors` section in all three
templates.

**`skill-author` Step 4 gains the one check the audit misses.** An empty body
reports `PASS SKILL.md body word count 0` — a word gate cannot tell "concise"
from "absent". Step 4 now hand-checks for a non-empty section, and its commit
verification is conditioned on actually being inside a git worktree, which a skill
under `~/.claude/skills/` is not.

Also here: absolute repo paths removed from `skill-author`'s SKILL.md and
contract.md in favour of naming the skill (`zoom-out`'s description is quoted
inline instead of pointed at), the boundary-target universe documented to match
the resolver, a two-hops-from-SKILL.md limit on reference chains, and
`new-agent.sh`'s next-steps output naming the description budget and the
deliberate absence of an agent body gate.

Refs: ADR-0020
This commit is contained in:
2026-08-16 16:40:51 +00:00
parent 2540e50fcc
commit 311e7cd22c
26 changed files with 332 additions and 108 deletions

View File

@@ -19,10 +19,9 @@ metadata:
## Gotchas
- A skill's `name` and `description` are preloaded into every agent's context every session, invoked or not; the body loads only on invocation. The description is the scarce budget.
- The word gates are two different measurements, not one rule with two tiers. The 2,770-word / 500-line spec backstop counts the whole file including frontmatter; Step 3's gate counts the body alone. A file can sit well inside one and fail the other, so never unify them.
- Never spawn a subagent to audit or recheck your own work here. Run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent can have its worktree torn down by concurrent cleanup, destroying an uncommitted draft.
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires transcript analysis that is out of scope here — flag the opportunity as a suggestion instead.
- The word gates are two measurements, not two tiers of one rule: the 2,770-word / 500-line spec backstop counts the whole file, Step 3's gate the body alone. Never unify them.
- Never spawn a subagent to audit or recheck your own work — run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent's worktree can be torn down by concurrent cleanup, destroying an uncommitted draft.
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires out-of-scope transcript analysis — flag the opportunity as a suggestion instead.
## Step 1 — Dispatch
@@ -32,15 +31,15 @@ metadata:
| Directory exists, at least one improvement signal present | Improve | `references/improve.md` |
| Directory exists, no signals | Stop and ask | — |
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask: "No improvement signals found. Did you mean to create a new skill, or do you have feedback to apply?"
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask whether the user meant to create a new skill or has feedback to apply.
Read only the reference matching the resolved flow — each is self-contained. Capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
Read only the reference matching the resolved flow — each is self-contained. If the target sits inside a git worktree, capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
## Step 2 — Invocation axis
Decide before writing any description: model-invoked or hand-invoked?
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Worked example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`. Skip Step 3's description rules.
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Skip Step 3's description rules.
- **Model-invoked** — the default.
## Step 3 — Contract
@@ -51,12 +50,12 @@ Gates `/skill-audit` enforces in both flows:
- **Description** — a trigger clause, at most one capability clause, and a boundary clause shaped `Not <thing> -> <skill-name>` whose target resolves to a real skill or agent. 250 characters SUGGESTION, 400 FAIL, value only.
- **Body** — decision procedure only: ordered steps, branches, gates, and which reference to load when. 600 words SUGGESTION, 900 FAIL, body only. At two or more mutually exclusive flows a dispatch table is mandatory and each flow gets its own self-contained `references/` file.
- **Gotchas** — at most five, each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL.
- **Gotchas** — each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL; over five entries is a SUGGESTION only.
## Step 4 — Validate and close
Run `/skill-audit` on the resolved skill directory. It checks name-to-directory match, description presence, leftover `FILL IN:` placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those first. Resolve every FAIL before reporting done.
Run `/skill-audit` on the resolved skill directory; resolve every FAIL before reporting done. It checks name-to-directory match, placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those. Hand-check the one thing it misses: an empty body reports `PASS SKILL.md body word count 0 (ADR-0020 target: 600)`, so confirm at least one non-empty section exists.
With `metadata.version` present, bump the **minor** version on create (new skills start at `0.1.0`) and the **patch** version on improve.
**Commit verification.** Once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from the one captured at Step 1. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up first. Report done only once the hash has changed.
**Commit verification.** Inside a git worktree: once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from Step 1's. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up. Report done only once the hash has changed. Outside a worktree (a skill under `~/.claude/skills/`, say) nothing is committable — report done on a clean audit, naming that as the reason.

View File

@@ -10,14 +10,22 @@ name: SKILL_NAME
# Examples: my-tool, data-analyzer, pdf-processor
description: >
Use when FILL IN: trigger — when should an agent activate this skill?
Use when FILL IN: trigger.
FILL IN: at most ONE capability clause, stated specifically
(e.g. "parses and validates OpenAPI specs", not "helps with APIs").
Not FILL IN: near-miss case -> FILL IN: real sibling skill name.
Not FILL IN: near-miss case -> FILL IN: real sibling skill.
# Required. Preloaded into EVERY session whether or not the skill is invoked.
# Exactly three parts, in this order: trigger clause, at most one capability
# clause, boundary clause. Drop the boundary line if no near-miss skill exists.
# Budget: 250 characters target, 400 hard ceiling (counting this value only).
# Trigger clause: when should an agent activate this skill? Describe the user's
# intent, not the skill's internal mechanics.
# Budget: 250 characters target, 400 hard ceiling (counting this value only,
# with YAML folding resolved). This scaffold sits at 214 — keep the fill-in
# under the target rather than growing past it.
# Boundary clauses may be plural: write one per genuine near-miss, and none
# where no sibling could steal activations.
# Never let a hyphenated skill name wrap across two lines of this folded block
# — folding turns the break into a space and the routing target stops resolving.
# Banned here: capability lists, output-format detail, composition notes,
# implementation detail, and restating one trigger twice in two registers.
# The boundary target must resolve to a real skill or agent — it is checked.

View File

@@ -47,11 +47,15 @@ explicitly" only where the user's natural phrasing genuinely omits the domain wo
for `git-commits`, where the user says "commit". Adding one everywhere is what inflated this
corpus, and it was deleted as a blanket rule.
**Boundary targets must resolve.** The name after the arrow is checked against real skill
directories under `plugins/*/.apm/skills/<name>/` and real agents under
`plugins/*/.apm/agents/<name>.agent.md`. A boundary clause naming a target that does not exist
sends the router nowhere and fails the audit. Check the target exists before writing it — do not
invent a plausible sibling name.
**Boundary targets must resolve.** Both forms are checked — the arrow and the prose form ("do not
use for X, use `y` instead") — so a typo dangles either way. Targets resolve against a universe
built by walking up **from the SKILL.md itself**: the nearest ancestor holding
`plugins/*/.apm/{skills,agents}` (or, failing that, the nearest ancestor holding `.git`) contributes
every skill and agent under `<root>/plugins/*/`, plus the skill's own apm package and the packages
that package declares in `apm.yml` under `dependencies.apm`. A sibling plugin in the same monorepo
therefore resolves; a skill in an unrelated repo does not. A boundary clause naming a target
outside that universe sends the router nowhere and fails the audit. Check the target exists before
writing it — do not invent a plausible sibling name.
**Length.** 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only
with YAML folding resolved. The agentskills.io 1,024-character spec limit is unchanged and sits
@@ -61,7 +65,13 @@ as the outlier stop.
**Hand-invoked skills are exempt.** A skill carrying `disable-model-invocation: true` is absent
from the model-visible listing and is reached only by the user typing `/name`. It takes one plain
human-facing sentence — no trigger clause, no boundary clause, no indirect triggers. Worked
example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`.
example — the whole description of the `zoom-out` skill, which carries `disable-model-invocation`:
````markdown
Tell the agent to zoom out and give broader context or a higher-level perspective. Use when
you're unfamiliar with a section of code or need to understand how it fits into the bigger
picture.
````
## Body
@@ -90,6 +100,12 @@ blocks, rationale prose, and any content only one branch reaches. Each reference
self-contained for its concern, and every one is wired from the body with the literal conditional
form:
**The one exception, stated once so it is not re-litigated:** an output schema stays in the body
only when it applies to *every* flow and is short — roughly 50 words or less, which is the "Output
format template" pattern below. An output schema that is longer than that, or that only one flow
produces, moves to `references/` like any other schema. No third option exists, and the two rules
do not disagree.
````markdown
If <condition>, read `references/<file>.md`.
````
@@ -98,8 +114,9 @@ A generic pointer ("see references/ for details") is a Vale error — the agent
**Dispatch is mandatory at two or more mutually exclusive flows.** The body carries the dispatch
table and the gates common to every branch; each flow gets its own self-contained `references/`
file. Exemplar: `plugins/kyberforge/.apm/skills/apm-workflow/SKILL.md` — a 554-word body
dispatching to 3,006 words of references.
file. Exemplar: the `apm-workflow` skill — a **421-word body** dispatching to 3,006 words of
references. Calibrate against 421: that file's whole-file count is 554 words, and aiming at that
number instead overshoots the body budget by ~30%.
**Length.** 600 words SUGGESTION, 900 words FAIL, counting the **body only** — everything after
the frontmatter's closing `---`.
@@ -108,7 +125,7 @@ the frontmatter's closing `---`.
- Each entry must state a fact that **contradicts a reasonable default** — something the agent
gets wrong by acting sensibly. "Never commit secrets" is not one; the agent already knows.
- Maximum five entries.
- More than five entries is a SUGGESTION — five is the guideline, not a ceiling.
- A Gotcha that paraphrases a step in the body below it is a **FAIL**. If the rule is already a
step, it is not a gotcha.
- A Gotchas section exceeding 25% of the body is a SUGGESTION.
@@ -164,7 +181,9 @@ Do not modify flags.
| <condition> | <flow> | `references/<file>.md` |
````
**Output format template** (when the skill produces structured output):
**Output format template** (when the skill produces structured output on *every* flow, and the
schema is roughly 50 words or less — see the exception under Body above; anything longer or
flow-specific belongs in `references/`):
````markdown
Output format:
@@ -173,7 +192,8 @@ Output format:
```
````
For longer templates, place them in `assets/<name>.md` and reference conditionally.
For longer templates, place them in `references/<topic>.md` or `assets/<name>.md` and reference
conditionally.
## Embedding org-specific policy

View File

@@ -136,10 +136,16 @@ If no scripts are needed, delete `scripts/README.md` and the `scripts/` director
## Step 5 — Add references, assets, and tests (if needed)
**`references/`** — additional documentation loaded on demand. One topic per file. Reference
conditionally from SKILL.md with the literal form ``If <condition>, read `references/<file>.md` ``.
Keep reference chains one level deep — a reference file that references another reference file is
rarely loaded correctly.
**`references/`** — additional documentation loaded on demand. One topic per file, named in
kebab-case after the topic. Reference conditionally from SKILL.md with the literal form
``If <condition>, read `references/<file>.md` ``.
**Two hops from `SKILL.md`, never three.** A flow file may route on to a shared contract or
sub-topic file — that is the shipped pattern here (`SKILL.md` → `references/create.md` → this
file's own pointers to `contract.md`, `scripts.md` and `deployment-modes.md`). What does not work
is a third hop: a file reachable only through two intermediates is rarely loaded at the moment it
is needed. Every hop past the first also needs the same literal conditional form, so the agent
knows when to take it.
**`assets/`** — static resources: templates, schemas, lookup tables. Reference by relative path
from SKILL.md.