feat(kyberforge): enforce the ADR-0020 context contract for skills and agents

Skill name+description pairs are preloaded into every session, costing
~6,200 tokens across 39 skills before any skill is invoked. The authoring
rules mandated that growth: skill-author:104 and description-quality.md:21
both required padding, while skill-author:102 (the deflating rule) had no
FAIL condition behind it.

Gates (blocking, no baseline file):
- description 250 chars SUGGESTION / 400 FAIL, measured on the folded
  YAML value
- body-only 600 words SUGGESTION / 900 FAIL, independent of the unchanged
  whole-file 2770-word / 500-line spec backstop
- every boundary-clause routing target must resolve to a real skill or
  agent; catches skill-improve, neuledge-context and gitea-labels
- agents take the description gates but deliberately no body gate; a test
  pins that absence

Vale: DescriptionOpener widened to ^This\b, new CompositionNote rule
banning architecture notes from descriptions. 10 hits, 0 false positives.

Kyberforge's own four skills retrofitted: descriptions 3,364 -> 938 chars
(-72%), bodies 8,306 -> 2,487 words (-70%), all via the apm-workflow
dispatch pattern. Fixes the skill-improve dangling route and the
agent-author misroute to manual review.

Also fixes a pre-existing false positive where any line-initial 'read '
was flagged as interactive input, which had already caused two scripts to
be rewritten around it.

Refs: ADR-0020
This commit is contained in:
2026-08-14 21:13:13 +00:00
parent 1c6eababb0
commit 4a5c3c0cff
104 changed files with 6272 additions and 1880 deletions

View File

@@ -0,0 +1,82 @@
---
source_keys:
- agentskills-best-practices
- agentskills-evaluating-skills
- agentskills-optimizing-descriptions
---
# Improving an existing skill
Return to `SKILL.md` Step 4 once Step 4 below is done — validation, versioning and commit
verification are shared with the create flow and are not repeated here.
## Step 1 — Verify inputs
Confirm the skill directory path exists and that at least one improvement signal is present in the
conversation or a referenced file.
If the skill directory is missing, ask for it. If no signals are present, stop: "This skill applies
existing signals to a skill. For a blind review without signals, use `/skill-audit` instead."
Signals can come from anywhere in the conversation or referenced files:
- Grill session output (most common predecessor in the factory sequence)
- `/skill-audit` findings (PASS/FAIL/SUGGESTION punch list)
- Human feedback (feedback.json, inline in conversation, PR or issue comments)
- Session context describing what went wrong
Also verify the `name` field in frontmatter matches the skill's directory name exactly.
## Step 2 — Gather and group signals
Read the current skill files (SKILL.md and any files in `scripts/`, `references/`, `assets/`,
`tests/`). Then collect all signals from the conversation and any file paths the user has
referenced.
Group signals by **root cause**, not symptom. Patching per symptom is the default failure mode:
three eval failures may all trace to one missing instruction. Ask: "What single gap in the skill
causes this cluster of failures?" One root cause → one fix. Do not make a separate edit for each
symptom.
```text
Example:
- Session context: output format is wrong on every run
- Audit finding: no output template defined
- User feedback: "I always have to ask it to format the output"
→ Root cause: SKILL.md has no output format specification → one fix: add an output template
```
## Step 3 — Announce planned changes
Before editing, state:
- Which root causes were identified and what evidence supports each
- Which files will be changed and what will change in each
Then proceed — edits are reversible via git, no approval checkpoint needed.
## Step 4 — Apply changes
Edit any file in the skill directory that the signals point to: SKILL.md, `scripts/`,
`references/`, `assets/`, `tests/`, README.md.
**Generalize, do not patch.** Find the underlying gap, not the specific example that failed. A fix
scoped only to the test cases you have seen will overfit and perform worse on new inputs.
**Keep it lean.** Remove instructions that are not pulling their weight. For every sentence you
add, ask: "Would the agent get this wrong without it?" A shorter, focused skill consistently
outperforms an exhaustive one.
**Explain the why.** Reasoning-based instructions outperform rigid directives. If you find yourself
writing a rule in all caps (ALWAYS/NEVER), reframe it: explain why the behavior matters so the
agent can apply judgment in edge cases.
**Retrofit before extending.** Any edit to a skill that predates ADR-0020 has to bring it into the
contract first — the gates are hot and carry no baseline file, so a one-line fix to a
non-compliant skill cannot be committed until the description and body meet
`references/contract.md`. Treat that retrofit as part of the same change, not a follow-up.
If a signal points to a script or reference file, edit that file directly rather than adding a
workaround in SKILL.md.
Then return to `SKILL.md` Step 4.