diff --git a/CONTEXT.md b/CONTEXT.md index fa44721..43351f4 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -26,6 +26,21 @@ A separate product (separate repo) for browsing, editing, and configuring AI dev ### Skills Reusable slash commands for AI coding tools, defined as `SKILL.md` files following the [Agent Skills open standard](https://agentskills.io). Authored at `plugins//.apm/skills//SKILL.md` and reaching a host by one of two install paths: `apm install`, which deploys the skill directory to `.claude/skills//` (this repo's own path — see "apm-consumed install"), or `claude plugin install @`, which caches the whole plugin (still supported for external consumers). Skills are self-contained — they cannot reference files outside the plugin directory after install-time caching. The two paths name skills differently: apm deploys a plain project skill (`skill-audit`), a plugin install namespaces it (`kyberforge:skill-audit`). +### Preload tax +The always-on context cost of every installed skill's `name` + `description`, which sit in the agent's context from the first token of every session whether or not the skill is invoked. Measured 2026-08-14 at 23,612 chars (~6,200 tokens) across 39 skills, plus 1,325 chars for 4 agents. Non-routing frontmatter (`metadata.source_keys`, `category`, `version`) is **not** part of it — the model-visible skill listing carries only `name` and `description`, which supersedes `LESSONS.md:63` on this host. Bodies are not part of it either; they are charged on invocation. + +### Skill context contract +The authoring rules that hold the preload tax and body size down, set by ADR-0020. A description carries a trigger clause, at most one capability clause, and a boundary clause of the form `Not → ` naming a resolvable target — nothing else. Capability enumeration, output formats, and composition notes ("composes X rather than duplicating Y") belong in the body or `README.md`; a description that summarises workflow is a correctness hazard, not just a cost, because agents act on it instead of reading the body. Sizes are two-tier and sit *below* the agentskills.io spec limits, which stay unchanged as conformance backstops: description 250 SUGGESTION / 400 FAIL (spec 1,024); body 600 SUGGESTION / 900 FAIL (spec 2,770 words / 500 lines). Conflating the quality gate with the spec ceiling is what let `skill-author` and `agent-author` grow to within twelve words of 2,770. + +### Dispatch body +The body pattern a skill with two or more mutually exclusive flows must use: the body carries only the dispatch table and the gates common to every branch, and each flow lives in its own self-contained `references/` file. Named for `apm-workflow` (554-word body, 3,006 words of references), which arrived at it independently and is the repo's exemplar. Its absence is the characteristic defect — `skill-author` inlines both its create and improve flows, and `agent-author` carries 50-60 lines marked inapplicable by their own headers on any single run. + +### Hand-invoked skill +A skill reached only by typing its slash command, declared with `disable-model-invocation: true`. The host withholds it from the model-visible skill listing entirely, so it pays no preload tax and its `description` becomes human-facing text rather than a trigger list. `zoom-out` is the worked example: apm passes the flag through verbatim to both install paths, and the skill is absent from the router while `/zoom-out` still works. Choosing model-invoked vs. hand-invoked is the first question `skill-author` asks, because it determines whether a description needs triggers at all. + +### Delegation discipline +The agent-side counterpart to the dispatch body. A plugin-scope agent is a single `.apm/agents/.agent.md` file with no sibling `references/` directory, so it cannot disclose to itself — it can only delegate to skills. Its characteristic defect is therefore restatement, not length: an agent body that spells out a procedure a skill it can invoke already owns creates a second copy that drifts. `agent-audit` fails that, with the fix being "invoke `` instead". Agents take the same description gates as skills but no body word gate — a skill body competes with the caller's live conversation, an agent body becomes the system prompt of a fresh context. + ### Plugin The deployable unit in the plugin marketplace. A plugin bundles one or more skills, agents, hooks, prompts, MCP servers, and optionally a `bin/` directory into a single installable directory. In this repo, plugins live under `plugins//`, each with its own `apm.yml` + `.apm/{skills,agents,hooks,...}` — this is the authoring source of truth for the plugin's content (ADR-0015). Two categories of tracked output are compiled from that source, never hand-edited: `.claude-plugin/plugin.json` (Claude Code) and `.github/plugin/plugin.json` (Copilot CLI) via `apm pack`/`apm compile`; and, alongside them, a flat `agents/`, `skills/`, `commands/`, `instructions/`, `extensions/` directory mirror at the plugin root plus a merged hooks file at `hooks/hooks.json`, generated by `scripts/sync-plugin-content.sh` — Claude Code's and Copilot's installers convention-scan only these flat paths (`hooks/hooks.json` is the convention path for hooks specifically; a root-level `hooks.json` is scanned by nothing and is deleted as stale by a sync — see ADR-0017's 2026-08-14 amendment) and have no awareness of `.apm/` nesting at all, so this mirror is what actually makes `.apm/` content discoverable at install time (ADR-0017). Plugins are copied to a cache on install — they cannot reference files outside their own directory. Install a plugin with `claude plugin install @`, or consume it as an apm dependency (see "apm-consumed install"). diff --git a/docs/adr/0020-skill-description-and-body-context-contract.md b/docs/adr/0020-skill-description-and-body-context-contract.md new file mode 100644 index 0000000..776e6fa --- /dev/null +++ b/docs/adr/0020-skill-description-and-body-context-contract.md @@ -0,0 +1,239 @@ +# Skills and agents are authored against a context budget, not a spec ceiling + +Every installed skill's `name` and `description` sits in every agent's context from the first token +of every session, whether or not the skill is ever invoked. Across this repo's 39 skills that is +23,612 characters — roughly 6,200 tokens — and the authoring rules that produced it optimised for +triggering reliability with no counter-pressure on size. This ADR sets the budget, the shape, and the +gates that hold them. + +**Status: accepted (2026-08-14).** + +## Context + +Measured before any change: + +| | | +|---|---| +| 39 skill `name` + `description` | 23,612 chars, ~6,200 tokens, **preloaded every session** | +| 4 agent `name` + `description` | 1,325 chars, ~350 tokens, preloaded every session | +| skill bodies | median 684 words, mean 815, p90 1,349 | +| `MAX_WORDS` gate (`skill-audit/scripts/validate.sh:147`) | **2,770** — 2× p90 | + +Three findings drove this, none of which is "the descriptions drifted". + +**The rules mandate the bloat.** `skill-author/SKILL.md:104` requires indirect triggers ("even if the +user doesn't mention X explicitly") and `skill-audit/references/description-quality.md:21` requires +authors to "err toward being pushy". Both are enforced. The one rule that would delete the waste — +`skill-author/SKILL.md:102`, "not the skill's internal mechanics" — is judgment-only and is absent +from the FAIL conditions at `description-quality.md:45-50`. The enforced rules inflate; the deflating +rule does not bite. The result is measurable: `gitea-files` spends 147 chars listing six verbs, then +301 chars re-quoting the same six as user phrasings, in the same order. `apm-workflow` does the same +with six capability clusters. Across the twelve longest descriptions, 30.7% is capability +enumeration and 11.6% is composition or implementation detail that cannot affect a routing decision. + +**Capability enumeration in a description is a correctness hazard, not only a token cost.** +`docs/research/examples/skill-write/writing-skills/SKILL.md:154-158` reports a measured failure: "when +a description summarizes the skill's workflow, an agent may follow the description instead of reading +the full skill content. A description saying 'code review between tasks' caused an agent to do ONE +review, even though the skill's flowchart clearly showed TWO reviews." `git-commits` is exactly that +shape — 74% of its description is capability enumeration, including a rules table (`header max 100 +chars, lowercase subject, no trailing periods, 11 standard types`) an agent can act on without ever +loading the body. + +**The upstream sources cannot settle this.** The four skill-writing references under +`docs/research/examples/skill-write/` disagree on what a description contains — when-only +(`writing-skills/SKILL.md:99`), what-and-when (`skill-creator/SKILL.md:67`, +`anthropic-best-practices.md:187`), triggers-only (`writing-great-skills/SKILL.md:28`), and +what-plus-when-plus-negative (`write-skill/SKILL-TEMPLATE.md:5-6`). `writing-skills` and the Anthropic +document it bundles contradict each other inside one skill directory. They also disagree on whether +500 lines is binding, on the inline-versus-bundle threshold, and on the TOC threshold (>100 lines vs +>300 lines). "Grounded in the research" is therefore not available as a tiebreaker; a house choice is +required and this is it. + +A fourth observation shaped the body half. The best progressive-disclosure ratio in the repo belongs +to `apm-workflow` — a 554-word body dispatching to 3,006 words of references — and the worst two +belong to the skills that define the house standard: `skill-author` (2,760 body / 1,247 references) +and `agent-author` (2,758 / 1,664). Both sit within twelve words of the 2,770 gate their own plugin +enforces. A ceiling that nothing approaches is not a constraint; a ceiling that two files have grown +into is a target. + +## Decision + +### Descriptions + +A description carries three things and nothing else: a **trigger clause**, at most one **capability +clause**, and a **boundary clause**. Capability enumeration, output-format detail, composition notes +("composes X rather than duplicating Y"), and implementation detail move to the body or to +`README.md`. + +- **250 characters SUGGESTION, 400 FAIL.** The agentskills.io 1,024-character limit remains as an + unchanged spec backstop. The SUGGESTION tier is what moves the average; the FAIL tier only stops + outliers. +- **Boundary clauses compress** to `Not → .` and must name a target that + resolves to a real skill under `plugins/*/.apm/skills/`. This is checked deterministically. +- **The blanket pushiness rules are deleted.** `skill-author/SKILL.md:104` and + `description-quality.md:21` are replaced by a conditional: add an indirect trigger only where the + user's natural phrasing genuinely omits the domain word — true for the `gitea-*` family, false for + `git-commits`. Stating the same trigger twice in two registers is a FAIL. + +### Bodies + +The body carries the **decision procedure only**: ordered steps, decision branches, gates, and which +reference to load when. Lookup tables, spec restatements, output schemas, templates, and rationale +prose move to `references/` behind an explicit "read X when Y" trigger. + +- **600 words SUGGESTION, 900 FAIL.** The 2,770-word / 500-line spec backstop is unchanged and keeps + its existing meaning: conformance, not quality. These are two different gates measuring two + different things, and conflating them is what produced the current state. +- **Dispatch is mandatory at two or more mutually exclusive flows.** The body carries the dispatch + table and the gates that apply to every branch; each flow lives in its own self-contained + `references/` file. This is `apm-workflow/SKILL.md:33-41` promoted from accident to rule. +- **Gotchas are constrained.** A Gotcha must state a fact that contradicts a reasonable default — + something the agent gets wrong by acting sensibly. Maximum five entries. A Gotcha that paraphrases + a step in the body below it is a FAIL. A Gotchas section exceeding 25% of the body is a + SUGGESTION. + +### Agents + +Agents take the same description gates — they are preloaded identically — and **no body word gate**. +A skill body is loaded into the caller's context, competing with the live conversation; an agent body +becomes the system prompt of a fresh context. The rationale for the 900-word FAIL does not transfer. + +A plugin-scope agent is a single file with no sibling `references/` directory, so it cannot disclose +to itself — it can only delegate to skills. `agent-audit` therefore gains a **delegation check**: an +agent body that restates a procedure owned by a skill it can invoke is a FAIL, with the fix being +"invoke `` instead". Length falls out of delegation rather than being gated directly. + +### Invocation as a design axis + +`skill-author` asks whether a skill is model-invoked or hand-invoked before writing a description. A +hand-invoked skill sets `disable-model-invocation: true` and carries one plain human-facing sentence +with no trigger list. + +Verified end-to-end rather than assumed: `plugins/bin/.apm/skills/zoom-out/SKILL.md:4` carries the +flag, apm passes it through verbatim to both `.claude/skills/zoom-out/SKILL.md:4` and the flat mirror +at `plugins/bin/skills/zoom-out/SKILL.md:4`, and `zoom-out` is the one installed skill absent from +the model-visible skill listing in a live session. It remains invocable as `/zoom-out`. + +### Merging siblings + +Two skills that share substantial content, name each other as near-misses, and differ only in the +type of input they take should be **one skill with a dispatch table**. This catches `skill-audit` + +`agent-audit` and is scoped to them; the author pair is explicitly excluded, because +`skill-author` and `agent-author` emit genuinely different artifacts (a skill directory versus a +one-or-two-file agent pair, per ADR-0005 and ADR-0016) and their overlap is in the improve flow +rather than the core job. + +### Enforcement and rollout + +Gates land where the existing gates already live — no new layer: + +| Check | Home | +|---|---| +| description and body counts, resolvable boundary targets | `skill-audit/scripts/validate.sh`, `scripts/skill-size-check.sh` | +| prose patterns (composition-note openers, restatement) | `plugins/kyberforge/.apm/skills/*/assets/vale/styles/Kyberforge/` | +| judgment calls | `references/description-quality.md`, `references/body-discipline.md` | + +**Blocking immediately, with no baseline file.** + +Three pre-existing contradictions are fixed in the same change, because they are the contract: + +- `skill-audit/SKILL.md:58` asks whether the description opens with an action verb ("Audits…", + "Reviews…") while `:101` and `DescriptionOpener.yml` require an imperative "Use when…" opener. The + criterion is unsatisfiable against the house's own skills, both of which open with "Use when". +- `DescriptionOpener.yml`'s regex is anchored to `^This (skill|agent)\b`, so `gitea-workflow` ("This + is the human-facing entry point…") and `gitea-labels-milestones` ("This is a cross-cutting shared + skill…") both violate the rule and pass the linter. +- `description-quality.md:45-50` has no FAIL condition for internal-mechanics content, which is why + `skill-author/SKILL.md:102` never bit. + +## Consequences + +**Editing any non-compliant skill now requires retrofitting it first.** 30 of 39 descriptions exceed +400 characters and 14 of 39 bodies exceed 900 words. With the gate hot and no baseline, a one-line +fix to `gitea-prs` cannot be committed until that skill meets the contract. This is deliberate — it +guarantees convergence and avoids a half-state — but it means the retrofit is lazy and *mandatory* +rather than deferred. The follow-up retrofit issue should be prioritised accordingly, and the risk it +carries is the ordinary one for hot gates: a gate expensive enough to be inconvenient gets bypassed +with `SKIP=` and loses its authority. + +**A ceiling does not produce an average.** If every author writes to the 400-character FAIL, the +preload lands at 15,600 chars — a 34% cut, not the ~50% intended. The halving depends entirely on the +250-character SUGGESTION tier being visible and respected. That tier works here in a way it does not +elsewhere in this repo: `skill-audit` already reports `PASS (N suggestions)` as a first-class +outcome. This is explicitly **not** the failure ADR-0013 records — Vale warnings are invisible +because vale's exit code keys on `error` alone, but these gates live in `validate.sh` and +`skill-audit`, where a SUGGESTION reaches the report. Realistic landing is 34-55% down, not a +guaranteed 50%. + +**A word gate cannot detect the defect it is standing in for.** `git-commits` carries thirteen +Gotchas of which four restate steps in its own Workflow (`:32` ≡ step 9, `:33` ≡ step 9, `:36` ≡ step +2, `:31` ≡ the description) — 1,217 words that pass any plausible gate. The counts are a backstop to +the dispatch rule and the Gotchas constraint, not a substitute for them, and should not be read as +the mechanism. + +**Some skills legitimately need more description budget than others.** A tiered limit keyed to +sibling density was considered and rejected as too clever; the flat 250/400 pair means the `gitea-*` +and `git-*` families — where every sibling shares a keyword and boundary clauses do real routing work +— are the ones most likely to sit at the FAIL tier permanently. If the retrofit shows that family +routing degrades, the tier is the first thing to revisit. + +**Four broken routing targets are live and are not fixed here.** `skill-audit` routes to +`/skill-improve` twice in its description plus `README.md:10`, and no such skill exists — the real +target is `skill-author`. `research` routes to `neuledge-context`, which exists only inside that +string. `agent-author` says "Do not use for read-only review — examine agent files manually", +routing away from `agent-audit`, the correct sibling. `gitea-issues` contains the literal string +`gitea-labels- milestones`, a stray space introduced by YAML folding, breaking the skill name in +preloaded text. The resolvable-target check added here will fail on all four the moment those files +are touched; fixing them is split into its own issue. + +**Duplication between `skill-author` and `agent-author` survives un-gated.** The merge rule +deliberately excludes the author pair, so the commit-verification argument in four near-copies, the +root-cause grouping rule in four copies, and the wholesale clone of the "Improving an existing X" +flow all remain. Cache isolation makes them structurally unavoidable +(`skill-audit/SKILL.md:95` forbids cross-skill references; `LESSONS.md:107` records why), so the +options are a sync gate or continued drift. This is an input to the kyberforge-bodies follow-up +issue, not a solved problem. + +**Provenance frontmatter is explicitly out of scope.** `LESSONS.md:63` asserts that non-routing +frontmatter (`source_keys`, `category`, `version`) is loaded at agent startup, which would make the +7,352 characters of it across the corpus a third again on top of the description tax. Measured +against a live session on this Claude Code version, it is not: the model-visible skill listing +contains only `name` and `description`. That is host-observed rather than spec-guaranteed and says +nothing about Copilot CLI, but it is sufficient to establish that cutting `source_keys` would break +the ADR-0009 provenance machinery for no runtime gain. The metadata was added deliberately and stays. + +## Alternatives considered + +- **Keep pushiness, raise the budget to ~500 chars.** Undertriggering is the worse failure mode — a + skill that never fires is worth nothing regardless of cost — and `skill-creator/SKILL.md:67` + explicitly recommends being "pushy" against an observed undertriggering tendency. Rejected because + that claim is an unmeasured assertion about an older model, and because the correctness hazard in + `writing-skills:154-158` cuts the other way: a fat description is not merely expensive, it is a + shortcut agents take instead of reading the body. Would have landed a 35% cut. +- **A trigger-eval loop to set lengths empirically.** `skill-creator/SKILL.md:337-404` specifies 20 + queries per skill, 8-10 positive and 8-10 near-miss, with a 60/40 train/test split selecting on + test score. This is the rigorous answer and the repo has deliberately never built it. Rejected + because it blocks the context cut behind a substantial new subsystem. +- **A repo-level aggregate preload budget** (≤12,000 chars across all skills, checked at pre-push). + The only option that measures the actual goal rather than a proxy. Rejected because it makes one + skill's edit fail on account of another skill's growth, and because it is meaningless for an + external consumer installing a subset of the plugins. +- **500-word body FAIL, matching `writing-skills/SKILL.md:217-221`.** Best-grounded in upstream and + would align this repo with the tightest source. Rejected because it fails 30 of 39 skills, and a + blunt gate gets satisfied by deleting content rather than relocating it. +- **A shrinking baseline file** recording each non-compliant skill's current numbers, failing only on + growth. Would have made the retrofit a visible burn-down instead of a wall. Rejected in favour of + hot gates. +- **A sync gate over the duplicated spans** instead of a merge rule — generalising + `scripts/check-vale-style-sync.sh` to cover shared prose so duplication persists but drift cannot. + Rejected for the audit pair in favour of merging, which removes the duplication rather than + policing it, and removes a mutually-excluding near-miss pair from the router at the same time. It + remains the only available answer for the author pair. +- **Merging `skill-author` + `agent-author` as well**, taking kyberforge from seven skills to five. + Largest cut available. Rejected because it reopens ADR-0005, ADR-0008 and ADR-0016 together, and a + merged author skill would carry both the skill-directory scaffold and the dual-provider agent + scaffold behind one dispatch. +- **Demoting Gotchas** to the end of the body or into `references/gotchas.md`, removing its + position-based exemption from the dispatch rule. Maximum saving on the largest body construct + (6,830 words, 21% of all body text). Rejected because a gotcha read after the mistake is worthless.