feat(kyberforge): enforce the ADR-0020 context contract for skills and agents
Skill name+description pairs are preloaded into every session, costing ~6,200 tokens across 39 skills before any skill is invoked. The authoring rules mandated that growth: skill-author:104 and description-quality.md:21 both required padding, while skill-author:102 (the deflating rule) had no FAIL condition behind it. Gates (blocking, no baseline file): - description 250 chars SUGGESTION / 400 FAIL, measured on the folded YAML value - body-only 600 words SUGGESTION / 900 FAIL, independent of the unchanged whole-file 2770-word / 500-line spec backstop - every boundary-clause routing target must resolve to a real skill or agent; catches skill-improve, neuledge-context and gitea-labels - agents take the description gates but deliberately no body gate; a test pins that absence Vale: DescriptionOpener widened to ^This\b, new CompositionNote rule banning architecture notes from descriptions. 10 hits, 0 false positives. Kyberforge's own four skills retrofitted: descriptions 3,364 -> 938 chars (-72%), bodies 8,306 -> 2,487 words (-70%), all via the apm-workflow dispatch pattern. Fixes the skill-improve dangling route and the agent-author misroute to manual review. Also fixes a pre-existing false positive where any line-initial 'read ' was flagged as interactive input, which had already caused two scripts to be rewritten around it. Refs: ADR-0020
This commit is contained in:
@@ -6,49 +6,112 @@ source_keys:
|
||||
|
||||
# Description Quality Reference
|
||||
|
||||
Source: agentskills.io — optimizing-descriptions
|
||||
Upstream source: agentskills.io — optimizing-descriptions, specification.
|
||||
House contract: ADR-0020, the context budget. The house contract is narrower than the spec
|
||||
rather than a reinterpretation of it: where both speak, both must be satisfied.
|
||||
|
||||
## How triggering works
|
||||
## Why the description is the expensive part
|
||||
|
||||
At startup, agents load only the `name` and `description` of each skill. When a user's task matches a description, the agent reads the full `SKILL.md` into context. **The description carries the entire triggering burden** — the body is never seen until after triggering.
|
||||
At startup an agent loads only the `name` and `description` of every installed skill. The body is
|
||||
never seen until the skill triggers. The description therefore carries the entire triggering
|
||||
burden **and** is paid for in every session, whether the skill fires or not.
|
||||
|
||||
Agents typically consult skills only for tasks requiring knowledge beyond their defaults. Specialized knowledge — unfamiliar APIs, domain-specific workflows, uncommon formats — is where description wording makes the difference.
|
||||
A second cost is less obvious and is a correctness hazard rather than a token cost: a description
|
||||
that summarises the workflow is a shortcut the agent takes *instead of* reading the body. A
|
||||
measured failure upstream — a description saying "code review between tasks" — produced one review
|
||||
where the body's flowchart specified two.
|
||||
|
||||
## What a good description does
|
||||
## Step 0 — establish which contract applies
|
||||
|
||||
- **Imperative phrasing** — "Use when..." not "This skill does...". The agent is deciding whether to act.
|
||||
- **User intent, not mechanics** — describe what the user is trying to achieve, not how the skill works internally.
|
||||
- **Err toward being pushy** — explicitly name contexts where the skill applies, including cases where the user doesn't name the domain: "even if they don't mention X explicitly."
|
||||
- **Specificity over vagueness** — "parses and validates OpenAPI specs" beats "helps with APIs."
|
||||
- **Near-miss exclusions** — add "Do not use when..." only if a near-miss skill exists that could steal activations. Use strong near-misses (queries that share keywords but need something different), not weak ones ("write a fibonacci function").
|
||||
- **Hard limit: 1024 characters** — descriptions grow during revision; check length before finalising.
|
||||
Read the frontmatter before judging a single word.
|
||||
|
||||
- **`disable-model-invocation: true`** — the skill is hand-invoked. Its description is never
|
||||
matched against user intent, so it is not a routing string. It carries **one plain human-facing
|
||||
sentence** stating what the skill does. Audit it for that and nothing else. Reporting a missing
|
||||
trigger clause, a missing boundary clause or absent indirect triggers on a hand-invoked skill is
|
||||
a wrong finding, not a strict one.
|
||||
- **No such flag** — the skill is model-invoked and the rest of this file applies.
|
||||
|
||||
## The three-part shape
|
||||
|
||||
A model-invoked description carries exactly three things:
|
||||
|
||||
1. **Trigger clause.** When to invoke, phrased imperatively: `Use when ...`. Not `This skill ...` —
|
||||
the agent is deciding whether to act, not reading a catalogue entry.
|
||||
2. **At most one capability clause.** What it does, in one clause. Never an enumeration.
|
||||
3. **Boundary clause.** Compressed form: `Not <thing> -> <skill-name>.` The target must resolve to
|
||||
a real skill directory or agent file in the authoring source; `validate.sh` checks that
|
||||
deterministically and a dangling target already surfaces as a Structure FAIL.
|
||||
|
||||
Everything else belongs in the body or in `README.md`.
|
||||
|
||||
## Indirect triggers — conditional, never blanket
|
||||
|
||||
Add "even if the user doesn't say X" **only where the user's natural phrasing genuinely omits the
|
||||
domain word.** True for the `gitea-*` family: people say "create an issue", not "create a Gitea
|
||||
issue". False for `git-commits`: nobody asks for a commit without saying commit. A blanket
|
||||
indirect-trigger clause on a skill whose domain word is unavoidable is padding charged to every
|
||||
session.
|
||||
|
||||
## Near-miss exclusions
|
||||
|
||||
Add a boundary clause only where a sibling skill could plausibly steal the activation. Use strong
|
||||
near-misses — queries that share keywords but need something different — not weak ones ("write a
|
||||
fibonacci function"). One boundary clause per genuine near-miss; a list of four is enumeration
|
||||
wearing a boundary's clothes.
|
||||
|
||||
## Before / after
|
||||
|
||||
```yaml
|
||||
# Weak
|
||||
description: Process CSV files.
|
||||
|
||||
# Strong
|
||||
# FAIL — enumeration first, mechanics as the opener, a blanket indirect trigger,
|
||||
# and 300+ characters of it preloaded into every session forever.
|
||||
description: >
|
||||
Analyze CSV and tabular data files — compute summary statistics,
|
||||
add derived columns, generate charts, and clean messy data. Use when
|
||||
the user has a CSV, TSV, or Excel file and wants to explore, transform,
|
||||
or visualize the data, even if they don't explicitly mention "CSV" or
|
||||
"analysis."
|
||||
Analyze CSV and tabular data files — compute summary statistics, add derived
|
||||
columns, generate charts, and clean messy data. Use when the user has a CSV,
|
||||
TSV, or Excel file and wants to explore, transform, or visualize the data,
|
||||
even if they don't explicitly mention "CSV" or "analysis."
|
||||
|
||||
# PASS — trigger, one capability clause, boundary. The four verbs the FAIL
|
||||
# version enumerates are the body's job; the router cannot act on them.
|
||||
description: >
|
||||
Use when the user has a CSV, TSV, or Excel file and wants it explored,
|
||||
transformed, or charted. Not schema design -> data-model.
|
||||
```
|
||||
|
||||
The strong version names capabilities precisely and broadens applicability beyond explicit keyword matches.
|
||||
(`data-model` is illustrative. In a real description the target has to resolve.)
|
||||
|
||||
## Auditing guidance
|
||||
|
||||
Flag as FAIL if:
|
||||
- Phrasing is descriptive ("This skill...") not imperative ("Use when...")
|
||||
- Capabilities are vague ("helps with APIs") — require precise verbs and nouns
|
||||
- No indirect trigger coverage when indirect cases clearly exist
|
||||
- No near-miss exclusions when a sibling skill could plausibly steal activations
|
||||
- Length exceeds 1024 characters
|
||||
|
||||
- **Over 400 characters.** Measured on the folded YAML value, not the raw source lines.
|
||||
`validate.sh` reports the number; do not re-derive it, but do point the Fix at what to cut.
|
||||
- **Internal mechanics appear in the description.** Any of:
|
||||
- capability enumeration or a feature list;
|
||||
- output-format detail ("Produces a compact findings report with Why and Fix per finding");
|
||||
- composition or architecture notes ("composes X rather than duplicating Y", "a cross-cutting
|
||||
shared skill", "the human-facing entry point", "replaces the old flat invocation");
|
||||
- implementation detail ("self-validates via a bundled deterministic script").
|
||||
|
||||
None of it can change a routing decision and all of it is preloaded.
|
||||
`Kyberforge.CompositionNote` catches the common phrasings deterministically; the rest is
|
||||
judgment. This is the rule that deflates a description, so apply it before reaching for length.
|
||||
- **The same trigger stated twice in two registers** — a verb list, then the same verbs re-quoted
|
||||
as user phrasings, usually in the same order. One register, whichever routes better.
|
||||
- **Descriptive rather than imperative phrasing** (`This skill ...`, `This is the ...`).
|
||||
`Kyberforge.DescriptionOpener` catches any opener matching `^This`.
|
||||
- **Vague capabilities** ("helps with APIs" where "parses and validates OpenAPI specs" was
|
||||
available). `Kyberforge.VagueWording` catches the known filler; imprecision outside that list is
|
||||
judgment.
|
||||
- **A boundary clause naming a target that does not resolve** to a real skill directory or agent
|
||||
file in the authoring source. `validate.sh` reports the unresolved name.
|
||||
- **Trigger-list, boundary or indirect-trigger content on a hand-invoked skill** — see Step 0.
|
||||
- **Over 1024 characters** — the agentskills.io specification ceiling, unchanged and independent
|
||||
of the 400-character house ceiling above.
|
||||
|
||||
Flag as SUGGESTION if:
|
||||
- Indirect trigger coverage exists but could be more specific
|
||||
- Near-miss exclusions are present but target weak near-misses only
|
||||
|
||||
- **Over 250 characters** but at or under 400. This tier is what moves the corpus average; the FAIL
|
||||
tier only stops outliers. Report it rather than treating a 399-character description as clean.
|
||||
- A near-miss exclusion is present but targets a weak near-miss.
|
||||
- An indirect trigger is present and warranted but could name the omitted phrasing more precisely.
|
||||
|
||||
Reference in New Issue
Block a user