Compare commits
7 Commits
76075223c7
...
e7ebc667b3
| Author | SHA1 | Date | |
|---|---|---|---|
| e7ebc667b3 | |||
| 64ffb9f35a | |||
| d02765d595 | |||
| 311e7cd22c | |||
| 2540e50fcc | |||
| a85bdbed42 | |||
| b6e68e9a2b |
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"name": "holocron",
|
||||
"description": "AI development skills for Claude Code and GitHub Copilot CLI — factory, design, implement, review, and cross-cutting workflows.",
|
||||
"version": "0.4.1",
|
||||
"version": "0.4.2",
|
||||
"owner": {
|
||||
"name": "Defame1297",
|
||||
"email": "defame1297@rkdr.net",
|
||||
@@ -11,7 +11,7 @@
|
||||
{
|
||||
"name": "kyberforge",
|
||||
"description": "Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.",
|
||||
"version": "1.5.0",
|
||||
"version": "1.6.0",
|
||||
"category": "Developer Tools",
|
||||
"source": "./plugins/kyberforge"
|
||||
},
|
||||
|
||||
4
.github/plugin/marketplace.json
vendored
4
.github/plugin/marketplace.json
vendored
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"name": "holocron",
|
||||
"description": "AI development skills for Claude Code and GitHub Copilot CLI — factory, design, implement, review, and cross-cutting workflows.",
|
||||
"version": "0.4.1",
|
||||
"version": "0.4.2",
|
||||
"owner": {
|
||||
"name": "Defame1297",
|
||||
"email": "defame1297@rkdr.net",
|
||||
@@ -11,7 +11,7 @@
|
||||
{
|
||||
"name": "kyberforge",
|
||||
"description": "Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.",
|
||||
"version": "1.5.0",
|
||||
"version": "1.6.0",
|
||||
"category": "Developer Tools",
|
||||
"source": "./plugins/kyberforge"
|
||||
},
|
||||
|
||||
@@ -38,14 +38,15 @@ Fall back to raw shell only when no skill covers it.
|
||||
## Setup and testing
|
||||
|
||||
- Run `apm install` to deploy this repo's own skills and agents into `.claude/skills/` and `.claude/agents/`. Both are gitignored install output, not authoring source — `plugins/<name>/.apm/` remains the only place to edit. The six dependencies in root `apm.yml` resolve from the holocron **remote**, unpinned against the default branch, so a `.apm/` edit is not visible to the running session until it is pushed and `apm update` re-runs (`apm install` deploys from `apm.lock.yaml` and does not re-resolve refs). Needs the network, and needs `apm_modules/` (which it materializes) left gitignored. `apm install` also configures the `obsidian` MCP server into the repo's `.mcp.json`, carried over from `plugins/bin/.mcp.json`.
|
||||
- Do not add repo-owned keys to `.claude/settings.json`. apm treats that file as its own deployed artifact: `apm audit --ci` replays the install into a scratch tree and diffs, so anything apm would not have written there — an `enabledPlugins` block, a real `hooks` entry — is permanent drift that fails the `apm-audit-ci` pre-push hook. Its committed content is whatever apm last wrote — `{"hooks": {}}` until kyberforge's `SessionStart` hook lands there, after which the merged hook entry is apm's output and belongs in the commit (ADR-0019). What does not change is that nothing repo-authored goes in the file. A hook you want in this repo is authored in `plugins/<name>/.apm/hooks/` and deployed by apm, never hand-written here. Machine-specific settings go in the gitignored `.claude/settings.local.json`, which apm does not deploy and the replay does not compare; shared enforcement belongs in `.pre-commit-config.yaml`.
|
||||
- Do not add repo-owned keys to `.claude/settings.json`. apm treats that file as its own deployed artifact: `apm audit --ci` replays the install into a scratch tree and diffs, so anything apm would not have written there — an `enabledPlugins` block, a real `hooks` entry — is permanent drift that fails the `apm-audit-ci` pre-push hook. Its committed content is whatever apm last wrote, which today is the merged `SessionStart` entry for kyberforge's `check-apm-current.sh` — apm's own output, and it belongs in the commit (ADR-0019). What does not change is that nothing repo-authored goes in the file. A hook you want in this repo is authored in `plugins/<name>/.apm/hooks/` and deployed by apm, never hand-written here. The file is also **excluded from `pretty-format-json`** in `.pre-commit-config.yaml` — the sixth and last alternation in that `exclude:` pattern, and the only one there for a reason other than "generated manifest". Mind which number you are quoting: six alternations, expanding to sixteen real files (3 root marketplace manifests, 2 per plugin × 6 plugins, plus this one). `pretty-format-json --autofix` sorts object keys while apm emits insertion order, so leaving the file in that hook's scope rewrites apm's output on the way into every commit and `apm audit --ci` then reports permanent drift on a file with an empty `git diff`. Do not tidy it out of that list; it is load-bearing (see `LESSONS.md`, 2026-08-14). Machine-specific settings go in the gitignored `.claude/settings.local.json`, which apm does not deploy and the replay does not compare; shared enforcement belongs in `.pre-commit-config.yaml`.
|
||||
- Keeping the install current is automatic but not free. Because the six dependencies are unpinned, deployed skills go stale whenever anyone merges. kyberforge ships a `SessionStart` hook that runs `apm outdated` at startup (~0.7s) and, when something is behind, runs `apm update --yes` and asks the host to re-scan skills (~10.4s). That rewrites `apm.lock.yaml`, so an unexplained modification to it after opening a session is expected, not a bug — commit or discard it deliberately. Note `apm install` alone will **not** pick up remote changes; it deploys from the lock. `apm update` is the command that re-resolves refs.
|
||||
- Install git hooks via `pc-run`, wiring all three stages — this repo's `.pre-commit-config.yaml` has no `default_install_hook_types`, so a plain install silently skips `commit-msg` (Conventional Commits) and `pre-push` (the 14-hook gate described below).
|
||||
- Install the `apm` CLI — four pre-push hooks shell out to it: `apm-marketplace-check`, `apm-audit-ci`, `apm-pack-check-clean`, and `check-plugin-content-sync` (via `scripts/sync-plugin-content.sh`, which wraps `apm pack`). `apm-marketplace-check` and `apm-pack-check-clean` are bare `apm …` hook entries and `apm-audit-ci` is a `bash -c` loop calling `apm` once per package, so without it the push dies with an unhelpful "command not found". Use `apm-install`, or `curl -sSL https://aka.ms/apm-unix | sh`; verify with `apm --version`.
|
||||
- Install `jq` — required by `scripts/check-manifests.sh` and `scripts/sync-plugin-content.sh`, both pre-push. These at least fail loudly (`Error: jq is required but not installed`).
|
||||
- Install `python3` — required by `scripts/skill-size-check.sh`, the `skill-size-check` pre-commit hook. It measures the *folded* `description` value: most descriptions here are `>`-block scalars, so a regex over the raw lines measures indentation and newlines instead of the value. Missing it fails the hook with an install pointer rather than skipping the ADR-0020 checks, which would be a vacuous green. In practice it is already present — pre-commit is itself a Python application. PyYAML is used when importable and is genuinely optional; a fallback reader covers the frontmatter shapes this corpus uses.
|
||||
- That hook enforces **two independent gate families** over `plugins/*/.apm/skills/*/SKILL.md`, and neither replaced the other. The agentskills.io spec backstop is unchanged: 500 lines and 2,770 words, counted over the **whole file including frontmatter**. ADR-0020 adds a context budget measured differently — `description` 250 chars SUGGESTION / 400 FAIL (it is preloaded into every session whether the skill fires or not), **body-only** word count 600 SUGGESTION / 900 FAIL (everything after the frontmatter's closing `---`), and every boundary-clause routing target resolving to a real skill or agent under `plugins/*/.apm/`. A file can sit well inside one family and fail the other. The hook is `verbose: true` so the SUGGESTION tier is audible — pre-commit prints nothing at all for a passing hook, and a SUGGESTION deliberately does not fail. `skill-audit`'s `validate.sh` holds a second copy of the four ADR-0020 constants; `tests/test-skill-size-check.sh` asserts the copies agree.
|
||||
- Install `python3` — required by `scripts/skill-size-check.sh`, the `skill-size-check` pre-commit hook. It measures the *folded* `description` value: most descriptions here are `>`-block scalars, so a regex over the raw lines measures indentation and newlines instead of the value. Missing it fails the hook with an install pointer rather than skipping the ADR-0020 checks, which would be a vacuous green. In practice it is already present — pre-commit is itself a Python application. **PyYAML is a hard requirement too**, not an optional accelerator: the hand-rolled fallback frontmatter reader has been removed, because a reader that mis-parses an unfamiliar scalar shape reports a clean pass on a file it never measured, which is the exact vacuous-green failure the `python3` check exists to avoid. `pip install pyyaml` if the hook reports it missing.
|
||||
- That hook enforces **two independent gate families** over `plugins/*/.apm/skills/*/SKILL.md`, and neither replaced the other. The agentskills.io spec backstop is unchanged: 500 lines and 2,770 words, counted over the **whole file including frontmatter**. ADR-0020 adds a context budget measured differently — `description` 250 chars SUGGESTION / 400 FAIL (it is preloaded into every session whether the skill fires or not), **body-only** word count 600 SUGGESTION / 900 FAIL (everything after the frontmatter's closing `---`), a missing, valueless or `null` `description:` (a hard FAIL, not a skip — a gate that declines to measure the one preloaded field reports green), every boundary-clause routing target resolving to a real skill or agent, and every `references/<file>.md` a body names actually existing. Target resolution walks up **from the file being checked** to an authoring root — the nearest ancestor holding `plugins/*/.apm/{skills,agents}`, falling back to the nearest `.git`, in two passes so a nested `.git` cannot beat a real monorepo root. The universe is then every skill and agent under `<root>/plugins/*/`, plus the checked file's own apm package and whatever that package declares in its own `apm.yml` `dependencies.apm`; the **root** manifest's `dependencies:` block is not read, and no plugin here declares a cross-plugin apm dependency. Deployed `.claude/`/`.agents/` trees are consulted only when no authoring root exists — the consumer case. That matters because those trees are gitignored `apm install` output: resolution used to reach the four cross-plugin `gitea-*` → `git-*` targets through `.claude/skills/` alone, so the same commit measured 2 dangling targets on a developer machine and 6 on a fresh clone. It no longer does — verified by running the hook over a tree holding only `plugins/` and the root `apm.yml`, which reports findings identical to the working tree (26 description / 9 body / 2 dangling / 0 missing references / 58 SUGGESTIONs). Three further checks are SUGGESTION-only: a description with no boundary clause at all, a `## Gotchas` section with more than five entries, and a `## Gotchas` section over 25% of the body. A file can sit well inside one family and fail the other. The hook is `verbose: true` so the SUGGESTION tier is audible — pre-commit prints nothing at all for a passing hook, and a SUGGESTION deliberately does not fail. `skill-audit`'s `validate.sh` holds a second copy of the four ADR-0020 constants; `tests/test-skill-size-check.sh` asserts the copies agree.
|
||||
- **Those ADR-0020 gates ship hot, with no baseline file.** 26 of 39 descriptions and 9 of 39 bodies currently exceed their FAIL tier, so editing one of those skills *for any reason* means retrofitting it to the contract first — a one-line fix to `gitea-prs` cannot be committed until that skill complies. This is deliberate, and the retrofit is tracked as Gitea issue #99. Check where a skill stands before starting: `pre-commit run skill-size-check --all-files`.
|
||||
- **A second gate ships hot alongside it, and `skill-size-check` will not warn you about it.** `Kyberforge.CompositionNote` — the ADR-0020 Vale rule banning composition and architecture prose from a description — currently fires **10 errors across four skills**: `gitea-issues`, `gitea-labels-milestones`, `gitea-prs` and `gitea-workflow`. Every Vale rule here is `level: error` with no ignorable tier, so touching any of those four means fixing its prose findings as well as its size findings. Scoping a retrofit off `skill-size-check` output alone will leave you blocked at the second gate. Check both: `pre-commit run --all-files`.
|
||||
- Install the `vale` binary — required by the `vale-audit-prefilter-skill`/`-agent` pre-commit hooks. Their `files:` patterns are `.apm/`-scoped: `^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$` and `^plugins/[^/]+/\.apm/agents/[^/]+\.agent\.md$`. Only the authoring source triggers them — a `SKILL.md` in the generated mirror matches neither pattern, so prose findings surface only when you edit the file you are supposed to be editing. Without the binary the hooks fail with a bare "command not found" and no install pointer. `brew install vale` (macOS), `snap install vale` (Linux), `choco install vale` (Windows), or see https://vale.sh/docs/vale-cli/installation/. No `vale sync` needed — the `Kyberforge` styles are committed under `plugins/kyberforge/.apm/skills/{skill-audit,agent-audit}/assets/vale/styles/`, not downloaded packages (see ADR-0014).
|
||||
- `vale` is also a **pre-push** dependency, not only pre-commit. `check-vale-style-sync` runs six glob-coverage probes by invoking `vale --config` — they are the only assertions in it that catch a `.vale.ini` glob typo, the failure mode where every text-level check stays clean while vale lints zero files. Missing `vale` is therefore a hard failure there. The opt-out is `CHECK_VALE_STYLE_SYNC_ALLOW_MISSING_VALE=1`, and it is **not** `SKIP=`: the hook still runs and still asserts everything verifiable from file text, but the six probes do not, and its summary says so explicitly — `Vale style sync check passed (text-level only, vale unavailable): … 0 glob probe(s) verified`. Use it only on a machine that genuinely cannot install `vale`, and read that summary line as "the glob axis was not checked", not as a pass.
|
||||
- Run `bash tests/run-tests.sh` before considering any change done — it runs every `test-*.sh` script in the repo plus the bats suite (`--bats-only` for just bats). First run auto-initializes the bats submodules; no manual `git submodule update` needed.
|
||||
@@ -53,7 +54,7 @@ Fall back to raw shell only when no skill covers it.
|
||||
- `tests/run-bats.sh` derives the set of `.bats` files it expects from `git ls-files`, so a `.bats` file deleted from the worktree but still tracked in the index fails the run rather than silently shrinking the suite. Remove one with `git rm` (or stage the deletion) when the removal is intentional; an untracked new `.bats` file is picked up and needs no ceremony. Both discovery walks (`tests/run-bats.sh` and `tests/run-tests.sh`) exclude `apm_modules/`: `apm install` materializes a full copy of every plugin there, and running a dependency's copy of a `.bats` file breaks its relative path to the bats helpers — 167 spurious failures before the exclusion landed.
|
||||
- Pushing runs 14 repo-defined pre-push hooks, not just the test suite — `run-tests` and `check-manifests`, plus generated-content drift gates (`check-plugin-content-sync`, `check-marketplace-mirror-sync`, `check-vale-style-sync`, `check-scope-walkup-sync`, `check-executables-allow-sync`), artifact validators (`check-apm-agents-valid`, which runs agent-audit's `validate.sh` over every real `plugins/*/.apm/agents/*.agent.md`), apm's own gates (`apm-marketplace-check`, `apm-audit-ci`, `apm-pack-check-clean`), host validators (`validate-plugins`, `validate-marketplace`, both needing the `claude` CLI), and `check-release-needed`. `check-executables-allow-sync` is the odd one in that first group — it guards a silent failure rather than drift in generated text. apm gates a package's `hooks/` and `bin/` on an exact `<package>#<version>` lookup in root `apm.yml`'s `executables.allow`, with no wildcard and no version-less form, so bumping `plugins/kyberforge/apm.yml`'s `version:` without bumping the key errors nowhere: the entry simply stops matching, kyberforge's `SessionStart` hook stops deploying, and the install goes quietly stale — the failure ADR-0019 records as live. Run `pre-commit run --hook-stage pre-push --all-files` locally — one command, the whole gate. That command reports **16**, not 14: pre-commit's own `meta` hooks, `check-hooks-apply` and `check-useless-excludes`, declare no `stages:` and so run at every stage including this one.
|
||||
- `apm-audit-ci` runs `apm audit --ci` once per manifest — the root one and each of the six plugin packages — because the root-only invocation audits the marketplace manifest and **nothing else**, and `apm-pack-check-clean` does not parse plugin `dependencies:` blocks either (verified: a malformed one passes `apm pack --check-versions --check-clean --dry-run` and fails `apm audit --ci` in that package's directory). It verifies two things and claims no more: each `apm.yml` parses as a valid APM manifest, and any package declaring dependencies has a consistent `apm.lock.yaml`. It does **not** enforce an org policy — apm discovers one from the git remote and only understands github.com and Azure DevOps, so against this repo's self-hosted Gitea remote it prints `No org policy found at unknown; enforcement skipped`. Do **not** "fix" that with `policy.fetch_failure_default: block` in `apm.yml`: it was tested and rejected, because with no reachable policy source it makes the hook exit 1 on every push forever.
|
||||
- `check-apm-agents-valid` derives its expected agent-file set from `git ls-files` (same pattern as `tests/run-bats.sh`), so an agent file deleted from the worktree but still tracked fails the run, and discovering zero agent files is an error rather than a pass. An untracked new agent file is still validated — the derivation is one-directional on purpose, so uncommitted work is not blocked but also cannot bypass the gate. Agents take the ADR-0020 description gates (`agent-audit`'s `validate.sh` holds its own copy of those two constants) and, deliberately, **no** body word gate: an agent body becomes the system prompt of a fresh context rather than competing with the caller's live conversation, so the 900-word FAIL does not transfer. A bats test pins that absence — adding a body gate there contradicts the ADR rather than fixing an inconsistency.
|
||||
- `check-apm-agents-valid` derives its expected agent-file set from `git ls-files` (same pattern as `tests/run-bats.sh`), so an agent file deleted from the worktree but still tracked fails the run, and discovering zero agent files is an error rather than a pass. An untracked new agent file is still validated — the derivation is one-directional on purpose, so uncommitted work is not blocked but also cannot bypass the gate. Agents take the ADR-0020 description gates (`agent-audit`'s `validate.sh` holds its own copy of those two constants) and, deliberately, **no** body word gate: an agent body becomes the system prompt of a fresh context rather than competing with the caller's live conversation, so the 900-word FAIL does not transfer. A bats test pins that absence in `agent-audit`'s validator — adding a body gate there contradicts the ADR rather than fixing an inconsistency. Be precise about the scope of that guarantee, though: it holds for the **validator**, not for the shared script. `scripts/skill-size-check.sh` applies its body gate to whatever path it is handed, and `bash scripts/skill-size-check.sh plugins/*/.apm/agents/*.agent.md` exits 1 today with 900-word body FAILs on `git-orchestrate` (933), `gitea-orchestrate` (1,199) and `apm-orchestrate` (1,080). Agent files escape only because the hook definitions filter on `SKILL.md` — a file-pattern accident that happens to implement the design, not the design itself. Do not "extend" that hook's `files:` pattern to cover agents on the assumption that the script already knows the difference.
|
||||
- **Two** pre-push hooks need the network, for one shared reason: root `apm.yml`'s `marketplace.packages[]` contains exactly one remote entry (`mattpocock-skills`, `source: mattpocock/skills`), and resolving it needs a `git ls-remote`. `apm-marketplace-check` resolves every entry and is `always_run`, so it fails with `No cached refs (offline)`. `apm-pack-check-clean` (`apm pack --check-versions --check-clean --dry-run`) re-resolves the same entry and fails with `Error: Git network timeout during ls-remote`. Pinning the entry to an exact version does **not** remove the call — an exact pin still ls-remotes. `--offline` rescues neither. To push without a network, skip both using pre-commit's own mechanism: `SKIP=apm-marketplace-check,apm-pack-check-clean git push`. Skip those two alone — verified under `unshare -rn`, the other twelve pre-push hooks pass offline because they are real local checks (`check-executables-allow-sync` landed after that run, but reads two local manifests and makes no network call), and adding one of them to `SKIP` disarms it silently. `apm-audit-ci` calls `apm` too but stays local: its org-policy discovery resolves nothing on this remote before any network call, so it does not join the pair above.
|
||||
- Author commits with `git-commits` — it validates Conventional Commits (enforced at `commit-msg`) for you.
|
||||
|
||||
|
||||
@@ -27,13 +27,13 @@ A separate product (separate repo) for browsing, editing, and configuring AI dev
|
||||
Reusable slash commands for AI coding tools, defined as `SKILL.md` files following the [Agent Skills open standard](https://agentskills.io). Authored at `plugins/<plugin-name>/.apm/skills/<skill-name>/SKILL.md` and reaching a host by one of two install paths: `apm install`, which deploys the skill directory to `.claude/skills/<skill-name>/` (this repo's own path — see "apm-consumed install"), or `claude plugin install <name>@<marketplace>`, which caches the whole plugin (still supported for external consumers). Skills are self-contained — they cannot reference files outside the plugin directory after install-time caching. The two paths name skills differently: apm deploys a plain project skill (`skill-audit`), a plugin install namespaces it (`kyberforge:skill-audit`).
|
||||
|
||||
### Preload tax
|
||||
The always-on context cost of every installed skill's `name` + `description`, which sit in the agent's context from the first token of every session whether or not the skill is invoked. Measured 2026-08-14 at 23,612 chars (~6,200 tokens) across 39 skills, plus 1,325 chars for 4 agents. Non-routing frontmatter (`metadata.source_keys`, `category`, `version`) is **not** part of it — the model-visible skill listing carries only `name` and `description`, which supersedes `LESSONS.md:63` on this host. Bodies are not part of it either; they are charged on invocation.
|
||||
The always-on context cost of every installed skill's `name` + `description`, which sit in the agent's context from the first token of every session whether or not the skill is invoked. Measured 2026-08-14 against base commit `f9b919d` at 23,427 chars (~5,900 tokens) across 39 skills, plus 1,325 chars for 4 agents. Method, so it can be re-run: sum `len(name) + len(description)` over each `plugins/*/.apm/skills/*/SKILL.md` frontmatter with `>` block scalars folded to the value the host loads, at ~4 characters per token. Non-routing frontmatter (`metadata.source_keys`, `category`, `version`) is **not** part of it — the model-visible skill listing carries only `name` and `description`, which supersedes `LESSONS.md:63` on this host. Bodies are not part of it either; they are charged on invocation.
|
||||
|
||||
### Skill context contract
|
||||
The authoring rules that hold the preload tax and body size down, set by ADR-0020. A description carries a trigger clause, at most one capability clause, and a boundary clause of the form `Not <thing> → <skill-name>` naming a resolvable target — nothing else. Capability enumeration, output formats, and composition notes ("composes X rather than duplicating Y") belong in the body or `README.md`; a description that summarises workflow is a correctness hazard, not just a cost, because agents act on it instead of reading the body. Sizes are two-tier and sit *below* the agentskills.io spec limits, which stay unchanged as conformance backstops: description 250 SUGGESTION / 400 FAIL (spec 1,024); body 600 SUGGESTION / 900 FAIL (spec 2,770 words / 500 lines). Conflating the quality gate with the spec ceiling is what let `skill-author` and `agent-author` grow to within twelve words of 2,770.
|
||||
The authoring rules that hold the preload tax and body size down, set by ADR-0020. A description carries a trigger clause, at most one capability clause, and a boundary clause of the form `Not <thing> → <skill-name>` naming a resolvable target — nothing else. "Resolvable" is decided by walking up *from the file being checked* to an **authoring root** — the nearest ancestor holding `plugins/*/.apm/{skills,agents}`, falling back to the nearest `.git`, in two passes so a nested `.git` cannot outrank a real monorepo root. The universe is then every skill and agent under `<root>/plugins/*/` (sibling plugins resolve against each other, which is what a monorepo means), plus the checked file's own apm package and that package's own declared `dependencies.apm`. The **root** manifest's dependency list is never consulted, and no plugin here declares a cross-plugin apm dependency. Deployed `.claude/`/`.agents/` trees count only when there is no authoring root at all — the consumer case. The property this buys is that one commit gets one verdict: those trees are gitignored `apm install` output, so resolving through them made the same commit report 2 dangling targets on a developer machine and 6 on a fresh clone, which a gate shipping hot with no baseline cannot do. A `${BASH_SOURCE}`-relative repo root is the other half of the same defect and is gone — it leaked this repo's 39-skill universe into consumer repos running the hook through pre-commit. A *missing* boundary clause is a SUGGESTION rather than a failure, for skills and agents alike — some skills genuinely have no near-miss sibling. A *missing or empty description* is the opposite: a hard FAIL in all three validators, because a gate that merely declines to measure the one preloaded field reports green. Capability enumeration, output formats, and composition notes ("composes X rather than duplicating Y") belong in the body or `README.md`; a description that summarises workflow is a correctness hazard, not just a cost, because agents act on it instead of reading the body. Sizes are two-tier and sit *below* the agentskills.io spec limits, which stay unchanged as conformance backstops: description 250 SUGGESTION / 400 FAIL (spec 1,024); body 600 SUGGESTION / 900 FAIL (spec 2,770 words / 500 lines). Conflating the quality gate with the spec ceiling is what let `skill-author` and `agent-author` grow to within twelve words of 2,770.
|
||||
|
||||
### Dispatch body
|
||||
The body pattern a skill with two or more mutually exclusive flows must use: the body carries only the dispatch table and the gates common to every branch, and each flow lives in its own self-contained `references/` file. Named for `apm-workflow` (554-word body, 3,006 words of references), which arrived at it independently and is the repo's exemplar. Its absence is the characteristic defect — `skill-author` inlines both its create and improve flows, and `agent-author` carries 50-60 lines marked inapplicable by their own headers on any single run.
|
||||
The body pattern a skill with two or more mutually exclusive flows must use: the body carries only the dispatch table and the gates common to every branch, and each flow lives in its own self-contained `references/` file. Named for `apm-workflow` (421-word body, 3,006 words of references), which arrived at it independently and is the repo's exemplar. Its absence was the characteristic defect at the time ADR-0020 was written: `skill-author` inlined both its create and improve flows, and `agent-author` carried 50-60 lines marked inapplicable by their own headers on any single run. Both were retrofitted to dispatch tables in the change that carries the ADR — `skill-author` went 2,623 body words to 595 and `agent-author` 2,582 to 616 — so they are now worked examples of the pattern rather than counter-examples of it. The 39-skill corpus at large is not: 9 bodies still exceed the 900-word FAIL (issue #99).
|
||||
|
||||
### Hand-invoked skill
|
||||
A skill reached only by typing its slash command, declared with `disable-model-invocation: true`. The host withholds it from the model-visible skill listing entirely, so it pays no preload tax and its `description` becomes human-facing text rather than a trigger list. `zoom-out` is the worked example: apm passes the flag through verbatim to both install paths, and the skill is absent from the router while `/zoom-out` still works. Choosing model-invoked vs. hand-invoked is the first question `skill-author` asks, because it determines whether a description needs triggers at all.
|
||||
@@ -100,7 +100,7 @@ Wiring Vale as a deterministic prefilter for `skill-audit`/`agent-audit`'s Descr
|
||||
|
||||
Both skills' Step 1, and the `vale-audit-prefilter-skill`/`-agent` pre-commit hooks, call each copy's own `scripts/vale-wrap.sh` rather than `vale` directly — a workaround for a confirmed Vale 3.15.2 limitation (see `vale-config`'s Gotchas): `text.frontmatter.description` silently stops matching on most — not all — multi-line descriptions. Verified by reproduction, not assumed: `>` folded scalars, plain (unquoted) continuation lines, and single- or double-quoted multi-line scalars all yield 0 alerts and exit 0 on a deliberately-bad fixture, while a `|` literal block spanning the same 2+ lines lints normally (alerts fire, exit 1). The wrapper flattens those three broken forms to one physical line in a scratch copy (padding with blank lines so every other line number is unchanged) before handing off to real `vale`; `|` literal blocks and single-line descriptions pass through untouched, already linting correctly. The plain and quoted forms previously passed silently — unflattened and unmatched — so a bad description in either sailed through the prefilter. Handed no `--config` at all, the wrapper falls back to its own sibling `assets/vale/.vale.ini`, located from `${BASH_SOURCE[0]}` rather than from the cwd — which is why both manifests' `entry:` is now the bare script path with no argument after it. pre-commit prefixes only `entry[0]` with the hook-repo clone path (`cmd = (prefix.path(cmd[0]), *cmd[1:])`), so every later argument resolves against the *consuming* repo's root: a `--config` in `.pre-commit-hooks.yaml` pointed at a path no consumer has and hard-failed every external run with `E100 [--config] Runtime error`. `.pre-commit-config.yaml` drops the argument too, deliberately keeping the two entries identical — the local `repo: local` hook resolved its `--config` correctly only because the consuming repo *was* this repo, and that divergence is why three review rounds exercised a path no external consumer takes and missed the defect. An explicit `--config` still wins, in all three argv forms (`--config X`, `--config=/abs`, `--config=rel`), and a relative one still resolves against the caller's cwd, matching bare `vale`, not the repo root. Both audit skills' Step 1 now passes no `--config` either: it resolves the script relative to the skill's own directory so the call works from an installed plugin cache, but a relative `--config` alongside it would still resolve against the cwd, yielding `E100 Runtime error ... does not exist` and exit 2 — which both skills' fallback misreads as "vale unavailable" and silently downgrades to full LLM judgment. `tests/test-vale-wrap.sh` regression-tests this against skill-audit's copy specifically (its fixtures are all `SKILL.md`-shaped, and only skill-audit's `.vale.ini` has that glob section). Each `.vale.ini`'s section globs are path-agnostic (`[**/SKILL.md]` for skill-audit's copy; `[**/agents/*.md]`/`[**/*.agent.md]` for agent-audit's) and do no scoping on their own: Vale's `*` crosses `/`. Scoping comes from each pre-commit hook's own `files:` regex and from the audit skills passing one explicit file per invocation. The two manifests scope differently on purpose: this repo's `.pre-commit-config.yaml` pins its own layout — `^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$` for `-skill`, `^plugins/[^/]+/\.apm/agents/[^/]+\.agent\.md$` for `-agent` — while the shipped `.pre-commit-hooks.yaml` stays layout-agnostic for external consumers whose skills live anywhere, using `(^|/)SKILL\.md$` and `(^|/)agents/[^/]+\.md$|\.agent\.md$`. Both manifests split the prefilter into two hooks precisely because one combined hook pointed at only one copy would silently 0-file-skip the other file type. A `SKILL.md` outside `plugins/` (e.g. project-scope `.claude/skills/foo/SKILL.md`) still matches `[**/SKILL.md]` and gets linted normally — the globs constrain filename shape, not location. Vale reports 0 files only when the path it is handed matches no glob section at all: a differently-named file, or a directory argument holding nothing that matches. That run prints `✔ 0 errors ... in 0 files.` and exits 0, indistinguishable from a clean pass, so both audits treat a 0-file Vale run as NOT RUN and fall back to full LLM judgment.
|
||||
|
||||
This scope expands per ADR-0013: one cherry-picked low-noise `write-good`/`alex` rule landed in `styles/Kyberforge`, `Kyberforge.SentenceOpenerThereIs` (22 held-out hits, both in-corpus hits clean rewrites, zero suppressions). A second, `Kyberforge.VagueQualifier`, was cherry-picked and then deleted: 2 hits across the 41 skill/agent files, one marginal and one an unfixable false positive (`caveman/SKILL.md` quotes `of course` as an example of filler — a mention, not a use) that forced the repo's only Vale suppression comments. Also new is a sibling pre-commit hook, `skill-size-check` (`scripts/skill-size-check.sh`), enforcing agentskills.io's `SKILL.md` ceiling as two blocking gates: `MAX_LINES=500` and `MAX_WORDS=2770` (a word-count proxy for the 5,000-token limit, calibrated to the densest prose measured in this repo — 1.81 tokens per word — so even a worst-case `SKILL.md` at the ceiling stays under 5,000 tokens). Both are inclusive, and `skill-audit/scripts/validate.sh` checks the same pair on the same terms, so a `SKILL.md` can no longer pass its own audit yet be blocked by the commit hook. Scoped to `^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$` only, same as `vale-audit-prefilter-skill`, so it never lints `docs/research/examples/` reference skills. It's also exposed in the root-level `.pre-commit-hooks.yaml` as `kyberforge-skill-size-check` — it has no external asset dependency, so it needed no relocation, only exposure to external consumers. File scope (`SKILL.md` + agent files) and enforcement model (rules land directly in `styles/Kyberforge`, blocking immediately, no trial tier) stay unchanged; governance.md/CONTROLS.md were evaluated and excluded as rule sources (nothing prose-pattern-matchable to mine). House convention: banned phrasing that must be mentioned rather than used goes in backticks or a fenced code block — Vale skips code spans and fences, so no suppression is needed; inline `<!-- vale Rule = NO -->` (HTML-comment form; the MDX `{/* */}` form does not work in plain Markdown) is the fallback only where backticking is impossible.
|
||||
This scope expands per ADR-0013: one cherry-picked low-noise `write-good`/`alex` rule landed in `styles/Kyberforge`, `Kyberforge.SentenceOpenerThereIs` (22 held-out hits, both in-corpus hits clean rewrites, zero suppressions). A second, `Kyberforge.VagueQualifier`, was cherry-picked and then deleted: 2 hits across the skill/agent corpus as it stood at the time of that measurement (2026-08-08, before the `.apm/` restructure), one marginal and one an unfixable false positive (`caveman/SKILL.md` quotes `of course` as an example of filler — a mention, not a use) that forced the repo's only Vale suppression comments. A third, `Kyberforge.CompositionNote`, landed with ADR-0020 and bans architecture and composition prose from a description; it is `level: error` like the rest, and it currently fires 10 times across `gitea-issues`, `gitea-labels-milestones`, `gitea-prs` and `gitea-workflow`, so `pre-commit run --all-files` is red on prose as well as on size until issue #99 lands. Also new is a sibling pre-commit hook, `skill-size-check` (`scripts/skill-size-check.sh`), which carries **two independent gate families that must not be conflated** (see "Skill context contract"). The agentskills.io spec backstop is `MAX_LINES=500` and `MAX_WORDS=2770`, both inclusive and both counting the **whole file including frontmatter** (2,770 is a word-count proxy for the 5,000-token limit, calibrated to the densest prose measured in this repo — 1.81 tokens per word — so even a worst-case `SKILL.md` at the ceiling stays under 5,000 tokens; it is not a percentile of the corpus). ADR-0020 adds a context budget measured differently: description characters 250 SUGGESTION / 400 FAIL, **body-only** words 600 SUGGESTION / 900 FAIL, plus deterministic checks that every boundary routing target resolves, that a body's named `references/<file>.md` all exist, and — SUGGESTION-tier — that a boundary clause is present at all, that `## Gotchas` holds at most five entries, and that it stays under 25% of the body. `skill-audit/scripts/validate.sh` and `agent-audit/scripts/validate.sh` hold their own copies of the shared constants and `tests/test-skill-size-check.sh` asserts the copies agree, so a `SKILL.md` can no longer pass its own audit yet be blocked by the commit hook. Agents take the description gates and no body word gate. `python3` **and PyYAML** are hard requirements — the earlier hand-rolled frontmatter fallback is gone, because a fallback that silently mis-parses a scalar shape reports a vacuous pass. Scoped to `^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$` only, same as `vale-audit-prefilter-skill`, so it never lints `docs/research/examples/` reference skills. It's also exposed in the root-level `.pre-commit-hooks.yaml` as `kyberforge-skill-size-check` — it has no external asset dependency, so it needed no relocation, only exposure to external consumers. File scope (`SKILL.md` + agent files) and enforcement model (rules land directly in `styles/Kyberforge`, blocking immediately, no trial tier) stay unchanged; governance.md/CONTROLS.md were evaluated and excluded as rule sources (nothing prose-pattern-matchable to mine). House convention: banned phrasing that must be mentioned rather than used goes in backticks or a fenced code block — Vale skips code spans and fences, so no suppression is needed; inline `<!-- vale Rule = NO -->` (HTML-comment form; the MDX `{/* */}` form does not work in plain Markdown) is the fallback only where backticking is impossible.
|
||||
|
||||
### LESSONS.md
|
||||
Long-loop feedback log for patterns observed across sessions. Three or more entries on the same pattern graduate to the relevant standing file (e.g. a coding convention, a governance rule). Updated by the session-handoff skill or directly by the human. Lives at the repo root.
|
||||
|
||||
53
LESSONS.md
53
LESSONS.md
@@ -194,7 +194,11 @@ is the broken multi-`raw:` form, which `tests/test-vale-hooks-consumer.sh` now f
|
||||
## 2026-08-14 — Un-anchoring a description rule to reach mid-sentence text is unshippable
|
||||
|
||||
Widening `DescriptionOpener` to catch `gitea-workflow`'s mid-description "This is the human-facing
|
||||
entry point…" looked like a one-character change. Under `scope: text.frontmatter.description`, `^`
|
||||
entry point…" looked like a one-character change. Both that skill and `gitea-labels-milestones`
|
||||
*open* with "Use when…" and satisfy the opener rule; the offending clause sits at character 377 and
|
||||
300 of the folded value respectively, so the rule was never violated and never silently passed — it
|
||||
simply had no jurisdiction, which is a different defect and takes a different fix.
|
||||
Under `scope: text.frontmatter.description`, `^`
|
||||
anchors to the start of the whole description value — and `vale-wrap.sh` has already flattened that
|
||||
value to one physical line, so `(?m)` changes nothing. Un-anchoring is therefore the only route to
|
||||
mid-description text, and measured across the corpus it scores 5 hits and 5 false positives: skills
|
||||
@@ -206,21 +210,38 @@ a new case does not fit.
|
||||
|
||||
## 2026-08-14 — A formatter in the commit path manufactures drift on a file with a clean git diff
|
||||
|
||||
`apm audit --ci` failed for weeks on `.claude/settings.json` while `git diff` on that file was empty —
|
||||
the worst possible pairing of signals, because the file matched HEAD exactly and every instinct says
|
||||
"nothing changed here". The content was identical to apm's output to the byte; only the JSON key
|
||||
order differed. `pretty-format-json --autofix` sorts object keys unless `--no-sort-keys` is passed,
|
||||
and its `exclude:` listed fifteen generated manifests but not this file, so from the commit that
|
||||
first wrote a hook entry there (`2e395a4`) onward, apm's insertion-ordered output was silently
|
||||
re-sorted on the way in. apm then replayed the install, produced its own order, and reported drift
|
||||
against a file no human had touched.
|
||||
`apm audit --ci` failed on `.claude/settings.json` while `git diff` on that file was empty — the worst
|
||||
possible pairing of signals, because the file matched HEAD exactly and every instinct says "nothing
|
||||
changed here". The content was identical to apm's output to the byte; only the JSON key order
|
||||
differed. `pretty-format-json --autofix` sorts object keys unless `--no-sort-keys` is passed, and its
|
||||
`exclude:` listed fifteen generated manifests but not this file, so from the commit that first wrote
|
||||
a hook entry there onward, apm's insertion-ordered output was silently re-sorted on the way in. apm
|
||||
then replayed the install, produced its own order, and reported drift against a file no human had
|
||||
touched.
|
||||
|
||||
Two general points. First, a tool-owned generated file that passes through an autofixing formatter is
|
||||
drifted by construction, and the diff that would reveal it never appears in `git diff` — it only
|
||||
The provenance matters as much as the mechanism, and the first account of this entry got it wrong in
|
||||
both directions. `git log --format='%h %ad %s' --date=iso` puts the introducing commit `2e395a4` at
|
||||
2026-08-14 18:47 and the fix `7607522` at 21:54 — roughly three hours, not "weeks". And `2e395a4` is
|
||||
the **first commit of the `refactor/trim-skills-agents-context` branch**, eleven minutes after the
|
||||
base merge `f9b919d`; `git branch -a --contains 2e395a4` returns only that branch and its own
|
||||
`remotes/origin/` tracking copy — two lines naming one branch, and `main` is not among them. So
|
||||
this was not a latent defect inherited from `main`, it was manufactured inside the same PR that
|
||||
diagnosed it, and the fixing commit's own message calling it "pre-existing … red at HEAD before
|
||||
ADR-0020 work began" is the mis-attribution rather than the record. Two cheap commands would have
|
||||
settled it before either sentence was written.
|
||||
|
||||
Three general points. First, a tool-owned generated file that passes through an autofixing formatter
|
||||
is drifted by construction, and the diff that would reveal it never appears in `git diff` — it only
|
||||
exists between the formatter's input and its output, which nothing stores. Second, the fix is
|
||||
self-undoing unless the exclude lands in the same commit: correcting the file alone means the hook
|
||||
re-breaks it as it is staged. Fix: when a tool declares ownership of a path, add that path to every
|
||||
autofixing hook's `exclude` at the moment ownership is declared, not when the drift is noticed. This
|
||||
repo gates marketplace-mirror, plugin-content and vale-style drift deterministically and has no
|
||||
equivalent gate asserting tool-owned paths stay out of formatter scope — `.claude/settings.json` was
|
||||
the sixteenth exclude and nothing prevents a seventeenth.
|
||||
re-breaks it as it is staged. Third — the one this entry had to learn twice — "pre-existing" is a
|
||||
claim about history, and history is queryable; a defect found while working on a branch feels
|
||||
inherited, and the feeling is not evidence. A three-hour-old self-inflicted bug and a months-old
|
||||
inherited one call for different responses, and writing the wrong one down converts a process failure
|
||||
into a story about someone else's neglect. Fix: when a tool declares ownership of a path, add that
|
||||
path to every autofixing hook's `exclude` at the moment ownership is declared, not when the drift is
|
||||
noticed — and before describing any defect as pre-existing, run `git log -S` or
|
||||
`git branch --contains` on the commit that introduced it. This repo gates marketplace-mirror,
|
||||
plugin-content and vale-style drift deterministically and has no equivalent gate asserting tool-owned
|
||||
paths stay out of formatter scope — `.claude/settings.json` was the sixteenth exclude and nothing
|
||||
prevents a seventeenth.
|
||||
|
||||
8
apm.yml
8
apm.yml
@@ -1,5 +1,5 @@
|
||||
name: holocron
|
||||
version: 0.4.1
|
||||
version: 0.4.2
|
||||
description: AI development skills for Claude Code and GitHub Copilot CLI — factory, design, implement, review, and cross-cutting workflows.
|
||||
license: MIT
|
||||
|
||||
@@ -42,7 +42,7 @@ dependencies:
|
||||
# after a kyberforge release, check this first.
|
||||
executables:
|
||||
allow:
|
||||
kyberforge#1.5.0:
|
||||
kyberforge#1.6.0:
|
||||
hooks: true
|
||||
bin: true
|
||||
|
||||
@@ -52,7 +52,7 @@ marketplace:
|
||||
# top-level apm.yml description:/version: above are NOT inherited into the
|
||||
# compiled output despite being used elsewhere (e.g. by `apm audit`).
|
||||
description: AI development skills for Claude Code and GitHub Copilot CLI — factory, design, implement, review, and cross-cutting workflows.
|
||||
version: 0.4.1
|
||||
version: 0.4.2
|
||||
owner:
|
||||
name: Defame1297
|
||||
email: defame1297@rkdr.net
|
||||
@@ -79,7 +79,7 @@ marketplace:
|
||||
- name: kyberforge
|
||||
description: Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.
|
||||
source: ./plugins/kyberforge
|
||||
version: 1.5.0
|
||||
version: 1.6.0
|
||||
category: Developer Tools
|
||||
|
||||
- name: bin
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
Every installed skill's `name` and `description` sits in every agent's context from the first token
|
||||
of every session, whether or not the skill is ever invoked. Across this repo's 39 skills that is
|
||||
23,612 characters — roughly 6,200 tokens — and the authoring rules that produced it optimised for
|
||||
23,427 characters — roughly 5,900 tokens — and the authoring rules that produced it optimised for
|
||||
triggering reliability with no counter-pressure on size. This ADR sets the budget, the shape, and the
|
||||
gates that hold them.
|
||||
|
||||
@@ -10,14 +10,32 @@ gates that hold them.
|
||||
|
||||
## Context
|
||||
|
||||
Measured before any change:
|
||||
Every `file:line` citation in this ADR is against the base commit the decision was taken on,
|
||||
`f9b919d7e3bd5e6b51fbdf88b32ace0438b313e0`, not against current `HEAD`. The change that carries this
|
||||
ADR rewrites several of the cited files, so a citation resolved against the worktree will land on
|
||||
unrelated text. Use `git show f9b919d:<path>` to follow one.
|
||||
|
||||
Measured before any change, at that commit. Method, so the figures are reproducible: sum
|
||||
`len(name) + len(description)` over the frontmatter of every `plugins/*/.apm/skills/*/SKILL.md`,
|
||||
folding `>` block scalars to the value the host actually loads (most descriptions here are folded
|
||||
scalars, so counting raw lines measures indentation instead); tokens at the standard
|
||||
~4-characters-per-token approximation `scripts/skill-size-check.sh` uses. Word counts are
|
||||
whitespace-separated tokens, and are stated as **body-only** or **whole-file** every time, never bare.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| 39 skill `name` + `description` | 23,612 chars, ~6,200 tokens, **preloaded every session** |
|
||||
| 4 agent `name` + `description` | 1,325 chars, ~350 tokens, preloaded every session |
|
||||
| skill bodies | median 684 words, mean 815, p90 1,349 |
|
||||
| `MAX_WORDS` gate (`skill-audit/scripts/validate.sh:147`) | **2,770** — 2× p90 |
|
||||
| 39 skill `name` + `description` | 23,427 chars, ~5,900 tokens, **preloaded every session** |
|
||||
| 4 agent `name` + `description` | 1,325 chars, ~330 tokens, preloaded every session |
|
||||
| skill bodies (body-only words) | median 684, mean 815, p90 1,349 |
|
||||
| skill files (whole-file words) | median 816, mean 927, p90 1,526 |
|
||||
| `MAX_WORDS` gate (`skill-audit/scripts/validate.sh:147`) | **2,770** whole-file — a density proxy, not a percentile |
|
||||
|
||||
That last row is worth stating plainly, because it is the first thing this ADR is about. 2,770 is not
|
||||
derived from the corpus distribution at all: per the derivation comment in
|
||||
`scripts/skill-size-check.sh`, it is 2,770 words at the densest observed 7.22 chars/word ≈ 20,000
|
||||
chars ≈ the agentskills.io ~5,000-token ceiling. Neither percentile reaches it — 2× the body-only p90
|
||||
is 2,698 and 2× the whole-file p90 is 3,052 — and reading it as "2× p90" would pair a whole-file gate
|
||||
against a body-only distribution, which is exactly the conflation this ADR exists to stop.
|
||||
|
||||
Three findings drove this, none of which is "the descriptions drifted".
|
||||
|
||||
@@ -32,7 +50,8 @@ with six capability clusters. Across the twelve longest descriptions, 30.7% is c
|
||||
enumeration and 11.6% is composition or implementation detail that cannot affect a routing decision.
|
||||
|
||||
**Capability enumeration in a description is a correctness hazard, not only a token cost.**
|
||||
`docs/research/examples/skill-write/writing-skills/SKILL.md:154-158` reports a measured failure: "when
|
||||
`plugins/kyberforge/docs/research/examples/skill-write/writing-skills/SKILL.md:154-158` reports a
|
||||
measured failure: "when
|
||||
a description summarizes the skill's workflow, an agent may follow the description instead of reading
|
||||
the full skill content. A description saying 'code review between tasks' caused an agent to do ONE
|
||||
review, even though the skill's flowchart clearly showed TWO reviews." `git-commits` is exactly that
|
||||
@@ -41,21 +60,25 @@ chars, lowercase subject, no trailing periods, 11 standard types`) an agent can
|
||||
loading the body.
|
||||
|
||||
**The upstream sources cannot settle this.** The four skill-writing references under
|
||||
`docs/research/examples/skill-write/` disagree on what a description contains — when-only
|
||||
(`writing-skills/SKILL.md:99`), what-and-when (`skill-creator/SKILL.md:67`,
|
||||
`anthropic-best-practices.md:187`), triggers-only (`writing-great-skills/SKILL.md:28`), and
|
||||
what-plus-when-plus-negative (`write-skill/SKILL-TEMPLATE.md:5-6`). `writing-skills` and the Anthropic
|
||||
document it bundles contradict each other inside one skill directory. They also disagree on whether
|
||||
`plugins/kyberforge/docs/research/examples/skill-write/` disagree on what a description contains —
|
||||
when-only (`writing-skills/SKILL.md:99`), what-and-when (`skill-creator/SKILL.md:67`,
|
||||
`writing-skills/anthropic-best-practices.md:187`), triggers-only
|
||||
(`writing-great-skills/SKILL.md:28`), and what-plus-when-plus-negative
|
||||
(`write-skill/SKILL-TEMPLATE.md:5-6`). Those four paths are relative to that directory.
|
||||
`writing-skills` and the Anthropic document it bundles contradict each other inside one skill
|
||||
directory. They also disagree on whether
|
||||
500 lines is binding, on the inline-versus-bundle threshold, and on the TOC threshold (>100 lines vs
|
||||
>300 lines). "Grounded in the research" is therefore not available as a tiebreaker; a house choice is
|
||||
required and this is it.
|
||||
|
||||
A fourth observation shaped the body half. The best progressive-disclosure ratio in the repo belongs
|
||||
to `apm-workflow` — a 554-word body dispatching to 3,006 words of references — and the worst two
|
||||
belong to the skills that define the house standard: `skill-author` (2,760 body / 1,247 references)
|
||||
and `agent-author` (2,758 / 1,664). Both sit within twelve words of the 2,770 gate their own plugin
|
||||
enforces. A ceiling that nothing approaches is not a constraint; a ceiling that two files have grown
|
||||
into is a target.
|
||||
to `apm-workflow` — a 421-word body dispatching to 3,006 words of references — and the worst two
|
||||
belong to the skills that define the house standard: `skill-author` (2,623-word body / 1,247 words of
|
||||
references) and `agent-author` (2,582 / 1,664). Measured the other way, whole-file, those two are
|
||||
2,760 and 2,758 words — ten and twelve words under the 2,770 gate their own plugin enforces. A
|
||||
ceiling that nothing approaches is not a constraint; a ceiling that two files have grown into is a
|
||||
target. The two numbers for one file are the point: 2,623 and 2,760 describe the same `skill-author`,
|
||||
and only one of them is what either gate measures.
|
||||
|
||||
## Decision
|
||||
|
||||
@@ -69,8 +92,39 @@ clause**, and a **boundary clause**. Capability enumeration, output-format detai
|
||||
- **250 characters SUGGESTION, 400 FAIL.** The agentskills.io 1,024-character limit remains as an
|
||||
unchanged spec backstop. The SUGGESTION tier is what moves the average; the FAIL tier only stops
|
||||
outliers.
|
||||
- **A missing, valueless or `null` `description:` is a hard FAIL** in all three validators. That
|
||||
reads as a trivial precondition and is not: a `description:` line with no value followed by
|
||||
`model: sonnet` let a line regex capture the *next* key, which looked non-empty, so the "missing or
|
||||
empty" branch never fired and every gate below it then early-returned on the genuinely empty folded
|
||||
value — exit 0, zero output, on a blocking pre-push gate. Presence is decided on the YAML-folded
|
||||
value and nowhere else. The field this contract is entirely about is the one field a gate must
|
||||
never fail to notice is absent.
|
||||
- **Boundary clauses compress** to `Not <thing> → <skill-name>.` and must name a target that
|
||||
resolves to a real skill under `plugins/*/.apm/skills/`. This is checked deterministically.
|
||||
resolves to a real skill or agent. Resolution walks up **from the file being checked** to an
|
||||
*authoring root* — the nearest ancestor holding `plugins/*/.apm/skills` or `plugins/*/.apm/agents`,
|
||||
falling back to the nearest ancestor holding `.git`. Two passes rather than one interleaved walk,
|
||||
so a nested `.git` (a submodule, a sub-package worktree) cannot beat a real monorepo root further
|
||||
up. When an authoring root is found the universe is every skill and agent under
|
||||
`<root>/plugins/*/`, plus the target's own apm package and the packages that package declares in
|
||||
its own `apm.yml` `dependencies.apm`. Sibling plugins resolve against each other, which is what a
|
||||
monorepo means. Deployed `.claude/`/`.agents/` trees are consulted **only** when no authoring root
|
||||
exists — the consumer case, where there is no monorepo to read. What the resolver must never do is
|
||||
derive the universe from its own location: a `${BASH_SOURCE}`-relative repo root leaked this repo's
|
||||
39-skill universe into every consumer repo running the hook through pre-commit, so a consumer skill
|
||||
routing to `skill-audit` resolved against a plugin it had never installed. Checked
|
||||
deterministically. A description carrying **no** boundary clause at all is a SUGGESTION, for skills
|
||||
and agents alike: most descriptions want one, some genuinely have no near-miss sibling to exclude,
|
||||
and that judgment is not a script's to make.
|
||||
- **The verdict must not depend on whether `apm install` has been run.** Deployed trees are
|
||||
gitignored install output, present only on a machine that has run it. Four cross-plugin targets
|
||||
here (`gitea-branches` → `git-branches`, `gitea-branches` → `git-history`, `gitea-issues` →
|
||||
`git-branches`, `gitea-workflow` → `git-workflow`) once resolved through `.claude/skills/` alone,
|
||||
so the same commit measured 2 dangling targets on a developer machine and 6 on a fresh clone. A
|
||||
gate shipping hot with no baseline cannot give two answers. Under the walk-up those four resolve
|
||||
because sibling plugins are in the universe — no plugin here declares a cross-plugin apm
|
||||
dependency, and none needs to. Verified: a tree holding only `plugins/` and the root `apm.yml`,
|
||||
with no `.claude/` or `.agents/` anywhere, now produces findings identical to the working tree —
|
||||
26 description FAILs, 9 body FAILs, 2 dangling targets, 0 missing references, 58 SUGGESTIONs.
|
||||
- **The blanket pushiness rules are deleted.** `skill-author/SKILL.md:104` and
|
||||
`description-quality.md:21` are replaced by a conditional: add an indirect trigger only where the
|
||||
user's natural phrasing genuinely omits the domain word — true for the `gitea-*` family, false for
|
||||
@@ -89,11 +143,16 @@ prose move to `references/` behind an explicit "read X when Y" trigger.
|
||||
current state.
|
||||
- **Dispatch is mandatory at two or more mutually exclusive flows.** The body carries the dispatch
|
||||
table and the gates that apply to every branch; each flow lives in its own self-contained
|
||||
`references/` file. This is `apm-workflow/SKILL.md:33-41` promoted from accident to rule.
|
||||
`references/` file. This is `apm-workflow/SKILL.md:33-41` promoted from accident to rule. "Two
|
||||
mutually exclusive flows" is not decidable from file text, so this rule is auditor judgment — see
|
||||
Enforcement below for what that means and does not mean.
|
||||
- **Every `references/<file>.md` a body names must exist.** A dispatch table pointing at a file that
|
||||
was never written is a silently dead branch. Checked deterministically.
|
||||
- **Gotchas are constrained.** A Gotcha must state a fact that contradicts a reasonable default —
|
||||
something the agent gets wrong by acting sensibly. Maximum five entries. A Gotcha that paraphrases
|
||||
a step in the body below it is a FAIL. A Gotchas section exceeding 25% of the body is a
|
||||
SUGGESTION.
|
||||
something the agent gets wrong by acting sensibly. More than five entries is a SUGGESTION, as is a
|
||||
Gotchas section exceeding 25% of the body; both are countable and both are checked
|
||||
deterministically. A Gotcha that paraphrases a step in the body below it is a FAIL, but a FAIL an
|
||||
auditor issues, not a script — semantic equivalence is not pattern-matchable.
|
||||
|
||||
### Agents
|
||||
|
||||
@@ -101,6 +160,14 @@ Agents take the same description gates — they are preloaded identically — an
|
||||
A skill body is loaded into the caller's context, competing with the live conversation; an agent body
|
||||
becomes the system prompt of a fresh context. The rationale for the 900-word FAIL does not transfer.
|
||||
|
||||
That exemption is expressed in `agent-audit/scripts/validate.sh`, which has no body constant, and in
|
||||
the `files:` pattern of the `skill-size-check` pre-commit hook, which is `SKILL.md`-only. It is *not*
|
||||
expressed in `scripts/skill-size-check.sh` itself, which measures whatever path it is handed —
|
||||
running it directly over `plugins/*/.apm/agents/*.agent.md` today reports 900-word body FAILs on
|
||||
`git-orchestrate` (933), `gitea-orchestrate` (1,199) and `apm-orchestrate` (1,080). Agents escape by
|
||||
file pattern, not by the script knowing the difference. Anyone widening that pattern to cover agents
|
||||
would silently enforce a gate this ADR declines to set.
|
||||
|
||||
A plugin-scope agent is a single file with no sibling `references/` directory, so it cannot disclose
|
||||
to itself — it can only delegate to skills. `agent-audit` therefore gains a **delegation check**: an
|
||||
agent body that restates a procedure owned by a skill it can invoke is a FAIL, with the fix being
|
||||
@@ -126,26 +193,84 @@ type of input they take should be **one skill with a dispatch table**. This catc
|
||||
one-or-two-file agent pair, per ADR-0005 and ADR-0016) and their overlap is in the improve flow
|
||||
rather than the core job.
|
||||
|
||||
**DEFERRED — not implemented in the change that carries this ADR. Tracked as issue #101.** Both
|
||||
skills still exist separately, and this change made the split deeper rather than shallower: retrofit
|
||||
to the dispatch pattern took `skill-audit` from 3 reference files to 7 and `agent-audit` from 4 to 8,
|
||||
and their two same-named `references/description-quality.md` files now differ on 100 of ~120 lines
|
||||
after normalising `skill`/`agent`, where before they were closer. The merge stays the decision; it
|
||||
reopens ADR-0008 (agent-audit's single-file invocation contract) and touches every call site in
|
||||
`skill-author`, `agent-author` and `forge`, which is why it is its own change and not a rider on
|
||||
this one. Recorded here rather than dropped, so the gap between the rule and the tree is deliberate
|
||||
and dated instead of discovered later.
|
||||
|
||||
### Enforcement and rollout
|
||||
|
||||
Gates land where the existing gates already live — no new layer:
|
||||
Gates land where the existing gates already live — no new layer. The table below is exhaustive about
|
||||
which tier each rule is in, because the failure this ADR is most exposed to is a rule filed under
|
||||
"Enforcement" that no validator implements:
|
||||
|
||||
| Check | Home |
|
||||
|---|---|
|
||||
| description and body counts, resolvable boundary targets | `skill-audit/scripts/validate.sh`, `scripts/skill-size-check.sh` |
|
||||
| prose patterns (composition-note openers, restatement) | `plugins/kyberforge/.apm/skills/*/assets/vale/styles/Kyberforge/` |
|
||||
| judgment calls | `references/description-quality.md`, `references/body-discipline.md` |
|
||||
| Check | Applies to | Tier | Home |
|
||||
|---|---|---|---|
|
||||
| description characters (250 SUGGESTION / 400 FAIL) | skills, agents | deterministic | `scripts/skill-size-check.sh`; constants mirrored in `skill-audit/scripts/validate.sh` and `agent-audit/scripts/validate.sh` |
|
||||
| body-only words (600 SUGGESTION / 900 FAIL) | skills | deterministic | `skill-size-check.sh`, `skill-audit/scripts/validate.sh` |
|
||||
| description present and non-empty (ERROR) | skills, agents | deterministic | same |
|
||||
| boundary target resolves to a real skill or agent (ERROR when written as `/name` or `-> name`, or when its own sentence names another target that resolves; SUGGESTION otherwise) | skills, agents | deterministic | same |
|
||||
| boundary clause absent (SUGGESTION) | skills, agents | deterministic | same |
|
||||
| Gotchas entry count over five (SUGGESTION) | skills | deterministic | same |
|
||||
| Gotchas over 25% of the body (SUGGESTION) | skills | deterministic | same |
|
||||
| every `references/<file>.md` a body names exists (ERROR) | skills | deterministic | same |
|
||||
| description opener, composition notes in a description | skills, agents | prose pattern | `plugins/kyberforge/.apm/skills/*/assets/vale/styles/Kyberforge/` |
|
||||
| a Gotcha paraphrasing a body step | skills | **auditor judgment** | `references/body-discipline.md` |
|
||||
| dispatch at two or more mutually exclusive flows | skills | **auditor judgment** | `references/body-discipline.md` |
|
||||
| delegation: an agent body restating a skill's procedure | agents | **auditor judgment** | `agent-audit` |
|
||||
| capability enumeration, restatement, trigger quality | skills, agents | **auditor judgment** | `references/description-quality.md` |
|
||||
|
||||
**Blocking immediately, with no baseline file.**
|
||||
The rows in bold are stated as FAILs in the Decision above and are FAILs an *auditor* issues. None of
|
||||
them is countable: "does this Gotcha paraphrase step 4", "are these two flows mutually exclusive" and
|
||||
"does this agent body restate what `git-commits` already owns" are semantic questions, and a script
|
||||
that guessed at them would be a worse gate than no gate, because it would be believed. They are not
|
||||
enforced, they are reviewed, and this table exists so that distinction is written down rather than
|
||||
inferred from whether a validator happens to have been written yet.
|
||||
|
||||
Two of the deterministic rows are tuned for **false positives over recall**, and what they decline to
|
||||
see is part of the contract. On target extraction: a bare hyphenated name counts only inside a
|
||||
boundary sentence, and a single-word name is never matchable bare — `research`, `triage`, `forge`,
|
||||
`prototype` and `tdd` are all real skill names *and* ordinary English, so it must be written
|
||||
`` `forge` `` or `/forge` to be seen at all. Grammar then decides whether a recognised target may
|
||||
raise an error: one followed by an ordinary lowercase noun is a compound **modifier**, not a route
|
||||
("use pre-commit hooks instead of ad-hoc scripts", "invoke the pull-request template"), so it is
|
||||
confirm-only — it still resolves and still counts as a route when the name exists, but it can never
|
||||
dangle. Only a *terminal* target can. The compressed arrow form `→ <name>` is exempt from that
|
||||
follower test and is always error-eligible, because nothing reads as a compound modifier after an
|
||||
arrow; a `/slash` target reached through a route verb is **not** exempt and takes the same test. The
|
||||
simpler rule — "only marked targets may dangle" — was available and would have been wrong here: both
|
||||
live true positives are bare, `research`'s "(use neuledge-context)" and the `gitea-labels-` /
|
||||
`milestones` fold. On the body-shape checks: a `## Gotchas` heading must *end* in "gotchas", not
|
||||
merely contain the word, so `## Gotcha handling` and `## Why gotchas matter` are prose sections and
|
||||
are skipped; fenced code blocks are masked out of heading detection and entry counting, so a fenced
|
||||
example list is not mistaken for the section; and a `references/` pointer named on a line
|
||||
that also says the file is gone ("removed", "deprecated", "no longer") is read as a historical
|
||||
mention rather than a dead dispatch entry. Note the 25% fraction is deliberately *not* fence-masked
|
||||
on either side — fenced lines are real body words, and the fraction is measured against the whole
|
||||
body.
|
||||
|
||||
**The deterministic tier blocks immediately, with no baseline file.**
|
||||
|
||||
Three pre-existing contradictions are fixed in the same change, because they are the contract:
|
||||
|
||||
- `skill-audit/SKILL.md:58` asks whether the description opens with an action verb ("Audits…",
|
||||
"Reviews…") while `:101` and `DescriptionOpener.yml` require an imperative "Use when…" opener. The
|
||||
criterion is unsatisfiable against the house's own skills, both of which open with "Use when".
|
||||
- `DescriptionOpener.yml`'s regex is anchored to `^This (skill|agent)\b`, so `gitea-workflow` ("This
|
||||
is the human-facing entry point…") and `gitea-labels-milestones` ("This is a cross-cutting shared
|
||||
skill…") both violate the rule and pass the linter.
|
||||
"Reviews…"), while `:56` defers the same question to `Kyberforge.DescriptionOpener` and
|
||||
`skill-author/SKILL.md:101` requires an imperative "Use when…" opener. The criterion is
|
||||
unsatisfiable against the house's own skills, both of which open with "Use when".
|
||||
- `DescriptionOpener.yml` is anchored to `^This (skill|agent)\b`, which misses a plain `This …`
|
||||
opener; it is widened here to `^This\b`. The anchor itself stays. Composition prose that sits
|
||||
*mid*-description — `gitea-workflow`'s "This is the human-facing entry point…" at character 377,
|
||||
`gitea-labels-milestones`'s "This is a cross-cutting shared skill…" at character 300 — was never in
|
||||
the opener rule's scope and correctly is not: under `scope: text.frontmatter.description` the `^`
|
||||
anchors to the start of the whole folded value, and un-anchoring to reach mid-description text was
|
||||
measured at 5 hits and 5 false positives and rejected (`LESSONS.md`, 2026-08-14). The real gap is
|
||||
that no rule covered that text at all, which a new token-list rule, `Kyberforge.CompositionNote`,
|
||||
closes: 10 alerts across four `gitea-*` skills, 0 false positives.
|
||||
- `description-quality.md:45-50` has no FAIL condition for internal-mechanics content, which is why
|
||||
`skill-author/SKILL.md:102` never bit.
|
||||
|
||||
@@ -159,24 +284,35 @@ retrofits kyberforge's own four author/audit skills, so the figures on landing a
|
||||
With the gate hot and no baseline, a one-line
|
||||
fix to `gitea-prs` cannot be committed until that skill meets the contract. This is deliberate — it
|
||||
guarantees convergence and avoids a half-state — but it means the retrofit is lazy and *mandatory*
|
||||
rather than deferred. The follow-up retrofit issue should be prioritised accordingly, and the risk it
|
||||
rather than deferred. Issue #99 tracks it and should be prioritised accordingly, and the risk it
|
||||
carries is the ordinary one for hot gates: a gate expensive enough to be inconvenient gets bypassed
|
||||
with `SKIP=` and loses its authority.
|
||||
|
||||
**A second hot gate ships alongside it, and it is easy to miss.** `Kyberforge.CompositionNote` is
|
||||
`level: error` like every other rule in that style, so `pre-commit run --all-files` is red on 10
|
||||
alerts across `gitea-issues`, `gitea-labels-milestones`, `gitea-prs` and `gitea-workflow`
|
||||
independently of anything `skill-size-check` reports. Someone scoping the #99 retrofit off the size
|
||||
findings alone will fix those and still be blocked. The two gates want fixing together.
|
||||
|
||||
**A ceiling does not produce an average.** If every author writes to the 400-character FAIL, the
|
||||
preload lands at 15,600 chars — a 34% cut, not the ~50% intended. The halving depends entirely on the
|
||||
preload lands at 39 × 400 = 15,600 chars — a 33% cut off 23,427, not the ~50% intended. Writing to
|
||||
the 250-character SUGGESTION instead lands at 9,750, a 58% cut. The halving depends entirely on the
|
||||
250-character SUGGESTION tier being visible and respected. That tier works here in a way it does not
|
||||
elsewhere in this repo: `skill-audit` already reports `PASS (N suggestions)` as a first-class
|
||||
outcome. This is explicitly **not** the failure ADR-0013 records — Vale warnings are invisible
|
||||
because vale's exit code keys on `error` alone, but these gates live in `validate.sh` and
|
||||
`skill-audit`, where a SUGGESTION reaches the report. Realistic landing is 34-55% down, not a
|
||||
guaranteed 50%.
|
||||
`skill-audit`, where a SUGGESTION reaches the report. Realistic landing is somewhere in that 33-58%
|
||||
band, not a guaranteed 50%.
|
||||
|
||||
**A word gate cannot detect the defect it is standing in for.** `git-commits` carries thirteen
|
||||
Gotchas of which four restate steps in its own Workflow (`:32` ≡ step 9, `:33` ≡ step 9, `:36` ≡ step
|
||||
2, `:31` ≡ the description) — 1,217 words that pass any plausible gate. The counts are a backstop to
|
||||
the dispatch rule and the Gotchas constraint, not a substitute for them, and should not be read as
|
||||
the mechanism.
|
||||
**A word gate cannot detect the defect it is standing in for.** `git-commits` carries twelve Gotchas
|
||||
of which four restate steps in its own Workflow (`:32` ≡ step 9, `:33` ≡ step 9, `:36` ≡ step 2,
|
||||
`:31` ≡ the description). Its body is 1,102 words and its whole file 1,217, so it does fail the
|
||||
900-word body FAIL — but for its length, not for the restatement. The four duplicated Gotchas are 114
|
||||
words between them; delete every one and the file still fails, while a skill 250 words shorter with
|
||||
the identical defect passes clean. The two properties are uncorrelated, which is why the counts are a
|
||||
backstop to the dispatch rule and the Gotchas constraint — both of which are auditor judgment for the
|
||||
semantic half, per the Enforcement table — and not a substitute for them. Reading the word gate as
|
||||
the mechanism is the specific mistake this paragraph exists to prevent.
|
||||
|
||||
**Some skills legitimately need more description budget than others.** A tiered limit keyed to
|
||||
sibling density was considered and rejected as too clever; the flat 250/400 pair means the `gitea-*`
|
||||
@@ -184,22 +320,34 @@ and `git-*` families — where every sibling shares a keyword and boundary claus
|
||||
— are the ones most likely to sit at the FAIL tier permanently. If the retrofit shows that family
|
||||
routing degrades, the tier is the first thing to revisit.
|
||||
|
||||
**Four broken routing targets are live and are not fixed here.** `skill-audit` routes to
|
||||
`/skill-improve` twice in its description plus `README.md:10`, and no such skill exists — the real
|
||||
target is `skill-author`. `research` routes to `neuledge-context`, which exists only inside that
|
||||
string. `agent-author` says "Do not use for read-only review — examine agent files manually",
|
||||
routing away from `agent-audit`, the correct sibling. `gitea-issues` contains the literal string
|
||||
`gitea-labels- milestones`, a stray space introduced by YAML folding, breaking the skill name in
|
||||
preloaded text. The resolvable-target check added here will fail on all four the moment those files
|
||||
are touched; fixing them is split into its own issue.
|
||||
**Four broken routing targets were found; two are fixed here and two are live.** Tracked as issue
|
||||
#100.
|
||||
|
||||
- `skill-audit` routed to `/skill-improve` twice in its description plus `README.md:10`, and no such
|
||||
skill exists — the real target is `skill-author`. **Fixed here**, as a side effect of retrofitting
|
||||
kyberforge's own skills.
|
||||
- `agent-author` said "Do not use for read-only review — examine agent files manually", routing away
|
||||
from `agent-audit`, the correct sibling. **Fixed here**, same way. Note this one was never
|
||||
detectable by the resolvable-target check and never will be: "examine agent files manually" names
|
||||
no target, and a check that resolves names cannot see a name that is absent. A misroute to nowhere
|
||||
is a review finding, not a gate finding.
|
||||
- `research` routes to `neuledge-context`, which exists only inside that string. **Live.**
|
||||
- `gitea-issues` carries the literal string `gitea-labels- milestones` in its folded description, a
|
||||
stray space introduced by YAML wrapping mid-token, breaking the skill name in preloaded text.
|
||||
**Live** — the check reports it as a dangling `gitea-labels`.
|
||||
|
||||
So the check fires on 3 of the 4 against the base commit and on 2 at the tip of this change, and
|
||||
`tests/test-skill-size-check.sh` probes exactly those three by name rather than asserting a count, so
|
||||
it degrades to SKIP as #100 lands rather than going stale.
|
||||
|
||||
**Duplication between `skill-author` and `agent-author` survives un-gated.** The merge rule
|
||||
deliberately excludes the author pair, so the commit-verification argument in four near-copies, the
|
||||
root-cause grouping rule in four copies, and the wholesale clone of the "Improving an existing X"
|
||||
flow all remain. Cache isolation makes them structurally unavoidable
|
||||
(`skill-audit/SKILL.md:95` forbids cross-skill references; `LESSONS.md:107` records why), so the
|
||||
options are a sync gate or continued drift. This is an input to the kyberforge-bodies follow-up
|
||||
issue, not a solved problem.
|
||||
options are a sync gate or continued drift. This is an input to issue #101, which carries both halves
|
||||
of the kyberforge duplication problem — the deferred audit-pair merge and this — not a solved
|
||||
problem.
|
||||
|
||||
**Provenance frontmatter is explicitly out of scope.** `LESSONS.md:63` asserts that non-routing
|
||||
frontmatter (`source_keys`, `category`, `version`) is loaded at agent startup, which would make the
|
||||
@@ -211,11 +359,14 @@ the ADR-0009 provenance machinery for no runtime gain. The metadata was added de
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
Upstream citations below are relative to
|
||||
`plugins/kyberforge/docs/research/examples/skill-write/`, as in Context above.
|
||||
|
||||
- **Keep pushiness, raise the budget to ~500 chars.** Undertriggering is the worse failure mode — a
|
||||
skill that never fires is worth nothing regardless of cost — and `skill-creator/SKILL.md:67`
|
||||
explicitly recommends being "pushy" against an observed undertriggering tendency. Rejected because
|
||||
that claim is an unmeasured assertion about an older model, and because the correctness hazard in
|
||||
`writing-skills:154-158` cuts the other way: a fat description is not merely expensive, it is a
|
||||
`writing-skills/SKILL.md:154-158` cuts the other way: a fat description is not merely expensive, it is a
|
||||
shortcut agents take instead of reading the body. Would have landed a 35% cut.
|
||||
- **A trigger-eval loop to set lengths empirically.** `skill-creator/SKILL.md:337-404` specifies 20
|
||||
queries per skill, 8-10 positive and 8-10 near-miss, with a 60/40 train/test split selecting on
|
||||
@@ -226,8 +377,9 @@ the ADR-0009 provenance machinery for no runtime gain. The metadata was added de
|
||||
skill's edit fail on account of another skill's growth, and because it is meaningless for an
|
||||
external consumer installing a subset of the plugins.
|
||||
- **500-word body FAIL, matching `writing-skills/SKILL.md:217-221`.** Best-grounded in upstream and
|
||||
would align this repo with the tightest source. Rejected because it fails 30 of 39 skills, and a
|
||||
blunt gate gets satisfied by deleting content rather than relocating it.
|
||||
would align this repo with the tightest source. Rejected because it fails 28 of 39 skills body-only
|
||||
(35 of 39 measured whole-file), and a blunt gate gets satisfied by deleting content rather than
|
||||
relocating it.
|
||||
- **A shrinking baseline file** recording each non-compliant skill's current numbers, failing only on
|
||||
growth. Would have made the retrofit a visible burn-down instead of a wall. Rejected in favour of
|
||||
hot gates.
|
||||
|
||||
@@ -76,6 +76,9 @@ Include what the fresh context lacks:
|
||||
- A direct role instruction opening the prompt: `You are a [role]. When invoked, [action].`
|
||||
- One bounded job, stated so the agent knows what it must refuse.
|
||||
- The dispatch, gates, inputs and outputs listed above.
|
||||
- **Error handling** — what the agent does on malformed, missing or contradictory input: stop and
|
||||
report, or degrade to a named fallback. Absent it, the agent invents a recovery, and a
|
||||
subagent's invented recovery is invisible to its caller until the output is wrong.
|
||||
- Non-obvious environment facts and project-specific conventions it cannot infer.
|
||||
- One default per decision point with one escape hatch.
|
||||
|
||||
@@ -111,6 +114,8 @@ Flag as FAIL if:
|
||||
Flag as SUGGESTION if:
|
||||
|
||||
- The body does not open with a direct role instruction
|
||||
- The body specifies no error handling — nothing tells the agent what to do with malformed,
|
||||
missing or contradictory input
|
||||
- The job the agent describes is unbounded, or bounded only implicitly
|
||||
- A rationale is missing from a rule the agent is expected to enforce — present but unexplained
|
||||
- Comments are useful but verbose enough to bury the field they annotate
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -35,6 +35,24 @@ You are a test agent. When invoked, do the thing.
|
||||
EOF
|
||||
}
|
||||
|
||||
# Helper: a description of EXACTLY <n> characters that carries a boundary
|
||||
# clause and names no routing target. ADR-0020's missing-boundary-clause
|
||||
# SUGGESTION fires on any description without one, so a fixture that omits it
|
||||
# is never "otherwise clean" and a test refuting SUGGESTION would be asserting
|
||||
# the boundary check's absence instead of the thing it names. The clause is
|
||||
# paid for out of the measured budget rather than appended to it, because
|
||||
# these tests measure the description LENGTH. "anything else" is not
|
||||
# hyphenated, so no routing target comes with it.
|
||||
desc_of_length() {
|
||||
python3 - "$1" <<'PY'
|
||||
import sys
|
||||
n = int(sys.argv[1])
|
||||
prefix = 'Use when doing the thing. Do not use for anything else. '
|
||||
assert n >= len(prefix), 'requested description shorter than the boundary clause'
|
||||
print(prefix + 'x' * (n - len(prefix)))
|
||||
PY
|
||||
}
|
||||
|
||||
# Helper: same shape as make_apm_agent, but the description is supplied
|
||||
# verbatim — used by the ADR-0020 description-budget tests.
|
||||
make_apm_agent_with_desc() {
|
||||
@@ -637,7 +655,7 @@ EOF
|
||||
|
||||
@test "ADR-0020: agent description of exactly 250 chars raises no suggestion" {
|
||||
local root="$TMPDIR/pkg"
|
||||
make_apm_agent_with_desc "$root" "my-agent" "$(python3 -c "print('x' * 250)")"
|
||||
make_apm_agent_with_desc "$root" "my-agent" "$(desc_of_length 250)"
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_success
|
||||
refute_output --partial "SUGGESTION"
|
||||
@@ -645,7 +663,7 @@ EOF
|
||||
|
||||
@test "ADR-0020: agent description of 251 chars raises a SUGGESTION and still exits 0" {
|
||||
local root="$TMPDIR/pkg"
|
||||
make_apm_agent_with_desc "$root" "my-agent" "$(python3 -c "print('x' * 251)")"
|
||||
make_apm_agent_with_desc "$root" "my-agent" "$(desc_of_length 251)"
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_success
|
||||
assert_output --partial "SUGGESTION"
|
||||
@@ -654,7 +672,7 @@ EOF
|
||||
|
||||
@test "ADR-0020: agent description of exactly 400 chars is a SUGGESTION, not a FAIL" {
|
||||
local root="$TMPDIR/pkg"
|
||||
make_apm_agent_with_desc "$root" "my-agent" "$(python3 -c "print('x' * 400)")"
|
||||
make_apm_agent_with_desc "$root" "my-agent" "$(desc_of_length 400)"
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_success
|
||||
assert_output --partial "SUGGESTION"
|
||||
@@ -662,7 +680,7 @@ EOF
|
||||
|
||||
@test "ADR-0020: agent description of 401 chars FAILs and exits non-zero" {
|
||||
local root="$TMPDIR/pkg"
|
||||
make_apm_agent_with_desc "$root" "my-agent" "$(python3 -c "print('x' * 401)")"
|
||||
make_apm_agent_with_desc "$root" "my-agent" "$(desc_of_length 401)"
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_failure
|
||||
assert_output --partial "description is 401 chars"
|
||||
@@ -732,10 +750,18 @@ EOF
|
||||
# body becomes the system prompt of a fresh context. ADR-0020 gates the
|
||||
# former at 900 words and explicitly declines to gate the latter. If a body
|
||||
# word gate is ever added here, it contradicts the ADR.
|
||||
#
|
||||
# The description carries a boundary clause so the ONLY thing this test can
|
||||
# go red on is a body finding. Without one, the missing-boundary-clause
|
||||
# SUGGESTION fires and the blanket `refute_output --partial "SUGGESTION"`
|
||||
# below trips for a reason that has nothing to do with body length — which
|
||||
# would look like the invariant breaking while proving nothing about it.
|
||||
# AGENTS.md cites this test as the pin for that invariant, so it has to fail
|
||||
# for one reason and one reason only.
|
||||
{
|
||||
echo "---"
|
||||
echo "name: my-agent"
|
||||
echo "description: A valid agent description."
|
||||
echo "description: A valid agent description. Do not use for anything else."
|
||||
echo "---"
|
||||
echo ""
|
||||
python3 -c "print(' '.join(['word'] * 1500))"
|
||||
@@ -743,7 +769,13 @@ EOF
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_success
|
||||
refute_output --partial "FAIL"
|
||||
# A 1,500-word body is 667% of the skill ceiling. Nothing may be said about
|
||||
# it at any tier: not a FAIL, not a SUGGESTION, and not the word-count
|
||||
# wording either tier would use if a gate were quietly added later.
|
||||
refute_output --partial "SUGGESTION"
|
||||
refute_output --partial "1500 words"
|
||||
refute_output --partial "900-word"
|
||||
refute_output --partial "body is"
|
||||
}
|
||||
|
||||
@test "a bare plugin.json with no apm.yml is no longer plugin scope — falls through to project scope" {
|
||||
|
||||
@@ -34,10 +34,13 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
|
||||
boundary clause naming a real sibling skill or agent.
|
||||
250 characters is the target, 400 the hard ceiling (ADR-0020).
|
||||
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was
|
||||
deleted. Add "Use proactively" only if the runtime should delegate here without
|
||||
the user naming this agent.
|
||||
deleted.
|
||||
Never write "Use proactively" here. It steers the Claude Code runtime and does
|
||||
nothing anywhere else, and this file compiles to a Copilot `.agent.md` too, where
|
||||
agent-audit's KyberforgeCopilot.ProactivePhrase rule grades it a hard FAIL.
|
||||
The phrase is CC-only; at this scope, a precise trigger clause does that job.
|
||||
Example: "Use when a diff needs checking for injected credentials before it
|
||||
merges. Not general code review -> code-reviewer." -->
|
||||
merges. Not prose or style linting -> `lint-runner`." -->
|
||||
|
||||
<!-- model: sonnet
|
||||
Optional. Aliases: sonnet, opus, haiku, fable. Or full model ID.
|
||||
@@ -80,3 +83,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
|
||||
## Output
|
||||
|
||||
FILL IN: What does the agent produce? Format, location, structure.
|
||||
|
||||
## Errors
|
||||
|
||||
FILL IN: What does the agent do on malformed, missing or contradictory input?
|
||||
State whether it stops and reports, or degrades to a named fallback — and what it
|
||||
tells the caller either way. An agent with no error handling invents a recovery,
|
||||
and an invented recovery is invisible until the output is wrong.
|
||||
|
||||
@@ -14,10 +14,14 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
|
||||
boundary clause naming a real sibling skill or agent.
|
||||
250 characters is the target, 400 the hard ceiling (ADR-0020).
|
||||
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was
|
||||
deleted. Add "Use proactively" only if the runtime should delegate here without
|
||||
the user naming this agent.
|
||||
deleted.
|
||||
"Use proactively" is valid HERE and only here: it steers the Claude Code runtime
|
||||
to offer this agent unprompted. Add it only if that is what you want. If you add
|
||||
it, leave it OUT of the Copilot half of the pair — the phrase does nothing there
|
||||
and agent-audit's KyberforgeCopilot.ProactivePhrase grades it a hard FAIL. The
|
||||
pair must describe the same job; it does not have to be byte-identical.
|
||||
Example: "Use when a diff needs checking for injected credentials before it
|
||||
merges. Not general code review -> code-reviewer." -->
|
||||
merges. Not prose or style linting -> `lint-runner`." -->
|
||||
|
||||
<!-- tools: Read, Bash, Grep
|
||||
Optional. Allowlist of tool names: a comma-separated string or a YAML list.
|
||||
@@ -99,3 +103,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
|
||||
## Output
|
||||
|
||||
FILL IN: What does the agent produce? Format, location, structure.
|
||||
|
||||
## Errors
|
||||
|
||||
FILL IN: What does the agent do on malformed, missing or contradictory input?
|
||||
State whether it stops and reports, or degrades to a named fallback — and what it
|
||||
tells the caller either way. An agent with no error handling invents a recovery,
|
||||
and an invented recovery is invisible until the output is wrong.
|
||||
|
||||
@@ -19,9 +19,13 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
|
||||
boundary clause naming a real sibling skill or agent.
|
||||
250 characters is the target, 400 the hard ceiling (ADR-0020).
|
||||
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was deleted.
|
||||
Keep it identical in wording to the Claude Code half of the pair.
|
||||
Never write "Use proactively" here. It steers the Claude Code runtime and does nothing
|
||||
in Copilot, and agent-audit's KyberforgeCopilot.ProactivePhrase grades it a hard FAIL.
|
||||
Otherwise keep the wording matched to the Claude Code half of the pair: agent-audit
|
||||
checks that both halves describe the same job, not that they are byte-identical, so
|
||||
dropping the CC-only phrase here is not a pair-consistency finding.
|
||||
Example: "Use when a diff needs checking for injected credentials before it
|
||||
merges. Not general code review -> code-reviewer." -->
|
||||
merges. Not prose or style linting -> `lint-runner`." -->
|
||||
|
||||
<!-- tools: ["read", "search", "edit"]
|
||||
Optional. Array of tool names. Omit = all available tools. [] = no tools.
|
||||
@@ -68,3 +72,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
|
||||
## Output
|
||||
|
||||
FILL IN: What does the agent produce? Format, location, structure.
|
||||
|
||||
## Errors
|
||||
|
||||
FILL IN: What does the agent do on malformed, missing or contradictory input?
|
||||
State whether it stops and reports, or degrades to a named fallback — and what it
|
||||
tells the caller either way. An agent with no error handling invents a recovery,
|
||||
and an invented recovery is invisible until the output is wrong.
|
||||
|
||||
@@ -44,15 +44,31 @@ Banned from a description; move it to the body or to `README.md`:
|
||||
and ADR-0020 deleted it: the opener is `Use when`, matching every skill in this corpus, so one
|
||||
router reads one shape.
|
||||
|
||||
**"Use proactively" is conditional.** Add it only where the runtime should delegate without the
|
||||
user naming the agent — an agent invoked by name does not need it, and it costs activations
|
||||
elsewhere when added by reflex. The same conditional governs indirect triggers ("even if the user
|
||||
doesn't say X"): add one only where the user's natural phrasing genuinely omits the domain word.
|
||||
**"Use proactively" is Claude Code-only, and conditional even there.** The phrase steers the
|
||||
Claude Code runtime to offer an agent unprompted and does nothing anywhere else, so where it may
|
||||
appear depends on the file:
|
||||
|
||||
**Boundary targets must resolve.** The name after the arrow is checked against real skills under
|
||||
`plugins/*/.apm/skills/<name>/` and real agents under `plugins/*/.apm/agents/<name>.agent.md`. A
|
||||
target that does not exist sends the router nowhere. Verify it before writing it — do not invent a
|
||||
plausible sibling.
|
||||
| File | Rule |
|
||||
|---|---|
|
||||
| Claude Code `.md` (project/user scope) | Allowed. Add it only where the runtime should delegate without the user naming the agent — an agent invoked by name does not need it, and it costs activations elsewhere when added by reflex. |
|
||||
| Copilot `.agent.md` (project/user scope) | **Never.** Inert there, and `KyberforgeCopilot.ProactivePhrase` grades it a hard FAIL. |
|
||||
| Vendor-neutral `.apm/agents/<name>.agent.md` (plugin/APM scope) | **Never.** Same Vale rule, same hard FAIL — the file matches the `**/*.agent.md` glob, and it compiles to a real Copilot agent downstream. |
|
||||
|
||||
A pair whose Claude Code half carries the phrase and whose Copilot half omits it is correct, not
|
||||
inconsistent: `agent-audit` checks that both halves describe the same job, not that they match
|
||||
word for word.
|
||||
|
||||
Indirect triggers ("even if the user doesn't say X") take a similar conditional at every scope:
|
||||
add one only where the user's natural phrasing genuinely omits the domain word.
|
||||
|
||||
**Boundary targets must resolve.** Both forms are checked — the arrow and the prose form ("do not
|
||||
use for X, use `y` instead") — so a typo dangles either way. Targets resolve against a universe
|
||||
built by walking up **from the agent file itself**: the nearest ancestor holding
|
||||
`plugins/*/.apm/{skills,agents}` (or, failing that, the nearest ancestor holding `.git`) contributes
|
||||
every skill and agent under `<root>/plugins/*/`, plus the agent's own apm package and the packages
|
||||
that package declares in `apm.yml` under `dependencies.apm`. A sibling plugin in the same monorepo
|
||||
therefore resolves; a skill in an unrelated repo does not. A target outside that universe sends the
|
||||
router nowhere. Verify it before writing it — do not invent a plausible sibling.
|
||||
|
||||
**Length.** 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only
|
||||
with YAML folding resolved. Treat 250 as the target: the SUGGESTION tier is what moves the corpus
|
||||
@@ -73,8 +89,18 @@ You are a <role>. When invoked, <primary action>.
|
||||
|
||||
## Output
|
||||
<what it produces: format, location, structure>
|
||||
|
||||
## Errors
|
||||
<what to do on malformed, missing or contradictory input: report and stop, or
|
||||
which fallback to take — and what to say to the caller either way>
|
||||
````
|
||||
|
||||
Four required elements: **inputs expected, process steps, output format, error handling.** The
|
||||
last is the one that gets dropped, and dropping it is not neutral: an agent given a malformed
|
||||
input and no instruction invents a recovery, and a subagent's invented recovery is invisible to
|
||||
the caller until the output is wrong. Say explicitly whether the agent stops and reports, or
|
||||
degrades to a named fallback.
|
||||
|
||||
One job per agent. An agent covering two jobs gets delegated to for the wrong one.
|
||||
|
||||
**Delegation discipline replaces the word gate.** A plugin/APM agent is a single file with no
|
||||
|
||||
@@ -63,5 +63,7 @@ to installed skills instead of transcribed procedure.
|
||||
- [ ] No `FILL IN:` placeholder and no `<!-- ... -->` template comment anywhere in the file
|
||||
- [ ] System prompt body non-empty, and a read-only agent says so in prose as well as in
|
||||
`disallowedTools`
|
||||
- [ ] Body covers all four required elements: inputs expected, process steps, output format,
|
||||
**error handling** — what the agent does on malformed, missing or contradictory input
|
||||
|
||||
Then return to the flow reference you came from.
|
||||
|
||||
@@ -95,6 +95,8 @@ Both files:
|
||||
|
||||
- [ ] `name` present and kebab-case; `description` written to `references/contract.md`
|
||||
- [ ] System prompt body present, non-empty and equivalent across the pair
|
||||
- [ ] Body covers all four required elements: inputs expected, process steps, output format,
|
||||
**error handling** — what the agent does on malformed, missing or contradictory input
|
||||
- [ ] No `FILL IN:` placeholder and no `<!-- ... -->` template comment left
|
||||
|
||||
Copilot file only:
|
||||
|
||||
@@ -285,6 +285,8 @@ else
|
||||
if [[ "$SCOPE" == "plugin" ]]; then
|
||||
echo " 1. Fill in $APM_FILE — replace every FILL IN: placeholder. Optional fields are" >&2
|
||||
echo " scaffolded there as commented blocks; uncomment the ones that apply." >&2
|
||||
echo " Description: 250 chars target / 400 ceiling (ADR-0020). The body has no" >&2
|
||||
echo " word gate — delegate to a skill instead of restating what it does." >&2
|
||||
echo " 2. Populate $SOURCES_DIR/sources.md with research sources, or delete it" >&2
|
||||
echo " 3. Validate: $VALIDATE_HINT $APM_FILE" >&2
|
||||
echo " It checks the frontmatter against the apm-agent-allowlist section of" >&2
|
||||
@@ -292,6 +294,8 @@ else
|
||||
else
|
||||
echo " 1. Fill in $CC_FILE — replace every FILL IN: placeholder. Optional fields are" >&2
|
||||
echo " scaffolded there as commented blocks; uncomment the ones that apply." >&2
|
||||
echo " Description: 250 chars target / 400 ceiling (ADR-0020). The body has no" >&2
|
||||
echo " word gate — delegate to a skill instead of restating what it does." >&2
|
||||
echo " 2. Fill in $CP_FILE — same, and heed its closing comment: the Claude Code-only" >&2
|
||||
echo " fields it names must not cross over from the file above." >&2
|
||||
echo " 3. Validate: run $VALIDATE_HINT on each file" >&2
|
||||
|
||||
@@ -6,10 +6,12 @@ Audit a skill directory against the agentskills.io specification and the house c
|
||||
|
||||
1. Runs `scripts/validate.sh` and `scripts/validate-provenance.sh` for structural and provenance checks, plus `scripts/vale-wrap.sh` — a Vale prefilter that deterministically flags non-imperative description openers, composition and architecture notes, vague wording, padding phrases, and "There is/are" sentence openers
|
||||
2. Reads all files in the skill directory
|
||||
3. Applies qualitative checks across six dimension groups, loading one rubric from `references/` per group
|
||||
3. Applies qualitative checks across five dimension groups, loading one rubric from `references/` per group
|
||||
4. Outputs a compact findings report — findings only, grouped by dimension, each with Why and Fix — and a result block with handoff to `skill-author`
|
||||
|
||||
`validate.sh` enforces two independent length families that must not be conflated: the agentskills.io spec conformance ceilings (500 lines, 2,770 words, both counting the whole file) and the ADR-0020 context budget (250/400 description characters, 600/900 body-only words, plus resolvable boundary targets).
|
||||
`validate.sh` enforces two independent length families that must not be conflated: the agentskills.io spec conformance ceilings (500 lines, 2,770 words, both counting the whole file) and the ADR-0020 context budget (250/400 description characters, 600/900 body-only words).
|
||||
|
||||
Alongside those it runs four shape checks that are not length measurements at all. Two are FAILs: every routing target named in the description — in the compressed `Not <thing> -> <name>` arrow **and** in the prose form — must resolve to a real skill or agent, and every `references/<file>.md` the body names must exist on disk. Three are SUGGESTIONs: a missing boundary clause, a Gotchas section over five entries, and a Gotchas section over 25% of the body. The resolution universe for boundary targets is derived by walking up from the audited `SKILL.md` — the authoring root above it, its own apm package, and that package's declared `apm.yml` dependencies — so a fresh clone and a machine that has run `apm install` return the same verdict. When no universe can be determined the check prints `INFO ... DID NOT RUN` and does not silently pass.
|
||||
|
||||
## Usage
|
||||
|
||||
@@ -24,7 +26,7 @@ Provide the path to the skill directory to audit when invoking.
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `SKILL.md` | Skill instructions for agents |
|
||||
| `scripts/validate.sh` | Structural validator — checks name format, name matches directory, description length, line count, placeholder detection, script executable bit, and interactive-prompt detection |
|
||||
| `scripts/validate.sh` | Structural validator — checks name format, name matches directory, description presence and length, body-only word count, line and whole-file word ceilings, boundary-clause presence, boundary-target resolution, `references/` pointer existence, Gotchas entry count and body share, placeholder detection, script executable bit, and interactive-prompt detection |
|
||||
| `scripts/validate-provenance.sh` | Provenance validator — checks sources.md completeness, source_keys/slug consistency, Contributing files existence, bidirectional linkage, Research doc: fields, and upstream research doc alignment |
|
||||
| `scripts/vale-wrap.sh` | Vale prefilter wrapper — runs the bundled `Kyberforge` Vale styles against SKILL.md and reports alerts as deterministic FAILs ahead of Step 3's qualitative review |
|
||||
| `assets/vale/.vale.ini` | Vale configuration — points Vale at the bundled `Kyberforge` style path, self-located relative to `vale-wrap.sh` |
|
||||
@@ -38,6 +40,7 @@ Provide the path to the skill directory to audit when invoking.
|
||||
| `references/patterns.md` | Rubric for the patterns dimension — which instruction construct fits which job, and how each is correctly formed |
|
||||
| `references/file-structure.md` | Rubric for the file-structure and internal-consistency dimensions — permitted directories, cross-plugin path rules and their two structural exemptions, README drift |
|
||||
| `references/formatting-and-scripts.md` | Rubric for the formatting and scripts dimensions — heading and fencing conventions, and the agentic-use criteria for bundled scripts |
|
||||
| `references/validation-scripts.md` | Step 1 troubleshooting — the manual structural fallback when `validate.sh` cannot run, and the script exit codes that are easy to misread (loaded only on a script failure) |
|
||||
| `references/sources.md` | Provenance record — agentskills.io sources that informed this skill and which files each contributed to |
|
||||
| `tests/validate.bats` | (source-only) Bats test suite for validate.sh |
|
||||
| `tests/validate-provenance.bats` | (source-only) Bats test suite for validate-provenance.sh |
|
||||
|
||||
@@ -33,7 +33,9 @@ bash scripts/validate-provenance.sh <skill-dir>
|
||||
scripts/vale-wrap.sh <skill-dir>/SKILL.md
|
||||
```
|
||||
|
||||
`validate.sh` findings become the `### Structure` dimension — its FAILs and its SUGGESTIONs both. If it cannot run at all (no `python3`, Bash denied), report that as an INFO finding rather than guessing; what it measures is not reproducible by reading.
|
||||
`validate.sh` findings become the `### Structure` dimension — its FAILs and its SUGGESTIONs both.
|
||||
|
||||
If any of the three fails, cannot run, or reports something needing interpretation, read `references/validation-scripts.md` — it carries the manual fallback and the misleading exit codes.
|
||||
|
||||
`validate-provenance.sh` prints nothing on success. Its FAIL and INFO findings become a separate `### Provenance` dimension, and it emits Why and Fix itself — surface those verbatim.
|
||||
|
||||
|
||||
@@ -69,8 +69,10 @@ table** plus the gates common to every branch, and each flow lives in its own se
|
||||
`references/` file. Inlining all of them is a FAIL regardless of word count, because every
|
||||
invocation then pays for every branch it did not take.
|
||||
|
||||
The reference shape in this repo is `apm-workflow`: a 554-word body dispatching to roughly 3,000
|
||||
words of references across five mutually exclusive invocations.
|
||||
The reference shape in this repo is `apm-workflow`: a **421-word body** dispatching to roughly
|
||||
3,000 words of references across five mutually exclusive invocations. Its whole-file count is 554
|
||||
words — cite 421 when calibrating a body, or the conflation this section warns against reappears
|
||||
in the finding itself.
|
||||
|
||||
## Gotchas sections
|
||||
|
||||
@@ -86,17 +88,20 @@ sensibly.
|
||||
|
||||
Constraints:
|
||||
|
||||
- **Maximum five entries.** Past five, the section is a summary of the body rather than a set of
|
||||
traps, and the agent stops reading it as a warning.
|
||||
- **More than five entries is a SUGGESTION** — five is the guideline, not a ceiling. Past five, the
|
||||
section is usually a summary of the body rather than a set of traps, and the agent stops reading
|
||||
it as a warning. It stays advisory because whether a given gotcha earns its place is judgment;
|
||||
`validate.sh` emits it through `suggest()` and the run still exits 0.
|
||||
- **A Gotcha that paraphrases a step in the body below it is a FAIL.** It has no independent
|
||||
content, and it teaches the agent that Gotchas can be skimmed because the real instruction is
|
||||
coming.
|
||||
coming. This one is the auditor's call — no script detects it.
|
||||
- **A Gotchas section exceeding 25% of the body is a SUGGESTION** — the body has been inverted into
|
||||
a preamble.
|
||||
a preamble. Same tier and same reasoning as the entry count, and independent of it: either can
|
||||
fire without the other.
|
||||
- Place the section near the top. A gotcha read after the mistake is worthless, which is also why
|
||||
Gotchas is the one construct exempt from moving to `references/`.
|
||||
|
||||
Worked negative example — `git-commits` carries thirteen entries, of which four restate content
|
||||
Worked negative example — `git-commits` carries twelve entries, of which four restate content
|
||||
that already appears below or in the description:
|
||||
|
||||
| Gotcha | Restates |
|
||||
@@ -106,8 +111,10 @@ that already appears below or in the description:
|
||||
| `:33` "Never skip hooks with `--no-verify`" | step 9 at `:52` |
|
||||
| `:36` "Never commit secrets" | step 2 at `:45` |
|
||||
|
||||
All four are FAILs under this rule, and the section as a whole breaches the five-entry maximum. It
|
||||
also passes every plausible word gate, which is the point of auditing the construct directly.
|
||||
All four are FAILs under the paraphrase rule. The entry count and the section's share of the body
|
||||
(387 of 1,102 words, 35%) are two further SUGGESTIONs on top — the script reports both, and neither
|
||||
fails the run on its own. What makes this worth auditing directly is that the four paraphrase FAILs
|
||||
pass every word gate there is; only reading the construct finds them.
|
||||
|
||||
## Calibrating control
|
||||
|
||||
@@ -144,7 +151,7 @@ Flag as FAIL if:
|
||||
- A sentence answers "no" to the core test — it is padding
|
||||
- The body exceeds 900 words counted body-only (`validate.sh` reports it)
|
||||
- Two or more mutually exclusive flows are inlined instead of dispatched
|
||||
- A Gotcha paraphrases a step in the body below it, or the section exceeds five entries
|
||||
- A Gotcha paraphrases a step in the body below it
|
||||
- A decision point presents a menu of options with no default
|
||||
- An instruction repeats content already in the description
|
||||
- A prescriptive sequence is used where flexibility is fine, or the reverse
|
||||
@@ -152,6 +159,7 @@ Flag as FAIL if:
|
||||
Flag as SUGGESTION if:
|
||||
|
||||
- The body exceeds 600 words counted body-only but stays at or under 900
|
||||
- The Gotchas section carries more than five entries
|
||||
- The Gotchas section exceeds 25% of the body
|
||||
- A rationale is missing from an include/exclude rule — present but unexplained
|
||||
- Gotchas are correct but placed late in the body rather than near the top
|
||||
|
||||
@@ -28,6 +28,14 @@ resolving there. Flag any `../`, `../../`, or absolute repo path (`plugins/<plug
|
||||
and its APM-native equivalent `.apm/skills/<other>/`) appearing in `SKILL.md`, `scripts/`,
|
||||
`references/` or `assets/`.
|
||||
|
||||
**Referring to another skill's file.** There is one sanctioned spelling, and it is possessive:
|
||||
`skill-audit's references/validation-scripts.md`. Write the skill by name and let the reader
|
||||
resolve it — do not spell the repo path. The full path is the thing this section forbids, and
|
||||
`references/validation-scripts.md` on its own is a hard ERROR from the ADR-0020 gate, which
|
||||
requires an unqualified `references/` pointer to exist in the skill's OWN directory. The
|
||||
possessive form is the only spelling both rules accept; the gate recognises it and skips the
|
||||
on-disk check. Flag any other spelling of a cross-skill reference.
|
||||
|
||||
Two directories are exempt, and the exemptions are structural rather than discretionary:
|
||||
|
||||
- **`references/sources.md`.** Its `Research doc:` fields are development-time provenance pointers,
|
||||
|
||||
@@ -15,7 +15,7 @@
|
||||
- **URL:** https://agentskills.io/specification.md
|
||||
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
|
||||
- **Description:** Complete SKILL.md format specification — frontmatter fields, constraints, body content, optional directories, progressive disclosure levels, file references, validation
|
||||
- **Contributing files:** SKILL.md, references/body-discipline.md, references/description-quality.md, references/patterns.md, references/file-structure.md, references/formatting-and-scripts.md
|
||||
- **Contributing files:** SKILL.md, references/body-discipline.md, references/description-quality.md, references/patterns.md, references/file-structure.md, references/formatting-and-scripts.md, references/validation-scripts.md
|
||||
- **Status:** `extracted`
|
||||
|
||||
## agentskills-best-practices
|
||||
@@ -47,7 +47,7 @@
|
||||
- **URL:** https://agentskills.io/skill-creation/using-scripts.md
|
||||
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
|
||||
- **Description:** Using scripts in skills — one-off commands, self-contained scripts with inline dependencies, designing scripts for agentic use (no interactive prompts, --help, structured output, idempotency)
|
||||
- **Contributing files:** SKILL.md, references/formatting-and-scripts.md
|
||||
- **Contributing files:** SKILL.md, references/formatting-and-scripts.md, references/validation-scripts.md
|
||||
- **Status:** `extracted`
|
||||
|
||||
## agentskills-quickstart
|
||||
|
||||
@@ -0,0 +1,118 @@
|
||||
---
|
||||
source_keys:
|
||||
- agentskills-spec
|
||||
- agentskills-using-scripts
|
||||
---
|
||||
|
||||
# Validation Scripts Reference
|
||||
|
||||
Read this when a Step 1 script fails, cannot run, or reports something that needs interpreting.
|
||||
Nothing here is needed on a clean run.
|
||||
|
||||
## Report the gap, do not guess
|
||||
|
||||
If a script cannot run at all — Bash denied, `python3` unavailable, PyYAML not importable, `vale`
|
||||
not installed — say so as an **INFO** finding naming the script and the missing dependency, then
|
||||
fall back to the manual checks below. An INFO never changes PASS/FAIL. Silently omitting the
|
||||
dimension a script would have covered reports a clean audit that checked less than it claims to
|
||||
have checked, and the Step 4 coverage line then names a dimension nothing actually examined.
|
||||
|
||||
## Manual structural fallback
|
||||
|
||||
`validate.sh` needs `python3` **and** PyYAML, and refuses to start without either — the description
|
||||
value has to be measured after YAML folding is resolved, so skipping the ADR-0020 gates would be a
|
||||
vacuous pass rather than a partial one. The two are checked separately, so the message already names
|
||||
the right one — report it verbatim rather than diagnosing further:
|
||||
|
||||
```text
|
||||
Error: python3 is required but was not found on PATH.
|
||||
Error: PyYAML is required but is not importable by python3.
|
||||
```
|
||||
|
||||
Without them — or with Bash denied, or on a permission error — work this list
|
||||
by hand and file the results under `### Structure` exactly as the script's output would have been:
|
||||
|
||||
- **`name`** present, 1–64 characters, kebab-case (lowercase letters, digits and hyphens; no
|
||||
leading, trailing or doubled hyphen), and **matching the skill's directory name** exactly.
|
||||
- **`description`** present and non-empty; no unfilled `FILL IN:` placeholder in it. An absent or
|
||||
empty description is a **FAIL**, never a silent skip — it is the one field preloaded into every
|
||||
session, so a skill without one can never be routed to.
|
||||
- **Description length**, measured on the folded YAML value with newlines collapsed to single
|
||||
spaces — not on the raw block scalar, which counts indentation. 250 characters SUGGESTION, 400
|
||||
FAIL (ADR-0020), 1,024 FAIL (agentskills.io spec).
|
||||
- **Body length**, counting everything after the frontmatter's closing `---`. 600 words
|
||||
SUGGESTION, 900 FAIL (ADR-0020).
|
||||
- **Whole-file ceilings**, counting the file including frontmatter: 500 lines FAIL, 2,770 words
|
||||
FAIL (agentskills.io spec). These are a different measurement from the two above — report them
|
||||
as separate findings, never merged.
|
||||
- **A boundary clause is present** — either the prose form (`do not` / `instead` / `rather than` /
|
||||
`not for`) or ADR-0020's compressed `Not <thing> -> <name>` arrow. **SUGGESTION**, not FAIL:
|
||||
the absence is deterministic, but whether this skill warrants one is the auditor's call.
|
||||
- **Boundary targets resolve** — **FAIL** on a name that resolves to nothing. See the section
|
||||
below; resolving these by hand is the one item on this list with a procedure of its own.
|
||||
- **Every `references/<file>.md` named in the body exists on disk** — **FAIL**, not a suggestion.
|
||||
A dispatch table or "read X" trigger naming a missing file sends the agent nowhere. Ignore
|
||||
mentions inside fenced code blocks, and ignore a mention whose own line says the file is gone
|
||||
(`removed`, `deleted`, `renamed`, `superseded`, `replaced`, `obsolete`, `deprecated`, `former`,
|
||||
`gone`, `no longer`, `used to`) — that is a historical note, not a dispatch entry.
|
||||
- **Gotchas discipline**, both **SUGGESTION**. Locate the section by a heading that *is* Gotchas
|
||||
(`## Common Gotchas` counts; `## Gotcha handling` and `## Why gotchas matter` do not), running to
|
||||
the next heading at the same level or shallower. More than five top-level entries is one
|
||||
suggestion; a section over 25% of the body word count is a second, independent one. Count
|
||||
entries at column 0 only — an indented child bullet is not an entry — and ignore fenced code
|
||||
blocks for both.
|
||||
- **No unfilled `FILL IN:` placeholder** anywhere in the body.
|
||||
- **Every file in `scripts/`** carries the executable bit and contains no interactive prompt —
|
||||
no bare `read`, no `select`, nothing that blocks on a TTY.
|
||||
|
||||
## Resolving boundary targets by hand
|
||||
|
||||
Targets are read from **both** boundary forms. The compressed `Not <thing> -> <name>` arrow and the
|
||||
prose form are each parsed *and* target-checked, so a typo in prose phrasing fails exactly as an
|
||||
arrow typo does — do not check only the names after an arrow.
|
||||
|
||||
Build the universe by walking up **from the `SKILL.md` under audit**, never from the validator's own
|
||||
location. The nearest ancestor holding `plugins/*/.apm/skills/` or `plugins/*/.apm/agents/` is the
|
||||
authoring root, falling back to the nearest ancestor holding `.git`. When one is found the universe
|
||||
is every skill and agent under `<root>/plugins/*/`, plus the skill's own apm package, plus the
|
||||
packages that package declares in its `apm.yml` under `dependencies.apm`. Deployed `.claude/` and
|
||||
`.agents/` trees are consulted **only** when no authoring root exists — they are gitignored
|
||||
`apm install` output, and reading them would make a fresh clone and a developer machine disagree.
|
||||
|
||||
Three ways to read the result wrong:
|
||||
|
||||
- **A hyphenated name used attributively is not a dangling target.** "Use pre-commit hooks instead
|
||||
of ad-hoc scripts" reads as a route to `pre-commit` on wording alone. What separates a route from
|
||||
prose is grammar: a route target is terminal — followed by punctuation, a conjunction, or a
|
||||
boundary word — whereas a compound modifier is followed by the noun it modifies. A name followed
|
||||
by an ordinary noun still *confirms* a route when it exists, but never raises a FAIL on its own.
|
||||
- **A SUGGESTION-tier unresolved target is not a FAIL you may promote.** Terminal position alone is
|
||||
not evidence of a route: "run `pre-commit` instead", "see `commit-msg`" and "use the clean-up
|
||||
instead" are all terminal and all prose. A prose-form target earns a FAIL only when its own
|
||||
sentence names another target that *does* resolve; otherwise the script reports it and moves on,
|
||||
and so should you. Route notation — `/name` and `-> name` — is exempt and always FAILs, and it is
|
||||
the fix to recommend when the author did mean a route.
|
||||
- **`INFO boundary-target resolution DID NOT RUN` is not a pass.** The script prints it, and exits
|
||||
0, when no universe could be determined for that path — the usual cause being a skill copy
|
||||
audited outside its package. Report it as an INFO naming the unchecked targets and re-run against
|
||||
the real directory; filing it as clean signs off targets nothing verified.
|
||||
|
||||
## Script-specific failures
|
||||
|
||||
- **`validate-provenance.sh` printed nothing.** That is a pass, not a skip. It also exits 0
|
||||
silently when the skill has no `source_keys` and no `references/sources.md` — nothing to
|
||||
validate is not a finding.
|
||||
- **`vale` reports `0 files`.** Treat the pass as NOT RUN, not as clean, and fall back to full
|
||||
Step 3 judgment for the dimensions it would have covered. The bundled `Kyberforge` style is
|
||||
scoped by glob in `assets/vale/.vale.ini`; a file outside those globs is silently not linted.
|
||||
- **`E100 Runtime error ... does not exist` (exit 2) from `vale-wrap.sh`.** An explicit relative
|
||||
`--config` was passed. Pass none: the wrapper locates its own `assets/vale/.vale.ini` from its
|
||||
own path, so a resolved script path plus an unresolved config path produces exactly this. Do not
|
||||
read this exit code as vale being unavailable — that misreading sends the audit down the
|
||||
fallback path while vale was installed and working the whole time.
|
||||
- **The `vale` binary is genuinely absent** (`command not found`). Report one INFO naming it, then
|
||||
fall back to full Step 3 judgment for the description, body-discipline and patterns dimensions —
|
||||
the prefilter's whole coverage. Judge those by rubric rather than dropping them.
|
||||
- **A path argument that does not exist is a hard error** in `vale-wrap.sh`, deliberately: bare
|
||||
`vale` would fall back to reading stdin and print a clean-looking `0 errors ... in stdin`, which
|
||||
the `0 files` guard above does not catch.
|
||||
File diff suppressed because it is too large
Load Diff
@@ -8,7 +8,14 @@ setup() {
|
||||
SCRIPT="$(cd "$BATS_TEST_DIRNAME/../scripts" && pwd)/validate.sh"
|
||||
TMPDIR="$(mktemp -d)"
|
||||
|
||||
# Helper: create a minimal valid skill directory
|
||||
# Helper: create a minimal valid skill directory.
|
||||
#
|
||||
# The description carries a boundary clause deliberately. ADR-0020's
|
||||
# missing-boundary-clause SUGGESTION fires on any description without one, so
|
||||
# a fixture that omits it is never "otherwise clean" — every test asserting
|
||||
# SUGGESTION-freedom would be asserting the boundary check's absence instead
|
||||
# of the thing it names. "anything else" is not hyphenated, so the clause adds
|
||||
# a boundary marker without adding a routing target to resolve.
|
||||
make_valid_skill() {
|
||||
local dir="$1"
|
||||
local name
|
||||
@@ -17,7 +24,7 @@ setup() {
|
||||
cat > "$dir/SKILL.md" <<EOF
|
||||
---
|
||||
name: $name
|
||||
description: A valid skill description that is well within the limit.
|
||||
description: A valid skill description that is well within the limit. Do not use for anything else.
|
||||
---
|
||||
|
||||
## Step 1
|
||||
@@ -26,6 +33,20 @@ Do the thing.
|
||||
EOF
|
||||
}
|
||||
|
||||
# Helper: a description of EXACTLY <n> characters that carries a boundary
|
||||
# clause and names no routing target. The tests below measure the description
|
||||
# LENGTH, so the clause has to be paid for out of the same budget rather than
|
||||
# appended to it — hence the padding arithmetic instead of a fixed suffix.
|
||||
desc_of_length() {
|
||||
python3 - "$1" <<'PY'
|
||||
import sys
|
||||
n = int(sys.argv[1])
|
||||
prefix = 'Use when doing the thing. Do not use for anything else. '
|
||||
assert n >= len(prefix), 'requested description shorter than the boundary clause'
|
||||
print(prefix + 'x' * (n - len(prefix)))
|
||||
PY
|
||||
}
|
||||
|
||||
# Helper: create a skill directory with an exact description length and an
|
||||
# exact body word count. <desc> is used verbatim; <body_words> "word"
|
||||
# tokens follow the frontmatter. Used by the ADR-0020 boundary tests.
|
||||
@@ -300,7 +321,7 @@ EOF
|
||||
|
||||
@test "ADR-0020: description of exactly 250 chars raises no suggestion" {
|
||||
local skill="$TMPDIR/my-skill"
|
||||
make_sized_skill "$skill" "$(python3 -c "print('x' * 250)")" 10
|
||||
make_sized_skill "$skill" "$(desc_of_length 250)" 10
|
||||
run bash "$SCRIPT" "$skill"
|
||||
assert_success
|
||||
refute_output --partial "SUGGESTION"
|
||||
@@ -308,7 +329,7 @@ EOF
|
||||
|
||||
@test "ADR-0020: description of 251 chars raises a SUGGESTION and still exits 0" {
|
||||
local skill="$TMPDIR/my-skill"
|
||||
make_sized_skill "$skill" "$(python3 -c "print('x' * 251)")" 10
|
||||
make_sized_skill "$skill" "$(desc_of_length 251)" 10
|
||||
run bash "$SCRIPT" "$skill"
|
||||
assert_success
|
||||
assert_output --partial "SUGGESTION"
|
||||
@@ -318,7 +339,7 @@ EOF
|
||||
|
||||
@test "ADR-0020: description of exactly 400 chars is a SUGGESTION, not a FAIL" {
|
||||
local skill="$TMPDIR/my-skill"
|
||||
make_sized_skill "$skill" "$(python3 -c "print('x' * 400)")" 10
|
||||
make_sized_skill "$skill" "$(desc_of_length 400)" 10
|
||||
run bash "$SCRIPT" "$skill"
|
||||
assert_success
|
||||
assert_output --partial "SUGGESTION"
|
||||
@@ -326,7 +347,7 @@ EOF
|
||||
|
||||
@test "ADR-0020: description of 401 chars FAILs and exits non-zero" {
|
||||
local skill="$TMPDIR/my-skill"
|
||||
make_sized_skill "$skill" "$(python3 -c "print('x' * 401)")" 10
|
||||
make_sized_skill "$skill" "$(desc_of_length 401)" 10
|
||||
run bash "$SCRIPT" "$skill"
|
||||
assert_failure
|
||||
assert_output --partial "description is 401 chars"
|
||||
@@ -364,7 +385,7 @@ EOF
|
||||
|
||||
@test "ADR-0020: body of exactly 600 words raises no suggestion" {
|
||||
local skill="$TMPDIR/my-skill"
|
||||
make_sized_skill "$skill" "A short valid description." 600
|
||||
make_sized_skill "$skill" "A short valid description. Do not use for anything else." 600
|
||||
run bash "$SCRIPT" "$skill"
|
||||
assert_success
|
||||
refute_output --partial "SUGGESTION"
|
||||
@@ -372,7 +393,7 @@ EOF
|
||||
|
||||
@test "ADR-0020: body of 601 words raises a SUGGESTION and still exits 0" {
|
||||
local skill="$TMPDIR/my-skill"
|
||||
make_sized_skill "$skill" "A short valid description." 601
|
||||
make_sized_skill "$skill" "A short valid description. Do not use for anything else." 601
|
||||
run bash "$SCRIPT" "$skill"
|
||||
assert_success
|
||||
assert_output --partial "body is 601 words"
|
||||
@@ -381,7 +402,7 @@ EOF
|
||||
|
||||
@test "ADR-0020: body of exactly 900 words is a SUGGESTION, not a FAIL" {
|
||||
local skill="$TMPDIR/my-skill"
|
||||
make_sized_skill "$skill" "A short valid description." 900
|
||||
make_sized_skill "$skill" "A short valid description. Do not use for anything else." 900
|
||||
run bash "$SCRIPT" "$skill"
|
||||
assert_success
|
||||
assert_output --partial "body is 900 words"
|
||||
@@ -389,7 +410,7 @@ EOF
|
||||
|
||||
@test "ADR-0020: body of 901 words FAILs and exits non-zero" {
|
||||
local skill="$TMPDIR/my-skill"
|
||||
make_sized_skill "$skill" "A short valid description." 901
|
||||
make_sized_skill "$skill" "A short valid description. Do not use for anything else." 901
|
||||
run bash "$SCRIPT" "$skill"
|
||||
assert_failure
|
||||
assert_output --partial "body is 901 words"
|
||||
@@ -425,15 +446,33 @@ EOF
|
||||
assert_output --partial "boundary target(s) resolve"
|
||||
}
|
||||
|
||||
@test "ADR-0020: a boundary target naming a non-existent skill FAILs" {
|
||||
@test "ADR-0020: a boundary target naming a non-existent skill FAILs when its sentence names one that resolves" {
|
||||
local skill
|
||||
skill="$(make_fixture_tree "$TMPDIR/tree" "my-skill")"
|
||||
make_sized_skill "$skill" "Use when doing the thing. Do not use for the other thing — use fixture-missing-skill instead." 10
|
||||
# `fixture-sibling-skill` is the corroborator: a prose-form target only earns
|
||||
# a FAIL when its own sentence proves it is a routing sentence. See the
|
||||
# shared resolver's CORROBORATION note, and the uncorroborated case below.
|
||||
make_sized_skill "$skill" "Use when doing the thing. Do not use for the other thing — use fixture-sibling-skill or fixture-missing-skill instead." 10
|
||||
run bash "$SCRIPT" "$skill"
|
||||
assert_failure
|
||||
assert_output --partial "routes to 'fixture-missing-skill'"
|
||||
}
|
||||
|
||||
@test "ADR-0020: a LONE boundary target naming a non-existent skill is a SUGGESTION, not a FAIL" {
|
||||
local skill
|
||||
skill="$(make_fixture_tree "$TMPDIR/tree" "my-skill")"
|
||||
# Same grammar as the case above and as "run \`pre-commit\` instead" — a
|
||||
# route verb, a hyphenated name, terminal position. Nothing local separates a
|
||||
# broken route from a tool name, so the target is named on every run but does
|
||||
# not block: this gate ships with no baseline and no suppression mechanism.
|
||||
make_sized_skill "$skill" "Use when doing the thing. Do not use for the other thing — use fixture-missing-skill instead." 10
|
||||
run bash "$SCRIPT" "$skill"
|
||||
assert_success
|
||||
assert_output --partial "SUGGESTION"
|
||||
assert_output --partial "routes to 'fixture-missing-skill'"
|
||||
refute_output --partial "FAIL description routes to"
|
||||
}
|
||||
|
||||
@test "ADR-0020: a boundary target naming an AGENT file resolves (agents are valid routing targets)" {
|
||||
local skill
|
||||
skill="$(make_fixture_tree "$TMPDIR/tree" "my-skill")"
|
||||
@@ -452,15 +491,27 @@ EOF
|
||||
assert_output --partial "routes to 'fixture-missing-improve'"
|
||||
}
|
||||
|
||||
@test "ADR-0020: a backticked name that does not resolve FAILs" {
|
||||
@test "ADR-0020: a backticked name that does not resolve FAILs when its sentence names one that resolves" {
|
||||
local skill
|
||||
skill="$(make_fixture_tree "$TMPDIR/tree" "my-skill")"
|
||||
make_sized_skill "$skill" "Use when doing the thing. Composes \`fixture-missing-helper\` for the shared part." 10
|
||||
make_sized_skill "$skill" "Use when doing the thing. Composes \`fixture-sibling-skill\` and \`fixture-missing-helper\` for the shared part." 10
|
||||
run bash "$SCRIPT" "$skill"
|
||||
assert_failure
|
||||
assert_output --partial "routes to 'fixture-missing-helper'"
|
||||
}
|
||||
|
||||
@test "ADR-0020: a /slash-command target is route NOTATION and FAILs on its own, uncorroborated" {
|
||||
local skill
|
||||
skill="$(make_fixture_tree "$TMPDIR/tree" "my-skill")"
|
||||
# The escape hatch from the SUGGESTION tier: `/name` and `-> name` are never
|
||||
# how prose cites a tool, so they are exempt from corroboration. An author
|
||||
# who wants a route checked unconditionally writes one of those two forms.
|
||||
make_sized_skill "$skill" "Use when doing the thing. Do not use for the other thing — use /fixture-missing-notation instead." 10
|
||||
run bash "$SCRIPT" "$skill"
|
||||
assert_failure
|
||||
assert_output --partial "routes to 'fixture-missing-notation'"
|
||||
}
|
||||
|
||||
@test "ADR-0020: a bare hyphenated word outside a boundary sentence is not read as a routing target" {
|
||||
local skill
|
||||
skill="$(make_fixture_tree "$TMPDIR/tree" "my-skill")"
|
||||
@@ -502,9 +553,18 @@ EOF
|
||||
}
|
||||
|
||||
@test "ADR-0020: the boundary check declines rather than false-FAILs when no authoring source is found" {
|
||||
# Deliberately NOT built with make_fixture_tree: this skill sits in a bare
|
||||
# temp directory with no plugins/*/.apm/ above it and no .git, so the resolver
|
||||
# legitimately has no universe. That is a real path (a skill being drafted
|
||||
# outside any repo), and the required behaviour is to DECLINE OUT LOUD rather
|
||||
# than either false-FAIL or pass in silence — silence is what let a whole gate
|
||||
# family go missing unnoticed. So the INFO text and the named unchecked target
|
||||
# are both asserted, not just the absence of a failure.
|
||||
local skill="$TMPDIR/orphan/my-skill"
|
||||
make_sized_skill "$skill" "Use when doing the thing. Do not use for the other thing — use some-other-skill instead." 10
|
||||
run bash "$SCRIPT" "$skill"
|
||||
assert_success
|
||||
refute_output --partial "routes to"
|
||||
assert_output --partial "boundary-target resolution DID NOT RUN"
|
||||
assert_output --partial "Unchecked target(s): some-other-skill"
|
||||
}
|
||||
|
||||
@@ -47,6 +47,7 @@ If the destination resolves inside an APM package, read `references/deployment-m
|
||||
| `references/create.md` | The create flow end to end — prerequisites, package-intent gate, scaffold, frontmatter, scripts, references, sources (loaded on demand) |
|
||||
| `references/improve.md` | The improve flow end to end — signal verification, root-cause grouping, announcement, edits (loaded on demand) |
|
||||
| `references/contract.md` | The ADR-0020 description and body contract, the Gotchas constraint, the two size gates, body patterns, and org-policy embedding (loaded on demand) |
|
||||
| `references/retrofit.md` | Bringing a pre-ADR-0020 skill into contract — ordered cut procedure, the mutually-exclusive-flows test, reference-file conventions, the collateral checklist, and a worked description retrofit (loaded from the improve flow when a budget is exceeded) |
|
||||
| `references/deployment-modes.md` | APM package vs standalone differences and self-containment/cache-isolation rules (loaded on demand) |
|
||||
| `references/scripts.md` | Package runners, inline dependency patterns, and full script contract (loaded on demand) |
|
||||
| `references/sources.md` | Upstream research sources and which skill files each contributed to |
|
||||
|
||||
@@ -19,10 +19,9 @@ metadata:
|
||||
|
||||
## Gotchas
|
||||
|
||||
- A skill's `name` and `description` are preloaded into every agent's context every session, invoked or not; the body loads only on invocation. The description is the scarce budget.
|
||||
- The word gates are two different measurements, not one rule with two tiers. The 2,770-word / 500-line spec backstop counts the whole file including frontmatter; Step 3's gate counts the body alone. A file can sit well inside one and fail the other, so never unify them.
|
||||
- Never spawn a subagent to audit or recheck your own work here. Run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent can have its worktree torn down by concurrent cleanup, destroying an uncommitted draft.
|
||||
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires transcript analysis that is out of scope here — flag the opportunity as a suggestion instead.
|
||||
- The word gates are two measurements, not two tiers of one rule: the 2,770-word / 500-line spec backstop counts the whole file, Step 3's gate the body alone. Never unify them.
|
||||
- Never spawn a subagent to audit or recheck your own work — run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent's worktree can be torn down by concurrent cleanup, destroying an uncommitted draft.
|
||||
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires out-of-scope transcript analysis — flag the opportunity as a suggestion instead.
|
||||
|
||||
## Step 1 — Dispatch
|
||||
|
||||
@@ -32,15 +31,15 @@ metadata:
|
||||
| Directory exists, at least one improvement signal present | Improve | `references/improve.md` |
|
||||
| Directory exists, no signals | Stop and ask | — |
|
||||
|
||||
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask: "No improvement signals found. Did you mean to create a new skill, or do you have feedback to apply?"
|
||||
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask whether the user meant to create a new skill or has feedback to apply.
|
||||
|
||||
Read only the reference matching the resolved flow — each is self-contained. Capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
|
||||
Read only the reference matching the resolved flow — each is self-contained. If the target sits inside a git worktree, capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
|
||||
|
||||
## Step 2 — Invocation axis
|
||||
|
||||
Decide before writing any description: model-invoked or hand-invoked?
|
||||
|
||||
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Worked example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`. Skip Step 3's description rules.
|
||||
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Skip Step 3's description rules.
|
||||
- **Model-invoked** — the default.
|
||||
|
||||
## Step 3 — Contract
|
||||
@@ -51,12 +50,12 @@ Gates `/skill-audit` enforces in both flows:
|
||||
|
||||
- **Description** — a trigger clause, at most one capability clause, and a boundary clause shaped `Not <thing> -> <skill-name>` whose target resolves to a real skill or agent. 250 characters SUGGESTION, 400 FAIL, value only.
|
||||
- **Body** — decision procedure only: ordered steps, branches, gates, and which reference to load when. 600 words SUGGESTION, 900 FAIL, body only. At two or more mutually exclusive flows a dispatch table is mandatory and each flow gets its own self-contained `references/` file.
|
||||
- **Gotchas** — at most five, each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL.
|
||||
- **Gotchas** — each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL; over five entries is a SUGGESTION only.
|
||||
|
||||
## Step 4 — Validate and close
|
||||
|
||||
Run `/skill-audit` on the resolved skill directory. It checks name-to-directory match, description presence, leftover `FILL IN:` placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those first. Resolve every FAIL before reporting done.
|
||||
Run `/skill-audit` on the resolved skill directory; resolve every FAIL before reporting done. It checks name-to-directory match, placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those. Hand-check the one thing it misses: an empty body reports `PASS SKILL.md body word count 0 (ADR-0020 target: 600)`, so confirm at least one non-empty section exists.
|
||||
|
||||
With `metadata.version` present, bump the **minor** version on create (new skills start at `0.1.0`) and the **patch** version on improve.
|
||||
|
||||
**Commit verification.** Once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from the one captured at Step 1. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up first. Report done only once the hash has changed.
|
||||
**Commit verification.** Inside a git worktree: once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from Step 1's. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up. Report done only once the hash has changed. Outside a worktree (a skill under `~/.claude/skills/`, say) nothing is committable — report done on a clean audit, naming that as the reason.
|
||||
|
||||
@@ -10,14 +10,22 @@ name: SKILL_NAME
|
||||
# Examples: my-tool, data-analyzer, pdf-processor
|
||||
|
||||
description: >
|
||||
Use when FILL IN: trigger — when should an agent activate this skill?
|
||||
Use when FILL IN: trigger.
|
||||
FILL IN: at most ONE capability clause, stated specifically
|
||||
(e.g. "parses and validates OpenAPI specs", not "helps with APIs").
|
||||
Not FILL IN: near-miss case -> FILL IN: real sibling skill name.
|
||||
Not FILL IN: near-miss case -> FILL IN: real sibling skill.
|
||||
# Required. Preloaded into EVERY session whether or not the skill is invoked.
|
||||
# Exactly three parts, in this order: trigger clause, at most one capability
|
||||
# clause, boundary clause. Drop the boundary line if no near-miss skill exists.
|
||||
# Budget: 250 characters target, 400 hard ceiling (counting this value only).
|
||||
# Trigger clause: when should an agent activate this skill? Describe the user's
|
||||
# intent, not the skill's internal mechanics.
|
||||
# Budget: 250 characters target, 400 hard ceiling (counting this value only,
|
||||
# with YAML folding resolved). This scaffold sits at 214 — keep the fill-in
|
||||
# under the target rather than growing past it.
|
||||
# Boundary clauses may be plural: write one per genuine near-miss, and none
|
||||
# where no sibling could steal activations.
|
||||
# Never let a hyphenated skill name wrap across two lines of this folded block
|
||||
# — folding turns the break into a space and the routing target stops resolving.
|
||||
# Banned here: capability lists, output-format detail, composition notes,
|
||||
# implementation detail, and restating one trigger twice in two registers.
|
||||
# The boundary target must resolve to a real skill or agent — it is checked.
|
||||
|
||||
@@ -47,11 +47,15 @@ explicitly" only where the user's natural phrasing genuinely omits the domain wo
|
||||
for `git-commits`, where the user says "commit". Adding one everywhere is what inflated this
|
||||
corpus, and it was deleted as a blanket rule.
|
||||
|
||||
**Boundary targets must resolve.** The name after the arrow is checked against real skill
|
||||
directories under `plugins/*/.apm/skills/<name>/` and real agents under
|
||||
`plugins/*/.apm/agents/<name>.agent.md`. A boundary clause naming a target that does not exist
|
||||
sends the router nowhere and fails the audit. Check the target exists before writing it — do not
|
||||
invent a plausible sibling name.
|
||||
**Boundary targets must resolve.** Both forms are checked — the arrow and the prose form ("do not
|
||||
use for X, use `y` instead") — so a typo dangles either way. Targets resolve against a universe
|
||||
built by walking up **from the SKILL.md itself**: the nearest ancestor holding
|
||||
`plugins/*/.apm/{skills,agents}` (or, failing that, the nearest ancestor holding `.git`) contributes
|
||||
every skill and agent under `<root>/plugins/*/`, plus the skill's own apm package and the packages
|
||||
that package declares in `apm.yml` under `dependencies.apm`. A sibling plugin in the same monorepo
|
||||
therefore resolves; a skill in an unrelated repo does not. A boundary clause naming a target
|
||||
outside that universe sends the router nowhere and fails the audit. Check the target exists before
|
||||
writing it — do not invent a plausible sibling name.
|
||||
|
||||
**Length.** 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only
|
||||
with YAML folding resolved. The agentskills.io 1,024-character spec limit is unchanged and sits
|
||||
@@ -61,7 +65,13 @@ as the outlier stop.
|
||||
**Hand-invoked skills are exempt.** A skill carrying `disable-model-invocation: true` is absent
|
||||
from the model-visible listing and is reached only by the user typing `/name`. It takes one plain
|
||||
human-facing sentence — no trigger clause, no boundary clause, no indirect triggers. Worked
|
||||
example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`.
|
||||
example — the whole description of the `zoom-out` skill, which carries `disable-model-invocation`:
|
||||
|
||||
````markdown
|
||||
Tell the agent to zoom out and give broader context or a higher-level perspective. Use when
|
||||
you're unfamiliar with a section of code or need to understand how it fits into the bigger
|
||||
picture.
|
||||
````
|
||||
|
||||
## Body
|
||||
|
||||
@@ -90,6 +100,12 @@ blocks, rationale prose, and any content only one branch reaches. Each reference
|
||||
self-contained for its concern, and every one is wired from the body with the literal conditional
|
||||
form:
|
||||
|
||||
**The one exception, stated once so it is not re-litigated:** an output schema stays in the body
|
||||
only when it applies to *every* flow and is short — roughly 50 words or less, which is the "Output
|
||||
format template" pattern below. An output schema that is longer than that, or that only one flow
|
||||
produces, moves to `references/` like any other schema. No third option exists, and the two rules
|
||||
do not disagree.
|
||||
|
||||
````markdown
|
||||
If <condition>, read `references/<file>.md`.
|
||||
````
|
||||
@@ -98,8 +114,9 @@ A generic pointer ("see references/ for details") is a Vale error — the agent
|
||||
|
||||
**Dispatch is mandatory at two or more mutually exclusive flows.** The body carries the dispatch
|
||||
table and the gates common to every branch; each flow gets its own self-contained `references/`
|
||||
file. Exemplar: `plugins/kyberforge/.apm/skills/apm-workflow/SKILL.md` — a 554-word body
|
||||
dispatching to 3,006 words of references.
|
||||
file. Exemplar: the `apm-workflow` skill — a **421-word body** dispatching to 3,006 words of
|
||||
references. Calibrate against 421: that file's whole-file count is 554 words, and aiming at that
|
||||
number instead overshoots the body budget by ~30%.
|
||||
|
||||
**Length.** 600 words SUGGESTION, 900 words FAIL, counting the **body only** — everything after
|
||||
the frontmatter's closing `---`.
|
||||
@@ -108,7 +125,7 @@ the frontmatter's closing `---`.
|
||||
|
||||
- Each entry must state a fact that **contradicts a reasonable default** — something the agent
|
||||
gets wrong by acting sensibly. "Never commit secrets" is not one; the agent already knows.
|
||||
- Maximum five entries.
|
||||
- More than five entries is a SUGGESTION — five is the guideline, not a ceiling.
|
||||
- A Gotcha that paraphrases a step in the body below it is a **FAIL**. If the rule is already a
|
||||
step, it is not a gotcha.
|
||||
- A Gotchas section exceeding 25% of the body is a SUGGESTION.
|
||||
@@ -164,7 +181,9 @@ Do not modify flags.
|
||||
| <condition> | <flow> | `references/<file>.md` |
|
||||
````
|
||||
|
||||
**Output format template** (when the skill produces structured output):
|
||||
**Output format template** (when the skill produces structured output on *every* flow, and the
|
||||
schema is roughly 50 words or less — see the exception under Body above; anything longer or
|
||||
flow-specific belongs in `references/`):
|
||||
|
||||
````markdown
|
||||
Output format:
|
||||
@@ -173,7 +192,8 @@ Output format:
|
||||
```
|
||||
````
|
||||
|
||||
For longer templates, place them in `assets/<name>.md` and reference conditionally.
|
||||
For longer templates, place them in `references/<topic>.md` or `assets/<name>.md` and reference
|
||||
conditionally.
|
||||
|
||||
## Embedding org-specific policy
|
||||
|
||||
|
||||
@@ -136,10 +136,16 @@ If no scripts are needed, delete `scripts/README.md` and the `scripts/` director
|
||||
|
||||
## Step 5 — Add references, assets, and tests (if needed)
|
||||
|
||||
**`references/`** — additional documentation loaded on demand. One topic per file. Reference
|
||||
conditionally from SKILL.md with the literal form ``If <condition>, read `references/<file>.md` ``.
|
||||
Keep reference chains one level deep — a reference file that references another reference file is
|
||||
rarely loaded correctly.
|
||||
**`references/`** — additional documentation loaded on demand. One topic per file, named in
|
||||
kebab-case after the topic. Reference conditionally from SKILL.md with the literal form
|
||||
``If <condition>, read `references/<file>.md` ``.
|
||||
|
||||
**Two hops from `SKILL.md`, never three.** A flow file may route on to a shared contract or
|
||||
sub-topic file — that is the shipped pattern here (`SKILL.md` → `references/create.md` → this
|
||||
file's own pointers to `contract.md`, `scripts.md` and `deployment-modes.md`). What does not work
|
||||
is a third hop: a file reachable only through two intermediates is rarely loaded at the moment it
|
||||
is needed. Every hop past the first also needs the same literal conditional form, so the agent
|
||||
knows when to take it.
|
||||
|
||||
**`assets/`** — static resources: templates, schemas, lookup tables. Reference by relative path
|
||||
from SKILL.md.
|
||||
|
||||
@@ -76,6 +76,12 @@ contract first — the gates are hot and carry no baseline file, so a one-line f
|
||||
non-compliant skill cannot be committed until the description and body meet
|
||||
`references/contract.md`. Treat that retrofit as part of the same change, not a follow-up.
|
||||
|
||||
If the skill's description exceeds 250 characters, or its body-only word count exceeds 600, read
|
||||
`references/retrofit.md` before editing. It carries the ordered cut procedure, the
|
||||
mutually-exclusive-flows test, the reference-file conventions this flow needs, the collateral
|
||||
checklist for `README.md` and `references/sources.md`, and a worked description retrofit. Do not
|
||||
improvise the cuts — four dry runs invented six to ten different answers to the same questions.
|
||||
|
||||
If a signal points to a script or reference file, edit that file directly rather than adding a
|
||||
workaround in SKILL.md.
|
||||
|
||||
|
||||
@@ -0,0 +1,152 @@
|
||||
---
|
||||
source_keys:
|
||||
- agentskills-best-practices
|
||||
- agentskills-optimizing-descriptions
|
||||
---
|
||||
|
||||
# Retrofitting a skill to the ADR-0020 contract
|
||||
|
||||
Read this when `references/improve.md` Step 4 sends you here: the skill you are editing is over
|
||||
the description or body budget and has to come into contract before any other change can be
|
||||
committed. The gates are hot and carry no baseline file, so a one-line fix to a non-compliant
|
||||
skill is blocked until this is done.
|
||||
|
||||
Measure first. Do not guess which gate fired: run `/skill-audit` on the directory and read its
|
||||
`### Structure` dimension, which reports the description characters and the **body-only** word
|
||||
count separately from the whole-file spec backstop. Retrofit against the number that actually
|
||||
fired — a skill can sit a thousand words inside the whole-file backstop while failing the body
|
||||
budget.
|
||||
|
||||
**Validate in place.** Audit the skill's real directory inside its package. Never audit a copy in a
|
||||
scratch directory, and never move a skill out to work on it: the boundary-target universe is built
|
||||
by walking up *from the file being checked*, so a copy with no authoring root above it resolves
|
||||
against nothing and the check declines rather than running —
|
||||
|
||||
```text
|
||||
INFO boundary-target resolution DID NOT RUN — no skill universe could be determined for
|
||||
this path ... Unchecked target(s): totally-fake-target
|
||||
```
|
||||
|
||||
The run still exits 0, so that line reads as a pass and is not one. Treat `DID NOT RUN` as **not
|
||||
checked**, always. A retrofit signed off on a scratch copy carries an unverified boundary target
|
||||
into the corpus, which is precisely the failure this gate exists to catch.
|
||||
|
||||
## Cut in this order
|
||||
|
||||
Work the list top down and stop as soon as the gate clears. The order is by ratio of tokens
|
||||
removed to behaviour lost — inverting it is how a retrofit ends up deleting the one instruction
|
||||
the skill existed to carry.
|
||||
|
||||
1. **Gotchas that paraphrase a step in the body below.** Zero information, and already a FAIL on
|
||||
its own. Delete the Gotcha, keep the step.
|
||||
2. **Spec restatements** — text that repeats a published specification, a tool's `--help`, or a
|
||||
ceiling the validator already enforces. The agent gets this right without it. Delete, or move
|
||||
the table to `references/` if a flow genuinely needs to look it up.
|
||||
3. **Capability enumeration** — in a description, the feature list after the trigger clause; in a
|
||||
body, the paragraph that recites what the skill can do. One capability clause survives in the
|
||||
description; the rest belongs in `README.md`.
|
||||
4. **Per-flow prose** — anything only one branch of the procedure ever reaches. This is the
|
||||
largest single win in most bodies, and it is a *move*, not a delete: each flow gets its own
|
||||
self-contained `references/` file, wired from a dispatch table.
|
||||
|
||||
If the body is still over after all four, the skill is doing two jobs. Split it, and say so
|
||||
rather than compressing prose until it stops being readable.
|
||||
|
||||
## What "mutually exclusive flows" means
|
||||
|
||||
Two or more flows that a single invocation cannot both take. The three-way test, copied verbatim
|
||||
from the body-discipline rubric `/skill-audit` judges against — nothing to load, it is quoted in
|
||||
full here:
|
||||
|
||||
> separate subcommands, separate input types, separate lifecycle stages
|
||||
|
||||
Any one of the three is enough. Two flows that differ only in a parameter value are one flow.
|
||||
At two or more mutually exclusive flows a dispatch table is **mandatory** regardless of word
|
||||
count, because every invocation otherwise pays for every branch it did not take.
|
||||
|
||||
## Reference-file conventions
|
||||
|
||||
The create flow owns these rules, and this flow is forbidden from reading `references/create.md`,
|
||||
so what a retrofit needs is restated here:
|
||||
|
||||
- **One topic per file.** A file mixing two concerns gets loaded for one of them and spends the
|
||||
caller's context on the other.
|
||||
- **Kebab-case filenames**, named after the topic rather than the flow that reads it —
|
||||
`body-discipline.md`, not `step-3.md`.
|
||||
- **Wire every file with the literal conditional form** ``If <condition>, read
|
||||
`references/<file>.md` ``. A generic pointer ("see `references/` for details") is a Vale error.
|
||||
- **Two hops from `SKILL.md`, never three.** A flow file may route on to a shared contract file;
|
||||
a file reachable only through two intermediates is rarely loaded when it is needed.
|
||||
- **`source_keys` frontmatter.** If the content you are moving drew on a research source, the new
|
||||
file needs top-level `source_keys:` frontmatter listing those slugs, and every slug must already
|
||||
exist as an `## <slug>` heading in `references/sources.md`. Moving sourced content out of
|
||||
`SKILL.md` without carrying its slugs across breaks the provenance chain, and `/skill-audit`
|
||||
reports the new file as an INFO with no `source_keys`.
|
||||
|
||||
## Collateral is mandatory, not optional
|
||||
|
||||
Moving content out of a `SKILL.md` leaves three files describing a structure that no longer
|
||||
exists. `/skill-audit`'s provenance check exits clean on all three of these, so nothing catches
|
||||
them for you. After every retrofit that adds, removes or renames a file:
|
||||
|
||||
- [ ] **`README.md` file table** — a row for every new `references/` file, and no row left for a
|
||||
file that is gone. Say what triggers the load, not just what the file contains.
|
||||
- [ ] **`references/README.md`**, where the skill has one — same update, same reason.
|
||||
- [ ] **`references/sources.md` → `Contributing files`** — add the new file to every slug whose
|
||||
content moved into it, and remove any file the retrofit deleted. This is the one that gets
|
||||
missed: `sources.md` keeps citing sections of `SKILL.md` that no longer exist, the
|
||||
provenance check still exits 0, and the stale claim survives review.
|
||||
- [ ] Re-run `/skill-audit` and confirm its `### Provenance` dimension does not report the new
|
||||
file as missing `source_keys`.
|
||||
|
||||
## Worked example — a description retrofit
|
||||
|
||||
`gitea-issues` before, 827 characters, the single most common shape in the corpus:
|
||||
|
||||
```text
|
||||
Use when reading or writing Gitea issues: listing repo issues, getting a single issue's details/
|
||||
comments/labels, creating an issue, updating its state, adding or editing comments, applying
|
||||
labels via issue_write, or searching issues/PRs across repositories. Triggers on "create an
|
||||
issue", "what issues are open", "get issue #N", "close issue #N", "comment on issue #N", "search
|
||||
issues for X" — even when the user doesn't say "Gitea" explicitly. Composes gitea-labels-
|
||||
milestones for all label inference/resolution and milestone lookup — do not use this skill to
|
||||
manage label or milestone definitions themselves (create/edit/delete a label, create/close a
|
||||
milestone), that's gitea-labels-milestones directly. Do not use for pull requests (use gitea-prs)
|
||||
or for local git branch/commit work (use gitea-branches or git-branches).
|
||||
```
|
||||
|
||||
After, 240 characters:
|
||||
|
||||
```text
|
||||
Use when reading or writing Gitea issues — list, read, create, comment on, label, close, or
|
||||
search — even when the user does not say "Gitea". Not pull requests -> `gitea-prs`. Not label or
|
||||
milestone definitions -> `gitea-labels-milestones`.
|
||||
```
|
||||
|
||||
What came out, and why:
|
||||
|
||||
| Removed | Why |
|
||||
|---|---|
|
||||
| The second trigger register — `Triggers on "create an issue", "what issues are open", …` | The same triggers restated as quoted user phrasings. Two registers of one trigger list is a FAIL, not a suggestion. |
|
||||
| `applying labels via issue_write` | Implementation detail. The router does not choose a skill by which MCP call it makes. |
|
||||
| `Composes gitea-labels-milestones for all label inference/resolution and milestone lookup` | A composition note. It changes no routing decision and belongs in `README.md`. |
|
||||
| The parenthetical `(create/edit/delete a label, create/close a milestone)` | Capability enumeration inside a boundary clause. The boundary needs the target, not its feature list. |
|
||||
| The `gitea-branches` / `git-branches` boundary | Dropped entirely. Neither was ever going to win an issue request, so the clause defended against nothing — an invented boundary costs characters and buys no routing accuracy. |
|
||||
| `Do not use for pull requests (use gitea-prs)` prose form | Kept, but rewritten as `Not pull requests -> \`gitea-prs\`.` The rewrite buys characters and one uniform shape for the router — not safety. Both forms are parsed **and** target-checked, so a typo in the prose form dangles exactly as an arrow typo does. |
|
||||
|
||||
What stayed: one trigger clause, one capability clause, the indirect trigger (genuinely warranted
|
||||
here — people say "create an issue", not "create a Gitea issue"), and the boundary clauses.
|
||||
|
||||
## Two rules the gates enforce but the prose does not spell out
|
||||
|
||||
**Boundary clauses may be plural.** Write one per genuine near-miss — the example above carries
|
||||
two, because two different skills could each steal activations. "A boundary clause" in the
|
||||
contract means *at least one*, not *exactly one*. What is banned is a boundary clause invented for
|
||||
a skill that was never going to compete, not a second real one.
|
||||
|
||||
**Never let a hyphenated routing target wrap across lines in a folded `>` scalar.** YAML folding
|
||||
replaces the newline with a space, so `gitea-labels-` at the end of one line and `milestones` at
|
||||
the start of the next fold into `gitea-labels- milestones`. `validate.sh` then reads the target as
|
||||
`gitea-labels`, finds no such skill, and reports a dangling boundary target — the live finding on
|
||||
`gitea-issues` today. Reflow the line so the whole name sits on one of them. The same applies to
|
||||
any backticked skill or agent name in a description.
|
||||
@@ -34,7 +34,7 @@ source_keys:
|
||||
- **URL:** https://agentskills.io/skill-creation/best-practices.md
|
||||
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
|
||||
- **Description:** Best practices for skill creators — starting from real expertise, spending context wisely, calibrating control, instruction patterns (gotchas, templates, checklists, validation loops)
|
||||
- **Contributing files:** SKILL.md, references/create.md, references/improve.md, references/contract.md
|
||||
- **Contributing files:** SKILL.md, references/create.md, references/improve.md, references/contract.md, references/retrofit.md
|
||||
- **Status:** `extracted`
|
||||
|
||||
## agentskills-optimizing-descriptions
|
||||
@@ -42,7 +42,7 @@ source_keys:
|
||||
- **URL:** https://agentskills.io/skill-creation/optimizing-descriptions.md
|
||||
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
|
||||
- **Description:** How to systematically test and improve skill descriptions for triggering accuracy — eval queries, trigger rate testing, train/validation splits, optimization loop
|
||||
- **Contributing files:** SKILL.md, references/improve.md, references/contract.md
|
||||
- **Contributing files:** SKILL.md, references/improve.md, references/contract.md, references/retrofit.md
|
||||
- **Status:** `extracted`
|
||||
|
||||
## agentskills-evaluating-skills
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "kyberforge",
|
||||
"version": "1.5.0",
|
||||
"version": "1.6.0",
|
||||
"description": "Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.",
|
||||
"author": {
|
||||
"name": "Defame1297",
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "kyberforge",
|
||||
"version": "1.5.0",
|
||||
"version": "1.6.0",
|
||||
"description": "Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.",
|
||||
"author": {
|
||||
"name": "Defame1297",
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
name: kyberforge
|
||||
version: 1.5.0
|
||||
version: 1.6.0
|
||||
description: Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.
|
||||
author:
|
||||
name: Defame1297
|
||||
|
||||
@@ -76,6 +76,9 @@ Include what the fresh context lacks:
|
||||
- A direct role instruction opening the prompt: `You are a [role]. When invoked, [action].`
|
||||
- One bounded job, stated so the agent knows what it must refuse.
|
||||
- The dispatch, gates, inputs and outputs listed above.
|
||||
- **Error handling** — what the agent does on malformed, missing or contradictory input: stop and
|
||||
report, or degrade to a named fallback. Absent it, the agent invents a recovery, and a
|
||||
subagent's invented recovery is invisible to its caller until the output is wrong.
|
||||
- Non-obvious environment facts and project-specific conventions it cannot infer.
|
||||
- One default per decision point with one escape hatch.
|
||||
|
||||
@@ -111,6 +114,8 @@ Flag as FAIL if:
|
||||
Flag as SUGGESTION if:
|
||||
|
||||
- The body does not open with a direct role instruction
|
||||
- The body specifies no error handling — nothing tells the agent what to do with malformed,
|
||||
missing or contradictory input
|
||||
- The job the agent describes is unbounded, or bounded only implicitly
|
||||
- A rationale is missing from a rule the agent is expected to enforce — present but unexplained
|
||||
- Comments are useful but verbose enough to bury the field they annotate
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -34,10 +34,13 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
|
||||
boundary clause naming a real sibling skill or agent.
|
||||
250 characters is the target, 400 the hard ceiling (ADR-0020).
|
||||
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was
|
||||
deleted. Add "Use proactively" only if the runtime should delegate here without
|
||||
the user naming this agent.
|
||||
deleted.
|
||||
Never write "Use proactively" here. It steers the Claude Code runtime and does
|
||||
nothing anywhere else, and this file compiles to a Copilot `.agent.md` too, where
|
||||
agent-audit's KyberforgeCopilot.ProactivePhrase rule grades it a hard FAIL.
|
||||
The phrase is CC-only; at this scope, a precise trigger clause does that job.
|
||||
Example: "Use when a diff needs checking for injected credentials before it
|
||||
merges. Not general code review -> code-reviewer." -->
|
||||
merges. Not prose or style linting -> `lint-runner`." -->
|
||||
|
||||
<!-- model: sonnet
|
||||
Optional. Aliases: sonnet, opus, haiku, fable. Or full model ID.
|
||||
@@ -80,3 +83,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
|
||||
## Output
|
||||
|
||||
FILL IN: What does the agent produce? Format, location, structure.
|
||||
|
||||
## Errors
|
||||
|
||||
FILL IN: What does the agent do on malformed, missing or contradictory input?
|
||||
State whether it stops and reports, or degrades to a named fallback — and what it
|
||||
tells the caller either way. An agent with no error handling invents a recovery,
|
||||
and an invented recovery is invisible until the output is wrong.
|
||||
|
||||
@@ -14,10 +14,14 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
|
||||
boundary clause naming a real sibling skill or agent.
|
||||
250 characters is the target, 400 the hard ceiling (ADR-0020).
|
||||
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was
|
||||
deleted. Add "Use proactively" only if the runtime should delegate here without
|
||||
the user naming this agent.
|
||||
deleted.
|
||||
"Use proactively" is valid HERE and only here: it steers the Claude Code runtime
|
||||
to offer this agent unprompted. Add it only if that is what you want. If you add
|
||||
it, leave it OUT of the Copilot half of the pair — the phrase does nothing there
|
||||
and agent-audit's KyberforgeCopilot.ProactivePhrase grades it a hard FAIL. The
|
||||
pair must describe the same job; it does not have to be byte-identical.
|
||||
Example: "Use when a diff needs checking for injected credentials before it
|
||||
merges. Not general code review -> code-reviewer." -->
|
||||
merges. Not prose or style linting -> `lint-runner`." -->
|
||||
|
||||
<!-- tools: Read, Bash, Grep
|
||||
Optional. Allowlist of tool names: a comma-separated string or a YAML list.
|
||||
@@ -99,3 +103,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
|
||||
## Output
|
||||
|
||||
FILL IN: What does the agent produce? Format, location, structure.
|
||||
|
||||
## Errors
|
||||
|
||||
FILL IN: What does the agent do on malformed, missing or contradictory input?
|
||||
State whether it stops and reports, or degrades to a named fallback — and what it
|
||||
tells the caller either way. An agent with no error handling invents a recovery,
|
||||
and an invented recovery is invisible until the output is wrong.
|
||||
|
||||
@@ -19,9 +19,13 @@ description: FILL IN: Use when <trigger>. <One capability clause.> Not <thing> -
|
||||
boundary clause naming a real sibling skill or agent.
|
||||
250 characters is the target, 400 the hard ceiling (ADR-0020).
|
||||
Do not open with an action verb ("Reviews...", "Analyzes...") — that rule was deleted.
|
||||
Keep it identical in wording to the Claude Code half of the pair.
|
||||
Never write "Use proactively" here. It steers the Claude Code runtime and does nothing
|
||||
in Copilot, and agent-audit's KyberforgeCopilot.ProactivePhrase grades it a hard FAIL.
|
||||
Otherwise keep the wording matched to the Claude Code half of the pair: agent-audit
|
||||
checks that both halves describe the same job, not that they are byte-identical, so
|
||||
dropping the CC-only phrase here is not a pair-consistency finding.
|
||||
Example: "Use when a diff needs checking for injected credentials before it
|
||||
merges. Not general code review -> code-reviewer." -->
|
||||
merges. Not prose or style linting -> `lint-runner`." -->
|
||||
|
||||
<!-- tools: ["read", "search", "edit"]
|
||||
Optional. Array of tool names. Omit = all available tools. [] = no tools.
|
||||
@@ -68,3 +72,10 @@ FILL IN: Steps the agent takes. Be specific about ordering if it matters.
|
||||
## Output
|
||||
|
||||
FILL IN: What does the agent produce? Format, location, structure.
|
||||
|
||||
## Errors
|
||||
|
||||
FILL IN: What does the agent do on malformed, missing or contradictory input?
|
||||
State whether it stops and reports, or degrades to a named fallback — and what it
|
||||
tells the caller either way. An agent with no error handling invents a recovery,
|
||||
and an invented recovery is invisible until the output is wrong.
|
||||
|
||||
@@ -44,15 +44,31 @@ Banned from a description; move it to the body or to `README.md`:
|
||||
and ADR-0020 deleted it: the opener is `Use when`, matching every skill in this corpus, so one
|
||||
router reads one shape.
|
||||
|
||||
**"Use proactively" is conditional.** Add it only where the runtime should delegate without the
|
||||
user naming the agent — an agent invoked by name does not need it, and it costs activations
|
||||
elsewhere when added by reflex. The same conditional governs indirect triggers ("even if the user
|
||||
doesn't say X"): add one only where the user's natural phrasing genuinely omits the domain word.
|
||||
**"Use proactively" is Claude Code-only, and conditional even there.** The phrase steers the
|
||||
Claude Code runtime to offer an agent unprompted and does nothing anywhere else, so where it may
|
||||
appear depends on the file:
|
||||
|
||||
**Boundary targets must resolve.** The name after the arrow is checked against real skills under
|
||||
`plugins/*/.apm/skills/<name>/` and real agents under `plugins/*/.apm/agents/<name>.agent.md`. A
|
||||
target that does not exist sends the router nowhere. Verify it before writing it — do not invent a
|
||||
plausible sibling.
|
||||
| File | Rule |
|
||||
|---|---|
|
||||
| Claude Code `.md` (project/user scope) | Allowed. Add it only where the runtime should delegate without the user naming the agent — an agent invoked by name does not need it, and it costs activations elsewhere when added by reflex. |
|
||||
| Copilot `.agent.md` (project/user scope) | **Never.** Inert there, and `KyberforgeCopilot.ProactivePhrase` grades it a hard FAIL. |
|
||||
| Vendor-neutral `.apm/agents/<name>.agent.md` (plugin/APM scope) | **Never.** Same Vale rule, same hard FAIL — the file matches the `**/*.agent.md` glob, and it compiles to a real Copilot agent downstream. |
|
||||
|
||||
A pair whose Claude Code half carries the phrase and whose Copilot half omits it is correct, not
|
||||
inconsistent: `agent-audit` checks that both halves describe the same job, not that they match
|
||||
word for word.
|
||||
|
||||
Indirect triggers ("even if the user doesn't say X") take a similar conditional at every scope:
|
||||
add one only where the user's natural phrasing genuinely omits the domain word.
|
||||
|
||||
**Boundary targets must resolve.** Both forms are checked — the arrow and the prose form ("do not
|
||||
use for X, use `y` instead") — so a typo dangles either way. Targets resolve against a universe
|
||||
built by walking up **from the agent file itself**: the nearest ancestor holding
|
||||
`plugins/*/.apm/{skills,agents}` (or, failing that, the nearest ancestor holding `.git`) contributes
|
||||
every skill and agent under `<root>/plugins/*/`, plus the agent's own apm package and the packages
|
||||
that package declares in `apm.yml` under `dependencies.apm`. A sibling plugin in the same monorepo
|
||||
therefore resolves; a skill in an unrelated repo does not. A target outside that universe sends the
|
||||
router nowhere. Verify it before writing it — do not invent a plausible sibling.
|
||||
|
||||
**Length.** 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only
|
||||
with YAML folding resolved. Treat 250 as the target: the SUGGESTION tier is what moves the corpus
|
||||
@@ -73,8 +89,18 @@ You are a <role>. When invoked, <primary action>.
|
||||
|
||||
## Output
|
||||
<what it produces: format, location, structure>
|
||||
|
||||
## Errors
|
||||
<what to do on malformed, missing or contradictory input: report and stop, or
|
||||
which fallback to take — and what to say to the caller either way>
|
||||
````
|
||||
|
||||
Four required elements: **inputs expected, process steps, output format, error handling.** The
|
||||
last is the one that gets dropped, and dropping it is not neutral: an agent given a malformed
|
||||
input and no instruction invents a recovery, and a subagent's invented recovery is invisible to
|
||||
the caller until the output is wrong. Say explicitly whether the agent stops and reports, or
|
||||
degrades to a named fallback.
|
||||
|
||||
One job per agent. An agent covering two jobs gets delegated to for the wrong one.
|
||||
|
||||
**Delegation discipline replaces the word gate.** A plugin/APM agent is a single file with no
|
||||
|
||||
@@ -63,5 +63,7 @@ to installed skills instead of transcribed procedure.
|
||||
- [ ] No `FILL IN:` placeholder and no `<!-- ... -->` template comment anywhere in the file
|
||||
- [ ] System prompt body non-empty, and a read-only agent says so in prose as well as in
|
||||
`disallowedTools`
|
||||
- [ ] Body covers all four required elements: inputs expected, process steps, output format,
|
||||
**error handling** — what the agent does on malformed, missing or contradictory input
|
||||
|
||||
Then return to the flow reference you came from.
|
||||
|
||||
@@ -95,6 +95,8 @@ Both files:
|
||||
|
||||
- [ ] `name` present and kebab-case; `description` written to `references/contract.md`
|
||||
- [ ] System prompt body present, non-empty and equivalent across the pair
|
||||
- [ ] Body covers all four required elements: inputs expected, process steps, output format,
|
||||
**error handling** — what the agent does on malformed, missing or contradictory input
|
||||
- [ ] No `FILL IN:` placeholder and no `<!-- ... -->` template comment left
|
||||
|
||||
Copilot file only:
|
||||
|
||||
@@ -285,6 +285,8 @@ else
|
||||
if [[ "$SCOPE" == "plugin" ]]; then
|
||||
echo " 1. Fill in $APM_FILE — replace every FILL IN: placeholder. Optional fields are" >&2
|
||||
echo " scaffolded there as commented blocks; uncomment the ones that apply." >&2
|
||||
echo " Description: 250 chars target / 400 ceiling (ADR-0020). The body has no" >&2
|
||||
echo " word gate — delegate to a skill instead of restating what it does." >&2
|
||||
echo " 2. Populate $SOURCES_DIR/sources.md with research sources, or delete it" >&2
|
||||
echo " 3. Validate: $VALIDATE_HINT $APM_FILE" >&2
|
||||
echo " It checks the frontmatter against the apm-agent-allowlist section of" >&2
|
||||
@@ -292,6 +294,8 @@ else
|
||||
else
|
||||
echo " 1. Fill in $CC_FILE — replace every FILL IN: placeholder. Optional fields are" >&2
|
||||
echo " scaffolded there as commented blocks; uncomment the ones that apply." >&2
|
||||
echo " Description: 250 chars target / 400 ceiling (ADR-0020). The body has no" >&2
|
||||
echo " word gate — delegate to a skill instead of restating what it does." >&2
|
||||
echo " 2. Fill in $CP_FILE — same, and heed its closing comment: the Claude Code-only" >&2
|
||||
echo " fields it names must not cross over from the file above." >&2
|
||||
echo " 3. Validate: run $VALIDATE_HINT on each file" >&2
|
||||
|
||||
@@ -6,10 +6,12 @@ Audit a skill directory against the agentskills.io specification and the house c
|
||||
|
||||
1. Runs `scripts/validate.sh` and `scripts/validate-provenance.sh` for structural and provenance checks, plus `scripts/vale-wrap.sh` — a Vale prefilter that deterministically flags non-imperative description openers, composition and architecture notes, vague wording, padding phrases, and "There is/are" sentence openers
|
||||
2. Reads all files in the skill directory
|
||||
3. Applies qualitative checks across six dimension groups, loading one rubric from `references/` per group
|
||||
3. Applies qualitative checks across five dimension groups, loading one rubric from `references/` per group
|
||||
4. Outputs a compact findings report — findings only, grouped by dimension, each with Why and Fix — and a result block with handoff to `skill-author`
|
||||
|
||||
`validate.sh` enforces two independent length families that must not be conflated: the agentskills.io spec conformance ceilings (500 lines, 2,770 words, both counting the whole file) and the ADR-0020 context budget (250/400 description characters, 600/900 body-only words, plus resolvable boundary targets).
|
||||
`validate.sh` enforces two independent length families that must not be conflated: the agentskills.io spec conformance ceilings (500 lines, 2,770 words, both counting the whole file) and the ADR-0020 context budget (250/400 description characters, 600/900 body-only words).
|
||||
|
||||
Alongside those it runs four shape checks that are not length measurements at all. Two are FAILs: every routing target named in the description — in the compressed `Not <thing> -> <name>` arrow **and** in the prose form — must resolve to a real skill or agent, and every `references/<file>.md` the body names must exist on disk. Three are SUGGESTIONs: a missing boundary clause, a Gotchas section over five entries, and a Gotchas section over 25% of the body. The resolution universe for boundary targets is derived by walking up from the audited `SKILL.md` — the authoring root above it, its own apm package, and that package's declared `apm.yml` dependencies — so a fresh clone and a machine that has run `apm install` return the same verdict. When no universe can be determined the check prints `INFO ... DID NOT RUN` and does not silently pass.
|
||||
|
||||
## Usage
|
||||
|
||||
@@ -24,7 +26,7 @@ Provide the path to the skill directory to audit when invoking.
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `SKILL.md` | Skill instructions for agents |
|
||||
| `scripts/validate.sh` | Structural validator — checks name format, name matches directory, description length, line count, placeholder detection, script executable bit, and interactive-prompt detection |
|
||||
| `scripts/validate.sh` | Structural validator — checks name format, name matches directory, description presence and length, body-only word count, line and whole-file word ceilings, boundary-clause presence, boundary-target resolution, `references/` pointer existence, Gotchas entry count and body share, placeholder detection, script executable bit, and interactive-prompt detection |
|
||||
| `scripts/validate-provenance.sh` | Provenance validator — checks sources.md completeness, source_keys/slug consistency, Contributing files existence, bidirectional linkage, Research doc: fields, and upstream research doc alignment |
|
||||
| `scripts/vale-wrap.sh` | Vale prefilter wrapper — runs the bundled `Kyberforge` Vale styles against SKILL.md and reports alerts as deterministic FAILs ahead of Step 3's qualitative review |
|
||||
| `assets/vale/.vale.ini` | Vale configuration — points Vale at the bundled `Kyberforge` style path, self-located relative to `vale-wrap.sh` |
|
||||
@@ -38,6 +40,7 @@ Provide the path to the skill directory to audit when invoking.
|
||||
| `references/patterns.md` | Rubric for the patterns dimension — which instruction construct fits which job, and how each is correctly formed |
|
||||
| `references/file-structure.md` | Rubric for the file-structure and internal-consistency dimensions — permitted directories, cross-plugin path rules and their two structural exemptions, README drift |
|
||||
| `references/formatting-and-scripts.md` | Rubric for the formatting and scripts dimensions — heading and fencing conventions, and the agentic-use criteria for bundled scripts |
|
||||
| `references/validation-scripts.md` | Step 1 troubleshooting — the manual structural fallback when `validate.sh` cannot run, and the script exit codes that are easy to misread (loaded only on a script failure) |
|
||||
| `references/sources.md` | Provenance record — agentskills.io sources that informed this skill and which files each contributed to |
|
||||
| `tests/validate.bats` | (source-only) Bats test suite for validate.sh |
|
||||
| `tests/validate-provenance.bats` | (source-only) Bats test suite for validate-provenance.sh |
|
||||
|
||||
@@ -33,7 +33,9 @@ bash scripts/validate-provenance.sh <skill-dir>
|
||||
scripts/vale-wrap.sh <skill-dir>/SKILL.md
|
||||
```
|
||||
|
||||
`validate.sh` findings become the `### Structure` dimension — its FAILs and its SUGGESTIONs both. If it cannot run at all (no `python3`, Bash denied), report that as an INFO finding rather than guessing; what it measures is not reproducible by reading.
|
||||
`validate.sh` findings become the `### Structure` dimension — its FAILs and its SUGGESTIONs both.
|
||||
|
||||
If any of the three fails, cannot run, or reports something needing interpretation, read `references/validation-scripts.md` — it carries the manual fallback and the misleading exit codes.
|
||||
|
||||
`validate-provenance.sh` prints nothing on success. Its FAIL and INFO findings become a separate `### Provenance` dimension, and it emits Why and Fix itself — surface those verbatim.
|
||||
|
||||
|
||||
@@ -69,8 +69,10 @@ table** plus the gates common to every branch, and each flow lives in its own se
|
||||
`references/` file. Inlining all of them is a FAIL regardless of word count, because every
|
||||
invocation then pays for every branch it did not take.
|
||||
|
||||
The reference shape in this repo is `apm-workflow`: a 554-word body dispatching to roughly 3,000
|
||||
words of references across five mutually exclusive invocations.
|
||||
The reference shape in this repo is `apm-workflow`: a **421-word body** dispatching to roughly
|
||||
3,000 words of references across five mutually exclusive invocations. Its whole-file count is 554
|
||||
words — cite 421 when calibrating a body, or the conflation this section warns against reappears
|
||||
in the finding itself.
|
||||
|
||||
## Gotchas sections
|
||||
|
||||
@@ -86,17 +88,20 @@ sensibly.
|
||||
|
||||
Constraints:
|
||||
|
||||
- **Maximum five entries.** Past five, the section is a summary of the body rather than a set of
|
||||
traps, and the agent stops reading it as a warning.
|
||||
- **More than five entries is a SUGGESTION** — five is the guideline, not a ceiling. Past five, the
|
||||
section is usually a summary of the body rather than a set of traps, and the agent stops reading
|
||||
it as a warning. It stays advisory because whether a given gotcha earns its place is judgment;
|
||||
`validate.sh` emits it through `suggest()` and the run still exits 0.
|
||||
- **A Gotcha that paraphrases a step in the body below it is a FAIL.** It has no independent
|
||||
content, and it teaches the agent that Gotchas can be skimmed because the real instruction is
|
||||
coming.
|
||||
coming. This one is the auditor's call — no script detects it.
|
||||
- **A Gotchas section exceeding 25% of the body is a SUGGESTION** — the body has been inverted into
|
||||
a preamble.
|
||||
a preamble. Same tier and same reasoning as the entry count, and independent of it: either can
|
||||
fire without the other.
|
||||
- Place the section near the top. A gotcha read after the mistake is worthless, which is also why
|
||||
Gotchas is the one construct exempt from moving to `references/`.
|
||||
|
||||
Worked negative example — `git-commits` carries thirteen entries, of which four restate content
|
||||
Worked negative example — `git-commits` carries twelve entries, of which four restate content
|
||||
that already appears below or in the description:
|
||||
|
||||
| Gotcha | Restates |
|
||||
@@ -106,8 +111,10 @@ that already appears below or in the description:
|
||||
| `:33` "Never skip hooks with `--no-verify`" | step 9 at `:52` |
|
||||
| `:36` "Never commit secrets" | step 2 at `:45` |
|
||||
|
||||
All four are FAILs under this rule, and the section as a whole breaches the five-entry maximum. It
|
||||
also passes every plausible word gate, which is the point of auditing the construct directly.
|
||||
All four are FAILs under the paraphrase rule. The entry count and the section's share of the body
|
||||
(387 of 1,102 words, 35%) are two further SUGGESTIONs on top — the script reports both, and neither
|
||||
fails the run on its own. What makes this worth auditing directly is that the four paraphrase FAILs
|
||||
pass every word gate there is; only reading the construct finds them.
|
||||
|
||||
## Calibrating control
|
||||
|
||||
@@ -144,7 +151,7 @@ Flag as FAIL if:
|
||||
- A sentence answers "no" to the core test — it is padding
|
||||
- The body exceeds 900 words counted body-only (`validate.sh` reports it)
|
||||
- Two or more mutually exclusive flows are inlined instead of dispatched
|
||||
- A Gotcha paraphrases a step in the body below it, or the section exceeds five entries
|
||||
- A Gotcha paraphrases a step in the body below it
|
||||
- A decision point presents a menu of options with no default
|
||||
- An instruction repeats content already in the description
|
||||
- A prescriptive sequence is used where flexibility is fine, or the reverse
|
||||
@@ -152,6 +159,7 @@ Flag as FAIL if:
|
||||
Flag as SUGGESTION if:
|
||||
|
||||
- The body exceeds 600 words counted body-only but stays at or under 900
|
||||
- The Gotchas section carries more than five entries
|
||||
- The Gotchas section exceeds 25% of the body
|
||||
- A rationale is missing from an include/exclude rule — present but unexplained
|
||||
- Gotchas are correct but placed late in the body rather than near the top
|
||||
|
||||
@@ -28,6 +28,14 @@ resolving there. Flag any `../`, `../../`, or absolute repo path (`plugins/<plug
|
||||
and its APM-native equivalent `.apm/skills/<other>/`) appearing in `SKILL.md`, `scripts/`,
|
||||
`references/` or `assets/`.
|
||||
|
||||
**Referring to another skill's file.** There is one sanctioned spelling, and it is possessive:
|
||||
`skill-audit's references/validation-scripts.md`. Write the skill by name and let the reader
|
||||
resolve it — do not spell the repo path. The full path is the thing this section forbids, and
|
||||
`references/validation-scripts.md` on its own is a hard ERROR from the ADR-0020 gate, which
|
||||
requires an unqualified `references/` pointer to exist in the skill's OWN directory. The
|
||||
possessive form is the only spelling both rules accept; the gate recognises it and skips the
|
||||
on-disk check. Flag any other spelling of a cross-skill reference.
|
||||
|
||||
Two directories are exempt, and the exemptions are structural rather than discretionary:
|
||||
|
||||
- **`references/sources.md`.** Its `Research doc:` fields are development-time provenance pointers,
|
||||
|
||||
@@ -15,7 +15,7 @@
|
||||
- **URL:** https://agentskills.io/specification.md
|
||||
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
|
||||
- **Description:** Complete SKILL.md format specification — frontmatter fields, constraints, body content, optional directories, progressive disclosure levels, file references, validation
|
||||
- **Contributing files:** SKILL.md, references/body-discipline.md, references/description-quality.md, references/patterns.md, references/file-structure.md, references/formatting-and-scripts.md
|
||||
- **Contributing files:** SKILL.md, references/body-discipline.md, references/description-quality.md, references/patterns.md, references/file-structure.md, references/formatting-and-scripts.md, references/validation-scripts.md
|
||||
- **Status:** `extracted`
|
||||
|
||||
## agentskills-best-practices
|
||||
@@ -47,7 +47,7 @@
|
||||
- **URL:** https://agentskills.io/skill-creation/using-scripts.md
|
||||
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
|
||||
- **Description:** Using scripts in skills — one-off commands, self-contained scripts with inline dependencies, designing scripts for agentic use (no interactive prompts, --help, structured output, idempotency)
|
||||
- **Contributing files:** SKILL.md, references/formatting-and-scripts.md
|
||||
- **Contributing files:** SKILL.md, references/formatting-and-scripts.md, references/validation-scripts.md
|
||||
- **Status:** `extracted`
|
||||
|
||||
## agentskills-quickstart
|
||||
|
||||
@@ -0,0 +1,118 @@
|
||||
---
|
||||
source_keys:
|
||||
- agentskills-spec
|
||||
- agentskills-using-scripts
|
||||
---
|
||||
|
||||
# Validation Scripts Reference
|
||||
|
||||
Read this when a Step 1 script fails, cannot run, or reports something that needs interpreting.
|
||||
Nothing here is needed on a clean run.
|
||||
|
||||
## Report the gap, do not guess
|
||||
|
||||
If a script cannot run at all — Bash denied, `python3` unavailable, PyYAML not importable, `vale`
|
||||
not installed — say so as an **INFO** finding naming the script and the missing dependency, then
|
||||
fall back to the manual checks below. An INFO never changes PASS/FAIL. Silently omitting the
|
||||
dimension a script would have covered reports a clean audit that checked less than it claims to
|
||||
have checked, and the Step 4 coverage line then names a dimension nothing actually examined.
|
||||
|
||||
## Manual structural fallback
|
||||
|
||||
`validate.sh` needs `python3` **and** PyYAML, and refuses to start without either — the description
|
||||
value has to be measured after YAML folding is resolved, so skipping the ADR-0020 gates would be a
|
||||
vacuous pass rather than a partial one. The two are checked separately, so the message already names
|
||||
the right one — report it verbatim rather than diagnosing further:
|
||||
|
||||
```text
|
||||
Error: python3 is required but was not found on PATH.
|
||||
Error: PyYAML is required but is not importable by python3.
|
||||
```
|
||||
|
||||
Without them — or with Bash denied, or on a permission error — work this list
|
||||
by hand and file the results under `### Structure` exactly as the script's output would have been:
|
||||
|
||||
- **`name`** present, 1–64 characters, kebab-case (lowercase letters, digits and hyphens; no
|
||||
leading, trailing or doubled hyphen), and **matching the skill's directory name** exactly.
|
||||
- **`description`** present and non-empty; no unfilled `FILL IN:` placeholder in it. An absent or
|
||||
empty description is a **FAIL**, never a silent skip — it is the one field preloaded into every
|
||||
session, so a skill without one can never be routed to.
|
||||
- **Description length**, measured on the folded YAML value with newlines collapsed to single
|
||||
spaces — not on the raw block scalar, which counts indentation. 250 characters SUGGESTION, 400
|
||||
FAIL (ADR-0020), 1,024 FAIL (agentskills.io spec).
|
||||
- **Body length**, counting everything after the frontmatter's closing `---`. 600 words
|
||||
SUGGESTION, 900 FAIL (ADR-0020).
|
||||
- **Whole-file ceilings**, counting the file including frontmatter: 500 lines FAIL, 2,770 words
|
||||
FAIL (agentskills.io spec). These are a different measurement from the two above — report them
|
||||
as separate findings, never merged.
|
||||
- **A boundary clause is present** — either the prose form (`do not` / `instead` / `rather than` /
|
||||
`not for`) or ADR-0020's compressed `Not <thing> -> <name>` arrow. **SUGGESTION**, not FAIL:
|
||||
the absence is deterministic, but whether this skill warrants one is the auditor's call.
|
||||
- **Boundary targets resolve** — **FAIL** on a name that resolves to nothing. See the section
|
||||
below; resolving these by hand is the one item on this list with a procedure of its own.
|
||||
- **Every `references/<file>.md` named in the body exists on disk** — **FAIL**, not a suggestion.
|
||||
A dispatch table or "read X" trigger naming a missing file sends the agent nowhere. Ignore
|
||||
mentions inside fenced code blocks, and ignore a mention whose own line says the file is gone
|
||||
(`removed`, `deleted`, `renamed`, `superseded`, `replaced`, `obsolete`, `deprecated`, `former`,
|
||||
`gone`, `no longer`, `used to`) — that is a historical note, not a dispatch entry.
|
||||
- **Gotchas discipline**, both **SUGGESTION**. Locate the section by a heading that *is* Gotchas
|
||||
(`## Common Gotchas` counts; `## Gotcha handling` and `## Why gotchas matter` do not), running to
|
||||
the next heading at the same level or shallower. More than five top-level entries is one
|
||||
suggestion; a section over 25% of the body word count is a second, independent one. Count
|
||||
entries at column 0 only — an indented child bullet is not an entry — and ignore fenced code
|
||||
blocks for both.
|
||||
- **No unfilled `FILL IN:` placeholder** anywhere in the body.
|
||||
- **Every file in `scripts/`** carries the executable bit and contains no interactive prompt —
|
||||
no bare `read`, no `select`, nothing that blocks on a TTY.
|
||||
|
||||
## Resolving boundary targets by hand
|
||||
|
||||
Targets are read from **both** boundary forms. The compressed `Not <thing> -> <name>` arrow and the
|
||||
prose form are each parsed *and* target-checked, so a typo in prose phrasing fails exactly as an
|
||||
arrow typo does — do not check only the names after an arrow.
|
||||
|
||||
Build the universe by walking up **from the `SKILL.md` under audit**, never from the validator's own
|
||||
location. The nearest ancestor holding `plugins/*/.apm/skills/` or `plugins/*/.apm/agents/` is the
|
||||
authoring root, falling back to the nearest ancestor holding `.git`. When one is found the universe
|
||||
is every skill and agent under `<root>/plugins/*/`, plus the skill's own apm package, plus the
|
||||
packages that package declares in its `apm.yml` under `dependencies.apm`. Deployed `.claude/` and
|
||||
`.agents/` trees are consulted **only** when no authoring root exists — they are gitignored
|
||||
`apm install` output, and reading them would make a fresh clone and a developer machine disagree.
|
||||
|
||||
Three ways to read the result wrong:
|
||||
|
||||
- **A hyphenated name used attributively is not a dangling target.** "Use pre-commit hooks instead
|
||||
of ad-hoc scripts" reads as a route to `pre-commit` on wording alone. What separates a route from
|
||||
prose is grammar: a route target is terminal — followed by punctuation, a conjunction, or a
|
||||
boundary word — whereas a compound modifier is followed by the noun it modifies. A name followed
|
||||
by an ordinary noun still *confirms* a route when it exists, but never raises a FAIL on its own.
|
||||
- **A SUGGESTION-tier unresolved target is not a FAIL you may promote.** Terminal position alone is
|
||||
not evidence of a route: "run `pre-commit` instead", "see `commit-msg`" and "use the clean-up
|
||||
instead" are all terminal and all prose. A prose-form target earns a FAIL only when its own
|
||||
sentence names another target that *does* resolve; otherwise the script reports it and moves on,
|
||||
and so should you. Route notation — `/name` and `-> name` — is exempt and always FAILs, and it is
|
||||
the fix to recommend when the author did mean a route.
|
||||
- **`INFO boundary-target resolution DID NOT RUN` is not a pass.** The script prints it, and exits
|
||||
0, when no universe could be determined for that path — the usual cause being a skill copy
|
||||
audited outside its package. Report it as an INFO naming the unchecked targets and re-run against
|
||||
the real directory; filing it as clean signs off targets nothing verified.
|
||||
|
||||
## Script-specific failures
|
||||
|
||||
- **`validate-provenance.sh` printed nothing.** That is a pass, not a skip. It also exits 0
|
||||
silently when the skill has no `source_keys` and no `references/sources.md` — nothing to
|
||||
validate is not a finding.
|
||||
- **`vale` reports `0 files`.** Treat the pass as NOT RUN, not as clean, and fall back to full
|
||||
Step 3 judgment for the dimensions it would have covered. The bundled `Kyberforge` style is
|
||||
scoped by glob in `assets/vale/.vale.ini`; a file outside those globs is silently not linted.
|
||||
- **`E100 Runtime error ... does not exist` (exit 2) from `vale-wrap.sh`.** An explicit relative
|
||||
`--config` was passed. Pass none: the wrapper locates its own `assets/vale/.vale.ini` from its
|
||||
own path, so a resolved script path plus an unresolved config path produces exactly this. Do not
|
||||
read this exit code as vale being unavailable — that misreading sends the audit down the
|
||||
fallback path while vale was installed and working the whole time.
|
||||
- **The `vale` binary is genuinely absent** (`command not found`). Report one INFO naming it, then
|
||||
fall back to full Step 3 judgment for the description, body-discipline and patterns dimensions —
|
||||
the prefilter's whole coverage. Judge those by rubric rather than dropping them.
|
||||
- **A path argument that does not exist is a hard error** in `vale-wrap.sh`, deliberately: bare
|
||||
`vale` would fall back to reading stdin and print a clean-looking `0 errors ... in stdin`, which
|
||||
the `0 files` guard above does not catch.
|
||||
File diff suppressed because it is too large
Load Diff
@@ -47,6 +47,7 @@ If the destination resolves inside an APM package, read `references/deployment-m
|
||||
| `references/create.md` | The create flow end to end — prerequisites, package-intent gate, scaffold, frontmatter, scripts, references, sources (loaded on demand) |
|
||||
| `references/improve.md` | The improve flow end to end — signal verification, root-cause grouping, announcement, edits (loaded on demand) |
|
||||
| `references/contract.md` | The ADR-0020 description and body contract, the Gotchas constraint, the two size gates, body patterns, and org-policy embedding (loaded on demand) |
|
||||
| `references/retrofit.md` | Bringing a pre-ADR-0020 skill into contract — ordered cut procedure, the mutually-exclusive-flows test, reference-file conventions, the collateral checklist, and a worked description retrofit (loaded from the improve flow when a budget is exceeded) |
|
||||
| `references/deployment-modes.md` | APM package vs standalone differences and self-containment/cache-isolation rules (loaded on demand) |
|
||||
| `references/scripts.md` | Package runners, inline dependency patterns, and full script contract (loaded on demand) |
|
||||
| `references/sources.md` | Upstream research sources and which skill files each contributed to |
|
||||
|
||||
@@ -19,10 +19,9 @@ metadata:
|
||||
|
||||
## Gotchas
|
||||
|
||||
- A skill's `name` and `description` are preloaded into every agent's context every session, invoked or not; the body loads only on invocation. The description is the scarce budget.
|
||||
- The word gates are two different measurements, not one rule with two tiers. The 2,770-word / 500-line spec backstop counts the whole file including frontmatter; Step 3's gate counts the body alone. A file can sit well inside one and fail the other, so never unify them.
|
||||
- Never spawn a subagent to audit or recheck your own work here. Run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent can have its worktree torn down by concurrent cleanup, destroying an uncommitted draft.
|
||||
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires transcript analysis that is out of scope here — flag the opportunity as a suggestion instead.
|
||||
- The word gates are two measurements, not two tiers of one rule: the 2,770-word / 500-line spec backstop counts the whole file, Step 3's gate the body alone. Never unify them.
|
||||
- Never spawn a subagent to audit or recheck your own work — run `/skill-audit` inline, in the same context as the edits. Clean-context recheck belongs to `/forge`'s outer loop, and a self-spawned subagent's worktree can be torn down by concurrent cleanup, destroying an uncommitted draft.
|
||||
- Do not create new scripts unless a signal explicitly calls for it. Writing one from scratch requires out-of-scope transcript analysis — flag the opportunity as a suggestion instead.
|
||||
|
||||
## Step 1 — Dispatch
|
||||
|
||||
@@ -32,15 +31,15 @@ metadata:
|
||||
| Directory exists, at least one improvement signal present | Improve | `references/improve.md` |
|
||||
| Directory exists, no signals | Stop and ask | — |
|
||||
|
||||
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask: "No improvement signals found. Did you mean to create a new skill, or do you have feedback to apply?"
|
||||
Signals: grill output, `/skill-audit` findings, inline feedback, eval results, session context describing what went wrong. With none, ask whether the user meant to create a new skill or has feedback to apply.
|
||||
|
||||
Read only the reference matching the resolved flow — each is self-contained. Capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
|
||||
Read only the reference matching the resolved flow — each is self-contained. If the target sits inside a git worktree, capture `git log --oneline -1` before touching the filesystem; Step 4 needs it.
|
||||
|
||||
## Step 2 — Invocation axis
|
||||
|
||||
Decide before writing any description: model-invoked or hand-invoked?
|
||||
|
||||
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Worked example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`. Skip Step 3's description rules.
|
||||
- **Hand-invoked** — the user types `/name` and no agent should route to it. Set `disable-model-invocation: true` and write one plain human-facing sentence: no trigger list, no boundary clause. Skip Step 3's description rules.
|
||||
- **Model-invoked** — the default.
|
||||
|
||||
## Step 3 — Contract
|
||||
@@ -51,12 +50,12 @@ Gates `/skill-audit` enforces in both flows:
|
||||
|
||||
- **Description** — a trigger clause, at most one capability clause, and a boundary clause shaped `Not <thing> -> <skill-name>` whose target resolves to a real skill or agent. 250 characters SUGGESTION, 400 FAIL, value only.
|
||||
- **Body** — decision procedure only: ordered steps, branches, gates, and which reference to load when. 600 words SUGGESTION, 900 FAIL, body only. At two or more mutually exclusive flows a dispatch table is mandatory and each flow gets its own self-contained `references/` file.
|
||||
- **Gotchas** — at most five, each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL.
|
||||
- **Gotchas** — each contradicting a reasonable default. A Gotcha paraphrasing a step below it is a FAIL; over five entries is a SUGGESTION only.
|
||||
|
||||
## Step 4 — Validate and close
|
||||
|
||||
Run `/skill-audit` on the resolved skill directory. It checks name-to-directory match, description presence, leftover `FILL IN:` placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those first. Resolve every FAIL before reporting done.
|
||||
Run `/skill-audit` on the resolved skill directory; resolve every FAIL before reporting done. It checks name-to-directory match, placeholders, both size budgets, boundary-target resolution and script hygiene — do not hand-check those. Hand-check the one thing it misses: an empty body reports `PASS SKILL.md body word count 0 (ADR-0020 target: 600)`, so confirm at least one non-empty section exists.
|
||||
|
||||
With `metadata.version` present, bump the **minor** version on create (new skills start at `0.1.0`) and the **patch** version on improve.
|
||||
|
||||
**Commit verification.** Once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from the one captured at Step 1. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up first. Report done only once the hash has changed.
|
||||
**Commit verification.** Inside a git worktree: once the audit is clean, run `git add` and `git commit` — do not stop at staging. Re-run `git log --oneline -1` and confirm the hash changed from Step 1's. A non-empty `git diff --stat` is not proof: staged-but-uncommitted work is part of no commit and is silently lost if the tree is cleaned up. Report done only once the hash has changed. Outside a worktree (a skill under `~/.claude/skills/`, say) nothing is committable — report done on a clean audit, naming that as the reason.
|
||||
|
||||
@@ -10,14 +10,22 @@ name: SKILL_NAME
|
||||
# Examples: my-tool, data-analyzer, pdf-processor
|
||||
|
||||
description: >
|
||||
Use when FILL IN: trigger — when should an agent activate this skill?
|
||||
Use when FILL IN: trigger.
|
||||
FILL IN: at most ONE capability clause, stated specifically
|
||||
(e.g. "parses and validates OpenAPI specs", not "helps with APIs").
|
||||
Not FILL IN: near-miss case -> FILL IN: real sibling skill name.
|
||||
Not FILL IN: near-miss case -> FILL IN: real sibling skill.
|
||||
# Required. Preloaded into EVERY session whether or not the skill is invoked.
|
||||
# Exactly three parts, in this order: trigger clause, at most one capability
|
||||
# clause, boundary clause. Drop the boundary line if no near-miss skill exists.
|
||||
# Budget: 250 characters target, 400 hard ceiling (counting this value only).
|
||||
# Trigger clause: when should an agent activate this skill? Describe the user's
|
||||
# intent, not the skill's internal mechanics.
|
||||
# Budget: 250 characters target, 400 hard ceiling (counting this value only,
|
||||
# with YAML folding resolved). This scaffold sits at 214 — keep the fill-in
|
||||
# under the target rather than growing past it.
|
||||
# Boundary clauses may be plural: write one per genuine near-miss, and none
|
||||
# where no sibling could steal activations.
|
||||
# Never let a hyphenated skill name wrap across two lines of this folded block
|
||||
# — folding turns the break into a space and the routing target stops resolving.
|
||||
# Banned here: capability lists, output-format detail, composition notes,
|
||||
# implementation detail, and restating one trigger twice in two registers.
|
||||
# The boundary target must resolve to a real skill or agent — it is checked.
|
||||
|
||||
@@ -47,11 +47,15 @@ explicitly" only where the user's natural phrasing genuinely omits the domain wo
|
||||
for `git-commits`, where the user says "commit". Adding one everywhere is what inflated this
|
||||
corpus, and it was deleted as a blanket rule.
|
||||
|
||||
**Boundary targets must resolve.** The name after the arrow is checked against real skill
|
||||
directories under `plugins/*/.apm/skills/<name>/` and real agents under
|
||||
`plugins/*/.apm/agents/<name>.agent.md`. A boundary clause naming a target that does not exist
|
||||
sends the router nowhere and fails the audit. Check the target exists before writing it — do not
|
||||
invent a plausible sibling name.
|
||||
**Boundary targets must resolve.** Both forms are checked — the arrow and the prose form ("do not
|
||||
use for X, use `y` instead") — so a typo dangles either way. Targets resolve against a universe
|
||||
built by walking up **from the SKILL.md itself**: the nearest ancestor holding
|
||||
`plugins/*/.apm/{skills,agents}` (or, failing that, the nearest ancestor holding `.git`) contributes
|
||||
every skill and agent under `<root>/plugins/*/`, plus the skill's own apm package and the packages
|
||||
that package declares in `apm.yml` under `dependencies.apm`. A sibling plugin in the same monorepo
|
||||
therefore resolves; a skill in an unrelated repo does not. A boundary clause naming a target
|
||||
outside that universe sends the router nowhere and fails the audit. Check the target exists before
|
||||
writing it — do not invent a plausible sibling name.
|
||||
|
||||
**Length.** 250 characters SUGGESTION, 400 characters FAIL, counting the frontmatter value only
|
||||
with YAML folding resolved. The agentskills.io 1,024-character spec limit is unchanged and sits
|
||||
@@ -61,7 +65,13 @@ as the outlier stop.
|
||||
**Hand-invoked skills are exempt.** A skill carrying `disable-model-invocation: true` is absent
|
||||
from the model-visible listing and is reached only by the user typing `/name`. It takes one plain
|
||||
human-facing sentence — no trigger clause, no boundary clause, no indirect triggers. Worked
|
||||
example: `plugins/bin/.apm/skills/zoom-out/SKILL.md`.
|
||||
example — the whole description of the `zoom-out` skill, which carries `disable-model-invocation`:
|
||||
|
||||
````markdown
|
||||
Tell the agent to zoom out and give broader context or a higher-level perspective. Use when
|
||||
you're unfamiliar with a section of code or need to understand how it fits into the bigger
|
||||
picture.
|
||||
````
|
||||
|
||||
## Body
|
||||
|
||||
@@ -90,6 +100,12 @@ blocks, rationale prose, and any content only one branch reaches. Each reference
|
||||
self-contained for its concern, and every one is wired from the body with the literal conditional
|
||||
form:
|
||||
|
||||
**The one exception, stated once so it is not re-litigated:** an output schema stays in the body
|
||||
only when it applies to *every* flow and is short — roughly 50 words or less, which is the "Output
|
||||
format template" pattern below. An output schema that is longer than that, or that only one flow
|
||||
produces, moves to `references/` like any other schema. No third option exists, and the two rules
|
||||
do not disagree.
|
||||
|
||||
````markdown
|
||||
If <condition>, read `references/<file>.md`.
|
||||
````
|
||||
@@ -98,8 +114,9 @@ A generic pointer ("see references/ for details") is a Vale error — the agent
|
||||
|
||||
**Dispatch is mandatory at two or more mutually exclusive flows.** The body carries the dispatch
|
||||
table and the gates common to every branch; each flow gets its own self-contained `references/`
|
||||
file. Exemplar: `plugins/kyberforge/.apm/skills/apm-workflow/SKILL.md` — a 554-word body
|
||||
dispatching to 3,006 words of references.
|
||||
file. Exemplar: the `apm-workflow` skill — a **421-word body** dispatching to 3,006 words of
|
||||
references. Calibrate against 421: that file's whole-file count is 554 words, and aiming at that
|
||||
number instead overshoots the body budget by ~30%.
|
||||
|
||||
**Length.** 600 words SUGGESTION, 900 words FAIL, counting the **body only** — everything after
|
||||
the frontmatter's closing `---`.
|
||||
@@ -108,7 +125,7 @@ the frontmatter's closing `---`.
|
||||
|
||||
- Each entry must state a fact that **contradicts a reasonable default** — something the agent
|
||||
gets wrong by acting sensibly. "Never commit secrets" is not one; the agent already knows.
|
||||
- Maximum five entries.
|
||||
- More than five entries is a SUGGESTION — five is the guideline, not a ceiling.
|
||||
- A Gotcha that paraphrases a step in the body below it is a **FAIL**. If the rule is already a
|
||||
step, it is not a gotcha.
|
||||
- A Gotchas section exceeding 25% of the body is a SUGGESTION.
|
||||
@@ -164,7 +181,9 @@ Do not modify flags.
|
||||
| <condition> | <flow> | `references/<file>.md` |
|
||||
````
|
||||
|
||||
**Output format template** (when the skill produces structured output):
|
||||
**Output format template** (when the skill produces structured output on *every* flow, and the
|
||||
schema is roughly 50 words or less — see the exception under Body above; anything longer or
|
||||
flow-specific belongs in `references/`):
|
||||
|
||||
````markdown
|
||||
Output format:
|
||||
@@ -173,7 +192,8 @@ Output format:
|
||||
```
|
||||
````
|
||||
|
||||
For longer templates, place them in `assets/<name>.md` and reference conditionally.
|
||||
For longer templates, place them in `references/<topic>.md` or `assets/<name>.md` and reference
|
||||
conditionally.
|
||||
|
||||
## Embedding org-specific policy
|
||||
|
||||
|
||||
@@ -136,10 +136,16 @@ If no scripts are needed, delete `scripts/README.md` and the `scripts/` director
|
||||
|
||||
## Step 5 — Add references, assets, and tests (if needed)
|
||||
|
||||
**`references/`** — additional documentation loaded on demand. One topic per file. Reference
|
||||
conditionally from SKILL.md with the literal form ``If <condition>, read `references/<file>.md` ``.
|
||||
Keep reference chains one level deep — a reference file that references another reference file is
|
||||
rarely loaded correctly.
|
||||
**`references/`** — additional documentation loaded on demand. One topic per file, named in
|
||||
kebab-case after the topic. Reference conditionally from SKILL.md with the literal form
|
||||
``If <condition>, read `references/<file>.md` ``.
|
||||
|
||||
**Two hops from `SKILL.md`, never three.** A flow file may route on to a shared contract or
|
||||
sub-topic file — that is the shipped pattern here (`SKILL.md` → `references/create.md` → this
|
||||
file's own pointers to `contract.md`, `scripts.md` and `deployment-modes.md`). What does not work
|
||||
is a third hop: a file reachable only through two intermediates is rarely loaded at the moment it
|
||||
is needed. Every hop past the first also needs the same literal conditional form, so the agent
|
||||
knows when to take it.
|
||||
|
||||
**`assets/`** — static resources: templates, schemas, lookup tables. Reference by relative path
|
||||
from SKILL.md.
|
||||
|
||||
@@ -76,6 +76,12 @@ contract first — the gates are hot and carry no baseline file, so a one-line f
|
||||
non-compliant skill cannot be committed until the description and body meet
|
||||
`references/contract.md`. Treat that retrofit as part of the same change, not a follow-up.
|
||||
|
||||
If the skill's description exceeds 250 characters, or its body-only word count exceeds 600, read
|
||||
`references/retrofit.md` before editing. It carries the ordered cut procedure, the
|
||||
mutually-exclusive-flows test, the reference-file conventions this flow needs, the collateral
|
||||
checklist for `README.md` and `references/sources.md`, and a worked description retrofit. Do not
|
||||
improvise the cuts — four dry runs invented six to ten different answers to the same questions.
|
||||
|
||||
If a signal points to a script or reference file, edit that file directly rather than adding a
|
||||
workaround in SKILL.md.
|
||||
|
||||
|
||||
152
plugins/kyberforge/skills/skill-author/references/retrofit.md
Normal file
152
plugins/kyberforge/skills/skill-author/references/retrofit.md
Normal file
@@ -0,0 +1,152 @@
|
||||
---
|
||||
source_keys:
|
||||
- agentskills-best-practices
|
||||
- agentskills-optimizing-descriptions
|
||||
---
|
||||
|
||||
# Retrofitting a skill to the ADR-0020 contract
|
||||
|
||||
Read this when `references/improve.md` Step 4 sends you here: the skill you are editing is over
|
||||
the description or body budget and has to come into contract before any other change can be
|
||||
committed. The gates are hot and carry no baseline file, so a one-line fix to a non-compliant
|
||||
skill is blocked until this is done.
|
||||
|
||||
Measure first. Do not guess which gate fired: run `/skill-audit` on the directory and read its
|
||||
`### Structure` dimension, which reports the description characters and the **body-only** word
|
||||
count separately from the whole-file spec backstop. Retrofit against the number that actually
|
||||
fired — a skill can sit a thousand words inside the whole-file backstop while failing the body
|
||||
budget.
|
||||
|
||||
**Validate in place.** Audit the skill's real directory inside its package. Never audit a copy in a
|
||||
scratch directory, and never move a skill out to work on it: the boundary-target universe is built
|
||||
by walking up *from the file being checked*, so a copy with no authoring root above it resolves
|
||||
against nothing and the check declines rather than running —
|
||||
|
||||
```text
|
||||
INFO boundary-target resolution DID NOT RUN — no skill universe could be determined for
|
||||
this path ... Unchecked target(s): totally-fake-target
|
||||
```
|
||||
|
||||
The run still exits 0, so that line reads as a pass and is not one. Treat `DID NOT RUN` as **not
|
||||
checked**, always. A retrofit signed off on a scratch copy carries an unverified boundary target
|
||||
into the corpus, which is precisely the failure this gate exists to catch.
|
||||
|
||||
## Cut in this order
|
||||
|
||||
Work the list top down and stop as soon as the gate clears. The order is by ratio of tokens
|
||||
removed to behaviour lost — inverting it is how a retrofit ends up deleting the one instruction
|
||||
the skill existed to carry.
|
||||
|
||||
1. **Gotchas that paraphrase a step in the body below.** Zero information, and already a FAIL on
|
||||
its own. Delete the Gotcha, keep the step.
|
||||
2. **Spec restatements** — text that repeats a published specification, a tool's `--help`, or a
|
||||
ceiling the validator already enforces. The agent gets this right without it. Delete, or move
|
||||
the table to `references/` if a flow genuinely needs to look it up.
|
||||
3. **Capability enumeration** — in a description, the feature list after the trigger clause; in a
|
||||
body, the paragraph that recites what the skill can do. One capability clause survives in the
|
||||
description; the rest belongs in `README.md`.
|
||||
4. **Per-flow prose** — anything only one branch of the procedure ever reaches. This is the
|
||||
largest single win in most bodies, and it is a *move*, not a delete: each flow gets its own
|
||||
self-contained `references/` file, wired from a dispatch table.
|
||||
|
||||
If the body is still over after all four, the skill is doing two jobs. Split it, and say so
|
||||
rather than compressing prose until it stops being readable.
|
||||
|
||||
## What "mutually exclusive flows" means
|
||||
|
||||
Two or more flows that a single invocation cannot both take. The three-way test, copied verbatim
|
||||
from the body-discipline rubric `/skill-audit` judges against — nothing to load, it is quoted in
|
||||
full here:
|
||||
|
||||
> separate subcommands, separate input types, separate lifecycle stages
|
||||
|
||||
Any one of the three is enough. Two flows that differ only in a parameter value are one flow.
|
||||
At two or more mutually exclusive flows a dispatch table is **mandatory** regardless of word
|
||||
count, because every invocation otherwise pays for every branch it did not take.
|
||||
|
||||
## Reference-file conventions
|
||||
|
||||
The create flow owns these rules, and this flow is forbidden from reading `references/create.md`,
|
||||
so what a retrofit needs is restated here:
|
||||
|
||||
- **One topic per file.** A file mixing two concerns gets loaded for one of them and spends the
|
||||
caller's context on the other.
|
||||
- **Kebab-case filenames**, named after the topic rather than the flow that reads it —
|
||||
`body-discipline.md`, not `step-3.md`.
|
||||
- **Wire every file with the literal conditional form** ``If <condition>, read
|
||||
`references/<file>.md` ``. A generic pointer ("see `references/` for details") is a Vale error.
|
||||
- **Two hops from `SKILL.md`, never three.** A flow file may route on to a shared contract file;
|
||||
a file reachable only through two intermediates is rarely loaded when it is needed.
|
||||
- **`source_keys` frontmatter.** If the content you are moving drew on a research source, the new
|
||||
file needs top-level `source_keys:` frontmatter listing those slugs, and every slug must already
|
||||
exist as an `## <slug>` heading in `references/sources.md`. Moving sourced content out of
|
||||
`SKILL.md` without carrying its slugs across breaks the provenance chain, and `/skill-audit`
|
||||
reports the new file as an INFO with no `source_keys`.
|
||||
|
||||
## Collateral is mandatory, not optional
|
||||
|
||||
Moving content out of a `SKILL.md` leaves three files describing a structure that no longer
|
||||
exists. `/skill-audit`'s provenance check exits clean on all three of these, so nothing catches
|
||||
them for you. After every retrofit that adds, removes or renames a file:
|
||||
|
||||
- [ ] **`README.md` file table** — a row for every new `references/` file, and no row left for a
|
||||
file that is gone. Say what triggers the load, not just what the file contains.
|
||||
- [ ] **`references/README.md`**, where the skill has one — same update, same reason.
|
||||
- [ ] **`references/sources.md` → `Contributing files`** — add the new file to every slug whose
|
||||
content moved into it, and remove any file the retrofit deleted. This is the one that gets
|
||||
missed: `sources.md` keeps citing sections of `SKILL.md` that no longer exist, the
|
||||
provenance check still exits 0, and the stale claim survives review.
|
||||
- [ ] Re-run `/skill-audit` and confirm its `### Provenance` dimension does not report the new
|
||||
file as missing `source_keys`.
|
||||
|
||||
## Worked example — a description retrofit
|
||||
|
||||
`gitea-issues` before, 827 characters, the single most common shape in the corpus:
|
||||
|
||||
```text
|
||||
Use when reading or writing Gitea issues: listing repo issues, getting a single issue's details/
|
||||
comments/labels, creating an issue, updating its state, adding or editing comments, applying
|
||||
labels via issue_write, or searching issues/PRs across repositories. Triggers on "create an
|
||||
issue", "what issues are open", "get issue #N", "close issue #N", "comment on issue #N", "search
|
||||
issues for X" — even when the user doesn't say "Gitea" explicitly. Composes gitea-labels-
|
||||
milestones for all label inference/resolution and milestone lookup — do not use this skill to
|
||||
manage label or milestone definitions themselves (create/edit/delete a label, create/close a
|
||||
milestone), that's gitea-labels-milestones directly. Do not use for pull requests (use gitea-prs)
|
||||
or for local git branch/commit work (use gitea-branches or git-branches).
|
||||
```
|
||||
|
||||
After, 240 characters:
|
||||
|
||||
```text
|
||||
Use when reading or writing Gitea issues — list, read, create, comment on, label, close, or
|
||||
search — even when the user does not say "Gitea". Not pull requests -> `gitea-prs`. Not label or
|
||||
milestone definitions -> `gitea-labels-milestones`.
|
||||
```
|
||||
|
||||
What came out, and why:
|
||||
|
||||
| Removed | Why |
|
||||
|---|---|
|
||||
| The second trigger register — `Triggers on "create an issue", "what issues are open", …` | The same triggers restated as quoted user phrasings. Two registers of one trigger list is a FAIL, not a suggestion. |
|
||||
| `applying labels via issue_write` | Implementation detail. The router does not choose a skill by which MCP call it makes. |
|
||||
| `Composes gitea-labels-milestones for all label inference/resolution and milestone lookup` | A composition note. It changes no routing decision and belongs in `README.md`. |
|
||||
| The parenthetical `(create/edit/delete a label, create/close a milestone)` | Capability enumeration inside a boundary clause. The boundary needs the target, not its feature list. |
|
||||
| The `gitea-branches` / `git-branches` boundary | Dropped entirely. Neither was ever going to win an issue request, so the clause defended against nothing — an invented boundary costs characters and buys no routing accuracy. |
|
||||
| `Do not use for pull requests (use gitea-prs)` prose form | Kept, but rewritten as `Not pull requests -> \`gitea-prs\`.` The rewrite buys characters and one uniform shape for the router — not safety. Both forms are parsed **and** target-checked, so a typo in the prose form dangles exactly as an arrow typo does. |
|
||||
|
||||
What stayed: one trigger clause, one capability clause, the indirect trigger (genuinely warranted
|
||||
here — people say "create an issue", not "create a Gitea issue"), and the boundary clauses.
|
||||
|
||||
## Two rules the gates enforce but the prose does not spell out
|
||||
|
||||
**Boundary clauses may be plural.** Write one per genuine near-miss — the example above carries
|
||||
two, because two different skills could each steal activations. "A boundary clause" in the
|
||||
contract means *at least one*, not *exactly one*. What is banned is a boundary clause invented for
|
||||
a skill that was never going to compete, not a second real one.
|
||||
|
||||
**Never let a hyphenated routing target wrap across lines in a folded `>` scalar.** YAML folding
|
||||
replaces the newline with a space, so `gitea-labels-` at the end of one line and `milestones` at
|
||||
the start of the next fold into `gitea-labels- milestones`. `validate.sh` then reads the target as
|
||||
`gitea-labels`, finds no such skill, and reports a dangling boundary target — the live finding on
|
||||
`gitea-issues` today. Reflow the line so the whole name sits on one of them. The same applies to
|
||||
any backticked skill or agent name in a description.
|
||||
@@ -34,7 +34,7 @@ source_keys:
|
||||
- **URL:** https://agentskills.io/skill-creation/best-practices.md
|
||||
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
|
||||
- **Description:** Best practices for skill creators — starting from real expertise, spending context wisely, calibrating control, instruction patterns (gotchas, templates, checklists, validation loops)
|
||||
- **Contributing files:** SKILL.md, references/create.md, references/improve.md, references/contract.md
|
||||
- **Contributing files:** SKILL.md, references/create.md, references/improve.md, references/contract.md, references/retrofit.md
|
||||
- **Status:** `extracted`
|
||||
|
||||
## agentskills-optimizing-descriptions
|
||||
@@ -42,7 +42,7 @@ source_keys:
|
||||
- **URL:** https://agentskills.io/skill-creation/optimizing-descriptions.md
|
||||
- **Research doc:** plugins/kyberforge/docs/research/docs/agentskillsio/sources.md
|
||||
- **Description:** How to systematically test and improve skill descriptions for triggering accuracy — eval queries, trigger rate testing, train/validation splits, optimization loop
|
||||
- **Contributing files:** SKILL.md, references/improve.md, references/contract.md
|
||||
- **Contributing files:** SKILL.md, references/improve.md, references/contract.md, references/retrofit.md
|
||||
- **Status:** `extracted`
|
||||
|
||||
## agentskills-evaluating-skills
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -36,6 +36,22 @@ STRICT=false
|
||||
if [[ "${RUN_TESTS_STRICT:-}" == "1" ]]; then
|
||||
STRICT=true
|
||||
fi
|
||||
# Latched, then REMOVED from the environment. The value has done its only job by
|
||||
# this line -- it is now held in the STRICT shell local -- and leaving it exported
|
||||
# makes strictness leak down the whole process tree: every test-*.sh dispatched
|
||||
# through batch_run below inherits it, and any of them that itself invokes
|
||||
# run-tests.sh (tests/test-run-tests.sh drives a copy of this script over fixture
|
||||
# trees) silently turns a deliberately non-strict fixture strict.
|
||||
#
|
||||
# That is not symmetric with `--strict`, which never leaked: the flag only ever
|
||||
# sets the shell local above, so `bash tests/run-tests.sh --strict` (the spelling
|
||||
# the run-tests pre-push hook uses) always gave children a clean environment. Only
|
||||
# the env-var spelling leaked, and it broke exactly two assertions in
|
||||
# tests/test-run-tests.sh -- its cases 10c and 10g. Unsetting here makes the two
|
||||
# documented invocations equivalent in what a CHILD sees, not just in the parent's
|
||||
# verdict, so no future suite has to defend itself the way test-run-tests.sh's
|
||||
# run_fake() does with `env -u`.
|
||||
unset RUN_TESTS_STRICT
|
||||
# A loop rather than the `[[ "${1:-}" == --bats-only ]]` test this used to be, so
|
||||
# the two flags compose and an unknown flag is rejected instead of ignored. A
|
||||
# silently-ignored `--strict` is the one typo that would turn the gate back off.
|
||||
|
||||
406
tests/test-adr0020-body-checks.sh
Executable file
406
tests/test-adr0020-body-checks.sh
Executable file
@@ -0,0 +1,406 @@
|
||||
#!/usr/bin/env bash
|
||||
# Regression test for the ADR-0020 body-shape checks and, just as importantly,
|
||||
# for the false-positive fixes each of them needed. Every check here was
|
||||
# completely untested.
|
||||
#
|
||||
# * Gotchas section over 5 entries — SUGGESTION
|
||||
# * Gotchas section over 25% of the body — SUGGESTION
|
||||
# * a references/<file>.md named but absent — ERROR (a broken pointer is not a
|
||||
# style opinion)
|
||||
# * description with no boundary clause — SUGGESTION
|
||||
#
|
||||
# The false-positive half is not optional extra coverage. Each of these checks
|
||||
# scans prose, and the first naive version of each one fired on ordinary writing:
|
||||
# a ```-fenced EXAMPLE of a Gotchas section became the section itself, indented
|
||||
# child bullets were counted as top-level entries, `## Gotcha handling` was read
|
||||
# as the Gotchas section, and a documented-then-removed references/ file became a
|
||||
# hard ERROR. The skills most likely to carry such an example are skill-author and
|
||||
# skill-audit — the two that DOCUMENT these conventions — so a gate that fires on
|
||||
# them is a gate nobody can turn on.
|
||||
#
|
||||
# Every case is a matched pair: the check fires just over its boundary, and stays
|
||||
# silent just under it (or on the shape it must not match). A test asserting only
|
||||
# that a bad file fails proves nothing about a check that fires on everything.
|
||||
set -euo pipefail
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
HOOK="$REPO_ROOT/scripts/skill-size-check.sh"
|
||||
PASS=0
|
||||
FAIL=0
|
||||
|
||||
pass() { echo " PASS: $1"; PASS=$((PASS + 1)); }
|
||||
fail() { echo " FAIL: $1"; FAIL=$((FAIL + 1)); }
|
||||
|
||||
TMPDIR_T="$(mktemp -d)"
|
||||
trap 'rm -rf "$TMPDIR_T"' EXIT
|
||||
|
||||
# A description with a boundary clause and no routing target — the neutral
|
||||
# default, so a fixture about Gotchas or references does not also trip the
|
||||
# missing-boundary-clause SUGGESTION and stop isolating what it names.
|
||||
CLEAN_DESC="Use when doing the thing. Do not use for anything else."
|
||||
|
||||
# make_skill <name> <desc> — SKILL.md with the body read from stdin. Echoes the
|
||||
# path. Each skill gets its own directory so references/ fixtures are isolated.
|
||||
make_skill() {
|
||||
local name="$1" desc="$2" dir
|
||||
dir="$TMPDIR_T/$name"
|
||||
mkdir -p "$dir"
|
||||
{
|
||||
echo "---"
|
||||
echo "name: $name"
|
||||
echo "description: $desc"
|
||||
echo "---"
|
||||
cat
|
||||
} > "$dir/SKILL.md"
|
||||
echo "$dir/SKILL.md"
|
||||
}
|
||||
|
||||
# expect <label> <file> <expect: silent|suggests|errors> [needle] [absent-needle]
|
||||
#
|
||||
# `silent` is the strict one: exit 0 AND completely empty output. It is what
|
||||
# makes every "must not fire" case below real — a check that fired with some
|
||||
# other wording would still be caught.
|
||||
expect() {
|
||||
local label="$1" file="$2" mode="$3" needle="${4:-}" absent="${5:-}" out status=0
|
||||
set +e
|
||||
out="$(bash "$HOOK" "$file" 2>&1)"
|
||||
status=$?
|
||||
set -e
|
||||
case "$mode" in
|
||||
silent)
|
||||
if [[ $status -eq 0 && -z "$out" ]]; then
|
||||
pass "$label"
|
||||
else
|
||||
fail "$label (exit $status, output: ${out:-<empty>})"
|
||||
fi
|
||||
;;
|
||||
suggests)
|
||||
if [[ $status -ne 0 ]]; then
|
||||
fail "$label — a SUGGESTION must never change the exit code (exit $status, output: $out)"
|
||||
elif [[ "$out" != *"SUGGESTION"* || "$out" != *"$needle"* ]]; then
|
||||
fail "$label (exit $status, output: ${out:-<empty>})"
|
||||
elif [[ -n "$absent" && "$out" == *"$absent"* ]]; then
|
||||
fail "$label — output also contained '$absent', which must not fire here: $out"
|
||||
else
|
||||
pass "$label"
|
||||
fi
|
||||
;;
|
||||
errors)
|
||||
if [[ $status -eq 0 ]]; then
|
||||
fail "$label — expected a hard ERROR, got exit 0 (output: ${out:-<empty>})"
|
||||
elif [[ "$out" != *"$needle"* ]]; then
|
||||
fail "$label (exit $status, output: ${out:-<empty>})"
|
||||
else
|
||||
pass "$label"
|
||||
fi
|
||||
;;
|
||||
quiet-about)
|
||||
# Exit 0 and the named text absent, but other output permitted. Used where
|
||||
# a second, unrelated finding legitimately fires.
|
||||
if [[ $status -ne 0 ]]; then
|
||||
fail "$label (exit $status, output: $out)"
|
||||
elif [[ "$out" == *"$needle"* ]]; then
|
||||
fail "$label — '$needle' fired when it must not: $out"
|
||||
else
|
||||
pass "$label"
|
||||
fi
|
||||
;;
|
||||
esac
|
||||
}
|
||||
|
||||
filler() { python3 -c "print(' '.join(['word'] * $1))"; }
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Gotchas: entry count (guideline 5)
|
||||
# ---------------------------------------------------------------------------
|
||||
# The filler after the section keeps the 25% fraction check well clear, so these
|
||||
# two cases isolate the ENTRY count. Without it a six-entry section in a short
|
||||
# body would fire both and the pair would not distinguish them.
|
||||
echo ""
|
||||
echo "--- Gotchas entry count: 5 is fine, 6 is a SUGGESTION ---"
|
||||
F_FIVE="$(make_skill gotchas-five "$CLEAN_DESC" <<EOF
|
||||
|
||||
## Gotchas
|
||||
|
||||
- first trap here
|
||||
- second trap here
|
||||
- third trap here
|
||||
- fourth trap here
|
||||
- fifth trap here
|
||||
|
||||
## Notes
|
||||
|
||||
$(filler 200)
|
||||
EOF
|
||||
)"
|
||||
expect "a Gotchas section with exactly 5 entries is silent" "$F_FIVE" silent
|
||||
|
||||
F_SIX="$(make_skill gotchas-six "$CLEAN_DESC" <<EOF
|
||||
|
||||
## Gotchas
|
||||
|
||||
- first trap here
|
||||
- second trap here
|
||||
- third trap here
|
||||
- fourth trap here
|
||||
- fifth trap here
|
||||
- sixth trap here
|
||||
|
||||
## Notes
|
||||
|
||||
$(filler 200)
|
||||
EOF
|
||||
)"
|
||||
expect "a Gotchas section with 6 entries raises a SUGGESTION and still exits 0" \
|
||||
"$F_SIX" suggests "Gotchas section has 6 entries" "over the 25% guideline"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Gotchas: share of the body (guideline 25%)
|
||||
# ---------------------------------------------------------------------------
|
||||
# Exact boundary arithmetic, not an approximation. The section is prose (no list
|
||||
# items) so the entry check cannot fire and confuse the result; body words are
|
||||
# then section + filler + the two two-word headings. At a body of 100 words a
|
||||
# 25-word section is exactly the guideline (inclusive — `>` is the comparison, so
|
||||
# it passes) and a 26-word section is one word past it.
|
||||
echo ""
|
||||
echo "--- Gotchas share of body: exactly 25% is fine, 26% is a SUGGESTION ---"
|
||||
F_AT="$(make_skill gotchas-at-fraction "$CLEAN_DESC" <<EOF
|
||||
|
||||
## Gotchas
|
||||
|
||||
$(filler 25)
|
||||
|
||||
## Notes
|
||||
|
||||
$(filler 71)
|
||||
EOF
|
||||
)"
|
||||
expect "a Gotchas section at exactly 25% of the body is silent" "$F_AT" silent
|
||||
|
||||
F_OVER="$(make_skill gotchas-over-fraction "$CLEAN_DESC" <<EOF
|
||||
|
||||
## Gotchas
|
||||
|
||||
$(filler 26)
|
||||
|
||||
## Notes
|
||||
|
||||
$(filler 70)
|
||||
EOF
|
||||
)"
|
||||
expect "a Gotchas section at 26% of the body raises a SUGGESTION" \
|
||||
"$F_OVER" suggests "Gotchas section is 26 of 100 body words (26%)" "entries"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Gotchas: false positives
|
||||
# ---------------------------------------------------------------------------
|
||||
echo ""
|
||||
echo "--- a ## Gotchas heading inside a fenced block is not the Gotchas section ---"
|
||||
# skill-author and skill-audit both document this convention by showing it. If a
|
||||
# fenced example counted, the two skills that define the rule would be the two
|
||||
# most likely to fail it.
|
||||
F_FENCED_HEADING="$(make_skill gotchas-fenced-heading "$CLEAN_DESC" <<EOF
|
||||
|
||||
## How to write one
|
||||
|
||||
\`\`\`markdown
|
||||
## Gotchas
|
||||
|
||||
- example one
|
||||
- example two
|
||||
- example three
|
||||
- example four
|
||||
- example five
|
||||
- example six
|
||||
- example seven
|
||||
\`\`\`
|
||||
|
||||
$(filler 200)
|
||||
EOF
|
||||
)"
|
||||
expect "a fenced ## Gotchas heading is not read as the section" \
|
||||
"$F_FENCED_HEADING" quiet-about "Gotchas section"
|
||||
|
||||
echo ""
|
||||
echo "--- list items inside a fenced block do not count as Gotchas entries ---"
|
||||
F_FENCED_ENTRIES="$(make_skill gotchas-fenced-entries "$CLEAN_DESC" <<EOF
|
||||
|
||||
## Gotchas
|
||||
|
||||
- a real trap
|
||||
- another real trap
|
||||
|
||||
Shown as an example of what NOT to write:
|
||||
|
||||
\`\`\`markdown
|
||||
- fake one
|
||||
- fake two
|
||||
- fake three
|
||||
- fake four
|
||||
- fake five
|
||||
- fake six
|
||||
- fake seven
|
||||
- fake eight
|
||||
\`\`\`
|
||||
|
||||
## Notes
|
||||
|
||||
$(filler 300)
|
||||
EOF
|
||||
)"
|
||||
expect "eight fenced bullets plus two real ones counts as two entries, not ten" \
|
||||
"$F_FENCED_ENTRIES" quiet-about "Gotchas section has"
|
||||
|
||||
echo ""
|
||||
echo "--- the heading must END in gotcha(s): '## Gotcha handling' is not the section ---"
|
||||
# `## Gotcha handling` and `## Why gotchas matter` are prose sections. Treating
|
||||
# one as the Gotchas section measures a span that was never a gotcha list.
|
||||
F_HANDLING="$(make_skill gotcha-handling "$CLEAN_DESC" <<EOF
|
||||
|
||||
## Gotcha handling
|
||||
|
||||
- item one
|
||||
- item two
|
||||
- item three
|
||||
- item four
|
||||
- item five
|
||||
- item six
|
||||
- item seven
|
||||
|
||||
## Notes
|
||||
|
||||
$(filler 200)
|
||||
EOF
|
||||
)"
|
||||
expect "'## Gotcha handling' is not matched as the Gotchas section" \
|
||||
"$F_HANDLING" quiet-about "Gotchas section"
|
||||
|
||||
# The control for the three cases above. Without it, "no Gotchas finding" could
|
||||
# equally mean the whole check is dead, and all three would still be green.
|
||||
F_CONTROL="$(make_skill gotchas-control "$CLEAN_DESC" <<EOF
|
||||
|
||||
## Common gotchas
|
||||
|
||||
- item one
|
||||
- item two
|
||||
- item three
|
||||
- item four
|
||||
- item five
|
||||
- item six
|
||||
- item seven
|
||||
|
||||
## Notes
|
||||
|
||||
$(filler 200)
|
||||
EOF
|
||||
)"
|
||||
expect "control: a real '## Common gotchas' heading with 7 entries IS matched" \
|
||||
"$F_CONTROL" suggests "Gotchas section has 7 entries"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# references/<file>.md pointers
|
||||
# ---------------------------------------------------------------------------
|
||||
echo ""
|
||||
echo "--- a references/ pointer that is not on disk is a hard ERROR ---"
|
||||
F_REF_MISSING="$(make_skill ref-missing "$CLEAN_DESC" <<EOF
|
||||
|
||||
If the caller needs the long form, read references/nowhere.md first.
|
||||
EOF
|
||||
)"
|
||||
expect "a body pointing at an absent references/nowhere.md ERRORs" \
|
||||
"$F_REF_MISSING" errors "points at references/nowhere.md"
|
||||
|
||||
F_REF_PRESENT="$(make_skill ref-present "$CLEAN_DESC" <<EOF
|
||||
|
||||
If the caller needs the long form, read references/here.md first.
|
||||
EOF
|
||||
)"
|
||||
mkdir -p "$TMPDIR_T/ref-present/references"
|
||||
echo "content" > "$TMPDIR_T/ref-present/references/here.md"
|
||||
expect "the same pointer is silent once the file exists" "$F_REF_PRESENT" silent
|
||||
|
||||
echo ""
|
||||
echo "--- a references/ pointer inside a fenced block is not a dispatch entry ---"
|
||||
F_REF_FENCED="$(make_skill ref-fenced "$CLEAN_DESC" <<EOF
|
||||
|
||||
Dispatch tables look like this:
|
||||
|
||||
\`\`\`markdown
|
||||
If X, read references/example-file.md.
|
||||
\`\`\`
|
||||
EOF
|
||||
)"
|
||||
expect "a fenced references/example-file.md does not ERROR" "$F_REF_FENCED" silent
|
||||
|
||||
echo ""
|
||||
echo "--- a references/ pointer in a same-line removal context is history, not dispatch ---"
|
||||
# Narrow on purpose: a live dispatch table never describes its own target as
|
||||
# removed, so the exemption costs no recall. Each phrasing is checked separately
|
||||
# because they are separate alternatives in one regex, and a typo in any one of
|
||||
# them turns ordinary prose back into a hard ERROR.
|
||||
#
|
||||
# The file names are deliberately NEUTRAL (detail-a.md, not gone-a.md). An
|
||||
# earlier draft of this block named them gone-*.md and every case passed for the
|
||||
# wrong reason: "gone" is itself one of the removal words, so the exemption fired
|
||||
# off the FILENAME and the phrase under test was never exercised. The control
|
||||
# below is what surfaced that — it is the assertion that keeps these five honest.
|
||||
i=0
|
||||
for phrase in \
|
||||
"The old references/detail-a.md was removed in v2." \
|
||||
"references/detail-b.md is no longer part of this skill." \
|
||||
"references/detail-c.md is deprecated and should not be read." \
|
||||
"references/detail-d.md was renamed, so nothing points at it now." \
|
||||
"references/detail-e.md has been superseded by the body itself."; do
|
||||
i=$((i + 1))
|
||||
F_REF_PAST="$(make_skill "ref-past-$i" "$CLEAN_DESC" <<EOF
|
||||
|
||||
$phrase
|
||||
EOF
|
||||
)"
|
||||
expect "removal-context pointer is not an ERROR: \"$phrase\"" "$F_REF_PAST" silent
|
||||
done
|
||||
|
||||
# The control: the SAME sentence shape without a removal word must still ERROR,
|
||||
# or the exemption above has swallowed the check rather than narrowed it.
|
||||
F_REF_LIVE="$(make_skill ref-live "$CLEAN_DESC" <<EOF
|
||||
|
||||
The details live in references/detail-a.md, which the agent should read first.
|
||||
EOF
|
||||
)"
|
||||
expect "control: the same pointer with no removal word still ERRORs" \
|
||||
"$F_REF_LIVE" errors "points at references/detail-a.md"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Missing boundary clause
|
||||
# ---------------------------------------------------------------------------
|
||||
echo ""
|
||||
echo "--- a description with no boundary clause raises a SUGGESTION ---"
|
||||
F_NO_BOUNDARY="$(make_skill no-boundary "Use when the user wants the thing done." <<EOF
|
||||
|
||||
Do the thing.
|
||||
EOF
|
||||
)"
|
||||
expect "a description with no boundary clause raises a SUGGESTION and still exits 0" \
|
||||
"$F_NO_BOUNDARY" suggests "description has no boundary clause"
|
||||
|
||||
# Both accepted shapes, asserted separately: the prose markers and ADR-0020's
|
||||
# compressed arrow form. Dropping either from the detector would leave the other
|
||||
# green.
|
||||
F_PROSE_BOUNDARY="$(make_skill prose-boundary "Use when the user wants the thing done. Do not use for anything else." <<EOF
|
||||
|
||||
Do the thing.
|
||||
EOF
|
||||
)"
|
||||
expect "the prose boundary form satisfies the check" "$F_PROSE_BOUNDARY" silent
|
||||
|
||||
F_ARROW_BOUNDARY="$(make_skill arrow-boundary "Use when the user wants the thing done. Not the other thing -> sibling-skill." <<EOF
|
||||
|
||||
Do the thing.
|
||||
EOF
|
||||
)"
|
||||
expect "ADR-0020's compressed 'Not X -> y' form satisfies the check" \
|
||||
"$F_ARROW_BOUNDARY" quiet-about "no boundary clause"
|
||||
|
||||
echo ""
|
||||
echo "Results: $PASS passed, $FAIL failed"
|
||||
[[ $FAIL -eq 0 ]]
|
||||
331
tests/test-adr0020-contract.sh
Executable file
331
tests/test-adr0020-contract.sh
Executable file
@@ -0,0 +1,331 @@
|
||||
#!/usr/bin/env bash
|
||||
# Regression test for the three STRUCTURAL claims the ADR-0020 gate family makes
|
||||
# about itself. None of them was pinned anywhere before this file, and each one
|
||||
# fails silently — which is the whole reason they need a test rather than a
|
||||
# comment:
|
||||
#
|
||||
# 1. "ONE resolver, embedded VERBATIM in three scripts." The block between the
|
||||
# BEGIN/END markers is copied, not imported, because a cache-installed
|
||||
# plugin's scripts cannot read files outside their own plugin directory.
|
||||
# Nothing but this file asserts the three copies are still identical, and a
|
||||
# one-line edit to a single copy is invisible: every constant-agreement
|
||||
# assertion in tests/test-skill-size-check.sh still passes, because the
|
||||
# CONSTANTS are not what drifted.
|
||||
# 2. Both interpreter preflights, in all three scripts. python3 and PyYAML are
|
||||
# declared HARD dependencies precisely so a missing one cannot turn into a
|
||||
# vacuous pass, and the two are checked separately so the message names the
|
||||
# thing to install rather than the wrong one.
|
||||
# 3. `verbose: true` on the skill-size-check hook. It is the ENTIRE delivery
|
||||
# mechanism for the SUGGESTION tier: pre-commit prints nothing at all for a
|
||||
# passing hook, and a SUGGESTION deliberately does not fail, so dropping
|
||||
# one word from the config silences the tier ADR-0020 depends on while
|
||||
# every test and every hook still reports green.
|
||||
set -euo pipefail
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
HOOK="$REPO_ROOT/scripts/skill-size-check.sh"
|
||||
SKILL_VALIDATE="$REPO_ROOT/plugins/kyberforge/.apm/skills/skill-audit/scripts/validate.sh"
|
||||
AGENT_VALIDATE="$REPO_ROOT/plugins/kyberforge/.apm/skills/agent-audit/scripts/validate.sh"
|
||||
PASS=0
|
||||
FAIL=0
|
||||
|
||||
pass() { echo " PASS: $1"; PASS=$((PASS + 1)); }
|
||||
fail() { echo " FAIL: $1"; FAIL=$((FAIL + 1)); }
|
||||
|
||||
TMPDIR_T="$(mktemp -d)"
|
||||
trap 'rm -rf "$TMPDIR_T"' EXIT
|
||||
|
||||
BEGIN_MARKER='# ===== BEGIN ADR-0020 SHARED BOUNDARY RESOLVER ====='
|
||||
END_MARKER='# ===== END ADR-0020 SHARED BOUNDARY RESOLVER ====='
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 1. The shared resolver block is byte-identical in all three scripts
|
||||
# ---------------------------------------------------------------------------
|
||||
echo ""
|
||||
echo "--- the ADR-0020 shared resolver block is byte-identical in all three scripts ---"
|
||||
|
||||
# Marker discipline first. An unbalanced or duplicated marker pair makes the
|
||||
# extraction below silently measure the wrong span — a sed range that never
|
||||
# closes swallows the rest of the file, and one that opens twice concatenates
|
||||
# two spans. Both would still compare "equal" if all three were mangled the
|
||||
# same way, so the shape is asserted before the contents.
|
||||
MARKERS_OK=true
|
||||
for f in "$HOOK" "$SKILL_VALIDATE" "$AGENT_VALIDATE"; do
|
||||
if [[ ! -f "$f" ]]; then
|
||||
fail "script not found: $f"
|
||||
MARKERS_OK=false
|
||||
continue
|
||||
fi
|
||||
b="$(grep -cFx "$BEGIN_MARKER" "$f" || true)"
|
||||
e="$(grep -cFx "$END_MARKER" "$f" || true)"
|
||||
if [[ "$b" == "1" && "$e" == "1" ]]; then
|
||||
pass "${f#"$REPO_ROOT/"} carries exactly one BEGIN and one END marker"
|
||||
else
|
||||
fail "${f#"$REPO_ROOT/"} has $b BEGIN and $e END markers, expected 1 and 1"
|
||||
MARKERS_OK=false
|
||||
fi
|
||||
done
|
||||
|
||||
if ! $MARKERS_OK; then
|
||||
fail "skipping the byte-identity comparison — the marker pairs are not well-formed, so any extraction would measure the wrong span"
|
||||
else
|
||||
HASHES=()
|
||||
LINECOUNTS=()
|
||||
for f in "$HOOK" "$SKILL_VALIDATE" "$AGENT_VALIDATE"; do
|
||||
out="$TMPDIR_T/block-$(echo "$f" | md5sum | cut -c1-8).txt"
|
||||
sed -n "/^${BEGIN_MARKER}\$/,/^${END_MARKER}\$/p" "$f" > "$out"
|
||||
HASHES+=("$(md5sum < "$out" | cut -d' ' -f1)")
|
||||
LINECOUNTS+=("$(wc -l < "$out" | tr -d ' ')")
|
||||
done
|
||||
if [[ "${HASHES[0]}" == "${HASHES[1]}" && "${HASHES[1]}" == "${HASHES[2]}" ]]; then
|
||||
pass "all three copies hash to ${HASHES[0]} (${LINECOUNTS[0]} lines) — agreement by construction, not by coincidence"
|
||||
else
|
||||
fail "the shared resolver has DRIFTED: skill-size-check=${HASHES[0]} (${LINECOUNTS[0]} lines), skill-audit=${HASHES[1]} (${LINECOUNTS[1]} lines), agent-audit=${HASHES[2]} (${LINECOUNTS[2]} lines). Edit one copy, then paste it over the other two."
|
||||
fi
|
||||
# A block that has been emptied out would hash equal in all three and pass the
|
||||
# comparison above while enforcing nothing. The resolver is ~570 lines; 100 is
|
||||
# a floor low enough never to need maintenance and high enough that a gutted
|
||||
# block cannot sneak past.
|
||||
if [[ "${LINECOUNTS[0]}" -gt 100 ]]; then
|
||||
pass "the extracted block is ${LINECOUNTS[0]} lines — the comparison is over real content, not an empty span"
|
||||
else
|
||||
fail "the extracted shared block is only ${LINECOUNTS[0]} lines — three identical empty spans would compare equal and assert nothing"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 2. Both interpreter preflights, in all three scripts
|
||||
# ---------------------------------------------------------------------------
|
||||
# The two are checked separately on purpose: `python3 -c 'import yaml'` fails
|
||||
# identically whether python3 is missing or PyYAML is, and naming the wrong one
|
||||
# sends the reader to install the wrong thing.
|
||||
REAL_PYTHON="$(command -v python3)"
|
||||
# Absolute path, deliberately. The no-python3 fixture below replaces PATH
|
||||
# wholesale, so a bare `bash` (or `/usr/bin/env bash`) would be resolved against
|
||||
# that stripped PATH and die with "No such file or directory" before the script
|
||||
# under test ever starts -- a 127 that looks like the preflight firing.
|
||||
BASH_BIN="$(command -v bash)"
|
||||
|
||||
# A PATH that genuinely has no python3 on it. Built by symlinking the handful of
|
||||
# binaries the three scripts touch before their own preflight rather than by
|
||||
# hiding python3 from a full PATH, because there is no portable way to subtract
|
||||
# one entry from a directory. `bash` is invoked by absolute path below so the
|
||||
# interpreter itself does not have to be on this PATH.
|
||||
NOPY_BIN="$TMPDIR_T/nopython-bin"
|
||||
mkdir -p "$NOPY_BIN"
|
||||
for b in awk cat cut dirname basename grep sed pwd rm mkdir tr; do
|
||||
src="$(command -v "$b" 2>/dev/null || true)"
|
||||
[[ -n "$src" ]] && ln -sf "$src" "$NOPY_BIN/$b"
|
||||
done
|
||||
|
||||
# A python3 that runs but cannot import yaml. A shim on PATH re-execs the real
|
||||
# interpreter with a PYTHONPATH entry holding a `yaml` module that raises on
|
||||
# import; PYTHONPATH precedes site-packages on sys.path, so it shadows a real
|
||||
# PyYAML install without touching it.
|
||||
SHADOW="$TMPDIR_T/shadow"
|
||||
mkdir -p "$SHADOW"
|
||||
printf 'raise ImportError("PyYAML deliberately unavailable in this fixture")\n' \
|
||||
> "$SHADOW/yaml.py"
|
||||
NOYAML_BIN="$TMPDIR_T/noyaml-bin"
|
||||
mkdir -p "$NOYAML_BIN"
|
||||
cat > "$NOYAML_BIN/python3" <<EOF
|
||||
#!/bin/sh
|
||||
PYTHONPATH="$SHADOW\${PYTHONPATH:+:\$PYTHONPATH}" exec "$REAL_PYTHON" "\$@"
|
||||
EOF
|
||||
chmod +x "$NOYAML_BIN/python3"
|
||||
|
||||
# Sanity-check the two fixtures themselves before trusting any verdict they
|
||||
# produce. A shim that silently still imports yaml would make every PyYAML
|
||||
# assertion below pass for the wrong reason.
|
||||
if PATH="$NOYAML_BIN:$PATH" python3 -c 'import yaml' 2>/dev/null; then
|
||||
fail "the no-PyYAML shim does not actually shadow PyYAML — every PyYAML assertion below would be vacuous"
|
||||
else
|
||||
pass "fixture check: the no-PyYAML shim makes 'import yaml' fail while python3 still runs"
|
||||
fi
|
||||
if PATH="$NOPY_BIN" command -v python3 > /dev/null 2>&1; then
|
||||
fail "the no-python3 PATH still resolves python3 — every python3 assertion below would be vacuous"
|
||||
else
|
||||
pass "fixture check: the no-python3 PATH resolves no python3"
|
||||
fi
|
||||
|
||||
# A minimal, entirely clean subject for each script. The preflight must fire
|
||||
# before any measurement, so the subject's own content is irrelevant — which is
|
||||
# exactly what makes a clean one the right choice: nothing else can produce the
|
||||
# non-zero exit these cases assert.
|
||||
SUBJECT_SKILL_DIR="$TMPDIR_T/subject/my-skill"
|
||||
mkdir -p "$SUBJECT_SKILL_DIR"
|
||||
cat > "$SUBJECT_SKILL_DIR/SKILL.md" <<'EOF'
|
||||
---
|
||||
name: my-skill
|
||||
description: A short valid description. Do not use for anything else.
|
||||
---
|
||||
|
||||
Do the thing.
|
||||
EOF
|
||||
SUBJECT_AGENT_ROOT="$TMPDIR_T/subject-agent"
|
||||
mkdir -p "$SUBJECT_AGENT_ROOT/.apm/agents"
|
||||
cat > "$SUBJECT_AGENT_ROOT/apm.yml" <<'EOF'
|
||||
name: test-package
|
||||
version: 0.1.0
|
||||
type: skill
|
||||
EOF
|
||||
cat > "$SUBJECT_AGENT_ROOT/.apm/agents/my-agent.agent.md" <<'EOF'
|
||||
---
|
||||
name: my-agent
|
||||
description: A short valid description. Do not use for anything else.
|
||||
---
|
||||
|
||||
You are a test agent. When invoked, do the thing.
|
||||
EOF
|
||||
|
||||
# probe_preflight <label> <env-kind: nopython|noyaml> <expect-needle> <cmd...>
|
||||
probe_preflight() {
|
||||
local label="$1" kind="$2" needle="$3"
|
||||
shift 3
|
||||
local out status=0
|
||||
set +e
|
||||
if [[ "$kind" == nopython ]]; then
|
||||
out="$(env -i PATH="$NOPY_BIN" HOME="$HOME" "$BASH_BIN" "$@" 2>&1)"
|
||||
else
|
||||
out="$(env PATH="$NOYAML_BIN:$PATH" "$BASH_BIN" "$@" 2>&1)"
|
||||
fi
|
||||
status=$?
|
||||
set -e
|
||||
if [[ $status -eq 0 ]]; then
|
||||
fail "$label exited 0 — a missing hard dependency became a vacuous pass (output: ${out:-<empty>})"
|
||||
elif [[ "$out" != *"$needle"* ]]; then
|
||||
# The needle is the DIAGNOSTIC ("python3 is required"), not the bare word.
|
||||
# Deleting the preflight entirely would still produce a non-zero exit and a
|
||||
# message mentioning python3 -- bash's own "python3: command not found" --
|
||||
# so a bare-word needle would go green on a script with no preflight at all.
|
||||
fail "$label exited $status but never produced the '$needle' diagnostic (output: ${out:-<empty>})"
|
||||
elif [[ "$kind" == nopython && "$out" == *PyYAML* ]]; then
|
||||
fail "$label reported PyYAML when python3 itself is missing — that sends the reader to install the wrong thing (output: $out)"
|
||||
else
|
||||
pass "$label"
|
||||
fi
|
||||
}
|
||||
|
||||
echo ""
|
||||
echo "--- a PATH with no python3 is a hard failure in all three scripts, naming python3 ---"
|
||||
probe_preflight "scripts/skill-size-check.sh reports missing python3" \
|
||||
nopython "python3 is required" \
|
||||
"$HOOK" "$SUBJECT_SKILL_DIR/SKILL.md"
|
||||
probe_preflight "skill-audit/scripts/validate.sh reports missing python3" \
|
||||
nopython "python3 is required" \
|
||||
"$SKILL_VALIDATE" "$SUBJECT_SKILL_DIR"
|
||||
probe_preflight "agent-audit/scripts/validate.sh reports missing python3" \
|
||||
nopython "python3 is required" \
|
||||
"$AGENT_VALIDATE" "$SUBJECT_AGENT_ROOT/.apm/agents/my-agent.agent.md"
|
||||
|
||||
echo ""
|
||||
echo "--- a python3 that cannot import yaml is a hard failure in all three scripts, naming PyYAML ---"
|
||||
probe_preflight "scripts/skill-size-check.sh reports missing PyYAML" \
|
||||
noyaml "PyYAML is required" \
|
||||
"$HOOK" "$SUBJECT_SKILL_DIR/SKILL.md"
|
||||
probe_preflight "skill-audit/scripts/validate.sh reports missing PyYAML" \
|
||||
noyaml "PyYAML is required" \
|
||||
"$SKILL_VALIDATE" "$SUBJECT_SKILL_DIR"
|
||||
probe_preflight "agent-audit/scripts/validate.sh reports missing PyYAML" \
|
||||
noyaml "PyYAML is required" \
|
||||
"$AGENT_VALIDATE" "$SUBJECT_AGENT_ROOT/.apm/agents/my-agent.agent.md"
|
||||
|
||||
# The control. Without it, "fails when the dependency is missing" is satisfied by
|
||||
# a script that fails unconditionally, and the two cases above would be green on
|
||||
# a gate that never runs at all.
|
||||
echo ""
|
||||
echo "--- control: with both dependencies present the same subjects pass ---"
|
||||
for probe in "$HOOK:$SUBJECT_SKILL_DIR/SKILL.md" \
|
||||
"$SKILL_VALIDATE:$SUBJECT_SKILL_DIR" \
|
||||
"$AGENT_VALIDATE:$SUBJECT_AGENT_ROOT/.apm/agents/my-agent.agent.md"; do
|
||||
script="${probe%%:*}"
|
||||
arg="${probe#*:}"
|
||||
set +e
|
||||
ctl_out="$(bash "$script" "$arg" 2>&1)"
|
||||
ctl_rc=$?
|
||||
set -e
|
||||
if [[ $ctl_rc -eq 0 ]]; then
|
||||
pass "${script#"$REPO_ROOT/"} exits 0 on a clean subject with python3 and PyYAML available"
|
||||
else
|
||||
fail "${script#"$REPO_ROOT/"} failed a clean subject (exit $ctl_rc): $ctl_out"
|
||||
fi
|
||||
done
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 3. verbose: true on the skill-size-check hook, in BOTH manifests
|
||||
# ---------------------------------------------------------------------------
|
||||
# .pre-commit-config.yaml governs this repo; .pre-commit-hooks.yaml is what a
|
||||
# CONSUMER repo gets when it points at this one. Dropping the flag from either
|
||||
# silences the SUGGESTION tier for that audience alone, which is the hardest
|
||||
# version of the defect to notice.
|
||||
echo ""
|
||||
echo "--- the skill-size-check hook declares verbose: true in both manifests ---"
|
||||
VERBOSE_REPORT="$(python3 - "$REPO_ROOT" <<'PY'
|
||||
import os
|
||||
import sys
|
||||
|
||||
import yaml
|
||||
|
||||
root = sys.argv[1]
|
||||
|
||||
|
||||
def emit(status, msg):
|
||||
print("%s\t%s" % (status, msg))
|
||||
|
||||
|
||||
# Repo config: nested repos[].hooks[].
|
||||
path = os.path.join(root, '.pre-commit-config.yaml')
|
||||
try:
|
||||
with open(path, encoding='utf-8') as fh:
|
||||
cfg = yaml.safe_load(fh) or {}
|
||||
except Exception as exc:
|
||||
emit('FAIL', '.pre-commit-config.yaml did not parse: %s' % exc)
|
||||
cfg = {}
|
||||
found = None
|
||||
for repo in cfg.get('repos') or []:
|
||||
for hook in (repo.get('hooks') or []):
|
||||
if hook.get('id') == 'skill-size-check':
|
||||
found = hook
|
||||
if found is None:
|
||||
emit('FAIL', '.pre-commit-config.yaml declares no hook with id skill-size-check')
|
||||
elif found.get('verbose') is True:
|
||||
emit('PASS', '.pre-commit-config.yaml: skill-size-check is verbose: true')
|
||||
else:
|
||||
emit('FAIL', '.pre-commit-config.yaml: skill-size-check has verbose=%r — '
|
||||
'pre-commit prints nothing for a passing hook, so every '
|
||||
'ADR-0020 SUGGESTION is swallowed' % (found.get('verbose'),))
|
||||
|
||||
# Consumer manifest: a flat list of hooks.
|
||||
path = os.path.join(root, '.pre-commit-hooks.yaml')
|
||||
try:
|
||||
with open(path, encoding='utf-8') as fh:
|
||||
hooks = yaml.safe_load(fh) or []
|
||||
except Exception as exc:
|
||||
emit('FAIL', '.pre-commit-hooks.yaml did not parse: %s' % exc)
|
||||
hooks = []
|
||||
found = None
|
||||
for hook in hooks:
|
||||
if isinstance(hook, dict) and hook.get('id') == 'kyberforge-skill-size-check':
|
||||
found = hook
|
||||
if found is None:
|
||||
emit('FAIL', '.pre-commit-hooks.yaml declares no hook with id kyberforge-skill-size-check')
|
||||
elif found.get('verbose') is True:
|
||||
emit('PASS', '.pre-commit-hooks.yaml: kyberforge-skill-size-check is verbose: true')
|
||||
else:
|
||||
emit('FAIL', '.pre-commit-hooks.yaml: kyberforge-skill-size-check has verbose=%r — '
|
||||
'a consumer repo would never see the SUGGESTION tier'
|
||||
% (found.get('verbose'),))
|
||||
PY
|
||||
)"
|
||||
while IFS=$'\t' read -r status msg; do
|
||||
[[ -n "$status" ]] || continue
|
||||
if [[ "$status" == PASS ]]; then
|
||||
pass "$msg"
|
||||
else
|
||||
fail "$msg"
|
||||
fi
|
||||
done <<< "$VERBOSE_REPORT"
|
||||
|
||||
echo ""
|
||||
echo "Results: $PASS passed, $FAIL failed"
|
||||
[[ $FAIL -eq 0 ]]
|
||||
365
tests/test-adr0020-differential.sh
Executable file
365
tests/test-adr0020-differential.sh
Executable file
@@ -0,0 +1,365 @@
|
||||
#!/usr/bin/env bash
|
||||
# Differential test: scripts/skill-size-check.sh (the pre-commit hook) and
|
||||
# skill-audit/scripts/validate.sh (the in-skill auditor) must reach the SAME
|
||||
# ADR-0020 verdict on the same file.
|
||||
#
|
||||
# Why this exists as a separate suite. tests/test-skill-size-check.sh already
|
||||
# asserts the two agree on their CONSTANTS, and that assertion is necessary but
|
||||
# demonstrably not sufficient: a previous review found the two scripts disagreeing
|
||||
# on real files while every constant matched perfectly. Constants are one of the
|
||||
# ways two hand-duplicated implementations diverge; comparison operators, message
|
||||
# wording, which value gets measured, and which branch runs first are the others,
|
||||
# and none of them is visible to a constant check.
|
||||
#
|
||||
# The consequence of divergence is specific and bad: skill-audit reports a skill
|
||||
# ready to ship and the commit hook then rejects it, or worse, the reverse. So the
|
||||
# comparison here is over VERDICTS on files, not over source text.
|
||||
#
|
||||
# Scope: the ADR-0020 axes the two scripts share — description length and tier,
|
||||
# body word count and tier, dangling routing targets, missing references/
|
||||
# pointers, the two Gotchas suggestions, the missing-boundary-clause suggestion,
|
||||
# a declined resolution, and an empty description. The two scripts legitimately
|
||||
# differ elsewhere (validate.sh also checks name/directory agreement, script
|
||||
# executability and the 1024-char spec backstop; the hook checks whole-file lines
|
||||
# and words), and those lines are ignored rather than being forced into a shared
|
||||
# shape they were never meant to have.
|
||||
#
|
||||
# Run over the real 39-skill corpus AND over purpose-built fixtures that sit ON
|
||||
# each boundary. The corpus alone is not enough — it happens not to contain a
|
||||
# file at exactly 900 body words, which is precisely where an inclusive/exclusive
|
||||
# comparison mismatch would hide.
|
||||
set -euo pipefail
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
HOOK="$REPO_ROOT/scripts/skill-size-check.sh"
|
||||
SKILL_VALIDATE="$REPO_ROOT/plugins/kyberforge/.apm/skills/skill-audit/scripts/validate.sh"
|
||||
|
||||
TMPDIR_T="$(mktemp -d)"
|
||||
trap 'rm -rf "$TMPDIR_T"' EXIT
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Fixtures: one per ADR-0020 axis, placed ON the boundary wherever there is one.
|
||||
# ---------------------------------------------------------------------------
|
||||
# Built inside a synthetic plugin monorepo so boundary-target resolution actually
|
||||
# runs for both scripts (in a bare temp dir both would decline, and "both
|
||||
# declined" is agreement about nothing).
|
||||
FIXTURE_ROOT="$TMPDIR_T/fixtures"
|
||||
FX="$FIXTURE_ROOT/plugins/fixture-plugin/.apm/skills"
|
||||
mkdir -p "$FX/sibling-skill" "$FIXTURE_ROOT/plugins/fixture-plugin/.apm/agents"
|
||||
|
||||
make_fx() {
|
||||
local name="$1" desc="$2" body_words="$3"
|
||||
mkdir -p "$FX/$name"
|
||||
{
|
||||
echo "---"
|
||||
echo "name: $name"
|
||||
echo "description: $desc"
|
||||
echo "---"
|
||||
echo ""
|
||||
python3 -c "print(' '.join(['word'] * $body_words))"
|
||||
} > "$FX/$name/SKILL.md"
|
||||
}
|
||||
desc_of_length() {
|
||||
python3 - "$1" <<'PY'
|
||||
import sys
|
||||
n = int(sys.argv[1])
|
||||
prefix = 'Use when doing the thing. Do not use for anything else. '
|
||||
print(prefix + 'x' * (n - len(prefix)))
|
||||
PY
|
||||
}
|
||||
CLEAN="Use when doing the thing. Do not use for anything else."
|
||||
|
||||
# Description tier boundaries, both sides of both thresholds.
|
||||
make_fx desc-249 "$(desc_of_length 249)" 10
|
||||
make_fx desc-250 "$(desc_of_length 250)" 10
|
||||
make_fx desc-251 "$(desc_of_length 251)" 10
|
||||
make_fx desc-400 "$(desc_of_length 400)" 10
|
||||
make_fx desc-401 "$(desc_of_length 401)" 10
|
||||
# Body tier boundaries, both sides of both thresholds.
|
||||
make_fx body-599 "$CLEAN" 599
|
||||
make_fx body-600 "$CLEAN" 600
|
||||
make_fx body-601 "$CLEAN" 601
|
||||
make_fx body-900 "$CLEAN" 900
|
||||
make_fx body-901 "$CLEAN" 901
|
||||
# Folding: the value has to be measured after YAML folding in both scripts.
|
||||
mkdir -p "$FX/folded-desc"
|
||||
{
|
||||
echo "---"
|
||||
echo "name: folded-desc"
|
||||
echo "description: >"
|
||||
python3 -c "print('\n'.join([' ' + 'x' * 40] * 11))"
|
||||
echo "---"
|
||||
echo ""
|
||||
echo "Do the thing."
|
||||
} > "$FX/folded-desc/SKILL.md"
|
||||
# Routing targets, one per tier the resolver can produce: resolves, dangles with
|
||||
# in-sentence corroboration (ERROR/FAIL), dangles alone (SUGGESTION on both
|
||||
# sides), route notation (ERROR/FAIL without corroboration), attributive
|
||||
# (silent). Each tier is here because the two scripts have to agree on the TIER,
|
||||
# not merely on the finding — a copy that promoted or demoted one of them would
|
||||
# otherwise pass this comparison.
|
||||
make_fx target-resolves "Use when doing the thing. Do not use for the other thing — use sibling-skill instead." 10
|
||||
make_fx target-dangles "Use when doing the thing. Do not use for the other thing — use sibling-skill or no-such-skill instead." 10
|
||||
make_fx target-dangles-lone "Use when doing the thing. Do not use for the other thing — use no-such-lone-skill instead." 10
|
||||
make_fx target-dangles-notation "Use when doing the thing. Do not use for the other thing — use /no-such-notation-skill instead." 10
|
||||
make_fx target-attributive "Use when doing the thing. Use pre-commit hooks instead of ad-hoc scripts." 10
|
||||
# No boundary clause at all.
|
||||
make_fx no-boundary "Use when the user wants the thing done." 10
|
||||
# A missing references/ pointer, and a present one.
|
||||
make_fx ref-missing "$CLEAN" 10
|
||||
printf '\nIf the caller needs detail, read references/absent.md first.\n' >> "$FX/ref-missing/SKILL.md"
|
||||
make_fx ref-present "$CLEAN" 10
|
||||
printf '\nIf the caller needs detail, read references/there.md first.\n' >> "$FX/ref-present/SKILL.md"
|
||||
mkdir -p "$FX/ref-present/references"
|
||||
echo "detail" > "$FX/ref-present/references/there.md"
|
||||
# Gotchas, over each guideline.
|
||||
make_fx gotchas-many "$CLEAN" 0
|
||||
cat >> "$FX/gotchas-many/SKILL.md" <<'EOF'
|
||||
|
||||
## Gotchas
|
||||
|
||||
- one trap here
|
||||
- two trap here
|
||||
- three trap here
|
||||
- four trap here
|
||||
- five trap here
|
||||
- six trap here
|
||||
- seven trap here
|
||||
|
||||
## Notes
|
||||
|
||||
EOF
|
||||
python3 -c "print(' '.join(['word'] * 200))" >> "$FX/gotchas-many/SKILL.md"
|
||||
# Gotchas over the 25% body-fraction guideline. Prose, not list items, so the
|
||||
# entry guideline cannot fire and the two suggestions stay separable: 26 section
|
||||
# words in a 100-word body is one word past the threshold.
|
||||
make_fx gotchas-fraction "$CLEAN" 0
|
||||
{
|
||||
echo ""
|
||||
echo "## Gotchas"
|
||||
echo ""
|
||||
python3 -c "print(' '.join(['word'] * 26))"
|
||||
echo ""
|
||||
echo "## Notes"
|
||||
echo ""
|
||||
python3 -c "print(' '.join(['word'] * 70))"
|
||||
} >> "$FX/gotchas-fraction/SKILL.md"
|
||||
|
||||
# Empty description — the shape that used to exit 0 in silence.
|
||||
mkdir -p "$FX/empty-desc"
|
||||
printf -- '---\nname: empty-desc\ndescription:\nmodel: sonnet\n---\n\nDo the thing.\n' \
|
||||
> "$FX/empty-desc/SKILL.md"
|
||||
# Every boundary shape at once, so a divergence that only appears when several
|
||||
# findings fire together is not missed.
|
||||
make_fx combined "$(desc_of_length 401)" 901
|
||||
printf '\nIf the caller needs detail, read references/absent.md first.\n' >> "$FX/combined/SKILL.md"
|
||||
|
||||
# A skill with routing targets and NO authoring root above it — deliberately
|
||||
# OUTSIDE the fixture plugin tree. Both scripts must decline out loud, and both
|
||||
# must decline identically; "both declined" is only meaningful as agreement if
|
||||
# the declining path is exercised on purpose somewhere.
|
||||
ORPHAN_ROOT="$TMPDIR_T/orphan"
|
||||
mkdir -p "$ORPHAN_ROOT/no-universe"
|
||||
{
|
||||
echo "---"
|
||||
echo "name: no-universe"
|
||||
echo "description: Use when doing the thing. Do not use for the other thing — use some-other-skill instead."
|
||||
echo "---"
|
||||
echo ""
|
||||
echo "Do the thing."
|
||||
} > "$ORPHAN_ROOT/no-universe/SKILL.md"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# The comparison
|
||||
# ---------------------------------------------------------------------------
|
||||
python3 - "$HOOK" "$SKILL_VALIDATE" "$REPO_ROOT" "$FX" "$ORPHAN_ROOT" <<'PYTHON'
|
||||
import glob
|
||||
import os
|
||||
import re
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
hook, validate, repo_root, fixture_dir, orphan_dir = sys.argv[1:6]
|
||||
|
||||
passes = 0
|
||||
failures = 0
|
||||
|
||||
|
||||
def ok(msg):
|
||||
global passes
|
||||
passes += 1
|
||||
print(" PASS: %s" % msg)
|
||||
|
||||
|
||||
def bad(msg):
|
||||
global failures
|
||||
failures += 1
|
||||
print(" FAIL: %s" % msg)
|
||||
|
||||
|
||||
# Tier prefixes. The hook writes `ERROR: ` / `SUGGESTION: ` / `INFO: `; the
|
||||
# auditor writes `FAIL ` / `SUGGESTION ` / `INFO ` and additionally `PASS `
|
||||
# lines, which carry no finding and are dropped.
|
||||
TIERS = (
|
||||
('ERROR', ('ERROR:', 'FAIL ')),
|
||||
('SUGGESTION', ('SUGGESTION:', 'SUGGESTION ')),
|
||||
('INFO', ('INFO:', 'INFO ')),
|
||||
)
|
||||
|
||||
# Each rule turns a finding line into a canonical token. Wording differs between
|
||||
# the two scripts by design (one addresses a committer, the other an auditor), so
|
||||
# the tokens deliberately capture the MEASUREMENT and not the sentence.
|
||||
RULES = (
|
||||
('DESC_CHARS', re.compile(r'description is (\d+) char')),
|
||||
('BODY_WORDS', re.compile(r'body is (\d+) words')),
|
||||
('ROUTE', re.compile(r"routes to '([^']+)'")),
|
||||
('MISSING_REF', re.compile(r'points at (references/[^\s,]+)')),
|
||||
('GOTCHA_ENTRIES', re.compile(r'Gotchas section has (\d+) entries')),
|
||||
('GOTCHA_FRACTION', re.compile(r'Gotchas section is (\d+) of (\d+) body words')),
|
||||
('NO_BOUNDARY_CLAUSE', re.compile(r'(description has no boundary clause)')),
|
||||
('RESOLUTION_DECLINED', re.compile(r'(boundary-target resolution DID NOT RUN)')),
|
||||
('DESC_EMPTY', re.compile(r'(description field is missing or empty)')),
|
||||
)
|
||||
|
||||
|
||||
def verdict(output):
|
||||
"""The set of ADR-0020 findings in a script's output, tier included.
|
||||
|
||||
Lines that match no rule are dropped rather than compared: the two scripts
|
||||
legitimately check different things outside ADR-0020 (name/directory
|
||||
agreement, script executability, the 1024-char spec backstop, whole-file
|
||||
line and word ceilings), and forcing those into the comparison would report
|
||||
a difference that is not a disagreement.
|
||||
"""
|
||||
found = set()
|
||||
for raw in output.splitlines():
|
||||
line = raw.strip()
|
||||
tier = None
|
||||
for name, prefixes in TIERS:
|
||||
if any(line.startswith(p) for p in prefixes):
|
||||
tier = name
|
||||
break
|
||||
if tier is None:
|
||||
continue
|
||||
for token, pattern in RULES:
|
||||
match = pattern.search(line)
|
||||
if match:
|
||||
found.add((tier, token) + tuple(match.groups()))
|
||||
return found
|
||||
|
||||
|
||||
def run(cmd):
|
||||
proc = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT)
|
||||
return proc.returncode, proc.stdout.decode('utf-8', 'replace')
|
||||
|
||||
|
||||
def compare(label, skill_dir):
|
||||
skill_md = os.path.join(skill_dir, 'SKILL.md')
|
||||
hook_rc, hook_out = run(['bash', hook, skill_md])
|
||||
audit_rc, audit_out = run(['bash', validate, skill_dir])
|
||||
hook_v = verdict(hook_out)
|
||||
audit_v = verdict(audit_out)
|
||||
|
||||
problems = []
|
||||
only_hook = sorted(hook_v - audit_v)
|
||||
only_audit = sorted(audit_v - hook_v)
|
||||
if only_hook:
|
||||
problems.append('only the hook reported %s' % (only_hook,))
|
||||
if only_audit:
|
||||
problems.append('only skill-audit reported %s' % (only_audit,))
|
||||
|
||||
# Exit codes are compared on the ADR-0020 axis only: an ERROR-tier ADR-0020
|
||||
# finding must make BOTH scripts non-zero, and neither may be turned
|
||||
# non-zero by a SUGGESTION. The raw codes cannot be compared directly —
|
||||
# validate.sh also fails on checks the hook does not run at all.
|
||||
hook_err = any(t == 'ERROR' for t, *_ in hook_v)
|
||||
audit_err = any(t == 'ERROR' for t, *_ in audit_v)
|
||||
if hook_err and hook_rc == 0:
|
||||
problems.append('the hook reported an ADR-0020 ERROR but exited 0')
|
||||
if audit_err and audit_rc == 0:
|
||||
problems.append('skill-audit reported an ADR-0020 FAIL but exited 0')
|
||||
if not hook_err and hook_rc != 0 and not _non_adr_hook_error(hook_out):
|
||||
problems.append('the hook exited %d with no ADR-0020 ERROR and no spec-ceiling ERROR'
|
||||
% hook_rc)
|
||||
|
||||
if problems:
|
||||
bad('%s: %s' % (label, '; '.join(problems)))
|
||||
else:
|
||||
return True
|
||||
return False
|
||||
|
||||
|
||||
def _non_adr_hook_error(output):
|
||||
"""True if the hook failed on a spec ceiling rather than an ADR-0020 gate.
|
||||
|
||||
MAX_LINES / MAX_WORDS are the hook's other ERROR sources and are outside
|
||||
this comparison, so a non-zero exit explained by one of them is not a
|
||||
disagreement.
|
||||
"""
|
||||
return bool(re.search(r'ERROR: .*(-line ceiling|-word ceiling \(~5,000 tokens)', output))
|
||||
|
||||
|
||||
# --- The real corpus -------------------------------------------------------
|
||||
corpus = sorted(glob.glob(os.path.join(repo_root, 'plugins', '*', '.apm', 'skills', '*')))
|
||||
corpus = [d for d in corpus if os.path.isfile(os.path.join(d, 'SKILL.md'))]
|
||||
print("")
|
||||
print("--- the two scripts agree on every skill in the live corpus (%d files) ---" % len(corpus))
|
||||
if len(corpus) < 30:
|
||||
bad('only %d corpus skills were discovered — the glob is wrong, so this leg '
|
||||
'proves nothing' % len(corpus))
|
||||
else:
|
||||
ok('discovered %d corpus skills to compare' % len(corpus))
|
||||
agreed = 0
|
||||
for skill_dir in corpus:
|
||||
rel = os.path.relpath(skill_dir, repo_root)
|
||||
if compare(rel, skill_dir):
|
||||
agreed += 1
|
||||
if agreed == len(corpus):
|
||||
ok('all %d corpus skills produce identical ADR-0020 verdicts from both scripts' % agreed)
|
||||
|
||||
# --- Boundary fixtures -----------------------------------------------------
|
||||
fixtures = sorted(d for d in (glob.glob(os.path.join(fixture_dir, '*'))
|
||||
+ glob.glob(os.path.join(orphan_dir, '*')))
|
||||
if os.path.isfile(os.path.join(d, 'SKILL.md')))
|
||||
print("")
|
||||
print("--- the two scripts agree on every boundary fixture (%d files) ---" % len(fixtures))
|
||||
if len(fixtures) < 15:
|
||||
bad('only %d fixtures were built — the fixture set is incomplete, so the '
|
||||
'boundaries the corpus does not cover are untested' % len(fixtures))
|
||||
else:
|
||||
ok('built %d boundary fixtures to compare' % len(fixtures))
|
||||
fx_agreed = 0
|
||||
for skill_dir in fixtures:
|
||||
if compare(os.path.basename(skill_dir), skill_dir):
|
||||
fx_agreed += 1
|
||||
if fx_agreed == len(fixtures):
|
||||
ok('all %d boundary fixtures produce identical ADR-0020 verdicts from both scripts' % fx_agreed)
|
||||
|
||||
# --- The comparison must not be vacuous ------------------------------------
|
||||
# Everything above would also pass if verdict() extracted nothing at all. So the
|
||||
# fixtures are required to have produced findings across every axis this suite
|
||||
# claims to compare — if a rule stops matching (a reworded message, say), that is
|
||||
# a silent loss of coverage and it fails here instead.
|
||||
print("")
|
||||
print("--- the comparison actually extracted findings on every axis it claims to cover ---")
|
||||
seen_tokens = set()
|
||||
for skill_dir in fixtures:
|
||||
_, out = run(['bash', hook, os.path.join(skill_dir, 'SKILL.md')])
|
||||
for entry in verdict(out):
|
||||
seen_tokens.add(entry[1])
|
||||
_, out = run(['bash', validate, skill_dir])
|
||||
for entry in verdict(out):
|
||||
seen_tokens.add(entry[1])
|
||||
expected_tokens = {t for t, _ in RULES}
|
||||
missing = sorted(expected_tokens - seen_tokens)
|
||||
if missing:
|
||||
bad('no fixture produced a finding for %s — verdict() may no longer match '
|
||||
'those messages, and any disagreement on them would go unseen' % missing)
|
||||
else:
|
||||
ok('every one of the %d compared axes was exercised by at least one fixture'
|
||||
% len(expected_tokens))
|
||||
|
||||
print("")
|
||||
print("Results: %d passed, %d failed" % (passes, failures))
|
||||
sys.exit(1 if failures else 0)
|
||||
PYTHON
|
||||
286
tests/test-adr0020-frontmatter.sh
Executable file
286
tests/test-adr0020-frontmatter.sh
Executable file
@@ -0,0 +1,286 @@
|
||||
#!/usr/bin/env bash
|
||||
# Regression test for the two ways an ADR-0020 gate can be made to check NOTHING
|
||||
# while still exiting 0. Both were live defects, both were silent, and both sit
|
||||
# in the shared resolver block that all three scripts embed verbatim — so every
|
||||
# case below runs against all three.
|
||||
#
|
||||
# 1. THE FRONTMATTER BLOCKER. The frontmatter matcher used to be `^---\n`. A
|
||||
# UTF-8 BOM, a leading blank line, a trailing space after either marker, or
|
||||
# CRLF line endings all defeated it, and the miss was not reported: every
|
||||
# ADR-0020 check was skipped and the file passed. Measured at the time: a
|
||||
# 550-character description with a 1,000-word body exited 0 behind a BOM.
|
||||
# So this file asserts two complementary things — that each of those four
|
||||
# shapes is now TOLERATED (the findings actually fire), and that
|
||||
# frontmatter which genuinely cannot be parsed is a hard ERROR rather than
|
||||
# a quiet skip. A file that cannot be measured must never report green.
|
||||
#
|
||||
# 2. THE VALUELESS DESCRIPTION. `description:` with no value, followed by
|
||||
# another key, let a line regex's `\s*` cross the newline and capture the
|
||||
# NEXT key. The value then looked present (so "missing or empty" never
|
||||
# fired) and was empty once folded (so every ADR-0020 gate early-returned).
|
||||
# An agent file with one exited 0 with zero output through a BLOCKING
|
||||
# pre-push gate. All five spellings of "no value" are pinned here.
|
||||
#
|
||||
# Both fixtures carry an over-ceiling description AND an over-ceiling body on
|
||||
# purpose: asserting a non-zero exit alone would be satisfied by the "cannot
|
||||
# parse" error itself, so the tolerated shapes are asserted on the CONTENT of
|
||||
# the findings, not on the exit code.
|
||||
set -euo pipefail
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
HOOK="$REPO_ROOT/scripts/skill-size-check.sh"
|
||||
SKILL_VALIDATE="$REPO_ROOT/plugins/kyberforge/.apm/skills/skill-audit/scripts/validate.sh"
|
||||
AGENT_VALIDATE="$REPO_ROOT/plugins/kyberforge/.apm/skills/agent-audit/scripts/validate.sh"
|
||||
PASS=0
|
||||
FAIL=0
|
||||
|
||||
pass() { echo " PASS: $1"; PASS=$((PASS + 1)); }
|
||||
fail() { echo " FAIL: $1"; FAIL=$((FAIL + 1)); }
|
||||
|
||||
TMPDIR_T="$(mktemp -d)"
|
||||
trap 'rm -rf "$TMPDIR_T"' EXIT
|
||||
|
||||
DESC_CHARS=450
|
||||
BODY_WORDS=1000
|
||||
|
||||
# write_fixture <kind> <path> <name> — one generator for both file shapes.
|
||||
#
|
||||
# Byte-level control is the point: BOM placement, line endings and trailing
|
||||
# whitespace are exactly what is under test, so the file is emitted in binary
|
||||
# mode rather than through a shell heredoc that would normalise them.
|
||||
write_fixture() {
|
||||
python3 - "$1" "$2" "$3" "$DESC_CHARS" "$BODY_WORDS" <<'PY'
|
||||
import sys
|
||||
|
||||
kind, path, name, desc_chars, body_words = sys.argv[1:6]
|
||||
desc = 'x' * int(desc_chars)
|
||||
body = ' '.join(['word'] * int(body_words))
|
||||
|
||||
# The default, well-formed shape. Variants below mutate it.
|
||||
open_marker = '---'
|
||||
close_marker = '---'
|
||||
prefix = ''
|
||||
newline = '\n'
|
||||
fm_lines = ['name: ' + name, 'description: ' + desc]
|
||||
|
||||
if kind == 'plain':
|
||||
pass
|
||||
elif kind == 'bom':
|
||||
prefix = ''
|
||||
elif kind == 'leading-blanks':
|
||||
prefix = '\n\n \n'
|
||||
elif kind == 'trailing-ws':
|
||||
open_marker = '--- '
|
||||
close_marker = '---\t '
|
||||
elif kind == 'crlf':
|
||||
newline = '\r\n'
|
||||
elif kind == 'no-close':
|
||||
close_marker = None
|
||||
elif kind == 'yaml-list':
|
||||
fm_lines = ['- one', '- two']
|
||||
elif kind == 'yaml-string':
|
||||
fm_lines = ['just a bare scalar, not a mapping']
|
||||
elif kind == 'yaml-none':
|
||||
fm_lines = []
|
||||
elif kind == 'yaml-malformed':
|
||||
fm_lines = ['name: ' + name, 'description: "unterminated', 'tabs:\t- a']
|
||||
elif kind == 'desc-no-value':
|
||||
# The exact shape that exited 0 with zero output: a line regex's `\s*`
|
||||
# crosses the newline and captures `model: sonnet` as the description.
|
||||
fm_lines = ['name: ' + name, 'description:', 'model: sonnet']
|
||||
elif kind == 'desc-null':
|
||||
fm_lines = ['name: ' + name, 'description: null']
|
||||
elif kind == 'desc-single-quoted-empty':
|
||||
fm_lines = ['name: ' + name, "description: ''"]
|
||||
elif kind == 'desc-double-quoted-empty':
|
||||
fm_lines = ['name: ' + name, 'description: ""']
|
||||
elif kind == 'desc-empty-fold':
|
||||
fm_lines = ['name: ' + name, 'description: >']
|
||||
else:
|
||||
raise SystemExit('unknown fixture kind: %s' % kind)
|
||||
|
||||
parts = [prefix, open_marker, newline]
|
||||
for line in fm_lines:
|
||||
parts.append(line)
|
||||
parts.append(newline)
|
||||
if close_marker is not None:
|
||||
parts.append(close_marker)
|
||||
parts.append(newline)
|
||||
parts.append(newline)
|
||||
parts.append(body)
|
||||
parts.append(newline)
|
||||
|
||||
with open(path, 'wb') as fh:
|
||||
fh.write(''.join(parts).encode('utf-8'))
|
||||
PY
|
||||
}
|
||||
|
||||
# Builds all three subjects for one fixture kind and echoes nothing; the paths
|
||||
# are fixed by convention so the probes below can find them.
|
||||
#
|
||||
# skill-audit takes a DIRECTORY (SKILL.md inside it, name matching the dir);
|
||||
# agent-audit takes a FILE inside an apm package. The hook takes the SKILL.md
|
||||
# directly, so it and skill-audit share one file.
|
||||
build_subjects() {
|
||||
local kind="$1" base="$TMPDIR_T/$1"
|
||||
rm -rf "$base"
|
||||
mkdir -p "$base/skill/my-skill" "$base/agent/.apm/agents"
|
||||
cat > "$base/agent/apm.yml" <<'EOF'
|
||||
name: test-package
|
||||
version: 0.1.0
|
||||
type: skill
|
||||
EOF
|
||||
write_fixture "$kind" "$base/skill/my-skill/SKILL.md" my-skill
|
||||
write_fixture "$kind" "$base/agent/.apm/agents/my-agent.agent.md" my-agent
|
||||
}
|
||||
|
||||
# probe_all <label> <kind> <needle>... — runs all three scripts over the fixture
|
||||
# and requires every one of them to exit non-zero AND report every needle. One
|
||||
# assertion per script would let two of them drift apart while the suite stayed
|
||||
# green; the whole point of the shared resolver block is that they cannot.
|
||||
#
|
||||
# A needle written `@skills:<text>` is asserted for the hook and skill-audit but
|
||||
# NOT for agent-audit. There is exactly one such needle in this file — the body
|
||||
# word ceiling — and the exemption is the ADR, not a workaround: ADR-0020 gives
|
||||
# agents the description gates and deliberately NO body word gate, because an
|
||||
# agent body becomes the system prompt of a fresh context rather than competing
|
||||
# with the caller's live conversation. Demanding a body finding from agent-audit
|
||||
# would be demanding the ADR be contradicted.
|
||||
probe_all() {
|
||||
local label="$1" kind="$2"
|
||||
shift 2
|
||||
local base="$TMPDIR_T/$kind"
|
||||
local -a targets=(
|
||||
"hook|$HOOK|$base/skill/my-skill/SKILL.md"
|
||||
"skill-audit|$SKILL_VALIDATE|$base/skill/my-skill"
|
||||
"agent-audit|$AGENT_VALIDATE|$base/agent/.apm/agents/my-agent.agent.md"
|
||||
)
|
||||
local problems=""
|
||||
for target in "${targets[@]}"; do
|
||||
local who="${target%%|*}" rest="${target#*|}"
|
||||
local script="${rest%%|*}" arg="${rest#*|}"
|
||||
local out status=0
|
||||
set +e
|
||||
out="$(bash "$script" "$arg" 2>&1)"
|
||||
status=$?
|
||||
set -e
|
||||
if [[ $status -eq 0 ]]; then
|
||||
problems="$problems [$who exited 0: ${out:-<no output>}]"
|
||||
continue
|
||||
fi
|
||||
for needle in "$@"; do
|
||||
if [[ "$needle" == @skills:* ]]; then
|
||||
# Spelled as a full `if`, not `[[ ... ]] && continue`. Under `set -e` the
|
||||
# short form's exit status is the test's when it is false, and relying on
|
||||
# the &&-list exemption to keep that from aborting the run is a footgun
|
||||
# one edit away from biting.
|
||||
if [[ "$who" == agent-audit ]]; then
|
||||
continue
|
||||
fi
|
||||
needle="${needle#@skills:}"
|
||||
fi
|
||||
if [[ "$out" != *"$needle"* ]]; then
|
||||
problems="$problems [$who never said '$needle': $out]"
|
||||
fi
|
||||
done
|
||||
done
|
||||
if [[ -z "$problems" ]]; then
|
||||
pass "$label"
|
||||
else
|
||||
fail "$label —$problems"
|
||||
fi
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 1a. Tolerated frontmatter shapes — the gates must RUN, not merely not-pass
|
||||
# ---------------------------------------------------------------------------
|
||||
# The needles are the FINDINGS, not the exit code. A script that rejected the BOM
|
||||
# outright would exit non-zero too, and would still be skipping every ADR-0020
|
||||
# measurement — which is the defect, one error message later.
|
||||
echo ""
|
||||
echo "--- a BOM, leading blanks, trailing marker whitespace and CRLF are all tolerated, and the gates still fire ---"
|
||||
for kind in plain bom leading-blanks trailing-ws crlf; do
|
||||
build_subjects "$kind"
|
||||
done
|
||||
probe_all "control: a well-formed over-ceiling file fails on BOTH the description and the body" \
|
||||
plain "description is $DESC_CHARS char" "@skills:body is $BODY_WORDS words"
|
||||
probe_all "a UTF-8 BOM does not hide an over-ceiling description or body" \
|
||||
bom "description is $DESC_CHARS char" "@skills:body is $BODY_WORDS words"
|
||||
probe_all "leading blank lines before the opening --- do not hide the findings" \
|
||||
leading-blanks "description is $DESC_CHARS char" "@skills:body is $BODY_WORDS words"
|
||||
probe_all "trailing whitespace after either --- marker does not hide the findings" \
|
||||
trailing-ws "description is $DESC_CHARS char" "@skills:body is $BODY_WORDS words"
|
||||
# Note on what this last one can and cannot detect. read_text() opens the file in
|
||||
# TEXT mode, so Python's universal-newline translation turns \r\n into \n before
|
||||
# the frontmatter matcher ever sees it — verified by mutation: reverting
|
||||
# FRONTMATTER_RE to the old `^---\n(.*?)\n---` breaks the leading-blanks and
|
||||
# trailing-whitespace cases above but NOT this one. So this case pins the
|
||||
# end-to-end behaviour (a CRLF file is measured, not skipped) rather than the
|
||||
# `\r?\n` alternations in the regex, and it would catch a future switch to binary
|
||||
# reads or to a newline='' open. Kept for that reason, and labelled so nobody
|
||||
# reads it as covering more than it does.
|
||||
probe_all "CRLF line endings do not hide the findings" \
|
||||
crlf "description is $DESC_CHARS char" "@skills:body is $BODY_WORDS words"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 1b. Unparseable frontmatter is a hard ERROR, never a quiet skip
|
||||
# ---------------------------------------------------------------------------
|
||||
echo ""
|
||||
echo "--- genuinely unparseable frontmatter exits non-zero with a message, rather than passing quietly ---"
|
||||
for kind in no-close yaml-list yaml-string yaml-none yaml-malformed; do
|
||||
build_subjects "$kind"
|
||||
done
|
||||
probe_all "frontmatter with no closing --- is reported, not skipped" \
|
||||
no-close "frontmatter"
|
||||
probe_all "frontmatter that parses to a LIST is reported, not skipped" \
|
||||
yaml-list "frontmatter"
|
||||
probe_all "frontmatter that parses to a STRING is reported, not skipped" \
|
||||
yaml-string "frontmatter"
|
||||
probe_all "frontmatter that parses to None (empty block) is reported, not skipped" \
|
||||
yaml-none "frontmatter"
|
||||
probe_all "malformed YAML in the frontmatter is reported, not skipped" \
|
||||
yaml-malformed "frontmatter"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 2. A valueless description is a hard FAIL in all three scripts
|
||||
# ---------------------------------------------------------------------------
|
||||
# All five spellings mean the same thing to a YAML parser — an empty value — and
|
||||
# all five have to be decided on the FOLDED value rather than on a line regex.
|
||||
# `description:` followed by `model: sonnet` is the one that shipped: it made the
|
||||
# value look present, skipped the "missing or empty" failure, and then
|
||||
# early-returned out of every ADR-0020 gate on the genuinely empty folded value.
|
||||
echo ""
|
||||
echo "--- every spelling of a valueless description hard-FAILs in all three scripts ---"
|
||||
for kind in desc-no-value desc-null desc-single-quoted-empty desc-double-quoted-empty desc-empty-fold; do
|
||||
build_subjects "$kind"
|
||||
done
|
||||
probe_all "'description:' with no value (next key not captured as the value) FAILs" \
|
||||
desc-no-value "description field is missing or empty"
|
||||
probe_all "'description: null' FAILs" \
|
||||
desc-null "description field is missing or empty"
|
||||
probe_all "\"description: ''\" FAILs" \
|
||||
desc-single-quoted-empty "description field is missing or empty"
|
||||
probe_all "'description: \"\"' FAILs" \
|
||||
desc-double-quoted-empty "description field is missing or empty"
|
||||
probe_all "'description: >' with nothing folded under it FAILs" \
|
||||
desc-empty-fold "description field is missing or empty"
|
||||
|
||||
# The specific regression, spelled out: the valueless-description agent file must
|
||||
# not merely fail — it must not be SILENT. Zero output on a blocking gate is what
|
||||
# made this un-diagnosable, so the output is asserted non-empty independently.
|
||||
echo ""
|
||||
echo "--- the valueless-description agent file produces output, not silence ---"
|
||||
build_subjects desc-no-value
|
||||
set +e
|
||||
SILENT_OUT="$(bash "$AGENT_VALIDATE" "$TMPDIR_T/desc-no-value/agent/.apm/agents/my-agent.agent.md" 2>&1)"
|
||||
SILENT_RC=$?
|
||||
set -e
|
||||
if [[ $SILENT_RC -ne 0 && -n "$SILENT_OUT" ]]; then
|
||||
pass "agent-audit reports a valueless description rather than exiting 0 with zero output"
|
||||
else
|
||||
fail "agent-audit exited $SILENT_RC with output '${SILENT_OUT:-<empty>}' — the original defect was exit 0 and total silence on a blocking pre-push gate"
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo "Results: $PASS passed, $FAIL failed"
|
||||
[[ $FAIL -eq 0 ]]
|
||||
409
tests/test-adr0020-targets.sh
Executable file
409
tests/test-adr0020-targets.sh
Executable file
@@ -0,0 +1,409 @@
|
||||
#!/usr/bin/env bash
|
||||
# Regression test for the two properties of ADR-0020 boundary-target resolution
|
||||
# that decide whether the gate can be trusted at all.
|
||||
#
|
||||
# 1. MACHINE INDEPENDENCE. The resolution universe is derived by walking up
|
||||
# FROM THE TARGET FILE to an authoring root, and when one is found the
|
||||
# deployed .claude/ and .agents/ trees are deliberately NOT consulted. Those
|
||||
# trees are `apm install` output — gitignored, and present only on a machine
|
||||
# that has run it. Four cross-plugin targets in this repo resolved through
|
||||
# .claude/skills/ alone, so the same commit measured 2 dangling targets on a
|
||||
# developer machine and 6 on a fresh clone. A gate shipping hot with no
|
||||
# baseline cannot give two answers, so this file asserts the verdict is
|
||||
# identical with and without a deployed tree — on a synthetic fixture AND on
|
||||
# the real 39-skill corpus.
|
||||
#
|
||||
# 2. THE BARE-TARGET GRAMMAR RULE. A hyphenated token used as a compound
|
||||
# MODIFIER ("pre-commit hooks", "pull-request template") is prose, not a
|
||||
# route; a terminal one is a real target. Getting this wrong in either
|
||||
# direction is fatal: firing on prose makes the gate untrustworthy and it
|
||||
# gets turned off, while suppressing too much deletes the only two true
|
||||
# positives the corpus has. Both live true positives are BARE, which is why
|
||||
# the rule keys on the FOLLOWER TOKEN rather than on marking, and why they
|
||||
# are pinned by name below — a future false-positive fix must not be able to
|
||||
# quietly take them with it.
|
||||
#
|
||||
# 3. IN-SENTENCE CORROBORATION. Terminal position alone is not evidence of a
|
||||
# route: "run `pre-commit` instead", "see `commit-msg`", "use the clean-up
|
||||
# instead" and "run unit-tests" are all terminal, all prose, and all were
|
||||
# hard FAILs with no suppression mechanism anywhere in the gate. A
|
||||
# prose-form target therefore blocks only when its own sentence names
|
||||
# another target that RESOLVES; otherwise it is reported at SUGGESTION tier
|
||||
# and the commit proceeds. Route NOTATION (`/name`, `-> name`) is exempt
|
||||
# and always blocks. Both halves are asserted below: the prose class must
|
||||
# report-not-block, and the notation and corroborated forms must still
|
||||
# ERROR, or the fix would have eaten the gate rather than narrowed it.
|
||||
set -euo pipefail
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
HOOK="$REPO_ROOT/scripts/skill-size-check.sh"
|
||||
PASS=0
|
||||
FAIL=0
|
||||
|
||||
pass() { echo " PASS: $1"; PASS=$((PASS + 1)); }
|
||||
fail() { echo " FAIL: $1"; FAIL=$((FAIL + 1)); }
|
||||
|
||||
TMPDIR_T="$(mktemp -d)"
|
||||
trap 'rm -rf "$TMPDIR_T"' EXIT
|
||||
|
||||
# write_skill <skill-dir> <name> <desc>
|
||||
write_skill() {
|
||||
mkdir -p "$1"
|
||||
{
|
||||
echo "---"
|
||||
echo "name: $2"
|
||||
echo "description: $3"
|
||||
echo "---"
|
||||
echo ""
|
||||
echo "Do the thing."
|
||||
} > "$1/SKILL.md"
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 1. Machine independence — synthetic fixture
|
||||
# ---------------------------------------------------------------------------
|
||||
# Two trees, identical except that one also carries a deployed .claude/ tree
|
||||
# holding a skill and an agent that exist NOWHERE in plugins/. The subject routes
|
||||
# to one name that lives in a sibling plugin (must resolve in both) and one that
|
||||
# lives only in .claude/ (must DANGLE in both — an authoring root exists, so the
|
||||
# deployed tree is not part of the universe).
|
||||
#
|
||||
# If the deployed tree were consulted, the second target would resolve on the
|
||||
# machine that has run `apm install` and dangle on a fresh clone. That is the
|
||||
# 2-vs-6 defect exactly, at fixture scale.
|
||||
echo ""
|
||||
echo "--- the same file gets the same verdict with and without a deployed .claude/ tree ---"
|
||||
build_tree() {
|
||||
local root="$1"
|
||||
write_skill "$root/plugins/other-plugin/.apm/skills/cross-plugin-skill" cross-plugin-skill \
|
||||
"Use when doing the other thing. Do not use for anything else."
|
||||
write_skill "$root/plugins/subject-plugin/.apm/skills/my-skill" my-skill \
|
||||
"Use when doing the thing. Do not use for the other thing — use cross-plugin-skill or deployed-only-skill instead."
|
||||
}
|
||||
build_tree "$TMPDIR_T/no-claude"
|
||||
build_tree "$TMPDIR_T/with-claude"
|
||||
# The deployed tree, present only in the second root. Both a skill and an agent,
|
||||
# because both are valid routing targets and both would leak.
|
||||
mkdir -p "$TMPDIR_T/with-claude/.claude/skills/deployed-only-skill" \
|
||||
"$TMPDIR_T/with-claude/.claude/agents"
|
||||
: > "$TMPDIR_T/with-claude/.claude/agents/deployed-only-agent.md"
|
||||
|
||||
run_subject() {
|
||||
local root="$1" out
|
||||
set +e
|
||||
out="$(bash "$HOOK" "$root/plugins/subject-plugin/.apm/skills/my-skill/SKILL.md" 2>&1)"
|
||||
set -e
|
||||
# Normalise the tree root out of the paths so the two runs are comparable.
|
||||
printf '%s\n' "$out" | sed "s#$root#<ROOT>#g"
|
||||
}
|
||||
NO_CLAUDE_OUT="$(run_subject "$TMPDIR_T/no-claude")"
|
||||
WITH_CLAUDE_OUT="$(run_subject "$TMPDIR_T/with-claude")"
|
||||
|
||||
if [[ "$NO_CLAUDE_OUT" == "$WITH_CLAUDE_OUT" ]]; then
|
||||
pass "identical output with and without a deployed .claude/ tree"
|
||||
else
|
||||
fail "the deployed .claude/ tree changed the verdict — without: [$NO_CLAUDE_OUT] with: [$WITH_CLAUDE_OUT]"
|
||||
fi
|
||||
# Identical-but-wrong is still possible (both could resolve everything, or
|
||||
# neither could resolve anything), so the CONTENT is asserted too: the
|
||||
# sibling-plugin name must resolve and the deployed-only name must not.
|
||||
if [[ "$WITH_CLAUDE_OUT" == *"routes to 'deployed-only-skill'"* ]]; then
|
||||
pass "a name that exists only in .claude/ still dangles when an authoring root is present"
|
||||
else
|
||||
fail "the deployed-only target did not dangle — the deployed tree is being read into the universe: $WITH_CLAUDE_OUT"
|
||||
fi
|
||||
if [[ "$WITH_CLAUDE_OUT" != *"routes to 'cross-plugin-skill'"* ]]; then
|
||||
pass "a name in a SIBLING PLUGIN resolves, so the comparison above is not 'nothing resolves'"
|
||||
else
|
||||
fail "the sibling-plugin target dangled — the monorepo universe is not being built: $WITH_CLAUDE_OUT"
|
||||
fi
|
||||
|
||||
# The other half of the rule: with NO authoring root, deployed trees ARE the
|
||||
# universe. That is the consumer case, and without this the rule above could be
|
||||
# implemented as "never read .claude/", which would leave consumers with no
|
||||
# resolution at all.
|
||||
echo ""
|
||||
echo "--- with no authoring root, a deployed .claude/ tree IS the universe ---"
|
||||
CONSUMER="$TMPDIR_T/consumer"
|
||||
mkdir -p "$CONSUMER/.claude/skills/deployed-only-skill"
|
||||
write_skill "$CONSUMER/.claude/skills/my-skill" my-skill \
|
||||
"Use when doing the thing. Do not use for the other thing — use deployed-only-skill instead."
|
||||
set +e
|
||||
CONSUMER_OUT="$(bash "$HOOK" "$CONSUMER/.claude/skills/my-skill/SKILL.md" 2>&1)"
|
||||
CONSUMER_RC=$?
|
||||
set -e
|
||||
if [[ $CONSUMER_RC -eq 0 && "$CONSUMER_OUT" != *"routes to"* && "$CONSUMER_OUT" != *"DID NOT RUN"* ]]; then
|
||||
pass "a sibling in a deployed .claude/skills/ tree resolves when there is no authoring root"
|
||||
else
|
||||
fail "the consumer path did not resolve through the deployed tree (exit $CONSUMER_RC): ${CONSUMER_OUT:-<empty>}"
|
||||
fi
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 1b. Machine independence — the real corpus
|
||||
# ---------------------------------------------------------------------------
|
||||
# The fixture above proves the rule; this proves it at the scale where it broke.
|
||||
#
|
||||
# The A/B is built rather than borrowed. plugins/ is copied TWICE: once bare (a
|
||||
# fresh clone), and once with a synthetic .claude/skills/ tree deployed beside it
|
||||
# holding a directory for every name the corpus currently reports as dangling. If
|
||||
# deployed trees leaked back into the universe, the second copy would resolve
|
||||
# those names and report an empty dangling set while the first reported two —
|
||||
# 2-vs-6, reproduced deterministically.
|
||||
#
|
||||
# Deliberately NOT keyed on whether THIS machine has run `apm install`. Doing that
|
||||
# would make the suite fail on a fresh clone (where there is no .claude/ to
|
||||
# contrast against) — a test of machine independence that is itself
|
||||
# machine-dependent. The live tree is still compared, as a third data point, but
|
||||
# nothing here requires it to be in either state.
|
||||
echo ""
|
||||
echo "--- the real corpus reports the same dangling targets with and without a deployed tree ---"
|
||||
dangling_set() {
|
||||
local -a files=()
|
||||
local f
|
||||
# Collected with a `while read` loop, not `mapfile`: macOS ships bash 3.2,
|
||||
# which has no `mapfile`, and tests/test-vale-wrap.sh scans tests/*.sh for
|
||||
# exactly that hazard. `find` rather than a glob so both roots walk identically.
|
||||
while IFS= read -r f; do
|
||||
files+=("$f")
|
||||
done < <(find "$1" -path '*/.apm/skills/*/SKILL.md' | sort)
|
||||
if [[ ${#files[@]} -eq 0 ]]; then
|
||||
echo "NO-FILES-FOUND"
|
||||
return
|
||||
fi
|
||||
set +e
|
||||
bash "$HOOK" ${files[@]+"${files[@]}"} 2>&1 \
|
||||
| grep -oE "routes to '[^']+'" \
|
||||
| sed "s/routes to '//; s/'//" \
|
||||
| sort -u
|
||||
set -e
|
||||
}
|
||||
LIVE_DANGLING="$(dangling_set "$REPO_ROOT/plugins")"
|
||||
|
||||
FRESH_ROOT="$TMPDIR_T/fresh-clone"
|
||||
mkdir -p "$FRESH_ROOT"
|
||||
cp -R "$REPO_ROOT/plugins" "$FRESH_ROOT/plugins"
|
||||
[[ -f "$REPO_ROOT/apm.yml" ]] && cp "$REPO_ROOT/apm.yml" "$FRESH_ROOT/apm.yml"
|
||||
FRESH_DANGLING="$(dangling_set "$FRESH_ROOT/plugins")"
|
||||
|
||||
DEPLOYED_ROOT="$TMPDIR_T/deployed-clone"
|
||||
mkdir -p "$DEPLOYED_ROOT/.claude/skills" "$DEPLOYED_ROOT/.claude/agents"
|
||||
cp -R "$REPO_ROOT/plugins" "$DEPLOYED_ROOT/plugins"
|
||||
[[ -f "$REPO_ROOT/apm.yml" ]] && cp "$REPO_ROOT/apm.yml" "$DEPLOYED_ROOT/apm.yml"
|
||||
# Deploy exactly the names that currently dangle. That is the strongest possible
|
||||
# bait: if the deployed tree were consulted, every one of them would resolve and
|
||||
# the dangling set would collapse to empty.
|
||||
DEPLOY_COUNT=0
|
||||
while IFS= read -r name; do
|
||||
[[ -n "$name" ]] || continue
|
||||
mkdir -p "$DEPLOYED_ROOT/.claude/skills/$name"
|
||||
DEPLOY_COUNT=$((DEPLOY_COUNT + 1))
|
||||
done <<< "$FRESH_DANGLING"
|
||||
DEPLOYED_DANGLING="$(dangling_set "$DEPLOYED_ROOT/plugins")"
|
||||
|
||||
if [[ "$DEPLOY_COUNT" -gt 0 ]]; then
|
||||
pass "precondition: $DEPLOY_COUNT dangling name(s) deployed into the contrast tree's .claude/skills/, so the A/B has something to distinguish"
|
||||
else
|
||||
fail "no dangling names to deploy — the corpus reports none, so this A/B distinguishes nothing. Deploy a known-absent name explicitly instead of deriving one."
|
||||
fi
|
||||
if [[ ! -d "$FRESH_ROOT/.claude" && ! -d "$FRESH_ROOT/.agents" ]]; then
|
||||
pass "precondition: the fresh-clone copy has no deployed tree of its own"
|
||||
else
|
||||
fail "the fresh-clone copy picked up a deployed tree — it is not a fresh-clone fixture"
|
||||
fi
|
||||
if [[ "$FRESH_DANGLING" == "$DEPLOYED_DANGLING" ]]; then
|
||||
pass "deploying every dangling name into .claude/skills/ changes nothing: $(echo "$FRESH_DANGLING" | tr '\n' ' ')"
|
||||
else
|
||||
fail "the corpus verdict depends on whether apm install has been run — fresh clone: [$(echo "$FRESH_DANGLING" | tr '\n' ' ')] with a deployed tree: [$(echo "$DEPLOYED_DANGLING" | tr '\n' ' ')]"
|
||||
fi
|
||||
# Third data point: whatever state THIS machine happens to be in, the live tree
|
||||
# must agree with a bare copy of the same plugins/. No precondition on that state
|
||||
# — see the section header.
|
||||
if [[ "$LIVE_DANGLING" == "$FRESH_DANGLING" ]]; then
|
||||
pass "the live tree agrees with a bare copy (this machine $( [[ -d "$REPO_ROOT/.claude/skills" ]] && echo "HAS" || echo "has no" ) deployed .claude/skills/ tree)"
|
||||
else
|
||||
fail "the live tree disagrees with a bare copy of the same plugins/ — live: [$(echo "$LIVE_DANGLING" | tr '\n' ' ')] fresh clone: [$(echo "$FRESH_DANGLING" | tr '\n' ' ')]"
|
||||
fi
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 1c. The two live true positives, pinned by name
|
||||
# ---------------------------------------------------------------------------
|
||||
# ADR-0020 records these as real broken routing targets and splits fixing them
|
||||
# into Gitea issue #100. Until that lands they are the ONLY evidence the dangling
|
||||
# check finds anything at all in real prose, so they are asserted as an exact set
|
||||
# rather than a "contains" — a false-positive fix that suppressed one of them
|
||||
# would otherwise land green.
|
||||
#
|
||||
# `gitea-labels` is the subtler of the two and is worth keeping: it is not
|
||||
# written anywhere as `gitea-labels`. gitea-issues' description says "Composes
|
||||
# `gitea-labels-\n milestones`" in a `>`-folded scalar, and the fold joins the
|
||||
# lines into "gitea-labels- milestones" — the trailing hyphen is what keeps the
|
||||
# token terminal and therefore danglable.
|
||||
#
|
||||
# WHEN ISSUE #100 IS FIXED: update EXPECTED_DANGLING to match. Do not delete the
|
||||
# assertion — an empty expected set is fine and still pins that no NEW dangling
|
||||
# target appeared.
|
||||
echo ""
|
||||
echo "--- the two live dangling targets in the corpus are exactly the two ADR-0020 records ---"
|
||||
EXPECTED_DANGLING="$(printf '%s\n' gitea-labels neuledge-context)"
|
||||
if [[ "$LIVE_DANGLING" == "$EXPECTED_DANGLING" ]]; then
|
||||
pass "the corpus dangling set is exactly {gitea-labels, neuledge-context}"
|
||||
else
|
||||
fail "the corpus dangling set changed — expected [$(echo "$EXPECTED_DANGLING" | tr '\n' ' ')], got [$(echo "$LIVE_DANGLING" | tr '\n' ' ')]. If a retrofit fixed one, update EXPECTED_DANGLING; if a false-positive fix silently deleted one, that is the regression this asserts."
|
||||
fi
|
||||
for probe in \
|
||||
"plugins/bin/.apm/skills/research/SKILL.md:neuledge-context" \
|
||||
"plugins/gitea/.apm/skills/gitea-issues/SKILL.md:gitea-labels"; do
|
||||
probe_file="$REPO_ROOT/${probe%%:*}"
|
||||
probe_name="${probe##*:}"
|
||||
if [[ ! -f "$probe_file" ]]; then
|
||||
fail "the true-positive fixture ${probe%%:*} no longer exists — this pin has become vacuous"
|
||||
continue
|
||||
fi
|
||||
set +e
|
||||
probe_out="$(bash "$HOOK" "$probe_file" 2>&1)"
|
||||
set -e
|
||||
if [[ "$probe_out" == *"routes to '$probe_name'"* ]]; then
|
||||
pass "detects the dangling '$probe_name' target in ${probe%%:*}"
|
||||
else
|
||||
fail "did not detect the dangling '$probe_name' target in ${probe%%:*} — a false-positive fix has taken a true positive with it: $probe_out"
|
||||
fi
|
||||
done
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 2. The bare-target grammar rule
|
||||
# ---------------------------------------------------------------------------
|
||||
# Every fixture is built inside a real plugin tree. In a bare temp directory the
|
||||
# resolver would decline ("DID NOT RUN") and every must-not-error case would pass
|
||||
# vacuously, proving nothing about extraction.
|
||||
echo ""
|
||||
echo "--- attributive compound modifiers are prose, not routing targets ---"
|
||||
GRAMMAR_ROOT="$TMPDIR_T/grammar"
|
||||
write_skill "$GRAMMAR_ROOT/plugins/p/.apm/skills/sibling-skill" sibling-skill \
|
||||
"Use when doing the other thing. Do not use for anything else."
|
||||
|
||||
# grammar_case <slug> <expect: silent|errors> <needle> <description>
|
||||
grammar_case() {
|
||||
local slug="$1" mode="$2" needle="$3" desc="$4" out status=0
|
||||
write_skill "$GRAMMAR_ROOT/plugins/p/.apm/skills/$slug" "$slug" "$desc"
|
||||
set +e
|
||||
out="$(bash "$HOOK" "$GRAMMAR_ROOT/plugins/p/.apm/skills/$slug/SKILL.md" 2>&1)"
|
||||
status=$?
|
||||
set -e
|
||||
if [[ "$out" == *"DID NOT RUN"* ]]; then
|
||||
fail "\"$desc\" — the resolver declined, so this case asserts nothing about extraction: $out"
|
||||
return
|
||||
fi
|
||||
case "$mode" in
|
||||
silent)
|
||||
if [[ $status -eq 0 && -z "$out" ]]; then
|
||||
pass "not a dangling target: \"$desc\""
|
||||
else
|
||||
fail "\"$desc\" (exit $status, output: ${out:-<empty>})"
|
||||
fi
|
||||
;;
|
||||
errors)
|
||||
if [[ $status -ne 0 && "$out" == *"$needle"* ]]; then
|
||||
pass "still a dangling target: \"$desc\""
|
||||
else
|
||||
fail "\"$desc\" should have ERRORed with $needle (exit $status, output: ${out:-<empty>})"
|
||||
fi
|
||||
;;
|
||||
suggests)
|
||||
# Reported, not blocking. Both halves matter: an ERROR here would be the
|
||||
# unsuppressable false positive this tier exists to remove, and silence
|
||||
# would mean the gate stopped noticing the target at all.
|
||||
if [[ $status -eq 0 && "$out" == *"SUGGESTION"*"$needle"* && "$out" != *"ERROR"* ]]; then
|
||||
pass "reported but not blocking: \"$desc\""
|
||||
else
|
||||
fail "\"$desc\" should have exited 0 with a SUGGESTION naming $needle (exit $status, output: ${out:-<empty>})"
|
||||
fi
|
||||
;;
|
||||
esac
|
||||
}
|
||||
|
||||
# The four phrasings that were hard dangling FAILs with no suppression. All four
|
||||
# are lifted from real descriptions in this corpus.
|
||||
grammar_case fp-precommit-hooks silent "" \
|
||||
"Use when running the linter. Use pre-commit hooks instead of ad-hoc scripts."
|
||||
grammar_case fp-pull-request silent "" \
|
||||
"Use when opening changes. Invoke the pull-request template instead of writing one by hand."
|
||||
grammar_case fp-conventional silent "" \
|
||||
"Use when writing history. Use conventional-commits formatting rather than free-form messages."
|
||||
grammar_case fp-prepush-backticked silent "" \
|
||||
"Use when checking a branch. Do not use for local edits — run the \`pre-push\` hooks instead."
|
||||
|
||||
echo ""
|
||||
echo "--- a lone unresolvable token in terminal position is REPORTED, not blocking ---"
|
||||
# The class this tier was added for. Every one of these is grammatically
|
||||
# identical to a real broken route — "route verb + name + terminal" is also how
|
||||
# prose cites a hook, a linter, a file format or an English compound — and every
|
||||
# one of them was a hard FAIL with no suppression mechanism anywhere in the gate.
|
||||
# The skills most exposed are exactly the ones the ADR-0020 retrofit sends
|
||||
# authors back to rewrite first: pc-run, pc-author, vale-run, vale-config and the
|
||||
# apm-* family are all ABOUT hyphenated tools.
|
||||
grammar_case fp-precommit-terminal suggests "routes to 'pre-commit'" \
|
||||
"Use when running the linter. Do not use for running hooks — run \`pre-commit\` instead."
|
||||
grammar_case fp-commit-msg suggests "routes to 'commit-msg'" \
|
||||
"Use when writing history. Do not use for the commit message — see \`commit-msg\`."
|
||||
grammar_case fp-type-check suggests "routes to 'type-check'" \
|
||||
"Use when compiling. Do not use for type errors — run \`type-check\` first."
|
||||
grammar_case fp-semantic-release suggests "routes to 'semantic-release'" \
|
||||
"Use when tagging a version. Instead, use \`semantic-release\`."
|
||||
# Single-word tool names are the same defect: `eslint` in terminal position hit
|
||||
# the marked-target path and hard-FAILed just as `pre-commit` did.
|
||||
grammar_case fp-eslint suggests "routes to 'eslint'" \
|
||||
"Use when linting JS. Do not use for style — run \`eslint\` instead."
|
||||
# Bare English compounds, which the backtick path never sees at all.
|
||||
grammar_case fp-clean-up suggests "routes to 'clean-up'" \
|
||||
"Use when doing the thing. Do not use for the old flow — use the clean-up instead."
|
||||
grammar_case fp-built-in suggests "routes to 'built-in'" \
|
||||
"Use when doing the thing. Do not use for the custom path — Instead, prefer the built-in."
|
||||
grammar_case fp-write-up suggests "routes to 'write-up'" \
|
||||
"Use when doing the thing. Do not use for the summary — see the write-up."
|
||||
grammar_case fp-front-end suggests "routes to 'front-end'" \
|
||||
"Use when doing the thing. Do not use for the API layer — use the front-end."
|
||||
grammar_case fp-unit-tests suggests "routes to 'unit-tests'" \
|
||||
"Use when testing. Do not run end-to-end, run unit-tests."
|
||||
|
||||
echo ""
|
||||
echo "--- terminal targets that resolve to nothing still ERROR ---"
|
||||
# The controls. Without them the cases above are satisfied by a check that never
|
||||
# fires, and the narrowing would have eaten the gate rather than sharpened it.
|
||||
# Three forms, three code paths:
|
||||
# * route NOTATION — `/name` and `-> name` — is exempt from corroboration and
|
||||
# blocks on its own. Nobody writes `/pre-commit` or `-> pre-commit` to mean
|
||||
# the hook, so there is no ambiguity to resolve, and an author who wants a
|
||||
# route checked unconditionally has two ways to say so.
|
||||
# * a PROSE-form target — backticked or bare — blocks when its own sentence
|
||||
# names another target that resolves. `sibling-skill` is that corroborator
|
||||
# here; it is the same shape as both live true positives, which sit beside
|
||||
# `write-docs` and `gitea-labels-milestones` respectively.
|
||||
grammar_case tp-arrow errors "routes to 'no-such-arrow-target'" \
|
||||
"Use when doing the thing. Not the other thing → no-such-arrow-target."
|
||||
grammar_case tp-slash errors "routes to 'no-such-slash-skill'" \
|
||||
"Use when doing the thing. Do not use for improvements — use /no-such-slash-skill instead."
|
||||
grammar_case tp-backticked errors "routes to 'no-such-backticked-skill'" \
|
||||
"Use when doing the thing. Do not use for improvements — use \`sibling-skill\` or \`no-such-backticked-skill\` instead."
|
||||
grammar_case tp-bare-terminal errors "routes to 'no-such-bare-skill'" \
|
||||
"Use when doing the thing. Do not use for improvements — use sibling-skill or no-such-bare-skill instead."
|
||||
|
||||
# And the confirming half of the grammar rule: a compound-modifier target is
|
||||
# CONFIRM-ONLY, not ignored. When the name does exist it still counts as a route
|
||||
# — the rule suppresses the ERROR, it does not delete the target.
|
||||
echo ""
|
||||
echo "--- an attributive target that DOES resolve is still a route, not a discarded token ---"
|
||||
write_skill "$GRAMMAR_ROOT/plugins/p/.apm/skills/attributive-subject" attributive-subject \
|
||||
"Use when doing the thing. Do not use for the other thing — use the sibling-skill helper instead."
|
||||
set +e
|
||||
ATTR_OUT="$(bash "$HOOK" "$GRAMMAR_ROOT/plugins/p/.apm/skills/attributive-subject/SKILL.md" 2>&1)"
|
||||
ATTR_RC=$?
|
||||
set -e
|
||||
if [[ $ATTR_RC -eq 0 && -z "$ATTR_OUT" ]]; then
|
||||
pass "a resolving attributive target neither errors nor is reported"
|
||||
else
|
||||
fail "an attributive target naming a REAL skill produced output (exit $ATTR_RC): $ATTR_OUT"
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo "Results: $PASS passed, $FAIL failed"
|
||||
[[ $FAIL -eq 0 ]]
|
||||
@@ -538,13 +538,24 @@ else
|
||||
fi
|
||||
|
||||
# --- 10g. An ambient RUN_TESTS_STRICT=1 must not reach a fixture that did not ask
|
||||
# for it. This suite is discovered and run by run-tests.sh itself, so under the
|
||||
# gate the variable is exported into every child here. That is not a hypothetical:
|
||||
# `bash tests/run-tests.sh` reported 19 passed while
|
||||
# `RUN_TESTS_STRICT=1 bash tests/run-tests.sh` reported this file as the one
|
||||
# failure, because case 10c's deliberately-non-strict run inherited strict and
|
||||
# went red. The gate was therefore RED for everyone, and the only way to see it
|
||||
# was to run the gate.
|
||||
# for it. This suite is discovered and run by run-tests.sh itself, so when the
|
||||
# outer run is launched with the env-var spelling the variable used to be exported
|
||||
# into every child here. That was not a hypothetical: `bash tests/run-tests.sh`
|
||||
# was green while `RUN_TESTS_STRICT=1 bash tests/run-tests.sh` reported this file
|
||||
# as the one failure, because case 10c's deliberately-non-strict run inherited
|
||||
# strict and went red.
|
||||
#
|
||||
# Two corrections to the record, because both were overstated before:
|
||||
# * The blast radius was TWO assertions, not six -- cases 10c and 10g here, and
|
||||
# nothing else in the repo reads RUN_TESTS_STRICT.
|
||||
# * The pre-push GATE was never red. It runs `bash tests/run-tests.sh --strict`
|
||||
# (see .pre-commit-config.yaml), and the flag sets a shell local that is never
|
||||
# exported, so the flag spelling never leaked. Only the env-var spelling did.
|
||||
#
|
||||
# run-tests.sh now `unset`s the variable immediately after latching it, so the
|
||||
# leak is closed at its source and the two spellings hand children an identical
|
||||
# environment (case 10i pins that directly). The `env -u` in run_fake() is kept as
|
||||
# this suite's own defence-in-depth rather than as the fix.
|
||||
#
|
||||
# The variable is exported here rather than passed as a prefix on purpose: a
|
||||
# prefix (`RUN_TESTS_STRICT=1 run_fake ...`) applies to the function call, and
|
||||
@@ -604,6 +615,60 @@ else
|
||||
pass "--strict reports skipped suites on stderr and suppresses the duplicate stdout list"
|
||||
fi
|
||||
|
||||
# --- 10i. The two documented invocations are equivalent in what a CHILD sees ---
|
||||
# `bash tests/run-tests.sh --strict` and `RUN_TESTS_STRICT=1 bash
|
||||
# tests/run-tests.sh` are documented as the same switch, and case 10b already
|
||||
# asserts they produce the same PARENT verdict. That is the weaker half: the two
|
||||
# differed in the ENVIRONMENT they handed every dispatched test-*.sh, because the
|
||||
# flag sets a shell local while the env var stayed exported down the whole process
|
||||
# tree. A dispatched suite could therefore behave differently depending on which
|
||||
# spelling launched the run above it -- which is how case 10c went red under one
|
||||
# invocation and green under the other.
|
||||
#
|
||||
# So this asks the children directly rather than reading the parent's summary. The
|
||||
# case script reports whether RUN_TESTS_STRICT is present in its own environment
|
||||
# AT ALL (`${VAR+set}`, not `${VAR:-}` -- an exported empty value is still a leak),
|
||||
# and both spellings must report it absent. Equivalence is asserted between the two
|
||||
# observations, not just against a hardcoded expectation, so the two cannot drift
|
||||
# apart in some future direction neither case anticipated.
|
||||
echo ""
|
||||
echo "--- --strict and RUN_TESTS_STRICT=1 hand children the same environment ---"
|
||||
DIR10I="$(make_fake_repo)"
|
||||
FIXTURES+=("$DIR10I")
|
||||
install_healthy_bats_runner "$DIR10I"
|
||||
add_case "$DIR10I" test-reports-its-env.sh <<'EOF'
|
||||
#!/usr/bin/env bash
|
||||
if [[ -n "${RUN_TESTS_STRICT+set}" ]]; then
|
||||
echo "CHILD-SAW-STRICT=[${RUN_TESTS_STRICT}]"
|
||||
else
|
||||
echo "CHILD-SAW-STRICT=<absent>"
|
||||
fi
|
||||
EOF
|
||||
# Flag spelling: run_fake scrubs the ambient variable first, so what the child
|
||||
# sees here is purely a function of what run-tests.sh itself exports.
|
||||
run_fake "$DIR10I" --strict
|
||||
FLAG_CHILD="$(echo "$FAKE_OUT" | grep -o 'CHILD-SAW-STRICT=.*' | head -n 1 || true)"
|
||||
FLAG_RC=$FAKE_RC
|
||||
# Env spelling: invoked directly, NOT through run_fake, because run_fake's `env -u`
|
||||
# would strip the very variable under test.
|
||||
I_PRIV="$(mktemp -d)"
|
||||
FIXTURES+=("$I_PRIV")
|
||||
ENV_RC=0
|
||||
ENV_OUT="$(TMPDIR="$I_PRIV" TEST_DIR="$DIR10I/cases" RUN_TESTS_STRICT=1 \
|
||||
bash "$DIR10I/tests/run-tests.sh" 2>&1)" || ENV_RC=$?
|
||||
ENV_CHILD="$(echo "$ENV_OUT" | grep -o 'CHILD-SAW-STRICT=.*' | head -n 1 || true)"
|
||||
if [[ -z "$FLAG_CHILD" || -z "$ENV_CHILD" ]]; then
|
||||
fail "the reporting case script never ran under one of the two invocations (flag: '${FLAG_CHILD:-<none>}', env: '${ENV_CHILD:-<none>}')"
|
||||
elif [[ "$FLAG_CHILD" != "$ENV_CHILD" ]]; then
|
||||
fail "the two documented invocations hand children different environments — flag: $FLAG_CHILD, env: $ENV_CHILD"
|
||||
elif [[ "$ENV_CHILD" != "CHILD-SAW-STRICT=<absent>" ]]; then
|
||||
fail "RUN_TESTS_STRICT is still exported to dispatched suites ($ENV_CHILD) — a suite that itself runs run-tests.sh inherits strictness it never asked for"
|
||||
elif [[ $FLAG_RC -ne 0 || $ENV_RC -ne 0 ]]; then
|
||||
fail "a clean fixture failed under one of the two invocations (flag rc=$FLAG_RC, env rc=$ENV_RC): $FAKE_OUT / $ENV_OUT"
|
||||
else
|
||||
pass "--strict and RUN_TESTS_STRICT=1 both dispatch children with RUN_TESTS_STRICT absent"
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo "Results: $PASS passed, $FAIL failed"
|
||||
[[ $FAIL -eq 0 ]]
|
||||
|
||||
@@ -294,6 +294,61 @@ make_budget_fixture() {
|
||||
echo "$file"
|
||||
}
|
||||
|
||||
# desc_of_length <n> — a description of EXACTLY n characters that carries a
|
||||
# boundary clause and names no routing target.
|
||||
#
|
||||
# ADR-0020's missing-boundary-clause SUGGESTION fires on every description
|
||||
# without one, so a fixture that omits it is never "otherwise clean": a test
|
||||
# asserting silence would be asserting the boundary check's ABSENCE rather than
|
||||
# the length boundary it names. The clause is paid for out of the same budget
|
||||
# being measured (padding arithmetic, not a fixed suffix) so the character count
|
||||
# stays exact. "anything else" is unhyphenated, so no routing target rides along.
|
||||
desc_of_length() {
|
||||
python3 - "$1" <<'PY'
|
||||
import sys
|
||||
n = int(sys.argv[1])
|
||||
prefix = 'Use when doing the thing. Do not use for anything else. '
|
||||
assert n >= len(prefix), 'requested description shorter than the boundary clause'
|
||||
print(prefix + 'x' * (n - len(prefix)))
|
||||
PY
|
||||
}
|
||||
|
||||
# make_tree_fixture <label> <desc> <body_words> — a SKILL.md inside a synthetic
|
||||
# apm plugin monorepo, so the boundary-target resolver has a universe.
|
||||
#
|
||||
# Resolution walks up FROM THE TARGET FILE to an authoring root (the nearest
|
||||
# ancestor holding plugins/*/.apm/{skills,agents}, falling back to .git); it is
|
||||
# never derived from the checker's own location, because deriving it from
|
||||
# ${BASH_SOURCE} leaked this repo's 39-skill universe into every consumer repo
|
||||
# running the hook. A fixture in a bare mktemp -d therefore has NO universe and
|
||||
# correctly reports "DID NOT RUN" — that is not a bug to paper over with a
|
||||
# looser assertion, it is why the fixture has to be a real tree:
|
||||
#
|
||||
# <root>/plugins/subject-plugin/.apm/skills/<label>/SKILL.md <- the subject
|
||||
# <root>/plugins/subject-plugin/.apm/skills/sibling-skill/ <- same package
|
||||
# <root>/plugins/subject-plugin/.apm/agents/sibling-agent.agent.md
|
||||
# <root>/plugins/other-plugin/.apm/skills/cross-plugin-skill/ <- sibling plugin
|
||||
#
|
||||
# The sibling plugin is what makes "every plugin in the monorepo contributes its
|
||||
# names" testable; without it a cross-plugin target and a typo are the same.
|
||||
make_tree_fixture() {
|
||||
local label="$1" desc="$2" body_words="$3" root apm
|
||||
root="$TMPDIR/tree-$label"
|
||||
apm="$root/plugins/subject-plugin/.apm"
|
||||
mkdir -p "$apm/skills/$label" "$apm/skills/sibling-skill" "$apm/agents" \
|
||||
"$root/plugins/other-plugin/.apm/skills/cross-plugin-skill"
|
||||
: > "$apm/agents/sibling-agent.agent.md"
|
||||
{
|
||||
echo "---"
|
||||
echo "name: $label"
|
||||
echo "description: $desc"
|
||||
echo "---"
|
||||
echo ""
|
||||
python3 -c "print(' '.join(['word'] * $body_words))"
|
||||
} > "$apm/skills/$label/SKILL.md"
|
||||
echo "$apm/skills/$label/SKILL.md"
|
||||
}
|
||||
|
||||
# expect_gate <label> <expected: pass|suggest|fail> <file> [needle]
|
||||
expect_gate() {
|
||||
local label="$1" expected="$2" file="$3" needle="${4:-}" out status
|
||||
@@ -323,15 +378,26 @@ expect_gate() {
|
||||
fail "$label (exit $status, output: ${out:-<empty>})"
|
||||
fi
|
||||
;;
|
||||
# A check that DECLINED to run must say so and must not fail the file. The
|
||||
# ERROR guard is the point: a declined check that also errored would satisfy
|
||||
# a bare "output contains INFO" assertion.
|
||||
info)
|
||||
if [[ $status -eq 0 && "$out" == *"INFO"* && "$out" == *"$needle"* \
|
||||
&& "$out" != *"ERROR"* ]]; then
|
||||
pass "$label"
|
||||
else
|
||||
fail "$label (exit $status, output: ${out:-<empty>})"
|
||||
fi
|
||||
;;
|
||||
esac
|
||||
}
|
||||
|
||||
echo ""
|
||||
echo "--- description budget: $DESC_SUGGEST_CHARS SUGGESTION / $DESC_MAX_CHARS FAIL, both inclusive ---"
|
||||
D_AT_SUGGEST="$(python3 -c "print('x' * $DESC_SUGGEST_CHARS)")"
|
||||
D_OVER_SUGGEST="$(python3 -c "print('x' * $((DESC_SUGGEST_CHARS + 1)))")"
|
||||
D_AT_MAX="$(python3 -c "print('x' * $DESC_MAX_CHARS)")"
|
||||
D_OVER_MAX="$(python3 -c "print('x' * $((DESC_MAX_CHARS + 1)))")"
|
||||
D_AT_SUGGEST="$(desc_of_length "$DESC_SUGGEST_CHARS")"
|
||||
D_OVER_SUGGEST="$(desc_of_length "$((DESC_SUGGEST_CHARS + 1))")"
|
||||
D_AT_MAX="$(desc_of_length "$DESC_MAX_CHARS")"
|
||||
D_OVER_MAX="$(desc_of_length "$((DESC_MAX_CHARS + 1))")"
|
||||
expect_gate "description at exactly $DESC_SUGGEST_CHARS chars is silent" \
|
||||
pass "$(make_budget_fixture desc-at-suggest "$D_AT_SUGGEST" 10)"
|
||||
expect_gate "description at $((DESC_SUGGEST_CHARS + 1)) chars suggests and exits 0" \
|
||||
@@ -359,18 +425,23 @@ FOLDED="$TMPDIR/folded.md"
|
||||
expect_gate "a >-folded 450-char description fails (raw first line would read as 1 char)" \
|
||||
fail "$FOLDED" "description is 450 characters"
|
||||
|
||||
# Every body fixture below carries a boundary clause for the same reason
|
||||
# desc_of_length() does: without one the missing-boundary-clause SUGGESTION
|
||||
# fires and a body-budget test that asserts silence stops isolating the body
|
||||
# budget. It is short, so the description gate stays quiet too.
|
||||
CLEAN_DESC="Short valid description. Do not use for anything else."
|
||||
echo ""
|
||||
echo "--- body budget: $BODY_SUGGEST_WORDS SUGGESTION / $BODY_MAX_WORDS FAIL, body only, both inclusive ---"
|
||||
expect_gate "body at exactly $BODY_SUGGEST_WORDS words is silent" \
|
||||
pass "$(make_budget_fixture body-at-suggest "Short valid description." "$BODY_SUGGEST_WORDS")"
|
||||
pass "$(make_budget_fixture body-at-suggest "$CLEAN_DESC" "$BODY_SUGGEST_WORDS")"
|
||||
expect_gate "body at $((BODY_SUGGEST_WORDS + 1)) words suggests and exits 0" \
|
||||
suggest "$(make_budget_fixture body-over-suggest "Short valid description." "$((BODY_SUGGEST_WORDS + 1))")" \
|
||||
suggest "$(make_budget_fixture body-over-suggest "$CLEAN_DESC" "$((BODY_SUGGEST_WORDS + 1))")" \
|
||||
"body is $((BODY_SUGGEST_WORDS + 1)) words"
|
||||
expect_gate "body at exactly $BODY_MAX_WORDS words suggests, does not fail" \
|
||||
suggest "$(make_budget_fixture body-at-max "Short valid description." "$BODY_MAX_WORDS")" \
|
||||
suggest "$(make_budget_fixture body-at-max "$CLEAN_DESC" "$BODY_MAX_WORDS")" \
|
||||
"body is $BODY_MAX_WORDS words"
|
||||
expect_gate "body at $((BODY_MAX_WORDS + 1)) words fails" \
|
||||
fail "$(make_budget_fixture body-over-max "Short valid description." "$((BODY_MAX_WORDS + 1))")" \
|
||||
fail "$(make_budget_fixture body-over-max "$CLEAN_DESC" "$((BODY_MAX_WORDS + 1))")" \
|
||||
"$BODY_MAX_WORDS-word ceiling"
|
||||
|
||||
# The two word gates measure different things and must stay separable: a file
|
||||
@@ -382,7 +453,7 @@ BODY_ONLY_DESC="$(python3 -c "print(' '.join(['w'] * 100))")"
|
||||
expect_gate "frontmatter words do not count toward the $BODY_MAX_WORDS-word body ceiling" \
|
||||
suggest "$(make_budget_fixture body-independent "$BODY_ONLY_DESC" "$((BODY_MAX_WORDS - 5))")" \
|
||||
"words"
|
||||
BIG_BODY="$(make_budget_fixture body-over-not-whole-file "Short valid description." "$((BODY_MAX_WORDS + 1))")"
|
||||
BIG_BODY="$(make_budget_fixture body-over-not-whole-file "$CLEAN_DESC" "$((BODY_MAX_WORDS + 1))")"
|
||||
BIG_BODY_WORDS="$(wc -w < "$BIG_BODY")"
|
||||
if [[ "$BIG_BODY_WORDS" -le "$MAX_WORDS" ]]; then
|
||||
pass "the body-gate fixture is $BIG_BODY_WORDS whole-file words, well under MAX_WORDS=$MAX_WORDS — it fails on the body gate alone"
|
||||
@@ -393,21 +464,37 @@ fi
|
||||
echo ""
|
||||
echo "--- resolvable boundary targets ---"
|
||||
# Resolution is against the AUTHORING SOURCE (plugins/*/.apm/skills/ and
|
||||
# plugins/*/.apm/agents/), found here via the script's own repo root — these
|
||||
# fixtures live in a temp dir with no plugin tree of their own, so a resolving
|
||||
# target proves the repo-root path works.
|
||||
expect_gate "a boundary target naming a real skill resolves" \
|
||||
pass "$(make_budget_fixture target-ok \
|
||||
"Use when doing the thing. Do not use for commits — use git-commits instead." 10)"
|
||||
expect_gate "a boundary target naming a real AGENT resolves (agents are valid targets)" \
|
||||
pass "$(make_budget_fixture target-agent-ok \
|
||||
"Use when doing the thing. Do not use when the caller is an agent — invoke git-orchestrate instead." 10)"
|
||||
expect_gate "a boundary target that resolves to nothing fails" \
|
||||
fail "$(make_budget_fixture target-missing \
|
||||
"Use when doing the thing. Do not use for improvements — use no-such-skill-anywhere instead." 10)" \
|
||||
# plugins/*/.apm/agents/), reached by walking up FROM THE SKILL FILE. These
|
||||
# fixtures therefore build their own synthetic monorepo (make_tree_fixture) and
|
||||
# name only fixture-local targets: they must not depend on this repo's live
|
||||
# skills, or renaming git-commits would break a test about extraction grammar.
|
||||
expect_gate "a boundary target naming a sibling skill in the same package resolves" \
|
||||
pass "$(make_tree_fixture target-ok \
|
||||
"Use when doing the thing. Do not use for commits — use sibling-skill instead." 10)"
|
||||
expect_gate "a boundary target naming a skill in a SIBLING PLUGIN resolves (that is what a monorepo means)" \
|
||||
pass "$(make_tree_fixture target-cross-plugin \
|
||||
"Use when doing the thing. Do not use for the other thing — use cross-plugin-skill instead." 10)"
|
||||
expect_gate "a boundary target naming an AGENT resolves (agents are valid targets)" \
|
||||
pass "$(make_tree_fixture target-agent-ok \
|
||||
"Use when doing the thing. Do not use when the caller is an agent — invoke sibling-agent instead." 10)"
|
||||
# CORROBORATED: `sibling-skill` resolves in the same sentence, which is what
|
||||
# promotes a prose-form target from "reported" to "blocking". A lone prose-form
|
||||
# target is deliberately not fatal — see the case below and the shared resolver's
|
||||
# CORROBORATION note.
|
||||
expect_gate "a boundary target that resolves to nothing fails when its sentence names one that does" \
|
||||
fail "$(make_tree_fixture target-missing \
|
||||
"Use when doing the thing. Do not use for improvements — use sibling-skill or no-such-skill-anywhere instead." 10)" \
|
||||
"routes to 'no-such-skill-anywhere'"
|
||||
# UNCORROBORATED: identical grammar to the case above, and identical grammar to
|
||||
# "run `pre-commit` instead". Reported at SUGGESTION tier, exit 0 — a gate that
|
||||
# ships hot with no baseline and no suppression mechanism must not block a commit
|
||||
# on a token it cannot tell from a tool name.
|
||||
expect_gate "a lone boundary target that resolves to nothing is reported, not fatal" \
|
||||
suggest "$(make_tree_fixture target-missing-lone \
|
||||
"Use when doing the thing. Do not use for improvements — use no-such-lone-skill instead." 10)" \
|
||||
"routes to 'no-such-lone-skill'"
|
||||
expect_gate "a /slash-command boundary target that resolves to nothing fails" \
|
||||
fail "$(make_budget_fixture target-missing-slash \
|
||||
fail "$(make_tree_fixture target-missing-slash \
|
||||
"Use when doing the thing. Do not use for improvements — use /no-such-slash-skill instead." 10)" \
|
||||
"routes to 'no-such-slash-skill'"
|
||||
# False-positive guards. These phrasings are lifted from real descriptions:
|
||||
@@ -415,16 +502,38 @@ expect_gate "a /slash-command boundary target that resolves to nothing fails" \
|
||||
# gitea-files says "(use Read/Write/Edit)", gitea-labels-milestones says
|
||||
# "through `issue_write`/`pull_request_write`". None of them is a routing
|
||||
# target, and reading any of them as one makes the gate untrustworthy.
|
||||
#
|
||||
# Each carries a boundary clause in a SEPARATE sentence. That is not decoration:
|
||||
# target extraction is decided per sentence, so the clause satisfies the
|
||||
# missing-boundary-clause SUGGESTION (keeping the expected output empty) while
|
||||
# leaving the sentence under test outside a boundary context, which is the exact
|
||||
# condition each of these is about. They are built as trees so a universe exists
|
||||
# — in a bare temp dir the resolver would decline and the guard would pass
|
||||
# vacuously, proving nothing about extraction.
|
||||
expect_gate "'run pre-commit hooks' outside a boundary sentence is not a routing target" \
|
||||
pass "$(make_budget_fixture fp-precommit \
|
||||
"Use when the user wants to run pre-commit hooks or install git hooks." 10)"
|
||||
pass "$(make_tree_fixture fp-precommit \
|
||||
"Use when the user wants to run pre-commit hooks or install git hooks. Do not use for anything else." 10)"
|
||||
expect_gate "an arrow chain outside a boundary clause is not a routing target" \
|
||||
pass "$(make_budget_fixture fp-arrow \
|
||||
"Reproduce → minimise → instrument → fix → regression-test. Use when a bug is reported." 10)"
|
||||
pass "$(make_tree_fixture fp-arrow \
|
||||
"Reproduce → minimise → instrument → fix → regression-test. Use when a bug is reported. Do not use for anything else." 10)"
|
||||
expect_gate "tool names and MCP tool names are not routing targets" \
|
||||
pass "$(make_budget_fixture fp-tools \
|
||||
pass "$(make_tree_fixture fp-tools \
|
||||
"Use when writing issues. Do not use for local files (use Read/Write/Edit) — that write goes through \`issue_write\`/\`pull_request_write\` instead." 10)"
|
||||
|
||||
echo ""
|
||||
echo "--- with NO authoring root the resolver declines OUT LOUD and does not fail the file ---"
|
||||
# The consumer/draft case, and a real one: a SKILL.md in a bare directory with no
|
||||
# plugins/*/.apm/ above it and no .git has no universe to resolve against. The
|
||||
# required behaviour is neither a false FAIL nor silence — silence is how a whole
|
||||
# gate family goes missing unnoticed — so the INFO and the named unchecked target
|
||||
# are both asserted, along with exit 0. This is the same path make_tree_fixture
|
||||
# exists to escape, kept pinned so a future "just use the repo root" shortcut
|
||||
# (the ${BASH_SOURCE} universe leak ADR-0020 removed) fails here.
|
||||
expect_gate "a fixture with no authoring root reports DID NOT RUN and exits 0" \
|
||||
info "$(make_budget_fixture no-universe \
|
||||
"Use when doing the thing. Do not use for improvements — use some-other-skill instead." 10)" \
|
||||
"Unchecked target(s): some-other-skill"
|
||||
|
||||
echo ""
|
||||
echo "--- the three live dangling routing targets are caught (issue #100) ---"
|
||||
# ADR-0020 records four broken routing targets and splits fixing them into its
|
||||
|
||||
Reference in New Issue
Block a user