refactor(skills): retrofit the corpus to the ADR-0020 context contract (#129)
Retrofits all 39 skills to ADR-0020's description/body context contract, then fixes what six rounds of independent review found in that retrofit — including four ways the hot gate itself failed open. Closes #99, #107, #108, #110, #111, #114, #115, #120. ## The retrofit (waves 1-5) | | Start | Now | |---|---|---| | Description FAILs (>400 chars) | 26 | **0** | | Body FAILs (>900 words, body-only) | 9 | **0** | | Dangling routing targets | 2 | **0** | | `Kyberforge.CompositionNote` | 10 | **0** | | Preload tax | 21,005 chars | **~10,500** | Under the 12,000-char success criterion. Per-wave detail is on #99. ## The review fixes **The gate failed open four ways, three of them found after the retrofit shipped.** An unrecognised follower token made a dangling target vanish. A skill directory with no `SKILL.md` resolved as a valid target, so a commit could be green locally and red in a fresh clone — three existing fixtures were relying on that, one of which made the install-leak A/B pass vacuously. Then the free-standing `/name` sweep turned out to be gated on the sentence carrying a boundary marker, so route notation in any other sentence was invisible — not an ERROR, not a SUGGESTION, not an INFO — which left the documented "`/name` always blocks" promise false from a second direction. All four fixed and pinned. **Two checks were silently not running.** `validate-provenance.sh` checks 7-8 were dead across nine skills. Waking them exposed a deeper problem: they assume `Research doc:` names a source index, but 30 of 121 entries point at topic content documents, so every new check-7 INFO was a false positive and check 8 was saved from a false-FAIL flood only by an *unannounced* skip. Checks 7/8 are now scoped to source indexes and every skip announces itself (#121). **The retrofit's own anti-goal, four times.** ADR-0020 warns that a blunt gate gets satisfied by deleting content rather than relocating it. `diagnose` and `skill-audit` relocated prose and then read it unconditionally; `prototype` and `vale-config` deleted rules outright that survived nowhere. All four addressed. ## Verification - `bash tests/run-tests.sh --strict` — 24 suites, 0 skipped, 0 failed - `bash tests/run-bats.sh` — 325 tests, 0 failures - `pre-commit run --all-files` — 17/17 - `pre-commit run --hook-stage pre-push --all-files` — 16/16, with `apm marketplace check` and `apm pack --check-clean` run against the remote, not skipped - `scripts/skill-size-check.sh` over all 39 skills — rc 0, 0 ERROR/FAIL, SUGGESTION-only - Preload tax measured at **10,498 chars**, max description 390 — both inside budget - Every new test proven non-vacuous by a deliberate mutation of the behaviour it covers **Per-commit sync, stated accurately:** the ten commits from the latest review round each pass `check-plugin-content-sync` in isolation, verified by checking each out in a detached worktree with a clean between. The earlier gitea window (`dfacf05..bedbd1d`, nine commits) does **not** — its mirror was regenerated in one batch at `bbc7300`. An earlier revision of this description claimed the property held for every commit; it does not, and a bisect through that window lands on a red commit. **Squash-merge** to collapse it, or accept that this range is not bisectable. ## Version bump Six plugins and the catalog take a **patch**, not a minor. The branch is **89 commits — 40 `fix` / 30 `refactor` / 12 `docs` / 5 `chore` / 2 `test` — zero `feat`, zero `!`, zero `BREAKING CHANGE`** — and adds no skill, agent, command or hook. (Two earlier revisions of this section cited a stale histogram, most recently 78 commits; the figures above are measured at HEAD.) Both rules this repo ships (`forge/references/version-bump.md`, landing in this PR, and `git-commits/references/conventional-commits-spec.md`) make that a patch, and the catalog set is unchanged at 7 entries. Not settled by that: four published files were removed from the installed tree, three moved, and `caveman` gained `disable-model-invocation`, retiring its old triggers. Under a strict reading those are major-class and currently ship under `refactor:` with no marker. Whether the deployed skill surface is a public contract is written down nowhere — worth deciding, but it outlives this PR. ## Deliberately not in scope #112 (cherry-pick ownership, now resolved in favour of `git-commits`), #113 (`rtk git` normalisation), #116 (research fan-out), #101 (audit-skill merge), #122 (non-spec skill-root files), #123 (no PRD producer) stay open. #117 is the one worth reading: the contract's remedy is to move prose into `references/`, which is exactly where neither the size gate nor Vale looks — and the blind spot is wider than #117 currently records, since there is no root `.vale.ini` at all, so every ADR, `CONTEXT.md` and `README.md` is unlinted too. That blind spot let this branch carry two `level: error` `Kyberforge.SentenceOpenerThereIs` violations into `references/` files it created — `provider-adapter-author/references/provider-matrix.md:31` and `agent-audit/references/finding-criteria.md:95`. Both are reworded in `afadaae`, confirmed by routing each file through the audit's own `vale-wrap.sh` (1 error each before, 0 after). Five further occurrences sit in `references/` files already on `main`; those are the pre-existing corpus and stay with #117, which is the real fix. Also unfixed and not this PR's: `apm install` appends a duplicate `SessionStart` entry to `.claude/settings.json`, so a fresh clone cannot get pre-push green without an edit AGENTS.md warns against. Reproduces identically on `main`. Co-authored-by: Defame1297 <gitea@rkdr.net> Reviewed-on: https://git.dev.rkdr.net/Defame1297/holocron/pulls/129 Co-authored-by: Claude Code AI - Gitea MCP <claude@noreply.git.dev.rkdr.net> Co-committed-by: Claude Code AI - Gitea MCP <claude@noreply.git.dev.rkdr.net>
This commit was merged in pull request #129.
This commit is contained in:
@@ -57,8 +57,9 @@ Pass the path to either agent file as the argument.
|
||||
| `assets/vale/styles/Kyberforge/VagueWording.yml` | Flags vague capability wording ("helps with", "utilize", "assists with", "used for") in descriptions |
|
||||
| `assets/vale/styles/KyberforgeCopilot/ProactivePhrase.yml` | Flags CC-specific "Use proactively" phrasing with no effect in Copilot descriptions |
|
||||
| `references/README.md` | Directory documentation for references/ |
|
||||
| `references/description-quality.md` | Rubric for the description dimension — three-part shape, the 250/400-character budget, the hand-invoked contract, and the internal-mechanics FAIL |
|
||||
| `references/body-and-delegation.md` | Rubric for the body, delegation and comment-discipline dimensions — the delegation FAIL and why agents take no body word gate |
|
||||
| `references/finding-criteria.md` | Every dimension's FAIL and SUGGESTION criteria — the one Step 3 file read on every run; it decides which rubrics below are worth loading |
|
||||
| `references/description-quality.md` | Rubric for the description dimension — why the description is the expensive part, the hand-invoked contract, the three-part shape, indirect triggers, and near-miss exclusions |
|
||||
| `references/body-and-delegation.md` | Rubric for the body, delegation and comment-discipline dimensions — the core test, the delegation FAIL, why agents take no body word gate, and what an agent body is for |
|
||||
| `references/scope-plugin-apm.md` | Scope contract for a single vendor-neutral APM agent file — allowlist, dimension routing, and the dimensions that do not apply |
|
||||
| `references/scope-project-user.md` | Scope contract for a CC / Copilot pair — counterpart derivation, provider field rules, pair consistency |
|
||||
| `references/validation-scripts.md` | Loaded only when a Step 1 script fails or cannot run — scope-detection walk-up, manual fallback checks, known script failures |
|
||||
|
||||
@@ -20,7 +20,7 @@ metadata:
|
||||
|
||||
- Do not narrate PASS/FAIL per check while auditing. Gather findings internally and surface them only in the Step 4 report. Narrating each check as you go is the default failure mode here.
|
||||
- Agents take the same 250/400-character description gates as skills and **no body word gate at all** — an agent body becomes the system prompt of a fresh context, so the 900-word skill ceiling does not transfer. Judge an over-long agent body through the delegation check, never by word count.
|
||||
- At plugin/APM scope the agent is a single vendor-neutral file by design: never raise a pair-consistency finding there, and provider safety stops meaning Claude-Code-versus-Copilot field leakage.
|
||||
- At plugin/APM scope the agent is a single vendor-neutral file by design, so provider safety stops meaning Claude-Code-versus-Copilot field leakage there.
|
||||
- Vale reporting `0 files` scanned means NOT RUN, not clean. Fall back to full Step 3 judgment for every dimension it would have covered.
|
||||
|
||||
## Step 1 — Deterministic checks
|
||||
@@ -30,14 +30,14 @@ Resolve all three paths against this skill's own directory so they work from a r
|
||||
```bash
|
||||
bash scripts/validate.sh <agent-file>
|
||||
bash scripts/validate-provenance.sh <agent-file>
|
||||
scripts/vale-wrap.sh <agent-file> [<counterpart-file>]
|
||||
bash scripts/vale-wrap.sh <agent-file> [<counterpart-file>]
|
||||
```
|
||||
|
||||
`validate.sh` takes either half of a project/user-scope pair or the single plugin/APM-scope file, detects the provider from the extension and the scope by walking up, then checks required fields, kebab-case `name`, `FILL IN:` placeholders, template HTML comments left in frontmatter, the ADR-0020 description budget (250 chars SUGGESTION, 400 FAIL, measured on the folded YAML value) and the fields that scope permits. Its findings become the `### Structure` dimension — its FAILs and its SUGGESTIONs both — except the ones the Step 2 scope contract re-routes.
|
||||
|
||||
If a validation script fails or cannot run — Bash denied, `python3` or `vale` absent, `references/field-inventory.md` missing — read `references/validation-scripts.md`; what these scripts measure is not reproducible by reading.
|
||||
|
||||
`validate-provenance.sh` prints nothing on success and runs at plugin/APM scope only, exiting 0 silently elsewhere. Its FAIL findings become a separate `### Provenance` dimension, and it emits Why and Fix itself — surface those verbatim.
|
||||
`validate-provenance.sh` prints nothing on success, so read its exit code before you read its silence. **0** is a genuine pass, including the silent exit 0 at project or user scope, where plugin-scope provenance does not apply. **1** means real findings: its FAILs and INFOs become a separate `### Provenance` dimension, and it emits Why and Fix itself — surface those verbatim. **2** means the check never ran — a bad argument or a missing dependency, reason on stderr, no findings and often no stdout at all. On a 2, report `### Provenance` as unverified and quote the stderr reason; never grade it as a clean pass. `validate.sh` uses the same 2 tier.
|
||||
|
||||
`vale-wrap.sh` applies the bundled `Kyberforge` style as a prefilter. Pass no `--config`; the wrapper locates its own. At project/user scope pass both files of the pair, not only the one you were handed. Every rule is graded `error`, so every alert is a FAIL. Report each one citing its rule ID, filed under the dimension it belongs to, and do not re-derive it by judgment:
|
||||
|
||||
@@ -57,14 +57,14 @@ Read the agent file end to end, and at project/user scope its counterpart too. A
|
||||
|
||||
## Step 3 — Qualitative audit
|
||||
|
||||
Load a dimension's rubric before judging that dimension.
|
||||
Read `references/finding-criteria.md` first — every dimension's FAIL and SUGGESTION criteria. Load the rubric below only for a dimension the criteria put in play: one carrying a candidate finding, or one where the criterion alone does not settle the call.
|
||||
|
||||
| Dimension | Read |
|
||||
| Dimension | Rubric |
|
||||
|---|---|
|
||||
| description | `references/description-quality.md` |
|
||||
| body, delegation, comment-discipline | `references/body-and-delegation.md` |
|
||||
|
||||
Cite file and line number for every finding.
|
||||
Each rubric is the reasoning behind its criteria, not a second copy of them. Cite file and line number for every finding.
|
||||
|
||||
## Step 4 — Report
|
||||
|
||||
|
||||
@@ -10,8 +10,9 @@ Additional documentation agents load on demand.
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `description-quality.md` | Rubric for the description dimension — three-part shape, the 250/400-character budget, the hand-invoked contract, and the internal-mechanics FAIL. |
|
||||
| `body-and-delegation.md` | Rubric for the body, delegation and comment-discipline dimensions — the delegation FAIL, why agents take no body word gate, and what an agent body is for. |
|
||||
| `finding-criteria.md` | Every dimension's FAIL and SUGGESTION criteria — the one Step 3 file read on every run; it decides which rubrics below are worth loading. |
|
||||
| `description-quality.md` | Rubric for the description dimension — why the description is the expensive part, the hand-invoked contract, the three-part shape, indirect triggers, and near-miss exclusions. |
|
||||
| `body-and-delegation.md` | Rubric for the body, delegation and comment-discipline dimensions — the core test, the delegation FAIL, why agents take no body word gate, and what an agent body is for. |
|
||||
| `scope-plugin-apm.md` | Contract for a single vendor-neutral `.apm/agents/<name>.agent.md` file — allowlist, dimension routing, and the dimensions that do not apply. |
|
||||
| `scope-project-user.md` | Contract for a Claude Code / Copilot file pair — counterpart derivation, provider field rules, and pair consistency. |
|
||||
| `validation-scripts.md` | Loaded only when a Step 1 script fails or cannot run — scope-detection walk-up, manual fallback checks, and known script failures. |
|
||||
|
||||
@@ -98,24 +98,8 @@ to a shipped file. At plugin/APM scope the stakes are higher than tidiness: `apm
|
||||
frontmatter verbatim to every target, `<!-- ... -->` is not valid YAML, and `validate.sh` FAILs a
|
||||
frontmatter block that still contains one.
|
||||
|
||||
## Auditing guidance
|
||||
## Where the criteria live
|
||||
|
||||
Flag as FAIL if:
|
||||
|
||||
- The body restates a procedure owned by a skill the agent can invoke — Fix: invoke `<skill>`
|
||||
instead
|
||||
- A sentence answers "no" to the core test — it is padding
|
||||
- A decision point presents a menu of options with no default
|
||||
- An instruction repeats content already in the description
|
||||
- Frontmatter comments are template scaffolding rather than instruction, or are HTML comments at
|
||||
plugin/APM scope
|
||||
- A prescriptive sequence is used where flexibility is fine, or the reverse
|
||||
|
||||
Flag as SUGGESTION if:
|
||||
|
||||
- The body does not open with a direct role instruction
|
||||
- The body specifies no error handling — nothing tells the agent what to do with malformed,
|
||||
missing or contradictory input
|
||||
- The job the agent describes is unbounded, or bounded only implicitly
|
||||
- A rationale is missing from a rule the agent is expected to enforce — present but unexplained
|
||||
- Comments are useful but verbose enough to bury the field they annotate
|
||||
Every FAIL and SUGGESTION criterion for these dimensions is in `references/finding-criteria.md`,
|
||||
which Step 3 reads on every run. This file is the reasoning behind them, loaded only when that file
|
||||
puts the body, delegation or comment-discipline dimension in play.
|
||||
|
||||
@@ -82,49 +82,8 @@ description: >
|
||||
dispatched and safety-gated. Not conversational git help -> git-workflow.
|
||||
```
|
||||
|
||||
## Auditing guidance
|
||||
## Where the criteria live
|
||||
|
||||
Flag as FAIL if:
|
||||
|
||||
- **Over 400 characters.** Measured on the folded YAML value, not the raw source lines.
|
||||
`validate.sh` reports the number; do not re-derive it, but do point the Fix at what to cut. Agent
|
||||
descriptions have no platform-documented ceiling of their own — unlike a skill's 1,024-character
|
||||
spec limit, the 400-character house ceiling is the only hard limit there is, so do not go looking
|
||||
for a backstop behind it.
|
||||
- **Internal mechanics appear in the description.** Any of:
|
||||
- capability enumeration or a feature list;
|
||||
- output-format detail ("Produces a compact findings report with Why and Fix per finding");
|
||||
- composition or architecture notes ("composes X rather than duplicating Y", "a cross-cutting
|
||||
shared agent", "the human-facing entry point", "replaces the old flat invocation");
|
||||
- implementation detail ("self-validates via a bundled deterministic script").
|
||||
|
||||
None of it can change a routing decision and all of it is preloaded.
|
||||
`Kyberforge.CompositionNote` catches the common phrasings deterministically; the rest is
|
||||
judgment. This is the rule that deflates a description, so apply it before reaching for length.
|
||||
- **The same trigger stated twice in two registers** — a verb list, then the same verbs re-quoted
|
||||
as user phrasings, usually in the same order. One register, whichever routes better.
|
||||
- **Descriptive rather than imperative phrasing** (`This agent ...`, `This is the ...`).
|
||||
`Kyberforge.DescriptionOpener` catches any opener matching `^This`. There is no action-verb rule
|
||||
here and never was a defensible one: an `Orchestrates ...` or `Audits ...` opener is a catalogue
|
||||
entry, not a trigger.
|
||||
- **Vague capabilities** ("helps with agents" where "audits an agent definition pair" was
|
||||
available). `Kyberforge.VagueWording` catches the known filler; imprecision outside that list is
|
||||
judgment.
|
||||
- **A boundary clause naming a target that does not resolve** to a real skill directory or agent
|
||||
file in the authoring source. `validate.sh` resolves this for agent files at both scopes and
|
||||
reports each unresolved target itself — take its verdict rather than re-resolving the name by
|
||||
hand, because a hand-walk over a different universe can contradict it. What is left to you is
|
||||
semantic and the script cannot reach it: whether a target that *does* resolve is the right
|
||||
sibling to exclude, and whether a clause naming no target at all ("examine the files manually")
|
||||
should have named one.
|
||||
- **`Use proactively` in a Copilot or vendor-neutral description.**
|
||||
`KyberforgeCopilot.ProactivePhrase` catches it. The phrase steers the Claude Code runtime and
|
||||
does nothing anywhere else, so in a `.agent.md` it is preloaded text that buys no behaviour.
|
||||
- **Trigger-list, boundary or indirect-trigger content on a hand-invoked agent** — see Step 0.
|
||||
|
||||
Flag as SUGGESTION if:
|
||||
|
||||
- **Over 250 characters** but at or under 400. This tier is what moves the corpus average; the FAIL
|
||||
tier only stops outliers. Report it rather than treating a 399-character description as clean.
|
||||
- A near-miss exclusion is present but targets a weak near-miss.
|
||||
- An indirect trigger is present and warranted but could name the omitted phrasing more precisely.
|
||||
Every FAIL and SUGGESTION criterion for this dimension is in `references/finding-criteria.md`,
|
||||
which Step 3 reads on every run. This file is the reasoning behind them, loaded only when that file
|
||||
puts the description dimension in play.
|
||||
|
||||
@@ -38,9 +38,9 @@ whose vocabulary differs per harness — Claude Code names its own tools, Copilo
|
||||
other, and `apm compile` copies frontmatter verbatim with no per-target integrator to reconcile
|
||||
them. `disallowedTools` is a **denylist**, and denying by name is safe under verbatim copy: a name
|
||||
the other harness does not recognise denies nothing, so the worst case is that the fence is absent
|
||||
there, never that the wrong capability is granted. Claude Code honours it for plugin subagents —
|
||||
`docs/research/docs/claude-code-plugins/agent-definition.md:99` names the fields plugin agents
|
||||
silently ignore (`hooks`, `mcpServers`, `permissionMode`) and `disallowedTools` is not among them.
|
||||
there, never that the wrong capability is granted. Claude Code honours it for plugin subagents: its
|
||||
plugin agent-definition reference names the fields plugin agents silently ignore (`hooks`,
|
||||
`mcpServers`, `permissionMode`), and `disallowedTools` is not among them.
|
||||
|
||||
`disallowedTools` also appears in `claude-code-only-fields` above, and that stays correct: at
|
||||
project/user scope it is still a Claude-only field and must not appear in a Copilot `.agent.md`.
|
||||
|
||||
@@ -0,0 +1,98 @@
|
||||
---
|
||||
source_keys:
|
||||
- context7-websites-code-claude
|
||||
- claude-code-plugins-docs
|
||||
- claude-code-subagents-docs
|
||||
- context7-github-en-copilot
|
||||
- github-custom-agents-configuration
|
||||
---
|
||||
|
||||
# Finding Criteria
|
||||
|
||||
Every FAIL and SUGGESTION criterion, for every qualitative dimension, and nothing else. The
|
||||
reasoning each criterion stands on, its worked examples and its house rules stay in that
|
||||
dimension's rubric, which Step 3 loads only for a dimension this file puts in play.
|
||||
|
||||
Two rules on using it:
|
||||
|
||||
- A criterion that plainly applies is a finding. Write it up citing file and line.
|
||||
- A criterion that might apply, or whose call the wording here does not settle, is a reason to load
|
||||
that dimension's rubric — never a reason to drop the candidate. This file decides which rubrics
|
||||
to read; it does not settle a close call on its own.
|
||||
|
||||
## description — `references/description-quality.md`
|
||||
|
||||
Flag as FAIL if:
|
||||
|
||||
- **Over 400 characters.** Measured on the folded YAML value, not the raw source lines.
|
||||
`validate.sh` reports the number; do not re-derive it, but do point the Fix at what to cut. Agent
|
||||
descriptions have no platform-documented ceiling of their own, so 400 is the only hard limit
|
||||
there is — do not go looking for a backstop behind it.
|
||||
- **Internal mechanics appear in the description.** Any of:
|
||||
- capability enumeration or a feature list;
|
||||
- output-format detail ("Produces a compact findings report with Why and Fix per finding");
|
||||
- composition or architecture notes ("composes X rather than duplicating Y", "a cross-cutting
|
||||
shared agent", "the human-facing entry point", "replaces the old flat invocation");
|
||||
- implementation detail ("self-validates via a bundled deterministic script").
|
||||
|
||||
None of it can change a routing decision and all of it is preloaded.
|
||||
`Kyberforge.CompositionNote` catches the common phrasings deterministically; the rest is
|
||||
judgment. This is the rule that deflates a description, so apply it before reaching for length.
|
||||
- **The same trigger stated twice in two registers** — a verb list, then the same verbs re-quoted
|
||||
as user phrasings, usually in the same order. One register, whichever routes better.
|
||||
- **Descriptive rather than imperative phrasing** (`This agent ...`, `This is the ...`).
|
||||
`Kyberforge.DescriptionOpener` catches any opener matching `^This`.
|
||||
- **Vague capabilities** ("helps with agents" where "audits an agent definition pair" was
|
||||
available). `Kyberforge.VagueWording` catches the known filler; imprecision outside that list is
|
||||
judgment.
|
||||
- **`Use proactively` in a Copilot or vendor-neutral description.**
|
||||
`KyberforgeCopilot.ProactivePhrase` catches it. The phrase steers the Claude Code runtime and
|
||||
does nothing anywhere else, so in a `.agent.md` it is preloaded text that buys no behaviour.
|
||||
- **Trigger-list, boundary or indirect-trigger content on a hand-invoked agent** — see Step 0 of
|
||||
`references/description-quality.md`.
|
||||
|
||||
Flag as SUGGESTION if:
|
||||
|
||||
- **Over 250 characters** but at or under 400. This tier is what moves the corpus average; the FAIL
|
||||
tier only stops outliers. Report it rather than treating a 399-character description as clean.
|
||||
- A near-miss exclusion is present but targets a weak near-miss.
|
||||
- An indirect trigger is present and warranted but could name the omitted phrasing more precisely.
|
||||
|
||||
**An unresolved boundary target is not graded here.** `validate.sh` resolves boundary targets for
|
||||
agent files at both scopes and tiers the verdict itself — route notation (`/name`, an arrow form)
|
||||
is an ERROR, the bare prose form a SUGGESTION unless a second target in the same sentence resolves.
|
||||
Step 1 has already filed it under `### Structure` at that tier. Take the script's verdict rather
|
||||
than re-resolving the name by hand, and do not re-grade it under description: a hand-walk over a
|
||||
different universe can contradict the script, and re-grading puts one target in the report twice.
|
||||
What is left to judgment is semantic and the script cannot reach it: whether a target that *does*
|
||||
resolve is the right sibling to exclude, and whether a clause naming no target at all ("examine the
|
||||
files manually") should have named one.
|
||||
|
||||
## body, delegation and comment-discipline — `references/body-and-delegation.md`
|
||||
|
||||
Flag as FAIL if:
|
||||
|
||||
- The body restates a procedure owned by a skill the agent can invoke — Fix: invoke `<skill>`
|
||||
instead
|
||||
- A sentence answers "no" to the core test — it is padding
|
||||
- A decision point presents a menu of options with no default
|
||||
- An instruction repeats content already in the description
|
||||
- Frontmatter comments are template scaffolding rather than instruction, or are HTML comments at
|
||||
plugin/APM scope
|
||||
- A prescriptive sequence is used where flexibility is fine, or the reverse
|
||||
|
||||
Flag as SUGGESTION if:
|
||||
|
||||
- The body does not open with a direct role instruction
|
||||
- The body specifies no error handling — nothing tells the agent what to do with malformed,
|
||||
missing or contradictory input
|
||||
- The job the agent describes is unbounded, or bounded only implicitly
|
||||
- A rationale is missing from a rule the agent is expected to enforce — present but unexplained
|
||||
- Comments are useful but verbose enough to bury the field they annotate
|
||||
|
||||
**Never report an agent body as too long on a word count.** ADR-0020 gates a skill body at
|
||||
600/900 words and deliberately gates an agent body at nothing, because an agent body *becomes* the
|
||||
system prompt of a fresh context rather than competing with a live conversation. No number exists
|
||||
to cite. The one length signal that applies is the Copilot runtime's 30,000-character body limit,
|
||||
which `validate.sh` already reports as a SUGGESTION. Length is judged through the delegation FAIL
|
||||
above instead.
|
||||
@@ -14,7 +14,7 @@ source_keys:
|
||||
- **URL:** context7:/websites/code_claude
|
||||
- **Research doc:** plugins/kyberforge/docs/research/docs/claude-code-plugins/sources.md
|
||||
- **Description:** Official Claude Code documentation site indexed by Context7 — plugin manifest schema, subagent definition types, marketplace JSON format, agent markdown file format
|
||||
- **Contributing files:** SKILL.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-project-user.md
|
||||
- **Contributing files:** SKILL.md, references/finding-criteria.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-project-user.md
|
||||
- **Status:** `extracted`
|
||||
|
||||
## claude-code-plugins-docs
|
||||
@@ -22,7 +22,7 @@ source_keys:
|
||||
- **URL:** https://code.claude.com/docs/en/plugins
|
||||
- **Research doc:** plugins/kyberforge/docs/research/docs/claude-code-plugins/sources.md
|
||||
- **Description:** Official Claude Code plugin authoring guide — plugin structure, manifest fields, loading methods, skill namespacing, agent activation, marketplace submission
|
||||
- **Contributing files:** SKILL.md, references/field-inventory.md, references/body-and-delegation.md, references/scope-plugin-apm.md, references/validation-scripts.md
|
||||
- **Contributing files:** SKILL.md, references/finding-criteria.md, references/field-inventory.md, references/body-and-delegation.md, references/scope-plugin-apm.md, references/validation-scripts.md
|
||||
- **Status:** `extracted`
|
||||
|
||||
## claude-code-subagents-docs
|
||||
@@ -30,7 +30,7 @@ source_keys:
|
||||
- **URL:** https://code.claude.com/docs/en/sub-agents
|
||||
- **Research doc:** plugins/kyberforge/docs/research/docs/claude-code-plugins/sources.md
|
||||
- **Description:** Official Claude Code subagent reference — definition format, all frontmatter fields, scope priority, built-in agents, CLI flags, environment variables, known limitations
|
||||
- **Contributing files:** SKILL.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-plugin-apm.md, references/scope-project-user.md, references/validation-scripts.md
|
||||
- **Contributing files:** SKILL.md, references/finding-criteria.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-plugin-apm.md, references/scope-project-user.md, references/validation-scripts.md
|
||||
- **Status:** `extracted`
|
||||
|
||||
## context7-github-en-copilot
|
||||
@@ -38,7 +38,7 @@ source_keys:
|
||||
- **URL:** context7:/websites/github_en_copilot
|
||||
- **Research doc:** plugins/kyberforge/docs/research/docs/github-copilot-plugins/sources.md
|
||||
- **Description:** Official GitHub Copilot documentation indexed by Context7; covers CLI plugins, custom agents, SDK, and marketplace
|
||||
- **Contributing files:** SKILL.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-project-user.md
|
||||
- **Contributing files:** SKILL.md, references/finding-criteria.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-project-user.md
|
||||
- **Status:** `extracted`
|
||||
|
||||
## github-custom-agents-configuration
|
||||
@@ -46,7 +46,7 @@ source_keys:
|
||||
- **URL:** https://docs.github.com/en/copilot/reference/custom-agents-configuration
|
||||
- **Research doc:** plugins/kyberforge/docs/research/docs/github-copilot-plugins/sources.md
|
||||
- **Description:** Reference for cloud and IDE custom agent definition format — frontmatter fields, tool aliases, MCP server config, secrets interpolation, scoping hierarchy
|
||||
- **Contributing files:** SKILL.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-plugin-apm.md, references/scope-project-user.md, references/validation-scripts.md
|
||||
- **Contributing files:** SKILL.md, references/finding-criteria.md, references/field-inventory.md, references/description-quality.md, references/body-and-delegation.md, references/scope-plugin-apm.md, references/scope-project-user.md, references/validation-scripts.md
|
||||
- **Status:** `extracted`
|
||||
|
||||
## github-cli-plugin-reference
|
||||
|
||||
@@ -34,8 +34,14 @@ first of these:
|
||||
`plugin.json` and no `apm.yml` falls through to project or user scope.
|
||||
|
||||
`validate-provenance.sh` exits 0 silently when that walk does not land on a package root, and again
|
||||
when the package has no provenance data. Silence from it is a pass, not a skip you need to
|
||||
investigate.
|
||||
when the package has no provenance data. Check the exit code before you believe the silence:
|
||||
|
||||
- **0** — a pass, not a skip you need to investigate. Both silent cases above land here.
|
||||
- **1** — real findings, on stdout with Why and Fix.
|
||||
- **2** — the check never ran. A missing, doubled, non-file or wrongly-named argument, an
|
||||
undecodable `apm.yml`, or an absent `python3`, each with a diagnostic on stderr and no findings
|
||||
at all. Report the `### Provenance` dimension as unverified and quote the reason. An exit 2 is
|
||||
never a clean pass: empty stdout there means nothing was checked, not that nothing was wrong.
|
||||
|
||||
## Manual fallback
|
||||
|
||||
|
||||
@@ -16,15 +16,48 @@ Arguments:
|
||||
Exit codes:
|
||||
0 All checks passed (or nothing to validate, or not plugin scope)
|
||||
1 One or more checks failed
|
||||
2 Script error (unrecognized file extension — expected .md or .agent.md)
|
||||
2 Usage error, or the argument is not an agent file this script can read
|
||||
|
||||
An exit code of 2 is NOT a finding. SKILL.md tells the auditor to surface a
|
||||
non-zero exit as findings, so a usage error leaving exit 1 with nothing on
|
||||
stdout was indistinguishable from a clean-but-failing run. Environment and
|
||||
argument problems exit 2; only real findings exit 1.
|
||||
|
||||
Exit 2 and the silent exit 0 answer two DIFFERENT questions, and neither may
|
||||
be spelled with the other's code:
|
||||
|
||||
exit 2 the argument is not something this script can audit at all — it is
|
||||
missing, doubled, not a file, or not named .md / .agent.md. Decided
|
||||
before the scope walk-up runs, from the argument alone.
|
||||
exit 0 the argument IS a readable agent file, and the scope walk-up found
|
||||
no type:-bearing apm.yml above it before hitting the \$HOME, .git or
|
||||
filesystem-root boundary. That is a real verdict about a real file —
|
||||
"this agent is user or project scope, so plugin-scope provenance
|
||||
does not apply to it" — not a rejected input.
|
||||
|
||||
scripts/check-scope-walkup-sync.sh's fixture 6 pins the second: a real agent
|
||||
file under a \$HOME with a type-bearing apm.yml ABOVE it must exit 0 with empty
|
||||
output. Widening exit 2 to cover "the walk-up found no package" would break
|
||||
that fixture AND would be wrong on its own terms, because new-agent.sh happily
|
||||
scaffolds exactly that layout.
|
||||
|
||||
Checks performed:
|
||||
0 source_keys present in agent pair but sources.md absent
|
||||
1 FILL IN: placeholders in sources.md
|
||||
2 source_keys in agent files → slug exists in sources.md
|
||||
3 Contributing files listed in sources.md exist on disk (plugin-root relative)
|
||||
3 Contributing files listed in sources.md exist on disk (plugin-root
|
||||
relative). An explicit '(none)' skips silently; a Contributing files block
|
||||
this parser cannot read is reported as an INFO saying checks 3 and 4 did
|
||||
not run, never skipped silently.
|
||||
4 Contributing files back-reference the parent slug in their source_keys
|
||||
5 Research doc field present and not placeholder
|
||||
|
||||
This script has no counterpart to skill-audit's checks 6, 7 and 8 (Research
|
||||
doc field / upstream forward / upstream reverse are numbered 6, 7, 8 there and
|
||||
5 here): an agent at plugin scope is a single file with a plugin-root
|
||||
sources.md, so there is no references/ tree to walk and no upstream research
|
||||
source index to cross-check. parse_status() and the sources.md-basename gate
|
||||
that those checks need exist only in the skill-audit copy.
|
||||
EOF
|
||||
}
|
||||
|
||||
@@ -33,26 +66,131 @@ if [[ "${1:-}" == "--help" || "${1:-}" == "-h" ]]; then
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# Usage and environment problems exit 2, findings exit 1. See the usage text
|
||||
# above for why the two must not share a code, and for why "not plugin scope"
|
||||
# is neither of them. This is a deliberate divergence from validate.sh, which
|
||||
# has no 2 tier for content: validate.sh always prints PASS lines, so a usage
|
||||
# error there is visibly not a findings report. This script prints NOTHING on a
|
||||
# clean run, so exit 1 plus empty stdout was the only signal a caller got
|
||||
# either way.
|
||||
if [[ $# -lt 1 ]]; then
|
||||
echo "Error: agent-file is required." >&2
|
||||
echo "" >&2
|
||||
usage >&2
|
||||
exit 1
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# Extra positional arguments were silently dropped, so a typo'd flag or a second
|
||||
# path looked like it had been honoured.
|
||||
if [[ $# -gt 1 ]]; then
|
||||
echo "Error: expected exactly one argument, got $#: $*" >&2
|
||||
echo "" >&2
|
||||
usage >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# python3 is a HARD dependency. Without this preflight a missing interpreter
|
||||
# produced 'line NN: python3: command not found' and exit 127 — an exit code no
|
||||
# caller maps to anything, from a message that names this script's line number
|
||||
# rather than the missing dependency.
|
||||
if ! command -v python3 > /dev/null 2>&1; then
|
||||
echo "Error: python3 is required but was not found on PATH." >&2
|
||||
echo " Why: skipping the provenance checks entirely would be a vacuous pass." >&2
|
||||
echo " Fix: install python3 (pre-commit itself is a Python application, so it is almost certainly already present)." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# A path that does not exist, or exists but is not a regular file, used to reach
|
||||
# the Python body, get os.path.dirname()'d into some ancestor directory and then
|
||||
# either report a silent exit 0 (no package above it) or — worse — audit a
|
||||
# DIFFERENT agent's package while naming the typo'd path. A typo'd target was
|
||||
# indistinguishable from a clean agent. vale-wrap.sh hard-errors on a
|
||||
# nonexistent path for exactly this reason.
|
||||
#
|
||||
# This is decided from the argument alone, before any walk-up runs, so it cannot
|
||||
# collide with the not-plugin-scope exit 0: that verdict is only ever reached by
|
||||
# a file that got past here.
|
||||
if [[ ! -e "$1" ]]; then
|
||||
echo "Error: no such file: $1" >&2
|
||||
echo " Why: a nonexistent target would otherwise report a silent pass." >&2
|
||||
echo " Fix: pass the path of the agent file to validate." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
if [[ ! -f "$1" ]]; then
|
||||
echo "Error: not a regular file: $1" >&2
|
||||
echo " Why: this script audits one agent file, not a directory of them, and reporting a directory as a pass hides the wrong-target mistake." >&2
|
||||
echo " Fix: pass the agent file itself — .apm/agents/<name>.agent.md — not its parent directory." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# The extension check used to live inside the Python body. It stays exit 2 and
|
||||
# keeps its wording; it moves up here so that every "this argument is not
|
||||
# auditable" verdict is reached in one place, before the interpreter starts and
|
||||
# before the scope walk-up can turn a bad argument into a silent exit 0.
|
||||
case "$1" in
|
||||
*.agent.md | *.md) ;;
|
||||
*)
|
||||
echo "Error: unrecognized extension '$(basename "$1")' — expected .md or .agent.md" >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
|
||||
python3 -u - "$1" <<'PYTHON'
|
||||
import sys
|
||||
import os
|
||||
import re
|
||||
|
||||
# Output is UTF-8 for the same reason input is: under LC_ALL=C the streams
|
||||
# default to ASCII, and every finding this script prints contains an em dash.
|
||||
# Pinning only the reads moved the crash from the read to the write — a
|
||||
# UnicodeEncodeError inside print_findings(), which loses the whole report
|
||||
# after all the checks have already run.
|
||||
for _stream in (sys.stdout, sys.stderr):
|
||||
try:
|
||||
_stream.reconfigure(encoding='utf-8')
|
||||
except AttributeError: # pragma: no cover — Python < 3.7
|
||||
pass
|
||||
|
||||
agent_file = os.path.abspath(sys.argv[1])
|
||||
fname = os.path.basename(agent_file)
|
||||
agent_dir = os.path.dirname(agent_file)
|
||||
|
||||
# --- Sanity-check extension (single vendor-neutral .agent.md file at plugin/APM scope) ---
|
||||
if not (fname.endswith('.agent.md') or fname.endswith('.md')):
|
||||
print(f"Error: unrecognized extension '{fname}' — expected .md or .agent.md", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
# --- Input ----------------------------------------------------------------
|
||||
# Ported from the skill-audit copy, where the same two problems were already
|
||||
# fixed.
|
||||
#
|
||||
# read_text() pins UTF-8 explicitly instead of inheriting
|
||||
# locale.getpreferredencoding(), which is ASCII under LC_ALL=C — an ordinary em
|
||||
# dash in an agent file or in sources.md then aborted the run with a bare
|
||||
# UnicodeDecodeError traceback, or, at the one call site that wrapped its read
|
||||
# in `except Exception: return []`, reported the unreadable file as having no
|
||||
# source_keys and therefore as clean. A file that genuinely is not UTF-8 still
|
||||
# fails; it just says which file and why.
|
||||
#
|
||||
# strip_bom() runs on every read because a leading BOM defeats
|
||||
# parse_frontmatter()'s `^---` anchor, which silently disabled check 2 on a
|
||||
# BOM-prefixed agent file: no frontmatter parsed means no source_keys parsed
|
||||
# means nothing to validate.
|
||||
|
||||
|
||||
class EncodingError(Exception):
|
||||
pass
|
||||
|
||||
|
||||
def strip_bom(text):
|
||||
return text[1:] if text.startswith(u'\ufeff') else text
|
||||
|
||||
|
||||
def read_text(path):
|
||||
"""File contents as text, UTF-8 and BOM-free, with a diagnostic instead of a traceback."""
|
||||
try:
|
||||
with open(path, encoding='utf-8') as fh:
|
||||
return strip_bom(fh.read())
|
||||
except UnicodeDecodeError as exc:
|
||||
raise EncodingError(
|
||||
"not valid UTF-8 (%s at byte %d) — re-save the file as UTF-8; "
|
||||
"this gate does not guess at other encodings"
|
||||
% (exc.reason, exc.start))
|
||||
|
||||
# Matches a top-level `type:` line whose value is exactly one of the four
|
||||
# package content types — identical to validate.sh's APM_TYPE_RE. Group 1's
|
||||
@@ -67,15 +205,31 @@ TYPE_RE = re.compile(r"^type:\s*(['\"]?)(instructions|skill|hybrid|prompts)\1(?:
|
||||
# keep walking. Stop at a $HOME boundary, a .git boundary, or the filesystem
|
||||
# root: none of these is plugin/APM scope, so this script has nothing to
|
||||
# check there.
|
||||
#
|
||||
# Returning None here means NOT PLUGIN SCOPE, which is a verdict, not an error:
|
||||
# the caller exits 0 silently, and scripts/check-scope-walkup-sync.sh fixture 6
|
||||
# pins that. It is deliberately NOT folded into the exit-2 tier above.
|
||||
def find_plugin_root(start_dir):
|
||||
home = os.path.expanduser('~')
|
||||
current = os.path.abspath(start_dir)
|
||||
while True:
|
||||
apm_yml = os.path.join(current, 'apm.yml')
|
||||
if os.path.isfile(apm_yml):
|
||||
with open(apm_yml) as f:
|
||||
if any(TYPE_RE.match(line) for line in f):
|
||||
return current
|
||||
# An apm.yml is a manifest this script must be able to READ to
|
||||
# classify scope at all. Under LC_ALL=C the old bare open() decoded
|
||||
# as ASCII, so a manifest with an accented author name raised
|
||||
# UnicodeDecodeError mid-walk and killed the run with a traceback.
|
||||
# It is an environment problem, not a finding, so it exits 2 rather
|
||||
# than being swallowed into a silent "no package here".
|
||||
try:
|
||||
content = read_text(apm_yml)
|
||||
except EncodingError as exc:
|
||||
print(
|
||||
"Error: %s is %s" % (apm_yml, exc),
|
||||
file=sys.stderr)
|
||||
sys.exit(2)
|
||||
if any(TYPE_RE.match(line) for line in content.splitlines()):
|
||||
return current
|
||||
# $HOME is a non-plugin-scope boundary — checked before the .git test
|
||||
# below (mirrors validate.sh's detect_scope ordering), so a
|
||||
# dotfiles-managed $HOME (yadm, chezmoi bare-repo, etc.) can't shadow
|
||||
@@ -101,7 +255,14 @@ if plugin_root is None:
|
||||
sources_md_path = os.path.join(plugin_root, 'sources.md')
|
||||
|
||||
# --- Helpers ---
|
||||
PLACEHOLDER_RE = re.compile(r'(?<!`)FILL IN:[^`\n]')
|
||||
|
||||
# The trailing character class used to be CONSUMING — `[^`\n]` — so a
|
||||
# `FILL IN:` at end of line matched nothing and escaped checks 1 and 5
|
||||
# entirely. `- **Description:** FILL IN:` is the most likely spelling of a
|
||||
# half-written entry, and it was the one spelling the placeholder gate could
|
||||
# not see. The exclusion it was really expressing is "not inside backticks",
|
||||
# which a lookahead states without eating a character.
|
||||
PLACEHOLDER_RE = re.compile(r'(?<!`)FILL IN:(?!`)')
|
||||
|
||||
def parse_frontmatter(content):
|
||||
m = re.match(r'^---\n(.*?)\n---', content, re.DOTALL)
|
||||
@@ -130,7 +291,56 @@ def parse_source_keys(fm):
|
||||
def parse_h2_slugs(content):
|
||||
return re.findall(r'^## (.+)$', content, re.MULTILINE)
|
||||
|
||||
# ===== BEGIN SHARED CONTRIBUTING-FILES PARSER =====
|
||||
# ONE parser, embedded VERBATIM in two scripts:
|
||||
# plugins/kyberforge/.apm/skills/skill-audit/scripts/validate-provenance.sh
|
||||
# plugins/kyberforge/.apm/skills/agent-audit/scripts/validate-provenance.sh
|
||||
# The block between these markers must stay byte-identical in both. It is
|
||||
# copied rather than imported because a cache-installed plugin's scripts cannot
|
||||
# read files outside their own plugin directory, so there is no single file both
|
||||
# can share — the same constraint that forces the ADR-0020 boundary resolver to
|
||||
# be duplicated across three scripts. Edit one copy, then paste it over the
|
||||
# other.
|
||||
#
|
||||
# tests/test-adr0020-contract.sh hashes both copies and fails on drift. Before
|
||||
# it did, the agent-audit copy's docstring merely ASSERTED the two were
|
||||
# "behaviourally identical" and nothing checked it — which is how the two
|
||||
# already-diverged spellings of the bullet loop went unnoticed.
|
||||
#
|
||||
# Requires: re (imported by the host script).
|
||||
|
||||
|
||||
def parse_contributing_files(content, slug):
|
||||
"""Find the Contributing files for a given slug H2 in content.
|
||||
|
||||
Both authored forms are accepted, because both are in use across the
|
||||
corpus and only recognising the first silently skipped the contributing-
|
||||
file checks on every sources.md written the other way:
|
||||
|
||||
- **Contributing files:** SKILL.md, references/a.md
|
||||
|
||||
**Contributing files:**
|
||||
- SKILL.md (what this source contributed)
|
||||
- references/a.md (what this source contributed)
|
||||
|
||||
Returns a list of paths with any trailing parenthetical note stripped.
|
||||
Note the bullet form's notes may themselves contain commas, so the list
|
||||
is built per bullet rather than by splitting the joined value.
|
||||
|
||||
The three return values are NOT interchangeable, and callers depend on
|
||||
the distinction:
|
||||
|
||||
[path, ...] the entry names contributing files
|
||||
[] the entry EXPLICITLY records "(none)"
|
||||
None the entry says nothing this parser can read
|
||||
|
||||
Only an explicit "(none)" yields []. A "Contributing files:" heading
|
||||
followed by a numbered list, by `*` bullets, or by prose parses nothing
|
||||
and returns None, never [] — a caller reads [] as a deliberate "no
|
||||
contributing files" record and SKIPS its check on that basis, so a parse
|
||||
failure returning [] would silently disable the check instead of leaving
|
||||
the unreadable entry exposed to it.
|
||||
"""
|
||||
pattern = re.compile(
|
||||
r'^## ' + re.escape(slug) + r'\s*\n(.*?)(?=^## |\Z)',
|
||||
re.MULTILINE | re.DOTALL
|
||||
@@ -139,72 +349,153 @@ def parse_contributing_files(content, slug):
|
||||
if not m:
|
||||
return None
|
||||
block = m.group(1)
|
||||
|
||||
def strip_note(entry):
|
||||
# "references/a.md (why)" -> "references/a.md"
|
||||
return re.sub(r'\s*\(.*$', '', entry).strip()
|
||||
|
||||
# Inline form: value on the same line, comma-separated, no notes.
|
||||
cf_m = re.search(r'^\- \*\*Contributing files:\*\* (.+)$', block, re.MULTILINE)
|
||||
if cf_m:
|
||||
value = cf_m.group(1).strip()
|
||||
if value.startswith("(none"):
|
||||
return []
|
||||
return [p for p in (strip_note(x) for x in value.split(","))
|
||||
if p] or None
|
||||
|
||||
# Bullet form: heading on its own line, one file per following bullet.
|
||||
cf_m = re.search(r'^\*\*Contributing files:\*\*\s*$', block, re.MULTILINE)
|
||||
if not cf_m:
|
||||
return None
|
||||
return cf_m.group(1).strip()
|
||||
files = []
|
||||
for line in block[cf_m.end():].splitlines():
|
||||
line = line.strip()
|
||||
if not line:
|
||||
if files:
|
||||
break
|
||||
continue
|
||||
if not line.startswith("- "):
|
||||
break
|
||||
entry = line[2:].strip()
|
||||
if entry.startswith("(none"):
|
||||
return []
|
||||
entry = strip_note(entry)
|
||||
if entry:
|
||||
files.append(entry)
|
||||
return files or None
|
||||
# ===== END SHARED CONTRIBUTING-FILES PARSER =====
|
||||
|
||||
def parse_research_doc(content, slug):
|
||||
def parse_research_docs(content, slug):
|
||||
"""Every Research doc value under a given slug H2, in document order.
|
||||
|
||||
The caller uses the first and reports the rest. Returning only the first —
|
||||
what this did before — meant a second '- **Research doc:**' line in one
|
||||
entry was silently ignored, so an author who added a doc rather than
|
||||
replacing one got check 5 run against the old value and no hint that the
|
||||
new one was never looked at.
|
||||
"""
|
||||
pattern = re.compile(
|
||||
r'^## ' + re.escape(slug) + r'\s*\n(.*?)(?=^## |\Z)',
|
||||
re.MULTILINE | re.DOTALL
|
||||
)
|
||||
m = pattern.search(content)
|
||||
if not m:
|
||||
return None
|
||||
return []
|
||||
block = m.group(1)
|
||||
rd_m = re.search(r'^\- \*\*Research doc:\*\* (.+)$', block, re.MULTILINE)
|
||||
if not rd_m:
|
||||
return None
|
||||
return rd_m.group(1).strip()
|
||||
return [v.strip() for v in
|
||||
re.findall(r'^\- \*\*Research doc:\*\* (.+)$', block, re.MULTILINE)]
|
||||
|
||||
findings = []
|
||||
has_fail = False
|
||||
|
||||
# A finding identical in every field is the same finding, and the same file is
|
||||
# now reached by more than one check — the agent file is read once for its own
|
||||
# source_keys and again as a contributing file, so an unreadable one would
|
||||
# otherwise be reported twice with the same words. Distinct findings about the
|
||||
# same file still both appear.
|
||||
def _record(entry):
|
||||
if entry not in findings:
|
||||
findings.append(entry)
|
||||
|
||||
def emit_fail(desc, fpath, why, fix):
|
||||
global has_fail
|
||||
has_fail = True
|
||||
findings.append(("FAIL", desc, fpath, why, fix))
|
||||
_record(("FAIL", desc, fpath, why, fix, None))
|
||||
|
||||
# INFO does not set has_fail and does not change the exit code. It is for a
|
||||
# check that could not RUN — an unverified entry, not a broken one — and it
|
||||
# exists so that "did not run" is never spelled the same way as "passed".
|
||||
def emit_info(desc, fpath, note):
|
||||
_record(("INFO", desc, fpath, None, None, note))
|
||||
|
||||
def print_findings():
|
||||
for kind, desc, fpath, why, fix in findings:
|
||||
print(f"FAIL {desc} — {fpath}")
|
||||
print(f" Why: {why}")
|
||||
print(f" Fix: {fix}")
|
||||
print()
|
||||
for entry in findings:
|
||||
kind = entry[0]
|
||||
desc = entry[1]
|
||||
fpath = entry[2]
|
||||
why = entry[3]
|
||||
fix = entry[4]
|
||||
note = entry[5]
|
||||
if kind == "FAIL":
|
||||
print(f"FAIL {desc} — {fpath}")
|
||||
print(f" Why: {why}")
|
||||
print(f" Fix: {fix}")
|
||||
print()
|
||||
else:
|
||||
print(f"INFO {desc} — {fpath}")
|
||||
print(f" Note: {note}")
|
||||
print()
|
||||
|
||||
def emit_unreadable(rel, exc):
|
||||
"""Report a file this script cannot decode. Never a silent skip."""
|
||||
emit_fail(
|
||||
f"File is {exc}",
|
||||
rel,
|
||||
f"'{rel}' cannot be decoded, so its frontmatter — and any source_keys in it — "
|
||||
f"cannot be read. This used to be swallowed by a bare 'except Exception: return []', "
|
||||
f"which reported the unreadable file as having no source_keys and therefore as clean.",
|
||||
f"Re-save '{rel}' as UTF-8."
|
||||
)
|
||||
|
||||
# --- Collect source_keys from agent pair ---
|
||||
def get_source_keys_from_file(fpath):
|
||||
def get_source_keys_from_file(fpath, rel):
|
||||
if not os.path.isfile(fpath):
|
||||
return []
|
||||
try:
|
||||
with open(fpath) as f:
|
||||
content = f.read()
|
||||
except Exception:
|
||||
content = read_text(fpath)
|
||||
except EncodingError as exc:
|
||||
emit_unreadable(rel, exc)
|
||||
return []
|
||||
fm, _ = parse_frontmatter(content)
|
||||
return parse_source_keys(fm)
|
||||
|
||||
# Plugin/APM scope is a single vendor-neutral file — no counterpart to merge.
|
||||
given_keys = get_source_keys_from_file(agent_file)
|
||||
rel_given = os.path.relpath(agent_file, plugin_root)
|
||||
given_keys = get_source_keys_from_file(agent_file, rel_given)
|
||||
all_source_keys = given_keys
|
||||
|
||||
sources_md_exists = os.path.isfile(sources_md_path)
|
||||
|
||||
# Early exit: nothing to validate
|
||||
# Early exit: nothing to validate. The read above can itself raise a finding —
|
||||
# an unreadable agent file — so print before leaving; the clean case still
|
||||
# prints nothing and exits 0.
|
||||
if not all_source_keys and not sources_md_exists:
|
||||
sys.exit(0)
|
||||
print_findings()
|
||||
sys.exit(1 if has_fail else 0)
|
||||
|
||||
sources_content = None
|
||||
sources_slugs = set()
|
||||
if sources_md_exists:
|
||||
with open(sources_md_path) as f:
|
||||
sources_content = f.read()
|
||||
try:
|
||||
sources_content = read_text(sources_md_path)
|
||||
except EncodingError as exc:
|
||||
emit_unreadable("sources.md", exc)
|
||||
print_findings()
|
||||
sys.exit(1)
|
||||
sources_slugs = set(parse_h2_slugs(sources_content))
|
||||
|
||||
# --- Check 0: source_keys present but sources.md absent ---
|
||||
if not sources_md_exists and all_source_keys:
|
||||
rel_given = os.path.relpath(agent_file, plugin_root)
|
||||
emit_fail(
|
||||
"source_keys declared but sources.md is absent",
|
||||
rel_given,
|
||||
@@ -240,11 +531,49 @@ for fpath, keys in [(agent_file, given_keys)]:
|
||||
)
|
||||
|
||||
# --- Checks 3, 4, 5: Per-slug checks in sources.md ---
|
||||
for slug in parse_h2_slugs(sources_content):
|
||||
# Check 3: Contributing files exist (paths relative to plugin root)
|
||||
cf_value = parse_contributing_files(sources_content, slug)
|
||||
if cf_value and not cf_value.startswith("(none"):
|
||||
cf_files = [p.strip() for p in cf_value.split(",") if p.strip()]
|
||||
|
||||
# Every per-slug parser below — parse_contributing_files, parse_research_docs —
|
||||
# locates its block with pattern.search(), so a slug written twice resolves to
|
||||
# the FIRST block every time. Iterating the raw heading list therefore checked
|
||||
# the first block's fields twice and the second block's never: a duplicated slug
|
||||
# is half-validated, and looked fully validated. The duplicate is announced and
|
||||
# the repeat visit dropped.
|
||||
all_slugs = parse_h2_slugs(sources_content)
|
||||
unique_slugs = []
|
||||
for _slug in all_slugs:
|
||||
if _slug in unique_slugs:
|
||||
continue
|
||||
unique_slugs.append(_slug)
|
||||
_count = all_slugs.count(_slug)
|
||||
if _count > 1:
|
||||
emit_info(
|
||||
f"Duplicate '## {_slug}' entry in sources.md — only the first block is checked",
|
||||
f"sources.md (## {_slug})",
|
||||
f"'## {_slug}' appears {_count} times. Every field parser here takes the first match, so the "
|
||||
f"second and later blocks' Contributing files and Research doc are never validated — "
|
||||
f"checks 3, 4 and 5 did not run for them. "
|
||||
f"Merge the blocks into one entry, or give each a distinct slug and reference it from source_keys."
|
||||
)
|
||||
|
||||
for slug in unique_slugs:
|
||||
# Checks 3 and 4: Contributing files exist (paths relative to plugin root),
|
||||
# and back-reference the slug. `[]` and None are NOT the same answer here.
|
||||
# `[]` is the author writing "(none)" — there is nothing to check and the
|
||||
# skip is correct. None is a Contributing-files block this parser cannot
|
||||
# read, and skipping THAT silently disables both checks on the one entry
|
||||
# least likely to be right, which is the failure mode
|
||||
# parse_contributing_files' own docstring warns about. Say so out loud.
|
||||
cf_files = parse_contributing_files(sources_content, slug)
|
||||
if cf_files is None:
|
||||
emit_info(
|
||||
f"Contributing-file checks skipped for '{slug}' — the Contributing files block could not be parsed",
|
||||
f"sources.md (## {slug})",
|
||||
f"The '## {slug}' entry has no Contributing files list this parser can read — a missing field, a bare heading, '*' bullets, a numbered list, or prose all read as unparsable rather than as an empty declaration. "
|
||||
f"Checks 3 and 4 did not run for this slug, so nothing verified that its contributing files exist or name it back. "
|
||||
f"Write the value as '- **Contributing files:** <comma-separated paths>', or as a '**Contributing files:**' heading followed by '- ' bullets — "
|
||||
f"or record '(none)' if this source contributed no files."
|
||||
)
|
||||
elif cf_files:
|
||||
for cf_rel in cf_files:
|
||||
cf_abs = os.path.join(plugin_root, cf_rel)
|
||||
if not os.path.isfile(cf_abs):
|
||||
@@ -256,8 +585,11 @@ for slug in parse_h2_slugs(sources_content):
|
||||
)
|
||||
else:
|
||||
# Check 4: Bidirectional — file should list slug in its source_keys
|
||||
with open(cf_abs) as f:
|
||||
cf_content = f.read()
|
||||
try:
|
||||
cf_content = read_text(cf_abs)
|
||||
except EncodingError as exc:
|
||||
emit_unreadable(cf_rel, exc)
|
||||
continue
|
||||
cf_fm, _ = parse_frontmatter(cf_content)
|
||||
cf_keys = parse_source_keys(cf_fm)
|
||||
if slug not in cf_keys:
|
||||
@@ -269,7 +601,17 @@ for slug in parse_h2_slugs(sources_content):
|
||||
)
|
||||
|
||||
# Check 5: Research doc field required
|
||||
rd_value = parse_research_doc(sources_content, slug)
|
||||
rd_values = parse_research_docs(sources_content, slug)
|
||||
if len(rd_values) > 1:
|
||||
emit_info(
|
||||
f"Multiple '- **Research doc:**' lines for '{slug}' — only the first is used",
|
||||
f"sources.md (## {slug})",
|
||||
f"The '## {slug}' entry has {len(rd_values)} Research doc lines; check 5 ran against the first "
|
||||
f"('{rd_values[0]}') and never looked at the rest. "
|
||||
f"Keep one Research doc line per entry — if a slug genuinely came from two documents, split it into two slugs, "
|
||||
f"or name the extra document inside the first value's annotation where it is at least visible."
|
||||
)
|
||||
rd_value = rd_values[0] if rd_values else None
|
||||
if rd_value is None:
|
||||
emit_fail(
|
||||
"Research doc field missing",
|
||||
|
||||
@@ -69,6 +69,24 @@ import glob
|
||||
|
||||
import yaml
|
||||
|
||||
# Output is UTF-8 for the same reason input is: under LC_ALL=C the streams
|
||||
# default to ASCII, and this script's own message text carries em dashes (the
|
||||
# ADR-0020 boundary SUGGESTION is one). Pinning only the reads moved the crash
|
||||
# from the read to the write — a UnicodeEncodeError raised while PRINTING, after
|
||||
# every check has already run, which loses the whole report and (here) flips a
|
||||
# clean exit 0 into a traceback and an exit 1. read_text() in the shared
|
||||
# resolver block below pins the reads; this pins the writes.
|
||||
#
|
||||
# Deliberately OUTSIDE the ADR-0020 shared boundary resolver block: the two
|
||||
# validate.sh copies print findings, skill-size-check.sh has its own top-level
|
||||
# equivalent, and tests/test-adr0020-contract.sh hashes that block for
|
||||
# byte-identity across all three.
|
||||
for _stream in (sys.stdout, sys.stderr):
|
||||
try:
|
||||
_stream.reconfigure(encoding='utf-8')
|
||||
except AttributeError: # pragma: no cover — Python < 3.7
|
||||
pass
|
||||
|
||||
agent_file = os.path.abspath(sys.argv[1])
|
||||
script_dir = sys.argv[2]
|
||||
|
||||
@@ -253,9 +271,27 @@ def _collect_package(pkg_dir, names):
|
||||
safe_dir = glob.escape(pkg_dir)
|
||||
for sub in ('.apm/skills/*/', 'skills/*/'):
|
||||
for path in glob.glob(os.path.join(safe_dir, sub)):
|
||||
names.add(os.path.basename(path.rstrip('/')).lower())
|
||||
# A directory is a skill only if it HOLDS a SKILL.md. An empty
|
||||
# leftover — a deleted skill whose directory survived, a scaffolding
|
||||
# stub, an editor's stray mkdir — is untracked by git, so it exists
|
||||
# on the machine that made it and nowhere else. Counting it made a
|
||||
# boundary target resolve locally and dangle in a fresh clone: the
|
||||
# same install-dependence the deployed-tree rule above exists to
|
||||
# remove, arriving through a different door.
|
||||
if os.path.isfile(os.path.join(path, 'SKILL.md')):
|
||||
names.add(os.path.basename(path.rstrip('/')).lower())
|
||||
for sub in ('.apm/agents/*.md', 'agents/*.md'):
|
||||
for path in glob.glob(os.path.join(safe_dir, sub)):
|
||||
# The same rule one directory over, which until now had no
|
||||
# counterpart here at all: the skills branch above tests for a
|
||||
# SKILL.md, the agents branch took every glob hit on trust. A
|
||||
# DIRECTORY named `ghost-agent.md` matches `*.md` and glob does not
|
||||
# tell the two apart, so a leftover of that shape resolved a routing
|
||||
# target on the machine holding it and dangled everywhere else —
|
||||
# identical install-dependence, arriving through the one door
|
||||
# nobody guarded.
|
||||
if not os.path.isfile(path):
|
||||
continue
|
||||
base = os.path.basename(path)
|
||||
if base.endswith('.agent.md'):
|
||||
base = base[:-len('.agent.md')]
|
||||
@@ -446,8 +482,17 @@ def known_targets(start_dir):
|
||||
# condition, pc-run's "run pre-commit hooks" reads as a route to a
|
||||
# non-existent `pre-commit` skill.
|
||||
# * A BARE arrow target counts only in ADR-0020's compressed boundary form,
|
||||
# `Not <thing> -> <skill-name>`. Without that, diagnose's process chain
|
||||
# "fix -> regression-test" reads as a route to `regression-test`.
|
||||
# `Not <thing> -> <skill-name>`. The example that motivated it is gone:
|
||||
# diagnose's process chain "fix -> regression-test", which without the
|
||||
# gate read as a route to a non-existent `regression-test` skill, was cut
|
||||
# when issue #99 retrofitted that description. So the gate is currently
|
||||
# UNEXERCISED — gating and not gating produce the same verdict corpus-wide.
|
||||
# Keep it anyway. It is a false-positive guard against prose no one has
|
||||
# written yet, and any new process chain re-arms it. Unexercised is not the
|
||||
# same as unnecessary, and the branch it guards is still load-bearing: the
|
||||
# bare-arrow rule is the sole extractor for three real targets in
|
||||
# kyberforge's audit skills (agent-audit -> agent-author, agent-audit ->
|
||||
# skill-audit, skill-audit -> skill-author), all written unbackticked.
|
||||
# * A backticked hyphenated token counts only inside a boundary sentence.
|
||||
# Unconditionally, `pre-push` or `commit-msg` in a TRIGGER clause is a hard
|
||||
# FAIL with no escape hatch. Gating it costs nothing (measured over this
|
||||
@@ -525,6 +570,66 @@ def known_targets(start_dir):
|
||||
# ambiguity to resolve, and an author who wants a route checked unconditionally
|
||||
# has two ways to say so.
|
||||
#
|
||||
# BOTH FORMS ARE SWEPT FOR ON THEIR OWN, and that is a repair of the promise
|
||||
# above rather than a widening of it. Until the sweeps existed, notation was
|
||||
# only ever seen as the OBJECT OF A ROUTE VERB (`use
|
||||
# /name`) or as the tail of a `not ... ->` clause with no `;` or sentence end in
|
||||
# between. Every one of these therefore exited 0 in total silence — no ERROR, no
|
||||
# SUGGESTION, not even the target's name:
|
||||
# Do not use for Y — /no-such-skill instead.
|
||||
# Do not use for Y; /no-such-skill handles that.
|
||||
# Do not use for Y (/no-such-skill covers it).
|
||||
# Do not use for Y — that is /no-such-skill's job.
|
||||
# Do not use for Y — defer to /no-such-skill.
|
||||
# Do not use for Y — /no-such-skill.
|
||||
# Do not use for Y; -> no-such-skill covers it.
|
||||
# For W, /no-such-skill is the right entry point.
|
||||
# The target was never EXTRACTED, so the notation-first rule in _add() had
|
||||
# nothing to apply itself to and the "always blocks" promise was false for the
|
||||
# ordinary way an author writes the thing. The SUGGESTION tier made it worse
|
||||
# than a gap: its printed remedy tells the author to "write it as `/name` or
|
||||
# `-> name` and it will be checked properly", and taking that advice turned a
|
||||
# visible SUGGESTION into silence — the gate teaching the one edit that blinds
|
||||
# it.
|
||||
#
|
||||
# THE TWO SWEEPS ARE GATED DIFFERENTLY, and the asymmetry is the whole point.
|
||||
# `/name` is Claude Code's invocation syntax and nothing else — no English
|
||||
# sentence contains one by accident — so the ADR-0020 amendment and
|
||||
# docs/spec/gates.md both promise it blocks UNCONDITIONALLY, for any name. So
|
||||
# NOTATION_SLASH is swept over every sentence, boundary marker or not. Gating it
|
||||
# on BOUNDARY_MARKER made that promise false for the last sentence of
|
||||
# Do not use for Z — use /real-skill instead.
|
||||
# For W, /no-such-skill is the right entry point.
|
||||
# which exited 0 in total silence: the boundary clause is one sentence up, so
|
||||
# the sweep never looked at the sentence carrying the broken route. Extraction is
|
||||
# per-sentence by design (corroboration is scoped to one sentence), which is
|
||||
# exactly what made the gap invisible.
|
||||
#
|
||||
# NOTATION_ARROW stays gated on BOUNDARY_MARKER, and so does the backtick sweep.
|
||||
# Neither form is unambiguous: `-> name` is also how a process chain is written
|
||||
# ("reproduce -> minimise -> regression-test") and a code span is how a tool, a
|
||||
# file and a skill are all cited. Ungating either would fire on prose that
|
||||
# carries no routing intent at all — the false-positive class this whole
|
||||
# extractor is tuned against.
|
||||
#
|
||||
# BOTH `/name` PATTERNS REFUSE A TOKEN THAT IS PART OF A PATH: a following `/`,
|
||||
# or a `.` followed by a non-space, means `references/foo.md`, `docs/a/b.md` or
|
||||
# `https://x/y`, not a route. A sentence's closing `.` is not followed by a
|
||||
# non-space, so `— /no-such-skill.` still counts.
|
||||
#
|
||||
# THAT GUARD IS WRITTEN `(?![\w-])` AND NOT `\b`, because `\b` is not a guard at
|
||||
# all here: it holds after a hyphen, so when the trailing lookahead rejected the
|
||||
# full segment the engine simply backtracked to a shorter hyphen-terminated
|
||||
# prefix and reported THAT as a route. Every one of these was a hard blocking
|
||||
# ERROR naming a skill nobody had written:
|
||||
# the config lives at /opt-tools/bin/thing. -> 'opt'
|
||||
# see /api-docs/v2.md for the schema. -> 'api' AND 'api-docs'
|
||||
# the file /no-such-skill.md documents it. -> 'no-such'
|
||||
# `(?![\w-])` forbids the shortened prefix outright, so the whole segment is
|
||||
# rejected as the path it is. MARKED_TARGET carries the same guard: it had no
|
||||
# trailing lookahead whatsoever, so `see /api-docs/v2.md` raised the second of
|
||||
# the two errors above through the route-verb path rather than the sweep.
|
||||
#
|
||||
# NAMESPACE: `plugin:skill` is live in this repo (native user-scope installs
|
||||
# still resolve `gitea:gitea-prs`), so the patterns admit an optional
|
||||
# `<plugin>:` prefix and normalize_target() strips it before resolution.
|
||||
@@ -535,7 +640,8 @@ ROUTE_VERB = (r"(?:use|uses|using|run|runs|invoke|invokes|invoking|try|see"
|
||||
r"|that'?s|compose|composes|call|calls"
|
||||
r"|routes?\s+to|delegates?\s+to|prefers?|switch(?:es)?\s+to"
|
||||
r"|hands?\s+off\s+to)")
|
||||
MARKED_TARGET = r"(?:`/?(%s)`|(?<![\w./*-])/(%s)\b)" % (NAME_ANY, NAME_ANY)
|
||||
MARKED_TARGET = (r"(?:`/?(%s)`|(?<![\w./*-])/(%s)(?![\w-])(?!/|\.\S))"
|
||||
% (NAME_ANY, NAME_ANY))
|
||||
ANY_TARGET = r"(?:%s|(%s)\b)" % (MARKED_TARGET, NAME_HYPH)
|
||||
ROUTE_MARKED = re.compile(r"\b%s\s+(?:the\s+|an?\s+)?%s" % (ROUTE_VERB, MARKED_TARGET), re.I)
|
||||
ROUTE_ANY = re.compile(r"\b%s\s+(?:the\s+|an?\s+)?%s" % (ROUTE_VERB, ANY_TARGET), re.I)
|
||||
@@ -548,12 +654,53 @@ ROUTE_ANY = re.compile(r"\b%s\s+(?:the\s+|an?\s+)?%s" % (ROUTE_VERB, ANY_TARGET)
|
||||
CONT_MARKED = re.compile(r"\s*(?:or|and|/|,)\s*%s" % MARKED_TARGET, re.I)
|
||||
CONT_ANY = re.compile(r"\s*(?:or|and|/|,)\s*%s" % ANY_TARGET, re.I)
|
||||
ARROW_MARKED = re.compile(r"(?:->|→)\s*%s" % MARKED_TARGET, re.I)
|
||||
ARROW_BOUNDARY = re.compile(r"\bnot\b[^.;]*?(?:->|→)\s*(%s)\b" % NAME_HYPH, re.I)
|
||||
# The two EXPLICIT ROUTE NOTATION sweeps. NOTATION_SLASH runs over EVERY
|
||||
# sentence; NOTATION_ARROW is scoped to a boundary sentence by its caller (see
|
||||
# the asymmetry note in the header). NOTATION_SLASH is deliberately not a reuse
|
||||
# of MARKED_TARGET's `/name` alternative: that one only ever runs behind a route
|
||||
# verb or an arrow, and it may match a namespaced or path-adjacent token in
|
||||
# positions this free-standing sweep must refuse.
|
||||
# NOTATION_ARROW is ARROW_BOUNDARY minus its leading `\bnot\b%s*?`, which is
|
||||
# what made `Do not use for Y; -> no-such-skill covers it.` invisible:
|
||||
# CLAUSE_BODY cannot cross the `;`, so the clause's own punctuation disarmed the
|
||||
# check. Dropping that prefix costs the one false positive the bare-arrow bullet
|
||||
# above names — a process chain ending in a hyphenated word, `Instead, reproduce
|
||||
# -> minimise -> regression-test.` — and costs it only in a sentence that already
|
||||
# carries a BOUNDARY_MARKER. That exposure is neither new nor larger: the same
|
||||
# chain written `Do not use for X — reproduce -> regression-test.` was already a
|
||||
# hard ERROR under ARROW_BOUNDARY, so this changes which boundary words reach the
|
||||
# arrow, not whether prose can. An author who means the chain and not a route
|
||||
# writes it in its own sentence, where neither pattern looks.
|
||||
NOTATION_SLASH = re.compile(
|
||||
r"(?<![\w./*-])/(%s)(?![\w-])(?!/|\.\S)" % NAME_ANY, re.I)
|
||||
NOTATION_ARROW = re.compile(r"(?:->|→)\s*(%s)\b" % NAME_HYPH, re.I)
|
||||
# CLAUSE_BODY is what may sit between `Not` and the arrow, and it is NOT
|
||||
# `[^.;]`. That class cannot cross a `.`, so every boundary clause naming a
|
||||
# DOTTED FILENAME between the two — `.pre-commit-config.yaml`, `AGENTS.md`,
|
||||
# `.vale.ini` — was invisible to both patterns below, and the two resulting
|
||||
# failures were different sizes (issue #110):
|
||||
# * with a BACKTICKED target the clause was MISDIAGNOSED. The backtick sweep
|
||||
# still extracted the target, so the route was checked, but the gate
|
||||
# reported "no boundary clause" on a clause that was present and working.
|
||||
# Three authors in two retrofit waves reworded a correct clause to satisfy
|
||||
# the regex, one of them stripping the very filename that discriminates the
|
||||
# skill from its neighbour.
|
||||
# * with a BARE target the clause was UNCHECKED. ARROW_BOUNDARY is the only
|
||||
# extractor for a bare arrow target, so `Not AGENTS.md -> no-such-skill`
|
||||
# produced no target, no dangling report and no missing-clause SUGGESTION.
|
||||
# Silence, not noise — the worse of the two failure modes.
|
||||
# A dot inside a filename is followed by a non-space; a sentence-ending dot is
|
||||
# followed by whitespace or by end of string. So the class admits a `.` only
|
||||
# when the next character is not whitespace, which crosses `AGENTS.md` and
|
||||
# still stops at a real sentence end.
|
||||
CLAUSE_BODY = r"(?:[^.;]|\.(?=\S))"
|
||||
ARROW_BOUNDARY = re.compile(
|
||||
r"\bnot\b%s*?(?:->|→)\s*(%s)\b" % (CLAUSE_BODY, NAME_HYPH), re.I)
|
||||
BACKTICK = re.compile(r"`(%s)`" % NAME_HYPH, re.I)
|
||||
# A boundary clause takes two shapes and BOTH count: the prose markers, and
|
||||
# ADR-0020's compressed arrow form `Not <thing> -> <name>`.
|
||||
BOUNDARY_MARKER = re.compile(r"\b(?:do\s+not|instead|rather\s+than|not\s+for)\b", re.I)
|
||||
BOUNDARY_ARROW = re.compile(r"\bnot\b[^.;]*?(?:->|→)", re.I)
|
||||
BOUNDARY_ARROW = re.compile(r"\bnot\b%s*?(?:->|→)" % CLAUSE_BODY, re.I)
|
||||
# Sentence boundaries decide the CORROBORATION scope above, so getting one wrong
|
||||
# is not cosmetic — it moves a target between SUGGESTION and blocking ERROR. Two
|
||||
# shapes common in these descriptions defeat the naive "period, space, capital"
|
||||
@@ -573,9 +720,17 @@ BOUNDARY_ARROW = re.compile(r"\bnot\b[^.;]*?(?:->|→)", re.I)
|
||||
# a lowercase letter. Verified zero-delta on the current corpus (37 ERROR / 58
|
||||
# SUGGESTION / 2 dangling before and after) — this protects the descriptions
|
||||
# issue #99 is about to rewrite, not the ones already measured.
|
||||
# re.I here too, and NOT as a tidy-up: this was the one pattern in the file
|
||||
# built without it, contradicting the uniformity note on CONT_*/ARROW_* above.
|
||||
# Without the flag `E.g.` and `I.e.` — the sentence-initial spellings, which is
|
||||
# where an abbreviation most often lands — matched none of the lookbehinds, so
|
||||
# the clause split at the abbreviation, the corroborating target was stranded on
|
||||
# the far side of the cut, and a genuinely dangling target silently demoted from
|
||||
# blocking ERROR to SUGGESTION. That is the OVER-SPLIT failure described
|
||||
# directly above, still live for exactly the capitalised half of the input.
|
||||
SENTENCE_SPLIT = re.compile(
|
||||
u'(?<!\\be\\.g\\.)(?<!\\bi\\.e\\.)(?<!\\betc\\.)(?<!\\bvs\\.)(?<!\\bcf\\.)'
|
||||
u'(?<=[.!?])\\s+(?=[A-Za-z`"“(])')
|
||||
u'(?<=[.!?])\\s+(?=[A-Za-z`"“(])', re.I)
|
||||
|
||||
# The token that may follow a route target without turning it into a compound
|
||||
# modifier: punctuation, end of sentence, a conjunction, a boundary word, or a
|
||||
@@ -634,11 +789,26 @@ def _notation(text, start, arrow):
|
||||
|
||||
|
||||
def _add(out, text, name, start, end, strict=None, arrow=False):
|
||||
"""Record one target as (name, may_dangle, notation).
|
||||
|
||||
NOTATION IS DECIDED FIRST, and when it is set the follower test is skipped.
|
||||
The header above promises that route notation "always blocks", and for the
|
||||
`/name` form that was false: `-> name` reached this function with
|
||||
strict=True from its two call sites, but `/name` did not, so it fell to
|
||||
_terminal() and a follower outside FOLLOWER_OK set may_dangle=False. The
|
||||
target then reached unresolved_targets() unblockable — and, before the
|
||||
companion fix there, unreported as well. `... use /no-such-skill
|
||||
afterwards.` exited 0 in total silence, on the one form ADR-0020 offers an
|
||||
author who wants a route checked unconditionally.
|
||||
"""
|
||||
if not name:
|
||||
return
|
||||
notation = _notation(text, start, arrow)
|
||||
if strict is None and notation:
|
||||
strict = True
|
||||
out.append((name,
|
||||
_terminal(text, end) if strict is None else strict,
|
||||
_notation(text, start, arrow)))
|
||||
notation))
|
||||
|
||||
|
||||
def _scan(text, route_re, cont_re, out):
|
||||
@@ -678,7 +848,19 @@ def _extract_sentence(sentence):
|
||||
for match in ARROW_BOUNDARY.finditer(sentence):
|
||||
_add(out, sentence, match.group(1), match.start(1), match.end(1),
|
||||
strict=True, arrow=True)
|
||||
# `/name` wherever it sits, in ANY sentence — not only where a route verb or
|
||||
# an arrow happens to precede it, and NOT only inside a boundary sentence.
|
||||
# See the EXPLICIT ROUTE NOTATION note in the header for the eight phrasings
|
||||
# this recovers and for why silence was the failure mode. The sweep takes no
|
||||
# follower test: _add() reads the notation first and marks it.
|
||||
for match in NOTATION_SLASH.finditer(sentence):
|
||||
_add(out, sentence, match.group(1), match.start(1), match.end(1))
|
||||
if boundary:
|
||||
# The arrow and backtick forms are ambiguous in ordinary prose, so they
|
||||
# stay scoped to a sentence that carries a boundary marker.
|
||||
for match in NOTATION_ARROW.finditer(sentence):
|
||||
_add(out, sentence, match.group(1), match.start(1), match.end(1),
|
||||
strict=True, arrow=True)
|
||||
for match in BACKTICK.finditer(sentence):
|
||||
_add(out, sentence, match.group(1), match.start(1), match.end(1))
|
||||
return out
|
||||
@@ -697,6 +879,85 @@ def boundary_targets(description):
|
||||
return sorted({name for name, _, _ in _extract(description)})
|
||||
|
||||
|
||||
def _arrow_targets(description):
|
||||
"""Names extracted from ARROW notation specifically.
|
||||
|
||||
Kept apart from boundary_targets() because the arrow form is the one shape
|
||||
that ALWAYS names a target: ADR-0020's `Not <thing> -> <name>`. A clause
|
||||
written that way from which nothing could be extracted is a parse failure
|
||||
that deserves its own message, and telling it apart needs the arrow targets
|
||||
alone rather than every target in the description.
|
||||
"""
|
||||
out = []
|
||||
for sentence in SENTENCE_SPLIT.split(description):
|
||||
for match in ARROW_MARKED.finditer(sentence):
|
||||
name, _, _ = _first(match)
|
||||
if name:
|
||||
out.append(name)
|
||||
for match in ARROW_BOUNDARY.finditer(sentence):
|
||||
out.append(match.group(1))
|
||||
return out
|
||||
|
||||
|
||||
def boundary_clause_status(description):
|
||||
"""'absent', 'unparsed' or 'present' — three outcomes, not two.
|
||||
|
||||
Issue #110's standing request: the gate must distinguish "no boundary
|
||||
clause" from "boundary clause I could not parse". Reporting the first for
|
||||
the second sends the author hunting for a problem that is not there, and
|
||||
three of them reworded a correct clause to satisfy a regex instead.
|
||||
|
||||
'unparsed' is the narrow, certain case: an ADR-0020 arrow clause was
|
||||
detected and NO target came out of it. The arrow form always names one, so
|
||||
zero targets means the name is written in a shape the extractor cannot see
|
||||
— a single-word bare target (`Not X -> forge`, which has to be written
|
||||
`` `forge` `` or `/forge`) is the live example, since single-word names are
|
||||
deliberately not matchable bare.
|
||||
|
||||
A PROSE clause yielding no target is NOT reported: "Do not use for anything
|
||||
else" is a complete and legitimate boundary clause that names nowhere to go.
|
||||
"""
|
||||
if BOUNDARY_ARROW.search(description) and not _arrow_targets(description):
|
||||
return 'unparsed'
|
||||
if has_boundary_clause(description):
|
||||
return 'present'
|
||||
return 'absent'
|
||||
|
||||
|
||||
def multi_target_arrow_clauses(description):
|
||||
"""[(first, second)] for arrow clauses naming more than one target.
|
||||
|
||||
Issue #107: only the FIRST target after an arrow is resolved. The
|
||||
conjunction continuation (CONT_*) is wired to the prose route verbs and
|
||||
never to arrows, so `Not X -> a or b` resolved `a`, left `b` neither
|
||||
resolved nor reported, and then printed "1 of 1 boundary target(s) resolve"
|
||||
on a clause naming two — a gate under-reporting its own coverage, which is
|
||||
the one failure mode ADR-0020 says a gate must not have.
|
||||
|
||||
The clause is REJECTED rather than the arrow scan extended. Extending it
|
||||
would widen the resolver's deliberately conservative false-positive tuning
|
||||
across every arrow in the corpus; rejecting costs nothing and makes the
|
||||
one-arrow-per-target convention — already what every retrofitted gitea
|
||||
skill does in practice — explicit instead of folkloric. The caller emits a
|
||||
SUGGESTION telling the author to split.
|
||||
"""
|
||||
hits = []
|
||||
for sentence in SENTENCE_SPLIT.split(description):
|
||||
matches = (list(ARROW_MARKED.finditer(sentence))
|
||||
+ list(ARROW_BOUNDARY.finditer(sentence)))
|
||||
for match in matches:
|
||||
first, _, _ = _first(match)
|
||||
if not first:
|
||||
continue
|
||||
cont = CONT_ANY.match(sentence, match.end())
|
||||
if not cont:
|
||||
continue
|
||||
second, _, _ = _first(cont)
|
||||
if second:
|
||||
hits.append((first, second))
|
||||
return hits
|
||||
|
||||
|
||||
def unresolved_targets(description, known):
|
||||
"""Targets resolving to nothing, split into (blocking, reported).
|
||||
|
||||
@@ -713,6 +974,17 @@ def unresolved_targets(description, known):
|
||||
Everything else is reported and left alone. `known` is the resolved
|
||||
universe from known_targets(); passing an empty set is not meaningful —
|
||||
callers check for that first and decline out loud instead.
|
||||
|
||||
A NON-TERMINAL target is reported, never dropped. FOLLOWER_OK is a closed
|
||||
whitelist of maybe eighty words, so the follower rule says "this token is
|
||||
outside a list I keep" and not "this is prose" — and the old `continue`
|
||||
turned that into invisibility at every tier. The gate then failed OPEN on
|
||||
its own unfamiliarity: any target followed by a word nobody thought to
|
||||
enumerate was neither blocked nor mentioned, so the check that did not run
|
||||
said nothing about not running. The follower rule may withdraw the power to
|
||||
BLOCK a commit — that is what it was added for, and the ATTRIBUTIVE USE note
|
||||
above is the argument for it — but it may not withdraw visibility, which is
|
||||
the same rule the corroboration tier already follows.
|
||||
"""
|
||||
blocking, reported = set(), set()
|
||||
for sentence in SENTENCE_SPLIT.split(description):
|
||||
@@ -721,7 +993,10 @@ def unresolved_targets(description, known):
|
||||
if normalize_target(name) in known}
|
||||
for name, may_dangle, notation in found:
|
||||
key = normalize_target(name)
|
||||
if key in known or not may_dangle:
|
||||
if key in known:
|
||||
continue
|
||||
if not may_dangle:
|
||||
reported.add(name)
|
||||
continue
|
||||
if notation or (resolved - {key}):
|
||||
blocking.add(name)
|
||||
@@ -801,6 +1076,47 @@ def description_value(fm_text):
|
||||
return re.sub(r'\s+', ' ', value).strip()
|
||||
|
||||
|
||||
def hand_invoked(fm_text):
|
||||
"""True when the frontmatter marks this file as reached only by hand.
|
||||
|
||||
`disable-model-invocation: true` removes a skill from the model-visible
|
||||
listing entirely — it is not preloaded, and the Skill tool refuses to call
|
||||
it — so its description is never matched against user intent. ADR-0020 and
|
||||
skill-author's contract give such a skill ONE plain human-facing sentence:
|
||||
no trigger list, no boundary clause. No validator knew the field existed
|
||||
(issue #108), so the boundary-clause SUGGESTION fired on exactly the shape
|
||||
the contract mandates, and its remedy — "add a boundary clause so the router
|
||||
knows where NOT to send this skill" — was addressed to a router that cannot
|
||||
see the skill at all. An author who followed the advice made the file worse.
|
||||
|
||||
Only the ROUTING rules are lifted. The body word budget still applies: the
|
||||
body is loaded on invocation like any other, and competes with the caller's
|
||||
live conversation the same way. So does the 400-character description FAIL —
|
||||
a hand-invoked description is not preloaded, but it is still the one line
|
||||
the user reads when choosing from the `/` menu, and the ceiling is the
|
||||
outlier stop rather than the style target.
|
||||
|
||||
A parse failure returns False rather than raising. This is a MODIFIER on
|
||||
other checks, not a check of its own: the frontmatter's validity is decided,
|
||||
and failed, by description_value() on the same text, and raising a second
|
||||
exception here would report one broken file twice with two different
|
||||
diagnoses.
|
||||
"""
|
||||
try:
|
||||
data = yaml.safe_load(fm_text)
|
||||
except Exception:
|
||||
return False
|
||||
if not isinstance(data, dict):
|
||||
return False
|
||||
value = data.get('disable-model-invocation')
|
||||
if isinstance(value, str):
|
||||
# PyYAML already resolves the unquoted YAML 1.1 booleans, so this only
|
||||
# catches a QUOTED "true" — which a host reads as truthy and which no
|
||||
# gate should treat as opting back in to the routing rules.
|
||||
return value.strip().lower() in ('true', 'yes', 'on')
|
||||
return value is True
|
||||
|
||||
|
||||
# --- Body-shape checks (skills only; agents have no references/ dir) -------
|
||||
# Deterministic and countable, so they are enforced here. Whether a given
|
||||
# gotcha is WARRANTED is semantic and stays the auditor's judgment, which is why
|
||||
@@ -915,7 +1231,15 @@ def missing_reference_pointers(body, skill_dir):
|
||||
end = masked.find('\n', match.end())
|
||||
if end < 0:
|
||||
end = len(masked)
|
||||
if REFERENCE_PAST.search(masked[start:end]):
|
||||
# The pointer's OWN SPAN is excised before the sweep. Run over the
|
||||
# whole line, the past-tense test matched the very path it was judging,
|
||||
# so a file exempted itself by its NAME: `references/deprecated-api.md`,
|
||||
# `references/removed-flags.md` and `references/gone.md` produced no
|
||||
# ERROR at all, while `references/missing.md` — an identical break —
|
||||
# errored. The exemption is about what the SENTENCE says about the
|
||||
# pointer, never about what the pointer is called.
|
||||
line = masked[start:match.start()] + masked[match.end():end]
|
||||
if REFERENCE_PAST.search(line):
|
||||
continue
|
||||
if REFERENCE_QUALIFIER.search(masked[start:match.start()]):
|
||||
continue
|
||||
@@ -968,8 +1292,14 @@ def agent_description(fm, local_fname):
|
||||
f"not run — {local_fname}")
|
||||
return None
|
||||
|
||||
def check_description_budget(value, local_fname):
|
||||
"""ADR-0020 description gates — identical for every scope."""
|
||||
def check_description_budget(value, local_fname, by_hand=False):
|
||||
"""ADR-0020 description gates — identical for every scope.
|
||||
|
||||
`by_hand` is ADR-0020's hand-invocation carve-out (issue #108): an agent
|
||||
carrying `disable-model-invocation: true` is absent from the model-visible
|
||||
listing, so the 250-character SUGGESTION — a routing-quality budget — has
|
||||
no listing to apply to. The 400-character ceiling is unaffected.
|
||||
"""
|
||||
if not value:
|
||||
return
|
||||
dlen = len(value)
|
||||
@@ -979,13 +1309,13 @@ def check_description_budget(value, local_fname):
|
||||
f"agent is invoked. Keep a trigger clause, at most one capability clause, "
|
||||
f"and a boundary clause; move capability enumeration, output-format detail, "
|
||||
f"composition notes and implementation detail to the body — {local_fname}")
|
||||
elif dlen > DESC_SUGGEST_CHARS:
|
||||
elif dlen > DESC_SUGGEST_CHARS and not by_hand:
|
||||
suggest(f"description is {dlen} chars — over the {DESC_SUGGEST_CHARS}-character "
|
||||
f"ADR-0020 target (hard fail at {DESC_MAX_CHARS}). The SUGGESTION tier is "
|
||||
f"what moves the corpus average; the FAIL tier only stops outliers "
|
||||
f"— {local_fname}")
|
||||
|
||||
def check_boundary(value, fpath, local_fname):
|
||||
def check_boundary(value, fpath, local_fname, by_hand=False):
|
||||
"""ADR-0020 boundary clause + resolvable boundary targets.
|
||||
|
||||
agent-author's SKILL.md states that an agent's boundary targets must
|
||||
@@ -1002,10 +1332,31 @@ def check_boundary(value, fpath, local_fname):
|
||||
# SUGGESTION, not FAIL: detecting the absence is deterministic, but whether
|
||||
# this particular agent warrants a boundary clause is judgment. All four
|
||||
# agents in this corpus currently lack one.
|
||||
if not has_boundary_clause(value):
|
||||
#
|
||||
# THREE outcomes, not two: "no boundary clause" and "boundary clause I could
|
||||
# not parse" are different findings (issue #110). And a hand-invoked agent is
|
||||
# exempt from the clause altogether (issue #108) — the boundary-target
|
||||
# resolution below still runs, because a target it DOES name should still
|
||||
# resolve.
|
||||
status = boundary_clause_status(value) if not by_hand else 'present'
|
||||
if status == 'absent':
|
||||
suggest(f"description has no boundary clause — add the prose form (\"Do not use "
|
||||
f"for X — use `y` instead\") or ADR-0020's compressed form (\"Not X -> y\") "
|
||||
f"so the router knows where NOT to send this agent — {local_fname}")
|
||||
elif status == 'unparsed':
|
||||
suggest(f"description has an arrow boundary clause (\"Not X -> y\") from which no "
|
||||
f"target could be read, so the dangling-target check did not run on it — "
|
||||
f"the clause is PRESENT and unparsed, not missing. Most often the target "
|
||||
f"is a single word, which is deliberately not matchable bare: write it as "
|
||||
f"`name` or /name — {local_fname}")
|
||||
if not by_hand:
|
||||
# One arrow, one target: a second name after the same arrow is resolved
|
||||
# by nothing and reported by nothing (issue #107).
|
||||
for first, second in multi_target_arrow_clauses(value):
|
||||
suggest(f"an arrow boundary clause names more than one target ('{first}', then "
|
||||
f"'{second}') and only the first is resolved — the second is checked by "
|
||||
f"nothing. Split it into one arrow per target: \"Not X -> {first}. "
|
||||
f"Not Y -> {second}.\" — {local_fname}")
|
||||
targets = boundary_targets(value)
|
||||
if not targets:
|
||||
return
|
||||
@@ -1240,8 +1591,9 @@ def check_apm_agent_file(fpath, allowlist, stem):
|
||||
else:
|
||||
if PLACEHOLDER_RE.search(folded):
|
||||
fail(f"description contains unfilled FILL IN: placeholder — {local_fname}")
|
||||
check_description_budget(folded, local_fname)
|
||||
check_boundary(folded, fpath, local_fname)
|
||||
by_hand = hand_invoked(fm)
|
||||
check_description_budget(folded, local_fname, by_hand)
|
||||
check_boundary(folded, fpath, local_fname, by_hand)
|
||||
|
||||
# body — required, non-empty, no placeholder; same Copilot truncation risk
|
||||
# applies since this file compiles verbatim into a real Copilot file downstream.
|
||||
@@ -1336,8 +1688,9 @@ def check_file(fpath, file_provider):
|
||||
else:
|
||||
if PLACEHOLDER_RE.search(folded):
|
||||
fail(f"description contains unfilled FILL IN: placeholder — {local_fname}")
|
||||
check_description_budget(folded, local_fname)
|
||||
check_boundary(folded, fpath, local_fname)
|
||||
by_hand = hand_invoked(fm)
|
||||
check_description_budget(folded, local_fname, by_hand)
|
||||
check_boundary(folded, fpath, local_fname, by_hand)
|
||||
|
||||
# body
|
||||
if not body.strip():
|
||||
|
||||
@@ -433,3 +433,477 @@ EOF
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_success
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Checks 3 and 4: None ("could not parse") is NOT [] ("explicitly (none)")
|
||||
#
|
||||
# parse_contributing_files returns three distinguishable answers and checks 3
|
||||
# and 4 have to honour all three. `[]` is the author writing "(none)" — the
|
||||
# skip is correct and silent. None is a Contributing files block the parser
|
||||
# cannot read, and skipping THAT silently disables both checks on the one entry
|
||||
# least likely to be right, which is the failure mode the parser's own
|
||||
# docstring warns about. The assertions below are therefore about the INFO
|
||||
# appearing; a silent exit 0 is exactly the bug.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@test "INFO: an unparsable Contributing files block names the slug instead of skipping checks 3 and 4 silently" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_agent_with_source_keys "$root"
|
||||
cat > "$root/sources.md" <<EOF
|
||||
# Sources
|
||||
|
||||
## my-source
|
||||
|
||||
- **URL:** https://example.com/my-source
|
||||
- **Description:** A test source.
|
||||
- **Research doc:** (none)
|
||||
**Contributing files:**
|
||||
* .apm/agents/ghost.agent.md (asterisk bullets are not the bullet form)
|
||||
- **Status:** \`extracted\`
|
||||
EOF
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_success
|
||||
assert_output --partial "INFO"
|
||||
assert_output --partial "Contributing-file checks skipped for 'my-source' — the Contributing files block could not be parsed"
|
||||
}
|
||||
|
||||
@test "INFO: an entry with no Contributing files field at all is reported, not skipped silently" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_agent_with_source_keys "$root"
|
||||
cat > "$root/sources.md" <<EOF
|
||||
# Sources
|
||||
|
||||
## my-source
|
||||
|
||||
- **URL:** https://example.com/my-source
|
||||
- **Description:** A test source.
|
||||
- **Research doc:** (none)
|
||||
- **Status:** \`extracted\`
|
||||
EOF
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_success
|
||||
assert_output --partial "INFO"
|
||||
assert_output --partial "Contributing-file checks skipped for 'my-source' — the Contributing files block could not be parsed"
|
||||
}
|
||||
|
||||
@test "checks 3 and 4 skipped silently: an explicit '(none)' emits no INFO" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_agent_with_source_keys "$root"
|
||||
make_sources_md "$root" "my-source" "(none — not used directly)"
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_success
|
||||
assert_output ""
|
||||
}
|
||||
|
||||
@test "checks 3 and 4 still run: a parseable Contributing files list is not diverted to the INFO" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_agent_with_source_keys "$root"
|
||||
make_sources_md "$root" "my-source" ".apm/agents/nonexistent.agent.md"
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_failure
|
||||
assert_output --partial "FAIL"
|
||||
assert_output --partial "Contributing file '.apm/agents/nonexistent.agent.md' does not exist"
|
||||
refute_output --partial "could not be parsed"
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# The INFO tier itself: kind-aware printing and a kind-aware exit code
|
||||
#
|
||||
# INFO is new here — before it, findings was a 5-tuple and print_findings
|
||||
# stamped every entry FAIL. The two cases below pin the tier rather than any
|
||||
# one check: an INFO must print under the INFO prefix and leave the exit code
|
||||
# at 0, and a real FAIL must keep printing under the FAIL prefix and still exit
|
||||
# non-zero even when an INFO is sitting in the same findings list.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@test "INFO tier: an INFO alone prints as INFO with a Note and does not set a failing exit code" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_agent_with_source_keys "$root"
|
||||
cat > "$root/sources.md" <<EOF
|
||||
# Sources
|
||||
|
||||
## my-source
|
||||
|
||||
- **URL:** https://example.com/my-source
|
||||
- **Description:** A test source.
|
||||
- **Research doc:** (none)
|
||||
- **Status:** \`extracted\`
|
||||
EOF
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_success
|
||||
assert_output --partial "INFO Contributing-file checks skipped for 'my-source'"
|
||||
assert_output --partial "Note:"
|
||||
refute_output --partial "FAIL"
|
||||
}
|
||||
|
||||
@test "FAIL tier: a genuine FAIL alongside an INFO still prints as FAIL and exits non-zero" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_agent_with_source_keys "$root"
|
||||
cat > "$root/sources.md" <<EOF
|
||||
# Sources
|
||||
|
||||
## my-source
|
||||
|
||||
- **URL:** https://example.com/my-source
|
||||
- **Description:** A test source.
|
||||
- **Contributing files:** .apm/agents/nonexistent.agent.md
|
||||
- **Research doc:** (none)
|
||||
- **Status:** \`extracted\`
|
||||
|
||||
## ghost-source
|
||||
|
||||
- **URL:** https://example.com/ghost-source
|
||||
- **Description:** Another test source.
|
||||
- **Research doc:** (none)
|
||||
- **Status:** \`extracted\`
|
||||
EOF
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_failure
|
||||
assert_output --partial "FAIL Contributing file '.apm/agents/nonexistent.agent.md' does not exist"
|
||||
assert_output --partial "Why:"
|
||||
assert_output --partial "INFO Contributing-file checks skipped for 'ghost-source'"
|
||||
assert_output --partial "Note:"
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# The exit-2 tier, and its boundary with the silent exit 0
|
||||
#
|
||||
# Ported from the skill-audit sibling, which had already split usage and
|
||||
# environment errors (exit 2) away from findings (exit 1). SKILL.md tells the
|
||||
# auditor to surface a non-zero exit, so a usage error leaving exit 1 with
|
||||
# nothing on stdout was indistinguishable from a clean-but-failing run.
|
||||
#
|
||||
# The reconciliation this script needs and the sibling does not: "the walk-up
|
||||
# found no type:-bearing apm.yml" is NOT bad input. It is a verdict about a
|
||||
# real, readable agent file — user or project scope, where plugin-scope
|
||||
# provenance does not apply — and scripts/check-scope-walkup-sync.sh fixture 6
|
||||
# pins it as exit 0 with empty output. Every exit-2 gate is therefore decided
|
||||
# from the ARGUMENT ALONE, before the walk-up runs, so the two can never
|
||||
# collide. The two tests at the end of this block assert both halves.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@test "exit 2: no arguments is a usage error, not a finding" {
|
||||
run bash "$SCRIPT"
|
||||
[ "$status" -eq 2 ]
|
||||
assert_output --partial "agent-file is required"
|
||||
}
|
||||
|
||||
@test "exit 2: a second positional argument is rejected instead of silently dropped" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_clean_agent "$root"
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md" --some-typo
|
||||
[ "$status" -eq 2 ]
|
||||
assert_output --partial "expected exactly one argument"
|
||||
}
|
||||
|
||||
@test "exit 2: a nonexistent path is an error, not a silent pass" {
|
||||
run bash "$SCRIPT" "$TMPDIR/no-such-agent.agent.md"
|
||||
[ "$status" -eq 2 ]
|
||||
assert_output --partial "no such file"
|
||||
}
|
||||
|
||||
@test "exit 2: a directory is not an agent file" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
run bash "$SCRIPT" "$root/.apm/agents"
|
||||
[ "$status" -eq 2 ]
|
||||
assert_output --partial "not a regular file"
|
||||
}
|
||||
|
||||
@test "exit 2: an unrecognized extension is rejected before the walk-up runs" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
echo "not an agent" > "$root/.apm/agents/my-agent.txt"
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.txt"
|
||||
[ "$status" -eq 2 ]
|
||||
assert_output --partial "unrecognized extension"
|
||||
}
|
||||
|
||||
@test "exit 2: a PATH with no python3 names the missing dependency instead of exiting 127" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_clean_agent "$root"
|
||||
local emptybin="$TMPDIR/emptybin"
|
||||
mkdir -p "$emptybin"
|
||||
local bash_bin
|
||||
bash_bin="$(command -v bash)"
|
||||
run env -i PATH="$emptybin" HOME="$HOME" "$bash_bin" "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
[ "$status" -eq 2 ]
|
||||
# The needle is the DIAGNOSTIC, not the bare word: with no preflight, bash's
|
||||
# own "python3: command not found" would satisfy a bare-word match.
|
||||
assert_output --partial "python3 is required"
|
||||
}
|
||||
|
||||
@test "reconciliation: a REAL agent file at non-plugin scope still exits 0 silently, never 2" {
|
||||
# scripts/check-scope-walkup-sync.sh fixture 6 in miniature. The exit-2 tier
|
||||
# must not widen to cover "find_plugin_root returned None": the file exists,
|
||||
# is readable and is correctly named — it is simply user/project scope.
|
||||
local dir="$TMPDIR/anc"
|
||||
mkdir -p "$dir"
|
||||
cat > "$dir/apm.yml" <<EOF
|
||||
name: outer-package
|
||||
version: 0.1.0
|
||||
type: skill
|
||||
EOF
|
||||
local fake_home="$dir/fakehome"
|
||||
mkdir -p "$fake_home/.apm/agents"
|
||||
cat > "$fake_home/.apm/agents/my-agent.agent.md" <<EOF
|
||||
---
|
||||
name: my-agent
|
||||
description: A valid agent description.
|
||||
source_keys:
|
||||
- my-source
|
||||
---
|
||||
|
||||
You are a test agent.
|
||||
EOF
|
||||
run env HOME="$fake_home" bash "$SCRIPT" "$fake_home/.apm/agents/my-agent.agent.md"
|
||||
[ "$status" -eq 0 ]
|
||||
assert_output ""
|
||||
}
|
||||
|
||||
@test "reconciliation: a nonexistent path inside a non-plugin-scope tree exits 2, not the old silent 0" {
|
||||
# The other half. Before the exit-2 tier, a typo'd path anywhere outside a
|
||||
# package took the not-plugin-scope exit and reported a silent pass, so the
|
||||
# typo and a clean agent produced identical output and identical status.
|
||||
local fake_home="$TMPDIR/plainhome"
|
||||
mkdir -p "$fake_home/.apm/agents"
|
||||
run env HOME="$fake_home" bash "$SCRIPT" "$fake_home/.apm/agents/typo.agent.md"
|
||||
[ "$status" -eq 2 ]
|
||||
assert_output --partial "no such file"
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# PLACEHOLDER_RE: the trailing character class was CONSUMING
|
||||
#
|
||||
# `(?<!\`)FILL IN:[^\`\n]` required a character after the colon, so a `FILL IN:`
|
||||
# at end of line matched nothing and escaped checks 1 and 5 entirely — and
|
||||
# `- **Description:** FILL IN:` is the most likely spelling of a half-written
|
||||
# entry. The lookahead states the same exclusion without eating a character.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@test "FAIL: a FILL IN: placeholder at end of line is caught, not skipped" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_agent_with_source_keys "$root"
|
||||
cat > "$root/sources.md" <<EOF
|
||||
# Sources
|
||||
|
||||
## my-source
|
||||
|
||||
- **URL:** https://example.com/my-source
|
||||
- **Description:** FILL IN:
|
||||
- **Contributing files:** .apm/agents/my-agent.agent.md
|
||||
- **Research doc:** (none)
|
||||
- **Status:** \`extracted\`
|
||||
EOF
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_failure
|
||||
assert_output --partial "Unfilled FILL IN: placeholder"
|
||||
}
|
||||
|
||||
@test "FAIL: a Research doc value that is a bare end-of-line FILL IN: is caught" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_agent_with_source_keys "$root"
|
||||
cat > "$root/sources.md" <<EOF
|
||||
# Sources
|
||||
|
||||
## my-source
|
||||
|
||||
- **URL:** https://example.com/my-source
|
||||
- **Description:** A test source.
|
||||
- **Contributing files:** .apm/agents/my-agent.agent.md
|
||||
- **Research doc:** FILL IN:
|
||||
- **Status:** \`extracted\`
|
||||
EOF
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_failure
|
||||
assert_output --partial "Research doc field is empty or placeholder"
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Encoding, read side: read_text() pins UTF-8 and strips a BOM
|
||||
#
|
||||
# The old code used bare open() calls inheriting locale.getpreferredencoding(),
|
||||
# which is ASCII under LC_ALL=C, and wrapped exactly one of them in
|
||||
# `except Exception: return []` — so an unreadable agent file was reported as
|
||||
# having no source_keys and therefore as CLEAN. The other call sites had no
|
||||
# handler at all and died with a traceback.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@test "FAIL: an undecodable agent file is reported, not swallowed into a clean pass" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
printf '\xff\xfe---\nname: my-agent\n---\n' > "$root/.apm/agents/my-agent.agent.md"
|
||||
make_sources_md "$root" "my-source" "(none)"
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_failure
|
||||
assert_output --partial "not valid UTF-8"
|
||||
refute_output --partial "Traceback"
|
||||
}
|
||||
|
||||
@test "FAIL: an undecodable contributing file is reported, not a traceback" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_agent_with_source_keys "$root"
|
||||
printf '\xff\xfe---\nname: other\n---\n' > "$root/.apm/agents/other.agent.md"
|
||||
make_sources_md "$root" "my-source" ".apm/agents/other.agent.md"
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_failure
|
||||
assert_output --partial "not valid UTF-8"
|
||||
refute_output --partial "Traceback"
|
||||
}
|
||||
|
||||
@test "exit 2: an undecodable apm.yml names the file instead of dying mid walk-up" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_clean_agent "$root"
|
||||
printf 'name: t\nversion: 0.1.0\ntype: skill\n# \xff\xfe\n' > "$root/apm.yml"
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
[ "$status" -eq 2 ]
|
||||
assert_output --partial "not valid UTF-8"
|
||||
refute_output --partial "Traceback"
|
||||
}
|
||||
|
||||
@test "a BOM-prefixed agent file still has its source_keys read (check 2 runs)" {
|
||||
# A leading BOM defeats parse_frontmatter()'s ^--- anchor, so no frontmatter
|
||||
# parsed means no source_keys parsed means nothing to validate — check 2
|
||||
# went silently missing on exactly the file it was pointed at.
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
printf '\xef\xbb\xbf---\nname: my-agent\ndescription: A valid agent description.\nsource_keys:\n - ghost-source\n---\n\nYou are a test agent.\n' \
|
||||
> "$root/.apm/agents/my-agent.agent.md"
|
||||
make_sources_md "$root" "my-source" "(none)"
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_failure
|
||||
assert_output --partial "source_keys slug 'ghost-source' not found in sources.md"
|
||||
}
|
||||
|
||||
@test "under LC_ALL=C a sources.md carrying an em dash is read, not a UnicodeDecodeError" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_agent_with_source_keys "$root"
|
||||
cat > "$root/sources.md" <<EOF
|
||||
# Sources
|
||||
|
||||
## my-source
|
||||
|
||||
- **URL:** https://example.com/my-source
|
||||
- **Description:** A test source — with an em dash.
|
||||
- **Contributing files:** .apm/agents/ghost.agent.md
|
||||
- **Research doc:** (none)
|
||||
- **Status:** \`extracted\`
|
||||
EOF
|
||||
run env LC_ALL=C PYTHONUTF8=0 bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_failure
|
||||
assert_output --partial "Contributing file '.apm/agents/ghost.agent.md' does not exist"
|
||||
refute_output --partial "Traceback"
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Encoding, write side: sys.stdout/stderr.reconfigure(encoding='utf-8')
|
||||
#
|
||||
# Pinning only the reads moved the crash from the read to the WRITE. Every
|
||||
# finding this script prints contains an em dash, so under LC_ALL=C
|
||||
# print_findings() died with UnicodeEncodeError after every check had already
|
||||
# run — losing the whole report at the last step.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@test "under LC_ALL=C the findings report is printed, not lost to a UnicodeEncodeError" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_agent_with_source_keys "$root"
|
||||
make_sources_md "$root" "my-source" ".apm/agents/ghost.agent.md"
|
||||
run env LC_ALL=C PYTHONUTF8=0 bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_failure
|
||||
assert_output --partial "FAIL Contributing file '.apm/agents/ghost.agent.md' does not exist"
|
||||
assert_output --partial "Why:"
|
||||
refute_output --partial "UnicodeEncodeError"
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Half-validated entries announced instead of passing silently
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@test "INFO: a duplicated '## slug' says only the first block was checked" {
|
||||
# Every per-slug parser locates its block with pattern.search(), so a slug
|
||||
# written twice resolves to the FIRST block every time: the second block's
|
||||
# fields are never validated, and the entry looked fully checked.
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_agent_with_source_keys "$root"
|
||||
cat > "$root/sources.md" <<EOF
|
||||
# Sources
|
||||
|
||||
## my-source
|
||||
|
||||
- **URL:** https://example.com/my-source
|
||||
- **Description:** A test source.
|
||||
- **Contributing files:** (none)
|
||||
- **Research doc:** (none)
|
||||
- **Status:** \`extracted\`
|
||||
|
||||
## my-source
|
||||
|
||||
- **URL:** https://example.com/dup
|
||||
- **Description:** A duplicate entry.
|
||||
- **Contributing files:** .apm/agents/ghost.agent.md
|
||||
- **Research doc:** (none)
|
||||
- **Status:** \`extracted\`
|
||||
EOF
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_success
|
||||
assert_output --partial "Duplicate '## my-source' entry in sources.md"
|
||||
# The second block's ghost contributing file is genuinely never checked —
|
||||
# the INFO is what makes that visible rather than a silent half-pass.
|
||||
refute_output --partial "does not exist"
|
||||
}
|
||||
|
||||
@test "INFO: a second '- **Research doc:**' line in one entry is announced, not ignored" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
make_agent_with_source_keys "$root"
|
||||
cat > "$root/sources.md" <<EOF
|
||||
# Sources
|
||||
|
||||
## my-source
|
||||
|
||||
- **URL:** https://example.com/my-source
|
||||
- **Description:** A test source.
|
||||
- **Contributing files:** (none)
|
||||
- **Research doc:** (none)
|
||||
- **Research doc:** docs/research/added-later.md
|
||||
- **Status:** \`extracted\`
|
||||
EOF
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_success
|
||||
assert_output --partial "Multiple '- **Research doc:**' lines for 'my-source'"
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Finding dedup
|
||||
#
|
||||
# The agent file is read once for its own source_keys and again as a
|
||||
# contributing file, so an unreadable one produced the identical finding twice.
|
||||
# Distinct findings about the same file still both appear.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@test "the same unreadable file reached by two checks is reported once, not twice" {
|
||||
local root="$TMPDIR/package"
|
||||
make_package "$root"
|
||||
printf '\xff\xfe---\nname: my-agent\n---\n' > "$root/.apm/agents/my-agent.agent.md"
|
||||
make_sources_md "$root" "my-source" ".apm/agents/my-agent.agent.md"
|
||||
run bash "$SCRIPT" "$root/.apm/agents/my-agent.agent.md"
|
||||
assert_failure
|
||||
local count
|
||||
count="$(printf '%s\n' "$output" | grep -c "^FAIL File is not valid UTF-8" || true)"
|
||||
[ "$count" -eq 1 ]
|
||||
}
|
||||
|
||||
@@ -943,3 +943,23 @@ EOF
|
||||
refute_output --partial "Traceback"
|
||||
refute_output --partial "FileNotFoundError"
|
||||
}
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Encoding, write side: sys.stdout/stderr.reconfigure(encoding='utf-8')
|
||||
#
|
||||
# read_text() in the shared resolver block pins the READS to UTF-8. That moved
|
||||
# the LC_ALL=C crash to the WRITE: this script's own message text carries em
|
||||
# dashes (the ADR-0020 boundary SUGGESTION is one), so the streams' ASCII
|
||||
# default raised UnicodeEncodeError while PRINTING — after every check had
|
||||
# already run. Here it also flipped a clean exit 0 into a traceback and exit 1.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@test "under LC_ALL=C the report is printed, not lost to a UnicodeEncodeError" {
|
||||
local root="$TMPDIR/locale-pkg"
|
||||
make_apm_agent "$root" "locale-agent"
|
||||
run env LC_ALL=C PYTHONUTF8=0 bash "$SCRIPT" "$root/.apm/agents/locale-agent.agent.md"
|
||||
assert_success
|
||||
assert_output --partial "description has no boundary clause"
|
||||
refute_output --partial "UnicodeEncodeError"
|
||||
refute_output --partial "Traceback"
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user