Files
holocron/docs/spec/gates.md
Defame1297 ffcbed6c41 fix(tests): replace pipefail-racy echo | grep -q with here-strings
Why

Two suites failed intermittently — tests/test-vale-wrap.sh case 21 and
tests/test-check-release-needed.sh cases 4 and 15 — on correct output, and never
when run alone. The cause is the `echo "$OUT" | grep -q P` idiom under
`set -o pipefail`: grep -q exits as soon as it has an answer, bash's echo can
hand a multi-line value to the pipe one line at a time, and a write after the
reader is gone kills echo with SIGPIPE. pipefail then reports the writer's
death, so output that DID match reads as "no match". Every observed failure had
lines after its match; case 15's match is on line 1 of 6, the widest window in
that file.

Forced with a pause before the writer's last line, the pipe form failed 50 of 50
runs; a here-string, a match on the last line, and the same pipe without
pipefail each passed 50 of 50. Unforced the rate is about 1 per 670 suite runs,
which is why it read as a flaky gate rather than a bug.

The failures at review time are consistent with this, but were not proven to be
it: the suite was running while agents edited live config files in place, and a
brief change to .vale.ini or .pre-commit-hooks.yaml would produce the same two
failures. The race is real and fixed either way.

Implementation Notes

`grep -q P <<< "$VAR"` has no separate writer process, so there is nothing to
race. It is not a retry or a sleep. 121 sites converted across 9 files, three of
them scripts rather than tests: new-agent.sh, new-skill.sh and
check-executables-allow-sync.sh. None ships via .pre-commit-hooks.yaml, so no
external consumer pins them, and all three are single-pipeline checks whose
verdict cannot change.

Left alone deliberately: 14 sites whose writer is a command, not a shell
builtin — they either absorb the writer's status with `|| true` or are python3
and awk, which write once at exit — and one file with no pipefail. `printf '%s'`
sites differ from a here-string only by a trailing newline, which no -q verdict
on a non-empty pattern depends on.

tests/test-no-pipefail-early-exit-grep.sh is a static guard against new
occurrences, discovered automatically by run-tests.sh. It only scans files that
set pipefail, joins continuation lines, skips comments, and flags only
echo/printf writers. Its first case proves the scanner can fail before its
second trusts a clean verdict on the tree.

A guard covers exactly the spellings its regex models, so the miss surface was
measured rather than assumed. Four were found and closed: pipefail declared as
`set -o errexit -o pipefail` (where the old pattern required pipefail to follow
the FIRST -o, and a file-level miss skips every site in that file); a writer
separated from grep by an intermediate stage; a pipeline wrapped on a trailing
`|` rather than a backslash; and readers spelled egrep, fgrep, /bin/grep,
`command grep` or with an env-var prefix. Segment characters exclude a bare `&`
so `echo ok && other | grep -q x`, whose writer is `other`, does not false-fire.
Widening surfaced 5 live sites invisible to the original scanner, all in
tests/test-apm-current-hook.sh, all `echo "$out" | json_field ... | grep -q`;
they are safe today only because json_field is python3, which reads to EOF and
writes once. Fixtures go 4 to 12 vulnerable spellings plus near-miss negatives.

Two `grep ... | head -1` sites (test-vale-wrap.sh) are the same race with a
different early-exiting reader, and are fixed by absorbing the writer. The
scanner deliberately does not model `head`, `sed -n 1p` or a bare `read`: most
legitimate uses in this tree are already absorbed with `|| true` and the scanner
cannot see absorption from pipeline text, so a high false-positive rate would be
how this guard gets weakened. Heredoc bodies are scanned as code; none in the
tree trips it today.

Impact

The bug predates the factory-audit merge: every converted site in
check-release-needed and case 21 dates to 4d018af and aa8cc22 (2026-08-09).

Test suites go 19 to 20. `run-tests.sh --strict` passes 20/20 with 0 skipped,
four consecutive runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YR2CjVumUbEGWcMikcoXBD
2026-09-16 09:14:01 +00:00

62 KiB
Raw Blame History

Enforcement gates

Reference for this repo's pre-commit and pre-push hooks: what each one guards, what its numbers mean, and which shapes were tried and rejected. Read it when a gate fails, before changing anything in .pre-commit-config.yaml, or before "fixing" something that looks like an inconsistency — several of the oddities documented here are load-bearing and have already been re-litigated once.

AGENTS.md carries only the operative rules an agent needs in the moment. The reasoning lives here.


Running the gates

Command Scope
pre-commit run --all-files the commit-stage hooks
pre-commit run --hook-stage pre-push --all-files the push gate, one command — with one caveat below
pre-commit run skill-size-check --all-files just the ADR-0020 size/context gates

Install hooks via pc-run, wiring all three stages. This repo's .pre-commit-config.yaml has no default_install_hook_types, so a plain install silently skips commit-msg (Conventional Commits) and pre-push (everything below).

The pre-push command reports 10 hooks, not 8. The extra two are pre-commit's own meta hooks, check-hooks-apply and check-useless-excludes: they declare no stages:, so they run at every stage including this one. Both are declared in this repo's .pre-commit-config.yaml like everything else — what separates them is repo: meta (pre-commit's own built-ins) from repo: local. Eight is the count of hooks this repo authors itself.

The caveat: one of those 8 is a silent no-op under that invocation. check-release-needed exits 0 immediately unless PRE_COMMIT_REMOTE_BRANCH equals refs/heads/main, and pre-commit exports that variable only from the real pre-push git hook during an actual git push. Running the stage by hand — or from a CI runner — therefore reports it Passed having checked nothing. That is by design for feature branches — pushing WIP must not be blocked on cutting a premature tag — but it means --hook-stage pre-push --all-files is a full rehearsal of 7 hooks and a skip of the eighth. The script's own header records the same gap for a PR merged through Gitea's merge button, where no local push happens at all.

The pre-push gate

Eight hooks, grouped below by what they guard rather than by the order .pre-commit-config.yaml declares them in.

Core checks

Hook Guards
run-tests bash tests/run-tests.sh --strict — the whole suite, skips fatal (see Tests)

Generated-content drift gates

Hook Guards
check-scope-walkup-sync validate.sh, validate-provenance.sh, new-agent.sh and new-skill.sh's four independent $HOME/.git/apm.yml walk-up ports still agree behaviorally
check-executables-allow-sync root apm.yml's executables.allow key names kyberforge's actual version (see apm gates)

check-executables-allow-sync is the odd one in this group: it guards a silent failure rather than drift in generated text.

Artifact validators

Hook Guards
check-apm-agents-valid runs factory-audit's validate.sh over every real plugins/*/.apm/agents/*.agent.md (see Agent files)

apm's own gates

Hook Guards
apm-audit-ci apm audit --ci once per manifest — root plus each of the six plugin packages
apm-pack-check-clean apm pack --check-versions --check-clean --dry-run — the compiled marketplace still matches what apm.yml + .apm/ would generate, and per-package versions agree with the per_package strategy

Host validators (needs the claude CLI on PATH)

Hook Guards
validate-marketplace claude plugin validate --strict on the root marketplace manifest

Release

Hook Guards
check-release-needed on a real git push to main only — fails if files exposed via .pre-commit-hooks.yaml changed since the last tag. A no-op everywhere else, including under pre-commit run --hook-stage pre-push (see the caveat above)

Two of these shell out to apm: apm-audit-ci and apm-pack-check-clean. The second is a bare apm … entry and the first is a bash -c loop calling apm once per package, so without the CLI the push dies with an unhelpful "command not found". Install with apm-install, or curl -sSL https://aka.ms/apm-unix | sh; verify with apm --version.

Skill and agent context gates (ADR-0020)

The skill-size-check pre-commit hook, scoped to ^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$, runs scripts/skill-size-check.sh. It is also shipped to external repos as kyberforge-skill-size-check (see External consumers). Besides the ADR-0020 gates below, it also asserts required frontmatter is present: name, a non-empty description, and a metadata.version matching three-part semver (1.0.0) — folded in from a formerly standalone skill-frontmatter hook that parsed the same fields with a shell script.

Two things fall outside that scope, both deliberately. The [^/]+/SKILL\.md$ tail admits only a SKILL.md sitting directly in a skill directory under .apm/skills/:

  • the plugins/kyberforge/docs/research/examples/ reference skills, which are vendored upstream corpus and not this repo's to gate;
  • plugins/kyberforge/.apm/skills/skill-author/assets/templates/SKILL.md — inside .apm/skills/, but two directories deeper. It is the FILL IN: scaffold skill-author copies, so its description: > is a comment block rather than a description and every ADR-0020 measurement over it would be meaningless. A reader adjusting the pattern needs to know it is there.

Everything else it matches exactly, with nothing over- or under-caught. Re-derive both halves:

git ls-files | grep -cE '^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$'          # the real skills
git ls-files | grep -E '^plugins/[^/]+/\.apm/skills/.*SKILL\.md$' \
             | grep -vE '^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$'          # the scaffold only

The first count equals the number of skill directories (ls -d plugins/*/.apm/skills/*/ | wc -l); the second returns exactly the template. Every other tracked SKILL.md in the tree is one of the four vendored plugins/kyberforge/docs/research/examples/skill-write/ corpus files, excluded by the .apm/skills/ segment — the first of the two deliberate exclusions above.

Two independent gate families, neither replaced the other

Family 1 — agentskills.io spec backstop (unchanged, conformance not quality):

Constant Value Measured over
MAX_LINES 500 whole file, frontmatter included
MAX_WORDS 2,770 whole file, frontmatter included

Family 2 — ADR-0020 context budget (measured differently, on purpose):

Check SUGGESTION FAIL Measured over
description characters 250 400 the YAML-folded value
body words 600 900 body only — everything after the frontmatter's closing ---

Plus two hard FAILs with no suggestion tier:

  • A missing, valueless or null description:. Not a skip. The description is the one field preloaded into every session, so a gate that declines to measure it reports green. (This is not hypothetical: description: with no value followed by model: sonnet let a line regex capture the next key, which looked non-empty, so the "missing or empty" branch never fired and every gate below early-returned on the genuinely empty folded value — exit 0, zero output, on a blocking gate.)
  • Every references/<file>.md a body names must exist on disk. A dispatch table pointing at a file that was never written is a silently dead branch, and nothing else in the gate/audit/vale stack notices it.

A file can sit well inside one family and fail the other. 2,770 whole-file words is a conformance backstop; 900 body-only words is a quality gate. Conflating them is what produced the current state.

An unresolved routing target is not automatically a FAIL

A boundary-clause target that resolves to no skill or agent has three possible verdicts, not one (unresolved_targets() in scripts/skill-size-check.sh):

Verdict When
SUGGESTION — the default the target does not resolve and neither promotion condition below holds
blocking ERROR the target is written in route notation — /name for any name, or any arrow form (a bare -> name only when the name is hyphenated, a backticked -> `name` for any — see the gap below); or it is a bare terminal name (not a compound modifier) corroborated by another target in the same sentence that does resolve
INFO, "DID NOT RUN" no skill universe could be determined for the path at all — the targets are named and left unchecked, exit 0

The default is deliberately soft because a hyphenated word in a boundary clause is as likely to be a tool, a file format or an English compound as a route: "pre-commit hooks" is prose about a tool and never reaches the check at all, being a compound modifier rather than a terminal name. The SUGGESTION text says how to opt in — write it as /name or -> name and it gets checked properly.

The two promotion conditions are not symmetric, and the order matters. _add() decides notation first: when the name is written /name, or reached through any arrow form, the target is marked error-eligible there and the terminal test is never run. Terminality gates only the bare path — a name in prose earns its error from corroboration, and a compound modifier can never dangle. Reading the row as "terminal AND (notation OR corroborated)" gets the notation half backwards: it predicts that … Do not use for Y — use /no-such-skill afterwards. is a SUGGESTION, because afterwards is a follower outside FOLLOWER_OK. It exits 1. That was the defect — -> name reached _add() with strict=True from both its call sites and /name did not, so the one spelling ADR-0020 offers an author who wants a route checked unconditionally was the one spelling a stray follower could silence.

Known gap: a BARE arrow target must be hyphenated. Target extraction is built on NAME_HYPH in scripts/skill-size-check.sh, which requires at least one hyphen, and ARROW_BOUNDARY inherits that. So Not X -> gitea-prs is extracted and checked, while Not X -> triage yields no target. The exclusion is deliberate, not an oversight: research, triage, forge, prototype and tdd are all real skill names and ordinary English, so a bare single-word rule would flag most of the corpus. The marked spellings carry no such restriction — `triage` and /triage are both extracted — and are the forms to prefer. Both arrow spellings are recognised: ARROW_MARKED, ARROW_BOUNDARY and BOUNDARY_ARROW are each built from (?:->|→), so the unicode arrow → behaves exactly like -> in every case below. Cite these constants by symbol name, never by line number: the script moves often enough that a pinned line lands a reader in an unrelated comment block and reads as plausible.

The gap is no longer silent. It used to be exactly that — no ERROR, no SUGGESTION, exit 0 — which made the dangling-target SUGGESTION's own advice unsafe for a single-word skill: taking it silenced the finding instead of checking it. boundary_clause_status() now separates the case out and reports it as unparsed (see below), naming the parse failure and the two spellings that fix it. The target is still not resolved; the author is now told so rather than left with a green gate. tests/test-adr0020-targets.sh covers both directions (arrow-single-word-target and the silent control arrow-single-word-marked).

Corroboration is what makes the soft default safe: a sentence whose other target resolves is demonstrably a routing sentence, so a sibling that does not resolve is a typo rather than a noun, and gets promoted.

Target resolution walk

Resolution walks up from the file being checked — never from the script's own location. Deriving it from ${BASH_SOURCE} leaked holocron's own skill universe into every consumer repo running the hook through pre-commit, so a consumer skill routing to a holocron skill (skill-audit at the time, now factory-audit) resolved against a plugin it had never installed.

The walk finds an authoring root: the nearest ancestor holding plugins/*/.apm/skills or plugins/*/.apm/agents, falling back to the nearest ancestor holding .git. Two passes, not one interleaved walk, so a nested .git (a submodule, a sub-package worktree) cannot beat a real monorepo root further up.

The universe is then:

  1. every skill and agent under <root>/plugins/*/ — sibling plugins resolve, which is what a monorepo means;
  2. the checked file's own apm package;
  3. the packages that package declares in its own apm.yml dependencies.apm.

The root manifest's dependencies: block is not read, and no plugin here declares a cross-plugin apm dependency — none needs to.

Deployed .claude/ / .agents/ trees are consulted only when the walk found no plugin monorepo root, whether it landed on a bare .git ancestor or on nothing at all. That is the consumer case.

The gate keys on which of the two passes matched, never on whether the root contributed a new name. A name-count delta looks equivalent and is not: _collect_authoring_root() re-collects the checked file's own plugin, whose names the earlier steps already added, so a single-plugin monorepo shows a delta of zero and would wrongly reach for the deployed trees — including the user's global ~/.claude/skills, making the verdict depend on what happens to be installed.

Why it matters: those trees are gitignored apm install output, present only on a machine that has run it. Four cross-plugin targets here (gitea-branches → git-branches, gitea-branches → git-history, gitea-issues → git-branches, gitea-workflow → git-workflow) once resolved through .claude/skills/ alone, so the same commit measured 2 dangling targets on a developer machine and 6 on a fresh clone. A gate shipping hot with no baseline cannot give two answers.

Verified: running the hook over a tree holding only plugins/ and the root apm.yml, with no .claude/ or .agents/ anywhere, produces findings identical to the working tree — confirming the two trees agree on the current corpus, independent of what happens to be installed locally.

Boundary-clause detection: three outcomes, not two

boundary_clause_status() returns one of three values, and the two findings get separate messages:

Status When Reported as
present a prose marker (do not, instead, rather than, not for) or an arrow clause was found nothing
absent neither was found SUGGESTION: add a boundary clause, in either form
unparsed an arrow clause was found and no target could be read out of it SUGGESTION: the clause is present — this is a parse failure, not a missing clause

The third had to be split out. Collapsing it into absent is a wrong finding, not a strict one: it sends the author to add a clause that is already there. Three of them instead reworded a correct clause until the regex accepted it, one stripping the very filename that discriminates the skill from its neighbour (#110).

unparsed is narrow and certain on purpose. It fires only on the arrow form, which always names a target, so zero targets means the name is written in a shape the extractor cannot see — in practice a bare single-word target, per the known gap above, and the message says to write it `name` or /name. A prose clause yielding no target is not reported at all: "Do not use for anything else" is a complete and legitimate boundary clause that names nowhere to go.

One arrow, one target. An arrow clause naming two or more targets draws its own SUGGESTION, quoting both names and asking for a split, because only the first is ever resolved: the conjunction continuation (CONT_MARKED / CONT_ANY) is wired to the prose route verbs and never to arrows. So Not X -> a or b resolved a, left b resolved by nothing and reported by nothing, and then let the audit print "1 of 1 boundary target(s) resolve" on a clause naming two — a gate under-reporting its own coverage, which is the one failure mode ADR-0020 says a gate must not have (#107). The clause is rejected rather than the arrow scan extended: extending it would widen the resolver's deliberately conservative false-positive tuning across every arrow in the corpus, where splitting costs the author one full stop. The convention is one arrow per target — Not X -> a. Not Y -> b. — already what every retrofitted gitea-* skill does in practice, now stated in skill-author's references/contract.md instead of being folklore.

Dotted filenames in a boundary clause now parse. CLAUSE_BODY — what may sit between Not and the arrow — used to be [^.;], a class that cannot cross a ., so every clause naming a dotted filename between the two (AGENTS.md, .vale.ini, .pre-commit-config.yaml) was invisible to both BOUNDARY_ARROW and ARROW_BOUNDARY. The two resulting failures were different sizes (#110):

  • with a backticked target the clause was misdiagnosed. The backtick sweep still extracted the target, so the route was checked, but the gate reported "no boundary clause" on a clause that was present and working. That is the misdiagnosis the three rewordings above came from.
  • with a bare target the clause was unchecked. ARROW_BOUNDARY is the only extractor for a bare arrow target, so Not AGENTS.md -> no-such-skill produced no target, no dangling report and no missing-clause SUGGESTION. Silence, not noise — the worse of the two.

CLAUSE_BODY is now (?:[^.;]|\.(?=\S)): a dot inside a filename is followed by a non-space, a sentence-ending dot by whitespace or end of string, so the class crosses AGENTS.md and still stops at a real sentence end. Read the second bullet forward as well as back: a bare target sitting after a dotted filename is now extracted, resolved, and a blocking ERROR when it dangles, where the same clause used to pass unchecked in silence.

SUGGESTION-only checks

Deterministic to measure, judgment to act on:

  • a description with no boundary clause at all (absent);
  • an arrow clause whose target could not be read (unparsed);
  • an arrow clause naming more than one target;
  • a ## Gotchas section with more than five entries;
  • a ## Gotchas section over 25% of the body.

Hand-invoked skills are exempt from the routing rules, and only those

A skill or agent whose frontmatter carries disable-model-invocation: true skips three checks:

  • the boundary-clause check, absent and unparsed alike;
  • the multi-target arrow check;
  • the 250-character description target (hand_invoked() in scripts/skill-size-check.sh).

It keeps the 400-character description FAIL and both body word tiers, and if its description does happen to name a target, that target is still resolved and can still dangle.

Why the exemption is right: disable-model-invocation: true removes the skill from the model-visible listing entirely — it is not preloaded, and the Skill tool refuses to call it — so its description is never matched against user intent. ADR-0020 and skill-author's contract therefore give such a skill one plain human-facing sentence: no trigger list, no boundary clause. No validator knew the field existed (#108), so the boundary-clause SUGGESTION fired on exactly the shape the contract mandates, and its remedy — "so the router knows where NOT to send this skill" — was addressed to a router that cannot see the skill at all. An author who followed the advice made the file worse. There is no router to inform.

The half that does not lift is the point. The body is still loaded on invocation and still competes with the caller's live conversation, so neither body tier moves. The 400-character ceiling stands too: a hand-invoked description is not preloaded, but it is still the one line the user reads when choosing from the / menu, and that ceiling is an outlier stop rather than a routing-quality budget — which is precisely why the 250-character target is the tier that lifts.

The field is read as a boolean, not as a mention of the key. PyYAML already resolves the unquoted YAML 1.1 booleans, so the extra handling catches a quoted "true", which a host reads as truthy; disable-model-invocation: false is the model-invoked case written out longhand and buys nothing. A frontmatter parse failure returns false rather than raising — the flag is a modifier on other checks, and description_value() on the same text already reports the broken frontmatter, so raising here would diagnose one file twice two different ways.

caveman and zoom-out are the two carriers here. tests/test-skill-size-check.sh pins both halves — what the carve-out lifts, each with a flag-removed control, and what it must not.

verbose: true is load-bearing

The hook is declared verbose: true so the SUGGESTION tier is audible. pre-commit prints nothing at all for a passing hook, and a SUGGESTION deliberately does not fail — without verbose every suggestion is swallowed, which is exactly the invisibility ADR-0013 records for Vale warnings. ADR-0020's preload arithmetic depends on it: writing to the 400-char FAIL delivers roughly half the cut that writing to the 250-char SUGGESTION does, so the intended saving depends entirely on that tier being visible. The numbers, and the measurement method behind them, are not restated here — they live in ADR-0020's Consequences section, under "A ceiling does not produce an average", whose figures are pinned to the base commit the decision was taken on (f9b919d). Quoting them here would just create a second copy to go stale. It costs nothing on a clean file — the script prints only findings.

Duplicated constants

factory-audit's validate.sh holds a second copy of the four ADR-0020 constants (DESC_SUGGEST_CHARS / DESC_MAX_CHARS / BODY_SUGGEST_WORDS / BODY_MAX_WORDS) — the two description constants apply to both artifact types it handles, the two body constants only to skills. They are copied rather than imported because a cache-installed plugin's scripts cannot read files outside their own plugin directory. tests/test-skill-size-check.sh asserts the copies agree, so drift fails CI rather than silently letting an audit bless a skill the commit hook then rejects.

The shared boundary resolver is now two copies, not three (ADR-0025). scripts/skill-size-check.sh still carries it embedded between BEGIN/END ADR-0020 SHARED BOUNDARY RESOLVER markers; the two plugin copies that used to sit inside skill-audit's and agent-audit's validate.sh collapsed into the single factory-audit/scripts/lib-boundary-resolver.sh, sourced by that skill's scripts. The two remaining copies must still stay byte-identical — a plugin script cannot source the root one, which is the constraint that forces a copy to exist at all.

python3 and PyYAML are hard requirements

Both, and neither is a best-effort accelerator.

python3 because the script measures the folded description value. Most descriptions here are >-block scalars, so a regex over the raw lines measures indentation and newlines instead of the value. Missing it fails the hook with an install pointer rather than skipping the ADR-0020 checks, which would be a vacuous green. In practice it is already present — pre-commit is itself a Python application.

PyYAML because the hand-rolled fallback frontmatter reader has been removed deliberately. It disagreed with a real parser across the FAIL boundary — one corpus description measured 270 characters parsed and 412 unparsed — and a quoted "description" key or an explicit description: null returned empty from it, silently skipping the description and routing checks. A reader that mis-parses an unfamiliar scalar shape reports a clean pass on a file it never measured, which is the exact vacuous-green failure the python3 check exists to avoid. pip install pyyaml (or python3 -m pip install PyYAML, or the distro's python3-yaml) if the hook reports it missing.

Neither requirement generalises to every hook in this repo. check-rtk-prefix needs python3 but not PyYAML: it reads the markdown body and never touches frontmatter, so it has no scalar to fold.

Agent files take the description gates, not the body gate

check-apm-agents-valid runs factory-audit's validate.sh over every real plugins/*/.apm/agents/*.agent.md. It derives its expected file set from git ls-files — the pattern tests/run-bats.sh established — so an agent file deleted from the worktree but still tracked fails the run, and discovering zero agent files is an error, not a pass. An untracked new agent file is still validated: the derivation is one-directional on purpose, so uncommitted work is not blocked but also cannot bypass the gate.

The hook exists because validate.sh was previously exercised only by check-scope-walkup-sync, against synthetic mktemp fixtures — it had never run against the agent files it governs. That is how ADR-0016 could be amended to bless a disallowedTools frontmatter field while validate.sh's allowlist still rejected it: spec and enforcer disagreed and every gate stayed green.

Agents take the ADR-0020 description gates (factory-audit's validate.sh holds its own copy of those two constants) and, deliberately, no body word gate. A skill body is loaded into the caller's context and competes with the live conversation; an agent body becomes the system prompt of a fresh context. The rationale for the 900-word FAIL does not transfer. A bats test pins that absence for the agent path of factory-audit's validator — adding a body gate there contradicts the ADR rather than fixing an inconsistency. The merge did not change this: the validator auto-detects the target type, and the body gate applies on the skill path only.

Be precise about the scope of that guarantee: it holds for the validator, not for the shared script. scripts/skill-size-check.sh applies its body gate to whatever path it is handed, and

bash scripts/skill-size-check.sh plugins/*/.apm/agents/*.agent.md

exits 1 today with 900-word body FAILs on git-orchestrate and gitea-orchestrate. (Counts are deliberately not pinned here — agent bodies are edited like any other file, and a figure in this paragraph goes stale the moment one is trimmed. Run the command.) Agent files escape only because the hook definitions filter on SKILL.md — a file-pattern accident that happens to implement the design, not the design itself. Do not "extend" that hook's files: pattern to cover agents on the assumption that the script already knows the difference; doing so silently enforces a gate ADR-0020 declines to set.

Current retrofit status

The ADR-0020 gates ship hot, with no baseline file — a shrinking baseline was considered and rejected. The corpus is currently clean on both: 0 of 38 descriptions/bodies exceed their FAIL tier, 0 dangling targets, 0 Kyberforge.CompositionNote (Vale) errors. History: issue #99.

Nothing is grandfathered — a new skill, or an edit that crosses a FAIL tier, is blocked on first commit. SUGGESTION counts are not pinned here; they move with every edit. Measure and check both gates before starting work on a skill:

bash scripts/skill-size-check.sh plugins/*/.apm/skills/*/SKILL.md | grep -c '^SUGGESTION'
pre-commit run --all-files    # size AND Vale — skill-size-check alone can pass while Vale still blocks

The rtk prefix gate (ADR-0023)

check-rtk-prefix is a repo: local pre-commit hook running scripts/check-rtk-prefix.sh over ^plugins/[^/]+/\.apm/(skills/.*\.md|agents/.*\.agent\.md)$, with README.md excluded. It enforces ADR-0023 clause 1 and nothing else: an executable, instructed local git command in plugin skill or agent content is written rtk git.

It is wider in file scope than the ADR-0020 hooks — every markdown file under a plugin's .apm/skills/ and .apm/agents/, not SKILL.md alone — because the rule it enforces is about commands an agent runs, and most of those live in references/, which the ADR-0020 gates do not reach (the references/ blind spot).

What it can decide, and what it declines to

ADR-0023 has three clauses and only the first is a pattern:

Clause Rule Gated
1 executable + instructed → rtk git yes
2 illustrative / referential → bare git no — undecidable
3 machine-parsed or interactive → bare git no — opt-out marker

Clause 2 is a judgement about what a sentence is doing. "Run git switch <branch>" and "git switch refuses rather than clobbering local edits" are the same token sequence. A gate that guessed would fire on correct prose, and a gate that fires on correct content gets added to SKIP — which disarms clause 1 along with it. So the hook looks only at the two contexts where a git mention is unambiguously an instruction to execute:

  • a line inside a fenced code block whose info string names a shell — bash, sh, shell, zsh, console, shell-session. Fences tagged text, yaml, json, or tagged with nothing, are not checked;
  • the opening backticked span of a "Run" column cell in a markdown dispatch table, and only the opening span.

That last narrowing is not fussiness. A Run cell routinely carries a command followed by prose about it, and the prose is clause 2. git-worktrees/SKILL.md has both shapes on adjacent rows — one cell reading `rtk git worktree add --track …` — always correct. `git worktree add <path> <branch>` expands to exactly this (instruction, then reference), and a **Never** … row whose Run cell is entirely explanation containing a bare git push. Checking every backticked span flags both; checking only a leading span flags neither, and still catches the ordinary | List | `git worktree list -v` | case the gate exists for.

The clause-3 opt-out

A command that is deliberately bare — because rtk rewrites the output the skill parses, or because the command is interactive — is exempted by putting the literal string ADR-0023 on the same line: in a shell comment for a code line, in the cell text for a table row.

Per line, never per block. A fenced procedure routinely mixes rtk git steps with one deliberately bare command (git-remotes/references/push.md does exactly that), and a block-level marker would silently disarm every checked line around the marked one. The cost is a repeated # bare per ADR-0023 in the three blocks of git-log-format.md where every line is deliberately bare; that repetition is the price of the marked line being the only line the marker speaks for.

The marker is a plain substring match, so a line that mentions ADR-0023 for an unrelated reason is also exempt. Accepted deliberately: the marker records an author's opt-out, it is not a security boundary, and a stricter form would only move the same trust to a different string.

What it deliberately does not cover

  • Clause 2. Nothing checks that an illustrative mention stayed bare. A sweep that re-prefixes a referential git passes this gate. The inline reasons ADR-0023 requires on clause-3 sites are the only defence, and they are prose.
  • Prose bullets. Most of branch-operations.md, merging.md and rewrite-history.md instruct in list items, not fences. Those are clause-1 sites the gate cannot see, because it cannot distinguish them from clause-2 mentions in the same list.
  • README.md, excluded by pattern. A skill-directory README is consumer-facing prose no agent loads, and the git clone https://github.com/bats-core/… lines in the six tests/README.md files are setup instructions for a third party who has no rtk. Prefixing those would be actively wrong, not merely noisy — see ADR-0023's consumer section.
  • Quoting. The line splitter breaks on ;, |, &&, ||, $( and backticks without tracking quotes, so a git command inside a quoted argument is decided by accident. rtk git submodule foreach 'git pull origin main || :' passes because the segment holding the inner command begins with rtk — the right answer for the wrong reason. Write foreach 'git a; git b' and the second inner command is a false positive needing the marker. ADR-0023 records this shape as one the rule itself does not decide.
  • Non-git commands. Only git is checked. rtk fronts gh, docker, kubectl and others; no gate covers those, and the corpus does not currently instruct them.

tests/test-check-rtk-prefix.sh pins all of it, including the false-positive cases. Its first case reconstructs the plugin corpus as it stood on main before the #113 sweep and asserts the gate fails there with at least 20 findings, one of them the gitea-* git remote get-url origin drift the sweep missed — a gate that only passes on the already-fixed tree proves nothing about the drift it was written for.

Vale

Install the vale binary — brew install vale (macOS), snap install vale (Linux), choco install vale (Windows), or see https://vale.sh/docs/vale-cli/installation/. No vale sync is needed: the Kyberforge styles are committed under plugins/kyberforge/.apm/skills/factory-audit/assets/vale/styles/, not downloaded packages (ADR-0014).

One copy, one config

Wiring Vale as a deterministic prefilter for factory-audit's Description dimension (motivation: issue #84) is repo-specific, not part of the generic lint plugin, so it does not live in plugins/lint/ — and per ADR-0014 it no longer lives at the repo root either. It lives once, under plugins/kyberforge/.apm/skills/factory-audit/assets/vale/, carrying both the Kyberforge and KyberforgeCopilot styles and a single .vale.ini with all three glob sections: [**/SKILL.md], [**/agents/*.md], [**/*.agent.md].

ADR-0014 split this into two skill-scoped copies because a plugin's cache-install copies only each skill's own files and skill-audit could not reach across the skill boundary into agent-audit's copy. Merging the two audit skills removed the boundary, so the copy went with it and the single-file .vale.ini ADR-0014 split apart is restored (ADR-0025). KyberforgeCopilot stays scoped to .agent.md files alone, for the Copilot-only "Use proactively has no effect" check.

With one copy there is nothing left to diff, so the check-vale-style-sync pre-push hook, scripts/check-vale-style-sync.sh and tests/test-check-vale-style-sync.sh are deleted — one hook off the push gate. Read what went with it, not just what became vacuous. The script had 17 assertion sites. ADR-0025 maps each one; the short version follows.

Genuinely moot (6):

  • the vale-wrap.sh diff and the styles/Kyberforge/ diff, which compared two copies that are now one;
  • the four hard-fail guards that located those copies (REPO_ROOT is a directory, the .apm/ paths are not stale, neither copy is missing).

Its StylesPath and BasedOnStyles checks were not diffs. They were per-file greps of each .vale.ini, so they survive: case 0 below checks that the config loads, and case 28 checks that the Kyberforge style is actually loaded.

Rehomed or ported (11). The largest group is the six-row glob-coverage probe table, which invoked vale --config on one representative path per file shape. It was the only assertion anywhere that catches a .vale.ini glob typo ([**/SKILL.md] → [**/SKILLS.md]), the failure mode where every other check stays clean while Vale lints zero files. One config does not make that impossible: a typo in any one of the three sections still 0-file-skips that shape.

Case 0 runs before any Vale-dependent case and needs no Vale binary. It asserts that the shipped .vale.ini exists and is readable, sets a StylesPath that resolves to a directory, and names only styles that ship. A config that cannot load used to surface as nine generic "vale printed no summary line" failures across cases 28–31. It now fails once, names the cause, and holds the Vale-dependent cases back.

The probes now live in tests/test-vale-wrap.sh (cases 28–30), rehomed against the merged config: one representative path per file shape, each asserted to produce a Vale scan of more than zero files and a Kyberforge alert (case 28). Case 28 also checks that each probe path is in scope of a published vale hook, and that every .vale.ini section has a probe row. Its Part B drops Kyberforge from each section's BasedOnStyles in a copy and requires that section's probes to fail as "style not loaded". Case 29 is a mutation case: it typos each section in a copy of the assets and requires that section's isolating probes to drop to zero. Case 30 asserts that KyberforgeCopilot reaches .agent.md files alone. Its Part B requires both an unload (dropped from [**/*.agent.md]) and a leak (added to [**/SKILL.md]) to fail. Case 31 is the third class that went with the script and is not a glob probe at all: the per-rule override allowlist, which pins every Kyberforge rule at a blocking bare YES/error. It is not redundant with the probes above — those key on Kyberforge.VagueWording and KyberforgeCopilot.ProactivePhrase, so the other four rules (DescriptionOpener, PaddingPhrase, SentenceOpenerThereIs, CompositionNote) can each be overridden out of error underneath a passing probe. That gap is closed.

Two cases cover the hook manifests.

Case 33 is the original's cross-manifest files: drift check, ported. It extracts each vale hook's files: regex from .pre-commit-hooks.yaml and from .pre-commit-config.yaml independently, compares them per hook and never as a union, and asserts that each shared probe path is in scope of both or neither. The original selected each hook's record by matching entry: against the owning skill's vale-wrap.sh path. After the merge both hook IDs share one entry:, so the port pairs them by id: from an explicit table: kyberforge-vale-audit-skill ↔ vale-audit-prefilter-skill, and kyberforge-vale-audit-agent ↔ vale-audit-prefilter-agent. A missing hook id or a class with no shared probe fails by name. Part B requires three mutations to fail: the skill hook narrowed to one plugin, the agent hook narrowed the same way, and a renamed local hook id.

This was briefly a real hole. Narrowing vale-audit-prefilter-skill from ^plugins/[^/]+/... to ^plugins/kyberforge/... left 6 of 38 skills prefiltered, and the whole suite green, before case 33 existed.

Case 32 covers the separate zero-match question on the local manifest alone. Each .pre-commit-config.yaml vale hook's files: regex must still match at least one tracked file, and every path it matches must be in that hook's own artifact class. A hook narrowed to zero files never runs, and pre-commit reports no error.

What Vale owns, and what stays LLM judgment

Six rule files, six distinct rules:

Rule Vale scope Bans From
Kyberforge.DescriptionOpener text.frontmatter.description non-imperative openers ("This skill/agent…") issue #84
Kyberforge.VagueWording text.frontmatter.description vague capability wording ("helps with", "utilize", …) issue #84
Kyberforge.PaddingPhrase text generic "see references/ for details" padding issue #84
KyberforgeCopilot.ProactivePhrase text.frontmatter.description Use proactively (no effect in Copilot) issue #84
Kyberforge.SentenceOpenerThereIs sentence "There is/are" sentence openers ADR-0013
Kyberforge.CompositionNote text.frontmatter.description architecture and composition prose in a description ADR-0020

Vale covers the pattern-matchable sub-checks named in issue #84 plus, per ADR-0013, one cherry-picked body-wide prose-pattern rule. Everything else stays LLM judgment: defaults-vs-menus, why-rationale, the non-pattern-matchable body-discipline calls, near-miss exclusion strength, and control calibration. New rules land directly in styles/Kyberforge and block immediately — there is no trial tier.

The cherry-pick record, so it is not re-litigated:

  • Kyberforge.SentenceOpenerThereIs landed — 22 held-out hits, both in-corpus hits clean rewrites, zero suppressions needed.
  • Kyberforge.VagueQualifier was cherry-picked and then deleted. 2 hits across the corpus as it stood on 2026-08-08 (before the .apm/ restructure): one marginal, and one unfixable false positive — caveman/SKILL.md quotes of course as an example of filler, a mention rather than a use — which forced the repo's only Vale suppression comments.
  • governance.md and CONTROLS.md were evaluated as rule sources and excluded: nothing prose-pattern-matchable to mine.

Why every rule is level: error

Every alert is a FAIL, with no ignorable tier — same all-or-nothing model as shellcheck, the test suite, and conventional-pre-commit. Graded severities do not work here: Vale's exit code keys on error alerts alone, so a warning or suggestion rule exits 0, and pre-commit swallows a passing hook's output. Such a rule would be invisible and would block nothing.

MinAlertLevel and --minAlertLevel are correspondingly absent from both .vale.ini files and from the hook definitions. Under this model they are no-ops; adding one is not a missing knob.

The verbose: true escape hatch that makes skill-size-check's SUGGESTION tier audible has no analogue here — Vale has no tier to make audible.

External consumers: the root .pre-commit-hooks.yaml

The root .pre-commit-hooks.yaml exposes two Vale hook IDs (kyberforge-vale-audit-skill, kyberforge-vale-audit-agent) plus kyberforge-skill-size-check, so any external repo can enforce the same rules with repo: <this-repo-url>, rev: <tag> in its own .pre-commit-config.yaml. pre-commit clones the pinned rev into its own cache, independent of whether Claude Code or the kyberforge plugin is installed at all; the same mechanism covers CI via pre-commit run --all-files. skill-size-check has no external asset dependency, so it needed no relocation under ADR-0014 — only exposure.

The two IDs survive the merge even though they now point at the same wrapper. Both kyberforge-vale-audit-skill and kyberforge-vale-audit-agent keep their IDs and their files: regexes, because an external repo pins them by name in its own .pre-commit-config.yaml and collapsing them to one would break every such consumer silently. What changed is only the entry: target: both now name factory-audit/scripts/vale-wrap.sh.

This repo's own vale-audit-prefilter-skill / -agent hooks consume the identical plugin-bundled copy via repo: local. Deliberately not a second root copy, and deliberately not a pinned self-reference — a pinned self-reference would lint working-tree edits against the last tagged release rather than against the change being made.

Pre-commit

Two prefilter hooks, with .apm/-scoped files: patterns:

Hook Pattern
vale-audit-prefilter-skill ^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$
vale-audit-prefilter-agent ^plugins/[^/]+/\.apm/agents/[^/]+\.agent\.md$

Only the authoring source triggers them. The skill-author scaffold template and the four vendored research-corpus SKILL.md files match neither pattern (see Skill and agent context gates, which uses the same scope for skill-size-check), so prose findings surface only when you edit a file this repo actually authors. Without the binary the hooks fail with a bare "command not found" and no install pointer.

Two hooks, not one combined hook — for a different reason than ADR-0014 gave. The original reason was mechanical: with a config per skill, a single hook could point at only one copy and would silently 0-file-skip the other file shape (see A 0-file Vale run is NOT RUN). One .vale.ini carrying all three sections removes that constraint. The split stays anyway because the two IDs are an exported contract external consumers pin by name, and because the files: regexes still have to differ — each hook hands Vale only the file shape it is scoped to.

The .vale.ini globs do no scoping

The .vale.ini's section globs are path-agnostic — [**/SKILL.md], [**/agents/*.md] and [**/*.agent.md] — and constrain filename shape, not location: Vale's * crosses /. A SKILL.md outside plugins/ (a project-scope .claude/skills/foo/SKILL.md, say) still matches [**/SKILL.md] and gets linted normally.

All scoping therefore comes from the pre-commit hook's own files: regex and from factory-audit passing one explicit file per invocation. The two manifests scope differently on purpose:

Manifest -skill -agent
.pre-commit-config.yaml (pins this repo's layout) ^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$ ^plugins/[^/]+/\.apm/agents/[^/]+\.agent\.md$
.pre-commit-hooks.yaml (layout-agnostic for consumers) (^|/)SKILL\.md$ (^|/)agents/[^/]+\.md$|\.agent\.md$

Narrowing a .vale.ini glob to a plugins/-shaped path to "tighten" it breaks the consumer case. check-vale-style-sync's probe set was built to catch exactly that; it moved to tests/test-vale-wrap.sh with the hook's deletion, and two of the six probes exist specifically to pin this location independence — see One copy, one config.

The blind spot: references/ is unlinted, for two independent reasons

Every references/*.md file in the corpus is outside the prose gate. Count them with git ls-files | grep -cE '^plugins/[^/]+/\.apm/skills/[^/]+/references/.*\.md$' rather than reading a figure here; it moves with every retrofit. This is the gap that matters most, because the context contract's own remedy for an over-long body is to move prose into references/ — the gate pushes text across its own boundary and then stops watching it. factory-audit is the live example. Its dispatch body keeps only the gotchas common to both flows, and the flow-specific gotchas live under ## Gotchas in references/skill-flow.md and references/agent-flow.md (ADR-0025). Handing both flow files to vale-wrap.sh prints 0 errors … in 0 files and exits 0.

Closing either cause alone changes nothing. There are two, and they are independent:

Cause Where Effect on a references/ file
the Kyberforge style is scoped [**/SKILL.md] factory-audit/assets/vale/.vale.ini matches no section, so Vale lints 0 files and exits 0
the hook's files: regex is ^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$ vale-audit-prefilter-skill in .pre-commit-config.yaml the file is never handed to Vale at all

Verified both ways. Handing factory-audit's vale-wrap.sh a reference file directly — bypassing pre-commit entirely, so only the style scope is in play — prints 0 errors … in 0 files and exits 0, where the same wrapper on a SKILL.md reports in 1 file. And the hook's files: regex, applied to git ls-files, selects only the skill-directory SKILL.md files scoped at the top of this page, so pre-commit never hands Vale a reference file to begin with. Widening the glob to [**/*.md] would still lint nothing through the hook; widening the hook's files: alone would hand Vale files its own config declines to match, which is the 0-file NOT RUN shape — a green run that measured nothing. Issue #117 records the style-scope half; the hook half has to land in the same change or the fix is cosmetic.

The consumer manifest is a third axis and does not rescue this either: .pre-commit-hooks.yaml's (^|/)SKILL\.md$ is layout-agnostic but still filename-shaped, so an external repo running kyberforge-vale-audit-skill has the same gap.

vale-wrap.sh, never bare vale

factory-audit's Step 1 and both pre-commit hooks call factory-audit/scripts/vale-wrap.sh, not vale. It works around a confirmed Vale 3.15.2 limitation: text.frontmatter.description silently stops matching on most — not all — multi-line descriptions.

Verified by reproduction on a deliberately-bad fixture, not assumed:

Description scalar spanning 2+ lines Vale's behaviour
> folded block 0 alerts, exit 0 — broken
plain (unquoted) continuation lines 0 alerts, exit 0 — broken
single- or double-quoted, wrapped 0 alerts, exit 0 — broken
| literal block alerts fire, exit 1 — lints normally

The wrapper flattens the three broken forms to a single-line scalar in a scratch copy — or, for the rare value no inline scalar can spell verbatim, a |- block with one content line — padding with blank lines so every other line number is unchanged. | literal blocks and single-line descriptions pass through untouched. Most descriptions in this repo are > blocks, so before the wrapper a bad description in any of the three broken forms sailed straight through the prefilter.

The --config argv defect

Handed no --config at all, the wrapper falls back to its own sibling assets/vale/.vale.ini, located from ${BASH_SOURCE[0]} rather than from the cwd. That is why both manifests' entry: is now the bare script path with no argument after it.

pre-commit prefixes only entry[0] with the hook-repo clone path (cmd = (prefix.path(cmd[0]), *cmd[1:])), so every later argument resolves against the consuming repo's root. A --config in .pre-commit-hooks.yaml therefore pointed at a path no consumer has and hard-failed every external run with E100 [--config] Runtime error.

.pre-commit-config.yaml drops the argument too, deliberately keeping the two entries identical. The local repo: local hook resolved its --config correctly only because the consuming repo was this repo — and that divergence is why three review rounds exercised a path no external consumer takes and missed the defect. Do not reintroduce a --config to either manifest to make the local run "explicit".

An explicit --config from any other caller still wins, in all three argv forms (--config X, --config=/abs, --config=rel), and a relative one resolves against the caller's cwd — matching bare vale, not the repo root.

Both audit skills' Step 1 passes no --config either. Step 1 resolves the script relative to the skill's own directory so the call works from an installed plugin cache; a relative --config alongside it would resolve against the cwd instead, yielding E100 Runtime error … does not exist and exit 2 — which both skills' fallback misreads as "vale unavailable" and silently downgrades to full LLM judgment.

tests/test-vale-wrap.sh regression-tests this against factory-audit's copy — the only one left. Its fixtures are all SKILL.md-shaped, and that copy's .vale.ini carries the matching glob section along with the two agent ones.

A 0-file Vale run is NOT RUN

Vale reports 0 files only when the path it is handed matches no glob section at all — a differently-named file, or a directory argument holding nothing that matches. That run prints

✔ 0 errors ... in 0 files.

and exits 0, indistinguishable from a clean pass. Both audits therefore treat a 0-file Vale run as NOT RUN and fall back to full LLM judgment rather than reporting the Description dimension clean.

Pre-push

vale is still a pre-push dependency, but no longer through a hook of its own. check-vale-style-sync — the hook that ran the six glob probes, and whose CHECK_VALE_STYLE_SYNC_ALLOW_MISSING_VALE=1 opt-out downgraded them audibly rather than skipping the hook — is deleted with the second Vale copy (ADR-0025). The six glob probes survive it inside test-vale-wrap.sh, so run-tests --strict is now the gate that runs them. That is also what keeps vale a pre-push requirement: test-vale-hooks-consumer.sh exits 77 without the binary, and so does test-vale-wrap.sh once its static cases pass, and a skip fails the push.

test-vale-wrap.sh without Vale skips only its Vale-dependent cases, not the whole suite. The cases that are plain greps and awk over the config and the two hook manifests still run: case 0, 16, 26, 27, the static halves of 28, 31 Parts A and B, 32 and 33. A static failure exits 1, because a real defect is not a setup error. Only an all-static-pass run exits 77.

Mentioning banned phrasing without tripping the rule

House convention: banned phrasing that must be mentioned rather than used goes in backticks or a fenced code block. Vale skips code spans and fences, so no suppression is needed — which is why this document quotes Use proactively and "There is/are" the way it does.

Inline <!-- vale Rule = NO --> is the fallback only where backticking is impossible. Use the HTML-comment form; the MDX {/* */} form does not work in plain Markdown. The one time a rule forced suppression comments, the rule was deleted instead (see the VagueQualifier entry above).

Tests

bash tests/run-tests.sh              # every test-*.sh plus the bats suite
bash tests/run-tests.sh --bats-only  # just bats

First run auto-initializes the bats submodules; no manual git submodule update needed.

Exit 77 = SKIPPED. A suite that skips because a dependency is missing does not fail an ad-hoc run. The pre-push hook invokes the same script as --strict (RUN_TESTS_STRICT=1 is equivalent), where a skip does fail the push: at pre-push a skip means one of the documented dependencies is absent on this machine, so the gate would otherwise report success having run fewer suites than it appears to. Without --strict the gate once went green having verified 15 of 17 suites on a vale-less PATH, with the skip list swallowed. Without vale, two suites skip — test-vale-hooks-consumer.sh and test-vale-wrap.sh — and the strict failure names each one and what to install. (It was three until test-check-vale-style-sync.sh was deleted with its hook; see One copy, one config.)

Output assertions use a here-string, never a pipe. Write grep -q PATTERN <<< "$OUT", not echo "$OUT" | grep -q PATTERN. Under set -o pipefail the pipe form fails depending on timing: grep -q exits on its first match, echo takes SIGPIPE on its next write, and pipefail reports that as the pipeline failing, so output that matched reads as "no match". It showed up as a push gate that failed about once in 670 runs, on a different suite each time. tests/test-no-pipefail-early-exit-grep.sh scans every tracked shell file that sets pipefail and fails on the pipe form. It does this for echo or printf piped into grep with -q, -m, -l, -L, --quiet or --silent. It checks its own scanner against fixtures before trusting a clean result. Pipes from other commands are out of scope. In practice they either absorb the writer's exit status with || true or write only once, at exit.

tests/run-bats.sh derives the set of .bats files it expects from git ls-files, so a .bats file deleted from the worktree but still tracked in the index fails the run rather than silently shrinking the suite. Remove one with git rm (or stage the deletion) when intentional; an untracked new .bats file is picked up and needs no ceremony.

Both discovery walks (tests/run-bats.sh and tests/run-tests.sh) exclude apm_modules/: apm install materializes a full copy of every plugin there, and running a dependency's copy of a .bats file breaks its relative path to the bats helpers — 167 spurious failures before the exclusion landed.

apm gates

apm-audit-ci

Runs apm audit --ci once per manifest — the root one and each of the six plugin packages — because the root-only invocation audits the marketplace manifest and nothing else, and apm-pack-check-clean does not parse plugin dependencies: blocks either. Verified: a malformed dependency entry passes apm pack --check-versions --check-clean --dry-run and fails apm audit --ci in that package's directory. Costs ~0.5s per package.

It verifies exactly two things per manifest and claims no more:

  • manifest-parse — each apm.yml parses as a valid APM manifest. Unconditional; verified to fire on a dependency entry missing its git/path/registry field (Cannot parse apm.yml).
  • lockfile-exists — any package declaring dependencies has a consistent apm.lock.yaml. Conditional, and vacuous while every plugin apm.yml declares dependencies: {apm: [], mcp: []}; it arms itself the moment one does not (verified by adding a git dependency to plugins/lint/apm.yml).

It does not enforce an org policy. apm discovers one from the git remote and only understands github.com and Azure DevOps, so against this repo's self-hosted Gitea remote it prints:

No org policy found at unknown; enforcement skipped

Do not "fix" that with policy.fetch_failure_default: block in apm.yml. apm's own message suggests it; it was tried on a scratch copy and rejected. With no reachable policy source it does not make the check meaningful, it makes it permanently red — apm audit --ci exits 1 with No org policy found at unknown (policy.fetch_failure_default=block) on every push, forever. A gate that can never go green is not a gate. Revisit only if this repo gains a policy source apm can reach.

It also does not scan for hidden Unicode: that scan is plain apm audit, a different mode (--ci refuses to combine with --file/--strip/--dry-run/PACKAGE), and plain apm audit here reports No apm.lock.yaml found -- nothing to scan and exits 0. Adding it would buy a second vacuous check.

check-executables-allow-sync

apm gates a package's hooks/ and bin/ on an exact <package>#<version> dictionary lookup in root apm.yml's executables.allow (apm_cli/security/executables.py, is_package_approved). There is no wildcard and no version-less form.

So bumping plugins/kyberforge/apm.yml's version: without bumping the key errors nowhere: the entry simply stops matching, the gate blocks the hook, kyberforge's SessionStart hook stops deploying, and the apm install goes quietly stale — the exact failure ADR-0019 exists to end, reintroduced through the mechanism meant to secure it. ADR-0019 records this as a live failure mode; the release that shipped the hook hit it immediately.

scripts/check-executables-allow-sync.sh parses version: out of plugins/kyberforge/apm.yml and asserts root apm.yml carries the matching kyberforge#<version> key. A comment in the executables: block stays as the human-facing pointer; the hook is what actually holds. It parses with PyYAML where importable and falls back to a two-shape scan otherwise, so a missing pip package cannot become the thing that blocks every push.

.claude/settings.json

apm owns this file. Nothing repo-authored goes in it.

apm audit --ci replays the install into a scratch tree and diffs the result byte-for-byte, so anything apm would not have written there — an enabledPlugins block, a real hooks entry — is permanent drift that fails apm-audit-ci. A hook you want in this repo is authored in plugins/<name>/.apm/hooks/ and deployed by apm, never hand-written here.

Its committed content is whatever apm last wrote, which today is the merged SessionStart entry for kyberforge's check-apm-current.sh. That is apm's own output and it belongs in the commit (ADR-0019; ADR-0018's statement that the committed content is exactly {"hooks": {}} is superseded on that point only). Machine-specific settings go in the gitignored .claude/settings.local.json, which apm does not deploy and the replay does not compare; shared enforcement belongs in .pre-commit-config.yaml.

Why it is excluded from pretty-format-json

It is the second and last alternation in that hook's exclude: pattern, and the only one there for a reason other than "generated manifest". Mind which number you are quoting: two alternations, expanding to two real files — .claude-plugin/marketplace.json, plus this one.

pretty-format-json --autofix sorts object keys unless --no-sort-keys is passed, while apm's hook integrator emits insertion order (matcher before hooks, type before command). Leaving the file in that hook's scope therefore rewrites apm's output into a form apm would never produce on the way into every commit, and apm-audit-ci then reports permanent drift on a file with an empty git diff — exactly what happened when the SessionStart hook first landed in 2e395a4. Re-running apm install fixes the file; leaving it in scope would re-break it on the very commit carrying the fix.

Load-bearing. Do not tidy it out of that list (see LESSONS.md, 2026-08-14).

Pushing without a network

No pre-push hook needs the network. Every entry in root apm.yml's marketplace.packages[] resolves from a local ./plugins/<name> path, so apm-pack-check-clean never calls git ls-remote.

apm-audit-ci calls apm too but was always local: its org-policy discovery resolves nothing on this remote before any network call.


See also

  • docs/adr/0020-skill-description-and-body-context-contract.md — the context contract, its enforcement table (deterministic vs. auditor judgment), and every rejected alternative
  • docs/adr/0019-session-start-hook-keeps-the-apm-install-current.md — the SessionStart hook, the executable-trust gate, and the version-pinned allow key
  • docs/adr/0024-apm-is-the-only-supported-install-path.md — apm as the sole install path, and the deletion of the flat content mirror and its check-plugin-content-sync gate. It supersedes docs/adr/0017-plugin-content-mirror-bridges-apm-to-host-discovery.md (superseded — plugin content sync, kept as the historical record)
  • docs/adr/0015-apm-replaces-plugin-marketplace-authoring.md, docs/adr/0014-vale-prefilter-ships-from-the-plugin.md — apm-generated manifests, committed Vale styles
  • docs/adr/0025-skill-audit-and-agent-audit-merge-into-factory-audit.md — the audit-pair merge that collapsed the two Vale copies to one, removed the check-vale-style-sync hook, and took the shared boundary resolver from three copies to two. It amends ADR-0014 and ADR-0020 on those points
  • docs/spec/architecture.md — directory structure, install pipeline, what is generated and what is hand-authored
  • .pre-commit-config.yaml — the hooks themselves, with inline rationale comments