5 Commits

Author SHA1 Message Date
4e22c4920a docs: reconcile LESSONS.md and VISION.md with what the branch removed
- LESSONS.md's 2026-06-22 test-placement entry told authors to put test
  files directly in scripts/ with a README row. The file-structure contract
  the repo now enforces permits tests/ as one of four directories, requires
  a tests/README.md when it exists, and FAILs test files in scripts/.
- LESSONS.md's 2026-08-16 entry described a dispatch chain ending at
  skill-author/references/retrofit.md in the present tense. This branch
  deleted that file. Sibling entries whose referents the branch removed
  were marked historical; this one was not.
- VISION.md's Phase 1 now puts stack, framework and deployment choices out
  of scope for this repo, while Phase 3 still named React Native and Tauri.

README.md was checked and needed no change: its offline guarantee already
carries the populated-apm_modules condition from 8cfd54f and agrees with
gates.md and AGENTS.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NwD8Egs5r4ndqeFLmhusX2
2026-09-20 18:38:21 +00:00
9b6f2b1583 docs(audit): freeze the simplification audit and strike what was never true
The document has been re-measured four times and each pass moved figures
the next pass had to chase -- three of the last five commits on this branch
were figure corrections to it, and correcting it changes the line counts it
reports about itself. It is now a dated record frozen at 1ec3e8a. Figures
stand as measured at the commit each one names and are not maintained.

Freezing covers staleness. It does not cover a figure that never
reproduced or a claim that says verified for a check that fails, so those
are struck:

- Finding 11's provenance-validator counts, 320/1,145/572/134 = 2,171, were
  true at no commit. The files are 324/1,152/576/134 = 2,186 and have been
  since 620f20b created them. The derived 5,380 and 7,136 follow.
- Finding 11's line citations into lib-provenance-skill.sh, stated as
  re-derived at HEAD, were uniformly seven low and none landed on the code
  named.
- The tests/ line total pinned to 1614bce is that commit's suite count with
  384756b's line count.
- Finding 33's "all five instruction-level citations still resolve at HEAD,
  verified with sed -n", dated 2026-09-19, is false. Four resolve.
  improve.md:82 stopped carrying the content at baa2f5d, three days before
  the verification was claimed.

Also reconciled: finding 11's 242-file effort total against its own struck
46, finding 16's two different deltas for ef27c97, a clause pinned to
baa2f5d carrying c07ca07's figures, and the preload-tax row's 39 skills
against the census row's 38.

Notes that date themselves "at HEAD" name no fixed commit, and this commit
moves HEAD under them, so the banner now says so rather than re-deriving
twenty of them.

This review round is recorded on the pull request, not here. A frozen
document that grows another section is not frozen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NwD8Egs5r4ndqeFLmhusX2
2026-09-20 18:38:13 +00:00
44bde9e9c9 docs(adr): correct the records the branch left describing deleted things
Seven ADRs described code that no longer exists or behaviour the gates do
not have. Where the wrong text came from main it carries a dated
correction; where this branch introduced it, it is fixed in place, because
main never published it and there is no record to preserve.

Fixed in place, branch-introduced:

- ADR-0021's 2026-09-14 correction asserted apm audit --ci "was never a
  drift gate at all". It is one: it replays the install and diffs. The
  claim contradicted this branch's own AGENTS.md and gates.md.
- ADR-0015 said unconditionally that no pre-push hook needs the network.
  The guarantee holds only once apm install has populated apm_modules/.
- ADR-0014's 2026-09-16 correction said restoring .pre-commit-hooks.yaml
  would ship a hook that fails for every consumer, because their checkout
  has no lib-boundary-resolver.sh. pre-commit clones the whole hook repo
  and skill-size-check.sh resolves the library from BASH_SOURCE, so the
  hook would work.
- ADR-0019's "twelve hooks pass under unshare -rn" matched neither HEAD
  (8) nor main (14), and stated the offline guarantee unconditionally.

Corrected, inherited from main:

- ADR-0022 and ADR-0013 named skill-frontmatter's pre-commit hook as the
  enforcer of mandatory metadata.version. That hook was deleted on this
  branch; the check lives in skill-size-check.sh.
- ADR-0022 enumerated the tip rule's carve-outs as a closed list and
  described a single merge-base. The gate also exempts a tree-identical
  skill and intersects every base from merge-base --all, and emits a third
  failure form. 8cfd54f said the documented behaviour did not change; it
  did. The gate is correct and is unchanged -- the record was not.
- ADR-0020's Decision still routed description overflow to README.md, its
  ADR-0025 amendment pointed the mirrored constants at validate.sh, which
  holds none, and its Enforcement table still named the two deleted
  validate.sh paths.
- ADR-0015's Status claimed every plugin's plugin.json is pack output;
  none exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NwD8Egs5r4ndqeFLmhusX2
2026-09-20 18:38:00 +00:00
e849a823f7 fix(gates): report an unparsed routing clause beside a parsing sibling
boundary_clause_status() ran BOUNDARY_ARROW.search() and _arrow_targets()
over the whole description, so one arrow clause that parsed suppressed the
diagnostic for every other clause in it. A backticked hyphenated routing
target wrapped across lines in a folded scalar was therefore silently
unchecked -- no error, no suggestion, exit 0 -- whenever the description
carried one other clause that parsed. Written bare, the same wrap errors
correctly. That is the shape #100 regressed on.

The check is now per clause. Nothing that passed starts failing: all 68
routing targets across the 38 SKILL.md files resolved before and still do.
26 of those descriptions carry more than one arrow clause, so the
suppression was live across two thirds of the corpus, not an edge case.

validate-skill.bats pins the shape. test-adr0020-targets.sh's comment
described the #100 regression as a backticked wrap; the historical text was
unbackticked, which is precisely the shape the gate did not catch.

Also closes three README misroutes the branch left in the enforcement
layer: CompositionNote.yml's message, agent-description-quality.md:58 and
vale-wrap.sh's header still sent overflow to a skill-root README.md and
named the two skills ADR-0025 merged away. 1ec3e8a fixed the prose and
missed the rules that enforce it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NwD8Egs5r4ndqeFLmhusX2
2026-09-20 18:37:45 +00:00
aa6586c0b6 docs(spec): correct the gate and hook descriptions that did not reproduce
Four claims in the spec and the hook config stated as fact what the tools
do not do:

- gates.md:485 said factory-audit's validate.sh "holds its own copy of"
  the ADR-0020 constants. gates.md:411-413, twenty lines earlier, said it
  carries none of them and named the mode libraries. The libraries are
  right: lib-checks-skill.sh:313-316 and lib-checks-agent.sh:164-165.
  architecture.md repeated the same error.
- gates.md stated the case count for test-adr0020-contract.sh as "29 at
  HEAD", explicitly presented as measured. Running it prints 44; 384756b
  added the hook-wiring assertions after the text was written.
- gates.md:913 and :916 described "Both audit skills'" behaviour in the
  present tense, three and six lines above :919 saying factory-audit's is
  the only copy left.
- The apm-audit-ci block named manifest-parse as a check, said the hook
  does not scan for hidden Unicode, and called root lockfile-exists
  vacuous. apm 0.28.0 runs ten checks, content-integrity does scan for
  hidden Unicode, and there is no manifest-parse row.

Also: the version-bump gate's baseline is documented as the single
merge-base it is not -- it resolves every base with merge-base --all,
intersects the changed-skill sets, exempts a tree-identical skill, and
emits a third sha-suffixed failure form. The gate's own header documents
this correctly; the spec did not. Behaviour is unchanged.

The "none of them need the network" line added on this branch cited a
README section that says the opposite for a fresh clone, and the
check-vale-style-sync rationale said 6 of 17 assertions diffed the Vale
copies where ADR-0025 says 2 diffed and 4 more only located them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NwD8Egs5r4ndqeFLmhusX2
2026-09-20 18:37:36 +00:00
20 changed files with 342 additions and 112 deletions

View File

@@ -97,48 +97,67 @@ repos:
- id: apm-audit-ci
name: apm audit --ci
description: Run apm's producer-side CI gate over the root manifest AND each of the six plugin packages. Verifies exactly two things per manifest -- apm.yml parses as a valid APM manifest (manifest-parse), and, if it declares dependencies, apm.lock.yaml exists and is consistent (lockfile-exists). It does NOT enforce an org policy and does NOT scan for hidden Unicode; see the comment below for why. Reference:plugins/kyberforge/.apm/skills/apm-workflow/references/audit.md
description: Run apm's producer-side CI gate over the root manifest AND each of the six plugin packages. On the root manifest it runs ten checks -- lockfile-exists, ref-consistency, deployment-ledger-owners, deployed-files-present, no-orphaned-packages, skill-subset-consistency, config-consistency, content-integrity, includes-consent, drift -- so it is both a hidden-Unicode scan and a drift gate that replays the install and diffs it. In a plugin package it runs one, lockfile-exists. It does NOT enforce an org policy; see the comment below for why. Reference:plugins/kyberforge/.apm/skills/apm-workflow/references/audit.md
entry: bash -c 'for d in . plugins/*/; do (cd "$d" && apm audit --ci) || { echo "apm audit --ci failed in $d" >&2; exit 1; }; done'
language: system
stages: [pre-push]
pass_filenames: false
always_run: true
# The description above deliberately claims less than this hook's old one
# did ("lockfile/policy/hidden-content integrity"), because two of those
# three were never happening:
# What this hook actually runs, read off apm 0.28.0's own compliance
# table by invoking `apm audit --ci` at the repo root and in
# plugins/lint/. Long form in docs/spec/gates.md, "apm-audit-ci".
#
# * POLICY. `apm audit --ci` discovers an org policy from the git remote,
# and apm's discovery only understands github.com and Azure DevOps.
# This repo's remote is a self-hosted Gitea, so discovery resolves
# nothing and the run prints `No org policy found at unknown;
# enforcement skipped`. apm's own message suggests
# `policy.fetch_failure_default=block` in apm.yml "to fail closed" --
# that was tried on a scratch copy and REJECTED: it does not make the
# check meaningful, it makes it permanently red. `apm audit --ci` then
# exits 1 with `No org policy found at unknown
# (policy.fetch_failure_default=block)` on every push, because there is
# no org policy to find and no supported way for this remote to serve
# one. A gate that can never go green is not a gate. Revisit if this
# repo ever gains a policy source apm can actually reach.
# * HIDDEN CONTENT. The hidden-Unicode scan is plain `apm audit`, not
# `apm audit --ci` (the two are different modes, and --ci refuses to
# combine with --file/--strip/--dry-run/PACKAGE). Plain `apm audit`
# here reports `No apm.lock.yaml found -- nothing to scan` and exits 0,
# so adding it would buy a second vacuous check, not coverage.
# * ROOT MANIFEST -- ten checks: lockfile-exists, ref-consistency,
# deployment-ledger-owners, deployed-files-present,
# no-orphaned-packages, skill-subset-consistency, config-consistency,
# content-integrity, includes-consent, drift. It is a drift gate: it
# replays the install cache-only and diffs the scratch result against
# the working tree. Root lockfile-exists is not vacuous -- the root
# declares dependencies, so it reports `Lockfile present`.
# * PLUGIN MANIFESTS -- one check: lockfile-exists. Conditional, and
# vacuous while every plugin apm.yml declares
# `dependencies: {apm: [], mcp: []}`: it reports `No dependencies
# declared -- lockfile not required` and arms itself the moment one
# does not (verified by adding a git dependency to
# plugins/lint/apm.yml). Everything else above is root-only, because
# only the root install has a lockfile, a deployment ledger and
# deployed files to check. Running the six plugin packages is what
# makes lockfile-exists reachable for them at all -- the root-only
# invocation audits the root manifest and nothing else.
# * HIDDEN CONTENT IS COVERED. content-integrity is that scan; it
# reports `No critical hidden Unicode or hash drift detected`. An
# earlier revision of this comment said the hook does NOT scan for
# hidden Unicode and that adding the scan would buy a second vacuous
# check. Both claims were wrong. What is true is that the STANDALONE
# mode differs: plain `apm audit` (--ci refuses to combine with
# --file/--strip/--dry-run/PACKAGE) run in a plugin directory reports
# `No apm.lock.yaml found -- nothing to scan` and exits 0, because
# only the root has a lockfile.
# * MANIFEST-PARSE IS NOT A CHECK in apm 0.28.0's table, and an earlier
# revision of this comment named it as one. Parsing is still
# enforced -- a dependency entry missing its git/path/registry field
# fails with `Cannot parse apm.yml` -- but it fails the invocation
# before the table is built, so it never appears as a row.
# * POLICY IS NOT ENFORCED. `apm audit --ci` discovers an org policy
# from the git remote, and apm's discovery only understands
# github.com and Azure DevOps. This repo's remote is a self-hosted
# Gitea, so discovery resolves nothing and the run prints `No org
# policy found at unknown; enforcement skipped`. apm's own message
# suggests `policy.fetch_failure_default=block` in apm.yml "to fail
# closed" -- that was tried on a scratch copy and REJECTED: it does
# not make the check meaningful, it makes it permanently red. `apm
# audit --ci` then exits 1 with `No org policy found at unknown
# (policy.fetch_failure_default=block)` on every push, because there
# is no org policy to find and no supported way for this remote to
# serve one. A gate that can never go green is not a gate. Revisit if
# this repo ever gains a policy source apm can actually reach.
#
# What IS left is worth keeping, and is now run against seven manifests
# instead of one. lockfile-exists is conditional -- it is vacuous while
# every apm.yml declares `dependencies: {apm: [], mcp: []}`, and it arms
# itself the moment one does not (verified: adding a git dependency to
# plugins/lint/apm.yml fails with `apm.yml declares dependencies but
# apm.lock.yaml is absent`). manifest-parse is unconditional and fires on
# any malformed manifest (verified: a dependency entry missing its
# git/path/registry field fails with `Cannot parse apm.yml`). Running the
# six plugin packages is what makes either reachable for them at all --
# the root-only invocation audits the marketplace manifest and nothing
# else. Costs ~0.5s per package, needs no network (checked under
# `unshare -rn`) -- consistent with every other pre-push hook: none of
# them need the network (see README.md's "Offline?" section).
# Costs ~0.5s per package. Needs no network ONCE `apm install` has
# populated apm_modules/ -- the root marketplace has no remote package
# entries, so the install replay is cache-only. On a FRESH CLONE there
# is no cache: deployed-files-present fails outright, and drift and
# config-consistency clone from the holocron remote. See README.md's
# "Offline?" section.
- id: check-apm-agents-valid
name: Validate real APM agent files
@@ -167,9 +186,12 @@ repos:
pass_filenames: false
always_run: true
# check-vale-style-sync was removed by ADR-0025. Only 6 of its 17
# assertions diffed skill-audit's Vale copy against agent-audit's; the
# merge into factory-audit leaves one copy, so those are moot. The other
# check-vale-style-sync was removed by ADR-0025. Of its 17 assertion
# sites only 2 actually diffed skill-audit's Vale copy against
# agent-audit's, and 4 more existed solely so the script could locate the
# two copies -- a real REPO_ROOT, non-stale .apm/ paths, both copies
# present (ADR-0025:285-287). The merge into factory-audit leaves one
# copy, so all 6 are moot. The other
# 11 moved into tests/test-vale-wrap.sh (case 0, cases 28-31, its
# Vale-absent skip, and case 32 for the one-plugin narrowing guard),
# which run-tests runs here at

View File

@@ -56,9 +56,9 @@ Skills sharing a resource (e.g. `validate.sh`) via a `shared/` directory and rel
`skill-audit`'s (now `factory-audit`'s skill flow, per ADR-0025: `references/skill-description-quality.md` and `references/skill-body-discipline.md`) description and body-discipline rubrics were derived from `skill-write`'s own conventions — circular, so drift in one silently propagated to the other. Fix: extract condensed reference files directly from the upstream spec (agentskills.io) into the audit skill, so the rubric is independent of in-repo convention drift.
## 2026-06-22 — Test files in scripts/ are dev tooling; document them in README as non-spec
## 2026-06-22 — Test files in scripts/ are dev tooling; document them in README as non-spec (historical)
The agentskills.io spec defines `scripts/` for bundled executables, not test infrastructure — bats files placed there are invisible to spec-following auditors and cause README drift. Fix: place test files directly in `scripts/` (no subdirectory), and add a README row noting each as "dev tooling, not shipped."
Superseded — the fix below is now itself a FAIL. `factory-audit`'s `references/skill-file-structure.md:14` permits `tests/` as one of the four allowed directories, `:21-22` fails a test file found in `scripts/`, and `:58-60` requires a `tests/README.md` when `tests/` exists. Skill-root READMEs are gone too, so there is no table left to add a row to. What survives is the reason: test infrastructure is dev tooling, not shipped content, and has to be declared where an auditor reads — which is now `tests/README.md`. Kept for reference: the agentskills.io spec defines `scripts/` for bundled executables, not test infrastructure — bats files placed there are invisible to spec-following auditors and cause README drift. Fix: place test files directly in `scripts/` (no subdirectory), and add a README row noting each as "dev tooling, not shipped."
## 2026-06-27 — Clean-context audit catches what biased forks miss
@@ -130,9 +130,9 @@ Widening a description-opener rule to also catch mid-sentence text looked like a
`apm audit --ci` failed on `.claude/settings.json` with an empty `git diff` — `pretty-format-json --autofix` silently re-sorts JSON keys, and this generated file was missing from its exclude list, so every commit re-sorted apm's insertion-ordered output before apm compared against it. Separately, a defect introduced 3 hours earlier on the same branch was first mis-described as "pre-existing," an unverified claim about history. Fix: add tool-owned paths to every autofixing hook's exclude the moment ownership is declared, and verify "pre-existing" claims with `git log -S` or `git branch --contains` before writing them down.
## 2026-08-16 — A rule reversed inside a retrofit leaves no trace unless someone writes it down
## 2026-08-16 — A rule reversed inside a retrofit leaves no trace unless someone writes it down (historical)
A retrofit replaced "keep reference chains one level deep" with "two hops, never three" — the opposite rule, needed because the new dispatch pattern requires `SKILL.md` → `improve.md` → `retrofit.md`. The ADR never mentioned chain depth, so the reversal was carried entirely by the diff with no sign a contradicting rule ever existed. Fix: when a change inverts a standing rule, record the inversion where the rule's rationale lives, or it reads as forgotten rather than overturned.
The chain named below no longer exists — `plugins/kyberforge/.apm/skills/skill-author/references/retrofit.md` was deleted, so the dispatch ends at `improve.md`. The reversed rule itself survives, in `skill-author/references/create.md:150`. Kept for reference: a retrofit replaced "keep reference chains one level deep" with "two hops, never three" — the opposite rule, needed because the new dispatch pattern requires `SKILL.md` → `improve.md` → `retrofit.md`. The ADR never mentioned chain depth, so the reversal was carried entirely by the diff with no sign a contradicting rule ever existed. Fix: when a change inverts a standing rule, record the inversion where the rule's rationale lives, or it reads as forgotten rather than overturned.
## 2026-09-15 — A rare flake in a pipefail suite is a race until proven otherwise

View File

@@ -60,4 +60,4 @@ Runtime orchestration: push config updates to machines, see running agents, mana
### Phase 3 — Native Apps
Mobile (React Native) and desktop (Tauri) wrappers over the Phase 1/2 web app. Deferred until the web app is mature.
Mobile and desktop wrappers over the Phase 1/2 web app. Deferred until the web app is mature; the wrapper technology is that product's own choice, on the same terms as the rest of its stack.

View File

@@ -26,7 +26,11 @@ stay out of this Vale-based harness because this repo already has dedicated tool
`skill-frontmatter` (required frontmatter fields), `validate-marketplace`
(`claude plugin validate --strict`, schema), and `gitleaks`/`detect-private-key` (secrets).
(ADR-0024 removed the companion `validate-plugins` gate along with the per-plugin manifests it
checked; the argument here is unaffected.)
checked, and the `skill-frontmatter` hook has since been removed as well — its required-field
checks were folded into `skill-size-check`, and the enforcer is now
`scripts/skill-size-check.sh:324-335` under the `skill-size-check` hook at
`.pre-commit-config.yaml:237`. The argument here is unaffected either way: a dedicated
non-Vale tool still owns required frontmatter fields.)
Duplicating those concerns as Vale rules would fight tools that already own them better.
**Governance docs are excluded as a rule source.** `docs/research/governance_principles/CONTROLS.md`

View File

@@ -277,11 +277,22 @@ release gate only once a server-side job can run it on merge.
`skill-size-check.sh`. `ef27c97` removed its embedded resolver copy: the hook now sources
`plugins/kyberforge/.apm/skills/factory-audit/scripts/lib-boundary-resolver.sh` by path and fails
closed without it (ADR-0020's 2026-09-16 amendment; `docs/spec/gates.md`, "Duplicated constants").
An external consumer's checkout has no such file, so restoring the manifest as described would ship
a hook that fails for every consumer — the defect recorded at `LESSONS.md:101`. Only `vale-wrap.sh`
still meets the `entry[0]`-only constraint. A return must first make `skill-size-check.sh`
self-contained again, by re-embedding the resolver or shipping the library beside the hook, and
restore a consumer test that proves it.
**Amended (2026-09-20): that is not a consumer-facing defect.** The correction above went on to say
that a consumer's checkout has no such file, so restoring the manifest would ship a hook that fails
for every consumer. Reproduced and found false. pre-commit's `script` language clones the **whole**
hook repo into its store and prefixes `entry[0]` with the clone directory: `clientlib.py` maps
`script` to `unsupported_script`, whose `run_hook` does `cmd = (prefix.path(cmd[0]), *cmd[1:])` over
`Prefix(store.clone(...))`, and `store.clone` checks out the full tree — shallow in depth, not in
content. `skill-size-check.sh` locates the library from `${BASH_SOURCE[0]}`
(`scripts/skill-size-check.sh:481-483`), which points into that same clone, so
`../plugins/kyberforge/.apm/skills/factory-audit/scripts/lib-boundary-resolver.sh` resolves beside
it. Verified end to end against pre-commit 4.6.1 with a probe hook of the same shape — bare script
entry, sibling file reached by climbing out of `scripts/` — and the file was found and sourced. Both
hook scripts therefore still meet the `entry[0]`-only constraint: `vale-wrap.sh` takes no `--config`,
and `skill-size-check.sh` passes no argv of its own. A return needs no re-embedding; restore
`test-vale-hooks-consumer.sh` with the manifest, extended to cover the sourced library, so the claim
stays checked rather than reasoned about.
**Superseded statements elsewhere.** ADR-0022's notes that the version-bump gate "is not exported
through `.pre-commit-hooks.yaml`" and that it shares its gaps with `check-release-needed`, and

View File

@@ -5,6 +5,12 @@ their authoring source; `.claude-plugin/marketplace.json` and every plugin's `pl
`apm pack`-compiled output. **Supersedes ADR-0001** ("Skills are distributed via plugins... each
plugin contains its own `skills/` directory") — in effect.
**Correction (2026-09-20): the present tense above has expired for `plugin.json`.** ADR-0024 made
apm the only supported install path and deleted per-plugin `plugin.json` with the native install
support that needed it. No plugin carries one at `HEAD` — `git ls-files | grep -c 'plugin\.json'`
returns 0 — so `.claude-plugin/marketplace.json` is the only `apm pack`-compiled output left. Read
the Status line as the state at execution, 2026-08-12.
This repo replaces its hand-maintained Claude Code plugin/marketplace authoring model
(`.claude-plugin/marketplace.json` + per-plugin `plugin.json`) with Microsoft APM (`apm.yml` +
`.apm/`) as the authoring source of truth — an outright replacement of the authoring layer, not an
@@ -180,7 +186,11 @@ correction) sorted what they document into three buckets:
`git ls-remote`, which is why two pre-push hooks needed the network (see `AGENTS.md`).
**Superseded 2026-09-13:** the `mattpocock-skills` entry has been removed from root `apm.yml`
entirely, along with the `codex` marketplace output profile. No pre-push hook needs the network
any longer.
any longer — **once `apm install` has populated `apm_modules/`**. The guarantee is a property of a
populated install, not of the hook set: on a fresh clone `apm-audit-ci`'s `deployed-files-present`
fails outright, and its `drift` and `config-consistency` install-replays have no cache to replay
from and clone from the holocron remote (`README.md:89`; `docs/spec/gates.md`, "Pushing without a
network").
- **Caveat on "Status: executed" above:** issue #90's own execution comment flagged, before merge,
that Claude Code's ability to actually load content out of `.apm/` was unverified — that caveat
turned out to be a real defect, not a formality: the native installer has zero awareness of

View File

@@ -156,6 +156,16 @@ via `url.<path>.insteadOf`, so the twelve-hooks-pass-under-`unshare -rn` propert
the real `apm outdated`, and replays its genuine output through the real hook. Reverting the grep
to plural-only fails it.
> **Correction (2026-09-20):** neither half of "the twelve-hooks-pass-under-`unshare -rn` property"
> is accurate. The count was never twelve: `main` declares 14 pre-push hooks and `HEAD` declares 8 —
> 10 counting the two `repo: meta` hooks, which set no `stages` and so run at every stage. And the
> property is conditional, not absolute: no pre-push hook needs the network **once `apm install` has
> populated `apm_modules/`**, but on a fresh clone `apm-audit-ci`'s `deployed-files-present` fails
> outright and its `drift` and `config-consistency` install-replays clone from the holocron remote
> (`README.md:89`; `docs/spec/gates.md`, "Pushing without a network"). What the probe itself
> establishes is unchanged and is the point of the sentence: staging the outdated dependency against
> a local git remote via `url.<path>.insteadOf` adds no network call of its own.
**The hook cannot install itself.** Dependencies resolve from the remote, so the hook does not
deploy until this change is merged and `apm update` has run once against the new default branch.
Until then the repo has the mechanism in source and not in effect.

View File

@@ -21,9 +21,12 @@ SHARED BOUNDARY RESOLVER` markers, and one plugin copy — extracted out of the
both. `validate-provenance.sh` is not a third reader: it sources `lib-contributing-files.sh` and one
of `lib-provenance-skill.sh`/`lib-provenance-agent.sh`, and never touches the resolver at all. The
Enforcement table's "constants mirrored in `skill-audit/scripts/validate.sh` and
`agent-audit/scripts/validate.sh`" is one path now, `factory-audit/scripts/validate.sh`, which
auto-detects the artifact type; the skills/agents columns are unaffected, since the merged validator
applies the body tiers on the skill path only. The two copies must still stay byte-identical — a
`agent-audit/scripts/validate.sh`" now means `factory-audit/scripts/lib-checks-skill.sh:313-316`
(all four constants) and `lib-checks-agent.sh:164-165` (the two description ones). It does **not**
mean `factory-audit/scripts/validate.sh`, which holds none of them: `validate.sh` auto-detects the
artifact type and sources the matching check suite (`validate.sh:231-233`, `:244-246`). The
skills/agents columns are unaffected — only the skill suite carries the body tiers. The two copies
must still stay byte-identical — a
plugin script cannot source the root one, which is why a second copy exists at all. Read every
"three" below as the count at the time of writing.
@@ -116,6 +119,12 @@ clause**, and a **boundary clause**. Capability enumeration, output-format detai
("composes X rather than duplicating Y"), and implementation detail move to the body or to
`README.md`.
**Correction (2026-09-20): not `README.md`.** The canonical destination for description overflow is
"the body or a `references/` file" (`plugins/kyberforge/.apm/skills/skill-author/references/contract.md:35`).
A skill-root `README.md` is no longer somewhere overflow can go: all 39 of them were deleted, and
`factory-audit/references/skill-file-structure.md:23` now FAILs a non-spec file at the skill root,
which a `README.md` is. Read every "or to `README.md`" below as "or to a `references/` file".
- **250 characters SUGGESTION, 400 FAIL.** The agentskills.io 1,024-character limit remains as an
unchanged spec backstop. The SUGGESTION tier is what moves the average; the FAIL tier only stops
outliers.
@@ -284,8 +293,8 @@ which tier each rule is in, because the failure this ADR is most exposed to is a
| Check | Applies to | Tier | Home |
|---|---|---|---|
| description characters (250 SUGGESTION † / 400 FAIL) | skills, agents | deterministic | `scripts/skill-size-check.sh`; constants mirrored in `skill-audit/scripts/validate.sh` and `agent-audit/scripts/validate.sh` |
| body-only words (600 SUGGESTION / 900 FAIL) | skills | deterministic | `skill-size-check.sh`, `skill-audit/scripts/validate.sh` (now `factory-audit/scripts/validate.sh`, see ADR-0025) |
| description characters (250 SUGGESTION † / 400 FAIL) | skills, agents | deterministic | `scripts/skill-size-check.sh`; constants mirrored in `skill-audit/scripts/validate.sh` and `agent-audit/scripts/validate.sh` (now `factory-audit/scripts/lib-checks-skill.sh:313-314` and `lib-checks-agent.sh:164-165`, see the ADR-0025 amendment — **not** `factory-audit/scripts/validate.sh`, which holds no constants) |
| body-only words (600 SUGGESTION / 900 FAIL) | skills | deterministic | `skill-size-check.sh`, `skill-audit/scripts/validate.sh` (now `factory-audit/scripts/lib-checks-skill.sh:315-316`, see the ADR-0025 amendment) |
| description present and non-empty (ERROR) | skills, agents | deterministic | same |
| boundary target resolves to a real skill or agent — **three** verdicts, not two (ERROR when written in route notation — `/name`, or any arrow form; or when a *terminal* bare name's own sentence names another target that resolves. SUGGESTION otherwise. INFO, "DID NOT RUN", exit 0, when no skill universe could be determined for the path at all — no authoring root above it, no apm package root, no declared apm dependencies, no deployed `.claude/` or `.agents/` tree: the targets are named and left unchecked) | skills, agents | deterministic | same |
| boundary clause absent — `absent` (SUGGESTION) † | skills, agents | deterministic | same |

View File

@@ -103,17 +103,20 @@ output against `apm.yml`, so their entire job is to propagate whatever the descr
those four files byte-for-byte and confirm they match. The `wiki` claim passed every one of the fourteen pre-push hooks, every day it
was published.
**Correction (2026-09-14): that gate list is down to one, and it was never two.**
**Correction (2026-09-14): that gate list is down to two.**
`scripts/sync-plugin-content.sh --check --all` does not exist — `718c79a` deleted the script and its
`check-plugin-content-sync` hook with the flat mirror (ADR-0024). Of the two names left,
`apm audit --ci` was never a drift gate at all: against this repo it checks only that each `apm.yml`
parses and that a manifest declaring dependencies has a consistent `apm.lock.yaml`, and it reads no
`description`. So the sole surviving gate that compares compiled output against `apm.yml` is
`check-plugin-content-sync` hook with the flat mirror (ADR-0024). The other two survive.
`apm audit --ci` at the repo root (apm 0.28.0) runs ten checks — `lockfile-exists`,
`ref-consistency`, `deployment-ledger-owners`, `deployed-files-present`, `no-orphaned-packages`,
`skill-subset-consistency`, `config-consistency`, `content-integrity`, `includes-consent` and
`drift` — and it *is* a drift gate: `drift` and `config-consistency` replay the install and diff the
result against the working tree, and `content-integrity` scans for hidden Unicode and hash drift.
(In a sub-package such as `plugins/lint` it runs one check, `lockfile-exists`.) The second is
`apm pack --check-versions --check-clean --dry-run`, run by the `apm-pack-check-clean` pre-push hook
— and with the per-plugin manifests gone it propagates a description into exactly one file,
`.claude-plugin/marketplace.json`, not four. This narrows the mechanism and changes nothing about
the finding: propagation is still not verification, and nothing anywhere reads the `description`
key for sense.
the finding: both gates compare bytes, neither reads the `description` key for sense, so propagation
is still not verification.
**And the obligation is unbounded.** Under enumeration, adding one skill to `bin`, `git` or `gitea`
means editing two copies of a prose string on top of the version bumps and regeneration any skill

View File

@@ -54,6 +54,13 @@ per-plugin choice.
`name:` or `description:`; a missing `metadata.version` is now the same class of failure, not a
style nit an audit might or might not catch.
> **Correction (2026-09-20): that hook no longer exists.** `skill-frontmatter` was removed and its
> required-field checks folded into `skill-size-check`. The enforcer is now
> `scripts/skill-size-check.sh:324-335`, declared under the `skill-size-check` hook at
> `.pre-commit-config.yaml:237`. It checks presence and three-part-semver shape, on the same
> `SKILL.md` glob and at the same pre-commit stage, so the decision is unaffected — only the name
> of the hook that holds it. The same substitution applies to the Consequences section below.
## Considered options
**Leave it per-plugin, document the split.** This was the initial framing of #127 and is coherent —
@@ -152,6 +159,38 @@ on the same skill. That is the case the rule exists for, and rebasing onto or me
shows the version to beat. The rule reads `origin/main` as last fetched, so a tip that moved since
the last fetch is not seen until the next one.
## Amendment (2026-09-20): two exemptions and a third failure form the rules above never stated
The carve-outs enumerated above read as a closed list, and the baseline above reads as a single
merge-base. `scripts/check-skill-version-bump.sh` as shipped has two further exemptions and emits a
third failure form. **The gate is right and is not changing; this ADR was behind it.** Its own header
comments (`:9-93`) have described all three correctly since it shipped.
- **The baseline is `git merge-base --all`, not one merge-base.** A criss-cross history — `main`
merges a branch while that branch merges a commit of `main` — has two merge bases, and which one
`git merge-base` prints is an implementation detail. The script takes every base (`:154-157`) and
**intersects** the changed-skill sets across them (`:203-217`): a skill matching any one base is
already shipped by that base and is exempt, and a skill that does reach the comparison must exceed
the version at every base it exists at (`:367-376`). Picking one base made the verdict a coin
flip — an already-merged bump failed the push it should have passed.
- **A skill whose directory tree object equals the tip's skips the tip comparison.**
`same_subtree()` (`:288-297`, applied at `:334-337`) compares tree object ids rather than diffing:
the same tree is the same content, whatever route the history took to it. A branch cut before a
fix landed on `main` and then cherry-picking that fix has one merge-base, predating the fix, so the
skill counts as changed against it and reaches the tip comparison carrying exactly the tip's
version — same content, same version. The merge-base intersection catches that only when some base
carries the content, which the criss-cross shape gives and a linear one does not. Without the skip
the only escapes are a spurious bump, leaving `main` carrying two versions of identical content,
or a rebase the push does not otherwise need.
- **A third failure form.** The amendment above lists `(not above merge-base)` and
`(not above origin/main tip)`. When there is more than one base, the merge-base line is
sha-suffixed — `(not above merge-base <sha>)` (`:372`, against the unsuffixed `:374`) — because
"which merge-base" is the one question a reader cannot answer from the branch alone.
`8cfd54f` recorded that "ADR-0022 is not amended: the documented behaviour does not change". That
was wrong for the tree-identical case: the `same_subtree` skip makes a push **pass** that this ADR as
written requires to **fail**, which is documented behaviour changing, not an implementation detail.
## Consequences
27 SKILL.md files gain `metadata.version: "1.0.0"`, and a 28th — `bin/write-docs` — reaches the same

View File

@@ -1,6 +1,6 @@
# Simplification audit
> **Status: complete (2026-09-16).** Every finding is closed at its own note except **22**, deferred with `bin`. See §7's status notes for the closing summary. This document is now a record; do not reopen it for new work — file an issue instead.
> **Frozen (2026-09-20) — a dated record, not a live document.** Status: complete; every finding is closed at its own note except **22**, deferred with `bin` (§7's status notes carry the closing summary). Every figure below is as measured at the commit it names, and none of them are maintained against HEAD; the document is frozen at `1ec3e8a`. Where a note dates itself "at HEAD", that means the branch tip on **that note's own date**, not the current tip — those figures were not re-derived for the freeze, and several are stale by construction because later commits moved what they measure. Do not re-measure it and do not reopen it for new work — file a Gitea issue instead.
Date: 2026-09-10. Read-only analysis; nothing has been changed. Purpose: a hand-off for deciding what to remove, merge, and shrink. Findings are ranked by payoff within each area; effort is S/M/L. Claims were independently re-verified against the repo by a clean reviewer; corrections have been applied.
@@ -23,7 +23,7 @@ Counting convention: line counts are hand-edited `.apm/` source unless marked "i
>
> > **Re-measured (2026-09-14, at `a6434e0`):** the right-hand column originally read 31,473 / 6,050 / 3,471 / 2,360 / 923 / 2,083 = 46,360 and was labelled "Today" against "the current working tree". It did not reconcile to its own commit's tree — at `061bb3d`, where it was written, the six plugins measured 31,435 / 6,048 / 3,474 / 2,358 / 926 / 2,087 = 46,328 — and "the current working tree" is a basis that goes stale silently. Re-counted at `a6434e0` and the column now names its SHA. The baseline column is confirmed exact against `9eb8bc7`. Commits after `061bb3d` (`c96ca9c`, which deleted the six plugin-root `.mcp.json` files) account for most of the remaining drift.
> **Re-derived (2026-09-16, at HEAD on `docs/simplification-audit`):** the 2026-09-15 notes recording finding 14's merge (~~`467bbd7`~~ → `620f20b`, ADR-0025) and the pipefail fix (~~`4059cb4`~~ → `ffcbed6`) were written without correcting the headlines they annotate, so this pass re-counted every figure those two commits could have moved and corrected each in place above and below. Everything re-measured here came from a command run at HEAD — `git ls-files`, `wc -l`, `grep -c`, and `bash tests/run-tests.sh --strict` — never from an earlier note. What moved: finding 2 (two surviving sync gates → one), the `.pre-commit-config.yaml` hook counts (27/9 → 26/8 → 27/9 → 26/8, the chain spelled out in §3's table note below; **26** `- id:` entries and **8** `stages: [pre-push]` at HEAD, `grep -c -- "- id:"` and `grep -c "stages: \[pre-push\]"`), the skill census (39 → 38 and everything derived from it), finding 11's validator and `sources.md` figures, finding 16's whole numeric basis, and the stale `skill-audit/`, `agent-audit/` and `formatting-and-scripts.md` paths in findings 18, 19 and 33. §1's three rows re-measured: ~~**469**~~ → **471** tracked files (~~465~~ → 467 regular plus the 4 submodule gitlinks) / ~~**74,594**~~ → **75,441** lines (pinned to `c07ca07`; see the note below); `plugins/` ~~**46,106** (62%)~~ → **46,127** (61%); the 38 `SKILL.md` bodies **2,409** (5.2% of plugin lines); enforcement ~~**20 `tests/test-*.sh` totalling 10,189 lines**~~ → ~~**21 `tests/test-*.sh` totalling 10,608 lines**~~ → **19 totalling 10,088** at `baa2f5d`, the two runners **502** (`run-tests.sh` 283 + `run-bats.sh` 219), and `scripts/` ~~**2,901**~~ → **3,139**; kyberforge's validator scripts and their bats tests ~~**5,861**~~ → **5,876** + **6,015** (the merge deduplicated scripts and left the test corpus larger, not smaller — `git ls-files 'plugins/kyberforge/.apm/skills/*/scripts/*.sh'` and `.../tests/*.bats`). `run-tests.sh --strict` reports ~~**20 passed, 0 skipped, 0 failed**~~ → ~~**21 passed, 0 skipped, 0 failed**~~ → **19 passed, 0 skipped, 0 failed** at `baa2f5d` (`4de5b6b` deleted two suites).
> **Re-derived (2026-09-16, at HEAD on `docs/simplification-audit`):** the 2026-09-15 notes recording finding 14's merge (~~`467bbd7`~~ → `620f20b`, ADR-0025) and the pipefail fix (~~`4059cb4`~~ → `ffcbed6`) were written without correcting the headlines they annotate, so this pass re-counted every figure those two commits could have moved and corrected each in place above and below. Everything re-measured here came from a command run at HEAD — `git ls-files`, `wc -l`, `grep -c`, and `bash tests/run-tests.sh --strict` — never from an earlier note. What moved: finding 2 (two surviving sync gates → one), the `.pre-commit-config.yaml` hook counts (27/9 → 26/8 → 27/9 → 26/8, the chain spelled out in §3's table note below; **26** `- id:` entries and **8** `stages: [pre-push]` at HEAD, `grep -c -- "- id:"` and `grep -c "stages: \[pre-push\]"`), the skill census (39 → 38 and everything derived from it), finding 11's validator and `sources.md` figures, finding 16's whole numeric basis, and the stale `skill-audit/`, `agent-audit/` and `formatting-and-scripts.md` paths in findings 18, 19 and 33. §1's three rows re-measured: ~~**469**~~ → **471** tracked files (~~465~~ → 467 regular plus the 4 submodule gitlinks) / ~~**74,594**~~ → **75,441** lines (pinned to `c07ca07`; see the note below); `plugins/` ~~**46,106** (62%)~~ → **46,127** (61%); the 38 `SKILL.md` bodies **2,409** (5.2% of plugin lines); enforcement ~~**20 `tests/test-*.sh` totalling 10,189 lines**~~ → ~~**21 `tests/test-*.sh` totalling 10,608 lines**~~ → **19 totalling 10,088** at `baa2f5d`, the two runners ~~**502** (`run-tests.sh` 283 + `run-bats.sh` 219)~~ → **514** (`run-tests.sh` 289 + `run-bats.sh` 225), and `scripts/` ~~**2,901**~~ → ~~**3,139**~~ → **1,926** (**corrected 2026-09-20**: the struck runner and `scripts/` figures are `c07ca07`'s, not `baa2f5d`'s, so this one clause carried two bases and contradicted §1's own row for the same commit; the replacements are `baa2f5d`'s and agree with that row); kyberforge's validator scripts and their bats tests ~~**5,861**~~ → **5,876** + **6,015** (the merge deduplicated scripts and left the test corpus larger, not smaller — `git ls-files 'plugins/kyberforge/.apm/skills/*/scripts/*.sh'` and `.../tests/*.bats`). `run-tests.sh --strict` reports ~~**20 passed, 0 skipped, 0 failed**~~ → ~~**21 passed, 0 skipped, 0 failed**~~ → **19 passed, 0 skipped, 0 failed** at `baa2f5d` (`4de5b6b` deleted two suites).
>
> > **Re-measured (2026-09-16, at `c07ca07`):** commit `8451169` added `check-skill-version-bump` — a pre-push hook, `scripts/check-skill-version-bump.sh` (238 lines) and `tests/test-skill-version-bump.sh` (410) — after the figures above were taken, so each was one short. `.pre-commit-config.yaml` now has **27** `- id:` entries and **9** `stages: [pre-push]` (`grep -c -- "- id:"`; `grep -c "stages: \[pre-push\]"`), all nine repo-authored. The struck figures are replaced from these commands. They were run against the working tree, and every figure reproduces exactly from the committed tree at `c07ca07`: `git ls-files | wc -l`; `cat` over every non-gitlink tracked path `| wc -l`; `git ls-files plugins | xargs cat | wc -l`; `git ls-files scripts | xargs wc -l` (no untracked files under `scripts/`); `ls tests/test-*.sh | wc -l` and `cat tests/test-*.sh | wc -l`; `bash tests/run-tests.sh --strict`. The earlier 469 / 74,594 / 46,106 did not reproduce exactly at `8451169^` either (469 / 74,638 / 46,121), so they were taken at an earlier commit than this note's "at HEAD" says. Re-checked and unchanged, so left alone: `docs/research/` inside plugins (19,030) and repo-level `docs/research/` + `docs/notes/` (4,488). Not re-measured, and still carrying their last stated basis: the preload-tax and commit-share rows, §2's timings, and the per-plugin table in the note above.
>
@@ -42,11 +42,11 @@ Counting convention: line counts are hand-edited `.apm/` source unless marked "i
| Of which the ~~39~~ → 38 `SKILL.md` files a model actually loads | ~~about 2,600 lines (under 4% of plugin lines)~~ → ~~2,509 lines (5.4% of plugin lines)~~ → 2,409 lines (5.2% of plugin lines; unchanged at `baa2f5d`) |
| Generated flat mirror files (byte copies of `.apm/`) | ~~263 files, ~22,000 lines~~ → 0 (deleted 2026-09-14, see below) |
| `docs/research/` vendored inside plugins | ~19,000 lines, nothing executable reads it |
| Repo-level `docs/research/` + `docs/notes/` | 4,500 lines, 47% of all prose words, 6 of 11 research files linked only from each other |
| Repo-level `docs/research/` + `docs/notes/` | 4,500 lines, 47% of all prose words, 6 of 11 research files linked only from each other (invalidated by this document's own move into `docs/notes/`, which adds 665 lines to the row it measures) |
| Enforcement: hook entries in `.pre-commit-config.yaml` / pre-push hooks | ~~33 / 14~~ → 26 / 8 (at `baa2f5d`; see the note below) |
| Enforcement: `tests/*.sh` + runners + `scripts/` | ~~12,400 + 475 + 4,500 lines~~ → ~~9,123 + 490 + 3,308 lines~~ → ~~10,189 + 502 + 2,901~~ → ~~10,608 + 502 + 3,139~~ → ~~10,000 + 502 + 1,924 (at `4b17703`)~~ → 10,088 + 514 + 1,926 (at `baa2f5d`) |
| Validator scripts inside kyberforge (+ their bats tests) | ~~6,800 + 5,300 lines~~ → ~~5,861 + 6,015~~ → ~~5,876 + 6,015~~ → 5,885 + 6,071 (at `baa2f5d`) |
| Preload tax (39 skill names + descriptions) | 10,987 chars, ~2,750 tokens per session |
| Preload tax (~~39~~ → 38 skill names + descriptions) | 10,987 chars, ~2,750 tokens per session (measured at 39 skills on 2026-09-10; never re-measured after ADR-0025's merge took the count to 38) |
| Commits since 2026-05-10 / share touching hook, test, gate, vale, or sync | 447 / ~25% |
> **Corrected then done (2026-09-14):** the mirror row's figure was wrong. The true mirror was **213 files / 20,061 lines**, not 263 / ~22,000 — the original count swept in files that were never mirror output. All 213 were deleted in commit `718c79a` on `docs/simplification-audit` (245 files changed, 298 insertions, 22,602 deletions across the whole change), so the row is now zero. The enforcement row is stale on **both** halves — it was correct at the 2026-09-10 baseline (`9eb8bc7`: 33 `- id:` entries, 14 repo-authored pre-push hooks), but `.pre-commit-config.yaml` today has ~~**27 entries and 9 `stages: [pre-push]`**~~ → ~~**26 entries and 8 `stages: [pre-push]`**~~ → ~~**27 entries and 9 `stages: [pre-push]`**~~ → **26 entries and 8 `stages: [pre-push]`** (~~`467bbd7`~~ → `620f20b` removed `check-vale-style-sync` with finding 14's merge; `8451169` then added `check-skill-version-bump`; `4de5b6b` then removed `check-release-needed`; re-measured 2026-09-16 at `4b17703` with `grep -c -- "- id:"` and `grep -c "stages: \[pre-push\]"` on `.pre-commit-config.yaml`). Like for like that is 14 → ~~9~~ → ~~8~~ → ~~9~~ → 8 repo-authored pre-push hooks. The stage *reports* ~~11~~ → ~~10~~ → ~~11~~ → 10, because the 2 pre-commit `meta` hooks also run there — a different counting basis; see the corrected §3 target, which states it the same way.
@@ -165,7 +165,7 @@ This is the area you named as hardest to understand and slowest. Root cause: mos
Pre-commit stays roughly as is minus `skill-frontmatter`, and minus `check-ast` once finding 9 removes the only `.py` files. ~~Tests 26 files to about 10 (12,400 to about 5,000 lines).~~ Keep bats and its three submodules; the 351 bats tests ship inside plugins and are the right tool there. ~~Do not port the bash suites to bats; delete them instead.~~ **Struck (2026-09-16, grill):** see finding 8's closing note — the suites are regression coverage (findings 3 and 5; finding 16 found the same of the validators they test).
> **Re-measured (2026-09-14, at `a6434e0`):** the tests target was stated against the 2026-09-10 baseline and both its numbers are stale. `tests/` now holds **20 `test-*.sh` suites totalling 9,123 lines** (plus the two runners, 490). Six suites have gone since the baseline: `test-check-manifests.sh` (`e647f14`), `test-skill-frontmatter.sh` (`c8a7c9e`), `test-governance-layer.sh` and `test-instructions-and-docs.sh` (`5f9f2b3`), `test-sync-marketplace-mirror.sh` (`0dffff3`), `test-sync-plugin-content.sh` (`718c79a`). ~~Restated on the same basis the target is **20 files to about 10, 9,123 to about 5,000 lines**~~ — **struck (2026-09-16):** the target itself is withdrawn (see the struck sentence above); for the record, `tests/` holds **19** suites totalling **10,000** lines at `4b17703`, after `4de5b6b` deleted `test-check-release-needed.sh` and `test-vale-hooks-consumer.sh`. Finding 9's `check-ast` clause is moot anyway, since finding 9 is not proceeding.
> > **Corrected (2026-09-20, at `1614bce`) — the deletion tally is nine, not ~~six~~ → ~~eight~~.** The six named above plus the two the 2026-09-16 strike adds come to eight, and a ninth was never folded into the running tally: **`test-check-vale-style-sync.sh`**, removed by `620f20b` with the `factory-audit` merge (finding 14) — the same commit finding 2's bullet already credits for deleting that gate's hook and script. The full `main...HEAD` set is nine: `test-check-manifests.sh` (`e647f14`), `test-check-release-needed.sh` (`4de5b6b`), `test-check-vale-style-sync.sh` (`620f20b`), `test-governance-layer.sh` and `test-instructions-and-docs.sh` (`5f9f2b3`), `test-skill-frontmatter.sh` (`c8a7c9e`), `test-sync-marketplace-mirror.sh` (`0dffff3`), `test-sync-plugin-content.sh` (`718c79a`), `test-vale-hooks-consumer.sh` (`4de5b6b`). Method: `git diff --name-status main...HEAD -- tests/ | grep '^D'`. The pinned "19 suites at `4b17703`" is unaffected — `620f20b` precedes that commit, so the file count already reflected the deletion even though the tally did not. At `1614bce` `tests/` holds **19** `test-*.sh` suites totalling **10,897** lines.
> > **Corrected (2026-09-20, at `1614bce`) — the deletion tally is nine, not ~~six~~ → ~~eight~~.** The six named above plus the two the 2026-09-16 strike adds come to eight, and a ninth was never folded into the running tally: **`test-check-vale-style-sync.sh`**, removed by `620f20b` with the `factory-audit` merge (finding 14) — the same commit finding 2's bullet already credits for deleting that gate's hook and script. The full `main...HEAD` set is nine: `test-check-manifests.sh` (`e647f14`), `test-check-release-needed.sh` (`4de5b6b`), `test-check-vale-style-sync.sh` (`620f20b`), `test-governance-layer.sh` and `test-instructions-and-docs.sh` (`5f9f2b3`), `test-skill-frontmatter.sh` (`c8a7c9e`), `test-sync-marketplace-mirror.sh` (`0dffff3`), `test-sync-plugin-content.sh` (`718c79a`), `test-vale-hooks-consumer.sh` (`4de5b6b`). Method: `git diff --name-status main...HEAD -- tests/ | grep '^D'`. The pinned "19 suites at `4b17703`" is unaffected — `620f20b` precedes that commit, so the file count already reflected the deletion even though the tally did not. At `1614bce` `tests/` holds **19** `test-*.sh` suites totalling ~~**10,897**~~ → **10,588** lines. (**Corrected 2026-09-20:** the suite count was right and the line total was not — 10,897 is the value at `384756b`, the commit that added the hook-wiring tests, and at `1ec3e8a`; at `1614bce` the nineteen suites total 10,588.)
## 4. Plugins
@@ -179,19 +179,19 @@ The shared pattern: per-skill `README.md` files no model reads, a `docs/research
10. [x] ~~**Delete per-skill `README.md` and `references/README.md` (48 files, 1,574 lines).** They restate the SKILL.md in narrative form. The pre-commit config itself notes a skill README "is consumer-facing prose that no agent ever loads". Keep one plugin-level README with one line per skill. Requires dropping the README criterion in `skill-audit/references/file-structure.md` and the README step in `new-skill.sh`. Effort S.~~
> **Done (2026-09-12):** see commit `edcc57c` on `docs/simplification-audit`. Deleted the 48 per-skill/reference READMEs plus 2 scaffold templates; dropped the README criterion from `skill-audit`'s `file-structure.md` and `finding-criteria.md` and the README-generation step from `new-skill.sh`; updated `new-skill.bats` to match. Plugin-root READMEs were kept, not part of this finding.
11. [x] **Drop the provenance chain: `sources.md`, `source_keys` frontmatter, `validate-provenance.sh`.** 32 plugin and skill `sources.md` files (about 1,300 lines) plus 9 research indexes, 216 source files with `source_keys`, ~~two copies of the validator (1,198 and 632 lines)~~ → **one validator, 2,171 lines across four files**, with ten checks, and ~~125 bats tests~~ → **138 bats tests** exist to track which upstream informed which file. Git blame and a URL in the README do the same job. This is more code than the content it tracks. Effort M (touches ~~skill-audit, both validator copies~~ → **`factory-audit`, its one provenance validator**, two repo tests, and every skill's frontmatter).
> **Re-measured (2026-09-16, at HEAD):** ADR-0025 merged the two copies, so the "two copies" arithmetic throughout this finding and its note below no longer resolves. The provenance validator is now `factory-audit/scripts/` `validate-provenance.sh` (320) + `lib-provenance-skill.sh` (1,145) + `lib-provenance-agent.sh` (572) + `lib-contributing-files.sh` (134) = **2,171** lines (`wc -l` on the four), against **3,209** bats lines (`validate-provenance-skill.bats` 2,062 + `validate-provenance-agent.bats` 1,147) carrying **138** cases (`grep -c '^@test'`). Note this is *more* than the 1,198 + 632 = 1,830 the finding counted, not less: the merge deduplicated the resolver and the Contributing-files parser, not the per-mode provenance checks, and the shared entry script added the exit-tier and library guards described in `docs/spec/gates.md`. The `sources.md` census also moved: **45 files / 1,756 lines** — 27 skill `references/sources.md` (1,207), 13 research indexes (435), 4 plugin-root (100), 1 scaffold template (14). The note below's 46 / 1,752 swept in `docs/adr/0013-vale-harness-scope-and-rule-sources.md`, which matches `sources\.md$` and is not one. Imbalance at HEAD: **5,380 validator+bats lines against 1,756 of metadata, 3.1:1** — worse than the 2.6:1 below, on the same direction of argument.
11. [x] **Drop the provenance chain: `sources.md`, `source_keys` frontmatter, `validate-provenance.sh`.** 32 plugin and skill `sources.md` files (about 1,300 lines) plus 9 research indexes, 216 source files with `source_keys`, ~~two copies of the validator (1,198 and 632 lines)~~ → **one validator, ~~2,171~~ → 2,186 lines across four files**, with ten checks, and ~~125 bats tests~~ → **138 bats tests** exist to track which upstream informed which file. Git blame and a URL in the README do the same job. This is more code than the content it tracks. Effort M (touches ~~skill-audit, both validator copies~~ → **`factory-audit`, its one provenance validator**, two repo tests, and every skill's frontmatter).
> **Re-measured (2026-09-16, at HEAD):** ADR-0025 merged the two copies, so the "two copies" arithmetic throughout this finding and its note below no longer resolves. The provenance validator is now `factory-audit/scripts/` `validate-provenance.sh` (~~320~~ → **324**) + `lib-provenance-skill.sh` (~~1,145~~ → **1,152**) + `lib-provenance-agent.sh` (~~572~~ → **576**) + `lib-contributing-files.sh` (134) = ~~**2,171**~~ → **2,186** lines (`wc -l` on the four; **corrected 2026-09-20** — the four struck figures never reproduced at any commit, and `wc -l` gives 324 / 1,152 / 576 / 134 at `620f20b`, the commit that created the files, and at every commit since, `1ec3e8a` included), against **3,209** bats lines (`validate-provenance-skill.bats` 2,062 + `validate-provenance-agent.bats` 1,147) carrying **138** cases (`grep -c '^@test'`). Note this is *more* than the 1,198 + 632 = 1,830 the finding counted, not less: the merge deduplicated the resolver and the Contributing-files parser, not the per-mode provenance checks, and the shared entry script added the exit-tier and library guards described in `docs/spec/gates.md`. The `sources.md` census also moved: **45 files / 1,756 lines** — 27 skill `references/sources.md` (1,207), 13 research indexes (435), 4 plugin-root (100), 1 scaffold template (14). The note below's 46 / 1,752 swept in `docs/adr/0013-vale-harness-scope-and-rule-sources.md`, which matches `sources\.md$` and is not one. Imbalance at HEAD: ~~**5,380**~~ → **5,395 validator+bats lines against 1,756 of metadata, 3.1:1** (2,186 + 3,209; the ratio is unchanged at 3.07) — worse than the 2.6:1 below, on the same direction of argument.
> **Verified (2026-09-14, at HEAD `062ca47`):** direction defensible, two scope figures wrong, and **blocked on a decision the finding never poses**. The `sources.md` census below is exact, and so are the finding's own validator and bats figures (1,198 / 632 lines, 125 bats tests); the scope errors are narrower than an earlier revision of this note claimed.
>
> Corrected figures: **46 `sources.md` files / 1,752 lines** in three distinct classes — 29 skill `references/sources.md` (1,217 lines), 13 research indexes (435), 4 plugin-root files (100, ADR-0010). The finding does **not** double-count: it states two disjoint classes additively ("32 plugin and skill `sources.md` files (about 1,300 lines) **plus** 9 research indexes"), and that plugin-and-skill subtotal is really **33 files / 1,317 lines**, matching its "about 1,300" exactly — had the 32 swept in the research indexes the figure would have been ~1,750. Its real errors there are an off-by-one (32 should be 33) and an omission: it missed the 4 vendored example indexes under `kyberforge/docs/research/examples/skill-write/`, so 9 should be 13. Carriers of `source_keys` in YAML frontmatter: **196** — 168 at column 0 and 28 nested two spaces under `metadata:` — so the finding's 216 is closer to the truth than it looks. (219 files merely *mention* the string. A naive `^[[:space:]]*source_keys:` grep returns 200, but 4 of those are heredoc or fixture text rather than frontmatter: both `validate-provenance.bats` copies, `scripts/check-scope-walkup-sync.sh`, and a fenced example in `plugins/bin/.apm/skills/research/references/file-format.md`.) Checks: **16 across the two copies** (skill-audit 0–9, agent-audit 0–5), not ten. Validator line counts (1,198 / 632) and 125 bats tests are exact.
>
> **"Touches every skill's frontmatter" is roughly right.** ~~**28 of the 39 real skills carry `source_keys` in frontmatter**~~ → **27 of the 38** (re-measured 2026-09-16 at HEAD; the audit-pair merge took one carrier skill with it), nested under `metadata:` — see `plugins/git/.apm/skills/git-commits/SKILL.md:10-17`, where `metadata:` → `source_keys:` carries four slugs. (~~44~~ → **43** tracked files match `*SKILL.md`; subtract `skill-author/assets/templates/SKILL.md` and the 4 vendored under `kyberforge/docs/research/examples/skill-write/`, leaving ~~39~~ → **38** real skills.) The 11 without it are exactly the `plugins/bin/` skills. Check 2 in the skill-side validator (SKILL.md `source_keys` → slug in `sources.md`) is correspondingly **live**, not dead code: `parse_source_keys()` at `plugins/kyberforge/.apm/skills/factory-audit/scripts/lib-provenance-skill.sh:277-305` handles both spellings explicitly — the metadata-nested branch at `:292`, the top-level branch at `:295`, and a docstring that says "handles metadata.source_keys and top-level" — check 2 at `:712` runs against all ~~28~~ → **27** carrier skills, every one of which has a `references/sources.md`, and bats pins it at `plugins/kyberforge/.apm/skills/factory-audit/tests/validate-provenance-skill.bats:222` ("FAIL: source_keys slug in SKILL.md not present as H2 in sources.md") and ~~`:1337`~~ → `:1338` (a BOM must not silently disable check 2). (Paths and line numbers re-derived at HEAD: ADR-0025's merge moved this code out of `skill-audit/scripts/validate-provenance.sh` into the shared skill-side library, so the figures this note carried at `062ca47` — `:242-270`, `:257`, `:260`, `:766`, `:1313` — no longer resolve.) The imbalance the finding names is real and **worse** than claimed: ~~4,641 validator+bats lines against 1,752 of metadata, a 2.6:1 ratio~~ → **5,380 against 1,756, a 3.1:1 ratio** (re-measured 2026-09-16 at HEAD; see the note under the headline).
> **"Touches every skill's frontmatter" is roughly right.** ~~**28 of the 39 real skills carry `source_keys` in frontmatter**~~ → **27 of the 38** (re-measured 2026-09-16 at HEAD; the audit-pair merge took one carrier skill with it), nested under `metadata:` — see `plugins/git/.apm/skills/git-commits/SKILL.md:10-17`, where `metadata:` → `source_keys:` carries four slugs. (~~44~~ → **43** tracked files match `*SKILL.md`; subtract `skill-author/assets/templates/SKILL.md` and the 4 vendored under `kyberforge/docs/research/examples/skill-write/`, leaving ~~39~~ → **38** real skills.) The 11 without it are exactly the `plugins/bin/` skills. Check 2 in the skill-side validator (SKILL.md `source_keys` → slug in `sources.md`) is correspondingly **live**, not dead code: `parse_source_keys()` at `plugins/kyberforge/.apm/skills/factory-audit/scripts/lib-provenance-skill.sh:`~~`277-305`~~ → `:284-312` handles both spellings explicitly — the metadata-nested branch at ~~`:292`~~ → `:299`, the top-level branch at ~~`:295`~~ → `:302`, and a docstring that says "handles metadata.source_keys and top-level" — check 2 at ~~`:712`~~ → `:719` runs against all ~~28~~ → **27** carrier skills, every one of which has a `references/sources.md`, and bats pins it at `plugins/kyberforge/.apm/skills/factory-audit/tests/validate-provenance-skill.bats:222` ("FAIL: source_keys slug in SKILL.md not present as H2 in sources.md") and ~~`:1337`~~ → `:1338` (a BOM must not silently disable check 2). (Paths and line numbers re-derived at HEAD: ADR-0025's merge moved this code out of `skill-audit/scripts/validate-provenance.sh` into the shared skill-side library, so the figures this note carried at `062ca47` — `:242-270`, `:257`, `:260`, `:766`, `:1313` — no longer resolve.) **Corrected (2026-09-20):** the four "re-derived at HEAD" citations into `lib-provenance-skill.sh` were themselves uniformly 7 lines low and never resolved at any commit; they are repointed above. The two `validate-provenance-skill.bats` citations (`:222`, `:1338`) do resolve and are left alone. The imbalance the finding names is real and **worse** than claimed: ~~4,641 validator+bats lines against 1,752 of metadata, a 2.6:1 ratio~~ → ~~**5,380**~~ → **5,395 against 1,756, a 3.1:1 ratio** (re-measured 2026-09-16 at HEAD; see the note under the headline).
>
> **Omitted entirely: the chain has a producer.** `plugins/bin/.apm/skills/research/` *specifies* the `sources.md` + `source_keys:` output format, and `plugins/bin/evals/research/research/eval.yaml` carries three criteria asserting it. **This is the blocking decision: does `research` keep emitting `sources.md`?** If yes, the chain is not dropped — only unenforced, and the finding collapses to "delete the validators." If no, the research skill's output contract and its evals need redesigning.
>
> Also breaks: `check-scope-walkup-sync` loses one of four walk-up ports (the hook exists because three scripts drifted); `tests/test-adr0020-contract.sh` loses its parser byte-identity assertion; `tests/test-check-scope-walkup-sync.sh` must re-base its fixture; ADR-0010 is superseded outright and ADR-0009/0016 need amending (`field-inventory.md`'s allowlist data line carries `source_keys`). `LESSONS.md:73` records this validator as the **only** thing that catches a skill authored outside `skill-author` — a failure that "recurred twice in one session" — so "git blame + a README URL do the same job" is false for the one thing the chain demonstrably catches. Side effect: 55 reference files have frontmatter containing *only* `source_keys:`, leaving empty `---\n---` blocks to delete.
>
> **Effort L, not M** (about ~~6,393~~ → **7,136** lines deleted across 242 files: the ~~4,641~~ → **5,380** validator and bats lines plus the ~~1,752~~ → **1,756** of `sources.md` measured above, across 196 `source_keys` carriers and ~~46~~ → **45** `sources.md` files. An earlier revision of this note said ~4,600 lines across ~230 files, which was internally inconsistent — 4,600 is validator-plus-bats only and silently drops the `sources.md` this same note measures, and ~230 inherited a carrier count of 172 that missed every `metadata:`-nested file.) Smaller alternative worth considering: scope the drop to the skill half only (~~1,217 lines, 1,198-line validator, 82 tests~~ → **1,207 lines of skill `sources.md`, the 1,145-line `lib-provenance-skill.sh`, 87 tests**, re-measured 2026-09-16 at HEAD) and leave the ADR-0010 plugin-root half alone — no ADR supersession needed.
> **Effort L, not M** (about ~~6,393~~ → ~~**7,136**~~ → **7,151** lines deleted across ~~242~~ → **241** files: the ~~4,641~~ → ~~**5,380**~~ → **5,395** validator and bats lines plus the ~~1,752~~ → **1,756** of `sources.md` measured above, across 196 `source_keys` carriers and ~~46~~ → **45** `sources.md` files — 196 + 45 = 241, and the struck 242 was consistent only with the struck 46. An earlier revision of this note said ~4,600 lines across ~230 files, which was internally inconsistent — 4,600 is validator-plus-bats only and silently drops the `sources.md` this same note measures, and ~230 inherited a carrier count of 172 that missed every `metadata:`-nested file.) Smaller alternative worth considering: scope the drop to the skill half only (~~1,217 lines, 1,198-line validator, 82 tests~~ → **1,207 lines of skill `sources.md`, the ~~1,145~~ → 1,152-line `lib-provenance-skill.sh`, 87 tests**, re-measured 2026-09-16 at HEAD) and leave the ADR-0010 plugin-root half alone — no ADR supersession needed.
>
> **Decision (2026-09-16):** Not proceeding — the human declined this finding. The provenance chain (`sources.md`, `source_keys:`, `validate-provenance.sh`) stays, and `research` keeps producing it. This also answers §8's provenance question.
@@ -229,7 +229,7 @@ The shared pattern: per-skill `README.md` files no model reads, a `docs/research
>
> **Re-measured (2026-09-16, at HEAD) — the basis of every figure below changed when ADR-0025 landed; the refutation is unaffected.** There are no longer three validators or two `vale-wrap.sh` copies. The headline's "ported twice" is void, and its `1,677` and `526` no longer name anything. At HEAD: `scripts/skill-size-check.sh` is **1,522** (the note below's 1,517 was correct at `a6434e0`); `factory-audit`'s validator is **2,663** lines across four files (`validate.sh` 255 + `lib-checks-skill.sh` 621 + `lib-checks-agent.sh` 683 + `lib-boundary-resolver.sh` 1,104); `vale-wrap.sh` is **535**, one copy. Validator total **4,185**, of which the resolver is **2,165** (the 1,061-line block still embedded in `skill-size-check.sh`, plus `lib-boundary-resolver.sh`'s 1,104 — the same 1,061 block wrapped in 43 lines of library preamble, which is why the byte-identity test compares the block and not the files). So the resolver is now **52%** of validator lines, not 65%, and **2,020** lines remain once it is excised, not 1,749. Tests: the six repo suites over `skill-size-check.sh` are **3,907** (was 3,619) and the two in-skill validator bats files **2,248** (`validate-skill.bats` 1,029 + `validate-agent.bats` 1,219), for **6,155**, not 5,506. The 200-line target is off by the same order of magnitude it was. (All figures `wc -l`; the resolver block by `awk '/BEGIN ADR-0020 SHARED BOUNDARY RESOLVER/,/END .../'`.)
>
> > **Superseded by `ef27c97` (re-measured 2026-09-19, at HEAD).** The paragraph above is a dated snapshot and its two load-bearing claims no longer hold. `scripts/skill-size-check.sh` is **509** lines, not 1,522 — it shrank by 1,013 — and the resolver is **no longer embedded in it**: `ef27c97` excised the 1,061-line block and the hook now sources `factory-audit`'s `lib-boundary-resolver.sh` by path (`RESOLVER_LIB` at `:483`, `. "$RESOLVER_LIB"` at `:492`), failing closed if the library is missing or defines no resolver. The single remaining `BEGIN ADR-0020 SHARED BOUNDARY RESOLVER` string in the hook is that fail-closed guard, not a copy. `factory-audit`'s four validator files now total **2,671** (`validate.sh` 255 + `lib-checks-skill.sh` 627 + `lib-checks-agent.sh` 685 + `lib-boundary-resolver.sh` 1,104) and `vale-wrap.sh` is **536**. So there is **one** resolver copy repo-wide, not two, and the "resolver is 52% of validator lines" arithmetic above is void along with its inputs. Only the refutation of finding 16 survives all of this unchanged.
> > **Superseded by `ef27c97` (re-measured 2026-09-19, at HEAD).** The paragraph above is a dated snapshot and its two load-bearing claims no longer hold. `scripts/skill-size-check.sh` is **509** lines, not ~~1,522~~ → **1,524** — it shrank by ~~1,013~~ → **1,015** — and the resolver is **no longer embedded in it**: `ef27c97` excised the 1,061-line block and the hook now sources `factory-audit`'s `lib-boundary-resolver.sh` by path (`RESOLVER_LIB` at `:483`, `. "$RESOLVER_LIB"` at `:492`), failing closed if the library is missing or defines no resolver. The single remaining `BEGIN ADR-0020 SHARED BOUNDARY RESOLVER` string in the hook is that fail-closed guard, not a copy. `factory-audit`'s four validator files now total **2,671** (`validate.sh` 255 + `lib-checks-skill.sh` 627 + `lib-checks-agent.sh` 685 + `lib-boundary-resolver.sh` 1,104) and `vale-wrap.sh` is **536**. So there is **one** resolver copy repo-wide, not two, and the "resolver is 52% of validator lines" arithmetic above is void along with its inputs. Only the refutation of finding 16 survives all of this unchanged. (**Corrected 2026-09-20:** this note originally read "not 1,522 — it shrank by 1,013", which contradicted the "−1,015" the closing note below states for the same commit. 1,522 was a stale pre-`ef27c97` reading: `git show ef27c97^:scripts/skill-size-check.sh | wc -l` is **1,524** and `ef27c97` is **509**, so the delta is **−1,015** in both places.)
>
> The three validators are **not three implementations**. They contain **one block, 1,061 lines, byte-identical in all three**, delimited by `# ===== BEGIN/END ADR-0020 SHARED BOUNDARY RESOLVER =====` and hashed by `tests/test-adr0020-contract.sh`. So 3,183 of 4,932 validator lines (65%) are that block × 3, and **what is left once the resolver is excised is 1,749 lines across all three** — 1,580 non-blank, 992 with comments and blanks both stripped. The duplication is forced by the self-containment constraint, which is why *merging* is the lever and *shrinking* is not.
>
@@ -421,9 +421,9 @@ Not covered by the area audits above; found on a final sweep of the root config
> Corrected headline: ~~**two** hand-maintained per-plugin locations (**three** for kyberforge)~~ → **one** hand-maintained per-plugin version location, `plugins/<name>/apm.yml` (**two** for kyberforge, adding the `executables.allow` key), not four. `2def060` deleted the root `packages[].version` lines (corrected 2026-09-16, review round). The root `packages[].description:` duplicates dropped in the same round are a separate duplication, not a version location, so they do not change this count — the audit's own "already done" note records the `plugin.json` deletion but never fixed the headline. Gitea skills drift across **six** values (`0.1.2, 0.1.3, 0.1.4, 0.1.5, 0.1.6, 1.0.1`), not five — ~~still six at HEAD on 2026-09-16~~ → **five** again at HEAD (`b426460`) on 2026-09-16 (`0.1.2, 0.1.4, 0.1.5, 0.1.6, 1.0.1`), because `8451169` bumped `gitea-branches` 0.1.3 → 0.1.4 under the new version-bump gate and it was the only skill at 0.1.3; re-derived by parsing `metadata.version` out of each `plugins/gitea/.apm/skills/*/SKILL.md` with PyYAML. ~~39 `SKILL.md` files ✓~~ → **38** carry it, and all 38 do (re-measured 2026-09-16; ADR-0025's merge took one). The `0.4.6` duplication between root `version:` and `marketplace.version:` is **forced by apm, not a repo choice** — deleting `marketplace.version` makes `--check-clean` go dirty.
>
> **"Nothing consumes `metadata.version`" is false twice over.** Machine enforcers: ~~`scripts/skill-size-check.sh:1365-1374`~~ → ~~`scripts/skill-size-check.sh:1370-1379`~~ → `scripts/skill-size-check.sh:323-335` and ~~`skill-audit/scripts/validate.sh:1292-1332`~~ → `plugins/kyberforge/.apm/skills/factory-audit/scripts/lib-checks-skill.sh:235-283`, both FAIL tier, the latter citing ADR-0022 by name, with four dedicated bats cases and ~10 fixture generators baking the field in.
> Instruction-level consumers: `skill-author/SKILL.md:60` (bump minor on create, patch on improve), `create.md:89,101`, `improve.md:82`, and `forge/SKILL.md:54` + `references/version-bump.md`. apm parses it for Chatmode/Instruction/Context primitives but not for Skills, and never emits it. Precise statement: the value is written, shape-validated, and never read *downstream* — it is an agent-visible revision counter, and the drift table shows the counter is not being maintained.
> Instruction-level consumers: `skill-author/SKILL.md:60` (bump minor on create, patch on improve), `create.md:89,101`, ~~`improve.md:82`~~ → `improve.md:105`, and `forge/SKILL.md:54` + `references/version-bump.md`. apm parses it for Chatmode/Instruction/Context primitives but not for Skills, and never emits it. Precise statement: the value is written, shape-validated, and never read *downstream* — it is an agent-visible revision counter, and the drift table shows the counter is not being maintained.
>
> > **Repointed (2026-09-16, at HEAD; re-verified and corrected 2026-09-19):** `skill-audit/scripts/validate.sh` no longer exists — ADR-0025's merge moved the ADR-0022 check into `factory-audit`'s skill-side check library, where it is the `SEMVER_RE` block: comment header at `:235`, `SEMVER_RE` itself at `:254`, `fail()` calls at ~~`:261` and `:275`~~ → `:265` and `:280`, the block running `:235-283` (the next section header, `# SKILL.md size ceilings`, is at `:285`). That library is **627** lines, not 621. In `skill-size-check.sh` the check is at `:323-335`; the earlier note said the file "grew by 5 lines above the block", which is the wrong direction by two orders of magnitude — `ef27c97` excised the embedded resolver and the file **shrank** from 1,522 to **509** lines, which is why the range moved from the 1,300s to the 320s. All five instruction-level citations still resolve at HEAD, verified with `sed -n`.
> > **Repointed (2026-09-16, at HEAD; re-verified and corrected 2026-09-19):** `skill-audit/scripts/validate.sh` no longer exists — ADR-0025's merge moved the ADR-0022 check into `factory-audit`'s skill-side check library, where it is the `SEMVER_RE` block: comment header at `:235`, `SEMVER_RE` itself at `:254`, `fail()` calls at ~~`:261` and `:275`~~ → `:265` and `:280`, the block running `:235-283` (the next section header, `# SKILL.md size ceilings`, is at `:285`). That library is **627** lines, not 621. In `skill-size-check.sh` the check is at `:323-335`; the earlier note said the file "grew by 5 lines above the block", which is the wrong direction by two orders of magnitude — `ef27c97` excised the embedded resolver and the file **shrank** from 1,522 to **509** lines, which is why the range moved from the 1,300s to the 320s. ~~All five instruction-level citations still resolve at HEAD, verified with `sed -n`.~~ → **Corrected (2026-09-20): four of the five resolve, not five.** `improve.md:82` stopped carrying the `metadata.version` content at `baa2f5d`, three days before the 2026-09-19 verification claim was written, so that claim was false when made; the content is at ~~`improve.md:82`~~ → `improve.md:105` ("A skill carrying no `metadata.version` is seeded at `"1.0.0"`, not bumped"). The other four — `skill-author/SKILL.md:60`, `create.md:89`, `create.md:101`, `forge/SKILL.md:54` — do resolve at `1ec3e8a`.
>
> **ADR-0022 already considered and rejected dropping the field**, on the grounds that `skill-author` depends on it to decide whether a pass owes a bump — a rationale still live today. Superseding costs: rewrite skill-author's bump rule, delete `forge`'s version-bump route premise, strip two scripts, delete four bats cases, fix ~10 fixture generators, edit the scaffold template, update ~~`gates.md:97`~~ → ~~`gates.md:145`~~ → `gates.md:146` — and re-open the "is this field present here?" question issue #127 closed, just from the other side. *(Repointed 2026-09-16, at HEAD `b426460`: the `metadata.version` frontmatter sentence formerly at `gates.md:97` was at `:143-146`, the field itself on `:145`, and is at `:143-147` / `:146` at `4b17703`; verified with `grep -n "metadata.version" docs/spec/gates.md`.)* **Recommendation: keep it and fix the actual defect, which is that nobody bumps it.** Either enforce the bump in the skill-author workflow or declare the values advisory in the ADR.
>

View File

@@ -27,7 +27,7 @@ Skills are **not** deployed by `install.sh`. They are distributed as plugins and
Skills, agents, MCP servers, and hooks are distributed as self-contained plugin units under `plugins/`, installed independently via `apm install`, here and in any consuming repo (ADR-0018). Each unit is an **apm package**: `plugins/<name>/apm.yml` plus a hand-authored `plugins/<name>/.apm/{skills,agents,hooks,commands,instructions,extensions}/` tree (ADR-0015). There is no per-plugin `plugin.json` at all — apm reads `apm.yml`, and the repo's one generated manifest, `.claude-plugin/marketplace.json`, is compiled from that source.
Self-contained is a hard constraint, not a description: a file reference inside `.apm/skills/<name>/` may not reach outside that skill's own directory, and there is no cross-skill sharing mechanism to reach for instead. That is why the Vale styles ship inside the one skill that uses them, `factory-audit/assets/vale/` (ADR-0014, ADR-0025), and why ADR-0020's constants are copied into two validators — the plugin's `validate.sh` and the repo's `scripts/skill-size-check.sh` — rather than sourced from one. The constraint used to be explained by Claude Code's plugin cache-install copying a plugin to a cache; that is no longer the reason and never was the only one. It is stated independently for APM package mode by the agentskills.io spec (`plugins/kyberforge/.apm/skills/skill-author/references/deployment-modes.md`), which is why ADR-0024 consequence 6 pins it as a negative result: ending native install did not relax it, and it is not to be re-litigated on the assumption that it did.
Self-contained is a hard constraint, not a description: a file reference inside `.apm/skills/<name>/` may not reach outside that skill's own directory, and there is no cross-skill sharing mechanism to reach for instead. That is why the Vale styles ship inside the one skill that uses them, `factory-audit/assets/vale/` (ADR-0014, ADR-0025), and why ADR-0020's constants are copied rather than sourced from one place: `scripts/skill-size-check.sh` carries them, and so do `factory-audit`'s mode libraries — `scripts/lib-checks-skill.sh:313-316` all four, `scripts/lib-checks-agent.sh:164-165` the two description ones. The plugin's `validate.sh` carries none of them; it sources the library its mode selects. The constraint used to be explained by Claude Code's plugin cache-install copying a plugin to a cache; that is no longer the reason and never was the only one. It is stated independently for APM package mode by the agentskills.io spec (`plugins/kyberforge/.apm/skills/skill-author/references/deployment-modes.md`), which is why ADR-0024 consequence 6 pins it as a negative result: ending native install did not relax it, and it is not to be re-litigated on the assumption that it did.
Which apm package a new skill belongs in follows from what each one is scoped to. The boundary that matters most in practice is `core` vs `kyberforge`: `core` is the home for cross-cutting, repo-agnostic utility skills that a consumer would want against *their* repo, while `kyberforge` is meta-tooling for the holocron marketplace itself. A skill that authors a target repo's `AGENTS.md` is `core`; a skill that audits a `SKILL.md` against this marketplace's contract is `kyberforge`.

View File

@@ -98,19 +98,32 @@ ADR-0022 makes `metadata.version` mandatory and says a skill change carries a bu
- **It runs on every push and under a manual `pre-commit run --hook-stage pre-push`.** It does not
read `PRE_COMMIT_REMOTE_BRANCH`, so the manual rehearsal really checks it. The pushed commit is `PRE_COMMIT_TO_REF`, or `HEAD` when that is unset.
- **"Changed" is measured from the merge-base of the pushed commit with `origin/main`** (local
`main` if `origin/main` does not resolve). Readers install from `main`, so "changed" means
changed against the `main` the branch started from. The remote branch tip is not the baseline:
a second push would excuse an unbumped change the first push already carried.
- **A changed skill's version must beat two baselines**: its version at that merge-base *and* its
- **"Changed" is measured from the merge-bases of the pushed commit with `origin/main`** (local
`main` if `origin/main` does not resolve), resolved with **`git merge-base --all`** — all of
them, not the single one git would otherwise pick. Readers install from `main`, so "changed"
means changed against the `main` the branch started from. The remote branch tip is not the
baseline: a second push would excuse an unbumped change the first push already carried.
- **With more than one base, the changed-skill sets are intersected.** A criss-cross history —
`main` merges a branch while that branch merges a `main` commit — has two merge-bases, and which
one a bare `git merge-base` prints is an implementation detail, so picking one made the verdict a
coin flip: a skill already identical to `main` was reported `(not above merge-base)` whenever the
losing base was chosen. A skill therefore counts as changed only when it differs from **every**
base; differing from none of them, or from only some, means a base already carries the pushed
content. A skill that does count as changed must then beat the version at every base it exists
at. Both directions are conservative: the intersection cannot exempt a skill that changed since
all of `main`'s reachable history, and requiring every base keeps the ratchet.
- **A changed skill's version must beat two baselines**: its version at each merge-base *and* its
version at the tip of the same `main` ref (ADR-0022's second 2026-09-16 amendment). The tip
check stops two branches that make the same bump (`1.0.0` → `1.0.1`) with different content from
both landing, since the identical version lines merge without a conflict. A skill absent at the
tip is held to the merge-base alone, and so is one whose directory at the pushed commit is the
*same tree object* as at the tip: it ships exactly what main ships, whatever route the history
took there — a criss-cross merge, a cherry-pick, a backport — so there is nothing for a bump to
announce. When `main` has not moved, the two baselines are the same commit. Each
failure line names the baseline it missed: `(not above merge-base)` or
tip is held to the merge-bases alone, and so is one whose directory at the pushed commit is the
**same tree object** as at the tip — compared as object ids, because a tree id *is* the content
whatever route the history took to it. That skill ships exactly what `main` ships, so there is
nothing for a bump to announce. The intersection does not already cover it: it exempts only when
some base carries the content, which a criss-cross history gives and a cherry-pick of a fix
`main` already has does not. When `main` has not moved, the tip is itself a base and the skill is
checked once. Each failure line names the baseline it missed: `(not above merge-base)`,
`(not above merge-base <sha>)` when there is more than one base to tell apart, or
`(not above origin/main tip)`. The tip is `origin/main` as last fetched.
- **It fails closed when it has no trustworthy baseline:** neither `origin/main` nor `main`
resolves; there is no merge-base (shallow clone, unrelated history); or only local `main`
@@ -428,7 +441,8 @@ if the library is missing or defines no resolver.
`tests/test-adr0020-contract.sh` pins that arrangement: the library carries the only marker pair,
the hook carries none, the hook fails closed without the library, and a sentinel planted in a copied
library proves the hook executes the library's text. One of its assertions was green on a
library proves the hook executes the library's text, and a later block pins every repo-authored
pre-commit hook's `entry` and `stages`. One of its assertions was green on a
defect it named. "`validate.sh` sources the resolver in **both mode branches**" was implemented as a
file-wide `grep -Ec … -ge 2`, which cannot see a branch at all: delete the `agent)` arm's source line
and duplicate the `skill)` arm's, and the file-wide count is still 2 and the assertion still passes,
@@ -436,11 +450,14 @@ with the agent path running no resolver or some other one. It is now a **per-arm
each arm of `validate.sh`'s `case "$MODE" in` block must carry exactly one `source` line inside its
own body, and the file must carry exactly those two — with a mutation self-test that performs that
exact count-preserving edit on a copy and requires the check to fail on it. The suite's case count
runs **28 → 27 → 29**, and is **29** at HEAD: 28 at `620f20b` (the ADR-0025 merge), 27 after
`4de5b6b` retired the `.pre-commit-hooks.yaml` export, and 29 after `ef27c97` replaced the two-copy
hash and its line-count floor with the six one-copy assertions above. There are two 2026-09-16
runs **28 → 27 → 29 → 44**, and is **44** at HEAD: 28 at `620f20b` (the ADR-0025 merge), 27 after
`4de5b6b` retired the `.pre-commit-hooks.yaml` export, 29 after `ef27c97` replaced the two-copy
hash and its line-count floor with the six one-copy assertions above, and 44 after `384756b` added
the hook-wiring block. An earlier revision of this section stopped the chain at 29 and called that
the figure at HEAD; it was written before `384756b`. There are two 2026-09-16
changes here, not one, which is what an earlier revision of this section conflated. Each figure is
`bash tests/test-adr0020-contract.sh` run in a worktree at that commit, reading its `Results:` line.
`bash tests/test-adr0020-contract.sh` run at that commit — in a worktree for the historical ones —
reading its `Results:` line.
An earlier revision also opened the chain at 25; that predates the branch squash, no reachable
commit reproduces it, and it is dropped as unverifiable rather than carried.
@@ -482,8 +499,10 @@ against synthetic `mktemp` fixtures — it had never run against the agent files
how ADR-0016 could be amended to bless a `disallowedTools` frontmatter field while `validate.sh`'s
allowlist still rejected it: spec and enforcer disagreed and every gate stayed green.
Agents take the ADR-0020 **description** gates (`factory-audit`'s `validate.sh` holds its own copy of
those two constants) and, deliberately, **no body word gate**. A skill body is loaded into the
Agents take the ADR-0020 **description** gates and, deliberately, **no body word gate**. The two
description constants `factory-audit` applies to an agent live in `scripts/lib-checks-agent.sh:164-165`;
an earlier revision of this line put them in its `validate.sh`, which carries none of them (see
[Duplicated constants](#duplicated-constants)). A skill body is loaded into the
caller's context and competes with the live conversation; an agent body becomes the system prompt of
a *fresh* context. The rationale for the 900-word FAIL does not transfer. A bats test pins that
absence for the agent path of `factory-audit`'s validator — adding a body gate there contradicts the
@@ -910,11 +929,12 @@ An explicit `--config` from any other caller still wins, in all three argv forms
`--config=/abs`, `--config=rel`), and a relative one resolves against the caller's cwd — matching
bare `vale`, not the repo root.
Both audit skills' Step 1 passes no `--config` either. Step 1 resolves the script relative to the
`factory-audit`'s Step 1 passes no `--config` either. Step 1 resolves the script relative to the
skill's own directory so the call works from an installed plugin cache; a relative `--config`
alongside it would resolve against the cwd instead, yielding `E100 Runtime error … does not exist`
and exit 2 — which both skills' fallback misreads as "vale unavailable" and silently downgrades to
full LLM judgment.
and exit 2 — which the skill's fallback misreads as "vale unavailable" and silently downgrades to
full LLM judgment. An earlier revision wrote this paragraph in the plural, for the `skill-audit` /
`agent-audit` pair ADR-0025 merged; there is one Step 1 now.
`tests/test-vale-wrap.sh` regression-tests this against `factory-audit`'s copy — the only one left.
Its fixtures are all `SKILL.md`-shaped, and that copy's `.vale.ini` carries the matching glob section

View File

@@ -7,7 +7,7 @@ description: >
fixes -> agent-author.
allowed-tools: Bash Read
metadata:
version: "1.0.2"
version: "1.0.3"
category: factory
source_keys:
- agentskills-home

View File

@@ -1,5 +1,5 @@
extends: existence
message: "Composition or architecture note in a description: '%s' — a description carries a trigger, one capability clause and a boundary clause only; move this to README.md"
message: "Composition or architecture note in a description: '%s' — a description carries a trigger, one capability clause and a boundary clause only; move this to the body or a references/ file"
level: error
scope: text.frontmatter.description
ignorecase: true

View File

@@ -55,7 +55,8 @@ A model-invoked description carries exactly three things:
3. **Boundary clause.** Compressed form: `Not <thing> -> <skill-name>.` The target must resolve to
a real skill directory or agent file in the authoring source.
Everything else belongs in the body or in the plugin's `README.md`.
Everything else belongs in the body or in a `references/` file. Not a `README.md`: `plugins/gitea/`
and `plugins/lint/` both ship agents this file governs and neither has one.
## Indirect triggers — conditional, never blanket

View File

@@ -732,8 +732,41 @@ def boundary_targets(description):
return sorted({name for name, _, _ in _extract(description)})
def _arrow_targets(description):
"""Names extracted from ARROW notation specifically.
def _clause_end(description, pos):
"""Where CLAUSE_BODY stops scanning forward from `pos`.
The same two stops the class itself encodes: a `;`, or a `.` that is not
followed by a non-space character (a sentence end rather than a dot inside
`AGENTS.md`).
"""
for index in range(pos, len(description)):
char = description[index]
if char == ';':
return index
if char == '.' and not description[index + 1:index + 2].strip():
return index
return len(description)
def _arrow_clause_spans(description):
"""(start, end) for EACH ADR-0020 arrow clause, one span per clause.
A clause runs from its `Not` to whichever comes first: the start of the
NEXT arrow clause, or the end of the clause body. Bounding on the next
clause is what keeps two clauses joined by a comma inside one sentence
apart — a sentence-scoped span would merge them and let the second clause's
target vouch for the first.
"""
starts = [match.start() for match in BOUNDARY_ARROW.finditer(description)]
spans = []
for index, start in enumerate(starts):
limit = starts[index + 1] if index + 1 < len(starts) else len(description)
spans.append((start, min(limit, _clause_end(description, start))))
return spans
def _arrow_clause_parses(clause):
"""True when either arrow extractor reads a target out of ONE clause.
Kept apart from boundary_targets() because the arrow form is the one shape
that ALWAYS names a target: ADR-0020's `Not <thing> -> <name>`. A clause
@@ -741,15 +774,11 @@ def _arrow_targets(description):
that deserves its own message, and telling it apart needs the arrow targets
alone rather than every target in the description.
"""
out = []
for sentence in SENTENCE_SPLIT.split(description):
for match in ARROW_MARKED.finditer(sentence):
name, _, _ = _first(match)
if name:
out.append(name)
for match in ARROW_BOUNDARY.finditer(sentence):
out.append(match.group(1))
return out
for match in ARROW_MARKED.finditer(clause):
name, _, _ = _first(match)
if name:
return True
return bool(ARROW_BOUNDARY.search(clause))
def boundary_clause_status(description):
@@ -761,16 +790,28 @@ def boundary_clause_status(description):
three of them reworded a correct clause to satisfy a regex instead.
'unparsed' is the narrow, certain case: an ADR-0020 arrow clause was
detected and NO target came out of it. The arrow form always names one, so
detected and NO target came out of IT. The arrow form always names one, so
zero targets means the name is written in a shape the extractor cannot see
— a single-word bare target (`Not X -> forge`, which has to be written
`` `forge` `` or `/forge`) is the live example, since single-word names are
deliberately not matchable bare.
The test is PER CLAUSE, and that is the whole point of the span walk. Both
operands used to take the whole description, so ONE arrow clause that
parsed suppressed the diagnostic for every other clause beside it: a
backticked hyphenated target wrapped across a line break inside a `>`
folded scalar — `` `git-`` / ``commits` `` — went unchecked with no ERROR
and no SUGGESTION, while the same wrap written bare was reported correctly.
26 of this corpus's 38 skill descriptions carry more than one arrow clause,
so the suppression covered most of it. This is the issue #100 regression
class, and a whole-description test cannot see it by construction.
A PROSE clause yielding no target is NOT reported: "Do not use for anything
else" is a complete and legitimate boundary clause that names nowhere to go.
"""
if BOUNDARY_ARROW.search(description) and not _arrow_targets(description):
spans = _arrow_clause_spans(description)
if spans and not all(_arrow_clause_parses(description[start:end])
for start, end in spans):
return 'unparsed'
if has_boundary_clause(description):
return 'present'

View File

@@ -38,7 +38,7 @@ set -euo pipefail
# Divergence 2: a path-shaped argument that does not exist is a hard error
# (exit 2). Bare vale drops it, falls back to reading stdin, and prints
# `0 errors ... in stdin` with exit 0 — a typo'd target is then indistinguishable
# from a clean run. Both audit skills treat a `0 files` report as NOT RUN rather
# from a clean run. factory-audit treats a `0 files` report as NOT RUN rather
# than clean, and `in stdin` does not match that guard, so the silent form would
# read as "prefilter clean" and skip the LLM fallback. Erroring is the only way
# to keep that guard honest. Linting prose piped on stdin is therefore

View File

@@ -945,6 +945,48 @@ make_hand_invoked_skill() {
assert_output --partial "no target could be read"
}
# ---------------------------------------------------------------------------
# ADR-0020 — the unparsed diagnostic is PER CLAUSE, not per description
#
# boundary_clause_status() used to test `BOUNDARY_ARROW.search(description) and
# not _arrow_targets(description)`. Both operands took the WHOLE description,
# so ONE arrow clause that parsed suppressed the diagnostic for every other
# clause beside it.
#
# The shape that hides there is a backticked hyphenated target wrapped across
# the line break of a `>` folded scalar: the fold turns `` `fixture-sibling- ``
# / `` skill` `` into `fixture-sibling- skill`, which no extractor can read.
# Written BARE the same wrap is reported correctly, so the two spellings
# disagreed. 26 of this repo's 38 skill descriptions carry more than one arrow
# clause, which is how wide the suppression was. This is the #100 regression
# class: no ERROR, no SUGGESTION, exit 0.
# ---------------------------------------------------------------------------
@test "ADR-0020: an unparsed arrow clause is reported even when a sibling clause parses" {
local skill
skill="$(make_fixture_tree "$TMPDIR/tree" "my-skill")"
# Written by hand rather than through make_sized_skill: the `>` folded
# scalar and the wrap INSIDE the backticks are the fixture. The second
# clause parses and resolves against the fixture sibling, and that is what
# used to silence the first.
cat > "$skill/SKILL.md" <<'EOF'
---
name: my-skill
description: >
Use when doing the thing. Not the other thing -> `fixture-sibling-
skill`. Not a third thing -> `fixture-sibling-skill`.
metadata:
version: "1.0.0"
---
word word word word word word word word word word
EOF
run bash "$SCRIPT" "$skill"
assert_success
assert_output --partial "no target could be read"
refute_output --partial "has no boundary clause"
}
# ---------------------------------------------------------------------------
# Encoding, write side: sys.stdout/stderr.reconfigure(encoding='utf-8')
#

View File

@@ -442,6 +442,24 @@ fi
# terminal and therefore danglable. The issue #99 retrofit cut that composition
# sentence and the dangling target went with it, so the set is down to one.
#
# Correction (2026-09-20): the historical text carried NO backticks. `7801589^`
# has gitea-issues' description as "Composes gitea-labels-\n milestones for all
# label inference/resolution and milestone lookup", bare, so the token was read
# by the route-verb path — `Composes` is a ROUTE_VERB and the name matched
# NAME_HYPH — and not by the backtick sweep. Everything the paragraph above says
# about the fold and the trailing hyphen holds; only the spelling is wrong.
#
# The spelling is the load-bearing part, because the two are not equally
# visible. Backticked, that same wrap reaches every extractor as
# `` `gitea-labels- milestones` ``, which none of them can read: the opening
# backtick blocks the bare NAME_HYPH alternative and the space inside blocks the
# backticked one. Bare, it was extracted and reported all along, which is the
# only reason this dangling target was ever measured. In an ARROW clause the
# backticked wrap was silent until the 2026-09-20 fix to
# boundary_clause_status() in the shared resolver made the unparsed diagnostic
# per clause: before it, one sibling clause that parsed suppressed the finding
# for the whole description.
#
# `neuledge-context` was the last one. The issue #99 wave-3 retrofit deleted that
# boundary clause outright — commit `6146120` had already deleted the skill it
# named, and nothing has owned MCP-server installation since — so the corpus