- factory-audit: hook events judged per deployed target after apm's
rename (Claude/Copilot event sets FAIL, others SUGGESTION); Claude
plugin layouts accepted as hook sources; interpreter options and
sh -c strings checked; bats 378 -> 386
- primitive-author: Must 4/5 match the audit; reference hand-back
points at the right steps; Step 4.2 --target all fallback
- skill-author: new-skill.sh repair only on the template marker line,
so complete skills stay a no-op; provenance and calibration text
- forge: restore "already named" qualifier; drop false HITL claim
- apm-workflow: token example uses an env var
- docs/hooks.md: the apm-hooks.json sidecar is committed, not ignored
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KkT7RSDwDbmrM9T34b6sTi
- factory-audit: no-op hooks, ./ after interpreters, split-quote and
spaced ${PLUGIN_ROOT} paths, camelCase events in Claude-targeted flat
files, case-insensitive routing stems, and non-string YAML keys are
now caught; input: forms and prompt boundary clauses align with
primitive-author; bats 347 -> 367
- primitive-author: routing forms, quoting guidance, install exit on
hidden Unicode, argument-hint exception
- forge: drop duplicated gotcha, fit description and body budgets (#143)
- skill-author: primitive-author boundary, Claude-only env vars
- hook: exit unless CLAUDE_PROJECT_DIR is set, so Copilot/Codex never
run apm update; ADR-0019 correction, ADR-0025 amendment, docs fixes
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KkT7RSDwDbmrM9T34b6sTi
lib-boundary-resolver.sh and the provider-adapter-author BOM fixtures
carried literal U+FEFF characters, which apm install reports as "files
contain hidden characters". Both sites feed the text to Python, which
interprets the '' escape identically, so behaviour is unchanged:
the resolver's strip_bom and all five validate-adapter BOM cases pass.
No core package bump: the fixture writes the same bytes, so the edit is
not substantive under the apm-workflow version policy.
Refs #94
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KkT7RSDwDbmrM9T34b6sTi
Second clean-context audit found author Must/Should and audit FAIL/SUGGESTION
tiers drifting apart, and author Musts the audit never checked.
- factory-audit: FAIL on absolute or bare relative hook script paths, an
applyTo present but empty, and unbalanced braces/brackets in applyTo;
judgment steps for dependency stem collisions, helper .json in hook dirs,
unresolvable instruction links, prompt model slugs and second-person
bodies; an unmatched glob drops to SUGGESTION; deliberate tier deviations
recorded in hook-flow.md; validate.sh --help lists the three new modes;
DescriptionOpener message no longer prescribes "Use when".
- primitive-author: deprecated routing, extra prompt keys and the prompt
description contract become Shoulds; hook Musts gain "contributes an
entry", no bare relative paths, and executable-when-run-directly;
prompt Must 1 covers hardlinks; Vale prose FAILs resolved at close.
- forge: say "hook, instruction or prompt" rather than "apm primitive".
Refs #94
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KkT7RSDwDbmrM9T34b6sTi
primitive-author:
- description excludes read-only review (-> factory-audit)
- validation Gotcha now matches the research: compile never reads
prompts, install fails only on a bad Copilot hook payload and warns on
prompt input names and dropped keys
- instruction fold-in into AGENTS.md/CLAUDE.md stated as conditional on
dedup and --force-instructions
- hook checklist gains the wrapped-shape Must, drops hardlinks, notes
why executable is stricter than the research, and states the
separate Copilot-targeted package route instead of a blanket "don't"
- prompt Must 5 keeps the research's Copilot-only-key exception; adds
model-slug and 250-char Shoulds; descriptions name skills or agents
- placeholder instruction covers both FILL IN and FILL_IN_ tokens
factory-audit: hardlink FAIL scoped to instructions and prompts
(find_hook_files skips symlinks only), with bats cases; prompt-flow
description rubric names skills or agents.
forge: version-bump, apm-routes and sources references updated for the
primitive route.
Refs #94
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KkT7RSDwDbmrM9T34b6sTi
factory-audit gains three Step 0 rows and flows for the apm primitives
that have no container of their own: a .json file under hooks/, a
*.instructions.md and a *.prompt.md. apm validates almost none of them
(invalid hook JSON is skipped silently, instruction validate() only
warns, input: names are never checked against ${input:x}), so the
deterministic checks live in a new scripts/lib-checks-primitive.sh,
wired into validate.sh's path-shape detection. Each check and tier
traces to the Authoring checklists in the microsoft-apm research docs.
- Hook: JSON/shape/event-list checks mirroring the Copilot payload
validator, never-firing event casing, missing/escaping/non-executable
scripts (FAIL); deprecated filename routing and ${CLAUDE_PLUGIN_ROOT}
(SUGGESTION).
- Instruction: location, frontmatter, description, body, stem clash
(FAIL); missing or list applyTo and unread keys (SUGGESTION).
- Prompt: location/name, frontmatter, description, input names, the
upstream `- name: x` docs bug, declared-vs-used ${input:x} (FAIL);
ADR-0029 description length and trigger clause, dropped keys,
camelCase aliases, argument-hint with input (SUGGESTION). Whether a
prompt carries procedure is judgment in prompt-flow.md, not a script
heuristic.
Vale now lints *.instructions.md and *.prompt.md with the Kyberforge
style; test-vale-wrap.sh gains their probe rows. New
tests/validate-primitive.bats (31 cases). kyberforge 2.0.1 -> 2.1.0 with
the executables.allow key, catalog 0.5.1 -> 0.5.2, marketplace.json
regenerated.
Refs #94
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KkT7RSDwDbmrM9T34b6sTi
The ADR-0020 boundary resolver (boundary_targets()/unresolved_targets())
only ever read a SKILL.md's description. A target named in the BODY -- a
dispatch table row, a "run X" step, both routine in a 900-word procedure
-- was checked by nothing. Two real instances shipped before either was
caught by reading rather than by a gate: bin/write-docs routed twice to a
deleted `to-prd` skill, and bin/triage told an agent to run a nonexistent
`/setup-matt-pocock-skills` (both fixed in 03abcff; that fix was the
symptom, this gate is the actual ask per #124).
Added a separate, narrower extractor -- body_targets() /
unresolved_body_targets() in the shared lib-boundary-resolver.sh -- rather
than reusing the description resolver at wider scope. The description
gate's sentence-level heuristics (BOUNDARY_MARKER, the follower test,
in-sentence corroboration) are tuned for a one-to-three-sentence routing
clause and misfire on dispatch-table/procedure prose in both directions,
so the body gate reads only explicit route notation (`/name`,
backticked-or-slash-prefixed `-> name` / `-> name`), already the
description gate's own unconditionally-blocking tier.
Three guards were added after running the extractor over the real
39-skill corpus and reading every hit rather than assuming the design was
correct:
- a target must be hyphenated, even in notation -- single-word citations
like `/fork` (forge, citing Claude Code's own /fork command) and
`/name` (skill-author, a placeholder) are not routes.
- a bare hyphenated word after any arrow is not notation -- only
ARROW_MARKED (backticked/slash-prefixed) is used, not NOTATION_ARROW's
bare form, so ordinary process-chain prose ("prop -> new ref ->
re-render", caveman) is not read as a route.
- a name immediately preceded by `<` is a closing tag
(`</what-to-do>`, grill-with-docs), not /name notation.
Wired into both consumers that must agree by contract: scripts/
skill-size-check.sh (the pre-commit hook) and factory-audit's
lib-checks-skill.sh (the audit). Verified identical findings across both
over the whole corpus.
tests/test-adr0020-targets.sh gains a dedicated section pinning the two
live true positives and all three guards. docs/spec/gates.md and
ADR-0020 get a matching amendment.
Fixes: #124
ADR: 0020
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EGHFJextYtVQseaHPDDhxB
Review of PR 139 found list-rejection and confinement holes that let the
exact malformed entries the grammar forbids pass check 7.
- Reject comma, space-separated and backticked path lists, so
`a/sources.md (x), b/topic.md` no longer exits 0 unchecked.
- FAIL absolute paths and any path whose realpath leaves the repo, for
both `Research doc:` and `Basis:`.
- Anchor `(removed in <sha>)` to the end of the value with a 7-40 hex
sha. The sha is format-checked only, not resolved with git cat-file.
- Read `* ` bullets and `- **X**` bullets correctly under a `**Basis:**`
header, and strip backticks from Basis paths.
- Stop the semicolon rule firing on annotation prose, and stop `none`
matching `none/foo.md`.
- Update the stale field messages to the new grammar and report an empty
field as empty, not missing.
- Skip a removed Basis silently when there is no repo root.
Adds 40 tests. Each guarded line was mutated in place and every mutant
is caught.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EGHFJextYtVQseaHPDDhxB
validate-provenance.sh assumed `Research doc:` names a research
sources.md whose H2 headings are the source slugs, but 29 corpus entries
named topic docs and 6 values were not a single path, so checks 7 and 8
reported INFO for 36 entries and nothing ever failed.
`Research doc:` now takes exactly one path. An entry with no registry
writes `none` plus one `- **Basis:** <path>` bullet per path; each Basis
path is existence-checked unless annotated `(removed in <sha>)`.
- Check 7 FAILs when a resolved registry lacks the slug, when the value
is a topic doc, or when it is a list. An unresolvable path stays INFO.
- Check 8 is retired: one registry serves many skills, so requiring
every registry slug in each skill's sources.md is unsatisfiable.
- The Research doc and Basis parsers accept the inline, bullet and
header-plus-bullets spellings, so a differently spelled field is no
longer read as absent.
Refs: #121
ADR: 0028
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EGHFJextYtVQseaHPDDhxB
boundary_clause_status() ran BOUNDARY_ARROW.search() and _arrow_targets()
over the whole description, so one arrow clause that parsed suppressed the
diagnostic for every other clause in it. A backticked hyphenated routing
target wrapped across lines in a folded scalar was therefore silently
unchecked -- no error, no suggestion, exit 0 -- whenever the description
carried one other clause that parsed. Written bare, the same wrap errors
correctly. That is the shape #100 regressed on.
The check is now per clause. Nothing that passed starts failing: all 68
routing targets across the 38 SKILL.md files resolved before and still do.
26 of those descriptions carry more than one arrow clause, so the
suppression was live across two thirds of the corpus, not an edge case.
validate-skill.bats pins the shape. test-adr0020-targets.sh's comment
described the #100 regression as a backticked wrap; the historical text was
unbackticked, which is precisely the shape the gate did not catch.
Also closes three README misroutes the branch left in the enforcement
layer: CompositionNote.yml's message, agent-description-quality.md:58 and
vale-wrap.sh's header still sent overflow to a skill-root README.md and
named the two skills ADR-0025 merged away. 1ec3e8a fixed the prose and
missed the rules that enforce it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NwD8Egs5r4ndqeFLmhusX2
Skills no longer carry a README.md, so the size advice in
skill-size-check and factory-audit's validator now names a references/
file instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Why: scripts/skill-size-check.sh embedded a byte-identical 1,061-line copy
of the ADR-0020 boundary resolver only because it was also exported
through .pre-commit-hooks.yaml, whose consumers could not reach a file
inside the plugin. 4de5b6b retired that export, so the hook now runs only
in this repo and can source factory-audit's lib-boundary-resolver.sh like
validate.sh does. One copy removes the edit-one-paste-the-other hazard.
Implementation Notes:
- The hook's Python program is assembled from its own preamble, the
library's resolver and its own checks, read from quoted here-docs. The
assembled program matches the old one line for line except one comment,
and the hook's stdout, stderr and exit code are identical over every
corpus SKILL.md and the 26 differential-suite fixtures.
- The hook fails closed, naming the library, when it is missing or
defines no resolver.
- test-adr0020-contract.sh assertion 1 now pins the single copy: one
marker pair in the library, none in the hook, fail-closed on a missing
or gutted library, and a sentinel planted in a copied library that must
appear in the hook's output. 1a expects exactly one authority. 27 -> 29
passes.
- ADR-0020 and ADR-0025 carry dated amendments; gates.md and the
library, hook and mode-library comments no longer describe two copies.
- factory-audit is new on this branch, so the version-bump gate exempts
it; kyberforge is already at 2.0.0 against main's 1.6.2.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Why: .pre-commit-hooks.yaml and its release-tag gate served external
consumers that do not exist. No repo on the Gitea instance pins these
hooks, and the README names apm as the only supported install path. The
mechanism was also already failing: skill-size-check.sh changed after
v2.0.1 with no tag cut, and the gate cannot fire through Gitea's merge
button. (Simplification audit finding 36.)
Implementation Notes:
- Delete .pre-commit-hooks.yaml, scripts/check-release-needed.sh,
tests/test-check-release-needed.sh and tests/test-vale-hooks-consumer.sh,
and remove the check-release-needed pre-push hook. The repo: local
skill-size-check and vale-audit-prefilter-* hooks are unchanged.
- ADR-0014 is amended, not retired: its runtime decision to bundle Vale
inside factory-audit stands. The amendment keeps the entry[0]-only
constraint (LESSONS.md:101,105) in case the export returns. ADR-0025
gets a pointer.
- test-vale-wrap.sh: drop case 33 (the cross-manifest drift check) and
case 28's hook-scope half, which read the published manifest. Case 32
now also requires each hook to select every tracked file of its class,
which keeps case 33's one-plugin-narrowing guard, with a mutation test.
- test-skill-size-check.sh and test-adr0020-contract.sh now assert the
hook contract and verbose: true on .pre-commit-config.yaml only.
- gates.md: pre-push count goes from 9 to 8 authored hooks (11 to 10
reported), and the Release table, the External consumers section and
the two-manifest scope table are removed. README and script/test
comments no longer describe the export as live. The resolver comment
is edited identically in both copies.
- The v1.0.0/v2.0.0/v2.0.1 tags are left in place; they are inert.
ADR: 0014
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Why: two branches that both bump a skill 1.0.0 -> 1.0.1 with different
content merge without a conflict, and each passed the gate against its own
merge-base, so main could ship two changes under one version.
Implementation Notes:
- check-skill-version-bump requires the pushed version to exceed both the
merge-base and the main tip; failures name the baseline they missed.
- Presence is read from the tree, so a blob missing from a partial clone is
a read failure instead of a silently exempt "new" skill.
- A leading UTF-8 BOM no longer reads as a missing version.
- Version parts reject leading zeros in all three validators
(check-skill-version-bump, skill-size-check, factory-audit).
- New tests cover equal bumps, moved files, major/minor ordering, bad refs,
unreadable blobs, mode-only changes, symlinks and tag peeling.
Impact: ADR-0022 amended (reverses "not main's current tip"); gates.md
updated to match, including pre-commit 4.6.1's exact ref selection.
ADR: 0022
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Why
The two audit skills carried 1,724 lines of byte-identical duplication: the ADR-0020 boundary
resolver (1,061), vale-wrap.sh (526), the Vale style rules (44) and the Contributing-files parser
(93). Nothing shared them — they were held in sync by a 413-line pre-push gate and its 797-line
test suite. Sync-by-gate had already failed once: at 484357a the two parser copies drifted into
different spellings of the bullet loop while a docstring asserted they were identical. That drift
was behaviour-neutral and was re-unified by hand at 598a7c3, so the copies were identical at merge
time — but nothing had caught it, and the next drift need not be neutral.
Implementation Notes
Self-containment binds BETWEEN skills, not within one. The agentskills.io spec forbids reaching
across skill directories, which is why two separate skills needed embedded copies; two files inside
ONE skill may source a third. That is the whole reason the merge removes duplication rather than
relocating it.
The union of both bodies measured 1,532 words against BODY_MAX_WORDS=900, and only 211 of those
words were shared, so SKILL.md is a dispatch body. Step 0 resolves the flow from the target path
before any validation, and its table mirrors validate.sh's detection exactly: a directory holding
SKILL.md or a SKILL.md file (skill); a *.agent.md, or a .md directly under an agents/ directory
(agent); anything else stops without running a validator. Steps 1-3 live in
references/skill-flow.md and references/agent-flow.md, and gotchas that apply to one flow live in
that flow's file, since it is loaded on every invocation anyway. If validate.sh reports on the
other artifact type, the body restarts at Step 0.
Named factory-audit rather than forge-audit because forge is a live skill, and a family prefix that
matches a live sibling reads as ownership rather than membership.
The description carries one arrow per boundary target, because ADR-0020 resolves only the first
target after an arrow. It drops the quoted "audit this skill"-style phrases, which restated
"audited" in a second register (ADR-0020's duplicate-register rule). 241 characters, Gotchas 16%
of the body: no size SUGGESTIONs.
The boundary resolver stays embedded in two files rather than imported: a cache-installed plugin
cannot read outside its own directory, and the repo-root hook resolves via .pre-commit-hooks.yaml
where entry[0] is the only token pre-commit rewrites, so no single file is reachable by both.
tests/test-adr0020-contract.sh hashes both copies for byte-identity, and asserts validate.sh sources
the resolver and that no third copy exists.
The entry scripts classify the target from its resolved parent directory, so a bare agent filename
typed inside agents/ works; resolve SCRIPT_DIR CDPATH-safely; and exit 2 when a lib-*.sh is
missing, rather than dying with exit 1, the tier the flows relay as real findings.
The provenance run functions stash their findings code in KYBERFORGE_PROV_RC and
return 0, so validate-provenance.sh calls them UNTESTED. Testing a function's
status (`f || RC=$?`) disables errexit for its entire body, and no subshell or
`set -e` inside can re-arm it once the call sits in a condition context
(measured, both spellings). Their error paths use `exit`, which is unaffected
either way; this keeps errexit armed for anything added later.
Case 0's readability guard reads the file instead of asking `[[ -r ]]`. `-r` is
access(2), which answers yes for uid 0 even on a mode-000 file, and this repo's
dev environment is root -- so the guard could never fire where it exists to fire.
A read attempt is also the stricter question, catching EIO. This is the reasoning
scripts/check-vale-style-sync.sh carried before this commit deleted it; the
hazard did not go with it.
All three entry scripts are CDPATH-safe, vale-wrap.sh included: both of its cd sites are cleared,
the --config resolution and the directory-mirror walk, where an exported CDPATH would otherwise
print a decoy path into the -print0 stream and build the mirror from the decoy's files. The two
remaining bare cd calls take absolute paths, which CDPATH is never consulted for.
Impact
BREAKING: skill-audit and agent-audit no longer exist as invocable skills. kyberforge goes to
2.0.0 (catalog 0.4.7).
Check logic is unchanged: differential runs of the old and new validators across every skill and
agent produced byte-identical stdout, stderr and exit codes, and the reconstructed Python payloads
differ only in comments and the references/field-inventory.md -> agent-field-inventory.md rename.
One doctrine governs the tiers: exit 0 is audited and clean, exit 1 is audited with findings OR a
target present but unreadable, exit 2 is that nothing was audited at all. Edge paths DID change,
deliberately (full table in ADR-0025):
- a missing target exits 2 (never ran), not 1, under its own "does not exist" message; detection is
by path shape, so a shape-matching path that is simply absent used to reach the validator and come
back as a FAIL against a file that never existed;
- an unshaped target exits 2 under the generic "matches neither" message, and a directory with no
SKILL.md under a third, distinct one -- three exit-2 messages, not one;
- a dangling symlink or a symlink loop stays exit 1: it is present but broken, which is a finding
about the artifact rather than a usage error;
- a SKILL.md file path is audited as its skill directory instead of refused;
- a .md agent outside an agents/ directory is refused rather than audited;
- a missing script library, a missing python3, a missing PyYAML, and no argument at all each exit 2.
validate-provenance.sh already exited 2 for the last two; validate.sh now matches it.
.pre-commit-hooks.yaml is a published contract consumed by external repos. Both hook IDs and both
files: regexes are unchanged; only entry: and description: moved.
scripts/check-vale-style-sync.sh (413), scripts/sync-vale-styles.sh (21),
tests/test-check-vale-style-sync.sh (797) and agent-audit/scripts/README.md (47) are deleted. The
checker made 17 assertions: 6 compared the two Vale copies and are moot; 10 are rehomed into
tests/test-vale-wrap.sh (case 0, cases 28-31, and the suite's Vale-absent skip); and the
cross-manifest files: agreement check, which selected hooks by entry: and so could not survive both
hooks sharing one, is ported as case 33 pairing hooks by id:. Cases 28, 30 and 33 carry mutation
self-tests; narrowing the local skill prefilter to 6 of 38 SKILL.md files now fails the suite.
Skills go 39 to 38. Pre-push goes 9 repo-authored hooks to 8.
ADR: 0025
BREAKING-CHANGE: the skill-audit and agent-audit skills are removed. Both flows are served by
factory-audit, which auto-detects whether it was handed a skill directory or an agent file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YR2CjVumUbEGWcMikcoXBD