Files
holocron/plugins/kyberforge/.apm/skills/factory-audit/references/skill-finding-criteria.md
Defame1297 620f20b0fd refactor(kyberforge)!: merge skill-audit and agent-audit into factory-audit
Why

The two audit skills carried 1,724 lines of byte-identical duplication: the ADR-0020 boundary
resolver (1,061), vale-wrap.sh (526), the Vale style rules (44) and the Contributing-files parser
(93). Nothing shared them — they were held in sync by a 413-line pre-push gate and its 797-line
test suite. Sync-by-gate had already failed once: at 484357a the two parser copies drifted into
different spellings of the bullet loop while a docstring asserted they were identical. That drift
was behaviour-neutral and was re-unified by hand at 598a7c3, so the copies were identical at merge
time — but nothing had caught it, and the next drift need not be neutral.

Implementation Notes

Self-containment binds BETWEEN skills, not within one. The agentskills.io spec forbids reaching
across skill directories, which is why two separate skills needed embedded copies; two files inside
ONE skill may source a third. That is the whole reason the merge removes duplication rather than
relocating it.

The union of both bodies measured 1,532 words against BODY_MAX_WORDS=900, and only 211 of those
words were shared, so SKILL.md is a dispatch body. Step 0 resolves the flow from the target path
before any validation, and its table mirrors validate.sh's detection exactly: a directory holding
SKILL.md or a SKILL.md file (skill); a *.agent.md, or a .md directly under an agents/ directory
(agent); anything else stops without running a validator. Steps 1-3 live in
references/skill-flow.md and references/agent-flow.md, and gotchas that apply to one flow live in
that flow's file, since it is loaded on every invocation anyway. If validate.sh reports on the
other artifact type, the body restarts at Step 0.

Named factory-audit rather than forge-audit because forge is a live skill, and a family prefix that
matches a live sibling reads as ownership rather than membership.

The description carries one arrow per boundary target, because ADR-0020 resolves only the first
target after an arrow. It drops the quoted "audit this skill"-style phrases, which restated
"audited" in a second register (ADR-0020's duplicate-register rule). 241 characters, Gotchas 16%
of the body: no size SUGGESTIONs.

The boundary resolver stays embedded in two files rather than imported: a cache-installed plugin
cannot read outside its own directory, and the repo-root hook resolves via .pre-commit-hooks.yaml
where entry[0] is the only token pre-commit rewrites, so no single file is reachable by both.
tests/test-adr0020-contract.sh hashes both copies for byte-identity, and asserts validate.sh sources
the resolver and that no third copy exists.

The entry scripts classify the target from its resolved parent directory, so a bare agent filename
typed inside agents/ works; resolve SCRIPT_DIR CDPATH-safely; and exit 2 when a lib-*.sh is
missing, rather than dying with exit 1, the tier the flows relay as real findings.

The provenance run functions stash their findings code in KYBERFORGE_PROV_RC and
return 0, so validate-provenance.sh calls them UNTESTED. Testing a function's
status (`f || RC=$?`) disables errexit for its entire body, and no subshell or
`set -e` inside can re-arm it once the call sits in a condition context
(measured, both spellings). Their error paths use `exit`, which is unaffected
either way; this keeps errexit armed for anything added later.

Case 0's readability guard reads the file instead of asking `[[ -r ]]`. `-r` is
access(2), which answers yes for uid 0 even on a mode-000 file, and this repo's
dev environment is root -- so the guard could never fire where it exists to fire.
A read attempt is also the stricter question, catching EIO. This is the reasoning
scripts/check-vale-style-sync.sh carried before this commit deleted it; the
hazard did not go with it.

All three entry scripts are CDPATH-safe, vale-wrap.sh included: both of its cd sites are cleared,
the --config resolution and the directory-mirror walk, where an exported CDPATH would otherwise
print a decoy path into the -print0 stream and build the mirror from the decoy's files. The two
remaining bare cd calls take absolute paths, which CDPATH is never consulted for.

Impact

BREAKING: skill-audit and agent-audit no longer exist as invocable skills. kyberforge goes to
2.0.0 (catalog 0.4.7).

Check logic is unchanged: differential runs of the old and new validators across every skill and
agent produced byte-identical stdout, stderr and exit codes, and the reconstructed Python payloads
differ only in comments and the references/field-inventory.md -> agent-field-inventory.md rename.
One doctrine governs the tiers: exit 0 is audited and clean, exit 1 is audited with findings OR a
target present but unreadable, exit 2 is that nothing was audited at all. Edge paths DID change,
deliberately (full table in ADR-0025):
- a missing target exits 2 (never ran), not 1, under its own "does not exist" message; detection is
  by path shape, so a shape-matching path that is simply absent used to reach the validator and come
  back as a FAIL against a file that never existed;
- an unshaped target exits 2 under the generic "matches neither" message, and a directory with no
  SKILL.md under a third, distinct one -- three exit-2 messages, not one;
- a dangling symlink or a symlink loop stays exit 1: it is present but broken, which is a finding
  about the artifact rather than a usage error;
- a SKILL.md file path is audited as its skill directory instead of refused;
- a .md agent outside an agents/ directory is refused rather than audited;
- a missing script library, a missing python3, a missing PyYAML, and no argument at all each exit 2.
  validate-provenance.sh already exited 2 for the last two; validate.sh now matches it.

.pre-commit-hooks.yaml is a published contract consumed by external repos. Both hook IDs and both
files: regexes are unchanged; only entry: and description: moved.

scripts/check-vale-style-sync.sh (413), scripts/sync-vale-styles.sh (21),
tests/test-check-vale-style-sync.sh (797) and agent-audit/scripts/README.md (47) are deleted. The
checker made 17 assertions: 6 compared the two Vale copies and are moot; 10 are rehomed into
tests/test-vale-wrap.sh (case 0, cases 28-31, and the suite's Vale-absent skip); and the
cross-manifest files: agreement check, which selected hooks by entry: and so could not survive both
hooks sharing one, is ported as case 33 pairing hooks by id:. Cases 28, 30 and 33 carry mutation
self-tests; narrowing the local skill prefilter to 6 of 38 SKILL.md files now fails the suite.

Skills go 39 to 38. Pre-push goes 9 repo-authored hooks to 8.

ADR: 0025
BREAKING-CHANGE: the skill-audit and agent-audit skills are removed. Both flows are served by
  factory-audit, which auto-detects whether it was handed a skill directory or an agent file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YR2CjVumUbEGWcMikcoXBD
2026-09-16 09:13:57 +00:00

6.6 KiB

source_keys
source_keys
agentskills-spec
agentskills-best-practices
agentskills-optimizing-descriptions
agentskills-using-scripts

Finding Criteria

Every FAIL and SUGGESTION criterion, for every qualitative dimension, and nothing else. The reasoning each criterion stands on, its worked examples and its house rules stay in that dimension's rubric, which Step 3 loads only for a dimension this file puts in play.

Two rules on using it:

  • A criterion that plainly applies is a finding. Write it up citing file and line.
  • A criterion that might apply, or whose call the wording here does not settle, is a reason to load that dimension's rubric — never a reason to drop the candidate. This file decides which rubrics to read; it does not settle a close call on its own.

description — references/skill-description-quality.md

Flag as FAIL if:

  • Over 400 characters. Measured on the folded YAML value, not the raw source lines. validate.sh reports the number; do not re-derive it, but do point the Fix at what to cut.

  • Internal mechanics appear in the description. Any of:

    • capability enumeration or a feature list;
    • output-format detail ("Produces a compact findings report with Why and Fix per finding");
    • composition or architecture notes ("composes X rather than duplicating Y", "a cross-cutting shared skill", "the human-facing entry point", "replaces the old flat invocation");
    • implementation detail ("self-validates via a bundled deterministic script").

    None of it can change a routing decision and all of it is preloaded. Kyberforge.CompositionNote catches the common phrasings deterministically; the rest is judgment. This is the rule that deflates a description, so apply it before reaching for length.

  • The same trigger stated twice in two registers — a verb list, then the same verbs re-quoted as user phrasings, usually in the same order. One register, whichever routes better.

  • Descriptive rather than imperative phrasing (This skill ..., This is the ...). Kyberforge.DescriptionOpener catches any opener matching ^This.

  • Vague capabilities ("helps with APIs" where "parses and validates OpenAPI specs" was available). Kyberforge.VagueWording catches the known filler; imprecision outside that list is judgment.

  • Trigger-list, boundary or indirect-trigger content on a hand-invoked skill — see Step 0 of references/skill-description-quality.md.

  • Over 1024 characters — the agentskills.io specification ceiling, unchanged and independent of the 400-character house ceiling above.

Flag as SUGGESTION if:

  • Over 250 characters but at or under 400. This tier is what moves the corpus average; the FAIL tier only stops outliers. Report it rather than treating a 399-character description as clean.
  • A near-miss exclusion is present but targets a weak near-miss.
  • An indirect trigger is present and warranted but could name the omitted phrasing more precisely.

An unresolved boundary target is not graded here. validate.sh owns that call and tiers it by notation — /name or an arrow form is an ERROR, the bare prose form a SUGGESTION unless a second target in the same sentence resolves — and Step 1 has already filed it under ### Structure at that tier. Re-grading it as a description FAIL puts one target in the report twice at two tiers. What is left to judgment here is semantic and the script cannot reach it: whether a target that does resolve is the right sibling to exclude, and whether a clause naming no target at all ("examine the files manually") should have named one.

body-discipline — references/skill-body-discipline.md

Flag as FAIL if:

  • A sentence answers "no" to the core test — it is padding
  • The body exceeds 900 words counted body-only (validate.sh reports it)
  • Two or more mutually exclusive flows are inlined instead of dispatched
  • A Gotcha paraphrases a step in the body below it that every branch reaching the Gotcha also reaches
  • A decision point presents a menu of options with no default
  • An instruction repeats content already in the description
  • A prescriptive sequence is used where flexibility is fine, or the reverse

Flag as SUGGESTION if:

  • The body exceeds 600 words counted body-only but stays at or under 900
  • The Gotchas section carries more than five entries
  • The Gotchas section exceeds 25% of the body
  • A rationale is missing from an include/exclude rule — present but unexplained
  • Gotchas are correct but placed late in the body rather than near the top
  • Content that only one branch reaches is inlined where a references/ file would serve

patterns — references/skill-patterns.md

Flag as FAIL if:

  • A Gotcha entry is a general tip or a reminder rather than a fact that defies a reasonable assumption
  • An inner code fence is unescaped inside a markdown block, breaking the render
  • A checklist wraps a single step
  • A conditional reference gives no trigger — Kyberforge.PaddingPhrase reports the common form
  • The agent must produce a specific format and no output template is given

Flag as SUGGESTION if:

  • Gotchas are correctly formed but placed late in the body
  • An output template is present but permissive where the consumer needs it exact
  • A conditional reference names a trigger that is real but broader than the branch it guards

file-structure and internal-consistency — references/skill-file-structure.md

Flag as FAIL if:

  • A directory outside the four permitted ones exists
  • Test files sit in scripts/
  • A non-spec file sits at the skill root
  • A path that resolves outside the skill directory appears outside the two exempt locations, in prose rather than in a fenced example
  • tests/ exists but tests/README.md is missing or does not document its repo-level dependency
  • SKILL.md describes a script invocation the script does not accept

Flag as SUGGESTION if:

  • An optional directory exists but holds only a placeholder README

formatting and scripts — references/skill-formatting-and-scripts.md

Flag as FAIL if:

  • A script prompts interactively, in any form
  • A script exposes no --help
  • A destructive script has no --dry-run
  • Data and diagnostics share a stream, so the output cannot be piped
  • A relative path named in the body does not resolve
  • Heading levels are inconsistent enough to break the document's structure

Flag as SUGGESTION if:

  • Exit codes are meaningful but undocumented in --help
  • A code block is untagged where a language applies
  • A script is idempotent in practice but does not say so, leaving a re-run's safety unclear
  • List indentation or section spacing is inconsistent without breaking the render