Files
holocron/tests/run-tests.sh
Defame1297 620f20b0fd refactor(kyberforge)!: merge skill-audit and agent-audit into factory-audit
Why

The two audit skills carried 1,724 lines of byte-identical duplication: the ADR-0020 boundary
resolver (1,061), vale-wrap.sh (526), the Vale style rules (44) and the Contributing-files parser
(93). Nothing shared them — they were held in sync by a 413-line pre-push gate and its 797-line
test suite. Sync-by-gate had already failed once: at 484357a the two parser copies drifted into
different spellings of the bullet loop while a docstring asserted they were identical. That drift
was behaviour-neutral and was re-unified by hand at 598a7c3, so the copies were identical at merge
time — but nothing had caught it, and the next drift need not be neutral.

Implementation Notes

Self-containment binds BETWEEN skills, not within one. The agentskills.io spec forbids reaching
across skill directories, which is why two separate skills needed embedded copies; two files inside
ONE skill may source a third. That is the whole reason the merge removes duplication rather than
relocating it.

The union of both bodies measured 1,532 words against BODY_MAX_WORDS=900, and only 211 of those
words were shared, so SKILL.md is a dispatch body. Step 0 resolves the flow from the target path
before any validation, and its table mirrors validate.sh's detection exactly: a directory holding
SKILL.md or a SKILL.md file (skill); a *.agent.md, or a .md directly under an agents/ directory
(agent); anything else stops without running a validator. Steps 1-3 live in
references/skill-flow.md and references/agent-flow.md, and gotchas that apply to one flow live in
that flow's file, since it is loaded on every invocation anyway. If validate.sh reports on the
other artifact type, the body restarts at Step 0.

Named factory-audit rather than forge-audit because forge is a live skill, and a family prefix that
matches a live sibling reads as ownership rather than membership.

The description carries one arrow per boundary target, because ADR-0020 resolves only the first
target after an arrow. It drops the quoted "audit this skill"-style phrases, which restated
"audited" in a second register (ADR-0020's duplicate-register rule). 241 characters, Gotchas 16%
of the body: no size SUGGESTIONs.

The boundary resolver stays embedded in two files rather than imported: a cache-installed plugin
cannot read outside its own directory, and the repo-root hook resolves via .pre-commit-hooks.yaml
where entry[0] is the only token pre-commit rewrites, so no single file is reachable by both.
tests/test-adr0020-contract.sh hashes both copies for byte-identity, and asserts validate.sh sources
the resolver and that no third copy exists.

The entry scripts classify the target from its resolved parent directory, so a bare agent filename
typed inside agents/ works; resolve SCRIPT_DIR CDPATH-safely; and exit 2 when a lib-*.sh is
missing, rather than dying with exit 1, the tier the flows relay as real findings.

The provenance run functions stash their findings code in KYBERFORGE_PROV_RC and
return 0, so validate-provenance.sh calls them UNTESTED. Testing a function's
status (`f || RC=$?`) disables errexit for its entire body, and no subshell or
`set -e` inside can re-arm it once the call sits in a condition context
(measured, both spellings). Their error paths use `exit`, which is unaffected
either way; this keeps errexit armed for anything added later.

Case 0's readability guard reads the file instead of asking `[[ -r ]]`. `-r` is
access(2), which answers yes for uid 0 even on a mode-000 file, and this repo's
dev environment is root -- so the guard could never fire where it exists to fire.
A read attempt is also the stricter question, catching EIO. This is the reasoning
scripts/check-vale-style-sync.sh carried before this commit deleted it; the
hazard did not go with it.

All three entry scripts are CDPATH-safe, vale-wrap.sh included: both of its cd sites are cleared,
the --config resolution and the directory-mirror walk, where an exported CDPATH would otherwise
print a decoy path into the -print0 stream and build the mirror from the decoy's files. The two
remaining bare cd calls take absolute paths, which CDPATH is never consulted for.

Impact

BREAKING: skill-audit and agent-audit no longer exist as invocable skills. kyberforge goes to
2.0.0 (catalog 0.4.7).

Check logic is unchanged: differential runs of the old and new validators across every skill and
agent produced byte-identical stdout, stderr and exit codes, and the reconstructed Python payloads
differ only in comments and the references/field-inventory.md -> agent-field-inventory.md rename.
One doctrine governs the tiers: exit 0 is audited and clean, exit 1 is audited with findings OR a
target present but unreadable, exit 2 is that nothing was audited at all. Edge paths DID change,
deliberately (full table in ADR-0025):
- a missing target exits 2 (never ran), not 1, under its own "does not exist" message; detection is
  by path shape, so a shape-matching path that is simply absent used to reach the validator and come
  back as a FAIL against a file that never existed;
- an unshaped target exits 2 under the generic "matches neither" message, and a directory with no
  SKILL.md under a third, distinct one -- three exit-2 messages, not one;
- a dangling symlink or a symlink loop stays exit 1: it is present but broken, which is a finding
  about the artifact rather than a usage error;
- a SKILL.md file path is audited as its skill directory instead of refused;
- a .md agent outside an agents/ directory is refused rather than audited;
- a missing script library, a missing python3, a missing PyYAML, and no argument at all each exit 2.
  validate-provenance.sh already exited 2 for the last two; validate.sh now matches it.

.pre-commit-hooks.yaml is a published contract consumed by external repos. Both hook IDs and both
files: regexes are unchanged; only entry: and description: moved.

scripts/check-vale-style-sync.sh (413), scripts/sync-vale-styles.sh (21),
tests/test-check-vale-style-sync.sh (797) and agent-audit/scripts/README.md (47) are deleted. The
checker made 17 assertions: 6 compared the two Vale copies and are moot; 10 are rehomed into
tests/test-vale-wrap.sh (case 0, cases 28-31, and the suite's Vale-absent skip); and the
cross-manifest files: agreement check, which selected hooks by entry: and so could not survive both
hooks sharing one, is ported as case 33 pairing hooks by id:. Cases 28, 30 and 33 carry mutation
self-tests; narrowing the local skill prefilter to 6 of 38 SKILL.md files now fails the suite.

Skills go 39 to 38. Pre-push goes 9 repo-authored hooks to 8.

ADR: 0025
BREAKING-CHANGE: the skill-audit and agent-audit skills are removed. Both flows are served by
  factory-audit, which auto-detects whether it was handed a skill directory or an agent file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YR2CjVumUbEGWcMikcoXBD
2026-09-16 09:13:57 +00:00

284 lines
13 KiB
Bash
Executable File

#!/usr/bin/env bash
# Run all test-*.sh files in the repo (including plugins) and the bats suite.
# Usage: bash tests/run-tests.sh [--bats-only] [--strict]
#
# A script exiting 77 (the automake convention) is reported as SKIPPED, not
# passed — a suite that can't run for lack of a binary must not read as green.
#
# --strict (or RUN_TESTS_STRICT=1) additionally makes any skip FAIL the run. Two
# different readings of a skip are both correct, and which one applies depends on
# who is running:
#
# * ad-hoc, on a laptop: skipping gracefully is the point. You are missing a
# dev binary, the other 15 suites still tell you something, and turning that
# into a red run would just train people to ignore red.
# * as a GATE (the run-tests pre-push hook): a skip is a SETUP ERROR, not a
# legitimate state. README.md's Prerequisites table documents vale, apm and
# python3/PyYAML -- the dependencies these suites actually guard on -- as
# required pre-push, so a suite that cannot run on the machine doing the
# pushing means the machine is misconfigured -- and
# pre-commit prints NOTHING for a passing hook, so the skip list below is
# swallowed entirely. On a vale-less PATH that once silently shipped a
# green gate having verified 15 of the 17 suites that existed then.
# Exactly the vacuous-pass class the rest of this file exists to close.
#
# Deliberately its own switch, scoped to this dispatcher alone: it governs
# whether an unrunnable suite is tolerated, nothing else. It was once kept
# separate from CHECK_VALE_STYLE_SYNC_ALLOW_MISSING_VALE, which governed whether
# check-vale-style-sync could downgrade itself; that gate is retired (ADR-0025
# merged the two Vale copies it diffed), but the rule that retired it does not
# apply here. Keep any future vale-related opt-out separate too — one flag
# disarming several gates is how an opt-out quietly grows blast radius.
#
# TEST_DIR — override root to search for test-*.sh (default: REPO_ROOT); used by tests.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
BATS="$REPO_ROOT/tests/run-bats.sh"
BATS_ONLY=false
STRICT=false
if [[ "${RUN_TESTS_STRICT:-}" == "1" ]]; then
STRICT=true
fi
# Latched, then REMOVED from the environment. The value has done its only job by
# this line -- it is now held in the STRICT shell local -- and leaving it exported
# makes strictness leak down the whole process tree: every test-*.sh dispatched
# through batch_run below inherits it, and any of them that itself invokes
# run-tests.sh (tests/test-run-tests.sh drives a copy of this script over fixture
# trees) silently turns a deliberately non-strict fixture strict.
#
# That is not symmetric with `--strict`, which never leaked: the flag only ever
# sets the shell local above, so `bash tests/run-tests.sh --strict` (the spelling
# the run-tests pre-push hook uses) always gave children a clean environment. Only
# the env-var spelling leaked, and it broke exactly two assertions in
# tests/test-run-tests.sh -- its cases 10c and 10g. Unsetting here makes the two
# documented invocations equivalent in what a CHILD sees, not just in the parent's
# verdict, so no future suite has to defend itself the way test-run-tests.sh's
# run_fake() does with `env -u`.
unset RUN_TESTS_STRICT
# A loop rather than the `[[ "${1:-}" == --bats-only ]]` test this used to be, so
# the two flags compose and an unknown flag is rejected instead of ignored. A
# silently-ignored `--strict` is the one typo that would turn the gate back off.
for arg in ${@+"$@"}; do
case "$arg" in
--bats-only) BATS_ONLY=true ;;
--strict) STRICT=true ;;
*)
echo "Usage: $0 [--bats-only] [--strict]" >&2
exit 2
;;
esac
done
SEARCH_ROOT="${TEST_DIR:-$REPO_ROOT}"
FAILED=()
SKIPPED=()
# Parallel array, index-matched to SKIPPED. Not an associative array: bash 3.2
# (macOS) has none, and tests/test-vale-wrap.sh's bash-3.2 scan rejects
# `declare -A` outright.
SKIP_REASONS=()
PASSED=0
SKIP_EXIT=77
# A missing or non-executable run-bats.sh is a hard error, never a silent skip.
# This was `if [[ -x "$BATS" ]]; then ... fi` with no else and no assertion that
# bats ran at all, so renaming, moving, or dropping the executable bit off
# run-bats.sh made the entire bats suite vanish with zero diagnostic and the run
# still printed "Summary: N passed, 0 failed" and exited 0 -- and --bats-only
# degraded to a no-op that printed nothing and exited 0. That is the same
# green-either-way hole run-bats.sh's own zero-count guard closes one level down;
# this closes it in the dispatcher that pre-push actually invokes.
#
# Present and executable is still not "it ran". `bash "$BATS"` on an EMPTY
# run-bats.sh exits 0 having printed nothing, and the dispatcher printed
# `=== bats ===`, a blank line, and a green summary -- the same green-either-way
# defect one spelling over. Truncation, a partial write, an editor saving an
# empty buffer, and a `set -e` abort in a future run-bats.sh preamble all land
# there. So the runner's own summary line is required, and its count must be
# non-zero: that line is run-bats.sh's contract with this script, and it is only
# emitted after run-bats.sh's own zero-count guard has passed.
#
# Stdout is captured (the summary is on stdout) while stderr passes straight
# through, so a failing runner's diagnostics still reach the terminal live. The
# capture costs no streaming that was not already lost: run-bats.sh buffers its
# per-file output and flushes it at the end regardless.
run_bats() {
if [[ ! -x "$BATS" ]]; then
echo "Error: bats runner not found or not executable at $BATS — the bats suite cannot be skipped silently" >&2
exit 1
fi
echo "=== bats ==="
local out rc=0 summary count
out="$(bash "$BATS")" || rc=$?
[[ -z "$out" ]] || printf '%s\n' "$out"
echo ""
if [[ $rc -ne 0 ]]; then
exit "$rc"
fi
summary="$(printf '%s\n' "$out" | grep -E '^[0-9]+ tests, [0-9]+ failures$' | tail -n 1 || true)"
if [[ -z "$summary" ]]; then
echo "Error: $BATS exited 0 without reporting an 'N tests, M failures' summary — it ran but produced nothing, so the bats suite was not verified" >&2
exit 1
fi
count="${summary%% *}"
if [[ "$count" -eq 0 ]]; then
echo "Error: $BATS reported 0 tests — the bats suite executed nothing" >&2
exit 1
fi
}
if $BATS_ONLY; then
run_bats
exit 0
fi
run_bats
# Collected with a `while read` loop rather than `mapfile` — macOS ships
# /bin/bash 3.2, which has no `mapfile`. Process substitution (not a pipe)
# keeps the loop in this shell so the appends survive. `sort` is still fed
# newline-delimited output, exactly as before.
#
# apm_modules/ is excluded for the same reason tests/run-bats.sh excludes it:
# `apm install` materializes dependency copies of this repo's own plugins there,
# and re-running a dependency's tests re-runs what plugins/ already covers.
#
# .claude/skills/ is excluded for the same reason, one install step later. apm
# deploys skills there straight from each plugin's `.apm/` tree, `tests/` dirs
# and all, so any `<skill>/tests/test-*.sh` file under plugins/ would also land
# at .claude/skills/<name>/tests/ and be discovered and run a second time. No
# such file exists today -- the skill suites are all .bats, where run-bats.sh
# already hit this -- so the exclusion is symmetry with its sibling, kept here
# so the first shell suite added under a skill does not reintroduce it.
SCRIPTS=()
while IFS= read -r script; do
SCRIPTS+=("$script")
done < <(
find "$SEARCH_ROOT" -name "test-*.sh" \
-not -path "*/.git/*" \
-not -path "*/.claude/worktrees/*" \
-not -path "*/apm_modules/*" \
-not -path "*/.claude/skills/*" \
| sort
)
# Each test-*.sh is independent (fixtures live under its own mktemp dir, none
# write back into the live repo tree -- verified before adding this), so they
# run concurrently in fixed-size batches instead of one at a time. Dispatch and
# throttling is scripts/lib/batch-run.sh's batch_run -- shared with
# tests/run-bats.sh so a batching bug fix only needs to land once; see that
# file for why this is batched rather than a rolling `wait -n` pool.
SCRATCH_ROOT="$(mktemp -d)"
trap 'rm -rf "$SCRATCH_ROOT"' EXIT
# The source= path below is repo-root-relative, matching scripts/install.sh:5 --
# NOT script-dir-relative. The source-path used to resolve it is the cwd
# pre-commit invokes the linter from, which is the repo root, so `lib/...`
# (resolving to a nonexistent tests/lib/) and `../scripts/...` (escaping the repo
# entirely) both fail. Both spellings were live until issue #97, and neither was
# visible: the resulting SC1091 is `info` while .pre-commit-config.yaml pins
# `--severity=warning`. A directive that does not resolve also blinds
# test-vale-wrap.sh's `sourced_files()` exemption, which reads these same
# directives to find array seeding that lives in the sourced file.
#
# Do not start a comment line here with the linter's name -- it is parsed as a
# directive and errors out (SC1073).
# shellcheck source=scripts/lib/batch-run.sh
source "$REPO_ROOT/scripts/lib/batch-run.sh"
declare -a batch_args=()
idx=0
for script in ${SCRIPTS[@]+"${SCRIPTS[@]}"}; do
idx=$((idx + 1))
cmd="$(printf 'rc=0; bash %q || rc=$?; echo "$rc" >%q' "$script" "$SCRATCH_ROOT/$idx.status")"
batch_args+=("$idx" "$cmd")
done
batch_run "$SCRATCH_ROOT" ${batch_args[@]+"${batch_args[@]}"}
idx=0
for script in ${SCRIPTS[@]+"${SCRIPTS[@]}"}; do
idx=$((idx + 1))
rel="${script#"$SEARCH_ROOT/"}"
echo "=== $rel ==="
cat "$SCRATCH_ROOT/$idx.log"
rc="$(cat "$SCRATCH_ROOT/$idx.status" 2>/dev/null || echo 1)"
# String comparison, not `-eq`. `-eq` is arithmetic, and bash evaluates an
# empty string as 0 there -- `[[ "" -eq 0 ]]` is true -- so an *empty* status
# file counted as a pass. The `|| echo 1` fallback above only covers a
# *missing* file; a file that exists but is empty is what you get when the job
# is killed between the `>` truncating it and the `echo` completing, or on
# ENOSPC. Under `==` an empty status falls through to FAILED, which is the only
# safe reading of "the job did not report a result".
if [[ "$rc" == "0" ]]; then
PASSED=$((PASSED + 1))
elif [[ "$rc" == "$SKIP_EXIT" ]]; then
SKIPPED+=("$rel")
# Capture WHY, not just that. The reason is printed by the suite itself and
# is otherwise swallowed with the rest of its log, which leaves the reader
# knowing something was skipped but not which binary to install. There is no
# single house format for it -- three suites print `SKIP: <reason>` on stdout
# and one prints `apm not installed -- skipping (...)` on stderr -- so this
# tries the shapes in decreasing order of confidence and falls back to the
# last thing the suite said before exiting 77, which for a guard that exits
# immediately is the reason by construction. batch-run.sh folds stderr into
# the same log, so the stderr spelling is reachable here.
reason="$(grep -E '^[[:space:]]*SKIP' "$SCRATCH_ROOT/$idx.log" 2>/dev/null | head -n 1 || true)"
if [[ -z "$reason" ]]; then
reason="$(grep -iE 'skip' "$SCRATCH_ROOT/$idx.log" 2>/dev/null | head -n 1 || true)"
fi
if [[ -z "$reason" ]]; then
reason="$(grep -vE '^[[:space:]]*$' "$SCRATCH_ROOT/$idx.log" 2>/dev/null | tail -n 1 || true)"
fi
if [[ -z "$reason" ]]; then
reason="(exited $SKIP_EXIT without printing a reason)"
fi
# Trimmed of leading whitespace so the reasons line up under their suite
# names regardless of how each suite indents its own message.
SKIP_REASONS+=("${reason#"${reason%%[![:space:]]*}"}")
else
FAILED+=("$rel")
fi
echo ""
done
echo "=== Summary: $PASSED passed, ${#SKIPPED[@]} skipped, ${#FAILED[@]} failed ==="
# Suppressed under --strict: the strict block below reports the same suites with
# the same reasons, and printing both left the reader scrolling past one list to
# reach an identical one. Under strict the failure block IS the list.
if [[ ${#SKIPPED[@]} -gt 0 && "$STRICT" != true ]]; then
echo "Skipped scripts:"
sidx=0
for s in ${SKIPPED[@]+"${SKIPPED[@]}"}; do
echo " $s"
echo " ${SKIP_REASONS[$sidx]}"
sidx=$((sidx + 1))
done
fi
RC=0
if [[ ${#FAILED[@]} -gt 0 ]]; then
echo "Failed scripts:"
for s in ${FAILED[@]+"${FAILED[@]}"}; do
echo " $s"
done
RC=1
fi
# Strict mode turns every skip into a failure. Reported separately from FAILED
# above rather than folded into it: a skipped suite did not fail, the machine
# did, and a message that says so points at the fix. Named with reasons again
# here (not just referenced) because this block goes to stderr and is what a
# pre-push reader actually gets handed.
if [[ "$STRICT" == true && ${#SKIPPED[@]} -gt 0 ]]; then
echo "Error: --strict and ${#SKIPPED[@]} suite(s) skipped. Run as a gate, a skip is a SETUP ERROR on this machine, not a legitimate state: README.md's Prerequisites table documents vale, apm and python3/PyYAML — what these suites guard on — as required pre-push dependencies, so every suite is expected to be runnable here. Install what each suite names below and re-run; do not skip the hook." >&2
sidx=0
for s in ${SKIPPED[@]+"${SKIPPED[@]}"}; do
echo " $s" >&2
echo " ${SKIP_REASONS[$sidx]}" >&2
sidx=$((sidx + 1))
done
RC=1
fi
exit "$RC"