fix(skill-audit): make check 9 reachable, wrap-safe and never silently skipped

Check 9 shipped in #130 to close #118, but three defects meant it could not do the job it was
added for.

Why:
- It is INFO-only, so it always exits 0 — and SKILL.md graded exit 0 "a genuine pass" and said the
  script "prints nothing on success". Every check-9 INFO was discarded before it reached a report,
  behind three further doors that only opened on a non-zero exit.
- `parse_field_raw()` matched `(.+)`, which does not span newlines, so only the first physical line
  of a wrapped value was compared. Rewriting only the continuation line of a wrapped Description
  from a hedge to a confident claim produced no finding at all — verbatim the regression #118 was
  filed about. The bullet branch had the same shape: a wrapped bullet broke the loop and dropped
  every later entry.
- A `git show` failure at the base ref was treated as "creation, nothing to flag" and skipped the
  whole skill with no output, collapsing "absent at that ref" with "not tracked under that name".
  A gitignored `.claude/skills/` copy reported clean while the authoring path reported four changed
  claims. The script's own usage text promises this is "never a silent skip".

Implementation notes:
- Exit-code guidance re-keyed on output as well as code: 0-and-silent passes, 0-with-output is
  INFO-only findings, 1 is FAILs, 2 never ran.
- `parse_field_raw()` is line-based and joins continuation lines; `normalize_field_text()`'s
  docstring is now true rather than aspirational. A reorder deliberately fires: the two fields share
  one parser, and order-insensitivity would mean splitting a prose Description on commas.
- The discarded `show_err` is now surfaced as one whole-check INFO naming both readings.
- `--base-ref=` given empty now beats the env var, as the usage text always claimed.

`validate.sh` gains an ADR-0022 `metadata.version` check at FAIL tier, because any lower tier lets
skill-author Step 4 report done on a file the commit gate then refuses. Its `read` heuristic now
skips here-doc bodies — reflowing the one offending line would have cleared the finding and left
the cause, since every usage() heredoc is one wrap from putting the English verb in column 0.

Impact: provenance tests 73 -> 82, validate tests 64 -> 72. Test 72 previously deleted origin/main
before asserting the override, so it proved the flag works with no default rather than that it beats
one; it now moves origin/main forward first.

Refs: #118
ADR: 0022
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EeH8SCbcrCAQrtymkNuhKP
This commit is contained in:
2026-09-09 05:15:05 +00:00
parent ed8c99efbd
commit 175ea89c0a
8 changed files with 704 additions and 65 deletions

View File

@@ -15,7 +15,8 @@ Arguments:
a long-lived branch, a mirror with a different remote name).
The VALIDATE_PROVENANCE_BASE_REF environment variable is an
equivalent, lower-precedence way to set it — the flag wins
if both are given.
if both are given, including when the flag is given empty
(\`--base-ref=\`), which selects the default resolution.
Exit codes:
0 All checks passed (or nothing to validate)
@@ -47,10 +48,13 @@ Checks performed:
8 Extracted non-(none) slug in research doc present in sources.md
9 Description or Contributing files text changed since --base-ref (INFO
only — a bash script cannot verify the claim is still TRUE, only that it
changed; the auditor reads the named files to check that). A slug absent
at the base ref is a creation, not a change, and is not flagged. When the
base ref cannot be resolved at all, this is announced as ONE INFO for the
whole check, never a silent skip.
changed; the auditor reads the named files to check that). Wrapped values
are joined before comparison, so a re-wrap alone is not a change and a
rewrite of any line of one is. A slug absent at the base ref is a
creation, not a change, and is not flagged; a field that WAS there and is
now gone is announced as a removal. When the base ref cannot be resolved,
or references/sources.md is not tracked under this path at that ref, this
is announced as ONE INFO for the whole check, never a silent skip.
Checks 7 and 8 apply ONLY when the Research doc value names a research SOURCE
INDEX — a file whose basename is sources.md, whose H2 headings ARE source
@@ -70,8 +74,15 @@ fi
# never counts against them — a caller passing it alongside skill-dir sees
# the same argument-count behaviour as one who does not pass it at all, and a
# genuinely extra positional argument is still rejected.
#
# BASE_REF_OVERRIDE is deliberately left UNSET here rather than initialised to
# the empty string. `--base-ref=` (given, but empty) and "no flag at all" are
# different instructions — the first says "use the default resolution, ignoring
# the environment", the second says "fall back to the environment" — and an
# empty-string initialiser collapsed them: `${BASE_REF_OVERRIDE:-$ENV}` treats
# an empty flag value as absent, so the environment variable won and the usage
# text's "the flag wins if both are given" was false for exactly that spelling.
declare -a POSITIONAL_ARGS=()
BASE_REF_OVERRIDE=""
for arg in "$@"; do
case "$arg" in
--base-ref=*)
@@ -107,10 +118,15 @@ fi
SKILL_DIR_ARG="${POSITIONAL_ARGS[0]}"
# The flag wins over the environment variable when both are given; either is
# empty-string when unset, and an empty string tells the Python body to fall
# The flag wins over the environment variable whenever the flag was GIVEN —
# `+x` tests for presence, not for a non-empty value, which is the distinction
# `:-` could not make. An empty result either way tells the Python body to fall
# back to `git merge-base HEAD origin/main`.
BASE_REF="${BASE_REF_OVERRIDE:-${VALIDATE_PROVENANCE_BASE_REF:-}}"
if [[ -n "${BASE_REF_OVERRIDE+x}" ]]; then
BASE_REF="$BASE_REF_OVERRIDE"
else
BASE_REF="${VALIDATE_PROVENANCE_BASE_REF:-}"
fi
# python3 is a HARD dependency. Without this preflight a missing interpreter
# produced 'line NN: python3: command not found' and exit 127 — an exit code no
@@ -491,10 +507,21 @@ def find_repo_root(start_dir):
# --- Check 9 helpers ---------------------------------------------------
# Check 9 needs a raw field VALUE (as text, to diff against an earlier
# version), not the parsed structure parse_contributing_files() and
# parse_status() return — a Contributing files list that reordered its
# entries without changing them is not what this check is looking for, but
# neither is normalizing so hard that a genuine rewrite disappears. Raw text,
# whitespace-normalized, is the middle ground.
# parse_status() return. The ONE normalization applied is whitespace
# collapsing, which is what makes a re-wrap or a re-indent invisible; nothing
# else is normalized away.
#
# In particular a REORDERED Contributing files list DOES fire this check, and
# that is deliberate — the header here used to claim the opposite, which the
# code never did. Order-insensitivity cannot be had for one field without
# distorting the other: the two fields share this parser, and the only way to
# ignore order is to split the value into items and sort them, which for a
# prose Description means splitting on commas and would then hide a genuine
# rewrite that merely permuted its clauses. Check 9 is always an INFO whose
# whole job is to point a human at a place to read; a reordered list costs
# that human one glance to dismiss, whereas a hidden rewrite is the exact
# failure #118 exists to catch. False positive over false negative, on this
# check, on purpose.
def run_git(args, cwd):
"""Run `git <args>` in cwd. Returns (returncode, stdout, stderr) — never
@@ -523,14 +550,44 @@ def find_slug_block(content, slug):
m = pattern.search(content)
return m.group(1) if m else None
# A field value ENDS at the next field, the next heading, or a blank line.
# Every other non-blank line is a continuation of the value the author wrapped
# across physical lines.
#
# This boundary is what the old `(.+)$` regex did not have. `.` does not cross
# a newline, so only the FIRST physical line of a wrapped value was ever
# compared — and a rewrite confined to a continuation line produced no finding
# at all. That is verbatim the hedge-to-confident-claim regression #118 exists
# to catch, invisible to the check written to catch it. The bullet branch had
# the same defect one level down: a wrapped bullet's continuation does not
# start with '- ', so the loop broke there and silently dropped every
# remaining bullet.
#
# A continuation line that itself opens with bold text ('**note** — ...') is
# read as a boundary and truncates the value. That is a known, narrow
# false-negative, accepted because the alternative — no boundary at all —
# is what produced the wide one above.
FIELD_BOUNDARY_RE = re.compile(r'^(?:- )?\*\*|^#{1,6} ')
def _is_field_boundary(stripped_line):
"""True when a stripped line starts a new field, bullet-less heading or H2."""
return bool(FIELD_BOUNDARY_RE.match(stripped_line))
def parse_field_raw(content, slug, field_name):
"""Raw text of a '**<field_name>:**' field under a slug H2.
"""Raw text of a '**<field_name>:**' field under a slug H2, wrapping joined.
Mirrors the two authored shapes parse_contributing_files() and
parse_status() already handle (inline value on the same line, or a
bare heading followed by '- ' bullets), but returns text rather than a
parsed structure, because check 9 diffs wording, not semantics.
Continuation lines are joined into the value they belong to before the
caller normalizes and compares, so a value wrapped across two lines and
the same value on one line are the same text — and a change made on any
line of a wrapped value is visible, not just one made on the first.
Returns None when the H2 itself is absent (the slug did not exist at
this content's revision) or the field is absent — both read as "no
earlier claim to compare against" to the caller, which is deliberate:
@@ -539,28 +596,49 @@ def parse_field_raw(content, slug, field_name):
block = find_slug_block(content, slug)
if block is None:
return None
inline_re = re.compile(r'^\- \*\*' + re.escape(field_name) + r':\*\* (.+)$', re.MULTILINE)
im = inline_re.search(block)
if im:
return im.group(1).strip()
heading_re = re.compile(r'^\*\*' + re.escape(field_name) + r':\*\*\s*$', re.MULTILINE)
hm = heading_re.search(block)
if not hm:
return None
lines = []
for line in block[hm.end():].splitlines():
line = line.strip()
if not line:
if lines:
break
continue
if not line.startswith("- "):
break
lines.append(line[2:].strip())
return ", ".join(lines) if lines else None
lines = block.splitlines()
inline_re = re.compile(r'^\- \*\*' + re.escape(field_name) + r':\*\*[ \t]*(.*)$')
heading_re = re.compile(r'^\*\*' + re.escape(field_name) + r':\*\*[ \t]*$')
for idx, line in enumerate(lines):
im = inline_re.match(line)
if im:
parts = [im.group(1).strip()]
for cont in lines[idx + 1:]:
stripped = cont.strip()
if not stripped or stripped.startswith("- ") or _is_field_boundary(stripped):
break
parts.append(stripped)
joined = " ".join(p for p in parts if p).strip()
return joined or None
if heading_re.match(line):
entries = []
for cont in lines[idx + 1:]:
stripped = cont.strip()
if not stripped:
if entries:
break
continue
if _is_field_boundary(stripped):
break
if stripped.startswith("- "):
entries.append(stripped[2:].strip())
elif entries:
# A wrapped bullet: fold it back into the bullet it
# continues rather than ending the list here.
entries[-1] = (entries[-1] + " " + stripped).strip()
else:
break
return ", ".join(e for e in entries if e) or None
return None
def normalize_field_text(value):
"""Collapse whitespace so reformatting alone never registers as a change."""
"""Collapse whitespace so reformatting alone never registers as a change.
True only because parse_field_raw() joins wrapped continuation lines
first: collapsing whitespace inside a value that had already been
truncated at its first newline normalized nothing a re-wrap could change.
"""
return re.sub(r'\s+', ' ', value).strip()
findings = []
@@ -1025,28 +1103,82 @@ else:
["show", f"{resolved_base_ref}:{sources_md_relpath}"], repo_root
)
if rc != 0:
# The base ref resolved fine, but references/sources.md did not
# exist there at all — the whole file is new. Every entry in it
# is therefore a creation, not a change: nothing to flag, and
# this is not a structural failure of the check, so no INFO
# either. Same reasoning applies per-slug below when the ref
# resolved but a given '## <slug>' heading did not exist yet.
# The base ref resolved fine but `git show <ref>:<path>` did not.
# That single return code covers two situations this check cannot
# tell apart, and only one of them is harmless:
#
# the file genuinely did not exist at the base ref — the whole
# sources.md is new, every entry in it is a creation, and there
# is nothing check 9 could have flagged;
#
# the path is not TRACKED under that name at the base ref — a
# renamed skill directory, or a copy of the skill living
# somewhere untracked or gitignored (an installed .claude/skills
# tree is the everyday case).
#
# Treating both as "creation, nothing to flag" made the second one
# a silent, whole-skill skip: the same directory audited at its
# authoring path reported changed claims and at its deployed path
# reported nothing, with no way to tell that from a clean run.
# That is the exact fail-open this script's own header forbids —
# "never a silent skip" — so announce it once for the whole check
# and hand over git's own stderr, which is the only diagnostic
# that separates the two cases.
detail = show_err.strip().splitlines()
detail = detail[0] if detail else "git gave no reason"
emit_info(
f"Check 9 skipped — '{sources_md_relpath}' is not tracked at {resolved_base_ref}",
"references/sources.md",
f"`git show {resolved_base_ref}:{sources_md_relpath}` failed ({detail}). "
f"Either the file did not exist at that ref — in which case every entry is a "
f"creation and there was nothing to flag — or this path is not tracked under "
f"that name there: a renamed skill directory, or an untracked or gitignored copy "
f"of the skill such as a deployed .claude/skills/ tree. "
f"Check 9 did not run for any slug in this skill. "
f"Re-run against the tracked authoring path, or pass --base-ref=<ref> naming a "
f"commit where this path exists."
)
old_sources_content = None
if old_sources_content is not None:
for slug in unique_slugs:
changed_fields = []
removed_fields = []
for field_name in ("Description", "Contributing files"):
old_value = parse_field_raw(old_sources_content, slug, field_name)
new_value = parse_field_raw(sources_content, slug, field_name)
if old_value is None or new_value is None:
if old_value is None and new_value is None:
continue
if old_value is None:
# No earlier claim to compare against — a brand-new
# entry, or a field that did not exist yet at the
# base ref. That is a creation, not a change, and is
# never flagged.
continue
if new_value is None:
# The field existed at the base ref and is gone now.
# This was folded into the creation skip above, which
# justified only the other half: deleting a whole
# '- **Description:**' line left NO finding anywhere —
# no other check in this script requires the field, so
# a claim could be removed as invisibly as it could be
# strengthened. Announce it; the auditor decides
# whether the removal was intended.
removed_fields.append(field_name)
continue
if normalize_field_text(old_value) != normalize_field_text(new_value):
changed_fields.append(field_name)
if removed_fields:
removed_list = " and ".join(removed_fields)
emit_info(
f"'{removed_list}' removed for '{slug}' since {resolved_base_ref}",
f"references/sources.md (## {slug})",
f"The '## {slug}' entry had {removed_list} at {resolved_base_ref} and has "
f"none now. Nothing else in this script requires the field, so the removal "
f"is otherwise invisible. Confirm it was deliberate — a provenance claim "
f"withdrawn is as much a change to the chain as one rewritten — and "
f"restore the field if it was lost to an edit."
)
if changed_fields:
field_list = " and ".join(changed_fields)
emit_info(

View File

@@ -1289,6 +1289,50 @@ else:
if desc:
ok("description has no unfilled placeholders")
# --- ADR-0022: metadata.version is mandatory -------------------------------
# FAIL, not SUGGESTION, and the tier is set by the gate rather than by taste.
# `.pre-commit-config.yaml`'s `skill-frontmatter` hook REJECTS a SKILL.md with
# no `metadata.version`, and rejects a value that is not three-part semver.
# skill-author's Step 4 says to run this audit and "resolve every FAIL", so any
# tier below FAIL lets that step report done on a skill the commit gate then
# refuses — the same audit-disagrees-with-the-gate failure the MAX_LINES note
# below warns about, arrived at from the other direction. Verified before this
# check existed: a SKILL.md with no `metadata:` block at all reported "All
# checks passed".
#
# The rule is DUPLICATED from that hook for the same cache-isolation reason as
# every other constant here — an installed plugin's scripts cannot read the
# repo-root config. Keep the two in step: this check must accept exactly what
# the hook accepts.
SEMVER_RE = re.compile(r'^\d+\.\d+\.\d+$')
try:
fm_data = yaml.safe_load(fm)
except Exception:
# Unreachable in practice: description_value() above parses the same text
# and hard-exits on a YAML error, so anything arriving here already parsed.
fm_data = None
metadata_block = fm_data.get('metadata') if isinstance(fm_data, dict) else None
if not isinstance(metadata_block, dict) or metadata_block.get('version') is None:
fail("frontmatter has no metadata.version — ADR-0022 makes it mandatory for "
"every skill, and the skill-frontmatter pre-commit hook rejects the file "
"without it. Add `metadata:` / ` version: \"1.0.0\"` (new skills start "
"at \"0.1.0\")")
else:
version_value = metadata_block['version']
# NOT str()-coerced blind: `version: 1.0` is a YAML float, and its "1.0"
# spelling is exactly the two-part value the hook rejects — coercing and
# then matching keeps this check and the hook agreeing on that case.
version_text = version_value if isinstance(version_value, str) else str(version_value)
version_text = version_text.strip()
if SEMVER_RE.match(version_text):
ok(f"metadata.version present: '{version_text}' (ADR-0022)")
else:
fail(f"metadata.version '{version_text}' is not three-part semver — the "
f"skill-frontmatter pre-commit hook rejects it. Use MAJOR.MINOR.PATCH, "
f"e.g. \"1.0.0\"")
# SKILL.md size ceilings (agentskills.io skill-authoring.md: 500 lines,
# ~5,000 tokens). Both constants are DUPLICATED from the repo-root pre-commit
# hook scripts/skill-size-check.sh — a plugin skill's scripts cannot read files
@@ -1515,11 +1559,66 @@ def stdin_redirected(line, prev_line):
unquoted = re.sub(r'"[^"]*"|\'[^\']*\'', '', line)
return '<' in unquoted or prev_line.rstrip().endswith('|')
# A here-doc body is DATA, not command position. Every script in this corpus
# carries a `usage() { cat <<EOF ... EOF; }`, and prose wrapped inside one puts
# ordinary English at the start of a line — "read is reported as an INFO ..."
# in this skill's own validate-provenance.sh, which made skill-audit hard-FAIL
# on its own script. Reflowing that one sentence would have cleared the finding
# and left the cause: every future usage text is one wrap away from the same
# false positive, and the remedy an author reaches for is contorting working
# source, which the note above records has already happened twice.
#
# Detection is deliberately conservative in the direction that matters. A
# here-doc body is skipped only when its terminator is actually found further
# down the file; an opener with no terminator — the shape a stray `<<` inside a
# string would produce — is ignored rather than allowed to swallow the tail,
# because swallowing the tail is a false NEGATIVE and this check exists to fail
# closed. `<<<` here-strings open nothing and are excluded by the lookbehind.
HEREDOC_START_RE = re.compile(r'(?<!<)<<-?\s*(["\']?)([A-Za-z_][A-Za-z0-9_]*)\1')
def heredoc_delimiter(line):
"""The here-doc terminator this line opens, or None."""
m = HEREDOC_START_RE.search(line)
return m.group(2) if m else None
def heredoc_body_indices(lines):
"""Line indices that are here-doc BODY (plus its terminator), not code."""
skip = set()
i, n = 0, len(lines)
while i < n:
stripped = lines[i].strip()
delim = None if stripped.startswith('#') else heredoc_delimiter(lines[i])
if delim:
# `<<-` allows an indented terminator, so compare stripped.
for j in range(i + 1, n):
if lines[j].strip() == delim:
skip.update(range(i + 1, j + 1))
i = j
break
i += 1
return skip
# The here-doc exemption applies to the `read` heuristic ONLY, and the
# asymmetry is the point. `read` is an ordinary English verb, so any prose a
# script prints is one line-wrap away from opening with it. `input(` is not a
# word — a line beginning `input(` inside a here-doc is an embedded Python
# program pausing for a keypress, which is exactly what this check is for, and
# these scripts embed Python in a here-doc as a matter of course. Exempting the
# whole body would have disarmed the check across every script in the corpus.
def interactive_reads(source):
hits = []
prev_line = ''
for line in source.splitlines():
lines = source.splitlines()
in_heredoc = heredoc_body_indices(lines)
for idx, line in enumerate(lines):
stripped = line.strip()
if idx in in_heredoc:
if re.match(r'input\(', stripped):
hits.append(stripped)
continue
if re.match(r'read(\s|$)', stripped):
if not stdin_redirected(line, prev_line):
hits.append(stripped)