refactor(skills): retrofit the corpus to the ADR-0020 context contract (#129)

Retrofits all 39 skills to ADR-0020's description/body context contract, then fixes what six rounds of independent review found in that retrofit — including four ways the hot gate itself failed open.

Closes #99, #107, #108, #110, #111, #114, #115, #120.

## The retrofit (waves 1-5)

| | Start | Now |
|---|---|---|
| Description FAILs (>400 chars) | 26 | **0** |
| Body FAILs (>900 words, body-only) | 9 | **0** |
| Dangling routing targets | 2 | **0** |
| `Kyberforge.CompositionNote` | 10 | **0** |
| Preload tax | 21,005 chars | **~10,500** |

Under the 12,000-char success criterion. Per-wave detail is on #99.

## The review fixes

**The gate failed open four ways, three of them found after the retrofit shipped.** An unrecognised follower token made a dangling target vanish. A skill directory with no `SKILL.md` resolved as a valid target, so a commit could be green locally and red in a fresh clone — three existing fixtures were relying on that, one of which made the install-leak A/B pass vacuously. Then the free-standing `/name` sweep turned out to be gated on the sentence carrying a boundary marker, so route notation in any other sentence was invisible — not an ERROR, not a SUGGESTION, not an INFO — which left the documented "`/name` always blocks" promise false from a second direction. All four fixed and pinned.

**Two checks were silently not running.** `validate-provenance.sh` checks 7-8 were dead across nine skills. Waking them exposed a deeper problem: they assume `Research doc:` names a source index, but 30 of 121 entries point at topic content documents, so every new check-7 INFO was a false positive and check 8 was saved from a false-FAIL flood only by an *unannounced* skip. Checks 7/8 are now scoped to source indexes and every skip announces itself (#121).

**The retrofit's own anti-goal, four times.** ADR-0020 warns that a blunt gate gets satisfied by deleting content rather than relocating it. `diagnose` and `skill-audit` relocated prose and then read it unconditionally; `prototype` and `vale-config` deleted rules outright that survived nowhere. All four addressed.

## Verification

- `bash tests/run-tests.sh --strict` — 24 suites, 0 skipped, 0 failed
- `bash tests/run-bats.sh` — 325 tests, 0 failures
- `pre-commit run --all-files` — 17/17
- `pre-commit run --hook-stage pre-push --all-files` — 16/16, with `apm marketplace check` and `apm pack --check-clean` run against the remote, not skipped
- `scripts/skill-size-check.sh` over all 39 skills — rc 0, 0 ERROR/FAIL, SUGGESTION-only
- Preload tax measured at **10,498 chars**, max description 390 — both inside budget
- Every new test proven non-vacuous by a deliberate mutation of the behaviour it covers

**Per-commit sync, stated accurately:** the ten commits from the latest review round each pass `check-plugin-content-sync` in isolation, verified by checking each out in a detached worktree with a clean between. The earlier gitea window (`dfacf05..bedbd1d`, nine commits) does **not** — its mirror was regenerated in one batch at `bbc7300`. An earlier revision of this description claimed the property held for every commit; it does not, and a bisect through that window lands on a red commit. **Squash-merge** to collapse it, or accept that this range is not bisectable.

## Version bump

Six plugins and the catalog take a **patch**, not a minor. The branch is **89 commits — 40 `fix` / 30 `refactor` / 12 `docs` / 5 `chore` / 2 `test` — zero `feat`, zero `!`, zero `BREAKING CHANGE`** — and adds no skill, agent, command or hook. (Two earlier revisions of this section cited a stale histogram, most recently 78 commits; the figures above are measured at HEAD.) Both rules this repo ships (`forge/references/version-bump.md`, landing in this PR, and `git-commits/references/conventional-commits-spec.md`) make that a patch, and the catalog set is unchanged at 7 entries.

Not settled by that: four published files were removed from the installed tree, three moved, and `caveman` gained `disable-model-invocation`, retiring its old triggers. Under a strict reading those are major-class and currently ship under `refactor:` with no marker. Whether the deployed skill surface is a public contract is written down nowhere — worth deciding, but it outlives this PR.

## Deliberately not in scope

#112 (cherry-pick ownership, now resolved in favour of `git-commits`), #113 (`rtk git` normalisation), #116 (research fan-out), #101 (audit-skill merge), #122 (non-spec skill-root files), #123 (no PRD producer) stay open. #117 is the one worth reading: the contract's remedy is to move prose into `references/`, which is exactly where neither the size gate nor Vale looks — and the blind spot is wider than #117 currently records, since there is no root `.vale.ini` at all, so every ADR, `CONTEXT.md` and `README.md` is unlinted too.

That blind spot let this branch carry two `level: error` `Kyberforge.SentenceOpenerThereIs` violations into `references/` files it created — `provider-adapter-author/references/provider-matrix.md:31` and `agent-audit/references/finding-criteria.md:95`. Both are reworded in `afadaae`, confirmed by routing each file through the audit's own `vale-wrap.sh` (1 error each before, 0 after). Five further occurrences sit in `references/` files already on `main`; those are the pre-existing corpus and stay with #117, which is the real fix.

Also unfixed and not this PR's: `apm install` appends a duplicate `SessionStart` entry to `.claude/settings.json`, so a fresh clone cannot get pre-push green without an edit AGENTS.md warns against. Reproduces identically on `main`.

Co-authored-by: Defame1297 <gitea@rkdr.net>
Reviewed-on: https://git.dev.rkdr.net/Defame1297/holocron/pulls/129
Co-authored-by: Claude Code AI - Gitea MCP <claude@noreply.git.dev.rkdr.net>
Co-committed-by: Claude Code AI - Gitea MCP <claude@noreply.git.dev.rkdr.net>
This commit was merged in pull request #129.
This commit is contained in:
Claude Code AI - Gitea MCP
2026-09-01 13:47:46 +00:00
committed by Defame1297
parent 0e91a3ae66
commit 598a7c326a
420 changed files with 15303 additions and 4740 deletions

View File

@@ -16,15 +16,48 @@ Arguments:
Exit codes:
0 All checks passed (or nothing to validate, or not plugin scope)
1 One or more checks failed
2 Script error (unrecognized file extension — expected .md or .agent.md)
2 Usage error, or the argument is not an agent file this script can read
An exit code of 2 is NOT a finding. SKILL.md tells the auditor to surface a
non-zero exit as findings, so a usage error leaving exit 1 with nothing on
stdout was indistinguishable from a clean-but-failing run. Environment and
argument problems exit 2; only real findings exit 1.
Exit 2 and the silent exit 0 answer two DIFFERENT questions, and neither may
be spelled with the other's code:
exit 2 the argument is not something this script can audit at all — it is
missing, doubled, not a file, or not named .md / .agent.md. Decided
before the scope walk-up runs, from the argument alone.
exit 0 the argument IS a readable agent file, and the scope walk-up found
no type:-bearing apm.yml above it before hitting the \$HOME, .git or
filesystem-root boundary. That is a real verdict about a real file —
"this agent is user or project scope, so plugin-scope provenance
does not apply to it" — not a rejected input.
scripts/check-scope-walkup-sync.sh's fixture 6 pins the second: a real agent
file under a \$HOME with a type-bearing apm.yml ABOVE it must exit 0 with empty
output. Widening exit 2 to cover "the walk-up found no package" would break
that fixture AND would be wrong on its own terms, because new-agent.sh happily
scaffolds exactly that layout.
Checks performed:
0 source_keys present in agent pair but sources.md absent
1 FILL IN: placeholders in sources.md
2 source_keys in agent files → slug exists in sources.md
3 Contributing files listed in sources.md exist on disk (plugin-root relative)
3 Contributing files listed in sources.md exist on disk (plugin-root
relative). An explicit '(none)' skips silently; a Contributing files block
this parser cannot read is reported as an INFO saying checks 3 and 4 did
not run, never skipped silently.
4 Contributing files back-reference the parent slug in their source_keys
5 Research doc field present and not placeholder
This script has no counterpart to skill-audit's checks 6, 7 and 8 (Research
doc field / upstream forward / upstream reverse are numbered 6, 7, 8 there and
5 here): an agent at plugin scope is a single file with a plugin-root
sources.md, so there is no references/ tree to walk and no upstream research
source index to cross-check. parse_status() and the sources.md-basename gate
that those checks need exist only in the skill-audit copy.
EOF
}
@@ -33,26 +66,131 @@ if [[ "${1:-}" == "--help" || "${1:-}" == "-h" ]]; then
exit 0
fi
# Usage and environment problems exit 2, findings exit 1. See the usage text
# above for why the two must not share a code, and for why "not plugin scope"
# is neither of them. This is a deliberate divergence from validate.sh, which
# has no 2 tier for content: validate.sh always prints PASS lines, so a usage
# error there is visibly not a findings report. This script prints NOTHING on a
# clean run, so exit 1 plus empty stdout was the only signal a caller got
# either way.
if [[ $# -lt 1 ]]; then
echo "Error: agent-file is required." >&2
echo "" >&2
usage >&2
exit 1
exit 2
fi
# Extra positional arguments were silently dropped, so a typo'd flag or a second
# path looked like it had been honoured.
if [[ $# -gt 1 ]]; then
echo "Error: expected exactly one argument, got $#: $*" >&2
echo "" >&2
usage >&2
exit 2
fi
# python3 is a HARD dependency. Without this preflight a missing interpreter
# produced 'line NN: python3: command not found' and exit 127 — an exit code no
# caller maps to anything, from a message that names this script's line number
# rather than the missing dependency.
if ! command -v python3 > /dev/null 2>&1; then
echo "Error: python3 is required but was not found on PATH." >&2
echo " Why: skipping the provenance checks entirely would be a vacuous pass." >&2
echo " Fix: install python3 (pre-commit itself is a Python application, so it is almost certainly already present)." >&2
exit 2
fi
# A path that does not exist, or exists but is not a regular file, used to reach
# the Python body, get os.path.dirname()'d into some ancestor directory and then
# either report a silent exit 0 (no package above it) or — worse — audit a
# DIFFERENT agent's package while naming the typo'd path. A typo'd target was
# indistinguishable from a clean agent. vale-wrap.sh hard-errors on a
# nonexistent path for exactly this reason.
#
# This is decided from the argument alone, before any walk-up runs, so it cannot
# collide with the not-plugin-scope exit 0: that verdict is only ever reached by
# a file that got past here.
if [[ ! -e "$1" ]]; then
echo "Error: no such file: $1" >&2
echo " Why: a nonexistent target would otherwise report a silent pass." >&2
echo " Fix: pass the path of the agent file to validate." >&2
exit 2
fi
if [[ ! -f "$1" ]]; then
echo "Error: not a regular file: $1" >&2
echo " Why: this script audits one agent file, not a directory of them, and reporting a directory as a pass hides the wrong-target mistake." >&2
echo " Fix: pass the agent file itself — .apm/agents/<name>.agent.md — not its parent directory." >&2
exit 2
fi
# The extension check used to live inside the Python body. It stays exit 2 and
# keeps its wording; it moves up here so that every "this argument is not
# auditable" verdict is reached in one place, before the interpreter starts and
# before the scope walk-up can turn a bad argument into a silent exit 0.
case "$1" in
*.agent.md | *.md) ;;
*)
echo "Error: unrecognized extension '$(basename "$1")' — expected .md or .agent.md" >&2
exit 2
;;
esac
python3 -u - "$1" <<'PYTHON'
import sys
import os
import re
# Output is UTF-8 for the same reason input is: under LC_ALL=C the streams
# default to ASCII, and every finding this script prints contains an em dash.
# Pinning only the reads moved the crash from the read to the write — a
# UnicodeEncodeError inside print_findings(), which loses the whole report
# after all the checks have already run.
for _stream in (sys.stdout, sys.stderr):
try:
_stream.reconfigure(encoding='utf-8')
except AttributeError: # pragma: no cover — Python < 3.7
pass
agent_file = os.path.abspath(sys.argv[1])
fname = os.path.basename(agent_file)
agent_dir = os.path.dirname(agent_file)
# --- Sanity-check extension (single vendor-neutral .agent.md file at plugin/APM scope) ---
if not (fname.endswith('.agent.md') or fname.endswith('.md')):
print(f"Error: unrecognized extension '{fname}' — expected .md or .agent.md", file=sys.stderr)
sys.exit(2)
# --- Input ----------------------------------------------------------------
# Ported from the skill-audit copy, where the same two problems were already
# fixed.
#
# read_text() pins UTF-8 explicitly instead of inheriting
# locale.getpreferredencoding(), which is ASCII under LC_ALL=C — an ordinary em
# dash in an agent file or in sources.md then aborted the run with a bare
# UnicodeDecodeError traceback, or, at the one call site that wrapped its read
# in `except Exception: return []`, reported the unreadable file as having no
# source_keys and therefore as clean. A file that genuinely is not UTF-8 still
# fails; it just says which file and why.
#
# strip_bom() runs on every read because a leading BOM defeats
# parse_frontmatter()'s `^---` anchor, which silently disabled check 2 on a
# BOM-prefixed agent file: no frontmatter parsed means no source_keys parsed
# means nothing to validate.
class EncodingError(Exception):
pass
def strip_bom(text):
return text[1:] if text.startswith(u'\ufeff') else text
def read_text(path):
"""File contents as text, UTF-8 and BOM-free, with a diagnostic instead of a traceback."""
try:
with open(path, encoding='utf-8') as fh:
return strip_bom(fh.read())
except UnicodeDecodeError as exc:
raise EncodingError(
"not valid UTF-8 (%s at byte %d) — re-save the file as UTF-8; "
"this gate does not guess at other encodings"
% (exc.reason, exc.start))
# Matches a top-level `type:` line whose value is exactly one of the four
# package content types — identical to validate.sh's APM_TYPE_RE. Group 1's
@@ -67,15 +205,31 @@ TYPE_RE = re.compile(r"^type:\s*(['\"]?)(instructions|skill|hybrid|prompts)\1(?:
# keep walking. Stop at a $HOME boundary, a .git boundary, or the filesystem
# root: none of these is plugin/APM scope, so this script has nothing to
# check there.
#
# Returning None here means NOT PLUGIN SCOPE, which is a verdict, not an error:
# the caller exits 0 silently, and scripts/check-scope-walkup-sync.sh fixture 6
# pins that. It is deliberately NOT folded into the exit-2 tier above.
def find_plugin_root(start_dir):
home = os.path.expanduser('~')
current = os.path.abspath(start_dir)
while True:
apm_yml = os.path.join(current, 'apm.yml')
if os.path.isfile(apm_yml):
with open(apm_yml) as f:
if any(TYPE_RE.match(line) for line in f):
return current
# An apm.yml is a manifest this script must be able to READ to
# classify scope at all. Under LC_ALL=C the old bare open() decoded
# as ASCII, so a manifest with an accented author name raised
# UnicodeDecodeError mid-walk and killed the run with a traceback.
# It is an environment problem, not a finding, so it exits 2 rather
# than being swallowed into a silent "no package here".
try:
content = read_text(apm_yml)
except EncodingError as exc:
print(
"Error: %s is %s" % (apm_yml, exc),
file=sys.stderr)
sys.exit(2)
if any(TYPE_RE.match(line) for line in content.splitlines()):
return current
# $HOME is a non-plugin-scope boundary — checked before the .git test
# below (mirrors validate.sh's detect_scope ordering), so a
# dotfiles-managed $HOME (yadm, chezmoi bare-repo, etc.) can't shadow
@@ -101,7 +255,14 @@ if plugin_root is None:
sources_md_path = os.path.join(plugin_root, 'sources.md')
# --- Helpers ---
PLACEHOLDER_RE = re.compile(r'(?<!`)FILL IN:[^`\n]')
# The trailing character class used to be CONSUMING — `[^`\n]` — so a
# `FILL IN:` at end of line matched nothing and escaped checks 1 and 5
# entirely. `- **Description:** FILL IN:` is the most likely spelling of a
# half-written entry, and it was the one spelling the placeholder gate could
# not see. The exclusion it was really expressing is "not inside backticks",
# which a lookahead states without eating a character.
PLACEHOLDER_RE = re.compile(r'(?<!`)FILL IN:(?!`)')
def parse_frontmatter(content):
m = re.match(r'^---\n(.*?)\n---', content, re.DOTALL)
@@ -130,7 +291,56 @@ def parse_source_keys(fm):
def parse_h2_slugs(content):
return re.findall(r'^## (.+)$', content, re.MULTILINE)
# ===== BEGIN SHARED CONTRIBUTING-FILES PARSER =====
# ONE parser, embedded VERBATIM in two scripts:
# plugins/kyberforge/.apm/skills/skill-audit/scripts/validate-provenance.sh
# plugins/kyberforge/.apm/skills/agent-audit/scripts/validate-provenance.sh
# The block between these markers must stay byte-identical in both. It is
# copied rather than imported because a cache-installed plugin's scripts cannot
# read files outside their own plugin directory, so there is no single file both
# can share — the same constraint that forces the ADR-0020 boundary resolver to
# be duplicated across three scripts. Edit one copy, then paste it over the
# other.
#
# tests/test-adr0020-contract.sh hashes both copies and fails on drift. Before
# it did, the agent-audit copy's docstring merely ASSERTED the two were
# "behaviourally identical" and nothing checked it — which is how the two
# already-diverged spellings of the bullet loop went unnoticed.
#
# Requires: re (imported by the host script).
def parse_contributing_files(content, slug):
"""Find the Contributing files for a given slug H2 in content.
Both authored forms are accepted, because both are in use across the
corpus and only recognising the first silently skipped the contributing-
file checks on every sources.md written the other way:
- **Contributing files:** SKILL.md, references/a.md
**Contributing files:**
- SKILL.md (what this source contributed)
- references/a.md (what this source contributed)
Returns a list of paths with any trailing parenthetical note stripped.
Note the bullet form's notes may themselves contain commas, so the list
is built per bullet rather than by splitting the joined value.
The three return values are NOT interchangeable, and callers depend on
the distinction:
[path, ...] the entry names contributing files
[] the entry EXPLICITLY records "(none)"
None the entry says nothing this parser can read
Only an explicit "(none)" yields []. A "Contributing files:" heading
followed by a numbered list, by `*` bullets, or by prose parses nothing
and returns None, never [] — a caller reads [] as a deliberate "no
contributing files" record and SKIPS its check on that basis, so a parse
failure returning [] would silently disable the check instead of leaving
the unreadable entry exposed to it.
"""
pattern = re.compile(
r'^## ' + re.escape(slug) + r'\s*\n(.*?)(?=^## |\Z)',
re.MULTILINE | re.DOTALL
@@ -139,72 +349,153 @@ def parse_contributing_files(content, slug):
if not m:
return None
block = m.group(1)
def strip_note(entry):
# "references/a.md (why)" -> "references/a.md"
return re.sub(r'\s*\(.*$', '', entry).strip()
# Inline form: value on the same line, comma-separated, no notes.
cf_m = re.search(r'^\- \*\*Contributing files:\*\* (.+)$', block, re.MULTILINE)
if cf_m:
value = cf_m.group(1).strip()
if value.startswith("(none"):
return []
return [p for p in (strip_note(x) for x in value.split(","))
if p] or None
# Bullet form: heading on its own line, one file per following bullet.
cf_m = re.search(r'^\*\*Contributing files:\*\*\s*$', block, re.MULTILINE)
if not cf_m:
return None
return cf_m.group(1).strip()
files = []
for line in block[cf_m.end():].splitlines():
line = line.strip()
if not line:
if files:
break
continue
if not line.startswith("- "):
break
entry = line[2:].strip()
if entry.startswith("(none"):
return []
entry = strip_note(entry)
if entry:
files.append(entry)
return files or None
# ===== END SHARED CONTRIBUTING-FILES PARSER =====
def parse_research_doc(content, slug):
def parse_research_docs(content, slug):
"""Every Research doc value under a given slug H2, in document order.
The caller uses the first and reports the rest. Returning only the first —
what this did before — meant a second '- **Research doc:**' line in one
entry was silently ignored, so an author who added a doc rather than
replacing one got check 5 run against the old value and no hint that the
new one was never looked at.
"""
pattern = re.compile(
r'^## ' + re.escape(slug) + r'\s*\n(.*?)(?=^## |\Z)',
re.MULTILINE | re.DOTALL
)
m = pattern.search(content)
if not m:
return None
return []
block = m.group(1)
rd_m = re.search(r'^\- \*\*Research doc:\*\* (.+)$', block, re.MULTILINE)
if not rd_m:
return None
return rd_m.group(1).strip()
return [v.strip() for v in
re.findall(r'^\- \*\*Research doc:\*\* (.+)$', block, re.MULTILINE)]
findings = []
has_fail = False
# A finding identical in every field is the same finding, and the same file is
# now reached by more than one check — the agent file is read once for its own
# source_keys and again as a contributing file, so an unreadable one would
# otherwise be reported twice with the same words. Distinct findings about the
# same file still both appear.
def _record(entry):
if entry not in findings:
findings.append(entry)
def emit_fail(desc, fpath, why, fix):
global has_fail
has_fail = True
findings.append(("FAIL", desc, fpath, why, fix))
_record(("FAIL", desc, fpath, why, fix, None))
# INFO does not set has_fail and does not change the exit code. It is for a
# check that could not RUN — an unverified entry, not a broken one — and it
# exists so that "did not run" is never spelled the same way as "passed".
def emit_info(desc, fpath, note):
_record(("INFO", desc, fpath, None, None, note))
def print_findings():
for kind, desc, fpath, why, fix in findings:
print(f"FAIL {desc} — {fpath}")
print(f" Why: {why}")
print(f" Fix: {fix}")
print()
for entry in findings:
kind = entry[0]
desc = entry[1]
fpath = entry[2]
why = entry[3]
fix = entry[4]
note = entry[5]
if kind == "FAIL":
print(f"FAIL {desc} — {fpath}")
print(f" Why: {why}")
print(f" Fix: {fix}")
print()
else:
print(f"INFO {desc} — {fpath}")
print(f" Note: {note}")
print()
def emit_unreadable(rel, exc):
"""Report a file this script cannot decode. Never a silent skip."""
emit_fail(
f"File is {exc}",
rel,
f"'{rel}' cannot be decoded, so its frontmatter — and any source_keys in it — "
f"cannot be read. This used to be swallowed by a bare 'except Exception: return []', "
f"which reported the unreadable file as having no source_keys and therefore as clean.",
f"Re-save '{rel}' as UTF-8."
)
# --- Collect source_keys from agent pair ---
def get_source_keys_from_file(fpath):
def get_source_keys_from_file(fpath, rel):
if not os.path.isfile(fpath):
return []
try:
with open(fpath) as f:
content = f.read()
except Exception:
content = read_text(fpath)
except EncodingError as exc:
emit_unreadable(rel, exc)
return []
fm, _ = parse_frontmatter(content)
return parse_source_keys(fm)
# Plugin/APM scope is a single vendor-neutral file — no counterpart to merge.
given_keys = get_source_keys_from_file(agent_file)
rel_given = os.path.relpath(agent_file, plugin_root)
given_keys = get_source_keys_from_file(agent_file, rel_given)
all_source_keys = given_keys
sources_md_exists = os.path.isfile(sources_md_path)
# Early exit: nothing to validate
# Early exit: nothing to validate. The read above can itself raise a finding —
# an unreadable agent file — so print before leaving; the clean case still
# prints nothing and exits 0.
if not all_source_keys and not sources_md_exists:
sys.exit(0)
print_findings()
sys.exit(1 if has_fail else 0)
sources_content = None
sources_slugs = set()
if sources_md_exists:
with open(sources_md_path) as f:
sources_content = f.read()
try:
sources_content = read_text(sources_md_path)
except EncodingError as exc:
emit_unreadable("sources.md", exc)
print_findings()
sys.exit(1)
sources_slugs = set(parse_h2_slugs(sources_content))
# --- Check 0: source_keys present but sources.md absent ---
if not sources_md_exists and all_source_keys:
rel_given = os.path.relpath(agent_file, plugin_root)
emit_fail(
"source_keys declared but sources.md is absent",
rel_given,
@@ -240,11 +531,49 @@ for fpath, keys in [(agent_file, given_keys)]:
)
# --- Checks 3, 4, 5: Per-slug checks in sources.md ---
for slug in parse_h2_slugs(sources_content):
# Check 3: Contributing files exist (paths relative to plugin root)
cf_value = parse_contributing_files(sources_content, slug)
if cf_value and not cf_value.startswith("(none"):
cf_files = [p.strip() for p in cf_value.split(",") if p.strip()]
# Every per-slug parser below — parse_contributing_files, parse_research_docs —
# locates its block with pattern.search(), so a slug written twice resolves to
# the FIRST block every time. Iterating the raw heading list therefore checked
# the first block's fields twice and the second block's never: a duplicated slug
# is half-validated, and looked fully validated. The duplicate is announced and
# the repeat visit dropped.
all_slugs = parse_h2_slugs(sources_content)
unique_slugs = []
for _slug in all_slugs:
if _slug in unique_slugs:
continue
unique_slugs.append(_slug)
_count = all_slugs.count(_slug)
if _count > 1:
emit_info(
f"Duplicate '## {_slug}' entry in sources.md — only the first block is checked",
f"sources.md (## {_slug})",
f"'## {_slug}' appears {_count} times. Every field parser here takes the first match, so the "
f"second and later blocks' Contributing files and Research doc are never validated — "
f"checks 3, 4 and 5 did not run for them. "
f"Merge the blocks into one entry, or give each a distinct slug and reference it from source_keys."
)
for slug in unique_slugs:
# Checks 3 and 4: Contributing files exist (paths relative to plugin root),
# and back-reference the slug. `[]` and None are NOT the same answer here.
# `[]` is the author writing "(none)" — there is nothing to check and the
# skip is correct. None is a Contributing-files block this parser cannot
# read, and skipping THAT silently disables both checks on the one entry
# least likely to be right, which is the failure mode
# parse_contributing_files' own docstring warns about. Say so out loud.
cf_files = parse_contributing_files(sources_content, slug)
if cf_files is None:
emit_info(
f"Contributing-file checks skipped for '{slug}' — the Contributing files block could not be parsed",
f"sources.md (## {slug})",
f"The '## {slug}' entry has no Contributing files list this parser can read — a missing field, a bare heading, '*' bullets, a numbered list, or prose all read as unparsable rather than as an empty declaration. "
f"Checks 3 and 4 did not run for this slug, so nothing verified that its contributing files exist or name it back. "
f"Write the value as '- **Contributing files:** <comma-separated paths>', or as a '**Contributing files:**' heading followed by '- ' bullets — "
f"or record '(none)' if this source contributed no files."
)
elif cf_files:
for cf_rel in cf_files:
cf_abs = os.path.join(plugin_root, cf_rel)
if not os.path.isfile(cf_abs):
@@ -256,8 +585,11 @@ for slug in parse_h2_slugs(sources_content):
)
else:
# Check 4: Bidirectional — file should list slug in its source_keys
with open(cf_abs) as f:
cf_content = f.read()
try:
cf_content = read_text(cf_abs)
except EncodingError as exc:
emit_unreadable(cf_rel, exc)
continue
cf_fm, _ = parse_frontmatter(cf_content)
cf_keys = parse_source_keys(cf_fm)
if slug not in cf_keys:
@@ -269,7 +601,17 @@ for slug in parse_h2_slugs(sources_content):
)
# Check 5: Research doc field required
rd_value = parse_research_doc(sources_content, slug)
rd_values = parse_research_docs(sources_content, slug)
if len(rd_values) > 1:
emit_info(
f"Multiple '- **Research doc:**' lines for '{slug}' — only the first is used",
f"sources.md (## {slug})",
f"The '## {slug}' entry has {len(rd_values)} Research doc lines; check 5 ran against the first "
f"('{rd_values[0]}') and never looked at the rest. "
f"Keep one Research doc line per entry — if a slug genuinely came from two documents, split it into two slugs, "
f"or name the extra document inside the first value's annotation where it is at least visible."
)
rd_value = rd_values[0] if rd_values else None
if rd_value is None:
emit_fail(
"Research doc field missing",

View File

@@ -69,6 +69,24 @@ import glob
import yaml
# Output is UTF-8 for the same reason input is: under LC_ALL=C the streams
# default to ASCII, and this script's own message text carries em dashes (the
# ADR-0020 boundary SUGGESTION is one). Pinning only the reads moved the crash
# from the read to the write — a UnicodeEncodeError raised while PRINTING, after
# every check has already run, which loses the whole report and (here) flips a
# clean exit 0 into a traceback and an exit 1. read_text() in the shared
# resolver block below pins the reads; this pins the writes.
#
# Deliberately OUTSIDE the ADR-0020 shared boundary resolver block: the two
# validate.sh copies print findings, skill-size-check.sh has its own top-level
# equivalent, and tests/test-adr0020-contract.sh hashes that block for
# byte-identity across all three.
for _stream in (sys.stdout, sys.stderr):
try:
_stream.reconfigure(encoding='utf-8')
except AttributeError: # pragma: no cover — Python < 3.7
pass
agent_file = os.path.abspath(sys.argv[1])
script_dir = sys.argv[2]
@@ -253,9 +271,27 @@ def _collect_package(pkg_dir, names):
safe_dir = glob.escape(pkg_dir)
for sub in ('.apm/skills/*/', 'skills/*/'):
for path in glob.glob(os.path.join(safe_dir, sub)):
names.add(os.path.basename(path.rstrip('/')).lower())
# A directory is a skill only if it HOLDS a SKILL.md. An empty
# leftover — a deleted skill whose directory survived, a scaffolding
# stub, an editor's stray mkdir — is untracked by git, so it exists
# on the machine that made it and nowhere else. Counting it made a
# boundary target resolve locally and dangle in a fresh clone: the
# same install-dependence the deployed-tree rule above exists to
# remove, arriving through a different door.
if os.path.isfile(os.path.join(path, 'SKILL.md')):
names.add(os.path.basename(path.rstrip('/')).lower())
for sub in ('.apm/agents/*.md', 'agents/*.md'):
for path in glob.glob(os.path.join(safe_dir, sub)):
# The same rule one directory over, which until now had no
# counterpart here at all: the skills branch above tests for a
# SKILL.md, the agents branch took every glob hit on trust. A
# DIRECTORY named `ghost-agent.md` matches `*.md` and glob does not
# tell the two apart, so a leftover of that shape resolved a routing
# target on the machine holding it and dangled everywhere else —
# identical install-dependence, arriving through the one door
# nobody guarded.
if not os.path.isfile(path):
continue
base = os.path.basename(path)
if base.endswith('.agent.md'):
base = base[:-len('.agent.md')]
@@ -446,8 +482,17 @@ def known_targets(start_dir):
# condition, pc-run's "run pre-commit hooks" reads as a route to a
# non-existent `pre-commit` skill.
# * A BARE arrow target counts only in ADR-0020's compressed boundary form,
# `Not <thing> -> <skill-name>`. Without that, diagnose's process chain
# "fix -> regression-test" reads as a route to `regression-test`.
# `Not <thing> -> <skill-name>`. The example that motivated it is gone:
# diagnose's process chain "fix -> regression-test", which without the
# gate read as a route to a non-existent `regression-test` skill, was cut
# when issue #99 retrofitted that description. So the gate is currently
# UNEXERCISED — gating and not gating produce the same verdict corpus-wide.
# Keep it anyway. It is a false-positive guard against prose no one has
# written yet, and any new process chain re-arms it. Unexercised is not the
# same as unnecessary, and the branch it guards is still load-bearing: the
# bare-arrow rule is the sole extractor for three real targets in
# kyberforge's audit skills (agent-audit -> agent-author, agent-audit ->
# skill-audit, skill-audit -> skill-author), all written unbackticked.
# * A backticked hyphenated token counts only inside a boundary sentence.
# Unconditionally, `pre-push` or `commit-msg` in a TRIGGER clause is a hard
# FAIL with no escape hatch. Gating it costs nothing (measured over this
@@ -525,6 +570,66 @@ def known_targets(start_dir):
# ambiguity to resolve, and an author who wants a route checked unconditionally
# has two ways to say so.
#
# BOTH FORMS ARE SWEPT FOR ON THEIR OWN, and that is a repair of the promise
# above rather than a widening of it. Until the sweeps existed, notation was
# only ever seen as the OBJECT OF A ROUTE VERB (`use
# /name`) or as the tail of a `not ... ->` clause with no `;` or sentence end in
# between. Every one of these therefore exited 0 in total silence — no ERROR, no
# SUGGESTION, not even the target's name:
# Do not use for Y — /no-such-skill instead.
# Do not use for Y; /no-such-skill handles that.
# Do not use for Y (/no-such-skill covers it).
# Do not use for Y — that is /no-such-skill's job.
# Do not use for Y — defer to /no-such-skill.
# Do not use for Y — /no-such-skill.
# Do not use for Y; -> no-such-skill covers it.
# For W, /no-such-skill is the right entry point.
# The target was never EXTRACTED, so the notation-first rule in _add() had
# nothing to apply itself to and the "always blocks" promise was false for the
# ordinary way an author writes the thing. The SUGGESTION tier made it worse
# than a gap: its printed remedy tells the author to "write it as `/name` or
# `-> name` and it will be checked properly", and taking that advice turned a
# visible SUGGESTION into silence — the gate teaching the one edit that blinds
# it.
#
# THE TWO SWEEPS ARE GATED DIFFERENTLY, and the asymmetry is the whole point.
# `/name` is Claude Code's invocation syntax and nothing else — no English
# sentence contains one by accident — so the ADR-0020 amendment and
# docs/spec/gates.md both promise it blocks UNCONDITIONALLY, for any name. So
# NOTATION_SLASH is swept over every sentence, boundary marker or not. Gating it
# on BOUNDARY_MARKER made that promise false for the last sentence of
# Do not use for Z — use /real-skill instead.
# For W, /no-such-skill is the right entry point.
# which exited 0 in total silence: the boundary clause is one sentence up, so
# the sweep never looked at the sentence carrying the broken route. Extraction is
# per-sentence by design (corroboration is scoped to one sentence), which is
# exactly what made the gap invisible.
#
# NOTATION_ARROW stays gated on BOUNDARY_MARKER, and so does the backtick sweep.
# Neither form is unambiguous: `-> name` is also how a process chain is written
# ("reproduce -> minimise -> regression-test") and a code span is how a tool, a
# file and a skill are all cited. Ungating either would fire on prose that
# carries no routing intent at all — the false-positive class this whole
# extractor is tuned against.
#
# BOTH `/name` PATTERNS REFUSE A TOKEN THAT IS PART OF A PATH: a following `/`,
# or a `.` followed by a non-space, means `references/foo.md`, `docs/a/b.md` or
# `https://x/y`, not a route. A sentence's closing `.` is not followed by a
# non-space, so `— /no-such-skill.` still counts.
#
# THAT GUARD IS WRITTEN `(?![\w-])` AND NOT `\b`, because `\b` is not a guard at
# all here: it holds after a hyphen, so when the trailing lookahead rejected the
# full segment the engine simply backtracked to a shorter hyphen-terminated
# prefix and reported THAT as a route. Every one of these was a hard blocking
# ERROR naming a skill nobody had written:
# the config lives at /opt-tools/bin/thing. -> 'opt'
# see /api-docs/v2.md for the schema. -> 'api' AND 'api-docs'
# the file /no-such-skill.md documents it. -> 'no-such'
# `(?![\w-])` forbids the shortened prefix outright, so the whole segment is
# rejected as the path it is. MARKED_TARGET carries the same guard: it had no
# trailing lookahead whatsoever, so `see /api-docs/v2.md` raised the second of
# the two errors above through the route-verb path rather than the sweep.
#
# NAMESPACE: `plugin:skill` is live in this repo (native user-scope installs
# still resolve `gitea:gitea-prs`), so the patterns admit an optional
# `<plugin>:` prefix and normalize_target() strips it before resolution.
@@ -535,7 +640,8 @@ ROUTE_VERB = (r"(?:use|uses|using|run|runs|invoke|invokes|invoking|try|see"
r"|that'?s|compose|composes|call|calls"
r"|routes?\s+to|delegates?\s+to|prefers?|switch(?:es)?\s+to"
r"|hands?\s+off\s+to)")
MARKED_TARGET = r"(?:`/?(%s)`|(?<![\w./*-])/(%s)\b)" % (NAME_ANY, NAME_ANY)
MARKED_TARGET = (r"(?:`/?(%s)`|(?<![\w./*-])/(%s)(?![\w-])(?!/|\.\S))"
% (NAME_ANY, NAME_ANY))
ANY_TARGET = r"(?:%s|(%s)\b)" % (MARKED_TARGET, NAME_HYPH)
ROUTE_MARKED = re.compile(r"\b%s\s+(?:the\s+|an?\s+)?%s" % (ROUTE_VERB, MARKED_TARGET), re.I)
ROUTE_ANY = re.compile(r"\b%s\s+(?:the\s+|an?\s+)?%s" % (ROUTE_VERB, ANY_TARGET), re.I)
@@ -548,12 +654,53 @@ ROUTE_ANY = re.compile(r"\b%s\s+(?:the\s+|an?\s+)?%s" % (ROUTE_VERB, ANY_TARGET)
CONT_MARKED = re.compile(r"\s*(?:or|and|/|,)\s*%s" % MARKED_TARGET, re.I)
CONT_ANY = re.compile(r"\s*(?:or|and|/|,)\s*%s" % ANY_TARGET, re.I)
ARROW_MARKED = re.compile(r"(?:->|→)\s*%s" % MARKED_TARGET, re.I)
ARROW_BOUNDARY = re.compile(r"\bnot\b[^.;]*?(?:->|→)\s*(%s)\b" % NAME_HYPH, re.I)
# The two EXPLICIT ROUTE NOTATION sweeps. NOTATION_SLASH runs over EVERY
# sentence; NOTATION_ARROW is scoped to a boundary sentence by its caller (see
# the asymmetry note in the header). NOTATION_SLASH is deliberately not a reuse
# of MARKED_TARGET's `/name` alternative: that one only ever runs behind a route
# verb or an arrow, and it may match a namespaced or path-adjacent token in
# positions this free-standing sweep must refuse.
# NOTATION_ARROW is ARROW_BOUNDARY minus its leading `\bnot\b%s*?`, which is
# what made `Do not use for Y; -> no-such-skill covers it.` invisible:
# CLAUSE_BODY cannot cross the `;`, so the clause's own punctuation disarmed the
# check. Dropping that prefix costs the one false positive the bare-arrow bullet
# above names — a process chain ending in a hyphenated word, `Instead, reproduce
# -> minimise -> regression-test.` — and costs it only in a sentence that already
# carries a BOUNDARY_MARKER. That exposure is neither new nor larger: the same
# chain written `Do not use for X — reproduce -> regression-test.` was already a
# hard ERROR under ARROW_BOUNDARY, so this changes which boundary words reach the
# arrow, not whether prose can. An author who means the chain and not a route
# writes it in its own sentence, where neither pattern looks.
NOTATION_SLASH = re.compile(
r"(?<![\w./*-])/(%s)(?![\w-])(?!/|\.\S)" % NAME_ANY, re.I)
NOTATION_ARROW = re.compile(r"(?:->|→)\s*(%s)\b" % NAME_HYPH, re.I)
# CLAUSE_BODY is what may sit between `Not` and the arrow, and it is NOT
# `[^.;]`. That class cannot cross a `.`, so every boundary clause naming a
# DOTTED FILENAME between the two — `.pre-commit-config.yaml`, `AGENTS.md`,
# `.vale.ini` — was invisible to both patterns below, and the two resulting
# failures were different sizes (issue #110):
# * with a BACKTICKED target the clause was MISDIAGNOSED. The backtick sweep
# still extracted the target, so the route was checked, but the gate
# reported "no boundary clause" on a clause that was present and working.
# Three authors in two retrofit waves reworded a correct clause to satisfy
# the regex, one of them stripping the very filename that discriminates the
# skill from its neighbour.
# * with a BARE target the clause was UNCHECKED. ARROW_BOUNDARY is the only
# extractor for a bare arrow target, so `Not AGENTS.md -> no-such-skill`
# produced no target, no dangling report and no missing-clause SUGGESTION.
# Silence, not noise — the worse of the two failure modes.
# A dot inside a filename is followed by a non-space; a sentence-ending dot is
# followed by whitespace or by end of string. So the class admits a `.` only
# when the next character is not whitespace, which crosses `AGENTS.md` and
# still stops at a real sentence end.
CLAUSE_BODY = r"(?:[^.;]|\.(?=\S))"
ARROW_BOUNDARY = re.compile(
r"\bnot\b%s*?(?:->|→)\s*(%s)\b" % (CLAUSE_BODY, NAME_HYPH), re.I)
BACKTICK = re.compile(r"`(%s)`" % NAME_HYPH, re.I)
# A boundary clause takes two shapes and BOTH count: the prose markers, and
# ADR-0020's compressed arrow form `Not <thing> -> <name>`.
BOUNDARY_MARKER = re.compile(r"\b(?:do\s+not|instead|rather\s+than|not\s+for)\b", re.I)
BOUNDARY_ARROW = re.compile(r"\bnot\b[^.;]*?(?:->|→)", re.I)
BOUNDARY_ARROW = re.compile(r"\bnot\b%s*?(?:->|→)" % CLAUSE_BODY, re.I)
# Sentence boundaries decide the CORROBORATION scope above, so getting one wrong
# is not cosmetic — it moves a target between SUGGESTION and blocking ERROR. Two
# shapes common in these descriptions defeat the naive "period, space, capital"
@@ -573,9 +720,17 @@ BOUNDARY_ARROW = re.compile(r"\bnot\b[^.;]*?(?:->|→)", re.I)
# a lowercase letter. Verified zero-delta on the current corpus (37 ERROR / 58
# SUGGESTION / 2 dangling before and after) — this protects the descriptions
# issue #99 is about to rewrite, not the ones already measured.
# re.I here too, and NOT as a tidy-up: this was the one pattern in the file
# built without it, contradicting the uniformity note on CONT_*/ARROW_* above.
# Without the flag `E.g.` and `I.e.` — the sentence-initial spellings, which is
# where an abbreviation most often lands — matched none of the lookbehinds, so
# the clause split at the abbreviation, the corroborating target was stranded on
# the far side of the cut, and a genuinely dangling target silently demoted from
# blocking ERROR to SUGGESTION. That is the OVER-SPLIT failure described
# directly above, still live for exactly the capitalised half of the input.
SENTENCE_SPLIT = re.compile(
u'(?<!\\be\\.g\\.)(?<!\\bi\\.e\\.)(?<!\\betc\\.)(?<!\\bvs\\.)(?<!\\bcf\\.)'
u'(?<=[.!?])\\s+(?=[A-Za-z`"“(])')
u'(?<=[.!?])\\s+(?=[A-Za-z`"“(])', re.I)
# The token that may follow a route target without turning it into a compound
# modifier: punctuation, end of sentence, a conjunction, a boundary word, or a
@@ -634,11 +789,26 @@ def _notation(text, start, arrow):
def _add(out, text, name, start, end, strict=None, arrow=False):
"""Record one target as (name, may_dangle, notation).
NOTATION IS DECIDED FIRST, and when it is set the follower test is skipped.
The header above promises that route notation "always blocks", and for the
`/name` form that was false: `-> name` reached this function with
strict=True from its two call sites, but `/name` did not, so it fell to
_terminal() and a follower outside FOLLOWER_OK set may_dangle=False. The
target then reached unresolved_targets() unblockable — and, before the
companion fix there, unreported as well. `... use /no-such-skill
afterwards.` exited 0 in total silence, on the one form ADR-0020 offers an
author who wants a route checked unconditionally.
"""
if not name:
return
notation = _notation(text, start, arrow)
if strict is None and notation:
strict = True
out.append((name,
_terminal(text, end) if strict is None else strict,
_notation(text, start, arrow)))
notation))
def _scan(text, route_re, cont_re, out):
@@ -678,7 +848,19 @@ def _extract_sentence(sentence):
for match in ARROW_BOUNDARY.finditer(sentence):
_add(out, sentence, match.group(1), match.start(1), match.end(1),
strict=True, arrow=True)
# `/name` wherever it sits, in ANY sentence — not only where a route verb or
# an arrow happens to precede it, and NOT only inside a boundary sentence.
# See the EXPLICIT ROUTE NOTATION note in the header for the eight phrasings
# this recovers and for why silence was the failure mode. The sweep takes no
# follower test: _add() reads the notation first and marks it.
for match in NOTATION_SLASH.finditer(sentence):
_add(out, sentence, match.group(1), match.start(1), match.end(1))
if boundary:
# The arrow and backtick forms are ambiguous in ordinary prose, so they
# stay scoped to a sentence that carries a boundary marker.
for match in NOTATION_ARROW.finditer(sentence):
_add(out, sentence, match.group(1), match.start(1), match.end(1),
strict=True, arrow=True)
for match in BACKTICK.finditer(sentence):
_add(out, sentence, match.group(1), match.start(1), match.end(1))
return out
@@ -697,6 +879,85 @@ def boundary_targets(description):
return sorted({name for name, _, _ in _extract(description)})
def _arrow_targets(description):
"""Names extracted from ARROW notation specifically.
Kept apart from boundary_targets() because the arrow form is the one shape
that ALWAYS names a target: ADR-0020's `Not <thing> -> <name>`. A clause
written that way from which nothing could be extracted is a parse failure
that deserves its own message, and telling it apart needs the arrow targets
alone rather than every target in the description.
"""
out = []
for sentence in SENTENCE_SPLIT.split(description):
for match in ARROW_MARKED.finditer(sentence):
name, _, _ = _first(match)
if name:
out.append(name)
for match in ARROW_BOUNDARY.finditer(sentence):
out.append(match.group(1))
return out
def boundary_clause_status(description):
"""'absent', 'unparsed' or 'present' — three outcomes, not two.
Issue #110's standing request: the gate must distinguish "no boundary
clause" from "boundary clause I could not parse". Reporting the first for
the second sends the author hunting for a problem that is not there, and
three of them reworded a correct clause to satisfy a regex instead.
'unparsed' is the narrow, certain case: an ADR-0020 arrow clause was
detected and NO target came out of it. The arrow form always names one, so
zero targets means the name is written in a shape the extractor cannot see
— a single-word bare target (`Not X -> forge`, which has to be written
`` `forge` `` or `/forge`) is the live example, since single-word names are
deliberately not matchable bare.
A PROSE clause yielding no target is NOT reported: "Do not use for anything
else" is a complete and legitimate boundary clause that names nowhere to go.
"""
if BOUNDARY_ARROW.search(description) and not _arrow_targets(description):
return 'unparsed'
if has_boundary_clause(description):
return 'present'
return 'absent'
def multi_target_arrow_clauses(description):
"""[(first, second)] for arrow clauses naming more than one target.
Issue #107: only the FIRST target after an arrow is resolved. The
conjunction continuation (CONT_*) is wired to the prose route verbs and
never to arrows, so `Not X -> a or b` resolved `a`, left `b` neither
resolved nor reported, and then printed "1 of 1 boundary target(s) resolve"
on a clause naming two — a gate under-reporting its own coverage, which is
the one failure mode ADR-0020 says a gate must not have.
The clause is REJECTED rather than the arrow scan extended. Extending it
would widen the resolver's deliberately conservative false-positive tuning
across every arrow in the corpus; rejecting costs nothing and makes the
one-arrow-per-target convention — already what every retrofitted gitea
skill does in practice — explicit instead of folkloric. The caller emits a
SUGGESTION telling the author to split.
"""
hits = []
for sentence in SENTENCE_SPLIT.split(description):
matches = (list(ARROW_MARKED.finditer(sentence))
+ list(ARROW_BOUNDARY.finditer(sentence)))
for match in matches:
first, _, _ = _first(match)
if not first:
continue
cont = CONT_ANY.match(sentence, match.end())
if not cont:
continue
second, _, _ = _first(cont)
if second:
hits.append((first, second))
return hits
def unresolved_targets(description, known):
"""Targets resolving to nothing, split into (blocking, reported).
@@ -713,6 +974,17 @@ def unresolved_targets(description, known):
Everything else is reported and left alone. `known` is the resolved
universe from known_targets(); passing an empty set is not meaningful —
callers check for that first and decline out loud instead.
A NON-TERMINAL target is reported, never dropped. FOLLOWER_OK is a closed
whitelist of maybe eighty words, so the follower rule says "this token is
outside a list I keep" and not "this is prose" — and the old `continue`
turned that into invisibility at every tier. The gate then failed OPEN on
its own unfamiliarity: any target followed by a word nobody thought to
enumerate was neither blocked nor mentioned, so the check that did not run
said nothing about not running. The follower rule may withdraw the power to
BLOCK a commit — that is what it was added for, and the ATTRIBUTIVE USE note
above is the argument for it — but it may not withdraw visibility, which is
the same rule the corroboration tier already follows.
"""
blocking, reported = set(), set()
for sentence in SENTENCE_SPLIT.split(description):
@@ -721,7 +993,10 @@ def unresolved_targets(description, known):
if normalize_target(name) in known}
for name, may_dangle, notation in found:
key = normalize_target(name)
if key in known or not may_dangle:
if key in known:
continue
if not may_dangle:
reported.add(name)
continue
if notation or (resolved - {key}):
blocking.add(name)
@@ -801,6 +1076,47 @@ def description_value(fm_text):
return re.sub(r'\s+', ' ', value).strip()
def hand_invoked(fm_text):
"""True when the frontmatter marks this file as reached only by hand.
`disable-model-invocation: true` removes a skill from the model-visible
listing entirely — it is not preloaded, and the Skill tool refuses to call
it — so its description is never matched against user intent. ADR-0020 and
skill-author's contract give such a skill ONE plain human-facing sentence:
no trigger list, no boundary clause. No validator knew the field existed
(issue #108), so the boundary-clause SUGGESTION fired on exactly the shape
the contract mandates, and its remedy — "add a boundary clause so the router
knows where NOT to send this skill" — was addressed to a router that cannot
see the skill at all. An author who followed the advice made the file worse.
Only the ROUTING rules are lifted. The body word budget still applies: the
body is loaded on invocation like any other, and competes with the caller's
live conversation the same way. So does the 400-character description FAIL —
a hand-invoked description is not preloaded, but it is still the one line
the user reads when choosing from the `/` menu, and the ceiling is the
outlier stop rather than the style target.
A parse failure returns False rather than raising. This is a MODIFIER on
other checks, not a check of its own: the frontmatter's validity is decided,
and failed, by description_value() on the same text, and raising a second
exception here would report one broken file twice with two different
diagnoses.
"""
try:
data = yaml.safe_load(fm_text)
except Exception:
return False
if not isinstance(data, dict):
return False
value = data.get('disable-model-invocation')
if isinstance(value, str):
# PyYAML already resolves the unquoted YAML 1.1 booleans, so this only
# catches a QUOTED "true" — which a host reads as truthy and which no
# gate should treat as opting back in to the routing rules.
return value.strip().lower() in ('true', 'yes', 'on')
return value is True
# --- Body-shape checks (skills only; agents have no references/ dir) -------
# Deterministic and countable, so they are enforced here. Whether a given
# gotcha is WARRANTED is semantic and stays the auditor's judgment, which is why
@@ -915,7 +1231,15 @@ def missing_reference_pointers(body, skill_dir):
end = masked.find('\n', match.end())
if end < 0:
end = len(masked)
if REFERENCE_PAST.search(masked[start:end]):
# The pointer's OWN SPAN is excised before the sweep. Run over the
# whole line, the past-tense test matched the very path it was judging,
# so a file exempted itself by its NAME: `references/deprecated-api.md`,
# `references/removed-flags.md` and `references/gone.md` produced no
# ERROR at all, while `references/missing.md` — an identical break —
# errored. The exemption is about what the SENTENCE says about the
# pointer, never about what the pointer is called.
line = masked[start:match.start()] + masked[match.end():end]
if REFERENCE_PAST.search(line):
continue
if REFERENCE_QUALIFIER.search(masked[start:match.start()]):
continue
@@ -968,8 +1292,14 @@ def agent_description(fm, local_fname):
f"not run — {local_fname}")
return None
def check_description_budget(value, local_fname):
"""ADR-0020 description gates — identical for every scope."""
def check_description_budget(value, local_fname, by_hand=False):
"""ADR-0020 description gates — identical for every scope.
`by_hand` is ADR-0020's hand-invocation carve-out (issue #108): an agent
carrying `disable-model-invocation: true` is absent from the model-visible
listing, so the 250-character SUGGESTION — a routing-quality budget — has
no listing to apply to. The 400-character ceiling is unaffected.
"""
if not value:
return
dlen = len(value)
@@ -979,13 +1309,13 @@ def check_description_budget(value, local_fname):
f"agent is invoked. Keep a trigger clause, at most one capability clause, "
f"and a boundary clause; move capability enumeration, output-format detail, "
f"composition notes and implementation detail to the body — {local_fname}")
elif dlen > DESC_SUGGEST_CHARS:
elif dlen > DESC_SUGGEST_CHARS and not by_hand:
suggest(f"description is {dlen} chars — over the {DESC_SUGGEST_CHARS}-character "
f"ADR-0020 target (hard fail at {DESC_MAX_CHARS}). The SUGGESTION tier is "
f"what moves the corpus average; the FAIL tier only stops outliers "
f"— {local_fname}")
def check_boundary(value, fpath, local_fname):
def check_boundary(value, fpath, local_fname, by_hand=False):
"""ADR-0020 boundary clause + resolvable boundary targets.
agent-author's SKILL.md states that an agent's boundary targets must
@@ -1002,10 +1332,31 @@ def check_boundary(value, fpath, local_fname):
# SUGGESTION, not FAIL: detecting the absence is deterministic, but whether
# this particular agent warrants a boundary clause is judgment. All four
# agents in this corpus currently lack one.
if not has_boundary_clause(value):
#
# THREE outcomes, not two: "no boundary clause" and "boundary clause I could
# not parse" are different findings (issue #110). And a hand-invoked agent is
# exempt from the clause altogether (issue #108) — the boundary-target
# resolution below still runs, because a target it DOES name should still
# resolve.
status = boundary_clause_status(value) if not by_hand else 'present'
if status == 'absent':
suggest(f"description has no boundary clause — add the prose form (\"Do not use "
f"for X — use `y` instead\") or ADR-0020's compressed form (\"Not X -> y\") "
f"so the router knows where NOT to send this agent — {local_fname}")
elif status == 'unparsed':
suggest(f"description has an arrow boundary clause (\"Not X -> y\") from which no "
f"target could be read, so the dangling-target check did not run on it — "
f"the clause is PRESENT and unparsed, not missing. Most often the target "
f"is a single word, which is deliberately not matchable bare: write it as "
f"`name` or /name — {local_fname}")
if not by_hand:
# One arrow, one target: a second name after the same arrow is resolved
# by nothing and reported by nothing (issue #107).
for first, second in multi_target_arrow_clauses(value):
suggest(f"an arrow boundary clause names more than one target ('{first}', then "
f"'{second}') and only the first is resolved — the second is checked by "
f"nothing. Split it into one arrow per target: \"Not X -> {first}. "
f"Not Y -> {second}.\" — {local_fname}")
targets = boundary_targets(value)
if not targets:
return
@@ -1240,8 +1591,9 @@ def check_apm_agent_file(fpath, allowlist, stem):
else:
if PLACEHOLDER_RE.search(folded):
fail(f"description contains unfilled FILL IN: placeholder — {local_fname}")
check_description_budget(folded, local_fname)
check_boundary(folded, fpath, local_fname)
by_hand = hand_invoked(fm)
check_description_budget(folded, local_fname, by_hand)
check_boundary(folded, fpath, local_fname, by_hand)
# body — required, non-empty, no placeholder; same Copilot truncation risk
# applies since this file compiles verbatim into a real Copilot file downstream.
@@ -1336,8 +1688,9 @@ def check_file(fpath, file_provider):
else:
if PLACEHOLDER_RE.search(folded):
fail(f"description contains unfilled FILL IN: placeholder — {local_fname}")
check_description_budget(folded, local_fname)
check_boundary(folded, fpath, local_fname)
by_hand = hand_invoked(fm)
check_description_budget(folded, local_fname, by_hand)
check_boundary(folded, fpath, local_fname, by_hand)
# body
if not body.strip():