feat(kyberforge): enforce the ADR-0020 context contract for skills and agents
Skill name+description pairs are preloaded into every session, costing ~6,200 tokens across 39 skills before any skill is invoked. The authoring rules mandated that growth: skill-author:104 and description-quality.md:21 both required padding, while skill-author:102 (the deflating rule) had no FAIL condition behind it. Gates (blocking, no baseline file): - description 250 chars SUGGESTION / 400 FAIL, measured on the folded YAML value - body-only 600 words SUGGESTION / 900 FAIL, independent of the unchanged whole-file 2770-word / 500-line spec backstop - every boundary-clause routing target must resolve to a real skill or agent; catches skill-improve, neuledge-context and gitea-labels - agents take the description gates but deliberately no body gate; a test pins that absence Vale: DescriptionOpener widened to ^This\b, new CompositionNote rule banning architecture notes from descriptions. 10 hits, 0 false positives. Kyberforge's own four skills retrofitted: descriptions 3,364 -> 938 chars (-72%), bodies 8,306 -> 2,487 words (-70%), all via the apm-workflow dispatch pattern. Fixes the skill-improve dangling route and the agent-author misroute to manual review. Also fixes a pre-existing false positive where any line-initial 'read ' was flagged as interactive input, which had already caused two scripts to be rewritten around it. Refs: ADR-0020
This commit is contained in:
@@ -1,18 +1,34 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
# Enforces agentskills.io's skill-authoring.md guidance: keep SKILL.md within
|
||||
# 500 lines and roughly 5,000 tokens, so the full body doesn't crowd out
|
||||
# conversation history and other active skills once loaded into context. Vale
|
||||
# can't express a whole-file length ceiling (its checks operate on text
|
||||
# patterns, not raw file size), so this is a plain script instead of a Vale
|
||||
# rule.
|
||||
# Two independent gate families live in this script. Do not conflate them.
|
||||
#
|
||||
# Both ceilings are inclusive: a file at exactly MAX_LINES or MAX_WORDS passes,
|
||||
# and only one past it fails. That matches skill-audit/scripts/validate.sh,
|
||||
# which has always used `line_count <= 500` as its pass condition — the two
|
||||
# previously disagreed at exactly 500 lines, so a SKILL.md could pass its own
|
||||
# audit and still be blocked by the commit hook.
|
||||
# 1. SPEC CONFORMANCE (MAX_LINES / MAX_WORDS, unchanged). Enforces
|
||||
# agentskills.io's skill-authoring.md guidance: keep SKILL.md within 500
|
||||
# lines and roughly 5,000 tokens, so the full body doesn't crowd out
|
||||
# conversation history and other active skills once loaded into context.
|
||||
# MAX_WORDS counts the WHOLE FILE, frontmatter included.
|
||||
#
|
||||
# 2. CONTEXT BUDGET (ADR-0020: description chars, body-only words, resolvable
|
||||
# boundary targets). A skill's name + description is preloaded into every
|
||||
# agent's context every session whether or not the skill is ever invoked;
|
||||
# the body is loaded only on invocation, and then competes with the caller's
|
||||
# live conversation. Those are different costs with different ceilings, so
|
||||
# they get their own numbers and their own measurements.
|
||||
#
|
||||
# The two families measure different things on purpose and neither replaces the
|
||||
# other: 2,770 whole-file words is a conformance backstop, 900 body-only words
|
||||
# is a quality gate, and a file can sit well inside one while failing the other.
|
||||
#
|
||||
# Vale can't express any of this (its checks operate on text patterns, not raw
|
||||
# file size), so this is a plain script instead of a Vale rule.
|
||||
#
|
||||
# Both spec ceilings are inclusive: a file at exactly MAX_LINES or MAX_WORDS
|
||||
# passes, and only one past it fails. That matches
|
||||
# skill-audit/scripts/validate.sh, which has always used `line_count <= 500` as
|
||||
# its pass condition — the two previously disagreed at exactly 500 lines, so a
|
||||
# SKILL.md could pass its own audit and still be blocked by the commit hook.
|
||||
# The ADR-0020 ceilings are inclusive the same way.
|
||||
#
|
||||
# Token counts aren't computed exactly here — word count (`wc -w`) is used as
|
||||
# a proxy. Measured over this repo's 39 in-scope SKILL.md files, characters per
|
||||
@@ -28,12 +44,18 @@ set -euo pipefail
|
||||
# such a file sit at exactly the ceiling and still spend ~5,240 tokens. A
|
||||
# median-density file at 2770 words spends ~4,700 tokens, so typical prose
|
||||
# gives up ~130 words of headroom to close that gap. The largest SKILL.md in
|
||||
# the repo is 2,489 words, so no current file is affected.
|
||||
# the repo is 2,760 words whole-file (skill-author), twelve words under the
|
||||
# ceiling — this is a gate two files have already grown into, not headroom.
|
||||
#
|
||||
# It is a one-sided proxy in the useful direction — nothing under the word
|
||||
# ceiling is wildly over the token ceiling — but it is not exact BPE
|
||||
# tokenization and does not replace one. Re-measure the corpus before treating
|
||||
# any of these numbers as still current.
|
||||
#
|
||||
# python3 is required for the ADR-0020 half. That is not a new dependency in
|
||||
# practice: pre-commit, which is how this script runs, is itself a Python
|
||||
# application. PyYAML is used when importable and is genuinely optional — the
|
||||
# fallback reader below recognises the frontmatter shapes this corpus uses.
|
||||
|
||||
# These constants are intentionally duplicated in
|
||||
# skill-audit/scripts/validate.sh (Python) rather than shared from one file:
|
||||
@@ -42,8 +64,23 @@ set -euo pipefail
|
||||
# vale-wrap.sh's per-plugin duplication — see its own header comment).
|
||||
# tests/test-skill-size-check.sh asserts both files agree on these values, so
|
||||
# drift between them fails CI rather than silently diverging.
|
||||
#
|
||||
# The ADR-0020 constants below are duplicated the same way and carry the same
|
||||
# warning: skill-audit/scripts/validate.sh holds a second copy of
|
||||
# DESC_SUGGEST_CHARS / DESC_MAX_CHARS / BODY_SUGGEST_WORDS / BODY_MAX_WORDS,
|
||||
# and agent-audit/scripts/validate.sh holds a third copy of the two
|
||||
# description constants (agents take the description gates and, per ADR-0020,
|
||||
# deliberately take NO body word gate). If they drift, this audit reports a
|
||||
# skill ready to ship that the commit hook then rejects.
|
||||
MAX_LINES=500
|
||||
MAX_WORDS=2770
|
||||
|
||||
# ADR-0020 context-budget gates. SUGGESTION does not fail; FAIL does.
|
||||
DESC_SUGGEST_CHARS=250
|
||||
DESC_MAX_CHARS=400
|
||||
BODY_SUGGEST_WORDS=600
|
||||
BODY_MAX_WORDS=900
|
||||
|
||||
FAIL=0
|
||||
|
||||
for f in "$@"; do
|
||||
@@ -68,4 +105,293 @@ for f in "$@"; do
|
||||
fi
|
||||
done
|
||||
|
||||
if ! command -v python3 > /dev/null 2>&1; then
|
||||
echo "ERROR: python3 is required for the ADR-0020 description/body/boundary-target gates but was not found on PATH." >&2
|
||||
echo " Why: skipping them would be a vacuous pass — the hook would go green having checked only the spec ceilings." >&2
|
||||
echo " Fix: install python3 (pre-commit itself is a Python application, so it is almost certainly already present)." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# REPO_ROOT is passed in so the boundary-target resolver has one guaranteed
|
||||
# place to look for the authoring source (plugins/*/.apm/skills/), independent
|
||||
# of the cwd pre-commit happens to invoke this hook from.
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
|
||||
if ! python3 -u - \
|
||||
"$REPO_ROOT" "$DESC_SUGGEST_CHARS" "$DESC_MAX_CHARS" \
|
||||
"$BODY_SUGGEST_WORDS" "$BODY_MAX_WORDS" "$MAX_WORDS" "$@" <<'PYTHON'
|
||||
import glob
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
|
||||
repo_root = sys.argv[1]
|
||||
DESC_SUGGEST_CHARS = int(sys.argv[2])
|
||||
DESC_MAX_CHARS = int(sys.argv[3])
|
||||
BODY_SUGGEST_WORDS = int(sys.argv[4])
|
||||
BODY_MAX_WORDS = int(sys.argv[5])
|
||||
MAX_WORDS = int(sys.argv[6])
|
||||
files = sys.argv[7:]
|
||||
|
||||
failed = False
|
||||
|
||||
|
||||
def error(msg):
|
||||
global failed
|
||||
failed = True
|
||||
print("ERROR: %s" % msg, file=sys.stderr)
|
||||
|
||||
|
||||
def suggest(msg):
|
||||
# stdout, not stderr, and never touches the exit code. The hook is declared
|
||||
# `verbose: true` in .pre-commit-config.yaml so this actually reaches a
|
||||
# human — pre-commit prints nothing at all for a passing hook otherwise,
|
||||
# which is the exact way ADR-0013 records Vale warnings going invisible.
|
||||
print("SUGGESTION: %s" % msg)
|
||||
|
||||
|
||||
# --- Frontmatter / description / body ------------------------------------
|
||||
# The description VALUE must be measured after YAML folding is resolved: this
|
||||
# corpus writes most descriptions as `>`-folded block scalars, so the raw lines
|
||||
# carry indentation and newlines that are not part of the value. Parse, don't
|
||||
# regex the raw text.
|
||||
|
||||
def normalize(value):
|
||||
return re.sub(r'\s+', ' ', value).strip()
|
||||
|
||||
|
||||
def fold_description_fallback(fm_text):
|
||||
"""Resolve `description:` without PyYAML.
|
||||
|
||||
Not a YAML parser — it recognises exactly the shapes this corpus uses:
|
||||
an inline scalar (optionally quoted, optionally continued on following
|
||||
indented lines) and a `>`/`|` block scalar with optional indentation and
|
||||
chomping indicators.
|
||||
"""
|
||||
lines = fm_text.splitlines()
|
||||
for i, line in enumerate(lines):
|
||||
m = re.match(r'^description:[ \t]*(.*)$', line)
|
||||
if not m:
|
||||
continue
|
||||
head = m.group(1).strip()
|
||||
block = bool(re.match(r'^[>|][0-9]*[-+]?$|^[>|][-+]?[0-9]*$', head))
|
||||
parts = [] if block else [head]
|
||||
for nxt in lines[i + 1:]:
|
||||
if not nxt.strip():
|
||||
parts.append('')
|
||||
continue
|
||||
if not re.match(r'^[ \t]', nxt):
|
||||
break
|
||||
parts.append(nxt.strip())
|
||||
value = ' '.join(parts)
|
||||
if not block:
|
||||
value = value.strip()
|
||||
if len(value) >= 2 and value[0] == value[-1] and value[0] in '"\'':
|
||||
value = value[1:-1]
|
||||
return value
|
||||
return ''
|
||||
|
||||
|
||||
def extract_description(fm_text):
|
||||
try:
|
||||
import yaml
|
||||
data = yaml.safe_load(fm_text)
|
||||
if isinstance(data, dict):
|
||||
value = data.get('description')
|
||||
if isinstance(value, str):
|
||||
return normalize(value)
|
||||
if value is not None:
|
||||
return normalize(str(value))
|
||||
return ''
|
||||
except Exception:
|
||||
pass
|
||||
return normalize(fold_description_fallback(fm_text))
|
||||
|
||||
|
||||
# --- Boundary-target resolution ------------------------------------------
|
||||
# A description's boundary clause names another skill (or an agent — agents are
|
||||
# legitimate routing targets: git-workflow routes to git-orchestrate). Resolve
|
||||
# every named target against the AUTHORING SOURCE, plugins/*/.apm/skills/ and
|
||||
# plugins/*/.apm/agents/, so the check works offline and before `apm install`
|
||||
# has deployed anything into .claude/skills/.
|
||||
|
||||
def known_targets(start_dir):
|
||||
names = set()
|
||||
|
||||
# Sibling skills/agents: covers a cache-installed plugin and a deployed
|
||||
# .claude/skills/ tree, neither of which has a plugins/ directory above it.
|
||||
parent = os.path.dirname(os.path.abspath(start_dir))
|
||||
if os.path.basename(parent) == 'skills' and os.path.isdir(parent):
|
||||
for entry in os.listdir(parent):
|
||||
if os.path.isdir(os.path.join(parent, entry)):
|
||||
names.add(entry)
|
||||
agents_dir = os.path.join(os.path.dirname(parent), 'agents')
|
||||
if os.path.isdir(agents_dir):
|
||||
for entry in os.listdir(agents_dir):
|
||||
if entry.endswith('.agent.md'):
|
||||
names.add(entry[:-len('.agent.md')])
|
||||
elif entry.endswith('.md'):
|
||||
names.add(entry[:-len('.md')])
|
||||
|
||||
# Walk up looking for a monorepo root (plugins/*/.apm/) or a plugin root
|
||||
# (.apm/). Capped so a pathological path can't turn this into a filesystem
|
||||
# crawl; ten levels covers every real layout by a wide margin.
|
||||
current = os.path.abspath(start_dir)
|
||||
for _ in range(10):
|
||||
# Never glob the filesystem root: a stray /.apm/skills/ (a scaffolding
|
||||
# test's leftover, say) would otherwise become part of every skill's
|
||||
# resolution universe on that machine.
|
||||
if os.path.dirname(current) == current:
|
||||
break
|
||||
for pattern in ('plugins/*/.apm/skills/*/', '.apm/skills/*/'):
|
||||
for path in glob.glob(os.path.join(current, pattern)):
|
||||
names.add(os.path.basename(path.rstrip('/')))
|
||||
for pattern in ('plugins/*/.apm/agents/*.agent.md', '.apm/agents/*.agent.md'):
|
||||
for path in glob.glob(os.path.join(current, pattern)):
|
||||
names.add(os.path.basename(path)[:-len('.agent.md')])
|
||||
current = os.path.dirname(current)
|
||||
return names
|
||||
|
||||
|
||||
NAME_ANY = r"[a-z0-9]+(?:-[a-z0-9]+)*"
|
||||
NAME_HYPH = r"[a-z0-9]+(?:-[a-z0-9]+)+"
|
||||
ROUTE_VERB = (r"(?:use|uses|using|run|runs|invoke|invokes|invoking|try|see"
|
||||
r"|that'?s|compose|composes|call|calls)")
|
||||
# A "marked" target is unambiguous on its own: backticked (`git-commits`) or
|
||||
# slash-command form (/skill-improve). A bare target is just a hyphenated word
|
||||
# and is only read as a routing target inside a boundary sentence — otherwise
|
||||
# "run pre-commit hooks" would be reported as a dangling route to `pre-commit`.
|
||||
MARKED_TARGET = r"(?:`/?(%s)`|(?<![\w./*-])/(%s)\b)" % (NAME_ANY, NAME_ANY)
|
||||
ANY_TARGET = r"(?:%s|(%s)\b)" % (MARKED_TARGET, NAME_HYPH)
|
||||
ROUTE_MARKED = re.compile(r"\b%s\s+(?:the\s+|an?\s+)?%s" % (ROUTE_VERB, MARKED_TARGET), re.I)
|
||||
ROUTE_ANY = re.compile(r"\b%s\s+(?:the\s+|an?\s+)?%s" % (ROUTE_VERB, ANY_TARGET), re.I)
|
||||
CONT_MARKED = re.compile(r"\s*(?:or|and|/|,)\s*%s" % MARKED_TARGET)
|
||||
CONT_ANY = re.compile(r"\s*(?:or|and|/|,)\s*%s" % ANY_TARGET)
|
||||
ARROW_MARKED = re.compile(r"(?:->|→)\s*%s" % MARKED_TARGET)
|
||||
# ADR-0020's compressed boundary form, `Not <thing> -> <skill-name>.`, with a
|
||||
# bare target. Anchored on "not" so diagnose's process arrow chain
|
||||
# ("fix -> regression-test") is not mistaken for a route.
|
||||
ARROW_BOUNDARY = re.compile(r"\bnot\b[^.;]*?(?:->|→)\s*(%s)\b" % NAME_HYPH, re.I)
|
||||
BACKTICK = re.compile(r"`(%s)`" % NAME_HYPH)
|
||||
BOUNDARY_MARKER = re.compile(r"\b(?:do\s+not|instead|rather\s+than|not\s+for)\b", re.I)
|
||||
SENTENCE_SPLIT = re.compile(u'(?<=[.!?])\\s+(?=[A-Z"“(])')
|
||||
|
||||
# Tool names (Read/Write/Edit) are excluded by construction: the name pattern is
|
||||
# lowercase-only. MCP tool names (issue_write, pull_request_write) are excluded
|
||||
# by construction too: the pattern admits no underscores. File names are
|
||||
# excluded because the pattern admits no dots or slashes inside the name.
|
||||
|
||||
|
||||
def _first(groups):
|
||||
for g in groups:
|
||||
if g:
|
||||
return g
|
||||
return None
|
||||
|
||||
|
||||
def _scan(text, route_re, cont_re, out):
|
||||
for m in route_re.finditer(text):
|
||||
name = _first(m.groups())
|
||||
if not name:
|
||||
continue
|
||||
out.append(name)
|
||||
# "use git-history or git-branches instead" / "use gitea-issues /
|
||||
# gitea-prs" — keep consuming conjoined targets after the first.
|
||||
pos = m.end()
|
||||
while True:
|
||||
cm = cont_re.match(text, pos)
|
||||
if not cm:
|
||||
break
|
||||
nxt = _first(cm.groups())
|
||||
if nxt:
|
||||
out.append(nxt)
|
||||
pos = cm.end()
|
||||
|
||||
|
||||
def boundary_targets(desc):
|
||||
out = []
|
||||
for sentence in SENTENCE_SPLIT.split(desc):
|
||||
boundary = bool(BOUNDARY_MARKER.search(sentence))
|
||||
_scan(sentence,
|
||||
ROUTE_ANY if boundary else ROUTE_MARKED,
|
||||
CONT_ANY if boundary else CONT_MARKED,
|
||||
out)
|
||||
for m in ARROW_MARKED.finditer(sentence):
|
||||
name = _first(m.groups())
|
||||
if name:
|
||||
out.append(name)
|
||||
for m in ARROW_BOUNDARY.finditer(sentence):
|
||||
out.append(m.group(1))
|
||||
out.extend(BACKTICK.findall(sentence))
|
||||
return sorted(set(out))
|
||||
|
||||
|
||||
# --- Per-file checks ------------------------------------------------------
|
||||
for path in files:
|
||||
if not os.path.isfile(path):
|
||||
continue
|
||||
with open(path) as fh:
|
||||
content = fh.read()
|
||||
|
||||
fm_match = re.match(r'^---\n(.*?)\n---', content, re.DOTALL)
|
||||
if not fm_match:
|
||||
# A missing frontmatter block is skill-frontmatter's / validate.sh's
|
||||
# finding, not this hook's — nothing to measure, so say nothing.
|
||||
continue
|
||||
|
||||
desc = extract_description(fm_match.group(1))
|
||||
body = content[fm_match.end():]
|
||||
|
||||
if desc:
|
||||
dlen = len(desc)
|
||||
if dlen > DESC_MAX_CHARS:
|
||||
error("%s: description is %d characters, exceeding the %d-character ceiling "
|
||||
"(ADR-0020). It is preloaded into every session. Keep a trigger clause, "
|
||||
"at most one capability clause, and a boundary clause; move capability "
|
||||
"enumeration, output-format detail, composition notes and implementation "
|
||||
"detail to the body or README.md."
|
||||
% (path, dlen, DESC_MAX_CHARS))
|
||||
elif dlen > DESC_SUGGEST_CHARS:
|
||||
suggest("%s: description is %d characters, over the %d-character target "
|
||||
"(ADR-0020, hard fail at %d)."
|
||||
% (path, dlen, DESC_SUGGEST_CHARS, DESC_MAX_CHARS))
|
||||
|
||||
body_words = len(body.split())
|
||||
if body_words > BODY_MAX_WORDS:
|
||||
error("%s: body is %d words, exceeding the %d-word ceiling (ADR-0020). This counts "
|
||||
"the body ONLY — it is a separate measurement from the %d-word whole-file "
|
||||
"spec ceiling above. Move lookup tables, spec restatements, output schemas, "
|
||||
"templates and rationale prose to references/ behind an explicit "
|
||||
"\"If X, read references/file.md\" trigger."
|
||||
% (path, body_words, BODY_MAX_WORDS, MAX_WORDS))
|
||||
elif body_words > BODY_SUGGEST_WORDS:
|
||||
suggest("%s: body is %d words, over the %d-word target (ADR-0020, hard fail at %d)."
|
||||
% (path, body_words, BODY_SUGGEST_WORDS, BODY_MAX_WORDS))
|
||||
|
||||
targets = boundary_targets(desc)
|
||||
if targets:
|
||||
# Two roots, unioned: the file's own directory (which finds siblings in
|
||||
# a cache install and walks up to a monorepo root in a checkout), and
|
||||
# this script's own repo root (which is authoritative when pre-commit
|
||||
# hands over a relative path from an unrelated cwd).
|
||||
known = known_targets(os.path.dirname(os.path.abspath(path)))
|
||||
known |= known_targets(repo_root)
|
||||
# An empty universe means the resolver found no authoring source at all
|
||||
# (a SKILL.md audited outside any plugin tree). Reporting every target
|
||||
# as dangling there would be noise, not a finding.
|
||||
if known:
|
||||
for target in targets:
|
||||
if target not in known:
|
||||
error("%s: description routes to '%s', which does not resolve to a skill "
|
||||
"under plugins/*/.apm/skills/ or an agent under "
|
||||
"plugins/*/.apm/agents/ (ADR-0020). A boundary clause that names a "
|
||||
"non-existent target sends the router nowhere."
|
||||
% (path, target))
|
||||
|
||||
sys.exit(1 if failed else 0)
|
||||
PYTHON
|
||||
then
|
||||
FAIL=1
|
||||
fi
|
||||
|
||||
exit $FAIL
|
||||
|
||||
Reference in New Issue
Block a user