Files
holocron/scripts/skill-size-check.sh
Defame1297 4a5c3c0cff feat(kyberforge): enforce the ADR-0020 context contract for skills and agents
Skill name+description pairs are preloaded into every session, costing
~6,200 tokens across 39 skills before any skill is invoked. The authoring
rules mandated that growth: skill-author:104 and description-quality.md:21
both required padding, while skill-author:102 (the deflating rule) had no
FAIL condition behind it.

Gates (blocking, no baseline file):
- description 250 chars SUGGESTION / 400 FAIL, measured on the folded
  YAML value
- body-only 600 words SUGGESTION / 900 FAIL, independent of the unchanged
  whole-file 2770-word / 500-line spec backstop
- every boundary-clause routing target must resolve to a real skill or
  agent; catches skill-improve, neuledge-context and gitea-labels
- agents take the description gates but deliberately no body gate; a test
  pins that absence

Vale: DescriptionOpener widened to ^This\b, new CompositionNote rule
banning architecture notes from descriptions. 10 hits, 0 false positives.

Kyberforge's own four skills retrofitted: descriptions 3,364 -> 938 chars
(-72%), bodies 8,306 -> 2,487 words (-70%), all via the apm-workflow
dispatch pattern. Fixes the skill-improve dangling route and the
agent-author misroute to manual review.

Also fixes a pre-existing false positive where any line-initial 'read '
was flagged as interactive input, which had already caused two scripts to
be rewritten around it.

Refs: ADR-0020
2026-08-14 21:13:13 +00:00

398 lines
17 KiB
Bash
Executable File

#!/usr/bin/env bash
set -euo pipefail
# Two independent gate families live in this script. Do not conflate them.
#
# 1. SPEC CONFORMANCE (MAX_LINES / MAX_WORDS, unchanged). Enforces
# agentskills.io's skill-authoring.md guidance: keep SKILL.md within 500
# lines and roughly 5,000 tokens, so the full body doesn't crowd out
# conversation history and other active skills once loaded into context.
# MAX_WORDS counts the WHOLE FILE, frontmatter included.
#
# 2. CONTEXT BUDGET (ADR-0020: description chars, body-only words, resolvable
# boundary targets). A skill's name + description is preloaded into every
# agent's context every session whether or not the skill is ever invoked;
# the body is loaded only on invocation, and then competes with the caller's
# live conversation. Those are different costs with different ceilings, so
# they get their own numbers and their own measurements.
#
# The two families measure different things on purpose and neither replaces the
# other: 2,770 whole-file words is a conformance backstop, 900 body-only words
# is a quality gate, and a file can sit well inside one while failing the other.
#
# Vale can't express any of this (its checks operate on text patterns, not raw
# file size), so this is a plain script instead of a Vale rule.
#
# Both spec ceilings are inclusive: a file at exactly MAX_LINES or MAX_WORDS
# passes, and only one past it fails. That matches
# skill-audit/scripts/validate.sh, which has always used `line_count <= 500` as
# its pass condition — the two previously disagreed at exactly 500 lines, so a
# SKILL.md could pass its own audit and still be blocked by the commit hook.
# The ADR-0020 ceilings are inclusive the same way.
#
# Token counts aren't computed exactly here — word count (`wc -w`) is used as
# a proxy. Measured over this repo's 39 in-scope SKILL.md files, characters per
# word runs min 5.97 / median 6.79 / mean 6.77 / max 7.22. At the standard
# ~4-characters-per-token English approximation that is 1.49 / 1.70 / 1.69 /
# 1.81 tokens per word.
#
# MAX_WORDS=2770 is therefore calibrated to the corpus WORST case rather than
# its median: 2770 words at the densest observed 7.22 chars/word is ~20,000
# characters, or ~5,000 tokens at the 4-characters-per-token approximation. So
# what this gate guarantees is "under 5,000 tokens even for the densest prose
# the corpus has produced" — the earlier median-calibrated MAX_WORDS=2900 let
# such a file sit at exactly the ceiling and still spend ~5,240 tokens. A
# median-density file at 2770 words spends ~4,700 tokens, so typical prose
# gives up ~130 words of headroom to close that gap. The largest SKILL.md in
# the repo is 2,760 words whole-file (skill-author), twelve words under the
# ceiling — this is a gate two files have already grown into, not headroom.
#
# It is a one-sided proxy in the useful direction — nothing under the word
# ceiling is wildly over the token ceiling — but it is not exact BPE
# tokenization and does not replace one. Re-measure the corpus before treating
# any of these numbers as still current.
#
# python3 is required for the ADR-0020 half. That is not a new dependency in
# practice: pre-commit, which is how this script runs, is itself a Python
# application. PyYAML is used when importable and is genuinely optional — the
# fallback reader below recognises the frontmatter shapes this corpus uses.
# These constants are intentionally duplicated in
# skill-audit/scripts/validate.sh (Python) rather than shared from one file:
# this script is a standalone bash pre-commit hook, that one is an in-skill
# Python validator invoked in a different context (same rationale as
# vale-wrap.sh's per-plugin duplication — see its own header comment).
# tests/test-skill-size-check.sh asserts both files agree on these values, so
# drift between them fails CI rather than silently diverging.
#
# The ADR-0020 constants below are duplicated the same way and carry the same
# warning: skill-audit/scripts/validate.sh holds a second copy of
# DESC_SUGGEST_CHARS / DESC_MAX_CHARS / BODY_SUGGEST_WORDS / BODY_MAX_WORDS,
# and agent-audit/scripts/validate.sh holds a third copy of the two
# description constants (agents take the description gates and, per ADR-0020,
# deliberately take NO body word gate). If they drift, this audit reports a
# skill ready to ship that the commit hook then rejects.
MAX_LINES=500
MAX_WORDS=2770
# ADR-0020 context-budget gates. SUGGESTION does not fail; FAIL does.
DESC_SUGGEST_CHARS=250
DESC_MAX_CHARS=400
BODY_SUGGEST_WORDS=600
BODY_MAX_WORDS=900
FAIL=0
for f in "$@"; do
[[ -f "$f" ]] || continue
# Single awk pass computes both line count and word count, avoiding a
# second read of the file. NR counts the final line even without a
# trailing newline, matching Python's splitlines() semantics (used by
# skill-audit/scripts/validate.sh for its own line count) — `wc -l`
# undercounts by 1 in that case. Word count uses awk's default
# whitespace-splitting NF, matching `wc -w` semantics.
read -r lines words <<< "$(awk '{w += NF} END{print NR, w+0}' "$f")"
if (( lines > MAX_LINES )); then
echo "ERROR: $f has $lines lines, exceeding the $MAX_LINES-line ceiling (agentskills.io skill-authoring.md)" >&2
FAIL=1
fi
if (( words > MAX_WORDS )); then
echo "ERROR: $f has $words words (proxy for tokens), exceeding the $MAX_WORDS-word ceiling (~5,000 tokens, agentskills.io skill-authoring.md)" >&2
FAIL=1
fi
done
if ! command -v python3 > /dev/null 2>&1; then
echo "ERROR: python3 is required for the ADR-0020 description/body/boundary-target gates but was not found on PATH." >&2
echo " Why: skipping them would be a vacuous pass — the hook would go green having checked only the spec ceilings." >&2
echo " Fix: install python3 (pre-commit itself is a Python application, so it is almost certainly already present)." >&2
exit 1
fi
# REPO_ROOT is passed in so the boundary-target resolver has one guaranteed
# place to look for the authoring source (plugins/*/.apm/skills/), independent
# of the cwd pre-commit happens to invoke this hook from.
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
if ! python3 -u - \
"$REPO_ROOT" "$DESC_SUGGEST_CHARS" "$DESC_MAX_CHARS" \
"$BODY_SUGGEST_WORDS" "$BODY_MAX_WORDS" "$MAX_WORDS" "$@" <<'PYTHON'
import glob
import os
import re
import sys
repo_root = sys.argv[1]
DESC_SUGGEST_CHARS = int(sys.argv[2])
DESC_MAX_CHARS = int(sys.argv[3])
BODY_SUGGEST_WORDS = int(sys.argv[4])
BODY_MAX_WORDS = int(sys.argv[5])
MAX_WORDS = int(sys.argv[6])
files = sys.argv[7:]
failed = False
def error(msg):
global failed
failed = True
print("ERROR: %s" % msg, file=sys.stderr)
def suggest(msg):
# stdout, not stderr, and never touches the exit code. The hook is declared
# `verbose: true` in .pre-commit-config.yaml so this actually reaches a
# human — pre-commit prints nothing at all for a passing hook otherwise,
# which is the exact way ADR-0013 records Vale warnings going invisible.
print("SUGGESTION: %s" % msg)
# --- Frontmatter / description / body ------------------------------------
# The description VALUE must be measured after YAML folding is resolved: this
# corpus writes most descriptions as `>`-folded block scalars, so the raw lines
# carry indentation and newlines that are not part of the value. Parse, don't
# regex the raw text.
def normalize(value):
return re.sub(r'\s+', ' ', value).strip()
def fold_description_fallback(fm_text):
"""Resolve `description:` without PyYAML.
Not a YAML parser — it recognises exactly the shapes this corpus uses:
an inline scalar (optionally quoted, optionally continued on following
indented lines) and a `>`/`|` block scalar with optional indentation and
chomping indicators.
"""
lines = fm_text.splitlines()
for i, line in enumerate(lines):
m = re.match(r'^description:[ \t]*(.*)$', line)
if not m:
continue
head = m.group(1).strip()
block = bool(re.match(r'^[>|][0-9]*[-+]?$|^[>|][-+]?[0-9]*$', head))
parts = [] if block else [head]
for nxt in lines[i + 1:]:
if not nxt.strip():
parts.append('')
continue
if not re.match(r'^[ \t]', nxt):
break
parts.append(nxt.strip())
value = ' '.join(parts)
if not block:
value = value.strip()
if len(value) >= 2 and value[0] == value[-1] and value[0] in '"\'':
value = value[1:-1]
return value
return ''
def extract_description(fm_text):
try:
import yaml
data = yaml.safe_load(fm_text)
if isinstance(data, dict):
value = data.get('description')
if isinstance(value, str):
return normalize(value)
if value is not None:
return normalize(str(value))
return ''
except Exception:
pass
return normalize(fold_description_fallback(fm_text))
# --- Boundary-target resolution ------------------------------------------
# A description's boundary clause names another skill (or an agent — agents are
# legitimate routing targets: git-workflow routes to git-orchestrate). Resolve
# every named target against the AUTHORING SOURCE, plugins/*/.apm/skills/ and
# plugins/*/.apm/agents/, so the check works offline and before `apm install`
# has deployed anything into .claude/skills/.
def known_targets(start_dir):
names = set()
# Sibling skills/agents: covers a cache-installed plugin and a deployed
# .claude/skills/ tree, neither of which has a plugins/ directory above it.
parent = os.path.dirname(os.path.abspath(start_dir))
if os.path.basename(parent) == 'skills' and os.path.isdir(parent):
for entry in os.listdir(parent):
if os.path.isdir(os.path.join(parent, entry)):
names.add(entry)
agents_dir = os.path.join(os.path.dirname(parent), 'agents')
if os.path.isdir(agents_dir):
for entry in os.listdir(agents_dir):
if entry.endswith('.agent.md'):
names.add(entry[:-len('.agent.md')])
elif entry.endswith('.md'):
names.add(entry[:-len('.md')])
# Walk up looking for a monorepo root (plugins/*/.apm/) or a plugin root
# (.apm/). Capped so a pathological path can't turn this into a filesystem
# crawl; ten levels covers every real layout by a wide margin.
current = os.path.abspath(start_dir)
for _ in range(10):
# Never glob the filesystem root: a stray /.apm/skills/ (a scaffolding
# test's leftover, say) would otherwise become part of every skill's
# resolution universe on that machine.
if os.path.dirname(current) == current:
break
for pattern in ('plugins/*/.apm/skills/*/', '.apm/skills/*/'):
for path in glob.glob(os.path.join(current, pattern)):
names.add(os.path.basename(path.rstrip('/')))
for pattern in ('plugins/*/.apm/agents/*.agent.md', '.apm/agents/*.agent.md'):
for path in glob.glob(os.path.join(current, pattern)):
names.add(os.path.basename(path)[:-len('.agent.md')])
current = os.path.dirname(current)
return names
NAME_ANY = r"[a-z0-9]+(?:-[a-z0-9]+)*"
NAME_HYPH = r"[a-z0-9]+(?:-[a-z0-9]+)+"
ROUTE_VERB = (r"(?:use|uses|using|run|runs|invoke|invokes|invoking|try|see"
r"|that'?s|compose|composes|call|calls)")
# A "marked" target is unambiguous on its own: backticked (`git-commits`) or
# slash-command form (/skill-improve). A bare target is just a hyphenated word
# and is only read as a routing target inside a boundary sentence — otherwise
# "run pre-commit hooks" would be reported as a dangling route to `pre-commit`.
MARKED_TARGET = r"(?:`/?(%s)`|(?<![\w./*-])/(%s)\b)" % (NAME_ANY, NAME_ANY)
ANY_TARGET = r"(?:%s|(%s)\b)" % (MARKED_TARGET, NAME_HYPH)
ROUTE_MARKED = re.compile(r"\b%s\s+(?:the\s+|an?\s+)?%s" % (ROUTE_VERB, MARKED_TARGET), re.I)
ROUTE_ANY = re.compile(r"\b%s\s+(?:the\s+|an?\s+)?%s" % (ROUTE_VERB, ANY_TARGET), re.I)
CONT_MARKED = re.compile(r"\s*(?:or|and|/|,)\s*%s" % MARKED_TARGET)
CONT_ANY = re.compile(r"\s*(?:or|and|/|,)\s*%s" % ANY_TARGET)
ARROW_MARKED = re.compile(r"(?:->|→)\s*%s" % MARKED_TARGET)
# ADR-0020's compressed boundary form, `Not <thing> -> <skill-name>.`, with a
# bare target. Anchored on "not" so diagnose's process arrow chain
# ("fix -> regression-test") is not mistaken for a route.
ARROW_BOUNDARY = re.compile(r"\bnot\b[^.;]*?(?:->|→)\s*(%s)\b" % NAME_HYPH, re.I)
BACKTICK = re.compile(r"`(%s)`" % NAME_HYPH)
BOUNDARY_MARKER = re.compile(r"\b(?:do\s+not|instead|rather\s+than|not\s+for)\b", re.I)
SENTENCE_SPLIT = re.compile(u'(?<=[.!?])\\s+(?=[A-Z"“(])')
# Tool names (Read/Write/Edit) are excluded by construction: the name pattern is
# lowercase-only. MCP tool names (issue_write, pull_request_write) are excluded
# by construction too: the pattern admits no underscores. File names are
# excluded because the pattern admits no dots or slashes inside the name.
def _first(groups):
for g in groups:
if g:
return g
return None
def _scan(text, route_re, cont_re, out):
for m in route_re.finditer(text):
name = _first(m.groups())
if not name:
continue
out.append(name)
# "use git-history or git-branches instead" / "use gitea-issues /
# gitea-prs" — keep consuming conjoined targets after the first.
pos = m.end()
while True:
cm = cont_re.match(text, pos)
if not cm:
break
nxt = _first(cm.groups())
if nxt:
out.append(nxt)
pos = cm.end()
def boundary_targets(desc):
out = []
for sentence in SENTENCE_SPLIT.split(desc):
boundary = bool(BOUNDARY_MARKER.search(sentence))
_scan(sentence,
ROUTE_ANY if boundary else ROUTE_MARKED,
CONT_ANY if boundary else CONT_MARKED,
out)
for m in ARROW_MARKED.finditer(sentence):
name = _first(m.groups())
if name:
out.append(name)
for m in ARROW_BOUNDARY.finditer(sentence):
out.append(m.group(1))
out.extend(BACKTICK.findall(sentence))
return sorted(set(out))
# --- Per-file checks ------------------------------------------------------
for path in files:
if not os.path.isfile(path):
continue
with open(path) as fh:
content = fh.read()
fm_match = re.match(r'^---\n(.*?)\n---', content, re.DOTALL)
if not fm_match:
# A missing frontmatter block is skill-frontmatter's / validate.sh's
# finding, not this hook's — nothing to measure, so say nothing.
continue
desc = extract_description(fm_match.group(1))
body = content[fm_match.end():]
if desc:
dlen = len(desc)
if dlen > DESC_MAX_CHARS:
error("%s: description is %d characters, exceeding the %d-character ceiling "
"(ADR-0020). It is preloaded into every session. Keep a trigger clause, "
"at most one capability clause, and a boundary clause; move capability "
"enumeration, output-format detail, composition notes and implementation "
"detail to the body or README.md."
% (path, dlen, DESC_MAX_CHARS))
elif dlen > DESC_SUGGEST_CHARS:
suggest("%s: description is %d characters, over the %d-character target "
"(ADR-0020, hard fail at %d)."
% (path, dlen, DESC_SUGGEST_CHARS, DESC_MAX_CHARS))
body_words = len(body.split())
if body_words > BODY_MAX_WORDS:
error("%s: body is %d words, exceeding the %d-word ceiling (ADR-0020). This counts "
"the body ONLY — it is a separate measurement from the %d-word whole-file "
"spec ceiling above. Move lookup tables, spec restatements, output schemas, "
"templates and rationale prose to references/ behind an explicit "
"\"If X, read references/file.md\" trigger."
% (path, body_words, BODY_MAX_WORDS, MAX_WORDS))
elif body_words > BODY_SUGGEST_WORDS:
suggest("%s: body is %d words, over the %d-word target (ADR-0020, hard fail at %d)."
% (path, body_words, BODY_SUGGEST_WORDS, BODY_MAX_WORDS))
targets = boundary_targets(desc)
if targets:
# Two roots, unioned: the file's own directory (which finds siblings in
# a cache install and walks up to a monorepo root in a checkout), and
# this script's own repo root (which is authoritative when pre-commit
# hands over a relative path from an unrelated cwd).
known = known_targets(os.path.dirname(os.path.abspath(path)))
known |= known_targets(repo_root)
# An empty universe means the resolver found no authoring source at all
# (a SKILL.md audited outside any plugin tree). Reporting every target
# as dangling there would be noise, not a finding.
if known:
for target in targets:
if target not in known:
error("%s: description routes to '%s', which does not resolve to a skill "
"under plugins/*/.apm/skills/ or an agent under "
"plugins/*/.apm/agents/ (ADR-0020). A boundary clause that names a "
"non-existent target sends the router nowhere."
% (path, target))
sys.exit(1 if failed else 0)
PYTHON
then
FAIL=1
fi
exit $FAIL