Files
holocron/docs/adr/0013-vale-harness-scope-and-rule-sources.md
Defame1297 44bde9e9c9 docs(adr): correct the records the branch left describing deleted things
Seven ADRs described code that no longer exists or behaviour the gates do
not have. Where the wrong text came from main it carries a dated
correction; where this branch introduced it, it is fixed in place, because
main never published it and there is no record to preserve.

Fixed in place, branch-introduced:

- ADR-0021's 2026-09-14 correction asserted apm audit --ci "was never a
  drift gate at all". It is one: it replays the install and diffs. The
  claim contradicted this branch's own AGENTS.md and gates.md.
- ADR-0015 said unconditionally that no pre-push hook needs the network.
  The guarantee holds only once apm install has populated apm_modules/.
- ADR-0014's 2026-09-16 correction said restoring .pre-commit-hooks.yaml
  would ship a hook that fails for every consumer, because their checkout
  has no lib-boundary-resolver.sh. pre-commit clones the whole hook repo
  and skill-size-check.sh resolves the library from BASH_SOURCE, so the
  hook would work.
- ADR-0019's "twelve hooks pass under unshare -rn" matched neither HEAD
  (8) nor main (14), and stated the offline guarantee unconditionally.

Corrected, inherited from main:

- ADR-0022 and ADR-0013 named skill-frontmatter's pre-commit hook as the
  enforcer of mandatory metadata.version. That hook was deleted on this
  branch; the check lives in skill-size-check.sh.
- ADR-0022 enumerated the tip rule's carve-outs as a closed list and
  described a single merge-base. The gate also exempts a tree-identical
  skill and intersects every base from merge-base --all, and emits a third
  failure form. 8cfd54f said the documented behaviour did not change; it
  did. The gate is correct and is unchanged -- the record was not.
- ADR-0020's Decision still routed description overflow to README.md, its
  ADR-0025 amendment pointed the mirrored constants at validate.sh, which
  holds none, and its Enforcement table still named the two deleted
  validate.sh paths.
- ADR-0015's Status claimed every plugin's plugin.json is pack output;
  none exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NwD8Egs5r4ndqeFLmhusX2
2026-09-20 18:38:00 +00:00

10 KiB

Vale audit prefilter expands into a plugin-content harness, scoped to prose-pattern rules only

Issue #84 wired Vale as a deterministic prefilter for skill-audit/agent-audit (merged into factory-audit's two flows by ADR-0025; read every mention of the pair below that way), scoped to exactly four pattern-matchable checks (imperative description opener, vague capability wording, generic reference-pointer padding, Copilot's dead Use proactively phrasing), documented only in CONTEXT.md's "Vale audit prefilter" section — never its own ADR — and explicitly excluding body discipline, near-miss exclusion strength, and control calibration as non-goals. This ADR records a deferred PR #85 review item to broaden that coverage, retroactively captures #84's own rationale (since it was never recorded as a decision in its own right), and layers the expansion on top without reversing or weakening the original four rules.

2026-08-17 amendment. The CONTEXT.md section named above no longer holds that documentation. CONTEXT.md was cut back to a glossary and the prefilter's mechanics — the two-copy style layout, vale-wrap.sh, the --config argv defect, the rule inventory, and the 0-files-means-NOT-RUN fallback — moved to docs/spec/gates.md. Read that file, not CONTEXT.md, for the harness itself; this ADR still owns the scope decision.

File scope stays the same. SKILL.md plus agent files (**/agents/*.md, **/*.agent.md) only — matching the existing prefilter's globs. Skill-level README.md files and plugin.json manifests are not added: README.md files are navigational, not spec-governed content, and plugin.json is JSON, not prose Vale can meaningfully lint.

Rule categories are prose-pattern-matchable only. Structural, schema, and security concerns stay out of this Vale-based harness because this repo already has dedicated tools for them: skill-frontmatter (required frontmatter fields), validate-marketplace (claude plugin validate --strict, schema), and gitleaks/detect-private-key (secrets). (ADR-0024 removed the companion validate-plugins gate along with the per-plugin manifests it checked, and the skill-frontmatter hook has since been removed as well — its required-field checks were folded into skill-size-check, and the enforcer is now scripts/skill-size-check.sh:324-335 under the skill-size-check hook at .pre-commit-config.yaml:237. The argument here is unaffected either way: a dedicated non-Vale tool still owns required frontmatter fields.) Duplicating those concerns as Vale rules would fight tools that already own them better.

Governance docs are excluded as a rule source. docs/research/governance_principles/CONTROLS.md and governance.md were investigated and found to contribute nothing minable: CONTROLS.md is org/CI-infrastructure controls (secret scanning, dependency/license scanning, agent permission scoping, audit logging, human approval gates, periodic reviews) — none of it is a prose pattern expressible as a Vale rule against SKILL.md/agent-file text, and what it does cover is either already handled elsewhere (gitleaks) or genuinely out of scope for a plugin-content prose harness (dependency/license scanning is a code-dependency concern, not skill authoring).

Spec-derived custom rules stay mostly as-is. Re-reading agentskills.io's optimizing-descriptions.md and skill-authoring.md, plus claude-code-plugins/agent-definition.md and github-copilot-plugins/agent-definition.md, found that the existing four Kyberforge rules already cover the pattern-matchable surface those specs describe. The remaining spec guidance — calibrating control vs. giving freedom, avoiding menus of options, coherent skill scope, moderate detail level — is semantic judgment, already skill-audit's job via LLM review (now factory-audit's skill flow, see ADR-0025), not new lintable rules. One confirmation surfaced: Claude Code's Use proactively phrasing is meaningful for .md agent files (it triggers auto-invocation), unlike Copilot's .agent.md files where it's dead phrasing — so KyberforgeCopilot/ProactivePhrase's existing .agent.md-only scope is correct and must not be extended to .md files.

write-good/alex are trialed, not adopted wholesale. These built-in/third-party Vale packages are tuned for general blog-style prose (passive voice, weasel words, wordy phrases) and are expected to be noisy against this repo's terse, imperative instruction-file corpus. Only individual rules proven low-noise against the existing corpus get cherry-picked into styles/Kyberforge; the packages are never referenced wholesale in BasedOnStyles.

A new non-Vale check closes a real gap. skill-authoring.md states SKILL.md should stay under 500 lines / 5,000 tokens — currently unenforced anywhere in this repo. This is a whole-file length ceiling, not a text pattern, so it isn't a Vale rule — it becomes a new deterministic script and pre-commit hook, sibling to the existing skill-frontmatter hook.

Rules land directly in styles/Kyberforge, enforcing immediately. No trial/report-only tier is introduced (see Considered Options). "Enforcing immediately" holds only because every rule in both styles is level: error: Vale's exit code keys on error-level alerts alone, so a warning- or suggestion-level rule prints an alert and still exits 0, and pre-commit suppresses output from hooks that pass — such a rule is invisible and blocks nothing. Every Vale alert is therefore a FAIL, in the audit skills and in the blocking pre-commit hook alike, with no ignorable tier; that matches every other gate in this repo (shellcheck, the test suite, conventional-pre-commit). The implementation pass finalizes the cherry-picked write-good/alex rules and any new spec-derived rule wording, runs the full set against the existing SKILL.md/agent-file corpus, fixes any resulting violations across that corpus, and lands the rule changes and the corpus fixes as one atomic commit — the same enforcement model as the original four rules, never a partial or opt-in state.

Considered options

Phased rollout via a separate trial style + config (rejected). A styles/KyberforgeTrial/ directory plus a parallel .vale.trial.ini (mirroring the root config's globs but with BasedOnStyles = Kyberforge, KyberforgeTrial) would let new rules be swept report-only via lint-runner/vale-run before promotion into the enforcing styles/Kyberforge + root .vale.ini. This was considered because BasedOnStyles = Kyberforge activates every rule file under that directory automatically — there's no partial/opt-in application within a style, so a rule dropped straight into styles/Kyberforge goes live in the blocking pre-commit hook immediately. Rejected in favor of finalizing rules directly and fixing violations via subagent before committing: simpler, no new trial-config machinery to build or maintain — at the cost of no standing report-only tier for future candidate rules. Note that the first implementation shipped graded severities (error/warning/suggestion) and thereby recreated the rejected option by accident: the five non-error rules never affected an exit code and never surfaced output through a passing pre-commit hook, so they were a report-only tier that reported to nobody. Flattening every rule to level: error is what actually implements this decision.

Consequences

  • styles/Kyberforge/ gained one new rule file, cherry-picked from write-good/alex as low-noise against this repo's corpus: SentenceOpenerThereIs.yml (22 hits across 273 held-out markdown files; both in-corpus hits were clean rewrites, needing no suppression).
  • A second candidate, VagueQualifier.yml, was cherry-picked and then dropped. Against the 41 skill/agent files it hit twice: one marginal real finding (prototype/SKILL.md, "very different" → "fundamentally different") and one false positive (caveman/SKILL.md, which quotes of course as an example of filler — a mention, not a use) that no rewrite could clear, forcing the repo's only Vale suppression comments. Of its 15 held-out hits, 9 were in docs/research/examples/ (out-of-scope upstream material) and the remaining 6 were the word "very" in two idioms in a single research doc, each already adjacent to the hard number carrying the fact. One marginal catch does not pay for a permanent suppression, so the rule is deleted and this ADR's "cherry-picked rules" is one rule, not two.
  • A new pre-commit hook, skill-size-check (scripts/skill-size-check.sh), enforces the 500-line/5,000-token SKILL.md ceiling, sibling to skill-frontmatter. Both halves of that ceiling are blocking gates, not just the line count: MAX_LINES=500, and MAX_WORDS=2770 as a word-count proxy for the 5,000-token limit (calibrated to the densest prose this repo measured, 1.81 tokens per word, so a worst-case SKILL.md at the ceiling still lands under 5,000 tokens — wc -w is not BPE tokenization). Either one exceeded fails the hook. Both are inclusive: a file at exactly 500 lines or exactly 2,770 words passes, and only one past a ceiling fails. skill-audit/scripts/validate.sh (now factory-audit/scripts/validate.sh, see ADR-0025) enforces the same pair on the same inclusive terms, so the audit and the commit hook cannot disagree about whether a given SKILL.md is over size.
  • styles/KyberforgeTrial/ and .vale.trial.ini were deliberately not created — noted here so a future reader doesn't wonder if a trial tier was forgotten.
  • The styles-portability question — whether styles/ and .vale.ini should move into plugins/lint/ so the prefilter also works for repos that install kyberforge@holocron as an external plugin, rather than living at this repo's root — was deliberately deferred, not fixed, in this pass. This repo-root placement remains intentional: this ADR's "File scope stays the same" framing is specific to Kyberforge's own authoring conventions in this repo, not a generic lint-plugin feature. Portability is a known limitation, tracked for a separate future session, not silently forgotten.

What this ADR's implementation pass did: synced and trialed write-good/alex against the existing SKILL.md/agent-file corpus, cherry-picked the one low-noise rule above into styles/Kyberforge, wrote scripts/skill-size-check.sh and its pre-commit hook, fixed the resulting corpus violations, and landed the rule changes and corpus fixes as one atomic commit — matching the enforcement model described above (no partial or opt-in state), with every rule at level: error so that model is real rather than nominal.