Three divergences between what the audit skills claim and what the hooks enforce, each of which fails silently rather than loudly: - `skill-size-check.sh` blocked at 2900 words while `validate.sh` checked only the 500-line ceiling, so `/skill-audit` could report a skill ready to ship that the commit hook then rejected. `validate.sh` now checks the same pair on the same inclusive terms; the constants are duplicated with a comment naming the other file, because a plugin skill's scripts cannot read outside the plugin directory once installed to the cache. - Both audit skills' Step 1 passed `--config assets/vale/.vale.ini`, which is redundant (the wrapper self-locates its sibling config) and fragile: an agent that resolves the script path against the skill directory but not the config path gets E100, exit 2, which the surrounding fallback clause misreads as "vale unavailable" and downgrades to full LLM judgment with no signal. - The external-consumer test registered only the two Vale hooks, never the third shipped hook, so a lost executable bit would have broken every consumer while the local suite stayed green. Verified by mutation: `chmod 644` on the copied script now turns three passes into two failures. Also corrects the size hook's calibration comment, which claimed ~5.7-6.5 characters per word against a corpus whose measured median is 6.79 — the stated upper bound sat below the median, so the "calibrated with margin" claim was inverted for prose-dense files. MAX_WORDS is unchanged pending a decision; the comment is now explicit that the gate holds under 5,000 tokens for typical prose density, not for any file. Refs: #85
9.9 KiB
name, description, allowed-tools, metadata
| name | description | allowed-tools | metadata | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| skill-audit | Use when the user wants to review a skill they wrote, says "audit this skill", "check if my skill follows best practices", "review my SKILL.md", or wants to know if a skill is ready to ship — even if they don't use the word "audit". Also invoke proactively after directly hand-editing a skill's files outside skill-author — an unaudited hand-edit is the same risk as unreviewed code. Audits a skill directory against the agentskills.io specification — structural checks plus qualitative review of description quality, body discipline, patterns, formatting, file structure, scripts, and internal consistency, plus a provenance chain check. Produces a compact findings report (findings only, no PASS noise) with Why and Fix per finding, suitable for agent handoff to /skill-improve or human auditability. Do not use to fix application code bugs or perform general code review unrelated to skill quality. Do not use when the user wants improvements applied — use /skill-improve instead. | Bash Read |
|
Gotchas
- Do not output PASS/FAIL per check while auditing — gather findings internally and surface them only in the Step 4 report. Narrating each check as you go is the default failure mode here.
Step 1 — Structural validation
bash scripts/validate.sh <skill-dir>
bash scripts/validate-provenance.sh <skill-dir>
scripts/vale-wrap.sh <skill-dir>/SKILL.md
Note any structural FAILs — they will appear in the report as a ### Structure dimension. If the script cannot execute (python3 unavailable, Bash denied, or permission error), perform structural checks manually: name format, name matches directory, description length ≤1024 chars, SKILL.md ≤500 lines, no unfilled FILL IN: placeholders, scripts executable and free of interactive prompts.
Note any Provenance FAILs and INFO findings from validate-provenance.sh — they surface in the report as a ### Provenance dimension (separate from ### Structure). The script embeds full FAIL/INFO format with Why and Fix per finding; surface them verbatim.
vale-wrap.sh ships inside this skill's own scripts/ — resolve it relative to this skill's directory the same way scripts/validate.sh is resolved above, so the invocation works whether this skill is running from this repo or from an installed plugin cache. Pass no --config: handed none, the wrapper loads its own sibling assets/vale/.vale.ini, located from the script's path rather than from the cwd. Adding an explicit relative --config breaks exactly the case the self-location covers — a resolved script path plus an unresolved config path yields E100 Runtime error ... does not exist, exit 2, which the fallback below then misreads as "vale unavailable". It applies that config's Kyberforge style — a deterministic prefilter for a subset of the Description/Patterns/Body dimensions below, not a replacement for Step 3. Every Vale alert is a FAIL — all rules are graded error — so report each one citing its rule ID (e.g. Kyberforge.DescriptionOpener). Skip and fall back to Step 3 judgment if the vale binary is unavailable. If Vale reports 0 files scanned, treat the pass as NOT RUN — not as clean — and fall back to full Step 3 judgment for the dimensions it would have covered.
Step 2 — Read all skill files
Read every file in the skill directory: SKILL.md, README.md (if present), all files in scripts/, references/, assets/, and tests/. Skip binary files only. Do not skip text files — internal consistency checks require the full picture.
Step 3 — Qualitative audit
Work through each dimension internally. Collect findings only; report them in Step 4. Cite file and line number for every finding.
Description
Vale's Kyberforge.DescriptionOpener ("This skill..." openers) and Kyberforge.VagueWording (filler like "helps with", "utilize") alerts from Step 1 — both FAILs — cover imperative phrasing and known vague-wording filler directly; report them as findings without re-deriving by judgment. The rest is still a judgment call:
- Specificity beyond the filler blocklist: are capabilities stated precisely ("parses OpenAPI specs") or genuinely vaguely ("handles files")?
- Indirect triggers: does it cover cases where the user doesn't name the domain directly?
- Near-miss exclusions: are "Do not use when..." clauses present if a near-miss skill could steal activations?
- Length: under 1024 characters?
If a description finding is borderline or the distinction between PASS and FAIL is unclear, read references/description-quality.md.
Body discipline
For each sentence in the body, apply: "Would the agent get this wrong without this sentence?" Flag any that answer "no" as padding.
- Defaults not menus: every decision point gives one default + one escape hatch, not a list of options
- Why rationale: include/exclude rules explain why, not just what
- Control calibration: prescriptive for fragile or critical sequences (e.g. a script invocation where flag order or exact arguments must not change); flexible where multiple approaches are valid
Vale's Kyberforge.SentenceOpenerThereIs alert from Step 1 (FAIL — sentences starting with "There is"/"There are") covers pattern-matchable body-wide filler directly; report it as a finding without re-deriving by judgment.
If uncertain whether a sentence is padding or whether a control decision is correctly calibrated, read references/body-discipline.md.
Patterns
Check each pattern is appropriate and correctly formed:
- Gotchas: placed near the top; each entry is a specific fact that defies a reasonable assumption — not a general tip
- Prescriptive sequence: inner code fences escaped as
\``` when nested inside a markdown block - Checklists: used for multi-step workflows, not single steps
- Conditional references: specific trigger stated ("If X, read
references/file.md") — not a generic "see references/". Vale'sKyberforge.PaddingPhrasealert from Step 1 flags the generic phrasing directly; other malformed conditional-reference forms still require judgment. - Output templates: present when the agent must produce a specific format; absent otherwise
File structure
- Permitted directories:
scripts/,references/,assets/,tests/; flag any other unlisted directory as FAIL — the spec allows additional dirs but this skill permits only these four to keep skills focused scripts/contains only executable code agents can run; test files (.bats,*_test.*,test_*.sh) inscripts/are a FAIL — they belong intests/- No non-spec files at the skill root (e.g. META.md, extra config files outside permitted directories)
- Optional directories contain real content — not just unfilled placeholder READMEs
README.mdpresent and accurately describes the skill and its files- No cross-plugin path references in SKILL.md, scripts/, references/, or assets/ — paths using
../,../../, or absolute repo paths (e.g.plugins/<plugin>/skills/<other-skill>/) break when the plugin is installed to a cache; flag any found references/sources.mdis exempt from the cross-plugin path check —Research doc:fields are development-only provenance pointers, not runtime references; they intentionally reference paths outside the skill directory and are expected to be non-resolvable after plugin install;validate-provenance.shhandles this gracefully by silently skipping upstream checks when those paths don't resolvetests/is exempt from the cross-plugin path check — test files are dev-only and may reference repo-level test infrastructure (e.g. a sharedtests/test_helper/). This dependency must be declared intests/README.md; flag if tests exist buttests/README.mdis absent or does not document the dependency
Formatting
- Heading levels consistent: H2 for main sections, H3 for subsections
- Code blocks fenced with a language tag where applicable (
bash,markdown,python) - Consistent whitespace: blank line between sections, consistent list indentation
- No broken relative paths in file references
Scripts
- No interactive TTY prompts (
read,input(),readline) --helpexposed with concise usage- Data to stdout, diagnostics to stderr
- Idempotent ("create if not exists")
- Meaningful exit codes documented in
--help --dry-runpresent for destructive operations
Internal consistency
- SKILL.md steps match what scripts actually do
README.mdfile table lists every file that exists — no missing entries, no stale entries- Placeholder READMEs in
scripts/,references/,assets/consistent with what SKILL.md says about each directory
Step 4 — Report
Open with a coverage line listing every dimension checked:
Checked: structure · description · body-discipline · patterns · file-structure · formatting · scripts · internal-consistency · provenance
Then output only dimensions that have findings, grouped under H3 headings, FAILs before SUGGESTIONs within each dimension. Omit clean dimensions entirely — their absence confirms they passed.
For each finding:
FAIL/SUGGESTION <finding> — file:line
Why: <why this is a problem>
Fix: <exact change — quote before/after where applicable>
Close with a result block:
## Result
PASS
PASS (N suggestions)
PASS · P info
PASS (N suggestions) · P info
FAIL (N fails · M suggestions)
FAIL (N fails · M suggestions) · P info
Run /skill-improve to address findings.
INFO findings are observational — do not affect PASS/FAIL. Omit · P info when there are no INFO findings. Omit the /skill-improve line when there are no findings at all. Do not apply fixes — report and propose only.