Records the upstream agentskills.io sources that informed skill-audit, continuing the research → docs → skill provenance chain. - New references/sources.md with 7 extracted sources attributed to skill files; agentskills-llms-txt demoted to discovery-only comment per skill-author precedent - source_keys frontmatter added to SKILL.md (5 slugs), references/body-discipline.md (agentskills-spec, agentskills-best-practices), and references/description-quality.md (agentskills-spec, agentskills-optimizing-descriptions) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
135 lines
6.9 KiB
Markdown
135 lines
6.9 KiB
Markdown
---
|
|
name: skill-audit
|
|
description: >
|
|
Use when the user wants to review a skill they wrote, says "audit this skill",
|
|
"check if my skill follows best practices", "review my SKILL.md", or wants to
|
|
know if a skill is ready to ship — even if they don't use the word "audit".
|
|
Audits a skill directory against the agentskills.io specification — structural
|
|
checks plus qualitative review of description quality, body discipline, formatting,
|
|
file structure, and internal consistency. Produces a compact findings report
|
|
(findings only, no PASS noise) with Why and Fix per finding, suitable for agent
|
|
handoff to /skill-improve or human auditability. Do not use to fix application
|
|
code bugs or perform general code review unrelated to skill quality.
|
|
Do not use when the user wants improvements applied — use /skill-improve instead.
|
|
allowed-tools: Bash Read
|
|
metadata:
|
|
category: factory
|
|
source_keys:
|
|
- agentskills-home
|
|
- agentskills-spec
|
|
- agentskills-best-practices
|
|
- agentskills-optimizing-descriptions
|
|
- agentskills-using-scripts
|
|
---
|
|
|
|
## Gotchas
|
|
|
|
- Do not output PASS/FAIL per check while auditing — gather findings internally and surface them only in the Step 4 report. Narrating each check as you go is the default failure mode here.
|
|
|
|
## Step 1 — Structural validation
|
|
|
|
```bash
|
|
bash scripts/validate.sh <skill-dir>
|
|
```
|
|
|
|
Note any structural FAILs — they will appear in the report as a `### Structure` dimension. If the script cannot execute (python3 unavailable, Bash denied, or permission error), perform structural checks manually: name format, name matches directory, description length ≤1024 chars, SKILL.md ≤500 lines, no unfilled `FILL IN:` placeholders, scripts executable and free of interactive prompts.
|
|
|
|
## Step 2 — Read all skill files
|
|
|
|
Read every file in the skill directory: `SKILL.md`, `README.md` (if present), all files in `scripts/`, `references/`, `assets/`, and `tests/`. Skip binary files only. Do not skip text files — internal consistency checks require the full picture.
|
|
|
|
## Step 3 — Qualitative audit
|
|
|
|
Work through each dimension internally. Collect findings only; report them in Step 4. Cite file and line number for every finding.
|
|
|
|
### Description
|
|
|
|
- **Imperative phrasing**: does it use "Use when..." not "This skill..."?
|
|
- **Specificity**: are capabilities stated precisely ("parses OpenAPI specs") or vaguely ("helps with APIs")?
|
|
- **Indirect triggers**: does it cover cases where the user doesn't name the domain directly?
|
|
- **Near-miss exclusions**: are "Do not use when..." clauses present if a near-miss skill could steal activations?
|
|
- **Length**: under 1024 characters?
|
|
|
|
If a description finding is borderline or the distinction between PASS and FAIL is unclear, read `references/description-quality.md`.
|
|
|
|
### Body discipline
|
|
|
|
For each sentence in the body, apply: *"Would the agent get this wrong without this sentence?"* Flag any that answer "no" as padding.
|
|
|
|
- **Defaults not menus**: every decision point gives one default + one escape hatch, not a list of options
|
|
- **Why rationale**: include/exclude rules explain why, not just what
|
|
- **Control calibration**: prescriptive for fragile or critical sequences (e.g. a script invocation where flag order or exact arguments must not change); flexible where multiple approaches are valid
|
|
|
|
If uncertain whether a sentence is padding or whether a control decision is correctly calibrated, read `references/body-discipline.md`.
|
|
|
|
### Patterns
|
|
|
|
Check each pattern is appropriate and correctly formed:
|
|
|
|
- **Gotchas**: placed near the top; each entry is a specific fact that defies a reasonable assumption — not a general tip
|
|
- **Prescriptive sequence**: inner code fences escaped as `\`\`\`` when nested inside a markdown block
|
|
- **Checklists**: used for multi-step workflows, not single steps
|
|
- **Conditional references**: specific trigger stated ("If X, read `references/file.md`") — not a generic "see references/"
|
|
- **Output templates**: present when the agent must produce a specific format; absent otherwise
|
|
|
|
### File structure
|
|
|
|
- Permitted directories: `scripts/`, `references/`, `assets/`, `tests/`; flag any other unlisted directory as FAIL — the spec allows additional dirs but this skill permits only these four to keep skills focused
|
|
- `scripts/` contains only executable code agents can run; test files (`.bats`, `*_test.*`, `test_*.sh`) in `scripts/` are a FAIL — they belong in `tests/`
|
|
- No non-spec files at the skill root (e.g. META.md, extra config files outside permitted directories)
|
|
- Optional directories contain real content — not just unfilled placeholder READMEs
|
|
- `README.md` present and accurately describes the skill and its files
|
|
- No cross-plugin path references in SKILL.md, scripts/, references/, or assets/ — paths using `../`, `../../`, or absolute repo paths (e.g. `plugins/<plugin>/skills/<other-skill>/`) break when the plugin is installed to a cache; flag any found
|
|
- `tests/` is exempt from the cross-plugin path check — test files are dev-only and may reference repo-level test infrastructure (e.g. a shared `tests/test_helper/`). This dependency must be declared in `tests/README.md`; flag if tests exist but `tests/README.md` is absent or does not document the dependency
|
|
|
|
### Formatting
|
|
|
|
- Heading levels consistent: H2 for main sections, H3 for subsections
|
|
- Code blocks fenced with a language tag where applicable (`bash`, `markdown`, `python`)
|
|
- Consistent whitespace: blank line between sections, consistent list indentation
|
|
- No broken relative paths in file references
|
|
|
|
### Scripts
|
|
|
|
- No interactive TTY prompts (`read`, `input()`, `readline`)
|
|
- `--help` exposed with concise usage
|
|
- Data to stdout, diagnostics to stderr
|
|
- Idempotent ("create if not exists")
|
|
- Meaningful exit codes documented in `--help`
|
|
- `--dry-run` present for destructive operations
|
|
|
|
### Internal consistency
|
|
|
|
- SKILL.md steps match what scripts actually do
|
|
- `README.md` file table lists every file that exists — no missing entries, no stale entries
|
|
- Placeholder READMEs in `scripts/`, `references/`, `assets/` consistent with what SKILL.md says about each directory
|
|
|
|
## Step 4 — Report
|
|
|
|
Open with a coverage line listing every dimension checked:
|
|
|
|
```text
|
|
Checked: structure · description · body-discipline · patterns · file-structure · formatting · scripts · internal-consistency
|
|
```
|
|
|
|
Then output only dimensions that have findings, grouped under H3 headings, FAILs before SUGGESTIONs within each dimension. Omit clean dimensions entirely — their absence confirms they passed.
|
|
|
|
For each finding:
|
|
|
|
```text
|
|
FAIL/SUGGESTION <finding> — file:line
|
|
Why: <why this is a problem>
|
|
Fix: <exact change — quote before/after where applicable>
|
|
```
|
|
|
|
Close with a result block:
|
|
|
|
```text
|
|
## Result
|
|
|
|
PASS / PASS (N suggestions) / FAIL (N fails · M suggestions)
|
|
Run /skill-improve to address findings.
|
|
```
|
|
|
|
Omit the `/skill-improve` line when there are no findings. Do not apply fixes — report and propose only.
|