- script: new-skill.sh now exits 0 when target already exists (idempotent retry-safe) instead of exit 1; --help updated to reflect narrowed error cases - test: updated bats test to assert success and "nothing to do" output - body: removed speculative "Extract the skill from a real task" advice (human-targeted, not agent-actionable) - formatting: converted H4 headings in Step 2 to bold text (H2/H3 two-tier model) - provenance: removed orphan agentskills-llms-txt entry from references/sources.md; added discovery-only comment Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
263 lines
12 KiB
Markdown
263 lines
12 KiB
Markdown
---
|
||
name: skill-author
|
||
description: >
|
||
Use when the user wants to create a new skill from scratch ("write a skill
|
||
for X", "build a skill that does Y", "create a SKILL.md for Z"), or improve
|
||
an existing one ("improve this skill", "fix based on feedback", "apply these
|
||
audit findings", "update based on grill output"). Also use when the user
|
||
provides inline feedback about a skill's behavior and wants it applied, or
|
||
when a grill session, eval run, or audit has produced findings the user wants
|
||
acted on — even if they don't say "improve" explicitly. Authors and refines
|
||
skills following the agentskills.io specification. Do not use for read-only review — use
|
||
/skill-audit instead. Do not use to author agent definition files.
|
||
allowed-tools: Bash Read Write Edit
|
||
metadata:
|
||
category: factory
|
||
source_keys:
|
||
- agentskills-home
|
||
- agentskills-spec
|
||
- agentskills-best-practices
|
||
- agentskills-optimizing-descriptions
|
||
- agentskills-evaluating-skills
|
||
- agentskills-using-scripts
|
||
- agentskills-quickstart
|
||
---
|
||
|
||
## Gotchas
|
||
|
||
- Patching per symptom is the default failure mode. Three eval failures may all trace to one missing instruction — always identify the root cause before editing.
|
||
- Do not create new scripts unless a signal explicitly calls for it. Writing scripts from scratch requires transcript analysis that is out of scope here; flag the opportunity as a suggestion instead.
|
||
|
||
## Route
|
||
|
||
Determine which flow to follow before touching the filesystem:
|
||
|
||
- **No skill directory at the target path** → follow **Creating a new skill**
|
||
- **Directory exists + at least one improvement signal present** → follow **Improving an existing skill**
|
||
- **Directory exists + no signals present** → ask: "No improvement signals found. Did you mean to create a new skill, or do you have feedback to apply?"
|
||
|
||
Signals include: grill session output, `/skill-audit` findings (PASS/FAIL punch list), inline user feedback, session context describing what went wrong.
|
||
|
||
## Creating a new skill
|
||
|
||
### Prerequisites
|
||
|
||
Run `/grill-me` on the skill's design and research the target domain first.
|
||
Share those outputs in this conversation: grill context, research docs, examples, constraints.
|
||
|
||
Design for one coherent user intent — skills too narrow force multiple loads per task; too broad are hard to activate precisely.
|
||
|
||
**Before touching the filesystem, verify you have:**
|
||
- [ ] A clear purpose — what specific task will this skill handle?
|
||
- [ ] Trigger scenarios — when should an agent activate it, including indirect cases?
|
||
- [ ] Skill name (kebab-case) and destination path
|
||
|
||
If any are missing, stop and ask the user before proceeding.
|
||
|
||
**Requires `/skill-audit`** — used in Step 6 for final validation. Both skills ship in the kyberforge plugin and are co-installed. If `/skill-audit` is unavailable, stop and ask the user to install the kyberforge plugin before continuing.
|
||
|
||
### Step 1 — Scaffold
|
||
|
||
Run the copy script with the skill name and destination directory:
|
||
|
||
```bash
|
||
bash scripts/new-skill.sh <skill-name> <destination-dir>
|
||
```
|
||
|
||
Examples:
|
||
```bash
|
||
bash scripts/new-skill.sh my-tool ~/.agents/skills/
|
||
bash scripts/new-skill.sh data-analyzer plugins/myplugin/skills/
|
||
```
|
||
|
||
This creates `<destination-dir>/<skill-name>/` with annotated templates ready to fill in.
|
||
|
||
If the destination is inside a plugin directory (path contains a `plugin.json`), read `references/deployment-modes.md` before adding any file references to SKILL.md.
|
||
|
||
### Step 2 — Fill in SKILL.md
|
||
|
||
Open `<destination-dir>/<skill-name>/SKILL.md`. Replace every `FILL IN:` placeholder.
|
||
|
||
**Frontmatter**
|
||
|
||
**`name`** — already set by the scaffold script. Must exactly match the directory name. Format: 1–64 characters, lowercase letters/numbers/hyphens only, no leading, trailing, or consecutive hyphens (`--`).
|
||
|
||
**`description`** — carries the entire triggering burden. Rules:
|
||
- Imperative: "Use when..." not "This skill..."
|
||
- Focus on user intent, not implementation — describe what the user is trying to achieve, not the skill's internal mechanics
|
||
- Specific about capabilities ("parses and validates OpenAPI specs", not "helps with APIs")
|
||
- Include indirect triggers: "even if the user doesn't mention X explicitly"
|
||
- Add "Do not use when..." only if a near-miss skill exists that could steal activations
|
||
- Hard limit: 1024 characters — count before finalizing
|
||
|
||
**Optional fields** — uncomment and fill in or remove entirely:
|
||
- `license` — include when distributing the skill externally
|
||
- `compatibility` — include if the skill requires specific tools, runtimes, or network access (max 500 characters)
|
||
- `metadata` — key-value map; use `author`, `version`, `category`
|
||
- `allowed-tools` — space-separated pre-approved tools; reduces permission prompts (experimental — support varies by client)
|
||
|
||
**Body — include only what the agent lacks**
|
||
|
||
Rename the placeholder section heading to one that fits the skill's structure — `## Step 1`, `## Workflow`, `## Instructions`, etc.
|
||
|
||
Ask of every sentence: "Would the agent get this wrong without it?" Cut anything that answers "no."
|
||
|
||
**Include:**
|
||
- Non-obvious sequences or ordering constraints — the agent may skip or reorder steps without this
|
||
- Domain conventions the agent cannot infer from general knowledge — this is the core value a skill adds
|
||
- One default per decision point, plus one escape hatch — never a menu; menus cause the agent to pause or pick arbitrarily
|
||
- Gotchas — facts that defy reasonable assumptions; the agent will get these wrong every time without them
|
||
|
||
**Exclude:**
|
||
- Concepts the agent already knows (what JSON is, how HTTP works) — adds tokens without changing behavior
|
||
- Exhaustive option lists — pick a default; the agent doesn't benefit from choosing
|
||
- Steps the agent handles independently — over-specifying leads agents to follow unproductive paths
|
||
- Restatements of the description — it's already in context; repeating it wastes the token budget
|
||
|
||
**Patterns**
|
||
|
||
**Gotchas** — highest value; place near the top:
|
||
````markdown
|
||
## Gotchas
|
||
- <Fact that defies a reasonable assumption>
|
||
- <Non-obvious naming discrepancy or hidden constraint>
|
||
````
|
||
|
||
**Default with escape hatch** (not a menu):
|
||
````markdown
|
||
Use <X> for <task>. For <edge case>, use <Y> instead.
|
||
````
|
||
|
||
**Prescriptive sequence** (when order is critical or fragile):
|
||
````markdown
|
||
Run exactly:
|
||
```bash
|
||
<command>
|
||
```
|
||
Do not modify flags.
|
||
````
|
||
|
||
**Checklist** (multi-step workflows):
|
||
````markdown
|
||
- [ ] Step 1: ...
|
||
- [ ] Step 2: ...
|
||
````
|
||
|
||
**Conditional reference** (progressive disclosure — load only when needed):
|
||
````markdown
|
||
If <condition>, read `references/<file>.md`.
|
||
````
|
||
|
||
**Output format template** (when the skill produces structured output):
|
||
````markdown
|
||
Output format:
|
||
```
|
||
<field>: <value>
|
||
<field>: <value>
|
||
```
|
||
````
|
||
For longer templates, place in `assets/<name>.md` and reference conditionally.
|
||
|
||
**Size budget**
|
||
|
||
Keep `SKILL.md` under 500 lines; 5,000 tokens is the recommended body budget. When approaching the limit:
|
||
- Move reference material to `references/<topic>.md` and load it conditionally
|
||
- Bundle repeated executable logic into `scripts/` rather than reinventing each run
|
||
|
||
### Step 3 — Add scripts (if needed)
|
||
|
||
Place executable scripts in `scripts/`. Critical rule: **no interactive prompts** — agents run non-interactive; blocking on TTY input hangs indefinitely. Accept all input via flags, env vars, or stdin.
|
||
|
||
Read `references/scripts.md` before writing any script — it covers the full contract: structured output, pinned versions, self-contained deps, idempotency, exit codes, dry-run, error messages, and output size limits.
|
||
|
||
If no scripts are needed, delete `scripts/README.md` and the `scripts/` directory.
|
||
|
||
### Step 4 — Add references, assets, and tests (if needed)
|
||
|
||
**`references/`** — additional documentation loaded on demand. One topic per file.
|
||
Reference conditionally from SKILL.md: `If <condition>, read references/<file>.md`.
|
||
Keep reference chains one level deep — a reference file that references another reference file is rarely loaded correctly.
|
||
|
||
**`assets/`** — static resources: templates, schemas, lookup tables.
|
||
Reference by relative path from SKILL.md.
|
||
|
||
**`tests/`** — test files for scripts in `scripts/`. Use when scripts are complex
|
||
enough to break silently. Test infrastructure (`.bats`, `*_test.*`) belongs here,
|
||
not in `scripts/`. See `tests/README.md` for setup instructions.
|
||
|
||
If not needed, delete the placeholder READMEs and their directories.
|
||
|
||
### Step 5 — Populate or delete `references/sources.md`
|
||
|
||
If a research `sources.md` is present in the conversation context:
|
||
|
||
1. Read it and filter to entries with `` `extracted` `` status only.
|
||
2. For each entry, determine which skill files it contributed to (SKILL.md and any files in references/ that drew from it). Update `Contributing files` accordingly — list skill files, not research topic files.
|
||
3. Write the updated content to `references/sources.md`.
|
||
4. Add `source_keys` to the frontmatter of `SKILL.md` (under `metadata`) listing the slugs of sources that informed it.
|
||
5. For each file in `references/` that was informed by research sources, add `source_keys` frontmatter (same format as research topic files) listing the relevant slugs.
|
||
|
||
If no research `sources.md` is in context, delete `references/sources.md`.
|
||
|
||
### Step 6 — Validate and close
|
||
|
||
Run `/skill-audit` on `<destination-dir>/<skill-name>`.
|
||
|
||
All FAIL findings must be resolved before the skill is considered done.
|
||
|
||
## Improving an existing skill
|
||
|
||
### Step 1 — Verify inputs
|
||
|
||
Confirm the skill directory path exists and that at least one improvement signal is present in the conversation or a referenced file.
|
||
|
||
If the skill dir is missing, ask for it. If no signals are present, stop: "This skill applies existing signals to a skill. For a blind review without signals, use `/skill-audit` instead."
|
||
|
||
Signals can come from anywhere in the conversation or referenced files:
|
||
- Grill session output (most common predecessor in the factory sequence)
|
||
- `/skill-audit` findings (PASS/FAIL/SUGGESTION punch list)
|
||
- Human feedback (feedback.json, inline in conversation, PR or issue comments)
|
||
- Session context describing what went wrong
|
||
|
||
Also verify the `name` field in frontmatter matches the skill's directory name exactly.
|
||
|
||
### Step 2 — Gather and group signals
|
||
|
||
Read the current skill files (SKILL.md and any files in scripts/, references/, assets/, tests/). Then collect all signals from the conversation and any file paths the user has referenced.
|
||
|
||
Group signals by **root cause**, not symptom. Ask: "What single gap in the skill causes this cluster of failures?" One root cause → one fix. Do not make a separate edit for each symptom.
|
||
|
||
```text
|
||
Example:
|
||
- Session context: output format is wrong on every run
|
||
- Audit finding: no output template defined
|
||
- User feedback: "I always have to ask it to format the output"
|
||
→ Root cause: SKILL.md has no output format specification → one fix: add an output template
|
||
```
|
||
|
||
### Step 3 — Announce planned changes
|
||
|
||
Before editing, state:
|
||
- Which root causes were identified and what evidence supports each
|
||
- Which files will be changed and what will change in each
|
||
|
||
Then proceed — edits are reversible via git, no approval checkpoint needed.
|
||
|
||
### Step 4 — Apply changes
|
||
|
||
Edit any file in the skill directory that the signals point to: SKILL.md, scripts/, references/, assets/, tests/, README.md.
|
||
|
||
**Generalize, don't patch.** Find the underlying gap, not the specific example that failed. A fix scoped only to the test cases you've seen will overfit and perform worse on new inputs.
|
||
|
||
**Keep it lean.** Remove instructions that aren't pulling their weight. For every sentence you add, ask: "Would the agent get this wrong without it?" A shorter, focused skill consistently outperforms an exhaustive one.
|
||
|
||
**Explain the why.** Reasoning-based instructions outperform rigid directives. If you find yourself writing a rule in all caps (ALWAYS/NEVER), reframe it: explain why the behavior matters so the agent can apply judgment in edge cases.
|
||
|
||
If a signal points to a script or reference file, edit that file directly rather than adding a workaround in SKILL.md.
|
||
|
||
**On scripts**: Fix and edit existing scripts freely when signals point to them.
|
||
|
||
### Step 5 — Validate and close
|
||
|
||
Run `/skill-audit` on the skill directory. Resolve any FAIL findings before considering the improvement complete.
|