Adds write-eval, the first factory meta-skill. Produces eval.yaml test files for skills following the two-section schema (trigger_tests + output_tests) with provider-agnostic string assertions and show-plan- then-merge-on-rerun behaviour. Hand-written bootstrap — subsequent skills will use write-eval to produce their own evals. Also tightens skill-implementation-workflow.md step 5b: per-section options walk-through is now a named gate before writing, separate from the synthesis grill. LESSONS.md entry added. HITL behavioral test pending. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
5.8 KiB
0017 — factory/write-eval (bootstrap skill)
Type: HITL
Parent PRD: docs/prd/chunk-3-skills-library.md
What to build
Build write-eval — the first factory meta-skill, bootstrapped with a hand-written eval for itself. Every subsequent skill in Chunk 3 gets its eval produced via this skill. This issue is the smallest unblocker: get write-eval and its own hand-crafted eval in place, then all later skill issues can use it.
Trigger description (from skills index): "Write evals for this skill, create eval.yaml for X, add tests for this skill"
Key constraints:
- Skill file (slash command):
.agents/skills/write-eval/SKILL.md— flat per ADR-0009;metadata.category: factory - Produces eval files at:
.agents/evals/<category>/<skill-name>/eval.yaml— nested by category (not skills; no discovery constraint) - Every eval must contain: ≥1 explicit trigger test, ≥1 implicit trigger test, ≥1 negative trigger test (adjacent task that must NOT activate), ≥2 deterministic output tests (schema/contains/regex), ≥1 LLM-rubric quality test
- For this first issue: write-eval's own eval is hand-crafted (write-eval cannot produce its own eval before it exists)
- Origin: new skill;
source:field populated only if upstream content is adopted (determine during implementation)
Process: follow docs/notes/skill-implementation-workflow.md. Bootstrap exception: steps 1–3 (source discovery, source review, conflict check) still apply; SKILL.md and eval.yaml are hand-written rather than factory-produced.
Implementation notes
Follow the per-skill workflow defined in docs/notes/skill-implementation-workflow.md (produced by issue 0016).
Known upstream sources to review:
mattpocock/skills— check for any eval-related content in the current set; record SHAs for any adopted contentbmad-method/bmad-method— check for QA/evaluation patterns relevant to skill testing- agentskills.io open standard — check whether an eval format is defined at the standard level before designing one from scratch; the eval schema in the PRD (5 test types) is derived from the factory design doc and may benefit from cross-referencing the standard
write-eval has no direct Pocock equivalent. Expect to synthesize from multiple upstreams or author original.
Acceptance criteria
.agents/skills/write-eval/SKILL.mdexists;metadata.category: factory; authoring standard met (frontmatter, role, when/when-not, required inputs, constraints, process, output format, failure handling)- Trigger description matches index or deviation is documented in SKILL.md with justification
.agents/evals/factory/write-eval/eval.yamlexists; hand-written; contains all 5 required test typesinstall.shdeployswrite-evalto~/.agents/skills/(confirm idempotent re-run)- HITL: human runs fresh-session behavioral test: invoke "write evals for this skill" and verify correct eval.yaml structure is produced
- HITL: human reviews hand-written eval.yaml for correctness before committing
- Per-skill process followed: source discovery (sub-agent) → source review with licence/security check (sub-agent) → conflict check against constitution + factory principles (sub-agent) → synthesis grill → co-write iteratively
- Trigger description tested against explicit, implicit, and negative queries before body was written
when:frontmatter field presentsource:field present only if upstream content adopted; absent if self-authoredreferences:field present if external citations used; absent otherwise- eval.yaml contains all 5 required test types: explicit trigger, implicit trigger, negative trigger, ≥2 deterministic output, ≥1 LLM-rubric quality
- Body ≤500 lines; XML tags used only if ≥3 logical sections and 500+ tokens
docs/spec/overview.mdupdated to reflectwrite-evaldeployed
Blocked by
- 0016 (grill defines the per-skill implementation workflow this issue must follow)
Handoff
Status: complete — pending HITL behavioral test (acceptance criteria step 5)
Files produced:
.agents/skills/write-eval/SKILL.md.agents/evals/factory/write-eval/eval.yaml
Key decisions:
- Two-section schema:
trigger_tests(explicit/implicit/negative,should_trigger: bool) +output_tests(deterministic/llm-rubric,type:field). Sources: BMAD-METHODtriggers.jsonsplit + darkrishabhtypes.ts. - Provider-agnostic string assertions — no tool-call assertions. Portable across runtimes.
- Show plan before writing; merge on re-run with conflict flagging (option B): NEW / IDENTICAL / CONFLICT classification; CONFLICT cases shown side-by-side, human resolves before write.
- Iteration loop (run evals → propose edits → apply) is out of scope — belongs to a future runner skill.
idas string slug (not integer);namefield as separate display label.
Workflow fix recorded:
docs/notes/skill-implementation-workflow.mdstep 5b updated: per-section options walk-through is now a named gate before writing. Synthesis grill answers schema questions; step 5b covers how upstream content maps to each SKILL.md section — these are separate conversations.LESSONS.mdentry added: "Synthesis grill and SKILL.md co-write are two separate conversations."
Open threads:
- HITL behavioral test: open a fresh Claude session, invoke "write evals for this skill" in this repo context, verify correct eval.yaml structure is produced at the right path with all 5 types.
write-eval's own eval.yaml is hand-written (bootstrap). Oncewrite-evalis behaviorally verified, it can be used to regenerate its own eval — a useful dogfood test.
Next session start:
- Load:
CONTEXT.md,docs/notes/skill-implementation-workflow.md,docs/issues/0018-factory-write-skill.md - First action: Step 1 (source discovery) for
write-skill