# 0017 — factory/write-eval (bootstrap skill) **Type:** HITL **Parent PRD:** `docs/prd/chunk-3-skills-library.md` ## What to build Build `write-eval` — the first factory meta-skill, bootstrapped with a hand-written eval for itself. Every subsequent skill in Chunk 3 gets its eval produced via this skill. This issue is the smallest unblocker: get `write-eval` and its own hand-crafted eval in place, then all later skill issues can use it. **Trigger description** (from skills index): "Write evals for this skill, create eval.yaml for X, add tests for this skill" **Key constraints:** - Skill file (slash command): `.agents/skills/write-eval/SKILL.md` — flat per ADR-0009; `metadata.category: factory` - Produces eval files at: `.agents/evals///eval.yaml` — nested by category (not skills; no discovery constraint) - Every eval must contain: ≥1 explicit trigger test, ≥1 implicit trigger test, ≥1 negative trigger test (adjacent task that must NOT activate), ≥2 deterministic output tests (schema/contains/regex), ≥1 LLM-rubric quality test - For this first issue: write-eval's own eval is hand-crafted (write-eval cannot produce its own eval before it exists) - Origin: new skill; `source:` field populated only if upstream content is adopted (determine during implementation) Process: follow `docs/notes/skill-implementation-workflow.md`. Bootstrap exception: steps 1–3 (source discovery, source review, conflict check) still apply; SKILL.md and eval.yaml are hand-written rather than factory-produced. ## Implementation notes Follow the per-skill workflow defined in `docs/notes/skill-implementation-workflow.md` (produced by issue 0016). **Known upstream sources to review:** - `mattpocock/skills` — check for any eval-related content in the current set; record SHAs for any adopted content - `bmad-method/bmad-method` — check for QA/evaluation patterns relevant to skill testing - agentskills.io open standard — check whether an eval format is defined at the standard level before designing one from scratch; the eval schema in the PRD (5 test types) is derived from the factory design doc and may benefit from cross-referencing the standard write-eval has no direct Pocock equivalent. Expect to synthesize from multiple upstreams or author original. ## Acceptance criteria - [x] `.agents/skills/write-eval/SKILL.md` exists; `metadata.category: factory`; authoring standard met (frontmatter, role, when/when-not, required inputs, constraints, process, output format, failure handling) - [x] Trigger description matches index or deviation is documented in SKILL.md with justification - [x] `.agents/evals/factory/write-eval/eval.yaml` exists; hand-written; contains all 5 required test types - [x] `install.sh` deploys `write-eval` to `~/.agents/skills/` (confirm idempotent re-run) - [x] **HITL (run HOTL):** subagent fresh-context behavioral test 2026-05-26 — invoked write-eval on caveman skill; correctly stopped on missing `metadata.category` before computing output path (failure handling PASS); after category supplied, produced complete eval with all 5 required test types; process followed correctly - [x] **HITL (run HOTL):** eval.yaml content reviewed by subagent auditor; 5 test types confirmed present and correctly structured; two caveman SKILL.md defects surfaced (missing category field, "be brief" trigger too broad) — deferred to upgrade-skill in 0028 - [x] Per-skill process followed: source discovery (sub-agent) → source review with licence/security check (sub-agent) → conflict check against constitution + factory principles (sub-agent) → synthesis grill → co-write iteratively - [x] Trigger description tested against explicit, implicit, and negative queries before body was written - [x] `when:` frontmatter field present - [x] `source:` field present only if upstream content adopted; absent if self-authored - [x] `references:` field present if external citations used; absent otherwise - [x] eval.yaml contains all 5 required test types: explicit trigger, implicit trigger, negative trigger, ≥2 deterministic output, ≥1 LLM-rubric quality - [x] Body ≤500 lines; XML tags used only if ≥3 logical sections and 500+ tokens - [x] `docs/spec/overview.md` updated to reflect `write-eval` deployed ## Blocked by - 0016 (grill defines the per-skill implementation workflow this issue must follow) ## Handoff **Status:** complete ✅ **Files produced:** - `.agents/skills/write-eval/SKILL.md` - `.agents/evals/factory/write-eval/eval.yaml` **Key decisions:** - Two-section schema: `trigger_tests` (explicit/implicit/negative, `should_trigger: bool`) + `output_tests` (deterministic/llm-rubric, `type:` field). Sources: BMAD-METHOD `triggers.json` split + darkrishabh `types.ts`. - Provider-agnostic string assertions — no tool-call assertions. Portable across runtimes. - Show plan before writing; merge on re-run with conflict flagging (option B): NEW / IDENTICAL / CONFLICT classification; CONFLICT cases shown side-by-side, human resolves before write. - Iteration loop (run evals → propose edits → apply) is out of scope — belongs to a future runner skill. - `id` as string slug (not integer); `name` field as separate display label. **Workflow fix recorded:** - `docs/notes/skill-implementation-workflow.md` step 5b updated: per-section options walk-through is now a named gate before writing. Synthesis grill answers schema questions; step 5b covers how upstream content maps to each SKILL.md section — these are separate conversations. - `LESSONS.md` entry added: "Synthesis grill and SKILL.md co-write are two separate conversations." **Open threads:** - `write-eval`'s own eval.yaml is hand-written (bootstrap). Now that write-eval is verified, it can be used to regenerate its own eval as a dogfood test — deferred to 0028. **Next session start:** - Load: `CONTEXT.md`, `docs/notes/skill-implementation-workflow.md`, `docs/issues/0018-factory-write-skill.md` - First action: Step 1 (source discovery) for `write-skill`