feat: implement issue 0017 — factory/write-eval bootstrap skill

Adds write-eval, the first factory meta-skill. Produces eval.yaml test
files for skills following the two-section schema (trigger_tests +
output_tests) with provider-agnostic string assertions and show-plan-
then-merge-on-rerun behaviour. Hand-written bootstrap — subsequent
skills will use write-eval to produce their own evals.

Also tightens skill-implementation-workflow.md step 5b: per-section
options walk-through is now a named gate before writing, separate from
the synthesis grill. LESSONS.md entry added.

HITL behavioral test pending.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
2026-05-17 17:52:02 +00:00
parent c705a38809
commit 83715018eb
7 changed files with 286 additions and 18 deletions

View File

@@ -31,21 +31,48 @@ write-eval has no direct Pocock equivalent. Expect to synthesize from multiple u
## Acceptance criteria
- [ ] `.agents/skills/write-eval/SKILL.md` exists; `metadata.category: factory`; authoring standard met (frontmatter, role, when/when-not, required inputs, constraints, process, output format, failure handling)
- [ ] Trigger description matches index or deviation is documented in SKILL.md with justification
- [ ] `.agents/evals/factory/write-eval/eval.yaml` exists; hand-written; contains all 5 required test types
- [ ] `install.sh` deploys `write-eval` to `~/.agents/skills/` (confirm idempotent re-run)
- [x] `.agents/skills/write-eval/SKILL.md` exists; `metadata.category: factory`; authoring standard met (frontmatter, role, when/when-not, required inputs, constraints, process, output format, failure handling)
- [x] Trigger description matches index or deviation is documented in SKILL.md with justification
- [x] `.agents/evals/factory/write-eval/eval.yaml` exists; hand-written; contains all 5 required test types
- [x] `install.sh` deploys `write-eval` to `~/.agents/skills/` (confirm idempotent re-run)
- [ ] **HITL:** human runs fresh-session behavioral test: invoke "write evals for this skill" and verify correct eval.yaml structure is produced
- [ ] **HITL:** human reviews hand-written eval.yaml for correctness before committing
- [ ] Per-skill process followed: source discovery (sub-agent) → source review with licence/security check (sub-agent) → conflict check against constitution + factory principles (sub-agent) → synthesis grill → co-write iteratively
- [ ] Trigger description tested against explicit, implicit, and negative queries before body was written
- [ ] `when:` frontmatter field present
- [ ] `source:` field present only if upstream content adopted; absent if self-authored
- [ ] `references:` field present if external citations used; absent otherwise
- [ ] eval.yaml contains all 5 required test types: explicit trigger, implicit trigger, negative trigger, ≥2 deterministic output, ≥1 LLM-rubric quality
- [ ] Body ≤500 lines; XML tags used only if ≥3 logical sections and 500+ tokens
- [ ] `docs/spec/overview.md` updated to reflect `write-eval` deployed
- [x] Per-skill process followed: source discovery (sub-agent) → source review with licence/security check (sub-agent) → conflict check against constitution + factory principles (sub-agent) → synthesis grill → co-write iteratively
- [x] Trigger description tested against explicit, implicit, and negative queries before body was written
- [x] `when:` frontmatter field present
- [x] `source:` field present only if upstream content adopted; absent if self-authored
- [x] `references:` field present if external citations used; absent otherwise
- [x] eval.yaml contains all 5 required test types: explicit trigger, implicit trigger, negative trigger, ≥2 deterministic output, ≥1 LLM-rubric quality
- [x] Body ≤500 lines; XML tags used only if ≥3 logical sections and 500+ tokens
- [x] `docs/spec/overview.md` updated to reflect `write-eval` deployed
## Blocked by
- 0016 (grill defines the per-skill implementation workflow this issue must follow)
## Handoff
**Status:** complete — pending HITL behavioral test (acceptance criteria step 5)
**Files produced:**
- `.agents/skills/write-eval/SKILL.md`
- `.agents/evals/factory/write-eval/eval.yaml`
**Key decisions:**
- Two-section schema: `trigger_tests` (explicit/implicit/negative, `should_trigger: bool`) + `output_tests` (deterministic/llm-rubric, `type:` field). Sources: BMAD-METHOD `triggers.json` split + darkrishabh `types.ts`.
- Provider-agnostic string assertions — no tool-call assertions. Portable across runtimes.
- Show plan before writing; merge on re-run with conflict flagging (option B): NEW / IDENTICAL / CONFLICT classification; CONFLICT cases shown side-by-side, human resolves before write.
- Iteration loop (run evals → propose edits → apply) is out of scope — belongs to a future runner skill.
- `id` as string slug (not integer); `name` field as separate display label.
**Workflow fix recorded:**
- `docs/notes/skill-implementation-workflow.md` step 5b updated: per-section options walk-through is now a named gate before writing. Synthesis grill answers schema questions; step 5b covers how upstream content maps to each SKILL.md section — these are separate conversations.
- `LESSONS.md` entry added: "Synthesis grill and SKILL.md co-write are two separate conversations."
**Open threads:**
- HITL behavioral test: open a fresh Claude session, invoke "write evals for this skill" in this repo context, verify correct eval.yaml structure is produced at the right path with all 5 types.
- `write-eval`'s own eval.yaml is hand-written (bootstrap). Once `write-eval` is behaviorally verified, it can be used to regenerate its own eval — a useful dogfood test.
**Next session start:**
- Load: `CONTEXT.md`, `docs/notes/skill-implementation-workflow.md`, `docs/issues/0018-factory-write-skill.md`
- First action: Step 1 (source discovery) for `write-skill`