Files
holocron/docs/issues/0017-factory-write-eval.md
Defame1297 83715018eb feat: implement issue 0017 — factory/write-eval bootstrap skill
Adds write-eval, the first factory meta-skill. Produces eval.yaml test
files for skills following the two-section schema (trigger_tests +
output_tests) with provider-agnostic string assertions and show-plan-
then-merge-on-rerun behaviour. Hand-written bootstrap — subsequent
skills will use write-eval to produce their own evals.

Also tightens skill-implementation-workflow.md step 5b: per-section
options walk-through is now a named gate before writing, separate from
the synthesis grill. LESSONS.md entry added.

HITL behavioral test pending.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-17 17:52:02 +00:00

5.8 KiB
Raw Blame History

0017 — factory/write-eval (bootstrap skill)

Type: HITL
Parent PRD: docs/prd/chunk-3-skills-library.md

What to build

Build write-eval — the first factory meta-skill, bootstrapped with a hand-written eval for itself. Every subsequent skill in Chunk 3 gets its eval produced via this skill. This issue is the smallest unblocker: get write-eval and its own hand-crafted eval in place, then all later skill issues can use it.

Trigger description (from skills index): "Write evals for this skill, create eval.yaml for X, add tests for this skill"

Key constraints:

  • Skill file (slash command): .agents/skills/write-eval/SKILL.md — flat per ADR-0009; metadata.category: factory
  • Produces eval files at: .agents/evals/<category>/<skill-name>/eval.yaml — nested by category (not skills; no discovery constraint)
  • Every eval must contain: ≥1 explicit trigger test, ≥1 implicit trigger test, ≥1 negative trigger test (adjacent task that must NOT activate), ≥2 deterministic output tests (schema/contains/regex), ≥1 LLM-rubric quality test
  • For this first issue: write-eval's own eval is hand-crafted (write-eval cannot produce its own eval before it exists)
  • Origin: new skill; source: field populated only if upstream content is adopted (determine during implementation)

Process: follow docs/notes/skill-implementation-workflow.md. Bootstrap exception: steps 1–3 (source discovery, source review, conflict check) still apply; SKILL.md and eval.yaml are hand-written rather than factory-produced.

Implementation notes

Follow the per-skill workflow defined in docs/notes/skill-implementation-workflow.md (produced by issue 0016).

Known upstream sources to review:

  • mattpocock/skills — check for any eval-related content in the current set; record SHAs for any adopted content
  • bmad-method/bmad-method — check for QA/evaluation patterns relevant to skill testing
  • agentskills.io open standard — check whether an eval format is defined at the standard level before designing one from scratch; the eval schema in the PRD (5 test types) is derived from the factory design doc and may benefit from cross-referencing the standard

write-eval has no direct Pocock equivalent. Expect to synthesize from multiple upstreams or author original.

Acceptance criteria

  • .agents/skills/write-eval/SKILL.md exists; metadata.category: factory; authoring standard met (frontmatter, role, when/when-not, required inputs, constraints, process, output format, failure handling)
  • Trigger description matches index or deviation is documented in SKILL.md with justification
  • .agents/evals/factory/write-eval/eval.yaml exists; hand-written; contains all 5 required test types
  • install.sh deploys write-eval to ~/.agents/skills/ (confirm idempotent re-run)
  • HITL: human runs fresh-session behavioral test: invoke "write evals for this skill" and verify correct eval.yaml structure is produced
  • HITL: human reviews hand-written eval.yaml for correctness before committing
  • Per-skill process followed: source discovery (sub-agent) → source review with licence/security check (sub-agent) → conflict check against constitution + factory principles (sub-agent) → synthesis grill → co-write iteratively
  • Trigger description tested against explicit, implicit, and negative queries before body was written
  • when: frontmatter field present
  • source: field present only if upstream content adopted; absent if self-authored
  • references: field present if external citations used; absent otherwise
  • eval.yaml contains all 5 required test types: explicit trigger, implicit trigger, negative trigger, ≥2 deterministic output, ≥1 LLM-rubric quality
  • Body ≤500 lines; XML tags used only if ≥3 logical sections and 500+ tokens
  • docs/spec/overview.md updated to reflect write-eval deployed

Blocked by

  • 0016 (grill defines the per-skill implementation workflow this issue must follow)

Handoff

Status: complete — pending HITL behavioral test (acceptance criteria step 5)

Files produced:

  • .agents/skills/write-eval/SKILL.md
  • .agents/evals/factory/write-eval/eval.yaml

Key decisions:

  • Two-section schema: trigger_tests (explicit/implicit/negative, should_trigger: bool) + output_tests (deterministic/llm-rubric, type: field). Sources: BMAD-METHOD triggers.json split + darkrishabh types.ts.
  • Provider-agnostic string assertions — no tool-call assertions. Portable across runtimes.
  • Show plan before writing; merge on re-run with conflict flagging (option B): NEW / IDENTICAL / CONFLICT classification; CONFLICT cases shown side-by-side, human resolves before write.
  • Iteration loop (run evals → propose edits → apply) is out of scope — belongs to a future runner skill.
  • id as string slug (not integer); name field as separate display label.

Workflow fix recorded:

  • docs/notes/skill-implementation-workflow.md step 5b updated: per-section options walk-through is now a named gate before writing. Synthesis grill answers schema questions; step 5b covers how upstream content maps to each SKILL.md section — these are separate conversations.
  • LESSONS.md entry added: "Synthesis grill and SKILL.md co-write are two separate conversations."

Open threads:

  • HITL behavioral test: open a fresh Claude session, invoke "write evals for this skill" in this repo context, verify correct eval.yaml structure is produced at the right path with all 5 types.
  • write-eval's own eval.yaml is hand-written (bootstrap). Once write-eval is behaviorally verified, it can be used to regenerate its own eval — a useful dogfood test.

Next session start:

  • Load: CONTEXT.md, docs/notes/skill-implementation-workflow.md, docs/issues/0018-factory-write-skill.md
  • First action: Step 1 (source discovery) for write-skill