Files
holocron/docs/issues/0017-factory-write-eval.md
Defame1297 9457efff36 docs: Chunk 3 issues 0015-0028 and doc updates
14 issues created covering AGENTS.md refactor prerequisite, skill
implementation workflow grill, bootstrap skills (write-eval, write-skill),
remaining factory skills, and one issue per skill category group through
to chunk closure. All HITL; acceptance criteria for 0017-0028 to be
refined after 0016 grill session.

Doc updates: CONTEXT.md PRD/issue scope clarified (HOW distribution
across architecture-review and issue design notes); ROADMAP.md housekeeping
updated with bootstrap order and issue range; spec/overview.md recent
changes entry added; PRD updated (skills-index: delete → update).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-17 15:18:36 +00:00

3.1 KiB

0017 — factory/write-eval (bootstrap skill)

Type: HITL
Parent PRD: docs/prd/chunk-3-skills-library.md

What to build

Build write-eval — the first factory meta-skill, bootstrapped with a hand-written eval for itself. Every subsequent skill in Chunk 3 gets its eval produced via this skill. This issue is the smallest unblocker: get write-eval and its own hand-crafted eval in place, then all later skill issues can use it.

Trigger description (from skills index): "Write evals for this skill, create eval.yaml for X, add tests for this skill"

Key constraints:

  • Skill file (slash command): .agents/skills/write-eval/SKILL.md — flat per ADR-0009; metadata.category: factory
  • Produces eval files at: .agents/evals/<category>/<skill-name>/eval.yaml — nested by category (not skills; no discovery constraint)
  • Every eval must contain: ≥1 explicit trigger test, ≥1 implicit trigger test, ≥1 negative trigger test (adjacent task that must NOT activate), ≥2 deterministic output tests (schema/contains/regex), ≥1 LLM-rubric quality test
  • For this first issue: write-eval's own eval is hand-crafted (write-eval cannot produce its own eval before it exists)
  • Origin: new skill; source: field populated only if upstream content is adopted (determine during implementation)

Process: upstream review → trigger description written and tested first → skill body → hand-write eval → behavioral test.

Implementation notes

Follow the per-skill workflow defined in docs/notes/skill-implementation-workflow.md (produced by issue 0016).

Known upstream sources to review:

  • mattpocock/skills — check for any eval-related content in the current set; record SHAs for any adopted content
  • bmad-method/bmad-method — check for QA/evaluation patterns relevant to skill testing
  • agentskills.io open standard — check whether an eval format is defined at the standard level before designing one from scratch; the eval schema in the PRD (5 test types) is derived from the factory design doc and may benefit from cross-referencing the standard

write-eval has no direct Pocock equivalent. Expect to synthesize from multiple upstreams or author original.

Acceptance criteria

  • .agents/skills/write-eval/SKILL.md exists; metadata.category: factory; authoring standard met (frontmatter, role, when/when-not, required inputs, constraints, process, output format, failure handling)
  • Trigger description matches index or deviation is documented in SKILL.md with justification
  • .agents/evals/factory/write-eval/eval.yaml exists; hand-written; contains all 5 required test types
  • install.sh deploys write-eval to ~/.agents/skills/ (confirm idempotent re-run)
  • HITL: human runs fresh-session behavioral test: invoke "write evals for this skill" and verify correct eval.yaml structure is produced
  • HITL: human reviews hand-written eval.yaml for correctness before committing
  • (Further criteria to be refined after issue 0016 grill session)

Blocked by

  • 0016 (grill defines the per-skill implementation workflow this issue must follow)