Files
holocron/docs/issues/0017-factory-write-eval.md
Defame1297 c705a38809 docs: issue 0016 — skill implementation workflow grill
Produces docs/notes/skill-implementation-workflow.md with agreed conventions
for all Chunk 3 skill issues (0017–0028). Key decisions:

- Per-skill process: source discovery (sub-agent) → source review with
  licence/security check (sub-agent) → conflict check vs constitution +
  factory principles (sub-agent) → synthesis grill → co-write iteratively
- Bootstrap: write-eval (hand-written) → write-skill (hand-written) →
  write-docs (first factory-authored, phase 2 of 0018) → everything else
- Upstream review changed from per-chunk-start to per-skill
- `when:` and `references:` frontmatter fields added to authoring standard
- Sub-agent usage prescribed as named steps in the workflow
- HITL: human reviewed and approved conventions

Updates: PRD implementation decisions; issues 0016–0028 with specific
acceptance criteria; docs/spec/overview.md; ROADMAP Chunk 3 housekeeping note
(bootstrap order, cadence, acceptance criteria status); CONTEXT.md Source field
(per-skill cadence, references: companion field); LESSONS.md with three patterns
from the grill session.

Post-grill additions (same session): Step 6 (session handoff) added to the
workflow; handoff section appended to issue 0016; handoff checklist item added
to Chunk 3 closure issue (0028).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-17 17:06:35 +00:00

4.0 KiB
Raw Blame History

0017 — factory/write-eval (bootstrap skill)

Type: HITL
Parent PRD: docs/prd/chunk-3-skills-library.md

What to build

Build write-eval — the first factory meta-skill, bootstrapped with a hand-written eval for itself. Every subsequent skill in Chunk 3 gets its eval produced via this skill. This issue is the smallest unblocker: get write-eval and its own hand-crafted eval in place, then all later skill issues can use it.

Trigger description (from skills index): "Write evals for this skill, create eval.yaml for X, add tests for this skill"

Key constraints:

  • Skill file (slash command): .agents/skills/write-eval/SKILL.md — flat per ADR-0009; metadata.category: factory
  • Produces eval files at: .agents/evals/<category>/<skill-name>/eval.yaml — nested by category (not skills; no discovery constraint)
  • Every eval must contain: ≥1 explicit trigger test, ≥1 implicit trigger test, ≥1 negative trigger test (adjacent task that must NOT activate), ≥2 deterministic output tests (schema/contains/regex), ≥1 LLM-rubric quality test
  • For this first issue: write-eval's own eval is hand-crafted (write-eval cannot produce its own eval before it exists)
  • Origin: new skill; source: field populated only if upstream content is adopted (determine during implementation)

Process: follow docs/notes/skill-implementation-workflow.md. Bootstrap exception: steps 1–3 (source discovery, source review, conflict check) still apply; SKILL.md and eval.yaml are hand-written rather than factory-produced.

Implementation notes

Follow the per-skill workflow defined in docs/notes/skill-implementation-workflow.md (produced by issue 0016).

Known upstream sources to review:

  • mattpocock/skills — check for any eval-related content in the current set; record SHAs for any adopted content
  • bmad-method/bmad-method — check for QA/evaluation patterns relevant to skill testing
  • agentskills.io open standard — check whether an eval format is defined at the standard level before designing one from scratch; the eval schema in the PRD (5 test types) is derived from the factory design doc and may benefit from cross-referencing the standard

write-eval has no direct Pocock equivalent. Expect to synthesize from multiple upstreams or author original.

Acceptance criteria

  • .agents/skills/write-eval/SKILL.md exists; metadata.category: factory; authoring standard met (frontmatter, role, when/when-not, required inputs, constraints, process, output format, failure handling)
  • Trigger description matches index or deviation is documented in SKILL.md with justification
  • .agents/evals/factory/write-eval/eval.yaml exists; hand-written; contains all 5 required test types
  • install.sh deploys write-eval to ~/.agents/skills/ (confirm idempotent re-run)
  • HITL: human runs fresh-session behavioral test: invoke "write evals for this skill" and verify correct eval.yaml structure is produced
  • HITL: human reviews hand-written eval.yaml for correctness before committing
  • Per-skill process followed: source discovery (sub-agent) → source review with licence/security check (sub-agent) → conflict check against constitution + factory principles (sub-agent) → synthesis grill → co-write iteratively
  • Trigger description tested against explicit, implicit, and negative queries before body was written
  • when: frontmatter field present
  • source: field present only if upstream content adopted; absent if self-authored
  • references: field present if external citations used; absent otherwise
  • eval.yaml contains all 5 required test types: explicit trigger, implicit trigger, negative trigger, ≥2 deterministic output, ≥1 LLM-rubric quality
  • Body ≤500 lines; XML tags used only if ≥3 logical sections and 500+ tokens
  • docs/spec/overview.md updated to reflect write-eval deployed

Blocked by

  • 0016 (grill defines the per-skill implementation workflow this issue must follow)