Files
holocron/docs/issues/0017-factory-write-eval.md
Defame1297 d98d0dae18 chore: migrate legacy pre-commit hook to .pre-commit-config.yaml
Replaces shell script (.git/hooks/pre-commit.legacy) with ecosystem-managed pre-commit framework:
- gitleaks/gitleaks: secret scanning
- jumanjihouse/pre-commit-hooks: shellcheck wrapper
- pre-commit/pre-commit-hooks: JSON/YAML validation, end-of-file-fixer, trailing-whitespace
- local hooks: SKILL.md frontmatter validation

Uses pinned versions for reproducibility across environments. Includes auto-fixes from hook runs (formatting, trailing whitespace, JSON beautification).

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-06-27 18:54:37 +00:00

78 lines
5.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 0017 — factory/write-eval (bootstrap skill)
**Type:** HITL
**Parent PRD:** `docs/prd/chunk-3-skills-library.md`
## What to build
Build `write-eval` — the first factory meta-skill, bootstrapped with a hand-written eval for itself. Every subsequent skill in Chunk 3 gets its eval produced via this skill. This issue is the smallest unblocker: get `write-eval` and its own hand-crafted eval in place, then all later skill issues can use it.
**Trigger description** (from skills index): "Write evals for this skill, create eval.yaml for X, add tests for this skill"
**Key constraints:**
- Skill file (slash command): `.agents/skills/write-eval/SKILL.md` — flat per ADR-0009; `metadata.category: factory`
- Produces eval files at: `.agents/evals/<category>/<skill-name>/eval.yaml` — nested by category (not skills; no discovery constraint)
- Every eval must contain: ≥1 explicit trigger test, ≥1 implicit trigger test, ≥1 negative trigger test (adjacent task that must NOT activate), ≥2 deterministic output tests (schema/contains/regex), ≥1 LLM-rubric quality test
- For this first issue: write-eval's own eval is hand-crafted (write-eval cannot produce its own eval before it exists)
- Origin: new skill; `source:` field populated only if upstream content is adopted (determine during implementation)
Process: follow `docs/notes/skill-implementation-workflow.md`. Bootstrap exception: steps 1–3 (source discovery, source review, conflict check) still apply; SKILL.md and eval.yaml are hand-written rather than factory-produced.
## Implementation notes
Follow the per-skill workflow defined in `docs/notes/skill-implementation-workflow.md` (produced by issue 0016).
**Known upstream sources to review:**
- `mattpocock/skills` — check for any eval-related content in the current set; record SHAs for any adopted content
- `bmad-method/bmad-method` — check for QA/evaluation patterns relevant to skill testing
- agentskills.io open standard — check whether an eval format is defined at the standard level before designing one from scratch; the eval schema in the PRD (5 test types) is derived from the factory design doc and may benefit from cross-referencing the standard
write-eval has no direct Pocock equivalent. Expect to synthesize from multiple upstreams or author original.
## Acceptance criteria
- [x] `.agents/skills/write-eval/SKILL.md` exists; `metadata.category: factory`; authoring standard met (frontmatter, role, when/when-not, required inputs, constraints, process, output format, failure handling)
- [x] Trigger description matches index or deviation is documented in SKILL.md with justification
- [x] `.agents/evals/factory/write-eval/eval.yaml` exists; hand-written; contains all 5 required test types
- [x] `install.sh` deploys `write-eval` to `~/.agents/skills/` (confirm idempotent re-run)
- [x] **HITL (run HOTL):** subagent fresh-context behavioral test 2026-05-26 — invoked write-eval on caveman skill; correctly stopped on missing `metadata.category` before computing output path (failure handling PASS); after category supplied, produced complete eval with all 5 required test types; process followed correctly
- [x] **HITL (run HOTL):** eval.yaml content reviewed by subagent auditor; 5 test types confirmed present and correctly structured; two caveman SKILL.md defects surfaced (missing category field, "be brief" trigger too broad) — deferred to upgrade-skill in 0028
- [x] Per-skill process followed: source discovery (sub-agent) → source review with licence/security check (sub-agent) → conflict check against constitution + factory principles (sub-agent) → synthesis grill → co-write iteratively
- [x] Trigger description tested against explicit, implicit, and negative queries before body was written
- [x] `when:` frontmatter field present
- [x] `source:` field present only if upstream content adopted; absent if self-authored
- [x] `references:` field present if external citations used; absent otherwise
- [x] eval.yaml contains all 5 required test types: explicit trigger, implicit trigger, negative trigger, ≥2 deterministic output, ≥1 LLM-rubric quality
- [x] Body ≤500 lines; XML tags used only if ≥3 logical sections and 500+ tokens
- [x] `docs/spec/overview.md` updated to reflect `write-eval` deployed
## Blocked by
- 0016 (grill defines the per-skill implementation workflow this issue must follow)
## Handoff
**Status:** complete ✅
**Files produced:**
- `.agents/skills/write-eval/SKILL.md`
- `.agents/evals/factory/write-eval/eval.yaml`
**Key decisions:**
- Two-section schema: `trigger_tests` (explicit/implicit/negative, `should_trigger: bool`) + `output_tests` (deterministic/llm-rubric, `type:` field). Sources: BMAD-METHOD `triggers.json` split + darkrishabh `types.ts`.
- Provider-agnostic string assertions — no tool-call assertions. Portable across runtimes.
- Show plan before writing; merge on re-run with conflict flagging (option B): NEW / IDENTICAL / CONFLICT classification; CONFLICT cases shown side-by-side, human resolves before write.
- Iteration loop (run evals → propose edits → apply) is out of scope — belongs to a future runner skill.
- `id` as string slug (not integer); `name` field as separate display label.
**Workflow fix recorded:**
- `docs/notes/skill-implementation-workflow.md` step 5b updated: per-section options walk-through is now a named gate before writing. Synthesis grill answers schema questions; step 5b covers how upstream content maps to each SKILL.md section — these are separate conversations.
- `LESSONS.md` entry added: "Synthesis grill and SKILL.md co-write are two separate conversations."
**Open threads:**
- `write-eval`'s own eval.yaml is hand-written (bootstrap). Now that write-eval is verified, it can be used to regenerate its own eval as a dogfood test — deferred to 0028.
**Next session start:**
- Load: `CONTEXT.md`, `docs/notes/skill-implementation-workflow.md`, `docs/issues/0018-factory-write-skill.md`
- First action: Step 1 (source discovery) for `write-skill`