0017 — factory/write-eval (bootstrap skill) #28

Closed
opened 2026-06-28 17:10:17 +00:00 by Claude · 0 comments
Collaborator

Type: HITL
Parent PRD: docs/prd/chunk-3-skills-library.md

What to build

Build write-eval — the first factory meta-skill, bootstrapped with a hand-written eval for itself. Every subsequent skill in Chunk 3 gets its eval produced via this skill. This issue is the smallest unblocker: get write-eval and its own hand-crafted eval in place, then all later skill issues can use it.

Trigger description (from skills index): "Write evals for this skill, create eval.yaml for X, add tests for this skill"

Key constraints:

  • Skill file (slash command): .agents/skills/write-eval/SKILL.md — flat per ADR-0009; metadata.category: factory
  • Produces eval files at: .agents/evals/<category>/<skill-name>/eval.yaml — nested by category (not skills; no discovery constraint)
  • Every eval must contain: ≥1 explicit trigger test, ≥1 implicit trigger test, ≥1 negative trigger test (adjacent task that must NOT activate), ≥2 deterministic output tests (schema/contains/regex), ≥1 LLM-rubric quality test
  • For this first issue: write-eval's own eval is hand-crafted (write-eval cannot produce its own eval before it exists)
  • Origin: new skill; source: field populated only if upstream content is adopted (determine during implementation)

Process: follow docs/notes/skill-implementation-workflow.md. Bootstrap exception: steps 1–3 (source discovery, source review, conflict check) still apply; SKILL.md and eval.yaml are hand-written rather than factory-produced.

Implementation notes

Follow the per-skill workflow defined in docs/notes/skill-implementation-workflow.md (produced by issue 0016).

Known upstream sources to review:

  • mattpocock/skills — check for any eval-related content in the current set; record SHAs for any adopted content
  • bmad-method/bmad-method — check for QA/evaluation patterns relevant to skill testing
  • agentskills.io open standard — check whether an eval format is defined at the standard level before designing one from scratch; the eval schema in the PRD (5 test types) is derived from the factory design doc and may benefit from cross-referencing the standard

write-eval has no direct Pocock equivalent. Expect to synthesize from multiple upstreams or author original.

Acceptance criteria

  • .agents/skills/write-eval/SKILL.md exists; metadata.category: factory; authoring standard met (frontmatter, role, when/when-not, required inputs, constraints, process, output format, failure handling)
  • Trigger description matches index or deviation is documented in SKILL.md with justification
  • .agents/evals/factory/write-eval/eval.yaml exists; hand-written; contains all 5 required test types
  • install.sh deploys write-eval to ~/.agents/skills/ (confirm idempotent re-run)
  • HITL (run HOTL): subagent fresh-context behavioral test 2026-05-26 — invoked write-eval on caveman skill; correctly stopped on missing metadata.category before computing output path (failure handling PASS); after category supplied, produced complete eval with all 5 required test types; process followed correctly
  • HITL (run HOTL): eval.yaml content reviewed by subagent auditor; 5 test types confirmed present and correctly structured; two caveman SKILL.md defects surfaced (missing category field, "be brief" trigger too broad) — deferred to upgrade-skill in 0028
  • Per-skill process followed: source discovery (sub-agent) → source review with licence/security check (sub-agent) → conflict check against constitution + factory principles (sub-agent) → synthesis grill → co-write iteratively
  • Trigger description tested against explicit, implicit, and negative queries before body was written
  • when: frontmatter field present
  • source: field present only if upstream content adopted; absent if self-authored
  • references: field present if external citations used; absent otherwise
  • eval.yaml contains all 5 required test types: explicit trigger, implicit trigger, negative trigger, ≥2 deterministic output, ≥1 LLM-rubric quality
  • Body ≤500 lines; XML tags used only if ≥3 logical sections and 500+ tokens
  • docs/spec/overview.md updated to reflect write-eval deployed

Blocked by

  • 0016 (grill defines the per-skill implementation workflow this issue must follow)

Handoff

Status: complete

Files produced:

  • .agents/skills/write-eval/SKILL.md
  • .agents/evals/factory/write-eval/eval.yaml

Key decisions:

  • Two-section schema: trigger_tests (explicit/implicit/negative, should_trigger: bool) + output_tests (deterministic/llm-rubric, type: field). Sources: BMAD-METHOD triggers.json split + darkrishabh types.ts.
  • Provider-agnostic string assertions — no tool-call assertions. Portable across runtimes.
  • Show plan before writing; merge on re-run with conflict flagging (option B): NEW / IDENTICAL / CONFLICT classification; CONFLICT cases shown side-by-side, human resolves before write.
  • Iteration loop (run evals → propose edits → apply) is out of scope — belongs to a future runner skill.
  • id as string slug (not integer); name field as separate display label.

Workflow fix recorded:

  • docs/notes/skill-implementation-workflow.md step 5b updated: per-section options walk-through is now a named gate before writing. Synthesis grill answers schema questions; step 5b covers how upstream content maps to each SKILL.md section — these are separate conversations.
  • LESSONS.md entry added: "Synthesis grill and SKILL.md co-write are two separate conversations."

Open threads:

  • write-eval's own eval.yaml is hand-written (bootstrap). Now that write-eval is verified, it can be used to regenerate its own eval as a dogfood test — deferred to 0028.

Next session start:

  • Load: CONTEXT.md, docs/notes/skill-implementation-workflow.md, docs/issues/0018-factory-write-skill.md
  • First action: Step 1 (source discovery) for write-skill
**Type:** HITL **Parent PRD:** `docs/prd/chunk-3-skills-library.md` ## What to build Build `write-eval` — the first factory meta-skill, bootstrapped with a hand-written eval for itself. Every subsequent skill in Chunk 3 gets its eval produced via this skill. This issue is the smallest unblocker: get `write-eval` and its own hand-crafted eval in place, then all later skill issues can use it. **Trigger description** (from skills index): "Write evals for this skill, create eval.yaml for X, add tests for this skill" **Key constraints:** - Skill file (slash command): `.agents/skills/write-eval/SKILL.md` — flat per ADR-0009; `metadata.category: factory` - Produces eval files at: `.agents/evals/<category>/<skill-name>/eval.yaml` — nested by category (not skills; no discovery constraint) - Every eval must contain: ≥1 explicit trigger test, ≥1 implicit trigger test, ≥1 negative trigger test (adjacent task that must NOT activate), ≥2 deterministic output tests (schema/contains/regex), ≥1 LLM-rubric quality test - For this first issue: write-eval's own eval is hand-crafted (write-eval cannot produce its own eval before it exists) - Origin: new skill; `source:` field populated only if upstream content is adopted (determine during implementation) Process: follow `docs/notes/skill-implementation-workflow.md`. Bootstrap exception: steps 1–3 (source discovery, source review, conflict check) still apply; SKILL.md and eval.yaml are hand-written rather than factory-produced. ## Implementation notes Follow the per-skill workflow defined in `docs/notes/skill-implementation-workflow.md` (produced by issue 0016). **Known upstream sources to review:** - `mattpocock/skills` — check for any eval-related content in the current set; record SHAs for any adopted content - `bmad-method/bmad-method` — check for QA/evaluation patterns relevant to skill testing - agentskills.io open standard — check whether an eval format is defined at the standard level before designing one from scratch; the eval schema in the PRD (5 test types) is derived from the factory design doc and may benefit from cross-referencing the standard write-eval has no direct Pocock equivalent. Expect to synthesize from multiple upstreams or author original. ## Acceptance criteria - [x] `.agents/skills/write-eval/SKILL.md` exists; `metadata.category: factory`; authoring standard met (frontmatter, role, when/when-not, required inputs, constraints, process, output format, failure handling) - [x] Trigger description matches index or deviation is documented in SKILL.md with justification - [x] `.agents/evals/factory/write-eval/eval.yaml` exists; hand-written; contains all 5 required test types - [x] `install.sh` deploys `write-eval` to `~/.agents/skills/` (confirm idempotent re-run) - [x] **HITL (run HOTL):** subagent fresh-context behavioral test 2026-05-26 — invoked write-eval on caveman skill; correctly stopped on missing `metadata.category` before computing output path (failure handling PASS); after category supplied, produced complete eval with all 5 required test types; process followed correctly - [x] **HITL (run HOTL):** eval.yaml content reviewed by subagent auditor; 5 test types confirmed present and correctly structured; two caveman SKILL.md defects surfaced (missing category field, "be brief" trigger too broad) — deferred to upgrade-skill in 0028 - [x] Per-skill process followed: source discovery (sub-agent) → source review with licence/security check (sub-agent) → conflict check against constitution + factory principles (sub-agent) → synthesis grill → co-write iteratively - [x] Trigger description tested against explicit, implicit, and negative queries before body was written - [x] `when:` frontmatter field present - [x] `source:` field present only if upstream content adopted; absent if self-authored - [x] `references:` field present if external citations used; absent otherwise - [x] eval.yaml contains all 5 required test types: explicit trigger, implicit trigger, negative trigger, ≥2 deterministic output, ≥1 LLM-rubric quality - [x] Body ≤500 lines; XML tags used only if ≥3 logical sections and 500+ tokens - [x] `docs/spec/overview.md` updated to reflect `write-eval` deployed ## Blocked by - 0016 (grill defines the per-skill implementation workflow this issue must follow) ## Handoff **Status:** complete **Files produced:** - `.agents/skills/write-eval/SKILL.md` - `.agents/evals/factory/write-eval/eval.yaml` **Key decisions:** - Two-section schema: `trigger_tests` (explicit/implicit/negative, `should_trigger: bool`) + `output_tests` (deterministic/llm-rubric, `type:` field). Sources: BMAD-METHOD `triggers.json` split + darkrishabh `types.ts`. - Provider-agnostic string assertions — no tool-call assertions. Portable across runtimes. - Show plan before writing; merge on re-run with conflict flagging (option B): NEW / IDENTICAL / CONFLICT classification; CONFLICT cases shown side-by-side, human resolves before write. - Iteration loop (run evals → propose edits → apply) is out of scope — belongs to a future runner skill. - `id` as string slug (not integer); `name` field as separate display label. **Workflow fix recorded:** - `docs/notes/skill-implementation-workflow.md` step 5b updated: per-section options walk-through is now a named gate before writing. Synthesis grill answers schema questions; step 5b covers how upstream content maps to each SKILL.md section — these are separate conversations. - `LESSONS.md` entry added: "Synthesis grill and SKILL.md co-write are two separate conversations." **Open threads:** - `write-eval`'s own eval.yaml is hand-written (bootstrap). Now that write-eval is verified, it can be used to regenerate its own eval as a dogfood test — deferred to 0028. **Next session start:** - Load: `CONTEXT.md`, `docs/notes/skill-implementation-workflow.md`, `docs/issues/0018-factory-write-skill.md` - First action: Step 1 (source discovery) for `write-skill`
Claude added this to the Legacy / Triage milestone 2026-06-28 17:10:17 +00:00
Claude added the Kind/Feature
Priority
Medium
3
labels 2026-06-28 17:10:17 +00:00
Sign in to join this conversation.