Files
holocron/docs/issues/0022-test-skills.md
Defame1297 c705a38809 docs: issue 0016 — skill implementation workflow grill
Produces docs/notes/skill-implementation-workflow.md with agreed conventions
for all Chunk 3 skill issues (0017–0028). Key decisions:

- Per-skill process: source discovery (sub-agent) → source review with
  licence/security check (sub-agent) → conflict check vs constitution +
  factory principles (sub-agent) → synthesis grill → co-write iteratively
- Bootstrap: write-eval (hand-written) → write-skill (hand-written) →
  write-docs (first factory-authored, phase 2 of 0018) → everything else
- Upstream review changed from per-chunk-start to per-skill
- `when:` and `references:` frontmatter fields added to authoring standard
- Sub-agent usage prescribed as named steps in the workflow
- HITL: human reviewed and approved conventions

Updates: PRD implementation decisions; issues 0016–0028 with specific
acceptance criteria; docs/spec/overview.md; ROADMAP Chunk 3 housekeeping note
(bootstrap order, cadence, acceptance criteria status); CONTEXT.md Source field
(per-skill cadence, references: companion field); LESSONS.md with three patterns
from the grill session.

Post-grill additions (same session): Step 6 (session handoff) added to the
workflow; handoff section appended to issue 0016; handoff checklist item added
to Chunk 3 closure issue (0028).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-17 17:06:35 +00:00

2.8 KiB

0022 — Test skills: write-tests, generate-test-data, review-test-coverage

Type: HITL
Parent PRD: docs/prd/chunk-3-skills-library.md

What to build

The 3 test phase skills. All are new. Authored via write-skill (0018), evals via write-eval (0017).

Skills and trigger descriptions:

Flat name Trigger description
write-tests Write tests, generate test cases, add unit tests
generate-test-data Generate test data, create fixtures, sample data
review-test-coverage Review test coverage, find untested paths, coverage gaps

Key constraints per skill:

  • write-tests: derives tests from spec (EARS acceptance criteria), NOT from implementation; uses pytest for Python, Vitest/Jest for TypeScript
  • generate-test-data: produces structurally valid, semantically unusual data; flags PII risk before generating
  • review-test-coverage: reports coverage gaps against spec acceptance criteria, not line coverage percentages

Implementation notes

Follow the per-skill workflow defined in docs/notes/skill-implementation-workflow.md.

Known upstream sources to review:

  • mattpocock/skills — check for any test-phase skills in the current set
  • bmad-method/bmad-method — BMAD QA role patterns
  • Search agentskills.io and GitHub for open-source test generation skills before writing from scratch

Acceptance criteria

  • All 3 SKILL.md files exist at .agents/skills/<skill-name>/SKILL.md; metadata.category: test; authoring standard met
  • write-tests includes explicit constraint: derives from spec, not from implementation
  • generate-test-data includes PII flag check before generating any data
  • source: fields populated for any adopted upstream content
  • Each skill has a co-located eval at .agents/evals/test/<skill-name>/eval.yaml via write-eval
  • install.sh deploys all 3 to ~/.agents/skills/
  • HITL: human runs behavioral test per skill
  • HITL: human reviews each SKILL.md and eval before committing
  • Per-skill process followed for all 3 skills: source discovery (sub-agent) → source review with licence/security check (sub-agent) → conflict check against constitution + factory principles (sub-agent) → synthesis grill → co-write iteratively
  • Trigger description for each skill tested against explicit, implicit, and negative queries before body written
  • when: frontmatter field present in all SKILL.md files
  • source: and references: fields correctly populated or absent
  • eval.yaml for each skill contains all 5 required test types
  • Body ≤500 lines for each skill
  • docs/spec/overview.md updated to reflect all 3 skills deployed

Blocked by

  • 0016 (per-skill workflow)
  • 0017 (write-eval)
  • 0018 (write-skill)