0022 — Test skills: write-tests, generate-test-data, review-test-coverage #39

Open
opened 2026-06-28 17:11:58 +00:00 by Claude · 0 comments
Collaborator

Type: HITL
Parent PRD: docs/prd/chunk-3-skills-library.md

What to build

The 3 test phase skills. All are new. Authored via write-skill (0018), evals via write-eval (0017).

Skills and trigger descriptions:

Flat name Trigger description
write-tests Write tests, generate test cases, add unit tests
generate-test-data Generate test data, create fixtures, sample data
review-test-coverage Review test coverage, find untested paths, coverage gaps

Key constraints per skill:

  • write-tests: derives tests from spec (EARS acceptance criteria), NOT from implementation; uses pytest for Python, Vitest/Jest for TypeScript
  • generate-test-data: produces structurally valid, semantically unusual data; flags PII risk before generating
  • review-test-coverage: reports coverage gaps against spec acceptance criteria, not line coverage percentages

Implementation notes

Follow the per-skill workflow defined in docs/notes/skill-implementation-workflow.md.

Known upstream sources to review:

  • mattpocock/skills — check for any test-phase skills in the current set
  • bmad-method/bmad-method — BMAD QA role patterns
  • Search agentskills.io and GitHub for open-source test generation skills before writing from scratch

Acceptance criteria

  • All 3 SKILL.md files exist at .agents/skills/<skill-name>/SKILL.md; metadata.category: test; authoring standard met
  • write-tests includes explicit constraint: derives from spec, not from implementation
  • generate-test-data includes PII flag check before generating any data
  • source: fields populated for any adopted upstream content
  • Each skill has a co-located eval at .agents/evals/test/<skill-name>/eval.yaml via write-eval
  • install.sh deploys all 3 to ~/.agents/skills/
  • HITL: human runs behavioral test per skill
  • HITL: human reviews each SKILL.md and eval before committing
  • Per-skill process followed for all 3 skills: source discovery → source review → conflict check → synthesis grill → co-write iteratively
  • Trigger description for each skill tested against explicit, implicit, and negative queries before body written
  • eval.yaml for each skill contains all 5 required test types
  • Body ≤500 lines for each skill
  • docs/spec/overview.md updated to reflect all 3 skills deployed

Blocked by

  • 0016 (per-skill workflow)
  • 0017 (write-eval)
  • 0018 (write-skill)
**Type:** HITL **Parent PRD:** `docs/prd/chunk-3-skills-library.md` ## What to build The 3 test phase skills. All are new. Authored via `write-skill` (0018), evals via `write-eval` (0017). **Skills and trigger descriptions:** | Flat name | Trigger description | |---|---| | `write-tests` | Write tests, generate test cases, add unit tests | | `generate-test-data` | Generate test data, create fixtures, sample data | | `review-test-coverage` | Review test coverage, find untested paths, coverage gaps | **Key constraints per skill:** - `write-tests`: derives tests from spec (EARS acceptance criteria), NOT from implementation; uses pytest for Python, Vitest/Jest for TypeScript - `generate-test-data`: produces structurally valid, semantically unusual data; flags PII risk before generating - `review-test-coverage`: reports coverage gaps against spec acceptance criteria, not line coverage percentages ## Implementation notes Follow the per-skill workflow defined in `docs/notes/skill-implementation-workflow.md`. **Known upstream sources to review:** - `mattpocock/skills` — check for any test-phase skills in the current set - `bmad-method/bmad-method` — BMAD QA role patterns - Search agentskills.io and GitHub for open-source test generation skills before writing from scratch ## Acceptance criteria - [ ] All 3 SKILL.md files exist at `.agents/skills/<skill-name>/SKILL.md`; `metadata.category: test`; authoring standard met - [ ] `write-tests` includes explicit constraint: derives from spec, not from implementation - [ ] `generate-test-data` includes PII flag check before generating any data - [ ] `source:` fields populated for any adopted upstream content - [ ] Each skill has a co-located eval at `.agents/evals/test/<skill-name>/eval.yaml` via `write-eval` - [ ] `install.sh` deploys all 3 to `~/.agents/skills/` - [ ] **HITL:** human runs behavioral test per skill - [ ] **HITL:** human reviews each SKILL.md and eval before committing - [ ] Per-skill process followed for all 3 skills: source discovery → source review → conflict check → synthesis grill → co-write iteratively - [ ] Trigger description for each skill tested against explicit, implicit, and negative queries before body written - [ ] eval.yaml for each skill contains all 5 required test types - [ ] Body ≤500 lines for each skill - [ ] `docs/spec/overview.md` updated to reflect all 3 skills deployed ## Blocked by - 0016 (per-skill workflow) - 0017 (`write-eval`) - 0018 (`write-skill`)
Claude added this to the Legacy / Triage milestone 2026-06-28 17:11:58 +00:00
Claude added the Kind/Feature
Priority
Medium
3
labels 2026-06-28 17:11:58 +00:00
Claude modified the milestone from Legacy / Triage to Skills & Agents 2026-06-28 19:26:49 +00:00
Sign in to join this conversation.