Files
holocron/docs/issues/0028-chunk-3-closure.md
Defame1297 5f355d664d chore: mark 0017 and 0018 HITL gates complete via HOTL subagent tests
Behavioral tests for write-eval, write-skill, and write-docs run via
fresh-context subagents (HOTL). All process steps verified correct.
Caveman defects surfaced during testing logged in 0028 for upgrade-skill.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-26 19:19:08 +00:00

3.3 KiB
Raw Blame History

0028 — Chunk 3 closure: update skills-index, update spec, behavioral tests

Type: HITL
Parent PRD: docs/prd/chunk-3-skills-library.md

What to build

Close out Chunk 3 once all 42 skills are complete: update the skills index to reflect the implemented state, update the living spec, and run the full behavioral acceptance test suite.

Tasks:

  1. Update docs/research/ai-coding-factory/ai-coding-factory-skills-index.md — replace the pre-implementation build reference with the as-implemented state: actual flat skill names, categories, trigger descriptions as deployed, any deviations from the original index noted
  2. Update docs/spec/overview.md — reflect the full 42-skill library as the current deployed state; remove "Chunk 3 target" language; mark Chunk 3 ✅ complete
  3. Update docs/spec/architecture.md — reflect the .agents/evals/ directory structure added in Chunk 3; any other structural changes from implementation
  4. Update docs/ROADMAP.md — mark Chunk 3 ✅ complete in the chunk table
  5. Run behavioral acceptance tests — for each skill, invoke with its trigger phrase in a fresh Claude session and verify the output meets the authoring standard; document results

Behavioral test scope: All 42 skills (including write-eval, write-skill, and the 4 preserved skills). The caveman skill is exempt — it has no content-generating behavior to verify.

Known caveman defects (surfaced during 0017 HOTL test, 2026-05-26): caveman is a pre-standard legacy skill pending adoption via upgrade-skill. Two defects to fix at that time: (1) missing metadata.category: cross-cutting in frontmatter — write-eval cannot compute output path without it; (2) "be brief" trigger is over-broad — fires on one-shot brevity requests, not just persistent mode activation. Negative test cases documenting the correct boundary are captured in the HOTL test output.

LESSONS.md: Extract any cross-session learnings from Chunk 3 implementation and add entries per the LESSONS.md format. Three or more observations on the same pattern graduate to the relevant standing file. 6. Review docs/notes/skill-implementation-workflow.md — verify the conventions are still accurate; update any entries that changed during implementation.

Acceptance criteria

  • docs/research/ai-coding-factory/ai-coding-factory-skills-index.md updated to reflect as-implemented state; deviations from original plan noted
  • docs/spec/overview.md updated; Chunk 3 marked ✅ complete; all 42 skills listed as deployed
  • docs/spec/architecture.md updated with .agents/evals/ structure
  • docs/ROADMAP.md Chunk 3 row updated to ✅
  • Behavioral test run completed; all skills pass their trigger test; failures documented as issues for resolution
  • LESSONS.md updated with Chunk 3 observations
  • docs/notes/skill-implementation-workflow.md reviewed and updated to reflect any workflow changes discovered during Chunk 3
  • All skill issues (0017–0027) have a ## Handoff section with status complete
  • HITL: human verifies the complete skills library in a fresh session before marking Chunk 3 done

Blocked by

  • 0020 (design skills)
  • 0021 (implement skills)
  • 0022 (test skills)
  • 0023 (review skills)
  • 0024 (deploy skills)
  • 0025 (operate skills)
  • 0026 (IaC skills)
  • 0027 (cross-cutting skills)