Adds write-eval, the first factory meta-skill. Produces eval.yaml test files for skills following the two-section schema (trigger_tests + output_tests) with provider-agnostic string assertions and show-plan- then-merge-on-rerun behaviour. Hand-written bootstrap — subsequent skills will use write-eval to produce their own evals. Also tightens skill-implementation-workflow.md step 5b: per-section options walk-through is now a named gate before writing, separate from the synthesis grill. LESSONS.md entry added. HITL behavioral test pending. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
5.2 KiB
Overview
Current deployed state of this repo — what you get if you run install.sh today. Updated at the close of each chunk and in the same PR as any behavior change.
Last updated: 2026-05-17 (issue 0017)
What is deployed
Skills
13 skills deployed to ~/.agents/skills/ via install.sh. Available as slash commands in Claude Code via ~/.claude/skills/ → ~/.agents/skills/ symlink. 12 are first-draft placeholders pending rebuild in Chunk 3; 1 is a new Chunk 3 factory skill (write-eval).
Factory bootstrap (Chunk 3): write-eval deployed — produces eval.yaml test files for skills. Hand-written (bootstrap skill). Eval at .agents/evals/factory/write-eval/eval.yaml.
Current skills: caveman, diagnose, grill-me, grill-with-docs, improve-codebase-architecture, prototype, tdd, to-issues, to-prd, triage, write-a-skill, zoom-out.
Chunk 3 target: 42 skills across 9 categories. PRD: docs/prd/chunk-3-skills-library.md. Canonical build reference: docs/research/ai-coding-factory/ai-coding-factory-skills-index.md (delete once all skills exist). Skills stored flat (skill-name/SKILL.md) per ADR-0009; category in metadata.category frontmatter. Categories: design, factory, implement, test, review, deploy, operate, cross-cutting, iac (2 skills only — docker-compose + iac-security-review). Role skills (6) deferred to Chunk 5. Gitea skills moved to providers/gitea/ provider adapter.
Skill implementation workflow: each skill follows the per-skill process in docs/notes/skill-implementation-workflow.md (produced by issue 0016). Sub-agents handle source discovery, source review, and conflict checking; synthesis grill and HITL test are human steps. Bootstrap: write-eval (hand-written) → write-skill (hand-written) → write-docs (first factory-authored) → all others via factory.
Claude Code configuration
~/.claude/CLAUDE.md— thin adapter; imports~/.agents/AGENTS.md(Communication + Behavior) andgovernance.md; content index pointers only~/.agents/AGENTS.md— global always-on rules (Communication + Behavior); provider-agnostic source of truth~/.claude/core/instructions/— coding, git, testing, governance instruction files~/.claude/settings.json— Claude Code settings
Governance layer
core/instructions/governance.md loads into every Claude Code session via @import in ~/.claude/CLAUDE.md. Covers: hard prohibitions on secrets and data, data classification tiers, HITL requirements, sycophancy resistance, deterministic execution preference.
What works end-to-end
install.shruns idempotently — safe to re-run after changes- Provider adapter pattern:
providers/*/provider-manifest.shauto-discovered byinstall.sh - Governance rules take effect at session start without any manual loading step
- Skills available as slash commands immediately after install
What is not yet deployed
sync.sh— pulls updates into existing projects (Chunk 6)init-project.sh— bootstraps a new project (Chunk 6)- Copilot provider adapter (Chunk 7)
- Formal CI/pre-commit enforcement of governance rules (Chunk 6)
For chunk planning and open questions, see docs/ROADMAP.md.
Recent changes
- 2026-05-17 — Issue 0017 complete:
write-evalbootstrap skill written and deployed. Two sections schema (trigger_tests+output_tests), provider-agnostic string assertions, show-plan-then-merge-on-rerun behaviour, conflict flagging (B model). Sources: agentskills/agentskills, darkrishabh/agent-skills-eval, bmad-code-org/BMAD-METHOD, mattpocock/skills. Hand-written eval at.agents/evals/factory/write-eval/eval.yaml. - 2026-05-17 — Issue 0016 complete: skill implementation workflow grill completed.
docs/notes/skill-implementation-workflow.mdwritten. All issues 0017–0028 updated with specific acceptance criteria. Key conventions: sub-agents prescribed at each research/writing step; conflict check against constitution + factory principles before synthesis grill;when:andreferences:fields added to authoring standard; write-docs moved to issue 0018 phase 2 (first factory-authored skill). - 2026-05-17 — Issue 0015 complete: AGENTS.md refactor implemented. Two AGENTS.md files created (
AGENTS.mdat repo root,core/AGENTS.mddeployed to~/.agents/AGENTS.md). Both CLAUDE.md files slimmed to thin adapters.deploy-manifest.shupdated.docs/spec/architecture.mdupdated with new structure. ADR-0012 in effect. - 2026-05-17 — Chunk 3 issues created (0015–0028): AGENTS.md refactor prerequisite, skill workflow grill, bootstrap skills (write-eval, write-skill), factory/design/implement/test/review/deploy/operate/IaC/cross-cutting skill groups, chunk closure; all HITL; acceptance criteria for 0017–0028 to be refined after issue 0016 grill session
- 2026-05-17 — behavioral tests fully resolved:
CONTEXT.mdnow always-loaded via@importin repoCLAUDE.md; standing rule added to checkdocs/adr/and ROADMAP resolved entries before answering design questions; communication/behavior and secrets rules tightened; Chunk 2 and Governance Phase 1 ✅ complete - 2026-05-17 — added
LESSONS.md(issue 0013) anddocs/spec/(issue 0014); refactoreddocs/VISION.mdto goals/intent only