Files
holocron/docs/ROADMAP.md
Defame1297 74f5e1840d test: run Chunk 2 and governance behavioral tests; fix failing rules
13 manual scenarios run across instructions and governance layers (two
rounds for failures). Fixed four rules that lost to RLHF defaults:

- Exploratory question format: tightened with boundary framing; added
  @import CONTEXT.md to repo CLAUDE.md and a standing rule to check
  docs/adr/ and ROADMAP resolved entries before answering design questions
  (3-round iteration to resolve)
- File-edit intent: added counter-example to stop clarification-seeking
- Push confirmation: reframed as "do not call the tool" not "ask first"
- Secrets rule: extended to cover credential reproduction in response
  text and usage examples, with explicit placeholder requirement

Scenario 4 (push confirmation) inconclusive — no remote configured.
Governance scenario 3 (HITL on real infra) untestable — Nginx not installed.
Both share the same root cause: agent delegates to permission system.

Also corrects stale skill list in docs/spec/overview.md (12 actual
deployed skills vs 16 names previously listed).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-17 11:44:56 +00:00

101 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Roadmap
## Chunk conventions
Content chunks (2–5) run in two phases, treated as separate sessions:
1. **Architecture + thin drafts** — define the format, schema, and loading model; populate every category with a minimal first draft. Mark speculative entries with `<!-- draft -->` so future sessions know what to trust. Architecture decisions must be stable before phase 2.
2. **Focused refinement** — work through each category properly, one at a time. Treated as ongoing rather than a hard deadline; refinement is triggered by real friction, not a schedule.
Phase 1 is the planned chunk. Phase 2 is ongoing.
Chunk 6 (tooling) is exempt — it is implementation-driven, not content-driven.
## Governance workstream
A parallel workstream (not a numbered chunk) that runs alongside the chunk sequence. Cross-cutting concern — governance rules apply to all chunks.
**Phase 1 — instruction and documentation layer** ✅ complete (before Chunk 3)
- `core/instructions/governance.md` — agent instruction file loaded via `@import` at every session start
- `docs/ai-constitution.md` — full evidence base and governance principles (human-facing)
- `docs/HUMANS.md` — practitioner checklist (human-facing)
- `CONTEXT.md` — extended with governance domain language (HITL, HOTL, sycophancy, data classification tiers, symbolic oversight)
- `docs/VISION.md`, `CLAUDE.md`, `docs/ROADMAP.md` — updated to reflect governance layer existence
- `tests/test-governance-layer.sh` — manual test plan verifying governance rules take effect in a fresh session
**Phase 2 — deterministic enforcement layer** (Chunk 6)
- Pre-commit hooks, CI gates, secret scanning, licence scanning, audit logging infrastructure, human approval gates in CI/CD
- Specification: `docs/research/governance_principles/CONTROLS.md`
## Chunk table
| Chunk | Scope | Why this order |
|---|---|---|
| ✅ 1 | Repo skeleton + `install.sh` — structure in place, Claude Code wired up | Nothing else can be built without the structure and install working |
| ✅ 2 | Core instructions — `coding.md`, `git.md` (incl. conventional commits), `testing.md`; communication rules in `providers/claude-code/CLAUDE.md` always-on section; retire `global.md`; migrate `docs/` to subdirectory-by-type naming | Instructions are the foundation everything else references; commit convention and doc naming must be in place before history accumulates |
| ⏳ 3 | Skills library rebuild — all 12 existing skills are first-draft placeholders; Chunk 3 rebuilds each from scratch following the full SKILL.md authoring standard (version field, category metadata, constraints section, self-check, failure handling, trigger-test-first discipline). Process per skill: check `docs/research/ai-coding-factory/ai-coding-factory-implementation-guidance.md` Section 4–5 (framework sourcing) and `docs/research/ai-coding-factory/ai-coding-factory-skills-index.md` (pre-researched trigger descriptions and constraints for 34 target skills) → research/inspect open-source implementations → grill design → implement. New skills: session-handoff (cross-cutting), governance-check (cross-cutting), git-guardrails (cross-cutting), write-adr (factory). IaC skills (global optional, iac category) and Gitea skills (global optional, gitea category) — scope defined in Chunk 3 PRD. **Infrastructure complete**: 12 skills deployed to `~/.agents/skills/` via `install.sh`; provider adapter pattern in place; all existing skills treated as drafts pending rebuild | Skills are the most immediately useful output; the rebuild is necessary because existing skills predate the authoring standard and the factory research |
| 4 | Workflows — formalize the workstream workflow (kick-off types → grill → artifact → issues → implement → QA → commit); feature, bug, architecture, improvement, feedback patterns. **Prerequisite:** WorkflowContext schema (what each skill in a chain receives and returns) must be designed before any workflow skill is written; `docs/spec/` must exist (implement-feature constraint: update spec in same PR as behavior change) | Higher-level patterns built on top of a working skills foundation; grill feedback intake design before starting |
| 5 | Agents — role skills (Architect, Developer, Reviewer, Security, QA, Ops) in `.agents/skills/` with `category: roles`; `core/agents/` for provider-agnostic subagent definitions needing isolated execution context (`context: fork`), translated to `.claude/agents/` by adapter; cross-project orchestration agents as use case | Role skills benefit from workflow patterns being established first; subagent definitions require the skills library to be stable |
| 6 | Sync + project init tooling — `sync.sh` and `init-project.sh` | Tooling only makes sense once there is content worth syncing and scaffolding |
| 7 | Copilot provider — adapter for GitHub Copilot. **Provider adapter pattern established**: `install.sh` auto-discovers `providers/*/provider-manifest.sh`; Copilot adapter is a new `providers/copilot/provider-manifest.sh` declaring a symlink if needed | Second provider comes after the first is fully proven |
## Development workflow
Every workstream follows this shape. Pick a kick-off type, grill it, then run the implementation loop per issue.
```
Kick-off (pick type)
├── Feature → /grill-with-docs → PRD → /to-issues
├── Bug → /grill-with-docs → Bug Brief → /to-issues → /diagnose
├── Architecture → /grill-with-docs → ARD (+ ADR later) → /to-issues
├── Improvement → /grill-with-docs → PRD or ARD → /to-issues
├── Feedback → /triage → PRD or Bug Brief → /to-issues
└── Ideation → /grill-me → Exploration Note → /to-issues (optional)
Per issue
└── /tdd → implement → automated QA → commit (conventional)
Manual QA — only for nuanced UI/UX or agent interaction behavior
/improve-codebase-architecture — ad hoc or at chunk/PR boundaries, not per issue
Ongoing (ad hoc, within any workstream)
├── /diagnose (unexpected breakage)
├── /prototype (design uncertainty)
└── /zoom-out (orientation)
Finalize (per workstream)
└── update docs → commit
```
This workflow is defined at convention level in Chunk 2. Chunk 4 formalizes it as a composable skill/workflow.
## Open questions / deferred decisions
Items consciously not resolved — to be addressed in the relevant chunk PRD or grill.
| Question | Deferred to |
|---|---|
| How project-level overrides are structured and what they can override | Chunk 6 PRD |
| ~~Deployment manifest seam — `install.sh` embeds source→target mappings implicitly; `sync.sh` will need the same mapping.~~ | ✅ Resolved in Chunk 2 architecture review — extracted to `scripts/deploy-manifest.sh`; `sync.sh` sources the same file in Chunk 6 |
| Feedback intake workflow — where does feedback arrive (GitHub issues, Slack, email)? | Grill before Chunk 4 (workflows) |
| QA agent design — what does automated agent testing look like in practice? | Grill before Chunk 5 (agents) |
| Automated deployment pipeline — CI/CD beyond gitops convention | Chunk 6 grill |
| Formal CI gate for `/improve-codebase-architecture` | Chunk 6 grill |
| Changelog tooling — which generator (git-cliff, conventional-changelog, etc.) and where it runs | Chunk 3 grill |
| Content index frontmatter — replace inline `when:` hints in CLAUDE.md content index with a `when:` field in each instruction/skill file so the agent discovers load conditions from the file itself. Cover before implementing Chunk 3 skills. | Chunk 3 grill |
| ~~Skill taxonomy — flat vs nested paths, category organisation~~ | ✅ Resolved — factory integration grill. Flat paths (Claude Code + agentskills.io standard enforce one-level-deep discovery). Categories via `metadata: category:` in SKILL.md frontmatter. See ADR-0009. |
| ~~Factory boundary — which factory features belong here vs project repos~~ | ✅ Resolved — factory integration grill. This repo is a provider (ADR-0008). LESSONS.md and docs/spec/ are exceptions: added here because this repo also develops itself. IaC and Gitea skills are global optional. Role skills in .agents/skills/; core/agents/ for subagent definitions (ADR-0010). |
| IaC and Gitea skill scope — which specific skills to include in the global optional set, and in what order | Chunk 3 PRD |
| Agent behavior confirmation model — writes/edits/git currently require stating intent + approval before acting. Loosen to autonomy-first once skills and workflows are proven and automated agents replace direct interaction. | Phase 2 refinement (post Chunk 4) |
| ~~CLAUDE.md always-on refinement — security floor (no credentials/auth URLs), scope discipline (no over-engineering), tool preference (Read/Edit over Bash); **plus instruction quality**: current rules are thin one-liners observed in practice to lose to RLHF-trained defaults (verbose responses, validating user positions); fix is specificity, counter-examples, and boundary framing — not accepting violations as expected. Needs its own grill session → PRD before implementation.~~ | ✅ Resolved — Governance workstream Phase 1. `core/instructions/governance.md` loaded via `@import` covers hard prohibitions, data classification, HITL, sycophancy resistance, and deterministic execution preference. Instruction quality principle documented in `CONTEXT.md`. |
## Housekeeping reminders
- **AI coding factory integration** — grill complete. Decision record: `docs/notes/factory-integration-decisions.md`. ADRs: 0008 (factory boundary), 0009 (flat taxonomy), 0010 (role skills vs subagents). Follow-on issues: ~~0013 (LESSONS.md)~~ ✅, ~~0014 (docs/spec/ + VISION.md refactor)~~ ✅. Chunk 3 scope substantially expanded — skills rebuild, new skills, IaC/Gitea skills. See updated chunk table above.
- **`.gitkeep` files** — placeholder files exist in `core/agents/`, `core/workflows/`, `core/prompts/`, `docs/ard/`, `docs/bug/`. Remove each when the first real file is added to that directory. Each `.gitkeep` names the chunk that will populate it. (`docs/notes/.gitkeep` already removed — directory has real content.)
- **Skills pipeline verified** — `install.sh` deploys 12 skills to `~/.agents/skills/` and creates `~/.claude/skills/ → ~/.agents/skills/` symlink adapter. Tested idempotent. `skills-lock.json` removed (was a manual artifact). If `~/.claude/skills/` exists as a real directory on a machine being migrated, remove it manually and re-run install.
- **Chunk 2 behavioral tests** — run and fully resolved 2026-05-17. 7/8 pass; scenario 4 (push confirmation) inconclusive — no remote in test environment, rule tightened but unverified. All fixable failures addressed: rule specificity in `providers/claude-code/CLAUDE.md`; context-loading guarantee via `@import CONTEXT.md` in repo CLAUDE.md; standing rule in CONTEXT.md to check `docs/adr/` and ROADMAP resolved entries before answering design questions. Chunk 2 ✅ complete.
- **Governance Phase 1 behavioral tests** — run 2026-05-17. 3/4 testable scenarios pass. Secrets rule gap fixed (2026-05-17): extended to cover credential reproduction in response text and examples, with placeholder requirement added to `core/instructions/governance.md`. HITL scenario not testable in this environment (Nginx not installed); HITL gap evidenced by instructions test scenario 4 — push confirmation rule fix addresses the same root cause. Governance Phase 1 ✅ complete.
- **AI ethics/security workstream** — `docs/notes/ai-ethics-security-principles.md` exploration note is superseded. Governance Phase 1 (`core/instructions/governance.md`) covers all planned scope: credentials, data classification, HITL, scope discipline, agent autonomy, transparency, and security code review. Tier-placement architectural question resolved by the `@import` always-on model. No separate workstream needed.