Files
holocron/plugins/bin/skills/improve-codebase-architecture
Defame1297 4011d149bc fix(bin): put five skills' boundaries where the router can read them
grill-me, grill-with-docs, improve-codebase-architecture, tdd and triage each had
their routing boundary written into README.md, which nothing loads at runtime,
while the gate still reported all five descriptions as boundary-less. The
boundaries move into the descriptions; write-docs' clause, which said 'those have
dedicated skills' without naming one, now names them.

research had moved its body out and then read both references unconditionally --
the anti-goal ADR-0020 names, where the word count moves and the per-run context
does not. Both loads are genuinely conditional now, with the topic list and the
four literal sources.md field names inlined, since the provenance validator
matches those literally.

Also restores the promote-the-prototype anti-pattern to prototype's ui.md, which
the gate does not measure, so deleting it bought nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EJJrm5YmacbwMdzZpXcoti
2026-08-31 19:47:00 +00:00
..

improve-codebase-architecture

Surface architectural friction and propose deepening opportunities — refactors that turn shallow modules into deep ones.

What it does

Looks for places where a codebase is hard to understand, hard to test, or hard for an agent to navigate, and proposes refactors that concentrate behaviour behind smaller interfaces. It runs in three stages:

  1. Explore. Reads the domain glossary and any ADRs in the area first, then walks the codebase with an Explore sub-agent — organically, noting friction rather than applying fixed heuristics. The deletion test is the filter: imagine deleting the module; if complexity vanishes it was a pass-through, if complexity reappears across N callers it was earning its keep.
  2. Present candidates. A numbered list, each with files, problem, solution and benefits — benefits stated in terms of locality and leverage and of how tests would improve. No interfaces are proposed yet; the user picks one.
  3. Grilling loop. Walks the design tree for the chosen candidate, with documentation side effects landing inline as decisions crystallise.

The skill is opinionated about vocabulary, and that is the point: module, interface, implementation, depth, seam, adapter, leverage, locality, used exactly, with no drift into "component", "service", "API" or "boundary". Domain nouns come from CONTEXT.md, architecture nouns from LANGUAGE.md — so a proposal reads as "the Order intake module", never "the FooBarHandler".

ADRs are treated as decisions not to be re-litigated. A candidate that contradicts one is surfaced only when the friction is real enough to warrant reopening it, and is marked as such.

Composition

diagnose hands off here when a bug's post-mortem concludes that no correct test seam exists, or that callers are tangled — the recommendation is made after the fix is in, not before. The grilling loop follows grill-with-docs's discipline for CONTEXT.md entries and ADR offers, and SKILL.md names that skill's format documents directly.

Usage

/improve-codebase-architecture

Point at a codebase or an area of one. Expect a numbered candidate list and a "which of these would you like to explore?" before any interface design happens.

Files

File Purpose
SKILL.md Condensed glossary, key principles, and the three-stage process
LANGUAGE.md Skill-root document, cited throughout SKILL.md: full definitions of every term, the words each one replaces, and the full principle list
INTERFACE-DESIGN.md Skill-root document, read at stage 3 when the user wants alternative interfaces explored: the parallel sub-agent "Design It Twice" pattern, framing the problem space, and the per-agent design constraints
DEEPENING.md Skill-root document, cited from INTERFACE-DESIGN.md: how to deepen a cluster of shallow modules safely, the four dependency categories (in-process, local-substitutable, remote-but-owned, true external), seam discipline, and the replace-don't-layer testing strategy