PR #133 renamed the skill's root-level LANGUAGE.md to references/language.md but missed a prose mention (not a markdown link) in the overview paragraph. Fix both the .apm/ source and its generated flat mirror. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PDj6F7SPXzh3FtPN78dZ88
improve-codebase-architecture
Surface architectural friction and propose deepening opportunities — refactors that turn shallow modules into deep ones.
What it does
Looks for places where a codebase is hard to understand, hard to test, or hard for an agent to navigate, and proposes refactors that concentrate behaviour behind smaller interfaces. It runs in three stages:
- Explore. Reads the domain glossary and any ADRs in the area first, then walks the codebase with an
Exploresub-agent — organically, noting friction rather than applying fixed heuristics. The deletion test is the filter: imagine deleting the module; if complexity vanishes it was a pass-through, if complexity reappears across N callers it was earning its keep. - Present candidates. A numbered list, each with files, problem, solution and benefits — benefits stated in terms of locality and leverage and of how tests would improve. No interfaces are proposed yet; the user picks one.
- Grilling loop. Walks the design tree for the chosen candidate, with documentation side effects landing inline as decisions crystallise.
The skill is opinionated about vocabulary, and that is the point: module, interface, implementation, depth, seam, adapter, leverage, locality, used exactly, with no drift into "component", "service", "API" or "boundary". Domain nouns come from CONTEXT.md, architecture nouns from references/language.md — so a proposal reads as "the Order intake module", never "the FooBarHandler".
ADRs are treated as decisions not to be re-litigated. A candidate that contradicts one is surfaced only when the friction is real enough to warrant reopening it, and is marked as such.
Composition
diagnose hands off here when a bug's post-mortem concludes that no correct test seam exists, or that callers are tangled — the recommendation is made after the fix is in, not before. The grilling loop follows grill-with-docs's discipline for CONTEXT.md entries and ADR offers, and SKILL.md names that skill's format documents directly.
Usage
/improve-codebase-architecture
Point at a codebase or an area of one. Expect a numbered candidate list and a "which of these would you like to explore?" before any interface design happens.
Files
| File | Purpose |
|---|---|
SKILL.md |
Condensed glossary, key principles, and the three-stage process |
references/language.md |
Cited throughout SKILL.md: full definitions of every term, the words each one replaces, and the full principle list |
references/interface-design.md |
Read at stage 3 when the user wants alternative interfaces explored: the parallel sub-agent "Design It Twice" pattern, framing the problem space, and the per-agent design constraints |
references/deepening.md |
Cited from references/interface-design.md: how to deepen a cluster of shallow modules safely, the four dependency categories (in-process, local-substitutable, remote-but-owned, true external), seam discipline, and the replace-don't-layer testing strategy |