refactor(bin): cut per-invocation load and repair broken skill references
diagnose read its feedback-loops reference unconditionally, so every invocation paid for guidance most runs never used; the read is conditional again and the per-invocation cost drops from 1,278 to 831 words. Its HITL template moves to assets/ because it is copied out, not read as reference. prototype's two branch flows move into references/ for the same reason — only one branch is ever taken. research could not search the codebase it was asked to research without Grep and Glob. caveman's description had grown into a paragraph where one sentence carries the trigger. Four references pointed at things that do not exist: a to-prd skill, a /setup-matt-pocock-skills command, two cross-skill ../ links that only resolve in the source tree, and two places calling this project's Gitea host GitHub. Addresses #114.
This commit is contained in:
@@ -18,7 +18,9 @@ When exploring the codebase, use the project's domain glossary to get a clear me
|
||||
|
||||
Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**
|
||||
|
||||
Read `references/feedback-loops.md` — even if you already have a signal. Ten ways to build a loop ordered by cost, how to sharpen the one you have, and what to do when the bug resists reproduction. An unsharpened loop is usually not good enough yet.
|
||||
**If you do not yet have such a signal, read `references/feedback-loops.md`** — ten ways to build one ordered by cost, and what to ask the user for when the bug resists reproduction entirely.
|
||||
|
||||
**If you do have one, it is probably not sharp enough yet.** Make it faster and more deterministic, and make it assert on the exact symptom rather than "didn't crash" — a 30-second flaky loop is barely better than no loop. If it stays slow or intermittent after that, read that file's "Iterate on the loop itself" and "Intermittent bugs" sections.
|
||||
|
||||
Do not proceed to Phase 2 until you have a loop you believe in. If you cannot build one, stop and say so explicitly, listing what you tried — never hypothesise without a signal.
|
||||
|
||||
@@ -62,13 +64,11 @@ Tool preference:
|
||||
|
||||
## Phase 5 — Fix + regression test
|
||||
|
||||
Write the regression test **before the fix** — but only if there is a **correct seam** for it.
|
||||
Write the regression test **before the fix** — but only at a **correct seam**: one where the test exercises the real bug pattern as it occurs at the call site. If the available seam looks too shallow, or you cannot tell whether it is, read `references/regression-seams.md`.
|
||||
|
||||
A correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site. If the only available seam is too shallow (single-caller test when the bug needs multiple callers, unit test that can't replicate the chain that triggered the bug), a regression test there gives false confidence.
|
||||
**If no correct seam exists, that itself is the finding.** Note it and carry it into Phase 6 — the architecture is preventing the bug from being locked down.
|
||||
|
||||
**If no correct seam exists, that itself is the finding.** Note it. The codebase architecture is preventing the bug from being locked down. Flag this for the next phase.
|
||||
|
||||
If a correct seam exists:
|
||||
At a correct seam:
|
||||
|
||||
1. Turn the Phase 1 loop into a failing test at that seam, narrowed to the symptom captured in Phase 2.
|
||||
2. Watch it fail.
|
||||
|
||||
@@ -13,7 +13,7 @@ A feedback loop is a fast, deterministic, agent-runnable pass/fail signal for th
|
||||
7. **Property / fuzz loop.** If the bug is "sometimes wrong output", run 1000 random inputs and look for the failure mode.
|
||||
8. **Bisection harness.** If the bug appeared between two known states (commit, dataset, version), automate "boot at state X, check, repeat" so you can `git bisect run` it.
|
||||
9. **Differential loop.** Run the same input through old-version vs new-version (or two configs) and diff outputs.
|
||||
10. **HITL bash script.** Last resort. If a human must click, drive _them_ with `scripts/hitl-loop.template.sh` so the loop is still structured. Captured output feeds back to you.
|
||||
10. **HITL bash script.** Last resort. If a human must click, drive _them_ with `assets/hitl-loop.template.sh` so the loop is still structured. Captured output feeds back to you.
|
||||
|
||||
## Iterate on the loop itself
|
||||
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
# Judging a regression-test seam
|
||||
|
||||
Read this when Phase 5 leaves you unsure whether the seam available for the regression test is the correct one — either because the obvious seam looks shallow, or because there appears to be no seam at all.
|
||||
|
||||
## What makes a seam correct
|
||||
|
||||
A correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site: the same entry point, the same participants, the same ordering, and the same state the real caller holds when it goes wrong.
|
||||
|
||||
## Seams that are too shallow
|
||||
|
||||
- A single-caller test when the bug only appears with multiple callers.
|
||||
- A unit test that cannot replicate the chain of calls that triggered the bug.
|
||||
- A test that reproduces the symptom by construction — asserting on a value the test itself set — rather than by driving the code path that produces it.
|
||||
- A test that mocks out the collaborator the bug actually lives in.
|
||||
|
||||
A regression test at a shallow seam gives false confidence. It passes forever, including after a change reintroduces the bug at the real call site, and it will be read by the next maintainer as proof the bug is locked down.
|
||||
|
||||
## When there is no correct seam
|
||||
|
||||
Do not force one, and do not settle for a shallow seam to have something green. Instead:
|
||||
|
||||
1. Apply the fix and verify it against the Phase 1 loop directly.
|
||||
2. Write down which seams you considered and why each was too shallow.
|
||||
3. Carry that into Phase 6's "what would have prevented this bug" question. A missing seam is an architecture finding — tangled callers, hidden coupling, or a module with no testable boundary — and the handoff is the `improve-codebase-architecture` skill, with those specifics attached.
|
||||
Reference in New Issue
Block a user