feat: implement issue 0017 — factory/write-eval bootstrap skill
Adds write-eval, the first factory meta-skill. Produces eval.yaml test files for skills following the two-section schema (trigger_tests + output_tests) with provider-agnostic string assertions and show-plan- then-merge-on-rerun behaviour. Hand-written bootstrap — subsequent skills will use write-eval to produce their own evals. Also tightens skill-implementation-workflow.md step 5b: per-section options walk-through is now a named gate before writing, separate from the synthesis grill. LESSONS.md entry added. HITL behavioral test pending. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -34,6 +34,10 @@ Behavioral tests (2026-05-17) showed three communication/behavior rules failing:
|
||||
|
||||
The secrets prohibition in `core/instructions/governance.md` fired correctly when asked to write a password to a file, but the agent then reproduced the literal credential in its response text (in a shell `export` example). The rule was interpreted as "don't write to files" not "don't output at all." Fix: the rule needs to explicitly state "never produce the credential value in any output" and give an example showing placeholder usage (`export DB_PASSWORD='<your-password>'`).
|
||||
|
||||
## 2026-05-17 — Synthesis grill and SKILL.md co-write are two separate conversations
|
||||
|
||||
The synthesis grill (step 4) answers schema-level questions: how to combine upstreams, which eval schema to use, merge behaviour. Step 5b is a different conversation: how upstream content maps to each SKILL.md body section, what options each section had, and which was chosen. Collapsing them — writing the SKILL.md immediately after the grill without a per-section walk-through — means the human never sees the upstream options for the body and has no opportunity to redirect before the file is written. Fix: step 5b is now a named gate in the workflow. Walk through every body section one at a time, cite the upstream source, present alternatives, get confirmation. Only then write. Applies to both hand-written (bootstrap) and write-skill-produced skills.
|
||||
|
||||
## 2026-05-17 — HITL gap: agent delegates confirmation to permission system
|
||||
|
||||
The agent-level HITL rule ("require explicit confirmation before irreversible shared-state operations") is being bypassed: the agent calls the tool and lets the permission dialog catch it. This means the rule is not firing in agent reasoning — it's the permission system acting as a safety net. If a user selects "don't ask again," the net disappears. Fix: the HITL rule needs to be framed as "do not call the tool" rather than "ask before proceeding" — the agent must ask first, then act only after explicit confirmation.
|
||||
|
||||
Reference in New Issue
Block a user