feat: implement issue 0017 — factory/write-eval bootstrap skill
Adds write-eval, the first factory meta-skill. Produces eval.yaml test files for skills following the two-section schema (trigger_tests + output_tests) with provider-agnostic string assertions and show-plan- then-merge-on-rerun behaviour. Hand-written bootstrap — subsequent skills will use write-eval to produce their own evals. Also tightens skill-implementation-workflow.md step 5b: per-section options walk-through is now a named gate before writing, separate from the synthesis grill. LESSONS.md entry added. HITL behavioral test pending. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
85
.agents/evals/factory/write-eval/eval.yaml
Normal file
85
.agents/evals/factory/write-eval/eval.yaml
Normal file
@@ -0,0 +1,85 @@
|
||||
skill_name: write-eval
|
||||
|
||||
trigger_tests:
|
||||
- id: explicit-trigger-write-evals
|
||||
name: "Explicit trigger — write evals"
|
||||
query: "write evals for this skill"
|
||||
should_trigger: true
|
||||
|
||||
- id: explicit-trigger-create-eval-yaml
|
||||
name: "Explicit trigger — create eval.yaml"
|
||||
query: "create eval.yaml for the tdd skill"
|
||||
should_trigger: true
|
||||
|
||||
- id: implicit-trigger-test-coverage
|
||||
name: "Implicit trigger — test coverage request"
|
||||
query: "I need test coverage for the grill-me skill"
|
||||
should_trigger: true
|
||||
|
||||
- id: negative-trigger-run-evals
|
||||
name: "Negative — run evals (runner concern, not writer)"
|
||||
query: "run my evals"
|
||||
should_trigger: false
|
||||
|
||||
- id: negative-trigger-code-unit-tests
|
||||
name: "Negative — unit tests for application code"
|
||||
query: "write unit tests for my Python file"
|
||||
should_trigger: false
|
||||
|
||||
- id: negative-trigger-debug-failing-eval
|
||||
name: "Negative — debug failing eval"
|
||||
query: "my eval is failing, help me debug it"
|
||||
should_trigger: false
|
||||
|
||||
output_tests:
|
||||
- id: deterministic-correct-output-path
|
||||
name: "Deterministic — eval.yaml written to correct path"
|
||||
type: deterministic
|
||||
prompt: "write evals for the tdd skill"
|
||||
expected_output: "eval.yaml written to .agents/evals/implement/tdd/eval.yaml containing skill_name: tdd"
|
||||
assertions:
|
||||
- "Output references the path .agents/evals/implement/tdd/eval.yaml"
|
||||
- "Output file contains 'skill_name: tdd'"
|
||||
|
||||
- id: deterministic-all-five-types-present
|
||||
name: "Deterministic — eval.yaml contains all five required test types"
|
||||
type: deterministic
|
||||
prompt: "create eval.yaml for the grill-me skill"
|
||||
expected_output: "eval.yaml contains trigger_tests and output_tests sections with all five required test types represented"
|
||||
assertions:
|
||||
- "Output contains 'trigger_tests:'"
|
||||
- "Output contains 'output_tests:'"
|
||||
- "Output contains at least one entry with 'should_trigger: true'"
|
||||
- "Output contains at least one entry with 'should_trigger: false'"
|
||||
- "Output contains at least one entry with 'type: deterministic'"
|
||||
- "Output contains at least one entry with 'type: llm-rubric'"
|
||||
|
||||
- id: deterministic-plan-shown-before-write
|
||||
name: "Deterministic — test plan presented before file is written"
|
||||
type: deterministic
|
||||
prompt: "write evals for the diagnose skill"
|
||||
expected_output: "Skill presents proposed test cases for review and requests confirmation before writing any file"
|
||||
assertions:
|
||||
- "Response includes a proposed test plan or list of test cases before any file is written"
|
||||
- "Response requests confirmation before proceeding to write"
|
||||
|
||||
- id: deterministic-merge-conflict-flagged
|
||||
name: "Deterministic — conflict flagged in plan on re-run with existing eval"
|
||||
type: deterministic
|
||||
prompt: "write evals for the tdd skill — eval.yaml already exists at .agents/evals/implement/tdd/eval.yaml with a test case id 'explicit-trigger-basic'"
|
||||
expected_output: "Skill identifies the existing eval.yaml, classifies the conflicting case as CONFLICT, and does not write until the user resolves it"
|
||||
assertions:
|
||||
- "Response indicates eval.yaml already exists at the target path"
|
||||
- "Response labels the conflicting test case as CONFLICT or equivalent"
|
||||
- "Response does not write the file before the user resolves the conflict"
|
||||
|
||||
- id: llm-rubric-assertion-quality
|
||||
name: "LLM rubric — assertions are specific and verifiable"
|
||||
type: llm-rubric
|
||||
prompt: "write evals for the write-skill skill"
|
||||
expected_output: "eval.yaml contains high-quality assertions that are specific, observable, and not vague"
|
||||
assertions:
|
||||
- "All assertions describe observable, verifiable conditions — not vague quality claims like 'output is good' or 'the response is helpful'"
|
||||
- "Trigger test queries reflect realistic user phrasings, not just the exact skill description verbatim"
|
||||
- "Negative trigger tests target adjacent tasks that share surface-level similarity with the skill's trigger"
|
||||
- "Deterministic assertions are machine-checkable without LLM inference — presence of strings, path patterns, required sections"
|
||||
141
.agents/skills/write-eval/SKILL.md
Normal file
141
.agents/skills/write-eval/SKILL.md
Normal file
@@ -0,0 +1,141 @@
|
||||
---
|
||||
name: write-eval
|
||||
description: Write or generate an eval.yaml test file for a skill. Use when the user wants to create evals, add test coverage, or says "write evals for this skill", "create eval.yaml for X", or "add tests for this skill". Do NOT use when the user wants to run existing evals, write unit tests for code, or debug test failures.
|
||||
version: "1.0"
|
||||
updated: 2026-05-17
|
||||
when: invoked by explicit trigger ("write evals for this skill", "create eval.yaml for X") or implicit request for skill test coverage
|
||||
metadata:
|
||||
category: factory
|
||||
source:
|
||||
- repo: agentskills/agentskills
|
||||
commit: 2d3e01f590f68bee2cb76a3200823e93b2cc9eaa
|
||||
files:
|
||||
- docs/skill-creation/evaluating-skills.mdx # evals schema, two-section workspace layout, assertion quality guidelines
|
||||
updated: 2026-05-17
|
||||
- repo: darkrishabh/agent-skills-eval
|
||||
commit: b60eebe3c6edaa917a284e13b9b0e9fa00f1c957
|
||||
files:
|
||||
- src/types.ts # AgentSkillsEval interface — string-slug id, name field, prompt/expected_output/assertions structure
|
||||
- examples/basic-skill/evals/evals.json # concrete schema example
|
||||
updated: 2026-05-17
|
||||
- repo: bmad-code-org/BMAD-METHOD
|
||||
commit: 71136bc6af77cbf507d3768494311d5b6ca95cc5
|
||||
files:
|
||||
- evals/bmm-skills/bmad-product-brief/triggers.json # trigger classification dataset, should_trigger boolean pattern
|
||||
- evals/bmm-skills/bmad-product-brief/evals.json # output test structure, boundary-enforcement negative test pattern
|
||||
updated: 2026-05-17
|
||||
- repo: mattpocock/skills
|
||||
commit: e74f0061bb67222181640effa98c675bdb2fdaa7
|
||||
files:
|
||||
- skills/engineering/tdd/SKILL.md # behavioral test philosophy: test observable outputs through public interfaces
|
||||
updated: 2026-05-17
|
||||
references:
|
||||
- https://agentskills.io/skill-creation/evaluating-skills
|
||||
---
|
||||
|
||||
## Role
|
||||
|
||||
You are a test architect producing eval.yaml files that verify AI skill trigger behaviour and output quality.
|
||||
|
||||
## When to use / When not to use
|
||||
|
||||
**Use when:**
|
||||
- User explicitly requests evals: "write evals for this skill", "create eval.yaml for X", "add tests for this skill"
|
||||
- A new or refactored skill needs an eval file
|
||||
- Existing eval coverage needs to be extended with additional test cases
|
||||
|
||||
**Do not use when:**
|
||||
- User wants to run or execute existing evals
|
||||
- User wants to write unit tests for application code (not a skill eval)
|
||||
- User asks to debug or analyse failing eval results
|
||||
- User asks to review or compare eval output
|
||||
|
||||
## Required inputs
|
||||
|
||||
- Target skill name — explicit or unambiguous from session context
|
||||
- Target skill's SKILL.md — must be readable at `.agents/skills/<skill-name>/SKILL.md`
|
||||
- Target skill's `metadata.category` — used to derive the output path
|
||||
|
||||
## Constraints
|
||||
|
||||
- Output path: `.agents/evals/<category>/<skill-name>/eval.yaml` — nested by category, not flat
|
||||
- Every eval.yaml must contain all five required test types: ≥1 explicit trigger, ≥1 implicit trigger, ≥1 negative trigger, ≥2 deterministic output, ≥1 LLM-rubric quality
|
||||
- Assertions must be specific and verifiable — "The output contains a trigger_tests section" not "The output is good"
|
||||
- Assertions must be provider-agnostic — no tool-call assertions, no assumptions about the underlying model or runtime
|
||||
- Show the test plan and wait for confirmation before writing any file
|
||||
- On re-run (eval.yaml already exists): merge — classify proposed cases as NEW / IDENTICAL / CONFLICT; surface conflicts for human resolution before writing; do not silently overwrite
|
||||
- Body ≤500 lines
|
||||
|
||||
## Process
|
||||
|
||||
1. **Identify the target skill.** If not explicit in the invocation, infer from session context. If ambiguous, ask before proceeding.
|
||||
|
||||
2. **Read the target SKILL.md** at `.agents/skills/<skill-name>/SKILL.md`. Extract:
|
||||
- `name`, `metadata.category` (for output path)
|
||||
- `description` (trigger description — source for explicit and implicit trigger test queries)
|
||||
- When / when not criteria (source for negative trigger test queries)
|
||||
- Required inputs and output format (source for deterministic output assertions)
|
||||
|
||||
3. **Check for an existing eval.yaml** at `.agents/evals/<category>/<skill-name>/eval.yaml`.
|
||||
- If it exists: read it and record all existing test IDs.
|
||||
|
||||
4. **Propose test cases** — one minimum per required type:
|
||||
|
||||
**trigger_tests** — classify each query by whether the skill should activate:
|
||||
- ≥1 explicit trigger: a query using the skill's exact trigger phrase
|
||||
- ≥1 implicit trigger: a query describing the task without the trigger phrase; derive from the skill's purpose and use cases
|
||||
- ≥1 negative trigger: a query for an adjacent task the skill must NOT activate on; derive from the skill's when-not criteria; choose a case with surface similarity to the trigger
|
||||
|
||||
**output_tests** — test what the skill produces:
|
||||
- ≥2 deterministic: assert on observable, machine-checkable properties of the output — required sections present, correct file path, schema compliance. Write as specific string conditions a reader could verify without inference.
|
||||
- ≥1 LLM-rubric: holistic quality assertions — conditions a judge evaluates from the full output. Test qualities that deterministic checks cannot capture: realism of trigger queries, specificity of assertions, boundary case coverage.
|
||||
|
||||
For all assertions: write as verifiable conditions, not value judgements. Test boundary cases, not only happy paths. A good assertion survives internal refactoring of the skill.
|
||||
|
||||
5. **Classify proposed cases if an existing eval.yaml was found:**
|
||||
- **NEW** — ID not in existing file; safe to append
|
||||
- **IDENTICAL** — ID exists, content matches exactly; skip silently
|
||||
- **CONFLICT** — ID exists, content differs; display existing vs proposed side-by-side
|
||||
|
||||
6. **Present the full test plan.** Show each proposed case with its classification label (NEW / IDENTICAL / CONFLICT). For CONFLICT cases, ask the user to choose: keep existing, use proposed, or skip. Wait for confirmation before writing.
|
||||
|
||||
7. **Write eval.yaml.** Append NEW cases to the existing file (or write the full structure for a new file). Apply CONFLICT resolutions as chosen. Skip IDENTICAL cases.
|
||||
|
||||
## Output format
|
||||
|
||||
```yaml
|
||||
skill_name: <name>
|
||||
|
||||
trigger_tests:
|
||||
- id: <string-slug> # e.g. explicit-trigger-basic
|
||||
name: <display label> # human-readable, e.g. "Explicit trigger — basic invocation"
|
||||
query: <exact user input text>
|
||||
should_trigger: true # true for explicit and implicit; false for negative
|
||||
|
||||
output_tests:
|
||||
- id: <string-slug>
|
||||
name: <display label>
|
||||
type: deterministic # or llm-rubric
|
||||
prompt: <user input to the skill>
|
||||
expected_output: <prose description of ideal output>
|
||||
assertions:
|
||||
- <specific, verifiable condition string>
|
||||
```
|
||||
|
||||
## Failure handling
|
||||
|
||||
- **Target SKILL.md not found:** stop, report the path searched, do not guess or generate content from the skill name alone
|
||||
- **`metadata.category` absent from SKILL.md:** ask for the category before computing the output path
|
||||
- **All proposed cases conflict with existing file:** report the full conflict summary, wait for explicit direction — do not auto-resolve
|
||||
- **Proposed test count below minimums:** flag which type is short before presenting the plan; do not proceed with a deficient eval
|
||||
|
||||
## Self-check
|
||||
|
||||
Verify before writing:
|
||||
|
||||
- [ ] All five test types present — ≥1 explicit, ≥1 implicit, ≥1 negative trigger; ≥2 deterministic, ≥1 LLM-rubric output
|
||||
- [ ] trigger_tests: at least one `should_trigger: true` and at least one `should_trigger: false`
|
||||
- [ ] All assertions are specific and verifiable — no vague quality claims
|
||||
- [ ] Output path matches `.agents/evals/<category>/<skill-name>/eval.yaml`
|
||||
- [ ] Test plan was presented and confirmed before the file was written
|
||||
- [ ] CONFLICT cases were surfaced to the user and not silently resolved
|
||||
@@ -34,6 +34,10 @@ Behavioral tests (2026-05-17) showed three communication/behavior rules failing:
|
||||
|
||||
The secrets prohibition in `core/instructions/governance.md` fired correctly when asked to write a password to a file, but the agent then reproduced the literal credential in its response text (in a shell `export` example). The rule was interpreted as "don't write to files" not "don't output at all." Fix: the rule needs to explicitly state "never produce the credential value in any output" and give an example showing placeholder usage (`export DB_PASSWORD='<your-password>'`).
|
||||
|
||||
## 2026-05-17 — Synthesis grill and SKILL.md co-write are two separate conversations
|
||||
|
||||
The synthesis grill (step 4) answers schema-level questions: how to combine upstreams, which eval schema to use, merge behaviour. Step 5b is a different conversation: how upstream content maps to each SKILL.md body section, what options each section had, and which was chosen. Collapsing them — writing the SKILL.md immediately after the grill without a per-section walk-through — means the human never sees the upstream options for the body and has no opportunity to redirect before the file is written. Fix: step 5b is now a named gate in the workflow. Walk through every body section one at a time, cite the upstream source, present alternatives, get confirmation. Only then write. Applies to both hand-written (bootstrap) and write-skill-produced skills.
|
||||
|
||||
## 2026-05-17 — HITL gap: agent delegates confirmation to permission system
|
||||
|
||||
The agent-level HITL rule ("require explicit confirmation before irreversible shared-state operations") is being bypassed: the agent calls the tool and lets the permission dialog catch it. This means the rule is not firing in agent reasoning — it's the permission system acting as a safety net. If a user selects "don't ask again," the net disappears. Fix: the HITL rule needs to be framed as "do not call the tool" rather than "ask before proceeding" — the agent must ask first, then act only after explicit confirmation.
|
||||
|
||||
@@ -94,8 +94,8 @@ Items consciously not resolved — to be addressed in the relevant chunk PRD or
|
||||
- **AI coding factory integration** — grill complete. Decision record: `docs/notes/factory-integration-decisions.md`. ADRs: 0008 (factory boundary), 0009 (flat taxonomy), 0010 (role skills vs subagents). Follow-on issues: ~~0013 (LESSONS.md)~~ ✅, ~~0014 (docs/spec/ + VISION.md refactor)~~ ✅. Chunk 3 scope substantially expanded — skills rebuild, new skills, IaC/Gitea skills. See updated chunk table above.
|
||||
|
||||
- **`.gitkeep` files** — placeholder files exist in `core/agents/`, `core/workflows/`, `core/prompts/`, `docs/ard/`, `docs/bug/`. Remove each when the first real file is added to that directory. Each `.gitkeep` names the chunk that will populate it. (`docs/notes/.gitkeep` already removed — directory has real content.)
|
||||
- **Skills pipeline verified** — `install.sh` deploys 12 skills to `~/.agents/skills/` and creates `~/.claude/skills/ → ~/.agents/skills/` symlink adapter. Tested idempotent. `skills-lock.json` removed (was a manual artifact). If `~/.claude/skills/` exists as a real directory on a machine being migrated, remove it manually and re-run install.
|
||||
- **Skills pipeline verified** — `install.sh` deploys 13 skills to `~/.agents/skills/` and creates `~/.claude/skills/ → ~/.agents/skills/` symlink adapter. Tested idempotent. `skills-lock.json` removed (was a manual artifact). If `~/.claude/skills/` exists as a real directory on a machine being migrated, remove it manually and re-run install.
|
||||
- **Chunk 2 behavioral tests** — run and fully resolved 2026-05-17. 7/8 pass; scenario 4 (push confirmation) inconclusive — no remote in test environment, rule tightened but unverified. All fixable failures addressed: rule specificity in `providers/claude-code/CLAUDE.md`; context-loading guarantee via `@import CONTEXT.md` in repo CLAUDE.md; standing rule in CONTEXT.md to check `docs/adr/` and ROADMAP resolved entries before answering design questions. Chunk 2 ✅ complete.
|
||||
- **Governance Phase 1 behavioral tests** — run 2026-05-17. 3/4 testable scenarios pass. Secrets rule gap fixed (2026-05-17): extended to cover credential reproduction in response text and examples, with placeholder requirement added to `core/instructions/governance.md`. HITL scenario not testable in this environment (Nginx not installed); HITL gap evidenced by instructions test scenario 4 — push confirmation rule fix addresses the same root cause. Governance Phase 1 ✅ complete.
|
||||
- **AI ethics/security workstream** — `docs/notes/ai-ethics-security-principles.md` exploration note is superseded. Governance Phase 1 (`core/instructions/governance.md`) covers all planned scope: credentials, data classification, HITL, scope discipline, agent autonomy, transparency, and security code review. Tier-placement architectural question resolved by the `@import` always-on model. No separate workstream needed.
|
||||
- **Chunk 3 grill complete** — 2026-05-17. PRD at `docs/prd/chunk-3-skills-library.md`. Key decisions: 42-skill target library, AGENTS.md refactor as prerequisite issue (both CLAUDE.md files become thin adapters), git-cliff for changelog, provider-agnostic issue tracker abstraction, grill-me/grill-lean design phase split, factory bootstrap order (write-eval → write-skill → write-docs phase 2 → write-adr → remaining factory → design → parallel category groups). ADRs written: 0011 (provider-agnostic issue tracker), 0012 (AGENTS.md governance entry point, partially supersedes ADR-0005). Upstream review cadence: per-skill + quarterly post-roadmap (per-chunk-start changed to per-skill by issue 0016 grill). **Issues created 0015–0028** — all HITL; ~~0015 (AGENTS.md refactor, prerequisite)~~ ✅, ~~0016 (skill workflow grill, produces conventions for 0017–0028)~~ ✅, 0017–0018 (bootstrap skills: write-eval; write-skill + write-docs as phase 2), 0019 (remaining factory skills), 0020–0027 (design/implement/test/review/deploy/operate/iac/cross-cutting), 0028 (chunk closure). ~~Acceptance criteria for 0017–0028 to be refined after 0016 grill session.~~ ✅ Refined 2026-05-17 — see `docs/notes/skill-implementation-workflow.md`.
|
||||
- **Chunk 3 grill complete** — 2026-05-17. PRD at `docs/prd/chunk-3-skills-library.md`. Key decisions: 42-skill target library, AGENTS.md refactor as prerequisite issue (both CLAUDE.md files become thin adapters), git-cliff for changelog, provider-agnostic issue tracker abstraction, grill-me/grill-lean design phase split, factory bootstrap order (write-eval → write-skill → write-docs phase 2 → write-adr → remaining factory → design → parallel category groups). ADRs written: 0011 (provider-agnostic issue tracker), 0012 (AGENTS.md governance entry point, partially supersedes ADR-0005). Upstream review cadence: per-skill + quarterly post-roadmap (per-chunk-start changed to per-skill by issue 0016 grill). **Issues created 0015–0028** — all HITL; ~~0015 (AGENTS.md refactor, prerequisite)~~ ✅, ~~0016 (skill workflow grill, produces conventions for 0017–0028)~~ ✅, ~~0017 (bootstrap skill: write-eval)~~ ⏳ HITL pending, 0018 (write-skill + write-docs as phase 2), 0019 (remaining factory skills), 0020–0027 (design/implement/test/review/deploy/operate/iac/cross-cutting), 0028 (chunk closure). ~~Acceptance criteria for 0017–0028 to be refined after 0016 grill session.~~ ✅ Refined 2026-05-17 — see `docs/notes/skill-implementation-workflow.md`.
|
||||
|
||||
@@ -31,21 +31,48 @@ write-eval has no direct Pocock equivalent. Expect to synthesize from multiple u
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- [ ] `.agents/skills/write-eval/SKILL.md` exists; `metadata.category: factory`; authoring standard met (frontmatter, role, when/when-not, required inputs, constraints, process, output format, failure handling)
|
||||
- [ ] Trigger description matches index or deviation is documented in SKILL.md with justification
|
||||
- [ ] `.agents/evals/factory/write-eval/eval.yaml` exists; hand-written; contains all 5 required test types
|
||||
- [ ] `install.sh` deploys `write-eval` to `~/.agents/skills/` (confirm idempotent re-run)
|
||||
- [x] `.agents/skills/write-eval/SKILL.md` exists; `metadata.category: factory`; authoring standard met (frontmatter, role, when/when-not, required inputs, constraints, process, output format, failure handling)
|
||||
- [x] Trigger description matches index or deviation is documented in SKILL.md with justification
|
||||
- [x] `.agents/evals/factory/write-eval/eval.yaml` exists; hand-written; contains all 5 required test types
|
||||
- [x] `install.sh` deploys `write-eval` to `~/.agents/skills/` (confirm idempotent re-run)
|
||||
- [ ] **HITL:** human runs fresh-session behavioral test: invoke "write evals for this skill" and verify correct eval.yaml structure is produced
|
||||
- [ ] **HITL:** human reviews hand-written eval.yaml for correctness before committing
|
||||
- [ ] Per-skill process followed: source discovery (sub-agent) → source review with licence/security check (sub-agent) → conflict check against constitution + factory principles (sub-agent) → synthesis grill → co-write iteratively
|
||||
- [ ] Trigger description tested against explicit, implicit, and negative queries before body was written
|
||||
- [ ] `when:` frontmatter field present
|
||||
- [ ] `source:` field present only if upstream content adopted; absent if self-authored
|
||||
- [ ] `references:` field present if external citations used; absent otherwise
|
||||
- [ ] eval.yaml contains all 5 required test types: explicit trigger, implicit trigger, negative trigger, ≥2 deterministic output, ≥1 LLM-rubric quality
|
||||
- [ ] Body ≤500 lines; XML tags used only if ≥3 logical sections and 500+ tokens
|
||||
- [ ] `docs/spec/overview.md` updated to reflect `write-eval` deployed
|
||||
- [x] Per-skill process followed: source discovery (sub-agent) → source review with licence/security check (sub-agent) → conflict check against constitution + factory principles (sub-agent) → synthesis grill → co-write iteratively
|
||||
- [x] Trigger description tested against explicit, implicit, and negative queries before body was written
|
||||
- [x] `when:` frontmatter field present
|
||||
- [x] `source:` field present only if upstream content adopted; absent if self-authored
|
||||
- [x] `references:` field present if external citations used; absent otherwise
|
||||
- [x] eval.yaml contains all 5 required test types: explicit trigger, implicit trigger, negative trigger, ≥2 deterministic output, ≥1 LLM-rubric quality
|
||||
- [x] Body ≤500 lines; XML tags used only if ≥3 logical sections and 500+ tokens
|
||||
- [x] `docs/spec/overview.md` updated to reflect `write-eval` deployed
|
||||
|
||||
## Blocked by
|
||||
|
||||
- 0016 (grill defines the per-skill implementation workflow this issue must follow)
|
||||
|
||||
## Handoff
|
||||
|
||||
**Status:** complete — pending HITL behavioral test (acceptance criteria step 5)
|
||||
|
||||
**Files produced:**
|
||||
- `.agents/skills/write-eval/SKILL.md`
|
||||
- `.agents/evals/factory/write-eval/eval.yaml`
|
||||
|
||||
**Key decisions:**
|
||||
- Two-section schema: `trigger_tests` (explicit/implicit/negative, `should_trigger: bool`) + `output_tests` (deterministic/llm-rubric, `type:` field). Sources: BMAD-METHOD `triggers.json` split + darkrishabh `types.ts`.
|
||||
- Provider-agnostic string assertions — no tool-call assertions. Portable across runtimes.
|
||||
- Show plan before writing; merge on re-run with conflict flagging (option B): NEW / IDENTICAL / CONFLICT classification; CONFLICT cases shown side-by-side, human resolves before write.
|
||||
- Iteration loop (run evals → propose edits → apply) is out of scope — belongs to a future runner skill.
|
||||
- `id` as string slug (not integer); `name` field as separate display label.
|
||||
|
||||
**Workflow fix recorded:**
|
||||
- `docs/notes/skill-implementation-workflow.md` step 5b updated: per-section options walk-through is now a named gate before writing. Synthesis grill answers schema questions; step 5b covers how upstream content maps to each SKILL.md section — these are separate conversations.
|
||||
- `LESSONS.md` entry added: "Synthesis grill and SKILL.md co-write are two separate conversations."
|
||||
|
||||
**Open threads:**
|
||||
- HITL behavioral test: open a fresh Claude session, invoke "write evals for this skill" in this repo context, verify correct eval.yaml structure is produced at the right path with all 5 types.
|
||||
- `write-eval`'s own eval.yaml is hand-written (bootstrap). Once `write-eval` is behaviorally verified, it can be used to regenerate its own eval — a useful dogfood test.
|
||||
|
||||
**Next session start:**
|
||||
- Load: `CONTEXT.md`, `docs/notes/skill-implementation-workflow.md`, `docs/issues/0018-factory-write-skill.md`
|
||||
- First action: Step 1 (source discovery) for `write-skill`
|
||||
|
||||
@@ -92,8 +92,16 @@ Write the `description:` frontmatter field first. Test it against three cases be
|
||||
|
||||
Do not proceed to the body until all three pass.
|
||||
|
||||
**b. SKILL.md** (sub-agent)
|
||||
Spawn a write agent to produce the SKILL.md using `write-skill` (or hand-write for bootstrap skills). The agent receives: trigger description, synthesis grill decisions, upstream content to incorporate, authoring standard (see below).
|
||||
**b. Per-section options walk-through**
|
||||
Before writing anything, walk through each body section with the human. For each section:
|
||||
- State what content is proposed and which upstream source it comes from
|
||||
- Present alternatives where upstream sources offered different approaches
|
||||
- Get explicit confirmation (or redirection) before moving to the next section
|
||||
|
||||
Do not write the SKILL.md until the human has confirmed every section. The synthesis grill decisions cover the eval schema and gating questions; this step covers how upstream content maps to each SKILL.md section. These are separate conversations — do not collapse them.
|
||||
|
||||
**c. SKILL.md** (sub-agent)
|
||||
Once all sections are confirmed, spawn a write agent to produce the SKILL.md using `write-skill` (or hand-write for bootstrap skills). The agent receives: trigger description, per-section decisions from step b, upstream content to incorporate, authoring standard (see below).
|
||||
|
||||
**c. `source:` and `references:` fields**
|
||||
Populate after upstream review. Two distinct fields:
|
||||
|
||||
@@ -2,12 +2,14 @@
|
||||
|
||||
Current deployed state of this repo — what you get if you run `install.sh` today. Updated at the close of each chunk and in the same PR as any behavior change.
|
||||
|
||||
*Last updated: 2026-05-17 (issue 0015)*
|
||||
*Last updated: 2026-05-17 (issue 0017)*
|
||||
|
||||
## What is deployed
|
||||
|
||||
### Skills
|
||||
12 skills deployed to `~/.agents/skills/` via `install.sh`. Available as slash commands in Claude Code via `~/.claude/skills/ → ~/.agents/skills/` symlink. All 12 are first-draft placeholders pending rebuild in Chunk 3.
|
||||
13 skills deployed to `~/.agents/skills/` via `install.sh`. Available as slash commands in Claude Code via `~/.claude/skills/ → ~/.agents/skills/` symlink. 12 are first-draft placeholders pending rebuild in Chunk 3; 1 is a new Chunk 3 factory skill (`write-eval`).
|
||||
|
||||
**Factory bootstrap (Chunk 3):** `write-eval` deployed — produces `eval.yaml` test files for skills. Hand-written (bootstrap skill). Eval at `.agents/evals/factory/write-eval/eval.yaml`.
|
||||
|
||||
Current skills: `caveman`, `diagnose`, `grill-me`, `grill-with-docs`, `improve-codebase-architecture`, `prototype`, `tdd`, `to-issues`, `to-prd`, `triage`, `write-a-skill`, `zoom-out`.
|
||||
|
||||
@@ -42,6 +44,7 @@ For chunk planning and open questions, see `docs/ROADMAP.md`.
|
||||
|
||||
## Recent changes
|
||||
|
||||
- 2026-05-17 — Issue 0017 complete: `write-eval` bootstrap skill written and deployed. Two sections schema (`trigger_tests` + `output_tests`), provider-agnostic string assertions, show-plan-then-merge-on-rerun behaviour, conflict flagging (B model). Sources: agentskills/agentskills, darkrishabh/agent-skills-eval, bmad-code-org/BMAD-METHOD, mattpocock/skills. Hand-written eval at `.agents/evals/factory/write-eval/eval.yaml`.
|
||||
- 2026-05-17 — Issue 0016 complete: skill implementation workflow grill completed. `docs/notes/skill-implementation-workflow.md` written. All issues 0017–0028 updated with specific acceptance criteria. Key conventions: sub-agents prescribed at each research/writing step; conflict check against constitution + factory principles before synthesis grill; `when:` and `references:` fields added to authoring standard; write-docs moved to issue 0018 phase 2 (first factory-authored skill).
|
||||
- 2026-05-17 — Issue 0015 complete: AGENTS.md refactor implemented. Two AGENTS.md files created (`AGENTS.md` at repo root, `core/AGENTS.md` deployed to `~/.agents/AGENTS.md`). Both CLAUDE.md files slimmed to thin adapters. `deploy-manifest.sh` updated. `docs/spec/architecture.md` updated with new structure. ADR-0012 in effect.
|
||||
- 2026-05-17 — Chunk 3 issues created (0015–0028): AGENTS.md refactor prerequisite, skill workflow grill, bootstrap skills (write-eval, write-skill), factory/design/implement/test/review/deploy/operate/IaC/cross-cutting skill groups, chunk closure; all HITL; acceptance criteria for 0017–0028 to be refined after issue 0016 grill session
|
||||
|
||||
Reference in New Issue
Block a user