docs: Chunk 3 issues 0015-0028 and doc updates

14 issues created covering AGENTS.md refactor prerequisite, skill
implementation workflow grill, bootstrap skills (write-eval, write-skill),
remaining factory skills, and one issue per skill category group through
to chunk closure. All HITL; acceptance criteria for 0017-0028 to be
refined after 0016 grill session.

Doc updates: CONTEXT.md PRD/issue scope clarified (HOW distribution
across architecture-review and issue design notes); ROADMAP.md housekeeping
updated with bootstrap order and issue range; spec/overview.md recent
changes entry added; PRD updated (skills-index: delete → update).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
2026-05-17 15:18:36 +00:00
parent 537c681bb9
commit 9457efff36
18 changed files with 681 additions and 4 deletions

View File

@@ -0,0 +1,51 @@
# 0025 — Operate skills: write-runbook, incident-diagnosis, post-mortem, inspect-deployment
**Type:** HITL
**Parent PRD:** `docs/prd/chunk-3-skills-library.md`
## What to build
The 4 operate phase skills. All are new. Authored via `write-skill` (0018), evals via `write-eval` (0017).
**Skills and trigger descriptions:**
| Flat name | Trigger description |
|---|---|
| `write-runbook` | Write runbook, operational guide, on-call playbook |
| `incident-diagnosis` | Diagnose this incident, analyse these logs, root cause analysis |
| `post-mortem` | Write post-mortem, incident review, after-action report |
| `inspect-deployment` | Check deployment health, container status, what's running |
**Key constraints per skill:**
- `write-runbook`: covers common failure modes, detection steps, remediation steps, and escalation path; written for on-call engineers under pressure
- `incident-diagnosis`: produces structured finding with confidence levels; never recommends production remediation directly — diagnosis only, human approves remediation
- `post-mortem`: blameless format; covers timeline, root cause analysis, and governance change (what process/rule changes prevent recurrence)
- `inspect-deployment`: read-only; uses Docker MCP and/or K8s MCP when configured; summarises health without modifying state
## Implementation notes
Follow the per-skill workflow defined in `docs/notes/skill-implementation-workflow.md` (produced by issue 0016).
**Known upstream sources to review:**
- `bmad-method/bmad-method` — BMAD ops role patterns
- Google SRE book patterns for blameless post-mortem and runbook formats (public domain principles)
- Search agentskills.io and GitHub for open-source ops/operate skill implementations
## Acceptance criteria
- [ ] All 4 SKILL.md files exist at `.agents/skills/<skill-name>/SKILL.md`; `metadata.category: operate`; authoring standard met
- [ ] `incident-diagnosis` explicitly states it produces diagnosis only and does not recommend production remediation
- [ ] `post-mortem` uses blameless format
- [ ] `inspect-deployment` is read-only; uses MCP when available
- [ ] `source:` fields populated for any adopted upstream content
- [ ] Each skill has a co-located eval at `.agents/evals/operate/<skill-name>/eval.yaml` via `write-eval`
- [ ] `install.sh` deploys all 4 to `~/.agents/skills/`
- [ ] **HITL:** human runs behavioral test per skill
- [ ] **HITL:** human reviews each SKILL.md and eval before committing
- [ ] _(Further criteria to be refined after issue 0016 grill session)_
## Blocked by
- 0016 (per-skill workflow)
- 0017 (`write-eval`)
- 0018 (`write-skill`)