Replaces shell script (.git/hooks/pre-commit.legacy) with ecosystem-managed pre-commit framework: - gitleaks/gitleaks: secret scanning - jumanjihouse/pre-commit-hooks: shellcheck wrapper - pre-commit/pre-commit-hooks: JSON/YAML validation, end-of-file-fixer, trailing-whitespace - local hooks: SKILL.md frontmatter validation Uses pinned versions for reproducibility across environments. Includes auto-fixes from hook runs (formatting, trailing whitespace, JSON beautification). Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
58 lines
3.1 KiB
Markdown
58 lines
3.1 KiB
Markdown
# 0025 — Operate skills: write-runbook, incident-diagnosis, post-mortem, inspect-deployment
|
|
|
|
**Type:** HITL
|
|
**Parent PRD:** `docs/prd/chunk-3-skills-library.md`
|
|
|
|
## What to build
|
|
|
|
The 4 operate phase skills. All are new. Authored via `write-skill` (0018), evals via `write-eval` (0017).
|
|
|
|
**Skills and trigger descriptions:**
|
|
|
|
| Flat name | Trigger description |
|
|
|---|---|
|
|
| `write-runbook` | Write runbook, operational guide, on-call playbook |
|
|
| `incident-diagnosis` | Diagnose this incident, analyse these logs, root cause analysis |
|
|
| `post-mortem` | Write post-mortem, incident review, after-action report |
|
|
| `inspect-deployment` | Check deployment health, container status, what's running |
|
|
|
|
**Key constraints per skill:**
|
|
- `write-runbook`: covers common failure modes, detection steps, remediation steps, and escalation path; written for on-call engineers under pressure
|
|
- `incident-diagnosis`: produces structured finding with confidence levels; never recommends production remediation directly — diagnosis only, human approves remediation
|
|
- `post-mortem`: blameless format; covers timeline, root cause analysis, and governance change (what process/rule changes prevent recurrence)
|
|
- `inspect-deployment`: read-only; uses Docker MCP and/or K8s MCP when configured; summarises health without modifying state
|
|
|
|
## Implementation notes
|
|
|
|
Follow the per-skill workflow defined in `docs/notes/skill-implementation-workflow.md`.
|
|
|
|
**Known upstream sources to review:**
|
|
- `bmad-method/bmad-method` — BMAD ops role patterns
|
|
- Google SRE book patterns for blameless post-mortem and runbook formats (public domain principles)
|
|
- Search agentskills.io and GitHub for open-source ops/operate skill implementations
|
|
|
|
## Acceptance criteria
|
|
|
|
- [ ] All 4 SKILL.md files exist at `.agents/skills/<skill-name>/SKILL.md`; `metadata.category: operate`; authoring standard met
|
|
- [ ] `incident-diagnosis` explicitly states it produces diagnosis only and does not recommend production remediation
|
|
- [ ] `post-mortem` uses blameless format
|
|
- [ ] `inspect-deployment` is read-only; uses MCP when available
|
|
- [ ] `source:` fields populated for any adopted upstream content
|
|
- [ ] Each skill has a co-located eval at `.agents/evals/operate/<skill-name>/eval.yaml` via `write-eval`
|
|
- [ ] `install.sh` deploys all 4 to `~/.agents/skills/`
|
|
- [ ] **HITL:** human runs behavioral test per skill
|
|
- [ ] **HITL:** human reviews each SKILL.md and eval before committing
|
|
- [ ] Per-skill process followed for all 4 skills: source discovery (sub-agent) → source review with licence/security check (sub-agent) → conflict check against constitution + factory principles (sub-agent) → synthesis grill → co-write iteratively
|
|
- [ ] Trigger description for each skill tested against explicit, implicit, and negative queries before body written
|
|
- [ ] `when:` frontmatter field present in all SKILL.md files
|
|
- [ ] `source:` and `references:` fields correctly populated or absent
|
|
- [ ] eval.yaml for each skill contains all 5 required test types
|
|
- [ ] Body ≤500 lines for each skill
|
|
- [ ] `docs/spec/overview.md` updated to reflect all 4 skills deployed
|
|
|
|
## Blocked by
|
|
|
|
- 0016 (per-skill workflow)
|
|
- 0017 (`write-eval`)
|
|
- 0018 (`write-skill`)
|