Files
holocron/docs/issues/0025-operate-skills.md
Defame1297 9457efff36 docs: Chunk 3 issues 0015-0028 and doc updates
14 issues created covering AGENTS.md refactor prerequisite, skill
implementation workflow grill, bootstrap skills (write-eval, write-skill),
remaining factory skills, and one issue per skill category group through
to chunk closure. All HITL; acceptance criteria for 0017-0028 to be
refined after 0016 grill session.

Doc updates: CONTEXT.md PRD/issue scope clarified (HOW distribution
across architecture-review and issue design notes); ROADMAP.md housekeeping
updated with bootstrap order and issue range; spec/overview.md recent
changes entry added; PRD updated (skills-index: delete → update).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-17 15:18:36 +00:00

2.6 KiB

0025 — Operate skills: write-runbook, incident-diagnosis, post-mortem, inspect-deployment

Type: HITL
Parent PRD: docs/prd/chunk-3-skills-library.md

What to build

The 4 operate phase skills. All are new. Authored via write-skill (0018), evals via write-eval (0017).

Skills and trigger descriptions:

Flat name Trigger description
write-runbook Write runbook, operational guide, on-call playbook
incident-diagnosis Diagnose this incident, analyse these logs, root cause analysis
post-mortem Write post-mortem, incident review, after-action report
inspect-deployment Check deployment health, container status, what's running

Key constraints per skill:

  • write-runbook: covers common failure modes, detection steps, remediation steps, and escalation path; written for on-call engineers under pressure
  • incident-diagnosis: produces structured finding with confidence levels; never recommends production remediation directly — diagnosis only, human approves remediation
  • post-mortem: blameless format; covers timeline, root cause analysis, and governance change (what process/rule changes prevent recurrence)
  • inspect-deployment: read-only; uses Docker MCP and/or K8s MCP when configured; summarises health without modifying state

Implementation notes

Follow the per-skill workflow defined in docs/notes/skill-implementation-workflow.md (produced by issue 0016).

Known upstream sources to review:

  • bmad-method/bmad-method — BMAD ops role patterns
  • Google SRE book patterns for blameless post-mortem and runbook formats (public domain principles)
  • Search agentskills.io and GitHub for open-source ops/operate skill implementations

Acceptance criteria

  • All 4 SKILL.md files exist at .agents/skills/<skill-name>/SKILL.md; metadata.category: operate; authoring standard met
  • incident-diagnosis explicitly states it produces diagnosis only and does not recommend production remediation
  • post-mortem uses blameless format
  • inspect-deployment is read-only; uses MCP when available
  • source: fields populated for any adopted upstream content
  • Each skill has a co-located eval at .agents/evals/operate/<skill-name>/eval.yaml via write-eval
  • install.sh deploys all 4 to ~/.agents/skills/
  • HITL: human runs behavioral test per skill
  • HITL: human reviews each SKILL.md and eval before committing
  • (Further criteria to be refined after issue 0016 grill session)

Blocked by

  • 0016 (per-skill workflow)
  • 0017 (write-eval)
  • 0018 (write-skill)