feat(factory-audit): audit hooks, instructions and prompts

factory-audit gains three Step 0 rows and flows for the apm primitives
that have no container of their own: a .json file under hooks/, a
*.instructions.md and a *.prompt.md. apm validates almost none of them
(invalid hook JSON is skipped silently, instruction validate() only
warns, input: names are never checked against ${input:x}), so the
deterministic checks live in a new scripts/lib-checks-primitive.sh,
wired into validate.sh's path-shape detection. Each check and tier
traces to the Authoring checklists in the microsoft-apm research docs.

- Hook: JSON/shape/event-list checks mirroring the Copilot payload
  validator, never-firing event casing, missing/escaping/non-executable
  scripts (FAIL); deprecated filename routing and ${CLAUDE_PLUGIN_ROOT}
  (SUGGESTION).
- Instruction: location, frontmatter, description, body, stem clash
  (FAIL); missing or list applyTo and unread keys (SUGGESTION).
- Prompt: location/name, frontmatter, description, input names, the
  upstream `- name: x` docs bug, declared-vs-used ${input:x} (FAIL);
  ADR-0029 description length and trigger clause, dropped keys,
  camelCase aliases, argument-hint with input (SUGGESTION). Whether a
  prompt carries procedure is judgment in prompt-flow.md, not a script
  heuristic.

Vale now lints *.instructions.md and *.prompt.md with the Kyberforge
style; test-vale-wrap.sh gains their probe rows. New
tests/validate-primitive.bats (31 cases). kyberforge 2.0.1 -> 2.1.0 with
the executables.allow key, catalog 0.5.1 -> 0.5.2, marketplace.json
regenerated.

Refs #94

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KkT7RSDwDbmrM9T34b6sTi
This commit is contained in:
2026-09-28 17:02:04 +00:00
parent 701e96d4b3
commit 70210d6a7e
14 changed files with 1062 additions and 29 deletions

View File

@@ -0,0 +1,54 @@
---
source_keys:
- apm-cli-installed-source
- apm-docs-llms-full
---
# Hook Flow
Steps 1 to 3 for an apm hook — the target Step 0 matched as a `.json` file directly under a
`hooks/` directory. Work them in order, then return to `SKILL.md` Step 4 to report.
## Gotchas
- apm checks almost nothing here. Invalid JSON is skipped without a word, an all-lowercase event deploys and never fires, and a missing script only warns — so `apm install` exiting 0 says nothing about whether the hook works. Never cite a clean install as evidence against a finding.
- Copilot receiving a Claude-shaped file is not a finding. apm renders one source for every target and documents that it owns the per-target shape; whether Copilot CLI honours a nested entry or `matcher` is unverified upstream, not a defect in the file.
## Step 1 — Deterministic checks
Resolve the path against this skill's own directory. Run exactly:
```bash
bash scripts/validate.sh <hook-file>
```
Its findings become the `### Structure` dimension, FAILs and SUGGESTIONs both, at the tier the script assigned: JSON validity, the wrapped-or-naked shape, event lists and nested handler lists (the checks whose failure makes the Copilot install fail), event names that never fire, referenced scripts that are missing, outside the package or not executable, deprecated filename routing, and `${CLAUDE_PLUGIN_ROOT}` where `${PLUGIN_ROOT}` would do. It exits **0** with no FAIL, **1** on real findings, **2** when it never ran — report that as `### Structure` unverified, quoting the stderr reason.
There is no provenance and no Vale step: a hook carries no `source_keys` and no prose.
## Step 2 — Read the hook and its scripts
Read the hook file, every script it references, and the package's `apm.yml` `targets:` — reach is narrowed there, never in the hook file.
## Step 3 — Qualitative audit
Cite file and line for every finding.
**purpose** — apm's own rule is to reach for a skill, instruction or prompt first; a hook is for "this must always happen at this event".
- FAIL: the script carries procedure the agent should follow — instructions printed to the model, a multi-step workflow — rather than a runtime callback. That is a skill.
- SUGGESTION: the behaviour is harness-specific (a Claude-only event, a Claude-only matcher value) in a package whose `targets:` includes other harnesses, and nothing records that the other targets receiving it was accepted. The apm-native fix is a separate package with its own `targets:`, not a routing filename.
**handlers** — the research checklist's Should and audit-only items, which apm never checks:
- SUGGESTION: a handler without `"type": "command"` or an explicit numeric `timeout` in seconds.
- SUGGESTION: a tool event (`PreToolUse`, `PostToolUse`) or `SessionStart` with no `matcher` — Claude receives `"*"`. A `matcher` on an event Claude ignores it for (`Stop`, `UserPromptSubmit`) is inert, not wrong.
- SUGGESTION: a PascalCase event name that is not a real Claude Code event (a misspelling deploys verbatim and never fires; the script cannot tell a typo from an event it does not know).
- SUGGESTION: `bash`/`powershell`/`timeoutSec` keys in a Claude-shaped file — they render, but leave stray keys in `settings.json`.
- SUGGESTION: an unquoted script path that could contain spaces.
Then return to `SKILL.md` Step 4, opening the report with this coverage line:
```text
Checked: structure · purpose · handlers
```

View File

@@ -0,0 +1,55 @@
---
source_keys:
- apm-cli-installed-source
- apm-docs-llms-full
---
# Instruction Flow
Steps 1 to 3 for an apm instruction — the target Step 0 matched as a `*.instructions.md` file.
Work them in order, then return to `SKILL.md` Step 4 to report.
## Gotchas
- `apm compile --validate` is not a gate. Every message `Instruction.validate()` produces is a warning, and it reports success on a file with no description and an empty body — never cite it as evidence against a finding.
- `description` never reaches Claude, and it is index text elsewhere, never a routing description. Do not hold it to the skill description contract: no trigger clause, no boundary clause. A Vale `Kyberforge.DescriptionOpener` alert here means rewrite it as a plain statement of what the rule covers.
## Step 1 — Deterministic checks
Resolve both paths against this skill's own directory. Run exactly:
```bash
bash scripts/validate.sh <instruction-file>
bash scripts/vale-wrap.sh <instruction-file>
```
`validate.sh` findings become the `### Structure` dimension, FAILs and SUGGESTIONs both, at the tier the script assigned. It exits **0** with no FAIL, **1** on real findings, **2** when it never ran — report that as `### Structure` unverified, quoting the stderr reason.
`vale-wrap.sh` applies the bundled `Kyberforge` style. Every alert is a FAIL under `### Prose`, cited by rule ID; do not re-derive it by judgment. `0 files` scanned means NOT RUN, not clean — say so and judge prose by reading.
There is no provenance step: an instruction carries no `source_keys`.
## Step 2 — Read the instruction and its context
Read the file, the package's `apm.yml`, and the repo's root `AGENTS.md`. For a scoped file, list the tracked files its `applyTo` matches (`rtk git ls-files` filtered by the glob).
## Step 3 — Qualitative audit
Cite file and line for every finding.
**scope** — an instruction applies when files matching `applyTo` are touched; with no `applyTo` it loads into every session of every repo that installs the package.
- FAIL: an always-on file whose content is a rule for this repo alone — it belongs in `AGENTS.md`, which is the repo's single always-on source, not in a package that ships it to every consumer.
- FAIL: an `applyTo` glob that matches no tracked file in any repo the package plausibly targets, so the rule never loads.
- SUGGESTION: an always-on file whose content is really file-type specific — narrow it with `applyTo`.
- SUGGESTION: a glob much broader than the content (`**` for a rule about Python).
**description**
- SUGGESTION: the description does not say what the rule covers, or contradicts the body. Any rationale Claude readers need belongs in the body, because Claude drops the description.
Then return to `SKILL.md` Step 4, opening the report with this coverage line:
```text
Checked: structure · prose · scope · description
```

View File

@@ -0,0 +1,55 @@
---
source_keys:
- apm-cli-installed-source
- apm-docs-llms-full
---
# Prompt Flow
Steps 1 to 3 for an apm prompt — the target Step 0 matched as a `*.prompt.md` file. Work them in
order, then return to `SKILL.md` Step 4 to report.
## Gotchas
- A prompt is judged against ADR-0029, not against apm's framing. apm calls a prompt "a callable program"; this repo holds it to a single-intent, user-triggered message that steers existing skills or agents by name and carries no procedure of its own.
- A prompt's description is not a skill description. It is one plain user-facing sentence with no "Use when" trigger clause and no boundary clause — so never raise a missing trigger or boundary as a finding. A Vale `Kyberforge.DescriptionOpener` alert here means rewrite it as an imperative action ("Review the current PR with …"), not add a trigger.
## Step 1 — Deterministic checks
Resolve both paths against this skill's own directory. Run exactly:
```bash
bash scripts/validate.sh <prompt-file>
bash scripts/vale-wrap.sh <prompt-file>
```
`validate.sh` findings become the `### Structure` dimension, FAILs and SUGGESTIONs both, at the tier the script assigned: path and name, frontmatter, `description` presence, length and trigger clause, keys Claude drops, `input:` names and shapes, and `${input:x}` references against `input:`. It exits **0** with no FAIL, **1** on real findings, **2** when it never ran — report that as `### Structure` unverified, quoting the stderr reason.
`vale-wrap.sh` applies the bundled `Kyberforge` style. Every alert is a FAIL under `### Prose`, cited by rule ID; do not re-derive it by judgment. `0 files` scanned means NOT RUN, not clean — say so and judge prose by reading.
There is no provenance step: a prompt carries no `source_keys`.
## Step 2 — Read the prompt and what it steers
Read the file end to end, then the description of every skill or agent its body names, and confirm each resolves in this repo or in a package the prompt's package declares.
## Step 3 — Qualitative audit
Cite file and line for every finding.
**role** — whether this is a prompt at all. Decide it by reading the body, not by its length or headings; there is no threshold.
- FAIL: the body clearly carries reusable procedure — steps, gotchas, domain know-how the agent could not act without — rather than steering skills or agents that hold it. Fix: move the procedure into a skill (new, or the one it belongs to) and reduce the prompt to the message that invokes it.
- FAIL: the body names a skill or agent that does not resolve, or one carrying `disable-model-invocation: true`, which the model cannot invoke.
- SUGGESTION: borderline — some how-to detail beyond steering, but not a full procedure.
- SUGGESTION: more than one intent in one prompt.
**description**
- SUGGESTION: the description does not read as one user-facing action, or does not name the skills the prompt steers. On Claude the description is model-visible and apm drops `disable-model-invocation`, so naming the steered skills keeps the router pointed at the capability rather than the wrapper.
Then return to `SKILL.md` Step 4, opening the report with this coverage line:
```text
Checked: structure · prose · role · description
```

View File

@@ -10,6 +10,8 @@ source_keys:
- claude-code-subagents-docs
- context7-github-en-copilot
- github-custom-agents-configuration
- apm-cli-installed-source
- apm-docs-llms-full
---
# Sources
@@ -151,3 +153,19 @@ source_keys:
- **Description:** SDK custom agent API — CustomAgentConfig fields in all five languages, session config, sub-agent lifecycle events, tool scoping, permission handling
- **Contributing files:** (none)
- **Status:** `extracted`
## apm-cli-installed-source
- **URL:** file:///root/.local/pipx/venvs/apm-cli/lib/python3.11/site-packages/apm_cli/
- **Research doc:** plugins/kyberforge/docs/research/docs/microsoft-apm/sources.md
- **Description:** Installed apm-cli 0.28.0 source — ground truth for what apm deploys from a hook, instruction or prompt file and what it silently skips, warns on, or fails the install for; every deterministic check in `scripts/lib-checks-primitive.sh` traces to it via the research docs' Authoring checklists
- **Contributing files:** SKILL.md, references/hook-flow.md, references/instruction-flow.md, references/prompt-flow.md
- **Status:** `extracted`
## apm-docs-llms-full
- **URL:** https://microsoft.github.io/apm/llms-full.txt
- **Research doc:** plugins/kyberforge/docs/research/docs/microsoft-apm/sources.md
- **Description:** Published apm docs bundle — the "Hooks and commands", "Instructions and agents" and "Author a prompt" guides: canonical hook shape and `${PLUGIN_ROOT}`, reach narrowed by `targets:` rather than filename routing, and "reach for a skill, instruction, or prompt first"
- **Contributing files:** SKILL.md, references/hook-flow.md, references/instruction-flow.md, references/prompt-flow.md
- **Status:** `extracted`