Compare commits
176 Commits
v1.0.0
...
c96ca9ca0d
| Author | SHA1 | Date | |
|---|---|---|---|
| c96ca9ca0d | |||
| 061bb3d5b4 | |||
| d2480b8191 | |||
| 718c79af70 | |||
| 0dffff3c21 | |||
| afa7187d33 | |||
| e2e957efdd | |||
| 124ce6eaa9 | |||
| 1b01e25f3b | |||
| a622200868 | |||
| af80d27b9b | |||
| 14f327409e | |||
| e21a5fb24c | |||
| 568ca749f0 | |||
| a35f5e889e | |||
| cc2553d5f6 | |||
| 60d3005e67 | |||
| bde9f7fdd8 | |||
| f11b6455ce | |||
| ace2d66343 | |||
| 5f9f2b33b0 | |||
| c8a7c9ea87 | |||
| e647f14535 | |||
| 6cfc3577e2 | |||
| f5e4d0d082 | |||
| 198eafd790 | |||
| 629320b8fd | |||
| edcc57c0d6 | |||
| 9eb8bc7e48 | |||
| a712f2c186 | |||
| 2fb1036329 | |||
| 5d7c76d797 | |||
| c613927fb4 | |||
| c75e4ef4f3 | |||
| 4f49b2a249 | |||
| 058fb5b748 | |||
| 29eefe7f70 | |||
| 6ba29b696c | |||
| b7bec71b8f | |||
| 0f2bb242ad | |||
| a3e721e937 | |||
| 3811f5481b | |||
| 175ea89c0a | |||
| ed8c99efbd | |||
| a6eedacfd8 | |||
| 4f4b55b0be | |||
| af8b46cd57 | |||
| 09eea5e7ab | |||
| 2c6ce438b6 | |||
| ffaa3afb41 | |||
| 60be7b3232 | |||
|
|
598a7c326a | ||
| 0e91a3ae66 | |||
| 967d3ade25 | |||
| 4c9d2d7751 | |||
| 68e08c2413 | |||
| d42f6368fe | |||
| c68e864159 | |||
| c7ba3d2ccf | |||
| 4d336bbf35 | |||
| 36596598ef | |||
| b1ea14df3e | |||
| de84d1b677 | |||
| 65bac15257 | |||
| b0ef503485 | |||
| bd2bf667c5 | |||
| ba7cec7672 | |||
| 56cc173f65 | |||
| b93af30750 | |||
| b9c7762463 | |||
| 1929ffd2da | |||
| 123ece2fb3 | |||
| 9385c77ac7 | |||
| 54d7bd80ba | |||
| 75a13c82f6 | |||
| 79c9089122 | |||
| ede3f06689 | |||
| b0d6d08239 | |||
| f7cc27908c | |||
| e7ebc667b3 | |||
| 64ffb9f35a | |||
| d02765d595 | |||
| 311e7cd22c | |||
| 2540e50fcc | |||
| a85bdbed42 | |||
| b6e68e9a2b | |||
| 76075223c7 | |||
| 36ba7a18f8 | |||
| 4a5c3c0cff | |||
| 1c6eababb0 | |||
| 2e395a4efa | |||
| f9b919d7e3 | |||
| c9fe2e8ab2 | |||
| ae178a95a2 | |||
| b4f5881973 | |||
| 099cf5846c | |||
| 3bfdf58960 | |||
| dee56c506a | |||
| 2e8732a8e5 | |||
| c3ec5f2d3d | |||
| a4a075b07e | |||
| b8dc400365 | |||
| cf625229f7 | |||
| a155af6827 | |||
| c16ec2d45a | |||
| f4bb1cf4e5 | |||
| a700b3771c | |||
| aa15fc850c | |||
| 874bf06b18 | |||
| 430f46b8e8 | |||
| 7ba3d9cf1d | |||
| 52bbd62286 | |||
| af085ed057 | |||
| 3f1ee47f1e | |||
| 9e612fd183 | |||
| 0f0ac5821f | |||
| c442f7eb85 | |||
| 5a61b417c9 | |||
| 73393b9d01 | |||
| 49d21bcb4d | |||
| 55d956b298 | |||
| 013b913bd4 | |||
| bb9158da22 | |||
| 413a750819 | |||
| d4fa4b7153 | |||
| f6cf83c841 | |||
| e79497b3cf | |||
| 23cef3627a | |||
| fd70c8d65e | |||
| 07ea0aeb17 | |||
| 925f04acdb | |||
| c6490096da | |||
| 911daddbe2 | |||
| b0b1470f2c | |||
| 560154c727 | |||
| 2c731eb476 | |||
| 7c3c867e00 | |||
| 4003c6a273 | |||
| 9c140efa2e | |||
| bff9662c52 | |||
| a873e93050 | |||
| a8beff7d2c | |||
| c3a56d89f0 | |||
| b0936ad386 | |||
| 5f42f57106 | |||
| 6e77c11474 | |||
| 38f1ba4e03 | |||
| 7910b8b12c | |||
| 5e232503c4 | |||
| 50d5c30a3c | |||
| eada85db99 | |||
| 044b2d3f08 | |||
| 6f6b70781d | |||
| f037d49b5c | |||
| 96bc946030 | |||
| ffebdc6584 | |||
| dc2a41034e | |||
| 099bdec1b2 | |||
| 239ea41842 | |||
| 675ba40238 | |||
| 8cd5c79c0a | |||
| 922eff3960 | |||
| 0dd044a782 | |||
| 5e296bcfef | |||
|
|
7a5c50fecc | ||
| 591b9cccb8 | |||
| d6fd9b6770 | |||
| 92e7ff26aa | |||
| 394052ff66 | |||
| e16c3dc95f | |||
| 2305f1c315 | |||
| 0aa66fe65d | |||
| fc69553ba7 | |||
| c48c9f5490 | |||
| 0e421acdbb | |||
| 1d07d1a76b |
@@ -1,49 +1,54 @@
|
||||
{
|
||||
"description": "AI development skills for Claude Code and GitHub Copilot CLI \u2014 factory, design, implement, review, and cross-cutting workflows.",
|
||||
"name": "holocron",
|
||||
"description": "AI development skills for Claude Code and GitHub Copilot CLI — factory, design, implement, review, and cross-cutting workflows.",
|
||||
"version": "0.4.6",
|
||||
"owner": {
|
||||
"name": "Defame1297",
|
||||
"email": "defame1297@rkdr.net",
|
||||
"name": "Defame1297"
|
||||
"url": "https://git.dev.rkdr.net/Defame1297/"
|
||||
},
|
||||
"plugins": [
|
||||
{
|
||||
"description": "Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.",
|
||||
"name": "kyberforge",
|
||||
"description": "Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.",
|
||||
"version": "1.6.2",
|
||||
"category": "Developer Tools",
|
||||
"source": "./plugins/kyberforge"
|
||||
},
|
||||
{
|
||||
"description": "A place for things to be binned",
|
||||
"name": "bin",
|
||||
"description": "Skills for everyday AI-assisted development work that is not tied to a single tool, forge or language, and has not yet been split into a focused plugin.",
|
||||
"version": "1.1.7",
|
||||
"category": "Utilities",
|
||||
"source": "./plugins/bin"
|
||||
},
|
||||
{
|
||||
"description": "Skills for working with Git \u2014 conventional commits, branch management, pull requests, and feature flow.",
|
||||
"name": "git",
|
||||
"description": "Skills and agents for working with a local Git clone over the git wire protocol, and for authoring and running the pre-commit hooks that guard it.",
|
||||
"version": "1.3.7",
|
||||
"category": "Version Control",
|
||||
"source": "./plugins/git"
|
||||
},
|
||||
{
|
||||
"description": "Skills for managing Gitea repositories \u2014 issues, pull requests, milestones, releases, and wikis.",
|
||||
"name": "gitea",
|
||||
"description": "Skills and agents for working with a Gitea forge through its HTTP API — the forge's own objects, as distinct from the local git clone.",
|
||||
"version": "1.3.8",
|
||||
"category": "Version Control",
|
||||
"source": "./plugins/gitea"
|
||||
},
|
||||
{
|
||||
"description": "Cross-cutting utility skills for everyday AI-assisted coding \u2014 triage, diagnosis, architecture review, and session navigation.",
|
||||
"name": "core",
|
||||
"description": "Skills for authoring and auditing a repo's AGENTS.md and the provider adapter files that defer to it.",
|
||||
"version": "1.1.2",
|
||||
"category": "Productivity",
|
||||
"source": "./plugins/core"
|
||||
},
|
||||
{
|
||||
"description": "Skills for Real Engineers \u2014 planning, TDD, architecture, and debugging workflows from Matt Pocock's .claude directory.",
|
||||
"name": "mattpocock-skills",
|
||||
"source": {
|
||||
"repo": "mattpocock/skills",
|
||||
"source": "github"
|
||||
}
|
||||
},
|
||||
{
|
||||
"description": "Skills and agents for configuring and running linters.",
|
||||
"name": "lint",
|
||||
"description": "Skills and agents for configuring and running linters.",
|
||||
"version": "1.1.7",
|
||||
"category": "Developer Tools",
|
||||
"source": "./plugins/lint"
|
||||
}
|
||||
],
|
||||
"version": "0.3.1"
|
||||
]
|
||||
}
|
||||
|
||||
@@ -1,12 +1,16 @@
|
||||
{
|
||||
"enabledPlugins": {
|
||||
"bin@holocron": true,
|
||||
"core@holocron": true,
|
||||
"git@holocron": true,
|
||||
"gitea@holocron": true,
|
||||
"kyberforge@holocron": true
|
||||
},
|
||||
"hooks": {
|
||||
"PreToolUse": []
|
||||
"SessionStart": [
|
||||
{
|
||||
"matcher": "startup",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "\"${CLAUDE_PROJECT_DIR}/.claude/hooks/kyberforge/.apm/hooks/check-apm-current.sh\"",
|
||||
"timeout": 380
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
49
.github/plugin/marketplace.json
vendored
49
.github/plugin/marketplace.json
vendored
@@ -1,49 +0,0 @@
|
||||
{
|
||||
"description": "AI development skills for Claude Code and GitHub Copilot CLI \u2014 factory, design, implement, review, and cross-cutting workflows.",
|
||||
"name": "holocron",
|
||||
"owner": {
|
||||
"email": "defame1297@rkdr.net",
|
||||
"name": "Defame1297"
|
||||
},
|
||||
"plugins": [
|
||||
{
|
||||
"description": "Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.",
|
||||
"name": "kyberforge",
|
||||
"source": "./plugins/kyberforge"
|
||||
},
|
||||
{
|
||||
"description": "A place for things to be binned",
|
||||
"name": "bin",
|
||||
"source": "./plugins/bin"
|
||||
},
|
||||
{
|
||||
"description": "Skills for working with Git \u2014 conventional commits, branch management, pull requests, and feature flow.",
|
||||
"name": "git",
|
||||
"source": "./plugins/git"
|
||||
},
|
||||
{
|
||||
"description": "Skills for managing Gitea repositories \u2014 issues, pull requests, milestones, releases, and wikis.",
|
||||
"name": "gitea",
|
||||
"source": "./plugins/gitea"
|
||||
},
|
||||
{
|
||||
"description": "Cross-cutting utility skills for everyday AI-assisted coding \u2014 triage, diagnosis, architecture review, and session navigation.",
|
||||
"name": "core",
|
||||
"source": "./plugins/core"
|
||||
},
|
||||
{
|
||||
"description": "Skills for Real Engineers \u2014 planning, TDD, architecture, and debugging workflows from Matt Pocock's .claude directory.",
|
||||
"name": "mattpocock-skills",
|
||||
"source": {
|
||||
"repo": "mattpocock/skills",
|
||||
"source": "github"
|
||||
}
|
||||
},
|
||||
{
|
||||
"description": "Skills and agents for configuring and running linters.",
|
||||
"name": "lint",
|
||||
"source": "./plugins/lint"
|
||||
}
|
||||
],
|
||||
"version": "0.3.1"
|
||||
}
|
||||
33
.gitignore
vendored
33
.gitignore
vendored
@@ -24,3 +24,36 @@ node_modules/
|
||||
|
||||
# Claude Code local settings (machine-specific)
|
||||
.claude/settings.local.json
|
||||
|
||||
# APM dependencies
|
||||
apm_modules/
|
||||
|
||||
# APM install output — deployed copies of released plugin content, regenerated
|
||||
# by `apm install`. The authoring source is plugins/<name>/.apm/; committing a
|
||||
# deployed copy would add a third mirror of the same skills to drift against.
|
||||
.claude/skills/
|
||||
.claude/agents/
|
||||
|
||||
# APM MCP deployment output — `apm install` writes the repo-root .mcp.json from
|
||||
# the MCP servers its dependencies declare, and regenerates it on every install.
|
||||
# It is apm's output, not repo content; nothing here is hand-authored.
|
||||
/.mcp.json
|
||||
|
||||
# APM hook deployment output — `apm install` copies each package's referenced
|
||||
# hook scripts here and tracks its own settings.json entries in the sidecar.
|
||||
# Regenerated on every install; the authoring source is
|
||||
# plugins/<name>/.apm/hooks/ (ADR-0019).
|
||||
.claude/hooks/
|
||||
.claude/apm-hooks.json
|
||||
|
||||
# `apm pack` bundle output. The pre-push gate runs pack with --dry-run, so this
|
||||
# only appears after a bare `apm pack` during a release; it is not repo content.
|
||||
build/
|
||||
|
||||
# `apm pack`'s manifest for the *root* package. Emitted beside the marketplace
|
||||
# manifest by a bare `apm pack`, and never tracked on any branch — the repo's
|
||||
# own paths hide it, since the apm-pack-check-clean pre-push hook, the only
|
||||
# thing that runs pack here, passes --dry-run. Scoped
|
||||
# to the file, not the directory: the sibling .claude-plugin/marketplace.json is
|
||||
# compiled output that IS committed and must stay tracked.
|
||||
/.claude-plugin/plugin.json
|
||||
|
||||
@@ -28,6 +28,36 @@ repos:
|
||||
- id: pretty-format-json
|
||||
stages: ['pre-commit']
|
||||
args: [--autofix]
|
||||
# Every generated manifest lives at a KNOWN path, so every alternative is
|
||||
# root-anchored and spells that path out. This was five `(^|/)`
|
||||
# any-depth alternatives plus one `^` root-only one -- a mixture with no
|
||||
# rationale, under which a fixture or vendored tree containing
|
||||
# `.../.claude-plugin/marketplace.json` would have been silently excluded
|
||||
# from formatting while an equivalent
|
||||
# `.../.agents/plugins/marketplace.json` would not. Only the one root
|
||||
# marketplace manifest matches now; anything else is hand-authored and
|
||||
# gets formatted. The twelve per-plugin `plugin.json` alternatives were
|
||||
# dropped with the plugin manifests themselves when native
|
||||
# `claude plugin install` support was removed (ADR-0024) -- apm probes
|
||||
# `apm.yml` and never reached them. The `.agents/plugins/` and
|
||||
# `.github/plugin/` marketplace mirrors went the same way, and their
|
||||
# alternations went with them: `check-useless-excludes` fails on a
|
||||
# pattern that matches no file.
|
||||
#
|
||||
# `.claude/settings.json` is the second and last alternation, and it is
|
||||
# the only one here for a reason other than "generated manifest":
|
||||
# apm OWNS that file (ADR-0018, ADR-0019), and
|
||||
# `apm audit --ci` replays the install into a scratch tree and diffs
|
||||
# the result byte-for-byte. `pretty-format-json` sorts object keys
|
||||
# unless `--no-sort-keys` is passed, while apm's hook integrator emits
|
||||
# insertion order (`matcher` before `hooks`, `type` before `command`).
|
||||
# Formatting the file therefore rewrites apm's output into a form apm
|
||||
# would never produce, and the `apm-audit-ci` pre-push hook reports it
|
||||
# as permanent drift on a file with no git diff -- exactly what
|
||||
# happened when the SessionStart hook first landed in 2e395a4.
|
||||
# Re-running `apm install` fixes the file; leaving it in scope here
|
||||
# would re-break it on the very commit that carries the fix.
|
||||
exclude: '^(\.claude-plugin/marketplace\.json|\.claude/settings\.json)$'
|
||||
- id: check-yaml
|
||||
stages: ['pre-commit']
|
||||
- id: trailing-whitespace
|
||||
@@ -45,17 +75,89 @@ repos:
|
||||
hooks:
|
||||
- id: run-tests
|
||||
name: Run test suite
|
||||
description: Run all test-*.sh files and bats suite
|
||||
entry: bash tests/run-tests.sh
|
||||
description: Run all test-*.sh files and bats suite. --strict because a suite that exits 77 (SKIPPED) at pre-push means a documented dependency is missing on this machine, and pre-commit prints nothing for a passing hook -- without it the gate went green having verified 15 of 17 suites on a vale-less PATH, with the skip list swallowed. Ad-hoc `bash tests/run-tests.sh` still skips gracefully.
|
||||
entry: bash tests/run-tests.sh --strict
|
||||
language: system
|
||||
stages: [pre-push]
|
||||
pass_filenames: false
|
||||
always_run: true
|
||||
|
||||
- id: check-manifests
|
||||
name: Check plugin manifests
|
||||
description: Validate marketplace.json and plugin.json paths
|
||||
entry: bash scripts/check-manifests.sh
|
||||
- id: check-executables-allow-sync
|
||||
name: Check executables allow key sync
|
||||
description: Verify root apm.yml's executables.allow key names kyberforge's actual version -- apm matches that key by exact "<package>#<version>" lookup, so a version bump on one side alone silently stops deploying kyberforge's hooks/ and bin/ and lets the apm install go stale (see ADR-0019)
|
||||
entry: bash scripts/check-executables-allow-sync.sh
|
||||
language: system
|
||||
stages: [pre-push]
|
||||
pass_filenames: false
|
||||
always_run: true
|
||||
|
||||
- id: apm-audit-ci
|
||||
name: apm audit --ci
|
||||
description: Run apm's producer-side CI gate over the root manifest AND each of the six plugin packages. Verifies exactly two things per manifest -- apm.yml parses as a valid APM manifest (manifest-parse), and, if it declares dependencies, apm.lock.yaml exists and is consistent (lockfile-exists). It does NOT enforce an org policy and does NOT scan for hidden Unicode; see the comment below for why. Reference:plugins/kyberforge/.apm/skills/apm-workflow/references/audit.md
|
||||
entry: bash -c 'for d in . plugins/*/; do (cd "$d" && apm audit --ci) || { echo "apm audit --ci failed in $d" >&2; exit 1; }; done'
|
||||
language: system
|
||||
stages: [pre-push]
|
||||
pass_filenames: false
|
||||
always_run: true
|
||||
# The description above deliberately claims less than this hook's old one
|
||||
# did ("lockfile/policy/hidden-content integrity"), because two of those
|
||||
# three were never happening:
|
||||
#
|
||||
# * POLICY. `apm audit --ci` discovers an org policy from the git remote,
|
||||
# and apm's discovery only understands github.com and Azure DevOps.
|
||||
# This repo's remote is a self-hosted Gitea, so discovery resolves
|
||||
# nothing and the run prints `No org policy found at unknown;
|
||||
# enforcement skipped`. apm's own message suggests
|
||||
# `policy.fetch_failure_default=block` in apm.yml "to fail closed" --
|
||||
# that was tried on a scratch copy and REJECTED: it does not make the
|
||||
# check meaningful, it makes it permanently red. `apm audit --ci` then
|
||||
# exits 1 with `No org policy found at unknown
|
||||
# (policy.fetch_failure_default=block)` on every push, because there is
|
||||
# no org policy to find and no supported way for this remote to serve
|
||||
# one. A gate that can never go green is not a gate. Revisit if this
|
||||
# repo ever gains a policy source apm can actually reach.
|
||||
# * HIDDEN CONTENT. The hidden-Unicode scan is plain `apm audit`, not
|
||||
# `apm audit --ci` (the two are different modes, and --ci refuses to
|
||||
# combine with --file/--strip/--dry-run/PACKAGE). Plain `apm audit`
|
||||
# here reports `No apm.lock.yaml found -- nothing to scan` and exits 0,
|
||||
# so adding it would buy a second vacuous check, not coverage.
|
||||
#
|
||||
# What IS left is worth keeping, and is now run against seven manifests
|
||||
# instead of one. lockfile-exists is conditional -- it is vacuous while
|
||||
# every apm.yml declares `dependencies: {apm: [], mcp: []}`, and it arms
|
||||
# itself the moment one does not (verified: adding a git dependency to
|
||||
# plugins/lint/apm.yml fails with `apm.yml declares dependencies but
|
||||
# apm.lock.yaml is absent`). manifest-parse is unconditional and fires on
|
||||
# any malformed manifest (verified: a dependency entry missing its
|
||||
# git/path/registry field fails with `Cannot parse apm.yml`). Running the
|
||||
# six plugin packages is what makes either reachable for them at all --
|
||||
# the root-only invocation audits the marketplace manifest and nothing
|
||||
# else. Costs ~0.5s per package, needs no network (checked under
|
||||
# `unshare -rn`) -- consistent with every other pre-push hook: none of
|
||||
# them need the network (see README.md's "Offline?" section).
|
||||
|
||||
- id: check-apm-agents-valid
|
||||
name: Validate real APM agent files
|
||||
description: Run agent-audit's validate.sh over every plugins/*/.apm/agents/*.agent.md file in this repo -- the artifacts it governs, not fixtures
|
||||
entry: bash scripts/check-apm-agents-valid.sh
|
||||
language: system
|
||||
stages: [pre-push]
|
||||
pass_filenames: false
|
||||
always_run: true
|
||||
# validate.sh was previously exercised only by check-scope-walkup-sync,
|
||||
# and only against synthetic mktemp fixtures -- it had never run against
|
||||
# the four agent files it governs. That is how ADR-0016 could be amended
|
||||
# to bless a `disallowedTools` frontmatter field while validate.sh's
|
||||
# allowlist still rejected it: the spec and its enforcer disagreed and
|
||||
# every gate stayed green. The expected file set is derived from
|
||||
# `git ls-files` (the pattern tests/run-bats.sh established) rather than
|
||||
# a hardcoded count, and discovering zero files is an error, not a pass.
|
||||
# Needs no network.
|
||||
|
||||
- id: apm-pack-check-clean
|
||||
name: apm pack --check-clean
|
||||
description: Release gate -- verify .claude-plugin/marketplace.json still matches what apm.yml + .apm/ would currently generate, and that per-package versions agree with the per_package versioning strategy. Closes issue #90's deferred item 3 (a check-clean-equivalent gate) using apm's own flag instead of custom drift logic.
|
||||
entry: apm pack --check-versions --check-clean --dry-run
|
||||
language: system
|
||||
stages: [pre-push]
|
||||
pass_filenames: false
|
||||
@@ -69,6 +171,25 @@ repos:
|
||||
stages: [pre-push]
|
||||
pass_filenames: false
|
||||
always_run: true
|
||||
# verbose so the DOWNGRADED run is audible. This hook can pass while
|
||||
# having verified strictly less than its name claims:
|
||||
# CHECK_VALE_STYLE_SYNC_ALLOW_MISSING_VALE=1 skips all six glob probes
|
||||
# and says so on a `passed (text-level only, vale unavailable)` line.
|
||||
# pre-commit prints nothing at all for a passing hook, so without this
|
||||
# the opt-out reinstated exactly the silent vacuous pass the script was
|
||||
# written to kill, one level up -- the run showed a bare `Passed` and
|
||||
# the documented instruction to read that summary line was impossible to
|
||||
# follow in the one situation the opt-out exists for. The script's clean
|
||||
# output is a single line, so this costs one line per push.
|
||||
|
||||
- id: check-scope-walkup-sync
|
||||
name: Check scope walk-up implementations agree
|
||||
description: Behaviorally cross-check validate.sh, validate-provenance.sh, new-agent.sh, and new-skill.sh's independent $HOME/.git/apm.yml walk-up ports against each other
|
||||
entry: bash scripts/check-scope-walkup-sync.sh
|
||||
language: system
|
||||
stages: [pre-push]
|
||||
pass_filenames: false
|
||||
always_run: true
|
||||
|
||||
- id: check-release-needed
|
||||
name: Check a release tag covers .pre-commit-hooks.yaml's paths
|
||||
@@ -79,15 +200,6 @@ repos:
|
||||
pass_filenames: false
|
||||
always_run: true
|
||||
|
||||
- id: validate-plugins
|
||||
name: Validate plugins
|
||||
description: Run claude plugin validate --strict on every plugin directory
|
||||
entry: bash -c 'for d in plugins/*/; do claude plugin validate --strict "$d" || exit 1; done'
|
||||
language: system
|
||||
stages: [pre-push]
|
||||
pass_filenames: false
|
||||
always_run: true
|
||||
|
||||
- id: validate-marketplace
|
||||
name: Validate marketplace manifest
|
||||
description: Run claude plugin validate --strict on the root marketplace manifest
|
||||
@@ -97,50 +209,55 @@ repos:
|
||||
pass_filenames: false
|
||||
always_run: true
|
||||
|
||||
- id: skill-frontmatter
|
||||
stages: ['pre-commit']
|
||||
name: SKILL.md frontmatter validation
|
||||
description: Ensure SKILL.md files have required frontmatter fields
|
||||
entry: bash
|
||||
language: system
|
||||
files: 'SKILL\.md$'
|
||||
args:
|
||||
- -c
|
||||
- |
|
||||
for f in "$@"; do
|
||||
if [[ -f "$f" ]]; then
|
||||
if ! grep -q "^name:" "$f" || ! grep -q "^description:" "$f"; then
|
||||
echo "ERROR: $f is missing required frontmatter fields (name: and description:)"
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
done
|
||||
|
||||
- id: skill-size-check
|
||||
stages: ['pre-commit']
|
||||
name: SKILL.md size ceiling
|
||||
description: Enforce agentskills.io's 500-line/5,000-token SKILL.md size ceiling
|
||||
name: SKILL.md size and context-budget ceilings
|
||||
description: Enforce agentskills.io's 500-line/2,770-whole-file-word spec ceilings AND ADR-0020's context budget -- description 250 chars SUGGESTION / 400 FAIL, body-only 600 words SUGGESTION / 900 FAIL, and every boundary-clause routing target resolving to a real skill or agent under plugins/*/.apm/
|
||||
entry: scripts/skill-size-check.sh
|
||||
language: script
|
||||
files: '^plugins/[^/]+/skills/[^/]+/SKILL\.md$'
|
||||
files: '^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$'
|
||||
pass_filenames: true
|
||||
verbose: true
|
||||
# verbose so the SUGGESTION tier is audible. ADR-0020 depends on it:
|
||||
# "A ceiling does not produce an average ... The halving depends
|
||||
# entirely on the 250-character SUGGESTION tier being visible and
|
||||
# respected." pre-commit prints nothing at all for a passing hook, and
|
||||
# a SUGGESTION deliberately does not fail, so without verbose every
|
||||
# suggestion would be swallowed -- the exact invisibility ADR-0013
|
||||
# records for Vale warnings. Costs nothing on a clean file: the script
|
||||
# prints only findings.
|
||||
|
||||
- id: check-rtk-prefix
|
||||
stages: ['pre-commit']
|
||||
name: ADR-0023 rtk prefix on executable git commands
|
||||
description: Enforce ADR-0023 clause 1 -- an executable, instructed git command in a shell code fence or a dispatch-table Run cell is written `rtk git`. Clauses 2 and 3 are not machine-decidable; a deliberately bare command opts out with the literal string ADR-0023 on its own line
|
||||
entry: scripts/check-rtk-prefix.sh
|
||||
language: script
|
||||
files: '^plugins/[^/]+/\.apm/(skills/.*\.md|agents/.*\.agent\.md)$'
|
||||
# README.md is excluded on purpose, not by oversight. A skill-directory
|
||||
# README is consumer-facing prose that no agent ever loads, and the
|
||||
# `git clone` lines in the seven tests/README.md files are setup
|
||||
# instructions for a third party who has no rtk installed. Prefixing
|
||||
# those would be actively wrong -- see ADR-0023's consumer section.
|
||||
exclude: '(^|/)README\.md$'
|
||||
pass_filenames: true
|
||||
|
||||
- id: vale-audit-prefilter-skill
|
||||
stages: ['pre-commit']
|
||||
name: Vale audit prefilter (SKILL.md)
|
||||
description: Run Vale against SKILL.md files as a deterministic prefilter for skill-audit, via skill-audit's own bundled copy
|
||||
entry: plugins/kyberforge/skills/skill-audit/scripts/vale-wrap.sh
|
||||
entry: plugins/kyberforge/.apm/skills/skill-audit/scripts/vale-wrap.sh
|
||||
language: script
|
||||
files: '^plugins/[^/]+/skills/[^/]+/SKILL\.md$'
|
||||
files: '^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$'
|
||||
pass_filenames: true
|
||||
|
||||
- id: vale-audit-prefilter-agent
|
||||
stages: ['pre-commit']
|
||||
name: Vale audit prefilter (agent files)
|
||||
description: Run Vale against agent markdown files as a deterministic prefilter for agent-audit, via agent-audit's own bundled copy
|
||||
entry: plugins/kyberforge/skills/agent-audit/scripts/vale-wrap.sh
|
||||
entry: plugins/kyberforge/.apm/skills/agent-audit/scripts/vale-wrap.sh
|
||||
language: script
|
||||
files: '^plugins/[^/]+/agents/[^/]+\.md$'
|
||||
files: '^plugins/[^/]+/\.apm/agents/[^/]+\.agent\.md$'
|
||||
pass_filenames: true
|
||||
|
||||
- repo: meta
|
||||
|
||||
@@ -1,20 +1,23 @@
|
||||
- id: kyberforge-vale-audit-skill
|
||||
name: Kyberforge Vale prose audit (SKILL.md)
|
||||
description: Deterministic prose-pattern prefilter for kyberforge's skill-audit, via its own bundled Vale config/styles
|
||||
entry: plugins/kyberforge/skills/skill-audit/scripts/vale-wrap.sh
|
||||
entry: plugins/kyberforge/.apm/skills/skill-audit/scripts/vale-wrap.sh
|
||||
language: script
|
||||
files: '(^|/)SKILL\.md$'
|
||||
|
||||
- id: kyberforge-vale-audit-agent
|
||||
name: Kyberforge Vale prose audit (agent files)
|
||||
description: Deterministic prose-pattern prefilter for kyberforge's agent-audit, via its own bundled Vale config/styles
|
||||
entry: plugins/kyberforge/skills/agent-audit/scripts/vale-wrap.sh
|
||||
entry: plugins/kyberforge/.apm/skills/agent-audit/scripts/vale-wrap.sh
|
||||
language: script
|
||||
files: '(^|/)agents/[^/]+\.md$|\.agent\.md$'
|
||||
|
||||
- id: kyberforge-skill-size-check
|
||||
name: SKILL.md size ceiling
|
||||
description: Enforce agentskills.io's 500-line/5,000-token SKILL.md size ceiling
|
||||
name: SKILL.md size and context-budget ceilings
|
||||
description: Enforce agentskills.io's 500-line/2,770-whole-file-word spec ceilings plus ADR-0020's context budget (description 250 chars SUGGESTION / 400 FAIL, body-only 600 words SUGGESTION / 900 FAIL, resolvable boundary-clause routing targets)
|
||||
entry: scripts/skill-size-check.sh
|
||||
language: script
|
||||
files: '(^|/)SKILL\.md$'
|
||||
# verbose so the SUGGESTION tier reaches a human -- pre-commit prints
|
||||
# nothing for a passing hook, and a SUGGESTION deliberately does not fail.
|
||||
verbose: true
|
||||
|
||||
42
AGENTS.md
42
AGENTS.md
@@ -1,39 +1,47 @@
|
||||
# Working in this repo
|
||||
|
||||
This repo is the global AI development configuration repository — the authoritative source for agent definitions, skills, workflows, and prompts across all projects. Built as a homelab tool intended to scale to professional environments.
|
||||
The global AI development configuration repository — the authoritative source for agent definitions, skills, workflows, and prompts across all projects.
|
||||
|
||||
This file carries only what applies to **every** session. Setup, prerequisites, and test commands are in `README.md`; the reasoning behind each enforcement gate is in `docs/spec/gates.md`.
|
||||
|
||||
## Structure
|
||||
|
||||
- `plugins/` — installable plugin units; each is self-contained (skills, agents, hooks, MCP servers, bundled assets); install separately via `claude plugin install <name>@holocron`
|
||||
- `providers/claude-code/` — Claude Code adapter (deployed to `~/.claude/` via `install.sh`)
|
||||
- `plugins/` — six installable plugin units, each an apm package (`apm.yml` + `.apm/`). Root `apm.yml` declares all six as `dependencies.apm`; `apm install` deploys them into `.claude/skills/` and `.claude/agents/`, both gitignored install output.
|
||||
- `providers/claude-code/` — Claude Code adapter, deployed to `~/.claude/` via `scripts/install.sh`.
|
||||
|
||||
## Prefer plugin skills over raw shell
|
||||
|
||||
This repo dogfoods its own plugins. Before shelling out to git, gitea, or lint tooling directly, check whether an installed skill already owns the operation — it usually does:
|
||||
This repo dogfoods its own plugins. Before shelling out, check whether a skill already owns the operation — it usually does:
|
||||
|
||||
- Commits, branches, history, worktrees, remotes → `git:git-commits`, `git:git-branches`, `git:git-history`, `git:git-worktrees`, `git:git-remotes`
|
||||
- Pre-commit hook install/config/troubleshooting → `git:pc-run` / `git:pc-author`
|
||||
- Issues, PRs, labels, milestones → `gitea:gitea-issues`, `gitea:gitea-prs`, `gitea:gitea-labels-milestones`; also `gitea:gitea-branches`, `gitea:gitea-files`, `gitea:gitea-releases`, or `gitea:gitea-workflow` when the domain is ambiguous
|
||||
- Vale prose linting → `lint:vale-config` / `lint:vale-run`
|
||||
- This repo's own AGENTS.md → `core:agentsmd-author` / `core:agentsmd-audit`
|
||||
- Commits, branches, history, worktrees, remotes → `git-commits`, `git-branches`, `git-history`, `git-worktrees`, `git-remotes`
|
||||
- Pre-commit hook install/config/troubleshooting → `pc-run` / `pc-author`
|
||||
- Issues, PRs, labels, milestones → `gitea-issues`, `gitea-prs`, `gitea-labels-milestones`; also `gitea-branches`, `gitea-files`, `gitea-releases`, or `gitea-workflow` when the domain is ambiguous
|
||||
- Vale prose linting → `vale-config` / `vale-run`
|
||||
- This repo's own AGENTS.md → `agentsmd-author` / `agentsmd-audit`
|
||||
|
||||
Use the bare, **unnamespaced** names. That is what `apm install` deploys and the only form this repo's own install produces — a project skill has no plugin to prefix (ADR-0018). Whether the `<plugin>:` form (`gitea:gitea-prs`) also resolves depends on native plugin installs at user scope, outside this repo; write the bare name either way.
|
||||
|
||||
Fall back to raw shell only when no skill covers it.
|
||||
|
||||
## Setup and testing
|
||||
## Session rules
|
||||
|
||||
- Install git hooks via `git:pc-run`, wiring all three stages — this repo's `.pre-commit-config.yaml` has no `default_install_hook_types`, so a plain install silently skips `commit-msg` (Conventional Commits) and `pre-push` (tests, manifest check).
|
||||
- Install the `vale` binary — required by the `vale-audit-prefilter-skill`/`-agent` pre-commit hooks, which run on every commit touching a `SKILL.md` or agent `.md` file. Without it the hooks fail with a bare "command not found" and no install pointer. `brew install vale` (macOS), `snap install vale` (Linux), `choco install vale` (Windows), or see https://vale.sh/docs/vale-cli/installation/. No `vale sync` needed — the `Kyberforge` styles are committed under `plugins/kyberforge/skills/{skill-audit,agent-audit}/assets/vale/styles/`, not downloaded packages (see ADR-0014).
|
||||
- Run `bash tests/run-tests.sh` before considering any change done — it runs every `test-*.sh` script in the repo plus the bats suite (`--bats-only` for just bats). First run auto-initializes the bats submodules; no manual `git submodule update` needed.
|
||||
- Pushing re-runs the full suite plus `scripts/check-manifests.sh` via the pre-push hook — same commands, so run them locally first.
|
||||
- Author commits with `git:git-commits` — it validates Conventional Commits (enforced at `commit-msg`) for you.
|
||||
- **Do not add repo-owned keys to `.claude/settings.json`.** apm treats it as its own deployed artifact and `apm audit --ci` replays the install and diffs, so anything apm would not have written is permanent drift that fails the `apm-audit-ci` pre-push hook. A hook you want here is authored in `plugins/<name>/.apm/hooks/` and deployed by apm, never hand-written into that file. The `SessionStart` entry already in it is exactly that: kyberforge authors it in `plugins/kyberforge/.apm/hooks/hooks.json` and apm merges it in, so it is apm's own output, it is what the replay expects, and it belongs in the commit — do not strip it (ADR-0019). Machine-specific settings go in the gitignored `.claude/settings.local.json`; shared enforcement goes in `.pre-commit-config.yaml`.
|
||||
- **`apm.lock.yaml` turning up modified is expected, not a bug.** kyberforge's `SessionStart` hook keeps the install current on launch and rewrites the lock in the process (ADR-0019). Commit or discard it deliberately.
|
||||
- **A `.apm/` edit is not live in this session until it is pushed.** The six dependencies resolve from the holocron remote, unpinned against the default branch. `apm install` deploys from the lock; `apm update` is what re-resolves refs.
|
||||
- **No pre-push hook needs the network.** Root `apm.yml`'s marketplace has no remote package entries, so every hook resolves locally.
|
||||
- **This repo and Gitea are the only source of truth.** All project state, decisions, and working conventions live here. Do not use an external memory system for this project — cached state diverges from the repo and you get a split brain. Before answering any design or architecture question, check `docs/adr/` for an existing decision.
|
||||
|
||||
## Key documents
|
||||
|
||||
Read CONTEXT.md at the start of every session in this repo.
|
||||
Read `CONTEXT.md` at the start of every session — it is this repo's domain glossary, and the terms it defines are used unglossed everywhere else. It is not exhaustive: terms it does not carry are defined at their point of use, mostly in `docs/spec/`.
|
||||
|
||||
Read these on demand:
|
||||
|
||||
- `docs/spec/architecture.md` — current directory structure, install pipeline, provider model
|
||||
- `README.md` — prerequisites, install, and test commands
|
||||
- `docs/VISION.md` — the phased roadmap and where this is going; read when a decision turns on product direction
|
||||
- `LESSONS.md` — patterns that went wrong once; read before repeating a class of change that has burned the repo before
|
||||
- `docs/spec/gates.md` — what each pre-commit and pre-push hook enforces and why; read when a gate fails or before changing hook config
|
||||
- `docs/spec/architecture.md` — directory structure, install pipeline, provider model
|
||||
- `docs/adr/` — architectural decisions; read before answering design questions or proposing structural changes
|
||||
- `docs/ai-constitution.md` — full governance evidence base; read when a governance decision needs justification
|
||||
- `docs/research/ai-coding-factory/ai-coding-factory-principles.md` — factory design rationale; read when implementing, auditing, or reviewing skills or factory structure
|
||||
|
||||
200
CONTEXT.md
200
CONTEXT.md
@@ -1,80 +1,180 @@
|
||||
---
|
||||
name: AI Development Repo
|
||||
description: Domain language and decisions for the global AI development config repository
|
||||
description: The domain language of the global AI development config repository
|
||||
---
|
||||
|
||||
# Context
|
||||
# AI Development Repo
|
||||
|
||||
## Principles
|
||||
The bounded context of this repo is **how agent instructions are authored, packaged, distributed, and
|
||||
kept small**. Terms here name concepts specific to that problem. Mechanics live elsewhere:
|
||||
`docs/spec/architecture.md` for structure, `docs/spec/gates.md` for enforcement, `docs/adr/` for
|
||||
decisions.
|
||||
|
||||
### CLAUDE.md index model
|
||||
`AGENTS.md` is the source of always-on universal rules (provider-agnostic). `providers/claude-code/CLAUDE.md` is a thin adapter: it imports `~/.agents/AGENTS.md` via `@~/.agents/AGENTS.md` and appends Claude Code-specific additions (`@import` for governance.md, content index). Deployed to `~/.claude/CLAUDE.md` via `install.sh`. Context size is kept minimal — only what is needed every session is loaded upfront; detailed content is pulled on demand. See ADR-0003.
|
||||
## Language
|
||||
|
||||
### Instruction file format
|
||||
`core/instructions/<topic>.md` files are plain markdown — no frontmatter, no schema. The agent decides when to read each file based on task context and the content index label in `providers/claude-code/CLAUDE.md`. Frontmatter is deferred until there is evidence that agents are loading the wrong files in practice.
|
||||
### Context cost
|
||||
|
||||
### Repo/gitea as source of truth
|
||||
All project state, decisions, context, and working conventions live in this repo or Gitea. External memory systems should not be used for this project — they create a split-brain risk where cached state diverges from the repo. At the start of every session, read `CLAUDE.md`, `CONTEXT.md`, and `docs/VISION.md`. Everything needed to orient is here.
|
||||
**Routing target**:
|
||||
The skill or agent name a boundary clause sends work to. It **resolves** when a skill or agent of
|
||||
that name is reachable from the file being checked, and **dangles** when none is — a route the router
|
||||
cannot take. Dangling is a blocking ERROR in route notation (`/name`, `→ name`) and a SUGGESTION for
|
||||
a bare name nothing else in the sentence corroborates. Verdicts and the resolution walk:
|
||||
`docs/spec/gates.md`.
|
||||
_Avoid_: route, pointer, cross-reference
|
||||
|
||||
Before answering any design or architecture question, check for existing decisions: `docs/adr/` (hard architectural decisions).
|
||||
**Dispatch body**:
|
||||
The body pattern a skill with two or more mutually exclusive flows must use — the body carries only
|
||||
the dispatch table and the gates common to every branch, and each flow lives in its own
|
||||
self-contained `references/` file. Exemplar: `apm-workflow`.
|
||||
_Avoid_: router body, thin body
|
||||
|
||||
## Glossary
|
||||
**Hand-invoked skill**:
|
||||
A skill reached only by typing its slash command, declared `disable-model-invocation: true`. The host
|
||||
withholds it from the model-visible listing entirely, so it pays no preload tax and its description
|
||||
becomes human-facing text. The flag also hard-blocks the Skill tool, so **no other skill can route to
|
||||
a hand-invoked skill** — a `` Call `x` `` step in another skill's body stops working the moment `x`
|
||||
takes the flag. Check inbound routes before declaring one. Exemplar: `zoom-out`.
|
||||
_Avoid_: manual skill, disabled skill
|
||||
|
||||
### Management Application
|
||||
A separate product (separate repo) for browsing, editing, and configuring AI development configs through a proper product UI. Git is the persistence layer, invisible to the user. The app is repo-agnostic — it works with any git repo that follows these conventions. This repo is the canonical default content (the official starter). See `docs/VISION.md` for the phased roadmap.
|
||||
**Delegation discipline**:
|
||||
The agent-side counterpart to the dispatch body. A plugin-scope agent is a single `.agent.md` file
|
||||
with no sibling `references/` directory, so it cannot disclose to itself — it can only delegate to
|
||||
skills. Its characteristic defect is therefore restatement, not length.
|
||||
_Avoid_: agent hygiene
|
||||
|
||||
### Skills
|
||||
Reusable slash commands for AI coding tools, defined as `SKILL.md` files following the [Agent Skills open standard](https://agentskills.io). Deployed via plugin — `plugins/<plugin-name>/skills/<skill-name>/SKILL.md`, available after the plugin is installed (`claude plugin install <name>@<marketplace>`). Skills are self-contained — they cannot reference files outside the plugin directory after install-time caching.
|
||||
### Distribution
|
||||
|
||||
### Plugin
|
||||
The deployable unit in the plugin marketplace. A plugin bundles one or more skills, agents, hooks, prompts, MCP servers, and optionally a `bin/` directory into a single installable directory. Each plugin has two manifests: `.claude-plugin/plugin.json` (Claude Code) and `plugin.json` at the plugin root (Copilot CLI). Plugins are copied to a cache on install — they cannot reference files outside their own directory. In this repo, plugins live under `plugins/<name>/`. Install a plugin with `claude plugin install <name>@<marketplace>`.
|
||||
**Skill**:
|
||||
A reusable slash command defined as a `SKILL.md` file following the
|
||||
[Agent Skills open standard](https://agentskills.io), authored at
|
||||
`plugins/<plugin>/.apm/skills/<skill>/SKILL.md`.
|
||||
_Avoid_: command, prompt, macro
|
||||
|
||||
### Plugin marketplace
|
||||
A Git repository with a `marketplace.json` manifest listing installable plugins. No backend, registry, or SaaS required — the Git repo is the marketplace. This repo is the `holocron` marketplace. The manifest lives at `.claude-plugin/marketplace.json` (read by both Claude Code and Copilot CLI) and is mirrored to `.github/plugin/marketplace.json`.
|
||||
**apm package**:
|
||||
The deployable unit apm builds and installs — one or more skills, agents, hooks, commands, and MCP
|
||||
servers under a single directory `plugins/<name>/`, consisting of that directory's `apm.yml` plus
|
||||
the hand-authored `plugins/<name>/.apm/` tree it deploys from (ADR-0015).
|
||||
_Avoid_: bundle, module, source tree; and bare "plugin" for the *installable artifact*, which since
|
||||
ADR-0024 is an apm package and not a Claude Code plugin. "Plugin" stays correct as a modifier in the
|
||||
repo's settled compounds — **Plugin marketplace**, "plugin units", `plugins/`.
|
||||
|
||||
### HITL (human-in-the-loop)
|
||||
Agent pauses before a consequential action; human approves before execution. Required for irreversible or high-stakes actions (architecture changes, production deployments, security configuration). The agent drafts the change plan and waits — it does not proceed autonomously. Contrast with HOTL.
|
||||
**Output profile**:
|
||||
A named ecosystem format `apm pack` can compile the marketplace manifest into, declared per profile
|
||||
under root `apm.yml`'s `marketplace.outputs:`. apm defines `claude` and `codex`; each writes to its
|
||||
own default path unless overridden. Mechanics: `docs/spec/architecture.md`.
|
||||
_Avoid_: build target, export format
|
||||
|
||||
### HOTL (human-on-the-loop)
|
||||
Agent acts; human monitors and can intervene after the fact. Acceptable for low-stakes, bounded, reversible actions where the cost of pausing for approval exceeds the blast radius of an error. The distinction between HITL and HOTL must be explicit and documented — defaulting to HOTL for convenience is not acceptable.
|
||||
**Plugin marketplace**:
|
||||
A Git repository carrying a `marketplace.json` manifest that lists installable plugins. There is no
|
||||
backend, registry, or SaaS — the Git repo is the marketplace.
|
||||
_Avoid_: registry, store, catalogue
|
||||
|
||||
### Sycophancy
|
||||
The failure mode where RLHF-trained models prioritise approval over accuracy. Treated as a first-class reliability risk: models change correct answers to wrong ones under user pressure in a majority of observed cases, then persist in the wrong answer. Designing against sycophancy is an explicit obligation, not a quality-of-life concern. Countermeasures: explicit pushback resistance instructions, prompting for dissent, cross-validating against independent sources. Never interpret AI agreement as AI accuracy.
|
||||
**holocron**:
|
||||
This repository, in its role as a plugin marketplace and as the remote the six plugin dependencies
|
||||
resolve against.
|
||||
_Avoid_: the marketplace, upstream
|
||||
|
||||
### AGENTS.md
|
||||
The provider-agnostic always-on instruction entry point. Two files:
|
||||
- **Repo-level `AGENTS.md`** — instructions for agents working inside this repo (structure, key rules); imported by repo `CLAUDE.md` via `@AGENTS.md`.
|
||||
- **Global `core/AGENTS.md`** — Communication and Behavior rules that apply across all projects; deployed to `~/.agents/AGENTS.md`; imported by `~/.claude/CLAUDE.md` via `@~/.agents/AGENTS.md`.
|
||||
**Provenance chain**:
|
||||
The three-stage traceability record linking a skill back to its research inputs: `/research` produces
|
||||
topic docs and a `sources.md`; the author skill records which sources informed which files in
|
||||
`references/sources.md` and `source_keys` frontmatter; `skill-audit` validates the chain is complete
|
||||
and internally consistent.
|
||||
_Avoid_: sources, citations, attribution
|
||||
|
||||
Contains always-on rules in plain markdown with no provider-specific syntax (no `@import`). Provider-specific files (`CLAUDE.md`) are thin adapters that import the relevant `AGENTS.md` and add only Claude Code-specific syntax. This pattern means a single source of truth can serve multiple providers without duplication. See ADR-0003.
|
||||
### Governance
|
||||
|
||||
### Skill composition
|
||||
A skill calling another skill by name to delegate a sub-task. The calling skill focuses on the orchestration decision ("when to do X"); the called skill owns the mechanics ("how to do X"). Established compositions: `grill-me` calls `write-adr` when a decision crystallises; `implement-feature` calls `tdd` as its implementation methodology; `forge` calls `grill-with-docs` to refine intent, classifies the target artifact type (skill / agent / plugin / marketplace entry), then routes to the matching `*-author` skill — which owns its own create/improve logic and, where applicable, its own inline audit closeout (`skill-author` runs `/skill-audit`, `agent-author` runs `kyberforge:agent-audit`, both in the same context as the authoring work). Reserve `forge` for genuinely undecided "which artifact type is this" questions — an already-fully-specified corrective edit (exact file, line, and fix already known) should call the target author skill directly instead (`skill-author`, `plugin-author`, `agentsmd-author`, etc.); routing a known fix through `forge`'s grill-and-classify layer adds unnecessary indirection and, in practice, has been observed to lose track of hard constraints handed down the chain (e.g. "don't commit yet," "edit in this worktree") because each hop re-derives instructions from a shorter brief. `forge` additionally runs its own independent recheck after a skill/agent route finishes: a clean-context subagent (not forked, no inherited context) re-runs the same audit skill against the finished artifact, as a distinct verification layer from the author skill's inline audit — the two can share blind spots since the inline audit runs in the same context as the work it checks. If the clean audit surfaces any unresolved finding, `forge` loops — re-invoke the author skill to resolve it, re-run the clean audit — until the clean audit comes back with nothing unresolved; only then is the route done. `plugin-author` and `marketplace-author` have no audit counterpart and get no recheck; their terminal check is `claude plugin validate`.
|
||||
**HITL** (human-in-the-loop):
|
||||
The agent pauses before a consequential action and a human approves before execution. Required for
|
||||
irreversible or high-stakes actions — architecture changes, production deployments, security
|
||||
configuration.
|
||||
_Avoid_: manual approval, gated action
|
||||
|
||||
### Provider-agnostic issue tracker
|
||||
Skills and workflows reference "linked issue" generically rather than a specific provider. Gitea is the canonical issue tracker for this repo (see ADR-0017). "Issue" is the cross-provider term (GitHub, GitLab, Gitea all use it).
|
||||
### Documents
|
||||
|
||||
### Provenance chain
|
||||
The three-stage traceability record linking a skill back to its research inputs: (1) `/research` produces topic docs and a `sources.md` in `plugins/<plugin>/docs/research/docs/<topic>/`; (2) `/skill-author` reads those docs and records which sources informed which skill files in `references/sources.md` (including a `Research doc:` back-pointer to the upstream research file) and `source_keys` frontmatter on `SKILL.md` and `references/*.md`; (3) `skill-audit` validates the chain is complete and internally consistent via `validate-provenance.sh`. A skill with research input but no `references/sources.md`, or with `source_keys` that don't match `references/sources.md` slugs, has a broken provenance chain.
|
||||
**AGENTS.md**:
|
||||
The provider-agnostic always-on instruction file, in plain markdown with no provider-specific syntax
|
||||
(ADR-0003). Two exist: repo-level, and the global `core/AGENTS.md` deployed to `~/.agents/AGENTS.md`.
|
||||
_Avoid_: instructions file, system prompt
|
||||
|
||||
### Bidirectional reference principle
|
||||
Files that reference other files should declare those references explicitly. The referencing file carries the forward reference (e.g. content index in `CLAUDE.md`, `references:` in frontmatter). The referenced file carries a `when:` field describing when it is loaded. Both sides should agree — divergence signals staleness. The reverse map ("what files reference this file?") is derived by a reference scanner script, not maintained manually. This principle applies to instruction files, skills, and workflow documents.
|
||||
**Thin adapter**:
|
||||
A provider-specific instruction file (`CLAUDE.md`, `.cursor/rules/*.mdc`, `copilot-instructions.md`)
|
||||
that imports its `AGENTS.md` and adds only that provider's syntax, carrying no original always-on
|
||||
content of its own (ADR-0002, ADR-0003).
|
||||
_Avoid_: wrapper, shim, provider file
|
||||
|
||||
### agentsmd-author / agentsmd-audit
|
||||
A skill pair in the `core` plugin for writing, updating, and reviewing a repo's `AGENTS.md` file(s) — the generic open-standard file (see the `AGENTS.md` entry above), including this repo's own. `agentsmd-author` creates/updates AGENTS.md content, supports nested monorepo placement (per the standard's nearest-file-wins precedence), and closes out by invoking `agentsmd-audit` inline. `agentsmd-audit` runs a single combined pass checking three mandatory baselines: secrets/credentials (governance.md hard prohibition — AGENTS.md is committed content), structural completeness (common-sections checklist from the agents.md spec), and accuracy/drift (do referenced commands and paths actually resolve against the repo). `agentsmd-audit` never inspects provider adapter files (see `provider-adapter-author`) — its scope is AGENTS.md content only. Chosen over folding this into `kyberforge` because kyberforge's scope is meta-tooling for the holocron marketplace itself, not generic target-repo documentation; `core` is the intended home for cross-cutting, repo-agnostic utility skills.
|
||||
**LESSONS.md**:
|
||||
The long-loop feedback log for patterns observed across sessions, at the repo root.
|
||||
_Avoid_: changelog, retro, postmortem
|
||||
|
||||
### provider-adapter-author
|
||||
A companion skill (`core` plugin) that detects a target repo's provider-specific instruction file (`CLAUDE.md`, `.cursor/rules/*.mdc`, `copilot-instructions.md`, etc.) and, where it duplicates content AGENTS.md should own, converts it into a thin adapter that imports AGENTS.md — mirroring this repo's own ADR-0002/ADR-0003 two-tier adapter pattern. Self-validates via its own bundled deterministic script (`scripts/validate-adapter.sh`: checks for an import reference, no duplicated headings, size threshold) rather than a separate paired audit skill — the check is mechanical, so a script suffices per governance.md's "prefer deterministic code for repeatable tasks." `agentsmd-author` calls this skill via skill composition when it detects an existing provider file with overlapping content.
|
||||
### Quality
|
||||
|
||||
### lint plugin
|
||||
A standalone, repo-agnostic plugin (`plugins/lint/`) for configuring and running linters — not scoped to kyberforge's own meta-tooling. First linter is Vale (prose style linting), split into two skills per the git/gitea per-concern pattern: `vale-config` (setup — `.vale.ini`, `StylesPath`, styles) and `vale-run` (invoke Vale, interpret/report findings). A `lint-runner` agent composes these for isolated-context lint sweeps; it is report-only (no `Edit` tool) — it flags findings, it does not rewrite prose. Vale's research docs (`docs/research/docs/vale/`) moved from `plugins/kyberforge/` to `plugins/lint/` to keep the provenance chain same-plugin.
|
||||
**Skill composition**:
|
||||
A skill calling another skill by name to delegate a sub-task — the caller owns the orchestration
|
||||
decision ("when to do X"), the callee owns the mechanics ("how to do X").
|
||||
_Avoid_: chaining, nesting, sub-skill
|
||||
|
||||
### Vale audit prefilter (skill-audit / agent-audit)
|
||||
Wiring Vale as a deterministic prefilter for `skill-audit`/`agent-audit`'s Description dimension (ADR motivation: issue #84) is repo-specific, not part of the generic `lint` plugin, so it doesn't live in `plugins/lint/` — but per ADR-0014 it also doesn't live at the repo root anymore. Two copies live inside `plugins/kyberforge/`, one per skill, since a plugin's cache-install only copies each skill's own files (no cross-skill sharing): `plugins/kyberforge/skills/agent-audit/assets/vale/` is canonical (`.vale.ini` plus a custom `Kyberforge` style covering description-opener banning ("This skill/agent..."), vague-capability wording ("helps with", "utilize", ...), and generic "see references/ for details" padding — and a `KyberforgeCopilot` style scoped only to `.agent.md` files for the Copilot-only "Use proactively has no effect" check), and `plugins/kyberforge/skills/skill-audit/assets/vale/` is a smaller duplicate (`Kyberforge` only, scoped to `SKILL.md`) kept in sync by `scripts/check-vale-style-sync.sh` (pre-push). A root-level `.pre-commit-hooks.yaml` exposes both copies (plus `skill-size-check`) so any external repo can enforce the same rules via `repo: <this-repo-url>, rev: <tag>` in its own `.pre-commit-config.yaml` — pre-commit clones the pinned rev into its own cache, independent of whether Claude Code or the `kyberforge` plugin is installed at all, and the same mechanism covers CI (`pre-commit run --all-files`). This repo's own `vale-audit-prefilter-skill`/`-agent` pre-commit hooks consume the identical plugin-bundled copies via `repo: local` (not a third root copy, and not a pinned self-reference — a pinned self-reference would lint working-tree edits against the last tagged release rather than the change being made). Every rule is `level: error` and every alert is a FAIL — no ignorable tier, same as shellcheck, the test suite, and conventional-pre-commit. Graded severities do not work here: Vale's exit code keys on `error` alerts alone, so `warning`/`suggestion` rules exit 0 and pre-commit swallows the output of a passing hook, leaving them invisible and blocking nothing. `MinAlertLevel` and `--minAlertLevel` are correspondingly absent from `.vale.ini` and the hook, being no-ops under this model. Vale covers the pattern-matchable sub-checks named in issue #84 (imperative opener, vague filler, `Use proactively`, generic reference-pointer padding) plus, per ADR-0013, one body-wide prose-pattern check ("There is/are" sentence openers) — everything else about body discipline (defaults-vs-menus, why-rationale, non-pattern-matchable judgment calls), near-miss exclusion strength, and control calibration stays LLM judgment.
|
||||
**Authoring root**:
|
||||
The directory a gate resolves against — the nearest ancestor of the file being checked holding
|
||||
`plugins/*/.apm/skills` or `plugins/*/.apm/agents`, falling back to the nearest ancestor holding
|
||||
`.git`. The walk: `docs/spec/gates.md`.
|
||||
_Avoid_: repo root, project root
|
||||
|
||||
Both skills' Step 1, and the `vale-audit-prefilter-skill`/`-agent` pre-commit hooks, call each copy's own `scripts/vale-wrap.sh` rather than `vale` directly — a workaround for a confirmed Vale 3.15.2 limitation (see `vale-config`'s Gotchas): `text.frontmatter.description` silently stops matching on most — not all — multi-line descriptions. Verified by reproduction, not assumed: `>` folded scalars, plain (unquoted) continuation lines, and single- or double-quoted multi-line scalars all yield 0 alerts and exit 0 on a deliberately-bad fixture, while a `|` literal block spanning the same 2+ lines lints normally (alerts fire, exit 1). The wrapper flattens those three broken forms to one physical line in a scratch copy (padding with blank lines so every other line number is unchanged) before handing off to real `vale`; `|` literal blocks and single-line descriptions pass through untouched, already linting correctly. The plain and quoted forms previously passed silently — unflattened and unmatched — so a bad description in either sailed through the prefilter. Handed no `--config` at all, the wrapper falls back to its own sibling `assets/vale/.vale.ini`, located from `${BASH_SOURCE[0]}` rather than from the cwd — which is why both manifests' `entry:` is now the bare script path with no argument after it. pre-commit prefixes only `entry[0]` with the hook-repo clone path (`cmd = (prefix.path(cmd[0]), *cmd[1:])`), so every later argument resolves against the *consuming* repo's root: a `--config` in `.pre-commit-hooks.yaml` pointed at a path no consumer has and hard-failed every external run with `E100 [--config] Runtime error`. `.pre-commit-config.yaml` drops the argument too, deliberately keeping the two entries identical — the local `repo: local` hook resolved its `--config` correctly only because the consuming repo *was* this repo, and that divergence is why three review rounds exercised a path no external consumer takes and missed the defect. An explicit `--config` still wins, in all three argv forms (`--config X`, `--config=/abs`, `--config=rel`), and a relative one still resolves against the caller's cwd, matching bare `vale`, not the repo root. Both audit skills' Step 1 now passes no `--config` either: it resolves the script relative to the skill's own directory so the call works from an installed plugin cache, but a relative `--config` alongside it would still resolve against the cwd, yielding `E100 Runtime error ... does not exist` and exit 2 — which both skills' fallback misreads as "vale unavailable" and silently downgrades to full LLM judgment. `tests/test-vale-wrap.sh` regression-tests this against skill-audit's copy specifically (its fixtures are all `SKILL.md`-shaped, and only skill-audit's `.vale.ini` has that glob section). Each `.vale.ini`'s section globs are path-agnostic (`[**/SKILL.md]` for skill-audit's copy; `[**/agents/*.md]`/`[**/*.agent.md]` for agent-audit's) and do no scoping on their own: Vale's `*` crosses `/`. Scoping comes from each pre-commit hook's own `files:` regex and from the audit skills passing one explicit file per invocation. The two manifests scope differently on purpose: this repo's `.pre-commit-config.yaml` pins its own layout — `^plugins/[^/]+/skills/[^/]+/SKILL\.md$` for `-skill`, `^plugins/[^/]+/agents/[^/]+\.md$` for `-agent` — while the shipped `.pre-commit-hooks.yaml` stays layout-agnostic for external consumers whose skills live anywhere, using `(^|/)SKILL\.md$` and `(^|/)agents/[^/]+\.md$|\.agent\.md$`. Both manifests split the prefilter into two hooks precisely because one combined hook pointed at only one copy would silently 0-file-skip the other file type. A `SKILL.md` outside `plugins/` (e.g. project-scope `.claude/skills/foo/SKILL.md`) still matches `[**/SKILL.md]` and gets linted normally — the globs constrain filename shape, not location. Vale reports 0 files only when the path it is handed matches no glob section at all: a differently-named file, or a directory argument holding nothing that matches. That run prints `✔ 0 errors ... in 0 files.` and exits 0, indistinguishable from a clean pass, so both audits treat a 0-file Vale run as NOT RUN and fall back to full LLM judgment.
|
||||
**Near-miss**:
|
||||
A query that shares keywords with this skill but needs a different one — and, by extension, the
|
||||
sibling that would wrongly answer it; boundary clauses exist to exclude genuine near-misses rather
|
||||
than to enumerate siblings. Detail: `skill-audit/references/description-quality.md`.
|
||||
_Avoid_: overlap, similar skill
|
||||
|
||||
This scope expands per ADR-0013: one cherry-picked low-noise `write-good`/`alex` rule landed in `styles/Kyberforge`, `Kyberforge.SentenceOpenerThereIs` (22 held-out hits, both in-corpus hits clean rewrites, zero suppressions). A second, `Kyberforge.VagueQualifier`, was cherry-picked and then deleted: 2 hits across the 41 skill/agent files, one marginal and one an unfixable false positive (`caveman/SKILL.md` quotes `of course` as an example of filler — a mention, not a use) that forced the repo's only Vale suppression comments. Also new is a sibling pre-commit hook, `skill-size-check` (`scripts/skill-size-check.sh`), enforcing agentskills.io's `SKILL.md` ceiling as two blocking gates: `MAX_LINES=500` and `MAX_WORDS=2770` (a word-count proxy for the 5,000-token limit, calibrated to the densest prose measured in this repo — 1.81 tokens per word — so even a worst-case `SKILL.md` at the ceiling stays under 5,000 tokens). Both are inclusive, and `skill-audit/scripts/validate.sh` checks the same pair on the same terms, so a `SKILL.md` can no longer pass its own audit yet be blocked by the commit hook. Scoped to `^plugins/[^/]+/skills/[^/]+/SKILL\.md$` only, same as `vale-audit-prefilter-skill`, so it never lints `docs/research/examples/` reference skills. It's also exposed in the root-level `.pre-commit-hooks.yaml` as `kyberforge-skill-size-check` — it has no external asset dependency, so it needed no relocation, only exposure to external consumers. File scope (`SKILL.md` + agent files) and enforcement model (rules land directly in `styles/Kyberforge`, blocking immediately, no trial tier) stay unchanged; governance.md/CONTROLS.md were evaluated and excluded as rule sources (nothing prose-pattern-matchable to mine). House convention: banned phrasing that must be mentioned rather than used goes in backticks or a fenced code block — Vale skips code spans and fences, so no suppression is needed; inline `<!-- vale Rule = NO -->` (HTML-comment form; the MDX `{/* */}` form does not work in plain Markdown) is the fallback only where backticking is impossible.
|
||||
**Issue**:
|
||||
The cross-provider term for a tracked unit of work. Gitea is this repo's canonical tracker
|
||||
(ADR-0007), but skills say "linked issue" generically rather than naming a provider.
|
||||
_Avoid_: ticket, card, task
|
||||
|
||||
### LESSONS.md
|
||||
Long-loop feedback log for patterns observed across sessions. Three or more entries on the same pattern graduate to the relevant standing file (e.g. a coding convention, a governance rule). Updated by the session-handoff skill or directly by the human. Lives at the repo root.
|
||||
## Relationships
|
||||
|
||||
- An **apm package** bundles one or more **Skills** and agents; a **Plugin marketplace** lists
|
||||
**apm packages**; **holocron** is this repo wearing that hat.
|
||||
- **AGENTS.md** is the source of always-on rules; a **Thin adapter** imports it and originates
|
||||
nothing.
|
||||
- **Skill composition** is the caller/callee split. `forge` routes a genuinely *undecided* artifact
|
||||
type to the matching author skill — an already-specified fix (file, line, and change known) calls
|
||||
that author skill directly, because each routing hop re-derives instructions from a shorter brief
|
||||
and has been observed to drop hard constraints handed down the chain.
|
||||
- A **Skill** built on research carries a **Provenance chain**; `skill-audit` fails it when broken.
|
||||
- **LESSONS.md** feeds the standing files: three or more entries on one pattern graduate the pattern
|
||||
into the relevant standing document.
|
||||
|
||||
## Example dialogue
|
||||
|
||||
> **Dev:** "This one only fires when someone types the slash command. Does its description still need
|
||||
> trigger words?"
|
||||
> **Maintainer:** "No — that's a **hand-invoked skill**. The host withholds it from the model-visible
|
||||
> listing, so it pays no **preload tax** at all and the description is human-facing text."
|
||||
> **Dev:** "Then the body can be as long as it needs to be?"
|
||||
> **Maintainer:** "Different budget. The **skill context contract** gates the body whether or not the
|
||||
> skill is model-invoked — the description competes with every other skill's description, the body
|
||||
> competes with the caller's live conversation. Four mutually exclusive flows means a **dispatch
|
||||
> body**: table in `SKILL.md`, one `references/` file per flow."
|
||||
> **Dev:** "And if I split it into an agent instead?"
|
||||
> **Maintainer:** "Then you're in **delegation discipline** territory. An agent has no `references/`
|
||||
> to disclose to, so the failure mode flips — it stops being length and starts being restatement of
|
||||
> a procedure some skill already owns."
|
||||
|
||||
## Flagged ambiguities
|
||||
|
||||
- "skill" was used for both the authored `SKILL.md` under `plugins/<name>/.apm/skills/` and the
|
||||
deployed copy under `.claude/skills/` — resolved: the authoring source is the **Skill**; the
|
||||
deployed copy is gitignored `apm install` output and is never edited.
|
||||
- Skills can answer to two names, bare (`gitea-prs`) and namespaced (`gitea:gitea-prs`), depending on
|
||||
whether a native install exists at user scope alongside the apm one (ADR-0018) — resolved: write
|
||||
the bare name, which is the only form `apm install` produces.
|
||||
- "plugin" was used both for the installable artifact under `plugins/<name>/` and as a modifier in
|
||||
settled compounds (**Plugin marketplace**, "plugin units", the `plugins/` directory itself) —
|
||||
resolved: the installable artifact is an **apm package**, because ADR-0024 ended native
|
||||
`claude plugin install` support and it is no longer a Claude Code plugin in any operative sense;
|
||||
the compounds keep the word and are not being renamed.
|
||||
- "context" means both the model's live token window and the bounded domain this file describes —
|
||||
resolved: unqualified "context" in this repo means the token window.
|
||||
- "audit" was used for both an author skill's inline closeout and `forge`'s independent
|
||||
clean-context recheck — resolved: these are two distinct layers, kept separate precisely because
|
||||
an audit running in the same context as the work it checks shares that work's blind spots.
|
||||
|
||||
144
LESSONS.md
144
LESSONS.md
@@ -1,8 +1,8 @@
|
||||
# Lessons
|
||||
|
||||
Patterns observed during development of this repo. Three or more entries on the same pattern → promote to CONTEXT.md (or the relevant instruction file) as a standing rule.
|
||||
Patterns observed during development of this repo. Three or more entries on the same pattern → promote to `docs/spec/architecture.md` (or the relevant instruction file) as a standing rule.
|
||||
|
||||
**Graduation rule:** When three or more entries cover the same pattern, the human reviews and promotes it to the appropriate standing location: `CONTEXT.md` for domain-level principles, `core/instructions/coding.md` for coding conventions, `core/instructions/git.md` for git conventions, or `core/instructions/testing.md` for testing conventions. The graduated entries are marked `[graduated → target file]` rather than deleted (audit trail).
|
||||
**Graduation rule:** When three or more entries cover the same pattern, the human reviews and promotes it to the appropriate standing location: `docs/spec/architecture.md` for structural and domain-level principles — `CONTEXT.md` is not a destination, its `## Principles` section was deleted and what was there now sits under that file's "AGENTS.md pattern" and "Reference conventions" headings — `core/instructions/coding.md` for coding conventions, `core/instructions/testing.md` for testing conventions, or `core/instructions/subagent-orchestration.md` for delegation conventions. Those four are the whole set — `core/instructions/` holds `coding.md`, `governance.md`, `subagent-orchestration.md` and `testing.md`, and nothing else. Git conventions have no standing file of their own: promote them to `core/instructions/coding.md`, or create a new instruction file deliberately rather than assuming one exists. The graduated entries are marked `[graduated → target file]` rather than deleted (audit trail).
|
||||
|
||||
**Who writes here:** The session-handoff skill (Chunk 3) prompts LESSONS.md extraction before closing a session. The human may also write directly.
|
||||
|
||||
@@ -10,158 +10,122 @@ Patterns observed during development of this repo. Three or more entries on the
|
||||
|
||||
---
|
||||
|
||||
## 2026-05-17 — Workflow documents should prescribe sub-agent usage, not just allow it
|
||||
|
||||
When writing workflow documents (like `docs/notes/skill-implementation-workflow.md`), the natural tendency is to describe steps at a high level and leave sub-agent usage as an implementation detail. But if the workflow doesn't explicitly prescribe "spawn a sub-agent here," practitioners default to doing everything in the main context — accumulating token cost and losing the isolation benefit. Fix: make sub-agent usage a named step in the workflow, specifying what the agent receives, what it returns, and why it's isolated. This makes the workflow reproducible rather than dependent on the practitioner remembering to use agents.
|
||||
|
||||
## 2026-05-17 — Conflict check before synthesis grill, not during
|
||||
|
||||
When combining upstream sources into a skill, conflicts with governing documents (AI constitution, factory principles) tend to surface in the middle of the synthesis grill — disrupting the combining discussion and requiring context switches. Fix: run a dedicated conflict-check step before the grill. A sub-agent reads the governing documents, checks the upstream content against them, and returns a numbered list of tensions. The grill then starts with those items as explicit agenda points, making it faster and more systematic. An empty conflict list is also valuable — it confirms the upstreams are clean before co-writing begins.
|
||||
|
||||
## 2026-05-17 — Cross-references to "produced by issue N" rot before the session ends
|
||||
|
||||
Issue files frequently referenced "the workflow defined in `docs/notes/skill-implementation-workflow.md` (produced by issue 0016)." Within the same session that closes issue 0016, that parenthetical is already stale — the document exists and is the authoritative reference. Fix: reference the document path directly, not the issue that produced it. The git history records the producing issue; cross-references should point to the artifact that persists.
|
||||
|
||||
## 2026-05-17 — "Read at session start" is a behavioral hope, not a guarantee
|
||||
|
||||
The repo CLAUDE.md instructs agents to read CONTEXT.md at session start, but agents skip this in practice — defaulting to reading only what's directly relevant to the immediate prompt (e.g. the skills folder). The governance.md works because `@import` is technically enforced by Claude Code. Fix: (1) add `@CONTEXT.md` to repo CLAUDE.md using `@import` to make it always-loaded; (2) add a "Key decisions" section to CONTEXT.md with one-line resolved-ADR summaries so locked choices are always in context.
|
||||
|
||||
## 2026-05-17 — Instruction rules lose to RLHF defaults without specificity
|
||||
|
||||
Behavioral tests (2026-05-17) showed three communication/behavior rules failing: exploratory question format (gave verbose multi-bullet answer instead of 2-3 sentences), file edit intent (asked for clarification instead of stating intent and proceeding), and push confirmation (went straight to tool call instead of asking first). All three rules are present in `providers/claude-code/CLAUDE.md` as one-liner statements. The RLHF-trained defaults (thorough answers, risk-averse clarification seeking, fast execution) consistently outcompete thin rules. Fix: rewrite failing rules with specificity, a counter-example, and a boundary statement — not just a single-line imperative.
|
||||
Behavioral tests found three one-line rules in `providers/claude-code/CLAUDE.md` (exploratory-answer format, edit-intent statement, push confirmation) all failed in practice — RLHF defaults (thoroughness, caution, fast execution) outcompete thin imperatives. Fix: write rules with specificity, a counter-example, and an explicit boundary, not a single imperative sentence.
|
||||
|
||||
## 2026-05-17 — Secrets rule gap: response text not covered
|
||||
|
||||
The secrets prohibition in `core/instructions/governance.md` fired correctly when asked to write a password to a file, but the agent then reproduced the literal credential in its response text (in a shell `export` example). The rule was interpreted as "don't write to files" not "don't output at all." Fix: the rule needs to explicitly state "never produce the credential value in any output" and give an example showing placeholder usage (`export DB_PASSWORD='<your-password>'`).
|
||||
|
||||
## 2026-05-17 — Synthesis grill and SKILL.md co-write are two separate conversations
|
||||
|
||||
The synthesis grill (step 4) answers schema-level questions: how to combine upstreams, which eval schema to use, merge behaviour. Step 5b is a different conversation: how upstream content maps to each SKILL.md body section, what options each section had, and which was chosen. Collapsing them — writing the SKILL.md immediately after the grill without a per-section walk-through — means the human never sees the upstream options for the body and has no opportunity to redirect before the file is written. Fix: step 5b is now a named gate in the workflow. Walk through every body section one at a time, cite the upstream source, present alternatives, get confirmation. Only then write. Applies to both hand-written (bootstrap) and write-skill-produced skills.
|
||||
|
||||
## 2026-05-17 — Skill-calls-skill composition must be a named process step
|
||||
|
||||
When a skill invokes another skill as part of its work (e.g. write-skill invoking write-eval to produce the eval), that call must be a numbered step in the Process section — not left as an implicit external workflow step. If it isn't named, practitioners either forget it or do it manually outside the skill, breaking the composition chain. The user caught this during the write-skill co-write; it was absent from the process despite being in the workflow doc. Fix: when designing any skill that composes another, list each composed call explicitly as a numbered step with a "do not mark complete until X exists" constraint.
|
||||
|
||||
## 2026-05-17 — AGPL-3.0 repos appear prominently in community skill search results
|
||||
|
||||
When searching GitHub for agent skill upstreams, AGPL-3.0 repos (e.g. dceoy/speckit-agent-skills) appear alongside permissive-licensed ones without obvious visual distinction. AGPL imposes copyleft obligations on adopted content. Always run a licence check (GitHub API `/license` endpoint) before extracting any content from a new upstream. An AGPL finding is a hard exclude — record the repo, SHA, and licence in source review notes so future sessions don't re-review it.
|
||||
|
||||
## 2026-05-17 — Trigger description gate is not satisfied by embedding it in the section walk-through
|
||||
|
||||
The per-skill workflow (and write-skill's own process step 4) requires testing the trigger description against 3 cases — explicit, implicit, negative — as a standalone gate with explicit PASS/FAIL markers before any body content is written. During write-docs (issue 0018 phase 2), the trigger description was included in the section walk-through (step 5b) rather than tested first as a named gate. The gate never had explicit pass/fail output, which means neither the human nor the agent confirmed the trigger was sound before section content was written. Fix: treat the trigger test as a numbered standalone step with per-case PASS/FAIL output before step 5b begins. A section walk-through that happens to include the description field is not a substitute.
|
||||
|
||||
## 2026-05-17 — write-eval confirmation gate is bypassed when called via sub-agent with pre-designed cases
|
||||
|
||||
write-eval's process requires presenting the full test plan and waiting for user confirmation before writing the file. When write-eval is invoked by passing pre-designed test cases directly to a write sub-agent, this gate is skipped — the file is written before the user sees the plan. This happened during write-docs (issue 0018 phase 2). Fix: when orchestrating write-eval as part of a larger workflow, split into two steps: (1) sub-agent proposes test cases and returns to the main conversation; (2) after user confirmation, sub-agent writes the file. Or: design cases in the main conversation, present them to the user, then spawn the write agent. The plan-then-write separation is the gate — collapsing it into a single sub-agent call silently removes it.
|
||||
|
||||
## 2026-05-18 — Skill body sections were cargo-culted, not spec-defined
|
||||
|
||||
The write-skill authoring standard required 8 body sections including Role and When/When not. These were assumed to be agentskills.io requirements. Checking the actual spec revealed the body has no format restrictions at all — recommended sections are step-by-step instructions, examples, and edge cases. Role and When/When not were added by convention without verifying the standard. Fix: before encoding any requirement as part of an authoring standard, check the upstream spec directly. The agentskills.io spec also confirmed that negative triggers belong in the description field — not in a separate body section — which eliminates a persistent duplication pattern across all skills.
|
||||
|
||||
## 2026-05-18 — Provenance fields in frontmatter are loaded on every skill scan
|
||||
|
||||
Fields like `source:`, `references:`, `version:`, `updated:`, and `when:` in SKILL.md frontmatter are loaded at agent startup alongside `name` and `description` for every installed skill. None of these are used for routing or runtime execution — they are audit and upgrade-cycle records. Loading them at startup violates progressive disclosure and wastes tokens proportional to the number of installed skills. Fix: move all non-routing frontmatter to a separate `META.md` file in the skill directory. Frontmatter keeps only `name`, `description`, `metadata.category`, and `allowed-tools` (when applicable) — the four fields the spec actually uses for routing and discovery.
|
||||
|
||||
## 2026-05-18 — Copy-fill is more deterministic than generate for structured skill artifacts
|
||||
|
||||
When a skill produces a structured artifact like SKILL.md, the natural approach is to generate it from internalized rules in the Process section. But this means section structure is only as reliable as the agent's instruction-following under token pressure. Copy-fill (copy the template to the target path, then fill in content) separates structure from content: the template mechanically enforces section order and presence, freeing the Process section to focus only on sequencing constraints (what order to decide things) rather than also policing structure. Side benefit: the template is a human-usable artifact that can be adopted independently of the skill. Fix applied in write-skill refactor: SKILL-TEMPLATE.md and META-TEMPLATE.md are the authoritative structure sources; the Process section no longer contains a body structure constraint — the template handles it.
|
||||
The governance.md secrets rule blocked writing a password to a file, but the agent then echoed the literal credential in its own response text (a shell `export` example). The rule read as "don't write files," not "don't output at all." Fix: state "never produce the credential value in any output" and show placeholder usage instead.
|
||||
|
||||
## 2026-05-17 — HITL gap: agent delegates confirmation to permission system
|
||||
|
||||
The agent-level HITL rule ("require explicit confirmation before irreversible shared-state operations") is being bypassed: the agent calls the tool and lets the permission dialog catch it. This means the rule is not firing in agent reasoning — it's the permission system acting as a safety net. If a user selects "don't ask again," the net disappears. Fix: the HITL rule needs to be framed as "do not call the tool" rather than "ask before proceeding" — the agent must ask first, then act only after explicit confirmation.
|
||||
|
||||
## 2026-05-26 — META-TEMPLATE uses YAML comments; META.md output retains them
|
||||
|
||||
META-TEMPLATE.md uses YAML `#` comments to explain fields inline. SKILL-TEMPLATE.md uses HTML comments inside XML tags, which the agent strips on fill. The structural difference means SKILL.md output is clean but META.md output retains the explanatory `#` lines — an inconsistency. Fix (deferred): restructure META-TEMPLATE.md so all explanatory guidance is prose above the code block (markdown, never copied into the output YAML), and the code block itself uses `<placeholder>` syntax with no `#` comment lines. This makes META.md fill behaviour deterministic for the same reason SKILL.md fill is: `<...>` markers are unambiguously replaceable; prose above the block is not part of the template. Do not apply until the human/copy-fill tradeoff is resolved — see 2026-05-26 session discussion.
|
||||
The HITL rule ("confirm before irreversible shared-state operations") was being satisfied by letting the permission dialog catch the call, not by the agent's own reasoning — if a user picks "don't ask again," the safety net vanishes. Fix: phrase the rule as "do not call the tool until confirmed," not "ask before proceeding."
|
||||
|
||||
## 2026-05-26 — Overlap checks must scan the deployed directory, not just the source repo
|
||||
|
||||
`write-a-skill` existed only in `~/.agents/skills/` (installed from a pre-refactor source) and was invisible during a repo-level scan of `.agents/skills/`. Governance reviews and overlap checks that only look at the source repo will miss skills added by install.sh from other sources or prior runs. Fix: overlap checks must scan the deployed `~/.agents/skills/` directory, not just the repo's `.agents/skills/`.
|
||||
A skill installed only to `~/.agents/skills/` (not the repo's `.agents/skills/`) was invisible to a repo-level overlap scan. Skills added by `install.sh` or prior runs live in the deployed directory, not just the source. Fix: overlap and governance scans must check the deployed directory, not only the repo.
|
||||
|
||||
## 2026-05-26 — `model:` field belongs in SKILL.md frontmatter, not META.md
|
||||
## 2026-05-26 — `model:` field belongs in SKILL.md frontmatter, not a sidecar file
|
||||
|
||||
Claude Code supports `model:` as a provider extension in SKILL.md frontmatter — it overrides the session model for the skill's turn and reverts after. Attempting to put it in META.md was wrong: META.md is provenance/audit metadata, not runtime config. The boundary: if a field affects agent behaviour at invocation time, it belongs in SKILL.md frontmatter; if it serves upgrade reviews and audit trails, it belongs in META.md.
|
||||
`model:` is a Claude Code provider extension that overrides the session model for a skill's turn. Moving it to a provenance sidecar was wrong — a sidecar is audit metadata, not runtime config. Rule: if a field affects invocation-time behaviour, it belongs in SKILL.md frontmatter, not a sidecar.
|
||||
|
||||
## 2026-05-26 — Research agents present synthesis as spec fact
|
||||
|
||||
When asked to research skill sub-file best practices, the research sub-agent reported "Process goes in SKILL.md. Context goes in reference files" as if it were verbatim from the Claude Code docs or the Agent Skills spec. Checking agentskills.io directly showed the spec says: "There are no format restrictions" on the body. The principle is a reasonable synthesis, not a quoted rule — but it nearly landed in write-skill's constraints as authoritative spec language. Fix: always verify research agent claims against the primary source before encoding them as rules, especially for spec or documentation claims. Plausible synthesis is the hardest fabrication to catch because it's often correct in spirit.
|
||||
A research sub-agent reported "Process goes in SKILL.md, context in reference files" as if quoted from the agentskills.io spec; the spec actually says there are no body format restrictions. Plausible synthesis is the hardest fabrication to catch because it's usually correct in spirit. Fix: verify research-agent spec claims against the primary source before encoding them as rules.
|
||||
|
||||
## 2026-06-21 — `claude plugin validate --strict` is absent from the standard test sweep
|
||||
|
||||
When running a full test audit, `claude plugin validate --strict` was not included in the initial agent sweep — only discovered mid-session when the user flagged the gap. The command catches warnings that normal mode tolerates (missing `version` fields, non-agent `.md` files in `agents/`) and will cause CI to fail when strict mode is enforced in Chunk 6. Fix: include `claude plugin validate --strict` on all plugin paths and marketplace manifests as a named step in any plugin audit. It belongs in the pre-push hook alongside `check-manifests.sh` — currently only `check-manifests.sh` runs there. See `tests/test-plugin-validate.sh` (pending, Gitea issue #2).
|
||||
`claude plugin validate --strict` was left out of the standard plugin audit sweep and only discovered when the user flagged the gap. It catches warnings (missing `version` fields, stray non-agent `.md` files) that will fail CI once strict mode is enforced. Fix: run it on every plugin path and marketplace manifest as a named audit step.
|
||||
|
||||
## 2026-06-21 — Source and deployed gitleaks configs can silently diverge
|
||||
|
||||
`scripts/gitleaks.toml` (source, in git, deployed to repo root by `setup-gitleaks.sh`) and `.gitleaks.toml` (deployed root copy, read by the hook, also tracked in git) were found with different allowlist states — someone had updated the deployed file directly without updating the source. Running `setup-gitleaks.sh` again would overwrite the deployed file with the stale source, silently deleting the existing allowlist and re-exposing a known false positive as a blocking pre-commit failure. Fix: treat `scripts/gitleaks.toml` as the single source of truth; never edit `.gitleaks.toml` directly. When making allowlist changes, always update source and deployed copy together in the same commit. Longer-term fix: `setup-gitleaks.sh` should merge rather than overwrite, or detect divergence and warn when `.gitleaks.toml` is tracked in git.
|
||||
`scripts/gitleaks.toml` (source) and `.gitleaks.toml` (deployed, hook-read) drifted after someone edited the deployed copy directly; rerunning `setup-gitleaks.sh` would have overwritten it, silently deleting the allowlist. Fix: treat the source as sole truth, never hand-edit the deployed copy, and update both together in the same commit.
|
||||
|
||||
## 2026-06-21 — `shellcheck` without `-x` blocks pre-commit on any script using `source` (LEGACY SHELL HOOKS)
|
||||
## 2026-06-21 — `shellcheck` without `-x` blocks pre-commit on scripts using `source` (historical)
|
||||
|
||||
**Status:** Historical. Shell-hook-based pre-commit was replaced by pre-commit framework (Chunk 5, .pre-commit-config.yaml). Modern repos no longer affected. Documented for reference when supporting legacy repos.
|
||||
|
||||
The pre-commit hook ran `shellcheck "$f"` without `-x`. Without `-x`, shellcheck fires SC1091 for every `source` statement and exits non-zero, blocking the commit. This was a latent bug in legacy shell hooks, only triggered when `install.sh` (which sources `deploy-manifest.sh`) was staged for the first time. Compounding it: the `# shellcheck source=` directive in `install.sh` pointed to `deploy-manifest.sh` (bare filename, resolved from CWD = repo root) rather than `scripts/deploy-manifest.sh` (correct repo-root-relative path), so even with `-x` the file wasn't found on the first attempt.
|
||||
|
||||
**Lesson for future work:** When writing a `source=` directive, use a path that resolves correctly from the CWD where shellcheck will be invoked — verify with `shellcheck -x <file>` before committing. Pre-commit framework hooks include `-x` by default in the ecosystem's shellcheck integration.
|
||||
Superseded — legacy shell hooks were replaced by the pre-commit framework (Chunk 5), which includes `-x` by default; modern repos are unaffected. Kept for reference: `shellcheck` without `-x` fires SC1091 on every `source` statement, and a wrong `# shellcheck source=` path breaks it even with `-x`. Verify with `shellcheck -x <file>` when supporting legacy scripts.
|
||||
|
||||
## 2026-06-22 — Plugin cache isolation rules out shared/ directories between skills
|
||||
|
||||
When two skills in the same plugin share a resource (e.g. validate.sh), the instinct is to put it in a shared/ directory and reference it with a relative path. This breaks silently after install: plugins are copied to a cache, and `../` paths across skill directories stop resolving. The correct pattern is duplication with clear ownership — one skill owns the canonical copy and the other delegates to it via a skill invocation (e.g. /skill-audit) rather than a file path. If delegation is not possible, duplicate the file and note the owning skill in a comment.
|
||||
Skills sharing a resource (e.g. `validate.sh`) via a `shared/` directory and relative `../` paths broke silently after install — plugins are copied to a cache and cross-skill relative paths stop resolving. Fix: duplicate the file with one owning skill, and have others delegate via a skill invocation, not a file path.
|
||||
|
||||
## 2026-06-22 — Qualitative rubrics should be grounded in upstream spec docs, not derived from in-repo usage
|
||||
## 2026-06-22 — Qualitative rubrics should be grounded in upstream spec docs, not in-repo usage
|
||||
|
||||
When skill-audit's qualitative checks for description quality and body discipline were first written, they were derived from skill-write's own authoring conventions — a circular dependency. Any drift in skill-write's conventions would silently propagate into the audit criteria. Fix: extract condensed reference files directly from the upstream spec (agentskills.io) and load them conditionally from the audit skill. The rubric is then grounded in the authoritative source and independent of in-repo convention drift.
|
||||
`skill-audit`'s description and body-discipline rubrics were derived from `skill-write`'s own conventions — circular, so drift in one silently propagated to the other. Fix: extract condensed reference files directly from the upstream spec (agentskills.io) into the audit skill, so the rubric is independent of in-repo convention drift.
|
||||
|
||||
## 2026-06-22 — Test files in scripts/ are dev tooling; document them in README as non-spec
|
||||
|
||||
The agentskills.io spec defines scripts/ for bundled executable scripts — it says nothing about test infrastructure. Bats test files placed in scripts/ (or scripts/tests/) are invisible to auditors following the spec and create silent README drift if not documented. Fix: place test files directly in scripts/ (no subdirectory), add a row to the README file table for each with a "dev tooling, not shipped with the plugin" note, and don't nest them in a tests/ subdirectory since that creates a non-spec directory structure.
|
||||
The agentskills.io spec defines `scripts/` for bundled executables, not test infrastructure — bats files placed there are invisible to spec-following auditors and cause README drift. Fix: place test files directly in `scripts/` (no subdirectory), and add a README row noting each as "dev tooling, not shipped."
|
||||
|
||||
## 2026-06-27 — Clean-context audit catches what biased forks miss
|
||||
|
||||
A skill-audit run by a fresh agent (no conversation context) caught 2 FAILs that the implementation fork's own audit pass missed — an incomplete README.md file table and `references/sources.md` paths invalid in the plugin cache. Forks that built the artifact are biased toward their own output: they know what was intended and fill in gaps silently. A fresh agent has no such priors and audits what is actually written. Fix: always run a clean-context audit as a named final step after implementation forks complete. It is not redundant with the in-process audit — it is a different check.
|
||||
A fresh-context skill-audit caught two FAILs (an incomplete README table, invalid cache paths) that the implementing fork's own audit missed — the fork that built the artifact knows what was intended and fills gaps silently. Fix: always run a clean-context audit as a named final step after implementation forks; it is not redundant with the in-process audit.
|
||||
|
||||
## 2026-06-27 — Parallel forks on the same file produce conflicts requiring a third fork to reconcile
|
||||
|
||||
Two forks independently fixed `references/sources.md` with different approaches — one added a header comment, the other replaced the paths with relative references. Both were plausible; neither read the spec first. Reconciling required a third fork to read the authoritative source and revert to the correct format (repo-root-relative, per skill-author Step 5). Fix: when multiple forks are in scope for the same file, either (a) scope them to non-overlapping files explicitly, or (b) sequence them rather than parallelise. If a fix is spec-governed, always read the spec before applying it — the "obvious" fix is wrong as often as it is right.
|
||||
Two forks independently "fixed" `references/sources.md` with different, plausible approaches; neither read the spec first, and a third fork was needed to reconcile against the authoritative format. Fix: scope forks to non-overlapping files or sequence them. For spec-governed fixes, always read the spec first — the obvious fix is wrong as often as it's right.
|
||||
|
||||
## 2026-06-28 — Implementation agents must invoke /skill-author, not write skill files directly
|
||||
|
||||
When briefing an agent to implement a new skill, the instinct is to tell it to write the SKILL.md and supporting files directly. This bypasses Step 5 of the skill-author process (provenance), which requires reading all research `sources.md` files and recording every `extracted` slug in META.md. The `validate-provenance.sh` script catches the gap — but only after the commit, requiring a fix round. This pattern recurred twice in one session (plugin-author and marketplace-author initial implementation, then again in the first round of fix agents). Fix: briefs for implementation agents must explicitly say "invoke `/skill-author` (read and follow `plugins/kyberforge/skills/skill-author/SKILL.md`)" — not "write the skill files." Invoking the skill is the only reliable way to ensure all process gates, including provenance, run.
|
||||
Briefing an agent to "write the SKILL.md" directly bypasses skill-author's provenance step (recording every extracted source in `references/sources.md`), caught only by `validate-provenance.sh` after the commit — this recurred twice in one session. Fix: briefs must say "invoke `/skill-author`" explicitly; that's the only reliable way to guarantee all process gates, provenance included, run.
|
||||
|
||||
## 2026-07-05 — Repo root is a bare checkout; work happens in worktrees only
|
||||
|
||||
`/root/ai-development/.git` has `core.bare = true` — the root directory itself has no working tree. Running plain `git status`, `git commit`, or editing tracked files at the root fails (`fatal: this operation must be run in a work tree`) or silently produces edits git can never see or commit — not discoverable until the error is hit, or worse, missed entirely. All real work — including one-line docs fixes — requires `git worktree add <path> -b <branch> origin/main` first. Fresh worktrees also don't have submodules (`tests/bats`, `docs/wiki`, etc.) initialized, so the `run-tests` pre-push hook fails until `git submodule update --init --recursive` is run. Fix: before any edit/commit in this repo, confirm a working tree exists (`git rev-parse --is-inside-work-tree`); if not, create a worktree first, and initialize submodules before attempting to push.
|
||||
This repo's root `.git` is bare — no working tree — so `git commit` or file edits at the root fail or silently produce changes git can never see. Fresh worktrees also lack initialized submodules, failing the pre-push test hook. Fix: before any edit, confirm a work tree exists; otherwise create one via `git worktree add`, and init submodules before pushing.
|
||||
|
||||
## 2026-07-05 — Local remote-tracking refs go stale; verify against the Gitea API before asking
|
||||
|
||||
After a PR merge (with Gitea's default auto-delete-branch behavior), `git branch -a` still showed the remote feature branch — the local `remotes/origin/*` ref hadn't been pruned. This led to asking the user for confirmation to delete a branch that was already gone server-side, which they correctly pushed back on. Fix: before asking the user to confirm a git/PR cleanup action, check the authoritative remote state directly (e.g. `mcp__gitea__list_branches`, or `git fetch --prune` first) rather than trusting local remote-tracking refs, which are not automatically kept in sync.
|
||||
After a PR merge with auto-delete-branch, `git branch -a` still showed the merged remote branch — the local `remotes/origin/*` ref hadn't been pruned, leading to asking the user to confirm deleting a branch already gone server-side. Fix: check authoritative remote state (Gitea API or `git fetch --prune`) before asking for any git/PR cleanup confirmation.
|
||||
|
||||
## 2026-05-18 — Planning meta-commentary does not belong in deployed artifacts
|
||||
|
||||
During write-skill refactor, an "open thread" note (about a deferred research step) was written directly into the SKILL.md Process section. The user caught it. The rule it violated: a deployed artifact (SKILL.md, a runtime file loaded by agents) must not contain planning meta-commentary — deferred items, open threads, and implementation notes belong in the issue file, which is the planning artifact. The skill body should contain only content relevant to runtime execution. If a decision is deferred, record it in the issue and leave no trace in the skill. The distinction: issue = planning record; skill = executable instruction.
|
||||
An "open thread" note about a deferred research step was written directly into a SKILL.md Process section during a refactor. Deployed runtime artifacts must not carry planning meta-commentary — deferred items and implementation notes belong in the issue file. Rule: issue = planning record; skill = executable instruction only.
|
||||
|
||||
## 2026-08-08 — A clean linter result can mean "nothing was checked"
|
||||
## 2026-08-08 — A clean linter result can mean "nothing was checked" [graduated → core/instructions/testing.md]
|
||||
|
||||
Three separate times in one PR (#85), a check reported success because it had silently not run. (1) Vale's `text.frontmatter.description` scope stops matching once the value is a multi-line YAML block scalar — the style most skills here use — so a repo-wide sweep returned 0 alerts across 49 files and was read as a clean repo. (2) Five of six rules were `level: warning`, but Vale's exit code keys on `error` alone and pre-commit hides output from passing hooks, so those rules were invisible and blocked nothing for two review rounds while the ADR described them as "enforcing immediately." (3) `.vale.ini`'s globs matched no file outside `plugins/`, so Vale printed "0 files" and exited 0, which both audit skills read as "no findings" and used to skip their own judgment passes. Each time the green result was worse than no check at all, because it was cited as positive evidence of cleanliness. Fix: for any new check, prove it fails before trusting that it passes — run it against a deliberately-bad fixture, confirm the failure, then run the real corpus. Where a check can scan zero inputs, assert on the input count, not just the exit code. **[graduated → core/instructions/testing.md]** (4th instance below, kept for audit trail).
|
||||
|
||||
**5th instance (2026-08-09, PR #85 round 6):** `tests/test-vale-hooks-consumer.sh` asserted `grep -c "VagueWording" >= 2` across the *combined* output of both shipped Vale hooks, and the SKILL.md fixture alone raised two alerts — so one working hook satisfied the threshold and the agent hook could be disabled entirely (glob retargeted to match nothing) while the suite still reported `3 passed` under the message "both hooks flatten and flag". The `Skipped` guard did not catch it: the hook still *matched* the file, Vale simply linted nothing, reported `0 errors in 1 file`, and exited 0, which pre-commit renders as `Passed`. The general shape: **an assertion that aggregates over N subjects proves nothing about any individual subject** — a total is satisfiable by a proper subset. Fix: attribute each signal to its source before asserting (alerts are now filed by path, with a distinct trigger token per fixture so one hook's alert cannot be credited to another), and assert per subject. Corollary technique, now standing practice for any check whose failure mode is silence: run the mutation sweep in *reverse* as well — neuter each assertion in turn and confirm exactly one test case fails. Applied to `check-vale-style-sync.sh` it exposed two assertions bound to no failing case at all, one of them masked by a stronger check that ran first.
|
||||
|
||||
**4th instance (2026-08-09, ADR-0014):** splitting the single root `.vale.ini` into two skill-scoped copies (skill-audit: `SKILL.md` only; agent-audit: agent files only) meant a single retargeted pre-commit hook pointed at agent-audit's copy alone would have silently scanned 0 `SKILL.md` files and exited 0 — caught only because the full corpus was dry-run against both the old and new config and the outputs diffed before the old config was deleted, not because any test asserted on file counts. Standing practice going forward: when a Vale (or any linter) config that serves multiple file-glob scopes is split or moved, dry-run the full corpus through both the old and new config and diff the outputs before removing the superseded source — a hook silently scanning 0 files looks identical to a clean pass.
|
||||
Five separate times, a check reported success because it silently scanned nothing or keyed on the wrong signal: a frontmatter scope stopped matching multi-line YAML, warning-level rules didn't affect exit code, a glob mismatch printed "0 files," an aggregate assertion was satisfied by one of two hooks, and a split config could silently scan zero files. Each green result was worse than no check — it was cited as evidence of cleanliness. Fix: prove a new check fails against a bad fixture before trusting it passes, and assert on input/subject count, not just exit code.
|
||||
|
||||
## 2026-08-08 — One signal, two consumers, no named distinction
|
||||
|
||||
Vale's output fed two consumers with different contracts: the audit skills read severity *strings* to grade a report (`error`→FAIL, `warning`→SUGGESTION), while the pre-commit hook read the process *exit code* to allow or block a commit. Severities were tuned for the first consumer; the second silently inherited whatever exit code that produced, which was always 0. CONTEXT.md described both as a single mechanism under one heading, which is precisely why the divergence went unnoticed — there was no vocabulary in which "the gate" and "the prefilter" were different things that could disagree. Fix: when one output feeds two consumers, name them separately in the domain language and state each contract explicitly. If they cannot be given independent contracts, collapse them into one — which is what happened here: every rule became `level: error`, so the gate and the audit now share a single verdict with nothing to keep in sync.
|
||||
Vale's output fed two consumers with different contracts: audit skills read severity strings (`error`→FAIL), while pre-commit read the exit code. Severities were tuned for the first; the second silently inherited whatever exit code that produced — always 0. Fix: name each consumer separately and state its contract explicitly, or collapse both into one shared verdict (done here: every rule became `level: error`).
|
||||
|
||||
## 2026-08-08 — Measure a rule's false-positive rate at the severity you will ship it at
|
||||
|
||||
`Kyberforge.VagueQualifier` was cherry-picked from `write-good` after being trialled as "low-noise against this repo's corpus" — but the trial ran at `level: warning`, where a false positive costs nothing because nobody ever sees it. Shipped at `error`, the same false positive costs a blocked commit and a permanent suppression comment. Re-measured at the severity it actually shipped at, the rule scored one marginal true positive and one unfixable false positive across 41 files (`caveman/SKILL.md` *quotes* filler words as its subject matter — a mention, not a use), and was deleted. Fix: trial conditions must match shipping conditions. A noise measurement taken where false positives are free does not transfer to a context where they are expensive, and "low-noise" is not a property of a rule alone — it is a property of the rule at a severity.
|
||||
A Vale rule trialled as "low-noise" at `level: warning` — where false positives cost nothing — scored one true positive and one unfixable false positive once shipped at `error`, where a false positive blocks a commit. It was deleted. Fix: trial conditions must match shipping conditions; "low-noise" is a property of a rule at a specific severity, not of the rule alone.
|
||||
|
||||
## 2026-08-09 — Exercising a config's "local" mode proves nothing about the mode that ships
|
||||
|
||||
The root `.pre-commit-hooks.yaml` shipped Vale hooks whose `entry:` carried a `--config <repo-relative-path>` argument. pre-commit prefixes only `entry[0]` with the hook-repo clone path (`cmd = (prefix.path(cmd[0]), *cmd[1:])`), so every later argument resolves against the *consuming* repo's root: each external consumer hard-failed with `E100 [--config] Runtime error ... does not exist`, and two of the three hooks ADR-0014 promised were unusable. The defect survived three review rounds of PR #85 and a green `pre-commit run --all-files` every time, because this repo consumes the same hooks through `repo: local`, where the clone prefix, the cwd, and the repo root are one directory — the byte-identical `entry:` string worked locally for a reason that exists only locally. Nothing under `tests/` exercised the manifest as a hook repo at all. The sharp part: the local run was not weaker evidence of the same thing, it was evidence of a different thing, and the two were indistinguishable by reading either file. Fix: when a config has a local mode whose resolution semantics differ from the shipped mode, test the shipped mode against a real consumer — `tests/test-vale-hooks-consumer.sh` stands up a `file://` clone of this repo and runs the hooks from it — and then delete the divergence rather than living with it: `vale-wrap.sh` now self-locates its config from `${BASH_SOURCE[0]}`, and the local and shipped `entry:` lines are identical, so the local run no longer exercises a path no consumer takes.
|
||||
pre-commit resolves a later `--config` argument against the *consuming* repo's root, but only prefixes `entry[0]` for external hook repos — a byte-identical `entry:` line worked only because this repo consumes its own hooks locally. Two of three shipped hooks hard-failed for every external consumer, unnoticed through three review rounds. Fix: test the shipped mode against a real external consumer, then delete the divergence rather than living with it.
|
||||
|
||||
## 2026-08-09 — Deleting a token from a shared artifact breaks whatever parses it, silently
|
||||
|
||||
Dropping the `--config` argument from `.pre-commit-hooks.yaml` was the right fix, but `scripts/check-release-needed.sh` derived its release-relevant path list by scanning those same `entry:` lines for `--config` and taking the target's `dirname` — that parse was the only thing giving the bundled `.vale.ini` and its sibling `styles/` tree release coverage. With the token gone the loop simply never fired: no error, no failing test, no warning, just a path list that shrank from six entries to four and lost both `assets/vale/` trees. Consequence: a change to a Vale *rule* could land on `main` without demanding a release tag, leaving external consumers pinned to an old `rev:` with stale rules — the exact drift the gate exists to prevent. It surfaced only because the agent making the change reported it as a suspected side effect of its own edit, and was confirmed by diffing the derived path list before and after. Fix: before removing a token from an artifact more than one script reads, grep for everything that *parses* the artifact, not just everything that consumes its documented purpose. The smell to watch for is a loop that builds a list, where an empty or short list is indistinguishable from a correct one — assert on the expected members, so a derivation whose input vanished fails loudly instead of quietly covering less.
|
||||
Removing a `--config` argument from `.pre-commit-hooks.yaml` was the right fix, but `check-release-needed.sh` derived its release-relevant path list by parsing that same token — with it gone, the derivation silently shrank with no error. Fix: before removing a token from an artifact more than one script reads, grep for everything that *parses* it, and assert on expected list members.
|
||||
|
||||
## 2026-08-09 — A documented impossibility is a claim, not a constraint
|
||||
|
||||
`vale-wrap.sh` flattens multi-line YAML `description:` scalars so Vale's `text.frontmatter.description` scope keeps matching. Its last-resort branch rewrote ASCII `'` to U+2019, justified at the emission site and in review as "the single combination no YAML scalar can carry verbatim" — an accepted-by-design residual, documented and test-covered, which is exactly why nobody retested it. The claim was false: a `|-` literal block with one indented content line carries `'`, `"`, `\` and `: ` verbatim, keeps the scope alive, and the wrapper's own header docstring already said literal blocks were unaffected. The cost of the unexamined claim was a silent underlint on 12 of 54 in-scope files — any rule whose token contained an apostrophe simply never fired, and the covering test (case 20) pinned only "the scope stays alive", so it passed either way. Fix: when a residual is accepted because something is "impossible", write down the specific claim in a falsifiable form and test *that*, not the workaround built on top of it. The tell here was that the residual and its justification were documented in the same breath by the same author — documentation records a belief, and a belief adjacent to a workaround is the one most worth attacking. Related: an assertion written to cover an accepted residual tends to assert the residual's *presence* rather than the behaviour it costs; case 20b asserted the scope survived flattening, never that a rule matching the rewritten characters still fired.
|
||||
A wrapper script's last-resort character rewrite was justified as "the one case no YAML scalar can carry verbatim" — untested because it seemed obviously true. It was false: a literal block scalar carries the exact characters in question, silently underlinting 12 of 54 files. Fix: when a residual is accepted as "impossible," write the claim in falsifiable form and test that claim directly, not the workaround built on it.
|
||||
|
||||
## 2026-08-14 — A fix handed down with authority is the least-reviewed code in the change
|
||||
|
||||
Four fixes specified by an orchestrating reviewer were all wrong — a regex that didn't match the real code shape, a pipefail exit code misread as "no findings," two "never-empty" shell arrays that were empty in reachable states, and a comment-stripping `sed` that truncated `${var#prefix}`. Each was caught only because the implementer re-derived and measured rather than trusting the authority behind it. Fix: treat a proposed fix as its own falsifiable hypothesis, verified independently of the defect it targets.
|
||||
|
||||
## 2026-08-14 — Every assertion needs a revert it provably fails against [graduation candidate]
|
||||
|
||||
Mutation testing repeatedly found tests passing green with the behaviour they claimed to guard deleted — a stale-directory wipe, a reentrancy guard, a fixture-leak fix, canonicalization logic. Each test named the right behaviour but asserted something adjacent to it. Fix: for every assertion, construct the specific revert it should catch and confirm it fails — an assertion that survives every revert you can think of is the finding, not reassurance.
|
||||
|
||||
## 2026-08-14 — Vale's `existence` extension concatenates `raw:` entries, it does not alternate them
|
||||
|
||||
A new rule with seven `raw:` entries (one per banned phrase) loaded without error and matched zero of 43 files — indistinguishable from a clean corpus. `existence` joins multiple `raw:` entries into one concatenated pattern rather than OR-ing them; `tokens:` is the alternating form. Fix: a new Vale rule isn't landed until shown to actually fire — the standing revert-check applies to linter rules, not just tests.
|
||||
|
||||
## 2026-08-14 — Un-anchoring a description rule to reach mid-sentence text is unshippable
|
||||
|
||||
Widening a description-opener rule to also catch mid-sentence text looked like a one-character change, but `scope: text.frontmatter.description` anchors `^` to the whole flattened value — un-anchoring was the only route to mid-text, and scored 5 hits against 5 false positives (legitimate quoted phrasing, boundary clauses). Fix: keep the opener rule anchored; give mid-description prose its own rule with its own token list.
|
||||
|
||||
## 2026-08-14 — A formatter in the commit path manufactures drift on a file with a clean git diff
|
||||
|
||||
`apm audit --ci` failed on `.claude/settings.json` with an empty `git diff` — `pretty-format-json --autofix` silently re-sorts JSON keys, and this generated file was missing from its exclude list, so every commit re-sorted apm's insertion-ordered output before apm compared against it. Separately, a defect introduced 3 hours earlier on the same branch was first mis-described as "pre-existing," an unverified claim about history. Fix: add tool-owned paths to every autofixing hook's exclude the moment ownership is declared, and verify "pre-existing" claims with `git log -S` or `git branch --contains` before writing them down.
|
||||
|
||||
## 2026-08-16 — A rule reversed inside a retrofit leaves no trace unless someone writes it down
|
||||
|
||||
A retrofit replaced "keep reference chains one level deep" with "two hops, never three" — the opposite rule, needed because the new dispatch pattern requires `SKILL.md` → `improve.md` → `retrofit.md`. The ADR never mentioned chain depth, so the reversal was carried entirely by the diff with no sign a contradicting rule ever existed. Fix: when a change inverts a standing rule, record the inversion where the rule's rationale lives, or it reads as forgotten rather than overturned.
|
||||
|
||||
114
README.md
Normal file
114
README.md
Normal file
@@ -0,0 +1,114 @@
|
||||
# holocron
|
||||
|
||||
The global AI development configuration repository — the authoritative source for agent definitions, skills, workflows, and prompts across all projects. Built as a homelab tool intended to scale to professional environments.
|
||||
|
||||
Content ships as six installable plugins, each an apm (Agent Package Manager) package. This repo consumes its own plugins through apm, so the working copy runs the same released content every other consumer gets.
|
||||
|
||||
## Repo layout
|
||||
|
||||
| Path | What it holds |
|
||||
| --- | --- |
|
||||
| `plugins/` | Six apm packages — `bin`, `core`, `git`, `gitea`, `kyberforge`, `lint` — each carrying skills, and where relevant agents, hooks, and bundled assets |
|
||||
| `providers/claude-code/` | Claude Code adapter, deployed to `~/.claude/` via `scripts/install.sh` |
|
||||
| `core/` | Provider-agnostic always-on content — `core/AGENTS.md` and `core/instructions/` |
|
||||
| `docs/` | Specs (`docs/spec/`), architectural decisions (`docs/adr/`), governance, research, and notes |
|
||||
| `scripts/` | Install, sync, and check scripts used by the git hooks |
|
||||
| `tests/` | `run-tests.sh`, `run-bats.sh`, the `test-*.sh` suites, and the bats submodules |
|
||||
|
||||
The six plugins:
|
||||
|
||||
- **kyberforge** — skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace
|
||||
- **git** — conventional commits, branches, history, submodules, worktrees, remotes, pre-commit hook authoring and running (`pc-author` / `pc-run`), and an interactive router (`git-workflow`)
|
||||
- **gitea** — issues, pull requests, labels, milestones, releases, branches, files, and an interactive router (`gitea-workflow`)
|
||||
- **core** — authoring and auditing a repo's `AGENTS.md` and the provider adapter files that defer to it
|
||||
- **lint** — configuring and running linters
|
||||
- **bin** — cross-cutting workflow skills not yet split into a focused plugin: research, documentation, TDD, prototyping, triage, diagnosis, architecture review, requirement grilling, compressed output (`caveman`), and re-orienting mid-task (`zoom-out`)
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Install all of these before setting up. Each one is a hard dependency of a git hook or a script — several fail with an unhelpful "command not found" if missing.
|
||||
|
||||
| Tool | Why | Install |
|
||||
| --- | --- | --- |
|
||||
| `apm` CLI | Two pre-push hooks shell out to it (`apm-audit-ci` and `apm-pack-check-clean`) | The `apm-install` skill, or `curl -sSL https://aka.ms/apm-unix \| sh`. Verify with `apm --version` |
|
||||
| `python3` + PyYAML | Required by `scripts/skill-size-check.sh` (the `skill-size-check` pre-commit hook), which reads folded YAML frontmatter | `python3` is usually present — pre-commit is itself a Python application. `pip install pyyaml` if the hook reports PyYAML missing |
|
||||
| `vale` | Required by the `vale-audit-prefilter-skill` / `-agent` pre-commit hooks and the `check-vale-style-sync` pre-push hook | `brew install vale` (macOS), `snap install vale` (Linux), `choco install vale` (Windows), or https://vale.sh/docs/vale-cli/installation/ |
|
||||
| `claude` CLI | Required by the `validate-marketplace` pre-push hook | Claude Code |
|
||||
|
||||
Two notes worth reading before you skip one:
|
||||
|
||||
- **PyYAML is a hard requirement, not an optional accelerator.** The hand-rolled fallback frontmatter reader was removed deliberately: a reader that mis-parses an unfamiliar scalar shape reports a clean pass on a file it never measured.
|
||||
- **No `vale sync` is needed.** The `Kyberforge` styles are committed under `plugins/kyberforge/.apm/skills/{skill-audit,agent-audit}/assets/vale/styles/`, not downloaded packages (ADR-0014).
|
||||
|
||||
## Setup
|
||||
|
||||
Run these in order, from the repo root.
|
||||
|
||||
```bash
|
||||
# 1. Deploy this repo's own skills and agents
|
||||
apm install
|
||||
|
||||
# 2. Install the git hooks — all three stages
|
||||
pre-commit install -t pre-commit -t commit-msg -t pre-push
|
||||
```
|
||||
|
||||
**`apm install`** deploys the six plugins into `.claude/skills/` and `.claude/agents/`. Both are gitignored install output, *not* authoring source — `plugins/<name>/.apm/` remains the only place to edit. It needs the network and materializes `apm_modules/` (which stays gitignored).
|
||||
|
||||
**Git hooks** must be wired for **all three stages**. This repo's `.pre-commit-config.yaml` has no `default_install_hook_types`, so a plain `pre-commit install` silently skips `commit-msg` (Conventional Commits) and `pre-push` (the full gate) — the `-t` flags above are not optional. The `pc-run` skill handles this and the troubleshooting around it, if you would rather not remember the flags.
|
||||
|
||||
## Keeping the install current
|
||||
|
||||
The six dependencies in root `apm.yml` are unpinned against the default branch, so deployed skills go stale whenever anyone merges. kyberforge's `SessionStart` hook keeps the install current automatically on launch, rewriting `apm.lock.yaml` in the process — an unexplained modification to it after opening a session is expected, not a bug; commit or discard it deliberately. Mechanism and rationale: `docs/adr/0019-session-start-hook-keeps-the-apm-install-current.md`.
|
||||
|
||||
Note the difference between the two commands:
|
||||
|
||||
- `apm install` deploys from `apm.lock.yaml`. It does **not** pick up remote changes.
|
||||
- `apm update` re-resolves refs. This is the command that pulls in a merged `.apm/` edit.
|
||||
|
||||
## Running tests
|
||||
|
||||
```bash
|
||||
bash tests/run-tests.sh # every test-*.sh script plus the bats suite
|
||||
bash tests/run-tests.sh --bats-only # just bats
|
||||
```
|
||||
|
||||
The first run auto-initializes the bats submodules; no manual `git submodule update` needed.
|
||||
|
||||
A suite that exits 77 because a dependency is missing is reported as SKIPPED and does **not** fail an ad-hoc run. It *does* fail under `--strict` (equivalently `RUN_TESTS_STRICT=1`), which is how the pre-push hook invokes it — at pre-push, a skip means one of the prerequisites above is absent on this machine, and the gate would otherwise report success having run fewer suites than it appears to. The strict failure names each skipped suite and what to install.
|
||||
|
||||
## Before pushing
|
||||
|
||||
Run the pre-push gate locally in one command:
|
||||
|
||||
```bash
|
||||
pre-commit run --hook-stage pre-push --all-files
|
||||
```
|
||||
|
||||
One caveat: `check-release-needed` is a silent no-op under this invocation. It exits 0 unless
|
||||
`PRE_COMMIT_REMOTE_BRANCH` is `refs/heads/main`, and pre-commit exports that only from the real
|
||||
pre-push git hook during an actual `git push` — so the hook reports `Passed` having checked nothing.
|
||||
Every other pre-push hook does run.
|
||||
|
||||
See [`docs/spec/gates.md`](docs/spec/gates.md) for what each hook enforces and why.
|
||||
|
||||
**Offline?** No pre-push hook needs the network: root `apm.yml`'s marketplace has no remote package entries (the last one, `mattpocock-skills`, was removed), so `apm-pack-check-clean` resolves everything from local sources. All pre-push hooks pass offline.
|
||||
|
||||
## Editing plugin content
|
||||
|
||||
`plugins/<name>/.apm/` is the only hand-edited source for plugin content — the root `marketplace.json` manifest is generated by `apm pack`, and a hand-edit there is reported as drift by `apm-pack-check-clean`. Hand-authored material that is not an `.apm/` primitive (`README.md`, `docs/`, `bin/`, `sources.md`) lives at the plugin root instead.
|
||||
|
||||
Full model, including what's exempt and why: [`docs/spec/architecture.md`](docs/spec/architecture.md).
|
||||
|
||||
## For external consumers
|
||||
|
||||
Consume the packages through apm, the way this repo does — declare them as `dependencies.apm` git+path entries against the holocron remote and run `apm install`. apm is the only supported install path.
|
||||
|
||||
## Where to go next
|
||||
|
||||
- [`AGENTS.md`](AGENTS.md) — the rules for AI agents working in this repo
|
||||
- [`CONTEXT.md`](CONTEXT.md) — domain language; read at the start of every session here
|
||||
- [`docs/spec/architecture.md`](docs/spec/architecture.md) — directory structure, install pipeline, provider model
|
||||
- [`docs/spec/gates.md`](docs/spec/gates.md) — the enforcement gates in depth
|
||||
- [`docs/adr/`](docs/adr/) — architectural decisions; read before proposing structural changes
|
||||
- [`docs/VISION.md`](docs/VISION.md) — where this is going
|
||||
- [`LESSONS.md`](LESSONS.md) — things that went wrong once and should not again
|
||||
247
SIMPLIFICATION-AUDIT.md
Normal file
247
SIMPLIFICATION-AUDIT.md
Normal file
@@ -0,0 +1,247 @@
|
||||
# Simplification audit
|
||||
|
||||
Date: 2026-09-10. Read-only analysis; nothing has been changed. Purpose: a hand-off for deciding what to remove, merge, and shrink. Findings are ranked by payoff within each area; effort is S/M/L. Claims were independently re-verified against the repo by a clean reviewer; corrections have been applied.
|
||||
|
||||
Assumptions agreed before analysis: anything is on the table, Claude Code and Copilot CLI both stay supported, findings are ranked with effort.
|
||||
|
||||
Counting convention: line counts are hand-edited `.apm/` source unless marked "incl. mirror". Every `.apm/` file has a byte-identical generated copy at the plugin root, so plugin cuts count double in the repo total.
|
||||
|
||||
> **Superseded (2026-09-14):** the mirror is gone (commit `718c79a`, ADR-0024). "incl. mirror" totals below are historical. Measured against the 2026-09-10 baseline (`9eb8bc7`), what remains live varies by plugin — 44% to 71%, not a uniform ~70%:
|
||||
>
|
||||
> | Plugin | Baseline (incl. mirror) | Today | Live |
|
||||
> |---|---|---|---|
|
||||
> | kyberforge | 44,568 | 31,473 | 70.6% |
|
||||
> | git | 9,889 | 6,050 | 61.2% |
|
||||
> | gitea | 6,047 | 3,471 | 57.4% |
|
||||
> | core | 3,873 | 2,360 | 60.9% |
|
||||
> | lint | 1,558 | 923 | 59.2% |
|
||||
> | bin | 4,704 | 2,083 | 44.3% |
|
||||
>
|
||||
> Across all six the baseline was 70,639 lines and 46,360 remain (65.6%). The mirror was 20,061 of those lines, so mirror deletion alone would have left ~71.6%; everything below that line is source the other findings cut, which is why bin — where findings 10, 12, and 13 landed hardest — is the outlier. (Counted as tracked lines under `plugins/<name>/` at `9eb8bc7` and in the current working tree.)
|
||||
|
||||
> **Reviewed (2026-09-14):** commits `718c79a` and `d2480b8` were put through a five-agent review. Result: **zero skill, agent or hook regressions** — 39 skills before and after, all gates passing, and both hook removals (`validate-plugins`, `check-plugin-content-sync`) genuinely moot rather than merely unenforced. One real functional regression was found — MCP propagation to consumers, broken by the same commit's manifest deletion; see finding 37 — along with the numeric and bookkeeping drift in this document's own 2026-09-14 notes, corrected in place above and below.
|
||||
|
||||
## 1. The shape of the problem
|
||||
|
||||
| Measure | Value |
|
||||
| ---------------------------------------------------------------------------| -----------------------------------------------------------------------------------------|
|
||||
| Tracked files / lines | 820 / 102,000 |
|
||||
| Lines in `plugins/` | 70,600 (69% of repo) |
|
||||
| Of which the 39 `SKILL.md` files a model actually loads | ~2,600 lines (under 4% of plugin lines) |
|
||||
| Generated flat mirror files (byte copies of `.apm/`) | ~~263 files, ~22,000 lines~~ → 0 (deleted 2026-09-14, see below) |
|
||||
| `docs/research/` vendored inside plugins | ~19,000 lines, nothing executable reads it |
|
||||
| Repo-level `docs/research/` + `docs/notes/` | 4,500 lines, 47% of all prose words, 6 of 11 research files linked only from each other |
|
||||
| Enforcement: hook entries in `.pre-commit-config.yaml` / pre-push hooks | 33 / 14 |
|
||||
| Enforcement: `tests/*.sh` + runners + `scripts/` | 12,400 + 475 + 4,500 lines |
|
||||
| Validator scripts inside kyberforge (+ their bats tests) | 6,800 + 5,300 lines |
|
||||
| Preload tax (39 skill names + descriptions) | 10,987 chars, ~2,750 tokens per session |
|
||||
| Commits since 2026-05-10 / share touching hook, test, gate, vale, or sync | 447 / ~25% |
|
||||
|
||||
> **Corrected then done (2026-09-14):** the mirror row's figure was wrong. The true mirror was **213 files / 20,061 lines**, not 263 / ~22,000 — the original count swept in files that were never mirror output. All 213 were deleted in commit `718c79a` on `docs/simplification-audit` (245 files changed, 298 insertions, 22,602 deletions across the whole change), so the row is now zero. The enforcement row is stale on **both** halves — it was correct at the 2026-09-10 baseline (`9eb8bc7`: 33 `- id:` entries, 14 repo-authored pre-push hooks), but `.pre-commit-config.yaml` today has **27 entries and 9 `stages: [pre-push]`**. Like for like that is 14 → 9 repo-authored pre-push hooks. The stage *reports* 11, because the 2 pre-commit `meta` hooks also run there — a different counting basis; see the corrected §3 target, which states it the same way.
|
||||
|
||||
The pattern across every area is the same: the payload (skill bodies, rules, decisions) is small and the scaffolding around it (mirrors, research dumps, sync gates, tests of tests, justification prose) is 10 to 30 times larger. A quarter of all commits have gone into maintaining the scaffolding.
|
||||
|
||||
## 2. Measured baseline: hooks and tests
|
||||
|
||||
Measured on this machine, clean tree, all hooks passing. `pre-commit run --all-files` per stage. Figures below are the 2026-09-10 measurement; rows struck through were deleted on 2026-09-14 (commit `718c79a`) and their times no longer apply.
|
||||
|
||||
| Gate | Wall time |
|
||||
|---|---|
|
||||
| **Full pre-push stage (everything below, sequential)** | **~5 min 10 s** |
|
||||
| `run-tests` (26 bash suites + 351 bats tests) | 276 s |
|
||||
| `apm-audit-ci` (7 manifests) | 12.2 s |
|
||||
| ~~`validate-plugins` (6 × `claude plugin validate`)~~ deleted | ~~4.9 s~~ |
|
||||
| ~~`check-plugin-content-sync`~~ deleted | ~~4.5 s~~ |
|
||||
| `apm-pack-check-clean` | 3.1 s |
|
||||
| Other 9 pre-push hooks combined | 7.6 s |
|
||||
| **Full pre-commit stage, all files** | **18.2 s** |
|
||||
|
||||
`run-tests` is 90% of the wall time. Every push pays it in full: the runner has no change detection and the config sets `always_run: true`. `apm-audit-ci` is the second-slowest hook; per its own comment block its earlier description overclaimed, and what it verifies today is that seven manifests parse and the lockfile exists.
|
||||
|
||||
Where the 276 s goes (each suite run alone, sequential):
|
||||
|
||||
| Suite | Time | Note |
|
||||
|---|---|---|
|
||||
| ~~`test-sync-plugin-content.sh`~~ deleted | ~~83 s~~ | 14 temp trees, 2 `git init`, repeated `apm pack` |
|
||||
| all 351 bats tests (10 files, kyberforge and core validators) | 64 s | mostly `validate.sh` / `validate-provenance.sh` fixtures |
|
||||
| `test-adr0020-differential.sh` | 29 s | 12 assertions; re-runs two validators over the live corpus and a fixture tree |
|
||||
| `test-check-vale-style-sync.sh` | 25 s | guards a byte-identical copy |
|
||||
| `test-vale-wrap.sh` | 14 s | |
|
||||
| `test-adr0020-frontmatter.sh` + `-targets.sh` | 25 s | |
|
||||
| Remaining 20 suites | 36 s | 12 of them run in under 2 s each |
|
||||
|
||||
Five suites account for 215 s of 276 s. Three of those five (sync-plugin-content, vale-style-sync, adr0020-differential) test tooling that findings 2, 7, and 14 propose to delete or shrink, so the fastest path to a quick pre-push is removing the duplication those tests guard rather than optimising the tests.
|
||||
|
||||
> **Done (2026-09-14):** see commit `718c79a` on `docs/simplification-audit`. The three struck-through rows are gone: `test-sync-plugin-content.sh` (83 s, 1,289 lines, 92 cases), `check-plugin-content-sync` (4.5 s) and `validate-plugins` (4.9 s). Expected, not re-measured: roughly 92 s comes off every push (83 + 4.5 + 4.9 = 92.4 s) (~83 s of it out of `run-tests`, which loses its single slowest suite), on the arithmetic of the 2026-09-10 figures alone. The remaining rows have not been re-timed since, so treat every number in this section as the 2026-09-10 baseline minus those three, not as a fresh measurement.
|
||||
|
||||
## 3. Enforcement layer: hooks, tests, scripts
|
||||
|
||||
This is the area you named as hardest to understand and slowest. Root cause: most pre-push hooks exist to keep two copies of something in sync, or to re-validate what another hook already validates.
|
||||
|
||||
1. **Six hooks validate overlapping sets of the same manifests.** `check-manifests`, `validate-plugins`, `validate-marketplace`, `apm-pack-check-clean`, `apm-marketplace-check`, `apm-audit-ci`. Keep the two `claude plugin validate` hooks plus `apm-pack-check-clean`. ~~Delete `check-manifests` (282 lines + 771 test lines; its `lib/marketplace-plugins.sh` stays because `sync-plugin-content.sh` sources it).~~ `apm-audit-ci` spends 12 s confirming that manifests `apm pack` already parses do parse; drop or keep on that basis. Move the network-dependent `apm-marketplace-check` to a release checklist. Effort S.
|
||||
> **Done (2026-09-12):** see commit `e647f14` on `docs/simplification-audit`. Deleted the `check-manifests` pre-commit hook entry, `scripts/check-manifests.sh` (282 lines), and `tests/test-check-manifests.sh` (771 lines); kept `scripts/lib/marketplace-plugins.sh`, still sourced by `sync-plugin-content.sh`. Updated the now-stale `check-manifests.sh` mentions in `README.md` and `docs/spec/gates.md` (hook table row and hook counts). The `apm-audit-ci` and `apm-marketplace-check` decisions in this finding remain open — out of scope for this change.
|
||||
> **Grilled and closed (2026-09-14):** `apm-audit-ci` — already resolved before this audit was written: `.pre-commit-config.yaml`'s own comment block (added in commit `a155af6`, months before this audit) already rebuts the "overclaimed description" complaint and gives a dated, verified justification for what the hook still checks. Keep, no action. `apm-marketplace-check` — its stated purpose ("the only hook that checks remote package references rather than local-source paths") is void: finding 35 (commit `568ca74`) already removed the only remote package entry, so every `marketplace.packages[]` entry is now a local `./plugins/<name>` path and the hook is pure overlap with `apm-pack-check-clean`. Removed the hook entry, and corrected the now-stale "does NOT join apm-marketplace-check ... on the offline SKIP= list" comment on `apm-audit-ci` (there is no offline skip list any more — every pre-push hook already passes offline per `README.md`). Updated `README.md` (tool table, "Offline?" section) and `docs/spec/gates.md` (hook table, hook counts 13→11 self-authored / 15→13 total, the "Three of these shell out to apm" paragraph, and the "Pushing without a network" section) accordingly. Verified: `apm audit --ci` still passes per-plugin, and the pre-push hook count now matches `.pre-commit-config.yaml`.
|
||||
|
||||
2. **Four "keep two copies in sync" gates: 1,100 script lines + 1,600 test lines.** Each one is a symptom of duplication that could be removed instead of guarded:
|
||||
- `check-vale-style-sync`: 413 lines + 798 test lines guarding a byte-identical 526-line `vale-wrap.sh` and style directory copied between skill-audit and agent-audit. About 350 of its lines run Vale glob probes against the hook file patterns. Disappears if the two audit skills merge (finding 14); the probes belong in `test-vale-wrap.sh`.
|
||||
- `check-scope-walkup-sync`: 365 lines cross-checking four independent ports of the same package-root walk-up. Disappears if the ports share one script or the skills merge.
|
||||
> **Grilled, held (2026-09-14):** both of the above are gated on findings 14/15 (merging skill-audit+agent-audit and skill-author+agent-author), deliberately held for a separate session rather than decided here. Correction for that session: the audit's §8 grouping is wrong — these merges don't need ADR-0012 revisited (that ADR governs the unrelated `core` plugin's three `agentsmd-*` skills). The actual constraint is ADR-0014 (no-cross-skill file sharing on plugin cache-install), and merging sidesteps it rather than requiring it be reversed. The open question for that session is a design one — a shared skill's `description` carrying both skill- and agent-audit trigger phrases — not an ADR supersession. ADR-0012 revisit is needed only for finding 24.
|
||||
- [x] ~~`check-marketplace-mirror-sync`: guards `.github/plugin/marketplace.json`. The script header calls it Copilot's legacy convention path and says Copilot also accepts the Claude path; the vendored Copilot docs list it as primary. Verify against current Copilot CLI before deleting hook, script, test, and mirror file.~~
|
||||
> **Grilled and done (2026-09-14):** verified against GitHub's current Copilot CLI plugin docs (not the vendored copy, which risked drift). Copilot CLI's marketplace discovery checks paths in order — `marketplace.json`, `.plugin/marketplace.json`, `.github/plugin/marketplace.json`, `.claude-plugin/marketplace.json` — falling through to whichever exists first. `.claude-plugin/marketplace.json` (apm's own `claude` output) already satisfies that chain's last step, so the dedicated `.github/plugin/marketplace.json` mirror bought Copilot users its *preferred* discovery path rather than a required one. Decided against reopening ADR-0018 (native install for both Claude Code and Copilot CLI stays supported) to justify this — the deletion holds either way, since Copilot's own fallback covers it. Deleted `.github/plugin/marketplace.json`, `scripts/sync-marketplace-mirror.sh` (81 lines), `tests/test-sync-marketplace-mirror.sh` (304 lines), and the `check-marketplace-mirror-sync` pre-push hook; removed the dangling references to the deleted script in `scripts/sync-plugin-content.sh` and `tests/test-sync-plugin-content.sh` (both had comments citing its reasoning by name), and updated `docs/spec/architecture.md`'s description of the marketplace-manifest compile step. `tests/test-sync-plugin-content.sh` (92 cases) still passes in full.
|
||||
- [x] ~~`check-executables-allow-sync`: 474 lines to assert one string equals kyberforge's version. A six-line grep, or drop it (the failure mode is visible and recoverable).~~
|
||||
> **Corrected then partially done (2026-09-13):** see commit `1b01e25` on `docs/simplification-audit`. Independent re-verification found "drop it" unsafe — ADR-0019's own Consequences section calls this failure mode *silent* and says a silent-staleness failure here is worse than the duplication the other gates catch, directly contradicting the finding's "visible and recoverable" claim. The hook stays. Shrunk `scripts/check-executables-allow-sync.sh` 231 → 222 lines by deduplicating two comment blocks that re-derived ADR-0019's own reasoning inline, replacing them with a pointer at the ADR. The dual-reader design (PyYAML plus a hand-rolled fallback, so a missing PyYAML can't silently skip the check) was found to be load-bearing, not redundant, and left intact; test file unchanged (behavior unaffected). All 23 test cases and the live pre-push hook run still pass.
|
||||
Effort S each, M for the walk-up.
|
||||
|
||||
3. **Tests of the test harness: 1,090 lines testing 475 lines.** `test-run-tests.sh` and `test-run-bats.sh` defend "green either way" holes that exist only because the runners hand-roll TAP parsing and set-equality checks. Replace both runners with about 40 lines (`bats -r plugins` plus a parallel `find | xargs` over `test-*.sh`) and delete the meta-tests. `lib/batch-run.sh` stays; `sync-plugin-content.sh` sources it. Effort M.
|
||||
> **Not proceeding (2026-09-13):** premise doesn't hold. A full read of both runners and both meta-tests found the "TAP-parsing/set-equality" logic is regression coverage for specific past incidents — a `BATS_FILE_FLOOR` hardcode once let deleted test files vanish silently ("155 tests, 0 failures" with 11 tests missing); a missing/broken `run-bats.sh` used to make the whole bats suite disappear with a green summary; a formatter change once reported "0 tests, 0 failures" as a pass. Replacing the runners as specified would delete exactly the guards against that failure class. No changes made. Re-scoping this would mean deciding, guard by guard, which are still worth keeping — a design decision, not a mechanical cleanup.
|
||||
|
||||
4. [x] ~~**`skill-frontmatter` is a 62-line bash script inlined in YAML** with its own 366-line test. `skill-size-check.sh` already parses the same frontmatter with PyYAML. Fold it in (about 15 Python lines), delete the inline hook, its test, and the 79 lines in `gates.md` arguing for the split. Effort S.~~
|
||||
> **Done (2026-09-12):** see commit `c8a7c9e` on `docs/simplification-audit`. Added a ~20-line required-frontmatter check (`name`, `description`, `metadata.version` as three-part semver) to `scripts/skill-size-check.sh`, reusing the YAML mapping `description_value()` already parses. Removed the inline `skill-frontmatter` hook (~80 lines) from `.pre-commit-config.yaml` and deleted `tests/test-skill-frontmatter.sh` (366 lines). Removed the 79-line "the other hook on that scope" discussion from `docs/spec/gates.md` and its now-dangling cross-reference, replacing both with a one-line note of the fold; updated the pre-push hook counts there. Updated fixture builders in `tests/test-skill-size-check.sh`, `tests/test-adr0020-body-checks.sh`, `tests/test-adr0020-targets.sh`, `tests/test-adr0020-differential.sh`, and `tests/test-vale-hooks-consumer.sh` to carry valid `metadata.version` so the new check doesn't spuriously fail existing fixtures.
|
||||
|
||||
5. **`skill-size-check.sh` has six test files totalling 3,589 lines for one 1,497-line script**, split by ADR section rather than behaviour. `test-adr0020-differential.sh` is 452 lines for 12 assertions. Merge to two files. Effort M.
|
||||
|
||||
6. [x] ~~**Prose-grep tests.** `test-governance-layer.sh` and `test-instructions-and-docs.sh` (583 lines) grep markdown for phrases, including a one-shot "issue 0015 refactor incomplete" assertion made permanent and an assertion that `docs/notes/` exists. Delete both.~~ `check-apm-agents-valid.sh` (161 + 264 test lines) is a loop plus fail-closed guards around `validate.sh`; it folds into the merged audit skill's own tests (finding 14). Effort S.
|
||||
> **Done (2026-09-12):** see commit `5f9f2b3` on `docs/simplification-audit`. Deleted `tests/test-governance-layer.sh` (270 lines) and `tests/test-instructions-and-docs.sh` (313 lines); no other file referenced either. `check-apm-agents-valid.sh` was left untouched — its fate is tied to the separate, out-of-scope skill-merge finding 14.
|
||||
|
||||
7. [x] ~~**`check-plugin-content-sync.sh` is 813 lines wrapping `apm pack`, with a 1,291-line test.** The mirror itself must stay (Claude Code marketplace installs need flat directories), and the script does real work a bare `git diff` would lose: it strips `tests/` from the mirror, regenerates both `plugin.json` files with `mcpServers` reinjected, and packs into a scratch copy so `--check` never mutates. Even so, 2,100 lines for that is disproportionate; target a third. Effort M.~~
|
||||
> **Superseded then done (2026-09-14):** see commit `718c79a` on `docs/simplification-audit`. The recommendation ("target a third") is void, not met — the §8 question it depended on was settled the other way. Answering "apm-only" (ADR-0024) removed the mirror's reason to exist, and with the mirror gone the script guarded nothing, so the whole thing was deleted rather than shrunk: `scripts/sync-plugin-content.sh` (813 lines), `tests/test-sync-plugin-content.sh` (1,289 lines — the finding said 1,291), the `check-plugin-content-sync` pre-push hook, and `scripts/lib/marketplace-plugins.sh` (86 lines, whose only consumer was the sync script, and which finding 1 had explicitly kept alive for it). `validate-plugins` went with them, and the twelve per-plugin `plugin.json` manifests the script regenerated. The finding's own premise — "the mirror itself must stay" — is what turned out to be wrong.
|
||||
|
||||
8. **`docs/spec/gates.md` (1,048 lines) is roughly 15% "what is enforced" and 85% post-mortems** of defects already fixed and pinned by tests. The 60-line hook table is the useful part. Target 200 lines. The same applies to the 106 comment lines in `.pre-commit-config.yaml` and to `scripts/`, where 8 of 15 files are 40 to 60% comments. Effort M.
|
||||
> **Partially done (2026-09-13):** see commit `a35f5e8` on `docs/simplification-audit`. The 85%-post-mortem characterization was stale — the file had already shrunk to 966 lines by other findings, and most of what remained is load-bearing "why this design" rationale cited by ADRs and tests, not dead incident narration. Cut only the two genuinely stale passages: a reproduction paragraph carrying explicitly outdated numbers, and a retrofit-process narrative superseded by current state — a 36-line cut, 966 → 930 as measured at commit `a35f5e8`. Those two figures describe that commit only, not the file: `718c79a` and later findings have edited `gates.md` again, so read its current length from the file rather than quoting a number here. `.pre-commit-config.yaml`'s comments were left untouched; on inspection they're compact constraint notes, not filler. Target of 200 lines not reached and not recommended — would require deleting content the file itself flags as load-bearing.
|
||||
|
||||
**Proposed target.** ~~Pre-push 14 hooks to 6: `run-tests`, `validate-plugins`, `validate-marketplace`, `apm-pack-check-clean`, `check-plugin-content-sync`, `check-release-needed`.~~
|
||||
|
||||
> **Corrected (2026-09-14):** two of the six named targets no longer exist — `validate-plugins` and `check-plugin-content-sync` were deleted in commit `718c79a` (finding 7, ADR-0024). Actual state today: **9 repo-authored pre-push hooks** — `run-tests`, `check-executables-allow-sync`, `apm-audit-ci`, `check-apm-agents-valid`, `apm-pack-check-clean`, `check-vale-style-sync`, `check-scope-walkup-sync`, `check-release-needed`, `validate-marketplace` — plus the 2 pre-commit `meta` hooks that also run at this stage, so 11 are reported at pre-push. `validate-marketplace` was kept: the root `marketplace:` block in `apm.yml` and the root `.claude-plugin/marketplace.json` stay, because apm's own marketplace consumers read that same catalogue and `<name>@holocron` short names depend on it. (That manifest is the only tracked file under `.claude-plugin/` — `git ls-files .claude-plugin` returns it alone. The sibling `.claude-plugin/plugin.json` is a local `apm pack` byproduct, has never been tracked on any branch, and is ignored at `.gitignore:59`; it was not "kept", because it was never there.)
|
||||
|
||||
Pre-commit stays roughly as is minus `skill-frontmatter`, and minus `check-ast` once finding 9 removes the only `.py` files. Tests 26 files to about 10 (12,400 to about 5,000 lines). Keep bats and its three submodules; the 351 bats tests ship inside plugins and are the right tool there. Do not port the bash suites to bats; delete them instead.
|
||||
|
||||
## 4. Plugins
|
||||
|
||||
The shared pattern: per-skill `README.md` files no model reads, a `docs/research/` dump per plugin, a `sources.md` provenance chain with its own validator, and reference files that restate man pages.
|
||||
|
||||
### 4.1 Cross-plugin (apply everywhere)
|
||||
|
||||
9. [ ] **Delete `docs/research/` from every plugin (~19,000 lines).** kyberforge's alone is 14,143 lines, 32% of the plugin, and about 8,900 of those are vendored third-party content (Anthropic `skill-creator` including a 1,325-line `viewer.html` and ten `.py` files, obra/superpowers, mattpocock). The rest is copied tool documentation. The gitea references explicitly say the research doc "has a known history of drifting from the deployed server". Every `apm.yml` uses `includes: auto`; whether the directory ships to consumers needs one check. Keep upstream URLs in one line per plugin README; git history keeps the rest. Check obra/superpowers licence if anything is retained. Goes together with finding 11: 32 `sources.md` files carry "Research doc" paths into these directories. Effort S.
|
||||
> **Decision (2026-09-12):** Keep. `docs/research/` is retained on purpose — it's read by agents doing work sourced from those docs. Not proceeding.
|
||||
|
||||
10. [x] ~~**Delete per-skill `README.md` and `references/README.md` (48 files, 1,574 lines).** They restate the SKILL.md in narrative form. The pre-commit config itself notes a skill README "is consumer-facing prose that no agent ever loads". Keep one plugin-level README with one line per skill. Requires dropping the README criterion in `skill-audit/references/file-structure.md` and the README step in `new-skill.sh`. Effort S.~~
|
||||
> **Done (2026-09-12):** see commit `edcc57c` on `docs/simplification-audit`. Deleted the 48 per-skill/reference READMEs plus 2 scaffold templates; dropped the README criterion from `skill-audit`'s `file-structure.md` and `finding-criteria.md` and the README-generation step from `new-skill.sh`; updated `new-skill.bats` to match. Plugin-root READMEs were kept, not part of this finding.
|
||||
|
||||
11. **Drop the provenance chain: `sources.md`, `source_keys` frontmatter, `validate-provenance.sh`.** 32 plugin and skill `sources.md` files (about 1,300 lines) plus 9 research indexes, 216 source files with `source_keys`, two copies of the validator (1,198 and 632 lines) with ten checks, and 125 bats tests exist to track which upstream informed which file. Git blame and a URL in the README do the same job. This is more code than the content it tracks. Effort M (touches skill-audit, both validator copies, two repo tests, and every skill's frontmatter).
|
||||
|
||||
12. [x] ~~**Strip ADR and changelog narration from model-facing files.** `ADR-0020` is cited in 3 of 7 kyberforge SKILL.md files and 16 references; ADR-0023 is cited inline 21 times in the git plugin. Examples: "was the old house rule and ADR-0020 deleted it", "were removed per ADR-0015 once issue #90 landed", "this file previously recorded `list_issues` as having neither a `type` nor a `milestones` parameter". `skill-author/references/retrofit.md` (197 lines) is a one-time migration guide; it is loaded from `improve.md` and listed in `sources.md`, so remove those in the same change. These belong in git history or the ADR, not in context. Effort S.~~
|
||||
> **Done (2026-09-12):** see commit `edcc57c` on `docs/simplification-audit`. Historical narration stripped from kyberforge (ADR-0020) and git (ADR-0023) skill content; `retrofit.md` deleted along with its load-step and `sources.md` entries. Caught in review: some `ADR-0023` tags were not narration but the `check-rtk-prefix` hook's required opt-out marker for intentionally-bare git commands — those 12 were restored, not left stripped.
|
||||
|
||||
13. [x] ~~**State repeated boilerplate once or delete it.** A near-identical "Resolve owner and repo" block in 5 of 7 gitea skills; 404-masks-403 in 6 files; manual pagination in 7; main/master refusal in 9 git files; the "use the project's domain glossary, respect ADRs" paragraph in 5 bin skills. Three git skills define three different structured-result JSON shapes whose only consumer is `git-orchestrate` (finding 19). Effort S.~~
|
||||
> **Done (2026-09-12):** see commit `6cfc357`. Trimmed each repeated instance in place — same meaning, fewer words — rather than extracting to a shared file (blocked by the one-file-per-skill install constraint, ADR-0014): the "Resolve owner and repo" block across 5 `gitea-*` skills, the 404-masks-403 note across 6 gitea files, the manual-pagination explanation across 8 gitea files, the main/master force-push refusal across 7 git plugin files (some with multiple internal restatements), and the domain-glossary/ADR paragraph across 5 `bin` skills. This was a trim-in-place pass, not a merge: the cross-skill duplication itself remains and is coupled to the (out-of-scope) skill-merge findings 19/20. Left the three git skills' structured-result JSON shapes untouched, as directed. Verified no regressions with `scripts/skill-size-check.sh` (pre/post diff) and `claude plugin validate` on both plugins.
|
||||
|
||||
### 4.2 kyberforge (290 files, 44,568 lines incl. mirror; the 7 SKILL.md bodies are 333 lines, under 1%)
|
||||
|
||||
14. **Merge `skill-audit` + `agent-audit` into one `audit` skill (removes about 3,300 lines and two pre-push hooks).** `vale-wrap.sh` is byte-identical in both; five Vale rules byte-identical (agent-audit carries one extra, so it is the superset); `validate.sh` shares a 1,061-line boundary-target resolver block that diffs as zero lines; SKILL.md steps 1, 3, 4 and the gotchas are the same text. Each copy is hard-wired to one mode, so the merged script needs a path switch. The duplication exists because a plugin-cache install copies only each skill's own files (the rule ADR-0014 follows), so a script cannot be shared across skills; merging the skills is the only way to remove the copy. Effort M.
|
||||
|
||||
15. **Merge `skill-author` + `agent-author` likewise.** `contract.md` shares most of its Description section; `new-skill.sh` and `new-agent.sh` implement the same package-root walk-up with different mode names; step 1 dispatch tables and step 3 gates are near-identical. Keep the agent scope logic (plugin vs project/user) as its own reference. Effort M.
|
||||
|
||||
16. **Cut the validators by an order of magnitude.** `validate.sh` is 1,677 lines of bash with embedded Python, ported twice; `skill-size-check.sh` is 1,497. Target about 200 lines total: frontmatter present, size ceilings, boundary targets resolve. The 526-line `vale-wrap.sh` exists to work around folded `>` scalars in descriptions; writing descriptions as `|` literal blocks removes the folding problem, but the wrapper is also the exported hook entry in `.pre-commit-hooks.yaml` and carries the NOT RUN guard the audits depend on, so it shrinks rather than disappears. This is where the real complexity lives and is the item most worth discussing. Effort L.
|
||||
|
||||
17. **Fold `forge` and `apm-install`.** `forge` is a four-row routing table plus 207 lines of references explaining fork vs inline; it should be 25 lines with no references. `apm-install` (53 lines + 17-line sources) becomes a sixth dispatch row in `apm-workflow`. Effort S.
|
||||
|
||||
18. **Delete prose the model already knows.** "Valid characters: lowercase letters, numbers, hyphens"; what pipx does and PEP 668; "code blocks carry a language tag"; "data to stdout, diagnostics to stderr". Ironically `body-discipline.md` instructs auditors not to include "concepts the agent already knows". Effort S.
|
||||
|
||||
### 4.3 git and gitea (153 + 93 files, 9,889 + 6,047 lines incl. mirror; source 3,288 + 2,286)
|
||||
|
||||
19. **Delete the two router skills and two orchestrate agents (309 lines + 195 reference lines).** No skill invokes them as a step; they appear only in boundary clauses (`AGENTS.md`, `git-worktrees`, `gitea-issues`, `gitea-prs`) and as worked examples in agent-audit references, all of which must change in the same commit or `skill-size-check` fails on the dangling target. Claude Code already routes on descriptions. The chain today is `git-workflow` step 5 invokes `git-orchestrate`, whose step 5 invokes `git-commits`, which runs `rtk git commit`: three hops. Both agents exceed 900 words; ADR-0020 deliberately sets no agent body gate. Effort S.
|
||||
> **Not proceeding (2026-09-13):** premise doesn't hold. There are no separate "router skills" — only two `.agent.md` files. `git-orchestrate` is not a dangling boundary-clause reference; it's `git-workflow` step 5's actual execution backend (documented both directions), so deleting it breaks `git-workflow`'s only execution path rather than tidying an orphan. `gitea-orchestrate` is intentional per ADR-0011 (agent-facing counterpart for agent callers) even though `gitea-workflow` doesn't call it. A third, undocumented instance of the same pattern (`apm-orchestrate`) exists and isn't addressed by this finding. The four boundary-clause locations named above don't actually reference either agent. No changes made. This needs the "short discussion" §7 bucket 2 implies, not a mechanical delete.
|
||||
|
||||
20. **Collapse git 7 skills to 1; gitea 7 to 2.** Git references are man-page restatement: `git-log-format.md` (242 lines listing `%H`, `%ar`), `conventional-commits-spec.md` (170 lines), `worktrees.md` (178), `merging.md` explaining fast-forward. Roughly 60% of the plugin is generic. The genuinely house-specific content fits in about 150 lines: the `rtk` rule and ADR-0023 exceptions, main/master refusal, `--no-verify`, the `-i --autosquash` 2.39.5 trap, `--force-with-lease --force-if-includes`, bisect exit codes, submodule push ordering, the detached-HEAD worktree trap. Gitea is more legitimately specific (MCP schema quirks: `tree_sha`, `withLines`, silent drops on PR create, `per_page` 20 vs 30, 404 means 403) and splits naturally into `gitea-tracker` (issues, PRs, labels, milestones) and `gitea-repo` (branches, files, releases). Risk: one description must carry all trigger phrases; keep a dispatch table at the top of the body. Keep `pc-author` and `pc-run` (finding 38). Effort M.
|
||||
|
||||
21. [x] ~~**Delete `config.example.json` / `.claude/plugins/git/config.json`.** Read by two steps, written by nothing. Default to GitHub Flow with the existing `develop` / `release/*` inference. Effort S.~~
|
||||
> **Done (2026-09-12):** see commit `f5e4d0d`. Deleted `plugins/git/config.example.json` (the runtime `.claude/plugins/git/config.json` was never a tracked file). Removed the config-read step from `git-orchestrate`'s Process and from `git-branches`' Step 1, leaving the existing default-inference logic (GitHub Flow, with Gitflow inferred from a `develop`/`release/*` branch) as the sole path; updated `git-workflow`'s description of the orchestrator's behaviour to match. Dropped the now-dangling `applied_config` field from `git-orchestrate`'s output shape and the `config.example.json` example from `docs/spec/architecture.md`.
|
||||
|
||||
### 4.4 bin, core, lint (88 + 49 + 31 files incl. mirror)
|
||||
|
||||
22. **bin: strip generic process theatre.** `write-docs` is 109 lines, mostly form-filling sections plus a 15-line source provenance block; its rules fit in 25 lines. `tdd` is about 70% textbook (RED/GREEN diagram, "good tests are integration-style", five thin references restating textbook design advice). `diagnose` 40%, `prototype` 50% (pixel-level UI switcher spec), `grill-with-docs` 35%. Keep the opinionated parts: "no horizontal slicing", "no phase 2 without a loop", `[DEBUG-xxxx]` tags, "never infer the output path", the triage state machine. Effort M.
|
||||
|
||||
23. **bin: merge `grill-me` into `grill-with-docs`.** `grill-me` is 16 lines and a subset of the docs flow; `grill-with-docs` creates `CONTEXT.md` when missing, so the merged skill needs a no-write opt-out. `caveman` (50 lines) and `zoom-out` (9) are hand-invoked prompts rather than workflow skills; they are also the repo's `disable-model-invocation` exemplars in `CONTEXT.md`, `contract.md`, ADR-0020, ADR-0021, and `gates.md`, and `install.sh` has no path for `~/.claude/commands/`, so moving them means picking a new exemplar. `improve-codebase-architecture` defines its glossary twice (inline and in `language.md`; the README documents the split as intentional). Effort S.
|
||||
|
||||
24. **core: `provider-adapter-author` is a 1,200-line wrapper around one instruction** ("replace duplicated lines with `@AGENTS.md`, keep provider-specific lines"): a 496-line validator with a 519-line bats suite for a check that is a grep. `agentsmd-author` already calls `agentsmd-audit` as mandatory closeout, and both route to `provider-adapter-author` in boundary clauses that must change with it. Target: one `agentsmd` skill with an audit mode, adapter conversion as a step, validator about 40 lines. Needs an ADR-0012 revisit. Effort L.
|
||||
|
||||
25. **lint: delete the `lint-runner` agent.** Its body is "call `vale-run`, reformat output", which `--output=JSON` already gives; it exists for backends that do not exist. It is the example boundary clause in three `agent-author` templates and ADR-0016, so those need a new example. About 40% of `vale-config` is install tables and settings lists the model can fetch from vale.sh. Keep the house-verified matrices (`E100`/`E201`, `Packages` below glob, frontmatter, ignore paths). `lint/docs/research/docs/vale/` overlaps the skill's own references by about two thirds. Effort S.
|
||||
|
||||
## 5. Prose and docs (9,600 lines, 109,000 words outside plugins)
|
||||
|
||||
26. [ ] **Move or delete `docs/research/` and `docs/notes/` (4,500 lines, 47% of prose words).** Six of eleven research files are linked only from each other; they are self-described session audit trails, agendas, and a "temporary build reference". `docs/notes/factory-research-gaps-conflicts.md` says "Status: Superseded"; `factory-integration-decisions.md` says "Complete" and its decisions already live in ADRs, yet `AGENTS.md` tells every session to read it. `archive/team-self-organisation-sprint-brief.md` (3,400 words) is unrelated to this repo. Archive or delete; drop the three `AGENTS.md` pointers. Moving `CONTROLS.md` to `docs/spec/` means updating its literal path in nine or more files including the deployed `governance.md`. Effort S.
|
||||
> **Decision (2026-09-12):** Keep. Same reasoning as finding 9 — these docs are intentional context for sourced work. Not proceeding.
|
||||
|
||||
27. **Four governance documents say one thing.** `core/instructions/governance.md` (949 words, always-on), `docs/ai-constitution.md` (2,906), `docs/wiki/HUMANS.md` (1,413), `CONTROLS.md` (1,224), with near-identical preambles and, in three of the four, a "what this file does not govern" block pointing at the others. The constitution repeats one of its own principle lead sentences. Keep `governance.md` as the operative file, trimmed to about 50 lines (drop the classification table that repeats the bullets above it, the footer, the non-governance block). Dedupe the constitution by about 20%. Effort M.
|
||||
|
||||
28. **ADRs: 2,740 lines, 72% in eight ADRs over 150 lines.** ADR-0020 is 513 lines with a 71-line measurement log as Context; ADR-0017 has 173 lines of amendments against 45 of decision. ADR-0001 is superseded and ADR-0006 moot, both keeping full text below the banner. ADR-0002 is three lines. Truncate superseded ones to the banner, fold amendments into the decision, cap Context at 20 lines, add a 25-line `docs/adr/README.md` index with status. The rules already live in `gates.md`; the ADRs need only decision and consequences. Effort M.
|
||||
|
||||
29. [x] ~~**The same facts are stated in full three or four times.** "Edit `.apm/`, never the mirror": README (2 paragraphs), AGENTS.md (2 paragraphs), architecture.md (2 paragraphs plus the lost-README anecdote), ADR-0017. The apm.lock / SessionStart story: README (11 lines), AGENTS.md, ADR-0018, ADR-0019, gates.md. The offline `SKIP=` command and the three-stage install each appear three times. Rule: README has the how-to, AGENTS.md has one-line rules with links, architecture.md has mechanics. Effort S.~~
|
||||
> **Corrected then partially done (2026-09-14):** independent re-verification found the "edit `.apm/`, never the mirror" and apm.lock/SessionStart clusters confirmed but the third overstated — no file documents an offline `SKIP=` command (the one `SKIP=`-adjacent mention in `gates.md` explicitly says a *different* opt-out "is not `SKIP=`"), and "three-stage install" appears twice, not three times, with no restatement worth trimming. Trimmed the two confirmed clusters: README's "Editing plugin content" and AGENTS.md's "Edit `.apm/`, never the flat mirror" sections cut to the how-to/one-line-plus-link split the finding itself proposed, full mechanics (the `rm -rf` behavior and the `plugins/kyberforge/hooks/README.md` anecdote) staying solely in `docs/spec/architecture.md`. README's "Keeping the install current" and AGENTS.md's apm.lock bullet trimmed to drop the restated `apm outdated`/`apm update --yes` timing narrative, pointing to ADR-0019 as the canonical mechanism instead. No test greps the trimmed wording (checked).
|
||||
|
||||
30. [x] ~~**`LESSONS.md`: 41 entries, 2 graduated, about 12 stale.** Twelve entries from 2026-05-17 describe a write-skill / write-eval workflow whose skills no longer exist. One entry is open work labelled "Status: neither part landed". The longest eight are 200 to 550-word incident reports. Delete the stale entries, move open work to an issue, cap entries at about 60 words, target 100 lines. Effort S.~~
|
||||
> **Done (2026-09-12):** see commit `629320b` on `docs/simplification-audit`. 255→131 lines, 41→30 entries. Kept 3 of the same-dated entries (RLHF defaults, secrets-rule gap, HITL gap) — judged unrelated to the defunct write-skill/write-eval workflow and still applicable, so 10 deleted rather than 12. The "neither part landed" open-work entry (CONTEXT.md not `@import`ed at session start) was removed rather than filed as an issue — full text preserved in this session's transcript if wanted later.
|
||||
|
||||
31. [x] ~~**`CONTEXT.md`: 28 terms, most used only by gates.md, scripts, or tests rather than by skills;** two (Preload tax, Skill context contract) are never used outside `CONTEXT.md` and ADR-0020. The preload-tax entry quotes two dated numbers then says not to quote them. The example dialogue and flagged-ambiguities sections are grill residue. Cut to about 20 one-line terms. Effort S.~~
|
||||
> **Corrected then done (2026-09-13):** see commits `124ce6e` and follow-up on `docs/simplification-audit`. Independent re-verification found "most used only by gates.md/scripts/tests" overstated: 13 of 28 terms are actually referenced from model-facing `references/*.md` files skills load in normal use (Routing target, Hand-invoked skill, Dispatch body, Near-miss, Thin adapter, Provenance chain, Output profile, apm package, Plugin marketplace, HITL, Skill composition, Delegation discipline, holocron) and were kept untouched. Only the 9 terms confirmed as true orphans were removed after a fresh independent grep: Content mirror, apm-consumed install, Vale audit prefilter, Vacuous green, Management Application, Sycophancy, HOTL, Preload tax, Skill context contract — 28 → 19 terms. The preload-tax self-contradiction (quotes 23,427/10,478-char figures then says not to quote either) was confirmed verbatim and resolved by the entry's own deletion. The "example dialogue" and "flagged ambiguities" sections were found to be mandated by `grill-with-docs/references/context-format.md`'s template spec, not grill residue — left untouched, except one dangling bolded cross-reference to the now-deleted "Preload tax" term in a Flagged-ambiguities line, which was unbolded/de-referenced in place (the ambiguity resolution itself still holds without a defined glossary entry to point at).
|
||||
|
||||
32. **Structure is described three ways** (README layout table, architecture.md plugin table, AGENTS.md structure bullets), and `VISION.md` carries a 35-line stack spec for a product that lives in another repo. One layout table in README; architecture.md keeps mechanics only; VISION drops the stack detail. Effort S.
|
||||
|
||||
## 6. Distribution, versioning, and session startup
|
||||
|
||||
Not covered by the area audits above; found on a final sweep of the root config and install pipeline. The install pipeline itself (`scripts/install.sh` 55 lines, `deploy-manifest.sh` 24, statusline 109) is fine and needs nothing.
|
||||
|
||||
33. **Every plugin version lives in four places (five for kyberforge), plus one per skill.** `plugins/<name>/apm.yml`, two generated `plugin.json` files, the root `apm.yml` packages list, the `executables.allow` key (`kyberforge#1.6.2`), and a `metadata.version` in all 39 SKILL.md files (ADR-0022) that nothing consumes and that drifts freely (gitea skills sit at five different values). Repo tags (`v2.0.1`) follow a third scheme that the declared `tagPattern: v{version}` can never match under `per_package` versioning. ADR-0006, ADR-0022, `check-executables-allow-sync`, `skill-frontmatter`, and `apm pack --check-versions` all exist to police this. Proposal: one version per plugin in its `apm.yml`; drop `metadata.version` and ADR-0022; let `apm pack` derive the rest. Effort M.
|
||||
> **Partially advanced (2026-09-14):** see commit `718c79a` on `docs/simplification-audit`. Two of the four locations per plugin are gone: the twelve generated `plugin.json` manifests (`plugins/*/.claude-plugin/` and `plugins/*/.github/plugin/`) were deleted with the mirror. ADR-0006 needed no action — it was already moot and governed only those two now-deleted manifests, so no version bumps were required by the change. **Not closed.** Still outstanding: `plugins/<name>/apm.yml`, the root `apm.yml` packages list, the `executables.allow` pin, and `metadata.version` in all 39 SKILL.md files (still unconsumed, still drifting), plus ADR-0022 and the `v{version}` `tagPattern` mismatch.
|
||||
|
||||
34. **The SessionStart hook auto-updates the install on every startup.** `check-apm-current.sh` runs `apm outdated` (network, 60 s timeout) and then `apm update --yes` (300 s timeout) at every session start, rewriting `apm.lock.yaml`. That is why the lock file is dirty at the start of this session and why `AGENTS.md` has to explain "commit or discard it deliberately". It is a 60-line script with a 368-line test, an ADR (0019), the `executables.allow` pin, and a sync hook behind it. For a repo that is its own source, the update belongs in `install.sh` or a manual `apm update`, not in session startup. Effort S to remove; the design question is whether auto-update at startup is wanted at all.
|
||||
|
||||
35. [x] ~~**Outputs and packages for consumers that do not exist.** The `codex` output profile generates `.agents/plugins/marketplace.json` (95 lines) although Codex is not a supported consumer. The `mattpocock-skills` remote package entry is the only reason `apm-marketplace-check` needs the network, and its pin is advanced by hand (ADR-0015). The `.github/plugin/marketplace.json` mirror is a legacy path (finding 2). Removing all three leaves one generated marketplace manifest (the per-plugin `plugin.json` pairs remain) and no network-dependent hook. Effort S.~~
|
||||
> **Done (2026-09-13):** see commit `568ca74` on `docs/simplification-audit`. Removed the `codex` output profile from root `apm.yml` and its compiled `.agents/plugins/marketplace.json` (95 lines), and the `mattpocock-skills` remote package entry — the only remote marketplace entry, so `apm-marketplace-check` and `apm-pack-check-clean` no longer need network access at all. Updated `README.md`, `AGENTS.md`, `docs/spec/gates.md`, and `docs/spec/architecture.md` accordingly; added one-line superseded/updated notes to ADR-0015 and ADR-0021. Left `.github/plugin/marketplace.json` untouched — that's the Copilot legacy-path question in finding 2/§8, out of scope here; only re-ran the sync script to keep it consistent. `apm.lock.yaml` unaffected (`marketplace.packages[]` isn't part of the lockfile). Verified via `apm install`, `apm pack --marketplace=claude --check-versions`, and all four affected pre-push hooks.
|
||||
|
||||
36. **The release-tag mechanism guards an external contract with no known consumer.** `.pre-commit-hooks.yaml` exports three hooks for other repos to pin by `rev: <tag>`. `check-release-needed` (242 lines + 442 test), `test-vale-hooks-consumer` (270 lines), ADR-0014, and three tags exist to serve that. If no other repo pins these hooks today, the whole mechanism can be deferred until one does. Effort S.
|
||||
|
||||
37. [x] ~~**Two `.mcp.json` files declare an Obsidian vault server over `docs/`** (root and `plugins/bin/`; the other five plugin `.mcp.json` files are empty stubs), while `AGENTS.md` forbids using an external memory system for this repo. If the Obsidian tools are unused, drop both and the `reinject_mcp_servers` explanation in the bin README; the bin `plugin.json` pair regenerates. Effort S.~~
|
||||
> **Not proceeding (2026-09-13):** premise doesn't hold. The server exposes the repo's own git-tracked `docs/` folder — not an external/off-repo store — so it isn't the "external memory system" AGENTS.md's rule targets. It was deliberately added and versioned (3 commits), is documented as current intended behavior in both READMEs, and ADR-0018 uses it as its only concrete worked example of apm's MCP-dependency propagation mechanism actually working. No skill invokes the Obsidian tools as a workflow step, but that alone doesn't make the config dead. No changes made; recommend a human confirm whether the vault tooling is still wanted before removing it.
|
||||
> **Confirmed and done (2026-09-14):** the human confirmed the vault tooling is not wanted — remove it entirely. All seven `.mcp.json` files deleted (the six plugin-root files and the repo-root one), and the repo-root path added to `.gitignore` so a local `apm` run cannot recreate it as tracked content. The bin README's `reinject_mcp_servers` explanation goes with it; the `plugin.json` pair the finding expected to regenerate no longer exists (deleted in `718c79a`, finding 7).
|
||||
>
|
||||
> What made this urgent is the substantive discovery, not the tidying: **deleting the per-plugin `plugin.json` manifests in `718c79a` had already broken MCP propagation silently.** `apm_cli/deps/plugin_parser.py` maps a plugin-root `.mcp.json` → `.apm/.mcp.json`, and that code path runs only for *marketplace* plugins — with no manifest, apm never reads the file. `plugins/bin/apm.yml` declares `dependencies.mcp: []`, so the supported mechanism was never used either. Proved on ref-pinned consumer clones: at the parent commit a consumer gets an `obsidian` server, at HEAD it gets none, and on upgrade apm prints `Removed stale MCP server 'obsidian' from .mcp.json` — which would in time have stripped the server from this repo's own tracked `.mcp.json` once the lock re-resolved. Deleting the files makes the intent match the behaviour instead of leaving a config that silently does nothing.
|
||||
|
||||
38. [x] ~~**`pc-author` / `pc-run` (689 lines) carry generic pre-commit documentation.** `hooks-by-language.md` (128 lines) and `failure-patterns.md` (133) restate pre-commit.com. Keep the skills, trim to the house-specific rules. Effort S.~~
|
||||
> **Corrected then done (2026-09-13):** see commit `a622200` on `docs/simplification-audit`. Independent re-verification found the 689-line figure overstated (actual combined size 598 lines) and the realistic cut smaller than a rewrite (~60-85 lines, concentrated in the two named reference files, not the SKILL.md files or the four short flow files, which are house-specific gates rather than restatement). Landed within that range: `hooks-by-language.md` 128 → 92 lines (collapsed six per-language tables repeating the same repo/rev/rationale into one shared-repo table plus a small other-repos table); `failure-patterns.md` 133 → 109 lines (removed generic SSH/proxy and shellcheck SC-code restatement, compressed generic schema-error bullets). Kept verbatim: both "Unverified — not in research corpus" flags, the rev-freshness caveat, the `rtk git add -u`/`rtk git commit` fix (ADR-0023), and the `pre-commit install -f` warning. Combined cut: 60 lines. Flat mirror regenerated and verified byte-identical.
|
||||
|
||||
## 7. Suggested order
|
||||
|
||||
1. Quick wins, all S, no design decisions needed: findings 9, 10, 26, 30, 31, 29, 12, 13, 1, 6, 4, 35, 37, 38, and the mirror-sync and executables-allow halves of 2. Removes roughly 25,000 to 30,000 lines and 6 hooks.
|
||||
2. Structural changes that need a short discussion: 14, 15, 19, 20, 23, 25, 17, 3, 5, 7, 33, 34, 36.
|
||||
3. The real complexity: 16 (validators), 11 (provenance), 24 (core), 8 and 28 (gates.md and ADRs).
|
||||
|
||||
Findings 9, 10, 11, and 12 are coupled through the provenance validator and the audit criteria; land them together or the audit gates start reporting the removals.
|
||||
|
||||
## 8. Questions to settle before starting
|
||||
|
||||
- [x] ~~**Native Claude Code marketplace install vs apm-only.** The flat mirror, `check-plugin-content-sync`, and ADR-0017 exist only for native `claude plugin install`. If apm install is the only supported path, the mirror and its 2,100 lines of tooling go away. Which install paths must work for consumers?~~
|
||||
> **Answered (2026-09-14):** apm-only. See ADR-0024 (`docs/adr/0024-apm-is-the-only-supported-install-path.md`) and commit `718c79a` on `docs/simplification-audit`. Native `claude plugin install` support is dropped; the flat mirror, the twelve per-plugin manifests, `sync-plugin-content.sh`, its test suite, `lib/marketplace-plugins.sh`, and the `check-plugin-content-sync` and `validate-plugins` hooks are all deleted (245 files changed, −22,602 lines). ADR-0017 carries a superseded banner. Kept deliberately: the root `marketplace:` block and the root `.claude-plugin/marketplace.json`, which apm's own consumers read. (`marketplace.json` is the only tracked file under `.claude-plugin/`; the root `plugin.json` beside it is untracked local `apm pack` output, ignored at `.gitignore:59`.) This answer is what voided finding 7's recommendation and closed §3's `check-plugin-content-sync` target.
|
||||
- **Copilot CLI legacy path.** Is `.github/plugin/marketplace.json` still read by any Copilot version you target? If not, finding 2c is a pure delete.
|
||||
- **Provenance chain.** Is "which upstream informed this file" a requirement you still want, or was it a governance experiment? Finding 11 hinges on this.
|
||||
- **ADR-0012 (three core skills) and the one-script-per-skill install constraint.** The merges in 14, 15, and 24 need the first revisited and are the only way around the second. Are you open to superseding ADR-0012?
|
||||
- **Granularity of git/gitea skills.** One `git` skill vs seven trades routing precision for size. Is one broad description acceptable?
|
||||
- **Auto-update at session start.** Do you want the install refreshed from the remote every time a session opens (finding 34), or is a manual `apm update` acceptable?
|
||||
- **External hook consumers.** Does any other repo pin this repo's `.pre-commit-hooks.yaml` by tag today? If not, finding 36 defers the release mechanism entirely.
|
||||
- [x] ~~**Obsidian MCP.** Are the Obsidian tools over `docs/` used by anyone? If not, finding 37 is a pure delete.~~
|
||||
> **Answered (2026-09-14):** not used — remove entirely. All seven `.mcp.json` files are deleted and the repo-root path is gitignored; see finding 37, which also records the functional regression this uncovered (since `718c79a` deleted the per-plugin manifests, apm no longer propagated the server to consumers at all).
|
||||
|
||||
## 9. Carried forward from the apm-only decision (2026-09-14)
|
||||
|
||||
Recorded here so they are not rediscovered as defects. All follow from commit `718c79a` / ADR-0024.
|
||||
|
||||
**Two accepted residuals.**
|
||||
|
||||
- **Native install still half-works, and cannot be prevented.** apm reuses Claude's catalogue format by design, so a Claude Code user can still register holocron natively and will install six plugins containing zero skills. Accepted, not overlooked: no schema change closes this, because the format that makes it possible is the format apm's own consumers need.
|
||||
- **Consumers now receive test fixtures.** apm installs from `.apm/`, which carries the `tests/` directories the mirror used to strip, so a consumer installing from this branch receives **10 `.bats` files across 6 skills**, plus those skills' 6 `tests/README.md` files — 16 files. (Repo-wide, 17 tracked paths contain `/tests/`: the 10 `.bats` and 7 `README.md`, one of which is a template asset under `skill-author/assets/templates/tests/` and is not a test fixture.) This is what consumers *receive*, not what this checkout shows: `.claude/skills/` here currently holds zero `.bats` files, because that deployed tree is stale and predates this branch. The mechanism was confirmed empirically on a ref-pinned consumer clone — the 16 files are absent at the parent commit and present at HEAD. Suppressing them means switching all six `apm.yml` files from `includes: auto` to explicit lists, where a wrong list silently drops content — worse failure mode than the noise. Deferred deliberately.
|
||||
|
||||
**Negative result — do not re-litigate.** Deleting native install does *not* relax the self-containment constraint. `plugins/kyberforge/.apm/skills/skill-author/references/deployment-modes.md`, sourced from the agentskills.io spec, states it independently for APM package mode: the spec defines no cross-skill sharing. So findings 14 and 15 still require *merging* skills; sharing one file between two skills remains impossible, and §8's "one-script-per-skill install constraint" bullet is unchanged by this decision.
|
||||
|
||||
**Accepted gap — symlinks under `.apm/`.** ADR-0017's `check_apm_symlinks()` was the only thing reporting that symlinks under `.apm/` do not survive to a consumer. It is gone, and no replacement guard is being added — the human decided to accept the gap.
|
||||
|
||||
The mechanism is not the bundle exporter, as ADR-0017 assumed; it is the **install** path, and it has since been verified. `apm_cli/security/gate.py`'s `ignore_non_content()` is a `shutil.copytree` ignore callback whose docstring says "Excludes symlinks (security)"; it is used at `apm_cli/integration/skill_integrator.py:424`, `:791` and `:1152`. Materialization into `apm_modules/` dereferences first, so symlinked content survives *there* and is dropped when skills are deployed out of it. ADR-0024 flagged the prediction as unverified; it holds, with that corrected attribution. No symlinks exist under any `.apm/` today, so nothing is broken now — but the next one added there will silently not reach consumers, and nothing will say so.
|
||||
3398
apm.lock.yaml
Normal file
3398
apm.lock.yaml
Normal file
File diff suppressed because it is too large
Load Diff
111
apm.yml
Normal file
111
apm.yml
Normal file
@@ -0,0 +1,111 @@
|
||||
name: holocron
|
||||
version: 0.4.6
|
||||
description: AI development skills for Claude Code and GitHub Copilot CLI — factory, design, implement, review, and cross-cutting workflows.
|
||||
license: MIT
|
||||
|
||||
# Consumer side: this repo installs its own published plugins from the holocron
|
||||
# remote, so the working copy runs the same released content every other
|
||||
# consumer gets. Addressed as git+path objects rather than <name>@holocron
|
||||
# marketplace aliases — an alias needs a `apm marketplace add` registration in
|
||||
# ~/.apm/marketplaces.json (user scope, outside this repo), the object form
|
||||
# needs nothing beyond this manifest.
|
||||
# Unpinned (default branch) on purpose: parity with the Claude Code plugin
|
||||
# install this replaced, which ran autoUpdate against main. Add `ref: <tag>`
|
||||
# per entry to pin.
|
||||
targets:
|
||||
- claude
|
||||
dependencies:
|
||||
apm:
|
||||
- git: git@git.dev.rkdr.net:Defame1297/holocron.git
|
||||
path: plugins/bin
|
||||
- git: git@git.dev.rkdr.net:Defame1297/holocron.git
|
||||
path: plugins/core
|
||||
- git: git@git.dev.rkdr.net:Defame1297/holocron.git
|
||||
path: plugins/git
|
||||
- git: git@git.dev.rkdr.net:Defame1297/holocron.git
|
||||
path: plugins/gitea
|
||||
- git: git@git.dev.rkdr.net:Defame1297/holocron.git
|
||||
path: plugins/kyberforge
|
||||
- git: git@git.dev.rkdr.net:Defame1297/holocron.git
|
||||
path: plugins/lint
|
||||
mcp: []
|
||||
|
||||
# Turns apm's executable-trust gate ON. Without this block the gate is disabled
|
||||
# and every hook, bin and MCP primitive a dependency ships deploys silently —
|
||||
# verified: `apm approve --list` reports "Executable-trust gate disabled -- all
|
||||
# executables deploy" until an `executables:` block exists.
|
||||
#
|
||||
# kyberforge ships the SessionStart hook that keeps this install level with the
|
||||
# remote (ADR-0019). The key is version-pinned by apm's own design, so a
|
||||
# kyberforge version bump makes this entry stop matching and the hook stops
|
||||
# deploying until the version here is bumped too. If skills silently go stale
|
||||
# after a kyberforge release, check this first.
|
||||
executables:
|
||||
allow:
|
||||
kyberforge#1.6.2:
|
||||
hooks: true
|
||||
bin: true
|
||||
|
||||
marketplace:
|
||||
# apm's Claude marketplace mapper only emits description:/version: into the
|
||||
# compiled marketplace.json when set explicitly here (an override) — the
|
||||
# top-level apm.yml description:/version: above are NOT inherited into the
|
||||
# compiled output despite being used elsewhere (e.g. by `apm audit`).
|
||||
description: AI development skills for Claude Code and GitHub Copilot CLI — factory, design, implement, review, and cross-cutting workflows.
|
||||
version: 0.4.6
|
||||
owner:
|
||||
name: Defame1297
|
||||
email: defame1297@rkdr.net
|
||||
url: https://git.dev.rkdr.net/Defame1297/
|
||||
|
||||
# Default tag pattern used to resolve version ranges for each package.
|
||||
build:
|
||||
tagPattern: "v{version}"
|
||||
|
||||
# Output targets (map form). Each output writes to its profile default
|
||||
# path; add 'path:' under a key to override.
|
||||
outputs:
|
||||
claude: {}
|
||||
|
||||
# CI tip: build a machine-readable manifest:
|
||||
# apm pack --marketplace=claude --json | jq -r '.marketplace.outputs[].path'
|
||||
|
||||
versioning:
|
||||
strategy: per_package
|
||||
|
||||
packages:
|
||||
- name: kyberforge
|
||||
description: Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace.
|
||||
source: ./plugins/kyberforge
|
||||
version: 1.6.2
|
||||
category: Developer Tools
|
||||
|
||||
- name: bin
|
||||
description: Skills for everyday AI-assisted development work that is not tied to a single tool, forge or language, and has not yet been split into a focused plugin.
|
||||
source: ./plugins/bin
|
||||
version: 1.1.7
|
||||
category: Utilities
|
||||
|
||||
- name: git
|
||||
description: Skills and agents for working with a local Git clone over the git wire protocol, and for authoring and running the pre-commit hooks that guard it.
|
||||
source: ./plugins/git
|
||||
version: 1.3.7
|
||||
category: Version Control
|
||||
|
||||
- name: gitea
|
||||
description: Skills and agents for working with a Gitea forge through its HTTP API — the forge's own objects, as distinct from the local git clone.
|
||||
source: ./plugins/gitea
|
||||
version: 1.3.8
|
||||
category: Version Control
|
||||
|
||||
- name: core
|
||||
description: Skills for authoring and auditing a repo's AGENTS.md and the provider adapter files that defer to it.
|
||||
source: ./plugins/core
|
||||
version: 1.1.2
|
||||
category: Productivity
|
||||
|
||||
- name: lint
|
||||
description: Skills and agents for configuring and running linters.
|
||||
source: ./plugins/lint
|
||||
version: 1.1.7
|
||||
category: Developer Tools
|
||||
@@ -1,5 +1,17 @@
|
||||
# Skills are distributed via plugins, not monolithic repo deployment
|
||||
|
||||
**Superseded by:** ADR-0015 (Microsoft APM replaces the hand-authored plugin/marketplace model
|
||||
as this repo's authoring source of truth) and, for plugin-scope agent files specifically,
|
||||
ADR-0016 (plugin-scope `.apm/agents/*.agent.md` drops provider-specific fields). Since issue
|
||||
#90's conversion executed, plugin content is authored under `plugins/<name>/apm.yml` +
|
||||
`.apm/{skills,agents,hooks}/` — not the flat `skills/`/`agents/` layout this ADR describes —
|
||||
and `.claude-plugin/plugin.json`/`.github/plugin/plugin.json` were compiled output of `apm pack`,
|
||||
not hand-authored — and as of ADR-0024 (2026-09-14) both are deleted, along with native
|
||||
`claude plugin install` support; `apm install` is the only route. This ADR's content is kept
|
||||
below as the historical record of the pre-APM decision; it is no longer the current model.
|
||||
|
||||
---
|
||||
|
||||
Skills (slash commands) are authored and distributed as part of **plugins** — each plugin contains its own `skills/` directory alongside agents and other artifacts. Plugins are installed via `claude plugin install <name>@holocron` rather than deployed from the repo's local tree. This decision decouples skill authoring cadence from core provider deployments and allows independent versioning per plugin.
|
||||
|
||||
## Context
|
||||
|
||||
@@ -44,3 +44,8 @@ separate single-provider skill, adding complexity with no benefit.
|
||||
file now lives at `<plugin-root>/sources.md`, outside the `agents/` directory, because
|
||||
`claude plugin validate --strict` auto-discovers every `.md` under `agents/` as an agent
|
||||
requiring frontmatter. See ADR-0010 for the empirical finding and rationale.
|
||||
|
||||
**Update (ADR-0016):** the plugin-scope clause above is superseded. Plugin scope is no longer
|
||||
detected via `plugin.json`, and no longer produces a Claude+Copilot file pair — a directory
|
||||
containing `apm.yml` now gets a single vendor-neutral `.apm/agents/<name>.agent.md` file with
|
||||
no provider-specific fields. Project scope and user scope are unaffected. See ADR-0016.
|
||||
|
||||
@@ -1,5 +1,23 @@
|
||||
# version field is present in both plugin manifests
|
||||
|
||||
**Moot as of ADR-0015.** This ADR addressed drift risk between two independently
|
||||
*hand-maintained* manifests. Since issue #90's conversion executed, `.claude-plugin/plugin.json`
|
||||
and `.github/plugin/plugin.json` are both **compiled output** of `apm pack`, generated in the
|
||||
same pass from a single `apm.yml` per plugin — there is no longer a second hand-authored file
|
||||
that could drift out of parity. The invariant this ADR required (`version` present and
|
||||
identical in both manifests) still holds in the compiled output, but structurally, not because
|
||||
a skill enforces it: both files are derived from the same `apm.yml` `version:` field, so
|
||||
divergence is no longer possible by construction. `plugin-author`, the skill that enforced this
|
||||
invariant, is deleted per ADR-0015 rather than adapted. Kept below as the historical record of
|
||||
the pre-APM decision.
|
||||
|
||||
**Fully void as of ADR-0024 (2026-09-14).** Both manifests are now deleted outright, so the two
|
||||
files this ADR was about no longer exist in any form, compiled or hand-authored. `apm.yml`'s
|
||||
`version:` is the only version field a plugin has. This ADR states no patch-bump rule and never
|
||||
did — ADR-0015 retired that rule explicitly; do not cite this ADR as the source of one.
|
||||
|
||||
---
|
||||
|
||||
Each plugin has two manifests: `plugin.json` (Copilot CLI) and `.claude-plugin/plugin.json` (Claude Code). Both tools support a `version` field. Prior to this decision, only the CC manifest carried `version`; the Copilot manifest omitted it.
|
||||
|
||||
We now require `version` in both manifests, always identical. A reader of `plugin.json` alone should be able to determine the plugin version without consulting the CC manifest. The `plugin-author` skill enforces this invariant on every create, update, and release operation.
|
||||
|
||||
@@ -2,6 +2,8 @@
|
||||
|
||||
**Supersedes:** ADR-0011 (provider-agnostic issue tracker with file-based default — archived during refactoring)
|
||||
|
||||
> **Note on the ADR-0011 number.** Every "ADR-0011" on this page means the *archived* provider-agnostic issue tracker ADR, which no longer exists in `docs/adr/` — it was removed when it was superseded, and the number 0011 was later reused for an unrelated decision, `docs/adr/0011-gitea-skill-deep-modules.md` (the gitea skill's split into deep modules). That file is not the ADR referenced below. The number is not renumbered here: these ADRs are a published record and renumbering would break every citation that already points at either one. The archived text is recoverable from git history.
|
||||
|
||||
ADR-0011 established a provider-agnostic model with `docs/issues/NNNN-<slug>.md` as the file-based default, switching to Gitea MCP at runtime when available. The interim model was justified because Gitea would not be configured until after Chunk 3, and the repo needed to work before then.
|
||||
|
||||
Gitea is now configured and in active use. The condition in ADR-0011 has been met. This ADR supersedes it.
|
||||
@@ -16,4 +18,4 @@ Three alternatives were rejected. Keeping the file-based fallback adds code comp
|
||||
|
||||
The file-based model also had a structural weakness: issues in `docs/issues/` were invisible from the Gitea UI, making it impossible to track work, assign milestones, or filter by label without opening the repo locally. Gitea provides all of that natively.
|
||||
|
||||
The "Provider-agnostic issue tracker" glossary entry in CONTEXT.md is updated in the same workstream to remove the file-based phase framing. The `providers/gitea/` adapter path described in ADR-0011 was never implemented — Gitea integration runs entirely via MCP, not a provider adapter.
|
||||
The "Provider-agnostic issue tracker" glossary entry in CONTEXT.md is updated in the same workstream to remove the file-based phase framing. (Amended 2026-08-17: the CONTEXT.md trim renamed that entry to **Issue**; it still records Gitea as this repo's canonical tracker and still tells skills to say "linked issue" generically.) The `providers/gitea/` adapter path described in ADR-0011 was never implemented — Gitea integration runs entirely via MCP, not a provider adapter.
|
||||
|
||||
@@ -14,3 +14,9 @@
|
||||
- Scope detection walks up from the input file: first directory containing `plugin.json` → plugin scope; first directory containing `.git` without `plugin.json` → project scope; path under `~` with neither → user scope.
|
||||
- At user scope the derivation crosses filesystem locations (`~/.claude/agents/` ↔ `~/.copilot/agents/`); the script must handle the home directory case explicitly.
|
||||
- The invocation signature is the public contract. Changing it is a breaking change to any caller — treat it as such.
|
||||
|
||||
**Update (ADR-0016):** the plugin-scope clause above is superseded. Plugin scope is no longer
|
||||
detected via `plugin.json`, and there is no counterpart to derive — a directory containing
|
||||
`apm.yml` produces a single `.apm/agents/<name>.agent.md` file, and `agent-audit` validates it
|
||||
directly with no pair-consistency check. Project scope and user scope keep the pair-derivation
|
||||
mechanism described above unchanged. See ADR-0016.
|
||||
|
||||
@@ -5,6 +5,21 @@ claim that "both files share a single `agents/sources.md` for provenance." The r
|
||||
ADR-0005 (dual-provider generation, scope detection, single-root script interface) is
|
||||
unaffected and remains in force.
|
||||
|
||||
**Path update per ADR-0016:** at plugin scope, agent files no longer live at
|
||||
`<plugin-root>/agents/<name>.md`. The authoring source is now
|
||||
`<plugin-root>/.apm/agents/<name>.agent.md` — a single vendor-neutral file (no dual Claude/
|
||||
Copilot pair) compiled to both targets via `apm pack`. See ADR-0016 for why (the field-dropping
|
||||
rationale, `tools:` incompatibility, the compiled-output mechanics) — not restated here. This
|
||||
ADR's own conclusion is unaffected by that move: the provenance file still belongs at
|
||||
`<plugin-root>/sources.md`, outside any directory `claude plugin validate --strict`
|
||||
auto-scans, and `.apm/agents/` is, if anything, further removed from plugin-root than the old
|
||||
flat `agents/` directory was, so the reasoning below still holds. References below to
|
||||
`<plugin-root>/agents/` describe the pre-APM layout in effect when this decision was made.
|
||||
**Scope boundary (per ADR-0016):** this path change is plugin scope only. Project scope
|
||||
(`.claude/agents/` + `.github/agents/`) and user scope (`~/.claude/agents/` +
|
||||
`~/.copilot/agents/`) are unaffected — they are not APM packages and keep the dual-file
|
||||
Claude+Copilot pair model this ADR originally described.
|
||||
|
||||
`claude plugin validate --strict` auto-discovers every `.md` file directly under a plugin's
|
||||
`agents/` directory and treats it as an agent definition requiring YAML frontmatter (`name`,
|
||||
`description`, etc.). A flat provenance file at `agents/sources.md` — no frontmatter, by
|
||||
|
||||
@@ -67,6 +67,14 @@ than being wired into the plugin manifest. This means the gitea plugin is not ye
|
||||
standalone via `claude plugin install gitea@holocron` without manual MCP setup. A follow-up Gitea
|
||||
issue tracks closing this gap.
|
||||
|
||||
**Correction (2026-09-14):** this gap is now closed by removal rather than by wiring. ADR-0024
|
||||
made apm the only supported install path, so `claude plugin install gitea@holocron` is no longer
|
||||
a route this repo supports, and the per-plugin manifests it needed are gone. Because apm reads a
|
||||
plugin-root `.mcp.json` only on the marketplace-plugin code path, that file became unreadable;
|
||||
every `plugins/*/.mcp.json` was deleted, including this one. A plugin that needs an MCP server
|
||||
declares it in `dependencies.mcp` in its `apm.yml` — the supported mechanism, which this repo has
|
||||
never used. The follow-up issue this paragraph anticipates is moot.
|
||||
|
||||
**Research backfill.** The existing research docs
|
||||
(`plugins/gitea/docs/research/docs/gitea/`) are 100% code-derived from gitea-mcp source with zero
|
||||
external/best-practice content (the original docs.gitea.com fetch timed out and was never
|
||||
|
||||
@@ -9,6 +9,12 @@ deferred PR #85 review item to broaden that coverage, retroactively captures #84
|
||||
(since it was never recorded as a decision in its own right), and layers the expansion on top
|
||||
without reversing or weakening the original four rules.
|
||||
|
||||
**2026-08-17 amendment.** The CONTEXT.md section named above no longer holds that documentation.
|
||||
CONTEXT.md was cut back to a glossary and the prefilter's mechanics — the two-copy style layout,
|
||||
`vale-wrap.sh`, the `--config` argv defect, the rule inventory, and the 0-files-means-NOT-RUN
|
||||
fallback — moved to `docs/spec/gates.md`. Read that file, not CONTEXT.md, for the harness itself;
|
||||
this ADR still owns the scope decision.
|
||||
|
||||
**File scope stays the same.** `SKILL.md` plus agent files (`**/agents/*.md`,
|
||||
`**/*.agent.md`) only — matching the existing prefilter's globs. Skill-level
|
||||
`README.md` files and `plugin.json` manifests are not added: README.md files are navigational, not
|
||||
@@ -16,8 +22,10 @@ spec-governed content, and `plugin.json` is JSON, not prose Vale can meaningfull
|
||||
|
||||
**Rule categories are prose-pattern-matchable only.** Structural, schema, and security concerns
|
||||
stay out of this Vale-based harness because this repo already has dedicated tools for them:
|
||||
`skill-frontmatter` (required frontmatter fields), `validate-plugins`/`validate-marketplace`
|
||||
`skill-frontmatter` (required frontmatter fields), `validate-marketplace`
|
||||
(`claude plugin validate --strict`, schema), and `gitleaks`/`detect-private-key` (secrets).
|
||||
(ADR-0024 removed the companion `validate-plugins` gate along with the per-plugin manifests it
|
||||
checked; the argument here is unaffected.)
|
||||
Duplicating those concerns as Vale rules would fight tools that already own them better.
|
||||
|
||||
**Governance docs are excluded as a rule source.** `docs/research/governance_principles/CONTROLS.md`
|
||||
|
||||
@@ -18,9 +18,9 @@ full LLM judgment every time outside this repo — the exact gap ADR-0013 named
|
||||
the plugin itself, following the no-cross-skill-path rule already established in
|
||||
`skill-author/references/deployment-modes.md` (a plugin's cache-install only copies each skill's
|
||||
own files; there is no plugin-level shared directory). `agent-audit` needs both `Kyberforge` and
|
||||
`KyberforgeCopilot` (it lints `.agent.md` files), so `plugins/kyberforge/skills/agent-audit/assets/vale/`
|
||||
`KyberforgeCopilot` (it lints `.agent.md` files), so `plugins/kyberforge/.apm/skills/agent-audit/assets/vale/`
|
||||
is the canonical, superset copy. `skill-audit` needs a second, smaller copy
|
||||
(`plugins/kyberforge/skills/skill-audit/assets/vale/`, `Kyberforge` only) since it cannot
|
||||
(`plugins/kyberforge/.apm/skills/skill-audit/assets/vale/`, `Kyberforge` only) since it cannot
|
||||
reference agent-audit's copy across the skill boundary. Both skills' Step 1 now resolve
|
||||
`scripts/vale-wrap.sh`/`assets/vale/.vale.ini` relative to their own directory, the same way
|
||||
`scripts/validate.sh <skill-dir>` already does — no new resolution mechanism, just applying the
|
||||
@@ -38,7 +38,7 @@ and gets all three, fully decoupled from Claude Code. CI is the identical `pre-c
|
||||
**This repo's own dev-time gate** consumes the same plugin-bundled copies instead of a third
|
||||
root-level copy — per explicit instruction, this repo should be set up like any other consumer
|
||||
would be, not dogfood a special root-only path. The existing `repo: local` hook is retargeted
|
||||
(not removed): `entry:` now points at `plugins/kyberforge/skills/{skill-audit,agent-audit}/scripts/vale-wrap.sh`.
|
||||
(not removed): `entry:` now points at `plugins/kyberforge/.apm/skills/{skill-audit,agent-audit}/scripts/vale-wrap.sh`.
|
||||
`repo: local` is kept rather than switching to a pinned self-reference
|
||||
(`repo: <own-url>, rev: <tag>`) — a pinned self-reference would lint working-tree edits against
|
||||
the *last tagged release*, not the change actually being made, which is wrong for the repo that
|
||||
@@ -58,15 +58,20 @@ single hook at agent-audit's copy silently scanned 0 SKILL.md files.)
|
||||
**The hook `entry:` is the wrapper alone; the wrapper self-locates its config.** pre-commit
|
||||
prefixes only `entry[0]` with the hook-repo clone path (`cmd = (prefix.path(cmd[0]), *cmd[1:])`);
|
||||
every later argument is handed to the process untouched and so resolves against the *consuming*
|
||||
repo's root. A `--config plugins/kyberforge/skills/…/assets/vale/.vale.ini` in
|
||||
repo's root. A `--config plugins/kyberforge/.apm/skills/…/assets/vale/.vale.ini` in
|
||||
`.pre-commit-hooks.yaml` therefore named a path no consumer has, and every external run died with
|
||||
`E100 [--config] Runtime error`. The external-consumer contract this ADR exists to establish
|
||||
cannot be expressed as a `--config` argument at all — the config path has to be derived inside
|
||||
the process, from the script's own location. `vale-wrap.sh` accordingly defaults to its sibling
|
||||
`assets/vale/.vale.ini`, resolved from `${BASH_SOURCE[0]}`, whenever no `--config` is supplied;
|
||||
an explicit `--config` from any other caller still wins and still resolves against the caller's
|
||||
cwd, so both audit skills' Step 1 (`--config assets/vale/.vale.ini`) is unaffected. Both
|
||||
manifests now carry the identical argument-free `entry:`. Keeping them identical is part of the
|
||||
cwd. Both audit skills' Step 1 passes no `--config` either, for the same reason and one more: a
|
||||
relative `--config assets/vale/.vale.ini` resolves against the cwd, not against the skill
|
||||
directory the wrapper path was resolved from, so it yields `E100 Runtime error … does not exist`
|
||||
and exit 2 — which both skills' fallback misreads as "vale unavailable" and silently downgrades
|
||||
to full LLM judgment, the exact failure the self-location exists to prevent. Both `SKILL.md` Step
|
||||
1 sections say so explicitly ("Pass no `--config`"), and both manifests now carry the identical
|
||||
argument-free `entry:`. Keeping them identical is part of the
|
||||
decision: the local `repo: local` hook resolved its `--config` correctly only because the
|
||||
consuming repo *was* this repo, and that one difference is why three review rounds exercised a
|
||||
code path no external consumer ever takes.
|
||||
@@ -105,11 +110,14 @@ doesn't wonder if it was overlooked.
|
||||
## Consequences
|
||||
|
||||
- Root `.vale.ini`, `styles/`, `scripts/vale-wrap.sh` are deleted. Two copies remain:
|
||||
`plugins/kyberforge/skills/agent-audit/assets/vale/` (canonical, superset) and
|
||||
`plugins/kyberforge/skills/skill-audit/assets/vale/` (subset, `Kyberforge` only).
|
||||
`plugins/kyberforge/.apm/skills/agent-audit/assets/vale/` (canonical, superset) and
|
||||
`plugins/kyberforge/.apm/skills/skill-audit/assets/vale/` (subset, `Kyberforge` only).
|
||||
- `plugins/kyberforge`'s `plugin.json` and `.claude-plugin/plugin.json` both patch-bump for every
|
||||
shipped content change (per ADR-0006's version-parity invariant): `1.2.5` for the relocation
|
||||
itself, `1.2.6` for the self-locating `vale-wrap.sh` that followed.
|
||||
**Amended 2026-09-14 (ADR-0024):** a record of what was done then, not current practice. Both
|
||||
manifests are deleted and `apm.yml`'s `version:` is a plugin's only version field; ADR-0015
|
||||
retired the parity/patch-bump rule this bullet invokes.
|
||||
- **`.pre-commit-hooks.yaml` entries are a bare script path and nothing else — a constraint, not a
|
||||
house style, and it binds every future hook here, not just the Vale two.** Since pre-commit
|
||||
rewrites only `entry[0]` into the hook-repo clone, no argument token in any entry can reference
|
||||
|
||||
186
docs/adr/0015-apm-replaces-plugin-marketplace-authoring.md
Normal file
186
docs/adr/0015-apm-replaces-plugin-marketplace-authoring.md
Normal file
@@ -0,0 +1,186 @@
|
||||
# Microsoft APM replaces the hand-authored plugin/marketplace model as this repo's authoring source of truth
|
||||
|
||||
**Status: executed (2026-08-12, issue #90).** All six plugins now carry `apm.yml` + `.apm/` as
|
||||
their authoring source; `.claude-plugin/marketplace.json` and every plugin's `plugin.json` are
|
||||
`apm pack`-compiled output. **Supersedes ADR-0001** ("Skills are distributed via plugins... each
|
||||
plugin contains its own `skills/` directory") — in effect.
|
||||
|
||||
This repo replaces its hand-maintained Claude Code plugin/marketplace authoring model
|
||||
(`.claude-plugin/marketplace.json` + per-plugin `plugin.json`) with Microsoft APM (`apm.yml` +
|
||||
`.apm/`) as the authoring source of truth — an outright replacement of the authoring layer, not an
|
||||
additive overlay. This ADR records the decision from a `grill-with-docs` session on issue #88.
|
||||
|
||||
## Context
|
||||
|
||||
Every plugin under `plugins/<name>/` currently ships two hand-maintained manifests
|
||||
(`.claude-plugin/plugin.json` for Claude Code, root `plugin.json` for Copilot CLI) plus a
|
||||
hand-maintained root `.claude-plugin/marketplace.json` listing all plugins. Adding a provider means
|
||||
hand-authoring a third manifest shape; keeping the two existing ones in parity is itself a tracked
|
||||
concern (ADR-0006).
|
||||
|
||||
Research on Microsoft APM (`plugins/kyberforge/docs/research/docs/microsoft-apm/`) found that its
|
||||
documented "monorepo-hybrid" repo shape maps directly onto this repo's existing `plugins/<name>/`
|
||||
layout: each plugin becomes its own `apm.yml` + `.apm/{skills,agents,hooks,prompts,instructions}/`
|
||||
package, listed from a root `apm.yml`'s `marketplace:` block. `apm compile`/`apm pack` generate
|
||||
per-target output — including a `.claude-plugin/marketplace.json` — from that vendor-neutral
|
||||
`.apm/` tree, so provider manifests become compiled artifacts instead of hand-authored files, and
|
||||
new providers (Copilot, Gemini, Codex — all supported by `apm runtime setup`) no longer require a
|
||||
new hand-maintained manifest format.
|
||||
|
||||
## Decision
|
||||
|
||||
- **The `plugins/<name>/` monorepo-hybrid directory layout survives.** `.claude-plugin/marketplace.json`
|
||||
and per-provider `plugin.json` files become **compiled output** via `apm compile`/`apm pack`,
|
||||
generated from `apm.yml` + `.apm/` per plugin, extensible to other `apm runtime`-supported
|
||||
providers without hand-maintaining a separate manifest per provider.
|
||||
- **This supersedes ADR-0001** ("Skills are distributed via plugins... each plugin
|
||||
contains its own `skills/` directory"). Executed in issue #90: skills and agents physically moved
|
||||
to `plugins/<name>/.apm/skills/` and `plugins/<name>/.apm/agents/*.agent.md`.
|
||||
- New operational tooling — `apm-install` (skill), `apm-workflow` (skill), `apm-orchestrate`
|
||||
(agent) — lands in `kyberforge`, tracked in issue #88
|
||||
(https://git.dev.rkdr.net/Defame1297/holocron/issues/88).
|
||||
- Adapting `skill-author`/`agent-author`'s routing to author `.apm/`-native content (retargeting to
|
||||
`.apm/skills/`, `.apm/agents/` paths — the content these two skills author is still meaningful
|
||||
post-conversion) is deferred to issue #89
|
||||
(https://git.dev.rkdr.net/Defame1297/holocron/issues/89). `forge` is out of scope for #89 — it
|
||||
stays untouched by this whole conversion effort and keeps routing to whatever the live author
|
||||
skills are at the time.
|
||||
- **`plugin-author`/`marketplace-author` are not adapted — they are superseded and deleted.**
|
||||
Unlike `skill-author`/`agent-author`, nothing in these two skills carries forward as authoring
|
||||
routing: `apm compile`/`apm pack` will generate `.claude-plugin/marketplace.json` and
|
||||
per-provider `plugin.json` directly from `apm.yml` + `.apm/`, so `apm-install`/`apm-workflow`/
|
||||
`apm-orchestrate` (issue #88, already landed on this branch) fully replace what these two skills
|
||||
did. `plugin-author`/`marketplace-author` were deleted in issue #90's execution.
|
||||
- Translating the existing plugins into `apm.yml` + `.apm/` and running the real conversion was
|
||||
executed under issue #90 (https://git.dev.rkdr.net/Defame1297/holocron/issues/90), which tracks
|
||||
that work through to merge.
|
||||
- `CONTEXT.md`'s "Plugin"/"Plugin marketplace" glossary entries were rewritten in issue #90 to
|
||||
describe the compiled-output model directly, rather than carrying a forward-pointer to this ADR.
|
||||
Superseded 2026-08-17: CONTEXT.md was cut back to one-line definitions, and the compiled-output
|
||||
model is now described in `docs/spec/architecture.md`. The same trim deleted the "lint plugin"
|
||||
entry cited under Considered options below; that pointer now reads `docs/spec/architecture.md`'s
|
||||
plugin scope table, which carries the repo-agnostic-versus-marketplace-specific argument.
|
||||
|
||||
## Considered options
|
||||
|
||||
**Additive/compile-layer only, no `apm.yml` (rejected).** Keep `plugin.json`/`marketplace.json`
|
||||
hand-authored and bolt APM on top as an optional extra. Rejected: doesn't achieve the multi-provider
|
||||
compile-reuse goal APM's package model provides, and leaves the existing dual-manifest hand
|
||||
maintenance in place unchanged.
|
||||
|
||||
**New standalone `plugins/apm/` plugin (rejected).** `plugins/lint/` was split out of `kyberforge`
|
||||
specifically because Vale tooling is generic and repo-agnostic, not holocron-marketplace-specific
|
||||
(see `docs/spec/architecture.md`'s plugin scope table) — the same argument applies to a generic `apm` CLI
|
||||
wrapper. The shipped `apm-install`/`apm-workflow` skills are, in fact, generic, repo-agnostic APM
|
||||
CLI documentation with no holocron-specific content, so a standalone `plugins/apm/` would have
|
||||
been a defensible split on artifact content alone. Rejected anyway, in favor of `kyberforge`,
|
||||
because holocron is currently the only repo that needs this tooling — standing up a separate
|
||||
plugin for a single consumer isn't worth it yet. Accepted as an explicit tradeoff (same pattern
|
||||
as ADR-0011's `gitea-workflow` naming tradeoff) — worth revisiting if this tooling is ever reused
|
||||
outside holocron's own conversion.
|
||||
|
||||
## Content migration out of `plugin-author`/`marketplace-author`
|
||||
|
||||
A content audit of `plugin-author`/`marketplace-author` (same `grill-with-docs` session as this
|
||||
correction) sorted what they document into three buckets:
|
||||
|
||||
- **Claude Code platform constraints — carried forward.** Facts that stay true regardless of
|
||||
authoring model (reserved plugin-name prefixes; the `agents/`-directory stray-`.md`-file
|
||||
validator gotcha, ADR-0010; `claude plugin validate` as a required terminal check) have been
|
||||
added into `apm-workflow`'s reference docs, since compiled output still has to satisfy these
|
||||
constraints post-conversion.
|
||||
- **Dual-manifest artifacts — obsolete, not carried forward.** Conventions that existed only
|
||||
because of hand-authored dual manifests (ADR-0006's version-parity/patch-bump rule, the
|
||||
CC-vs-Copilot field-placement split, dual-file mirroring) are obsolete under `apm.yml`'s
|
||||
single-manifest model and were deliberately dropped.
|
||||
- **Holocron policy choice — resolved in #90.** `marketplace-author`'s catalog-version convention
|
||||
(minor bump for package add/remove, patch bump for field-only updates) isn't an APM mechanic —
|
||||
`apm` doesn't enforce it, and has no native version-bump automation at all — so rather than
|
||||
building a new script, the convention is now documented as guidance inside `apm-workflow`'s
|
||||
reference docs (`references/marketplace.md` for the root catalog version rule,
|
||||
`references/configure.md` for the per-package version-bump-on-content-edit rule), applied
|
||||
manually by whoever edits `apm.yml`.
|
||||
|
||||
## Consequences
|
||||
|
||||
- ADR-0001 is superseded (issue #90).
|
||||
- ADR-0006 (plugin-version-parity) is moot (issue #90): `plugin.json`/`marketplace.json` are now
|
||||
compiled output of a single `apm.yml`, so there's no second hand-authored file left to keep in
|
||||
parity, and `plugin-author` — the skill that enforced ADR-0006 — was deleted rather than adapted
|
||||
(see "Content migration" above).
|
||||
- ADR-0010 (agent sources relocated outside agents dir) was updated (issue #90) for agents now
|
||||
living at `plugins/<name>/.apm/agents/*.agent.md` — the directory path changed; the pre-existing
|
||||
`.agent.md` extension convention (ADR-0005/ADR-0010) and project/user scope are unaffected, per
|
||||
ADR-0016.
|
||||
- ADR-0014 (Vale prefilter ships from the plugin) had its hardcoded `plugins/<name>/skills/...`
|
||||
paths (the Vale prefilter is skill-scoped only; ADR-0014 never referenced a
|
||||
`plugins/<name>/agents/...` path) updated for the `.apm/` nesting as part of issue #90's
|
||||
execution.
|
||||
- `kyberforge` gained three new artifacts (issue #88) before any conversion of existing content
|
||||
happened, then lost two (`plugin-author`/`marketplace-author`, deleted once issue #90 verified
|
||||
parity) — net version bump 1.3.1 → 1.4.0. The root marketplace catalog bumped 0.3.1 → 0.3.2 to
|
||||
match.
|
||||
- ADR-0016 (a narrower decision discovered while designing issue #89) turned out to gate how
|
||||
issue #90 had to re-author plugin-scope agents: `.apm/agents/*.agent.md` compiles verbatim to
|
||||
both Claude and Copilot, so those files carry only the fields in the `apm-agent-allowlist` section
|
||||
of `plugins/kyberforge/.apm/skills/agent-audit/references/field-inventory.md` (as amended
|
||||
2026-08-14: `name`/`description`/`model`/`source_keys`/`disallowedTools`) — existing dual-file
|
||||
`<name>.md`+`<name>.agent.md` pairs could not be raw-moved, only re-authored.
|
||||
- Two follow-up issues tracked the remaining work: #89 (`skill-author`/`agent-author` routing
|
||||
adaptation — closed, merged in #93) and #90 (the actual repo conversion, which also deleted
|
||||
`plugin-author`/`marketplace-author` — tracked through to merge; treat #90's own state as the
|
||||
authority on whether it has landed, not this line).
|
||||
- **`displayName` is gone from all six compiled `plugin.json` files — accepted, not overlooked.**
|
||||
`apm.yml` has no key that compiles to it: `synthesize_plugin_json_from_apm_yml`
|
||||
(`apm_cli/deps/plugin_parser.py`) emits only `name`, `version`, `description`, `author`,
|
||||
`license`, `homepage`, `repository` and `keywords`, and nothing in `plugin_manifest.py` adds
|
||||
`displayName` afterwards. So every `plugins/<name>/.claude-plugin/plugin.json` now carries
|
||||
`author`/`description`/`homepage`/`keywords`/`license`/`name`/`repository`/`version` (plus
|
||||
`mcpServers` for `bin`) and no `displayName`. The field is optional —
|
||||
`plugins/kyberforge/docs/research/docs/claude-code-plugins/api-reference.md:14` lists
|
||||
`displayName` as `Required: No`, "Human-readable name shown in plugin manager" — which is why
|
||||
`claude plugin validate --strict` still passes on all six. The visible cost is that the plugin
|
||||
manager falls back to the bare `name` as each plugin's label. Accepted as the price of `apm.yml`
|
||||
being the single authoring source: re-injecting `displayName` post-compile would mean a second
|
||||
`reinject_*` workaround of the kind ADR-0017's amendment reserves for fields apm strips on a
|
||||
factually wrong premise, and apm's premise here is simply that the key does not exist in its
|
||||
schema.
|
||||
- **`owner.email` was dropped by mistake and has been restored (2026-08-14).** An earlier revision
|
||||
of this ADR listed `owner.email` alongside `displayName` as a field `apm.yml` "has no key that
|
||||
compiles to." That was wrong. `apm_cli/marketplace/yml_schema.py:186` defines
|
||||
`_AUTHOR_OBJECT_KEYS = frozenset({"name", "email", "url"})`, and an `email:` under root
|
||||
`apm.yml`'s `marketplace.owner` block was empirically confirmed to compile straight through into
|
||||
`.claude-plugin/marketplace.json`'s `owner`. The key is declared in root `apm.yml` again and the
|
||||
compiled `owner` block is `{name, email, url}`. Only `displayName` is a genuine schema gap; this
|
||||
one was a documentation error that removed working configuration.
|
||||
- **`mattpocock-skills` is pinned to an exact version, and the pin is advanced by hand.**
|
||||
Pre-conversion the entry was `{"repo": "mattpocock/skills", "source": "github"}` — an unpinned
|
||||
reference that tracked the upstream default branch, so consumers got whatever was on it at
|
||||
install time. The conversion first replaced that with `version: "^1.2.0"`, which was still not a
|
||||
pin: a caret range has nothing to freeze it, because there is no lockfile for
|
||||
`marketplace.packages[]`. `apm pack` re-resolved the range against upstream on **every** run, so
|
||||
an upstream `v1.2.4` would immediately invalidate the committed `ref`/`sha` and fail
|
||||
`apm-pack-check-clean` with exit 4 — blocking every push in the repo, triggered by a third party
|
||||
at an unrelated moment, with no local change to explain it. Root `apm.yml` therefore declares an
|
||||
exact `version: "1.2.3"`, which `apm pack` freezes into `.claude-plugin/marketplace.json` as
|
||||
`ref: v1.2.3` + an explicit `sha`. Two consequences, both intended: the committed ref/sha is
|
||||
genuinely reproducible and cannot move under the repo, and picking up a new upstream release is a
|
||||
deliberate act — a human edits the `version:` string in root `apm.yml` and re-runs `apm pack`.
|
||||
apm has no version-bump automation (established under "Versioning" in issue #90's plan), so an
|
||||
ageing pin is the accepted cost of a push gate that only fires on this repo's own changes.
|
||||
Note the pin does not make the entry offline-resolvable: an exact version still requires a
|
||||
`git ls-remote`, which is why two pre-push hooks needed the network (see `AGENTS.md`).
|
||||
**Superseded 2026-09-13:** the `mattpocock-skills` entry has been removed from root `apm.yml`
|
||||
entirely, along with the `codex` marketplace output profile. No pre-push hook needs the network
|
||||
any longer.
|
||||
- **Caveat on "Status: executed" above:** issue #90's own execution comment flagged, before merge,
|
||||
that Claude Code's ability to actually load content out of `.apm/` was unverified — that caveat
|
||||
turned out to be a real defect, not a formality: the native installer has zero awareness of
|
||||
`.apm/` and reported `Skills (0) Agents (0) Hooks (0)` on every plugin installed from this
|
||||
marketplace. The manifest-compilation deliverable this ADR describes was genuinely complete;
|
||||
runtime discoverability was not. Fixed in ADR-0017 (a second, compiled flat-directory content
|
||||
mirror at each plugin root, generated by `scripts/sync-plugin-content.sh`) — see that ADR for
|
||||
the root cause and the fix. **Superseded 2026-09-14:** ADR-0024 deleted that mirror and its
|
||||
generator along with native `claude plugin install` support. The discoverability gap this
|
||||
paragraph describes is therefore no longer bridged — it is no longer a gap this repo has, because
|
||||
apm is now the only supported install path and apm reads `.apm/` directly.
|
||||
@@ -0,0 +1,174 @@
|
||||
# Plugin-scope agent-author omits `tools:` and all Claude-only fields from `.apm/agents/*.agent.md`
|
||||
|
||||
This ADR is a narrower, downstream consequence discovered while designing issue #89's
|
||||
implementation under ADR-0015's broader direction (Microsoft APM replaces hand-authored
|
||||
plugin/marketplace authoring). It does not restate ADR-0015's rationale — see that ADR for
|
||||
the parent decision.
|
||||
|
||||
## Context
|
||||
|
||||
APM's agent primitive (`.apm/agents/<name>.agent.md`) has no per-target integrator in
|
||||
`apm compile` — confirmed via APM's own Python source (`integration/targets.py` and related
|
||||
files, cited in `plugins/kyberforge/docs/research/docs/microsoft-apm/agent-primitive-schema.md`).
|
||||
Compilation does a naive verbatim copy of the whole frontmatter and body to both the Claude
|
||||
Code and Copilot CLI targets. This is unlike:
|
||||
|
||||
- The **skill** primitive, which is also a straight copy (confirmed in the same research doc)
|
||||
but has no field semantics to conflict — `SKILL.md`'s content is target-agnostic already.
|
||||
- The **prompt**, **instructions**, and **hooks** primitives, which each get real per-target
|
||||
reconstruction through a dedicated integrator (field allowlisting, key renaming, dropped-field
|
||||
warnings).
|
||||
|
||||
Because the agent primitive ships the same frontmatter unchanged to both harnesses, two
|
||||
concrete incompatibilities surface:
|
||||
|
||||
1. **`tools:`** — Claude Code expects tool names drawn from its own vocabulary, as a
|
||||
comma-separated string or a YAML list (`agent-definition.md:37`); Copilot CLI expects a list
|
||||
drawn from a different alias vocabulary (`execute`/`read`/`edit`/`search`/`agent`/`web`). The
|
||||
incompatibility is the vocabulary, not the punctuation: a value correct for one harness names
|
||||
tools the other does not have.
|
||||
2. **Claude-only knobs with no Copilot equivalent** — `isolation`, `maxTurns`, `effort`,
|
||||
`memory`, `permissionMode`. Writing any of these means Copilot's copy carries frontmatter
|
||||
keys it doesn't recognize at all. Whether Copilot's agent loader ignores unknown keys or
|
||||
errors on them is unconfirmed by research. *(Still unconfirmed as of the 2026-08-14 amendment
|
||||
below, which admits `disallowedTools` as an explicitly accepted risk rather than by resolving
|
||||
this question.)*
|
||||
|
||||
## Decision
|
||||
|
||||
At **plugin scope only** (destination package has an `apm.yml` at its root — an APM producer
|
||||
package compiled via `apm compile`), `.apm/agents/<name>.agent.md` carries only `name`,
|
||||
`description`, `model`, and the prose body. No `tools:` field, no Claude-only fields, at all.
|
||||
|
||||
*(Narrowed by the 2026-08-14 amendment below: `disallowedTools` is admitted as a fifth allowed
|
||||
field. `tools:` and every other Claude-only knob remain excluded on the reasoning given here.)*
|
||||
|
||||
Absent `tools:` means inherit-all-tools on both harnesses — the one value that is never wrong
|
||||
on either target, unlike a present, harness-specific value that is guaranteed wrong on at least
|
||||
one of them.
|
||||
|
||||
`agent-audit`, at plugin scope, is intended to flag — as a **SUGGESTION**, not a FAIL, since
|
||||
this is an upstream schema limitation rather than an authoring mistake — any agent whose
|
||||
description or body implies a need for tool restriction or a Claude-only behavior the
|
||||
frontmatter can no longer express. This would give visibility into the gap without pretending
|
||||
the schema can do something it can't. **Not yet implemented**: `check_apm_agent_file()` in
|
||||
`validate.sh` currently validates only the field allowlist, `name`, `description`, and
|
||||
body-emptiness/length — it has no heuristic for this case. Tracked as follow-up work.
|
||||
|
||||
### Scope boundary
|
||||
|
||||
This decision applies to **plugin-scope `agent-author` only**. Project scope (`.claude/agents/`
|
||||
+ `.github/agents/`) and user scope (`~/.claude/agents/` + `~/.copilot/agents/`) are not APM
|
||||
packages — neither goes through `apm compile` — so both keep today's dual-file Claude+Copilot
|
||||
pair model exactly as ADR-0005 and ADR-0008 already describe. Those two ADRs remain fully
|
||||
authoritative for project and user scope; only their plugin-scope clauses are affected by this
|
||||
ADR (see the update notes appended to each).
|
||||
|
||||
## Considered options
|
||||
|
||||
**Pick one harness's vocabulary and accept breakage on the other (rejected).** E.g. always
|
||||
write Claude's space-separated `tools:` string. Rejected because it ships a value that is
|
||||
silently wrong (or possibly a hard error) on Copilot, and which harness "wins" would be an
|
||||
arbitrary, undocumented asymmetry.
|
||||
|
||||
**Same as above, but `agent-audit` flags the cross-harness breakage as a tracked finding
|
||||
(rejected).** Rejected for the same core reason — it still ships a wrong value to a real
|
||||
harness. Tracking the breakage doesn't prevent it, and the chosen decision already gets
|
||||
equivalent visibility (a SUGGESTION finding) without ever shipping the wrong value in the first
|
||||
place.
|
||||
|
||||
## Amendment (2026-08-14): the write fence comes back as a denylist
|
||||
|
||||
The decision above generalised from `tools:` to "no tool restriction at all". That over-reached.
|
||||
The unportability argument is specific to the **allowlist**: Claude Code reads `tools:` as a
|
||||
delimited string of its own tool names, Copilot CLI reads it as a list drawn from its
|
||||
alias vocabulary (`execute`/`read`/`edit`/`search`/`agent`/`web`), so one value is wrong on one
|
||||
harness. That reasoning stands, and `tools:` stays out of every plugin-scope agent.
|
||||
|
||||
A **denylist** has no such conflict. The evidence for that splits three ways, and this amendment
|
||||
states which part is which rather than asserting the whole as settled.
|
||||
|
||||
**Confirmed — Claude Code honours it for plugin subagents.**
|
||||
`plugins/kyberforge/docs/research/docs/claude-code-plugins/agent-definition.md:39` documents
|
||||
`disallowedTools` as a "Denylist applied before `tools`… Takes precedence over `tools`", and — the
|
||||
part that matters here — it is **not** in that document's plugin-subagent ignore list. Line 99
|
||||
names exactly three fields plugin agents silently ignore: `hooks`, `mcpServers`, `permissionMode`.
|
||||
`disallowedTools` is absent from that list. Claude Code is also the harness where the fence is
|
||||
actually wanted, so the field earns its place on this evidence alone.
|
||||
|
||||
**Inferred — the field is very likely inert on Copilot CLI, but by analogy, not by documentation.**
|
||||
`plugins/kyberforge/docs/research/docs/github-copilot-plugins/troubleshooting.md:50` and `:53`
|
||||
record Copilot *silently ignoring* two agent frontmatter fields it does not process (`mcp-servers`
|
||||
and `metadata` outside the cloud runtime) rather than erroring on them. That is a documented
|
||||
tolerance for *known-but-unprocessed* keys, which is adjacent to, not identical to, tolerance for
|
||||
an *unknown* key. No stronger evidence exists: a sweep of the vendored Copilot corpus
|
||||
(`agent-definition.md`, `api-reference.md`, `troubleshooting.md`, `configuration.md`) documents
|
||||
unknown-key handling nowhere.
|
||||
|
||||
**Unverified — Copilot's loader behaviour on an unrecognised key.** Context item 2 above says this
|
||||
is unconfirmed by research and that remains true; nothing found since changes it. An earlier
|
||||
revision of this amendment claimed "an unrecognised frontmatter key is inert" as settled fact and
|
||||
attributed it to apm's verbatim-copy behaviour. That attribution was a non-sequitur — verbatim copy
|
||||
describes what *apm* does at compile time and says nothing about what *Copilot* does at load time —
|
||||
and the claim contradicted this ADR's own Context section.
|
||||
|
||||
**So this is an accepted risk, stated as one.** Blast radius if the inference is wrong and Copilot
|
||||
errors on the key: the three affected plugin-scope agents fail to load under Copilot CLI. It is
|
||||
loud, not silent; it is confined to three agents in three plugins; no other primitive and no Claude
|
||||
Code path is affected; and the remedy is a one-line frontmatter deletion. What the denylist shape
|
||||
*does* rule out categorically — independent of loader behaviour — is the failure mode that motivated
|
||||
dropping `tools:` in the first place: a denied name the other harness does not recognise denies
|
||||
nothing, so a mis-shaped value can never grant or misroute a capability. The risk is a load failure,
|
||||
never a silent over-grant. That asymmetry is why the same verbatim copy that makes `tools:`
|
||||
unshippable makes `disallowedTools` worth shipping.
|
||||
|
||||
So the read-only orchestrator agents regain their write fence: `gitea-orchestrate`,
|
||||
`apm-orchestrate` and `lint-runner` each carry `disallowedTools: Edit, Write, NotebookEdit` plus
|
||||
explicit prose in the body stating the agent does not edit files. `git-orchestrate` is deliberately
|
||||
excluded — it legitimately declared `edit` before the conversion and still needs to write.
|
||||
|
||||
**Residual — the fence is partial, and the prose is doing more of the work than the field is.**
|
||||
`disallowedTools: Edit, Write, NotebookEdit` denies exactly those three tools. It does not deny
|
||||
`Bash`, and at plugin scope these agents carry no `tools:` and therefore inherit it, so
|
||||
`bash -c 'echo … > f'` remains unfenced by frontmatter. Only the body prose covers that path. This
|
||||
is not a regression introduced here — the pre-conversion `tools:` allowlists also granted `Bash`,
|
||||
so the shell route was open then too — but the ADR should not credit the mechanism with more than
|
||||
it delivers. Closing it would need a `disallowedTools` entry for `Bash`, which these agents cannot
|
||||
take because they legitimately shell out.
|
||||
|
||||
Net position: the allowlist stays dropped for the reason originally given, and the denylist is
|
||||
admitted as the portable-by-construction half of what was lost. It restores a real, Claude-Code-
|
||||
confirmed write fence against the tool-call path, not a complete write sandbox. The consequence
|
||||
below is narrowed accordingly.
|
||||
|
||||
Enforcement follows the decision: `agent-audit`'s plugin-scope validator reads its allowlist as
|
||||
data from the `apm-agent-allowlist` section of
|
||||
`plugins/kyberforge/.apm/skills/agent-audit/references/field-inventory.md`, and that line now reads
|
||||
`name description model source_keys disallowedTools`. `disallowedTools` also stays in that file's
|
||||
`claude-code-only-fields` list, which is not a contradiction — that list governs whether a field
|
||||
may cross the CC/Copilot boundary in a real project/user-scope *pair*, a different question from
|
||||
whether a field is safe under verbatim copy in a single vendor-neutral file.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Every plugin-scope APM agent loses per-agent tool *allowlisting* and any Claude-only capability
|
||||
(isolation, maxTurns, effort, memory, permissionMode) until APM ships a real per-target
|
||||
integrator for the agent primitive. This is a known, accepted regression, not an oversight.
|
||||
Tool **denial** is not part of that loss — see the 2026-08-14 amendment above.
|
||||
- **ADR-0005 is partially superseded** — its plugin-scope clause ("directory containing
|
||||
`plugin.json` is plugin scope → both files land in `<root>/agents/`") no longer applies.
|
||||
Plugin scope is now "directory containing `apm.yml` → single vendor-neutral file lands in
|
||||
`<root>/.apm/agents/`." Project and user scope, and the rest of ADR-0005, are unaffected.
|
||||
- **ADR-0008 is partially superseded** — its counterpart-derivation/pair-validation mechanism
|
||||
no longer applies at plugin scope; `agent-audit` takes the single file directly there. Project
|
||||
and user scope, where a real pair still exists, are unaffected.
|
||||
- **ADR-0009 is not superseded.** The mechanism it established — `agent-audit` reading field
|
||||
lists from `references/field-inventory.md` rather than hardcoding them, with a `source_keys`
|
||||
provenance chain — survives and is reused. Only the *content shape* changes for plugin scope:
|
||||
`field-inventory.md` shifts from two side-by-side CC-only/Copilot-only blocklists to one
|
||||
vendor-neutral allowlist for plugin-scope agents, while continuing to serve its original
|
||||
two-blocklist role for project/user-scope validation. That file's `apm-agent-allowlist` section
|
||||
is the authoritative list and is read as data by `validate.sh`; as amended on 2026-08-14 it holds
|
||||
`name`/`description`/`model`/`source_keys`/`disallowedTools` — `source_keys` for provenance
|
||||
tracking, validated separately by `validate-provenance.sh` against `sources.md` rather than being
|
||||
a provider-specific field, and `disallowedTools` per the amendment above.
|
||||
@@ -0,0 +1,360 @@
|
||||
# Plugin roots gain a compiled flat-directory mirror of `.apm/` content so Claude Code can discover it
|
||||
|
||||
**Superseded by:** ADR-0024 (apm is the only supported install path; the flat content mirror is
|
||||
deleted). The mirror this ADR created — `plugins/<name>/{skills,agents,hooks}/` — has been deleted,
|
||||
along with `scripts/sync-plugin-content.sh`, its test suite, and the `check-plugin-content-sync`
|
||||
pre-push gate. Native `claude plugin install` is no longer a supported path, so the host discovery
|
||||
contract this ADR bridged is no longer one this repo satisfies. The diagnosis below is still
|
||||
accurate about how Claude Code's installer works; what changed is that nothing consumes it. The
|
||||
`mcpServers`, `hooks`-pointer and `hooks/hooks.json` amendments below are moot with the artifacts
|
||||
they governed; the symlink amendment's underlying gap is not — see ADR-0024's consequences. This
|
||||
ADR's content is kept below as the historical record; it is no longer the current model.
|
||||
|
||||
---
|
||||
|
||||
This ADR is a follow-on correction to ADR-0015 (Microsoft APM replaces hand-authored
|
||||
plugin/marketplace authoring), discovered during issue #90's post-execution review. It does not
|
||||
restate ADR-0015's rationale for adopting `.apm/` as the authoring source of truth — see that ADR
|
||||
for the parent decision. It resolves the one question ADR-0015's own execution flagged as open but
|
||||
did not block on: whether Claude Code's installer can actually load content out of `.apm/`. It
|
||||
could not.
|
||||
|
||||
**Status: executed (2026-08-13, issue #90).** `scripts/sync-plugin-content.sh` has been run
|
||||
against all 6 plugins; flat `agents/`, `skills/`, `commands/` (etc., wherever `.apm/` populates
|
||||
them), and a merged hooks file now exist at each plugin root as tracked, generated files. The
|
||||
merged hooks file lands at `hooks/hooks.json`, not at the plugin root itself — see the second
|
||||
amendment below, which corrects the path this ADR originally recorded.
|
||||
|
||||
## Context
|
||||
|
||||
ADR-0015's execution comment on issue #90 (2026-08-12) flagged, before merge: "it's currently
|
||||
unverified whether Claude Code can actually discover any skill/agent content in these plugins...
|
||||
This needs to be checked... before treating this conversion as functionally complete, not just
|
||||
manifest-complete." That caveat did not block ADR-0015 from shipping "Status: executed" — the
|
||||
manifest-compilation deliverable (`.claude-plugin/marketplace.json`/`plugin.json` generated from
|
||||
`apm.yml` + `.apm/`) was genuinely complete, and every automated gate (`apm audit --ci`,
|
||||
`claude plugin validate --strict` ×6, `apm marketplace check`) passed clean — so the ADR merged
|
||||
with the caveat noted but unresolved.
|
||||
|
||||
The caveat turned out to be a real defect, not a formality. `claude plugin install` against all
|
||||
three plugins tested (`git@holocron`, `gitea@holocron`, `kyberforge@holocron`) reported
|
||||
`Skills (0) Agents (0) Hooks (0)`. Root cause, confirmed two independent ways:
|
||||
|
||||
1. **Claude Code's installer scans flat convention directories only.** `strings` on the installed
|
||||
`claude` binary finds zero references to `.apm/` or `apm.yml` anywhere. The installed plugin
|
||||
cache (`~/.claude/plugins/cache/holocron/kyberforge/1.3.1/`) mirrors the pre-conversion flat
|
||||
`skills/`/`agents/`/`hooks/` layout verbatim — that is what the installer actually copies and
|
||||
reads. `plugins/kyberforge/docs/research/docs/claude-code-plugins/configuration.md`'s own
|
||||
"Plugin Directory Layout" table documents the same flat convention (`skills/<name>/SKILL.md`,
|
||||
`agents/`, `hooks/hooks.json`, all "at the plugin root, not inside `.claude-plugin/`") — this
|
||||
was accurate before ADR-0015 and never stopped being accurate; ADR-0015 moved plugin content
|
||||
without adding a bridge to it.
|
||||
2. **apm's own manifest compiler has no `.apm/` → host-path bridge, by design.**
|
||||
`apm_cli/core/plugin_manifest.py`'s `build_plugin_manifest` docstring states directly:
|
||||
"Convention directories (`agents/`, `skills/`, `commands/`) are auto-discovered by the host, so
|
||||
they are never listed explicitly in the manifest." apm's Claude/Copilot compiler assumes plugin
|
||||
content already lives in those flat root-level directories; it has no model of `.apm/` nesting
|
||||
being host-visible at all, so it never emits anything that would point a host at `.apm/`.
|
||||
|
||||
Separately, `apm_cli/bundle/plugin_exporter.py`'s `export_plugin_bundle` (the engine behind
|
||||
`apm pack --format plugin`) *does* implement the correct mapping — `.apm/agents` → `agents/`,
|
||||
`.apm/skills` → `skills/` (subdirs preserved), `.apm/prompts` + `.apm/commands` → `commands/`
|
||||
(`*.prompt.md` renamed to `*.md`), `.apm/instructions` → `instructions/`, `.apm/extensions` →
|
||||
`extensions/`, and `.apm/hooks/*.json` merged into one `hooks.json`. But it was only ever wired to
|
||||
produce a distributable bundle under `build/<name>-<version>/` — a path nothing in root
|
||||
`apm.yml`'s per-package `marketplace.packages[].source:` fields (e.g. `./plugins/bin`) or
|
||||
`marketplace.json`'s equivalent points at. The correct mapping existed in apm's own codebase the
|
||||
whole time; it was simply never connected to the path this repo's marketplace actually installs
|
||||
plugins from.
|
||||
|
||||
## Decision
|
||||
|
||||
Each plugin root gains a second, generated content category, produced by
|
||||
`scripts/sync-plugin-content.sh` (wraps `apm pack --format plugin`, copies the resulting bundle's
|
||||
`agents/`, `skills/`, `commands/`, `instructions/`, `extensions/`, and merged hooks file back to
|
||||
the plugin root — the hooks file to `hooks/hooks.json`, per the second amendment below) — same
|
||||
governance status as `.claude-plugin/plugin.json`/`marketplace.json`:
|
||||
**compiled output of `.apm/`, never hand-edited.**
|
||||
|
||||
- `.apm/` remains the sole hand-edited authoring source, unchanged from ADR-0015.
|
||||
- The flat mirror is what Claude Code's (and Copilot's) installer actually convention-scans at
|
||||
install time — it exists purely to satisfy the host's discovery contract, a contract apm's own
|
||||
manifest compiler deliberately does not bridge.
|
||||
- `plugin.json`/`apm.lock.yaml`/`.mcp.json` from the bundle are excluded from the copy:
|
||||
`plugin.json` is already correctly generated by a separate, already-verified apm code path
|
||||
(`build_plugin_manifest`, run in the same `apm pack` invocation); `.mcp.json` is hand-authored
|
||||
at the plugin root per ADR-0015 and is not an `.apm/` primitive.
|
||||
- Dev-fixture `tests/` directories are excluded too — they are dev-time fixtures no plugin host
|
||||
ever needs to discover, and several reference their own repo root through a hardcoded relative
|
||||
walk-up sized for `.apm/`-nested depth, so a copy one directory level shallower breaks the
|
||||
duplicate and double-runs the original under repo-wide bats discovery. The exclusion is
|
||||
**depth-scoped to `<category>/<name>/tests`**, deliberately: a skill may legitimately ship a
|
||||
directory literally named `tests` as a template asset it scaffolds *from*
|
||||
(`skills/skill-author/assets/templates/tests`, at depth 4). A depth-agnostic `-name tests`
|
||||
matched that too and stripped it, making the mirrored `new-skill.sh` die mid-run on
|
||||
`sed: can't read .../tests/README.md` — the scaffolder seds its way through the template tree
|
||||
file by file. Scaffolding assets survive; fixtures do not.
|
||||
- Drift is enforced by a pre-push gate (`scripts/sync-plugin-content.sh --check --all`, wired into
|
||||
`.pre-commit-config.yaml` as hook id `check-plugin-content-sync` by a parallel workstream on
|
||||
issue #90) — the same enforcement model `check-manifests.sh` already applies to the other
|
||||
compiled-output category. `--check` alone is not the gate: the script requires either `--all` or
|
||||
an explicit list of plugin directories, and run bare it prints usage and exits 1. `--all` derives
|
||||
its work list from `marketplace.json`, a generated file, so it asserts its own coverage against
|
||||
that list: it fails if it verified fewer plugins than the marketplace declares, not merely if it
|
||||
verified none. A listed plugin whose `.apm/` has gone missing is skipped by the per-plugin sync
|
||||
and would otherwise let the gate report success over a shrinking work list.
|
||||
- Verified two ways before landing: `claude plugin validate --strict` passes on all 6 real
|
||||
(non-scratch) plugin directories, and a live behavioral test
|
||||
(`claude --plugin-dir plugins/kyberforge -p "list your skills and agents"`) against the real
|
||||
committed directory confirms `kyberforge:*` skills and the `kyberforge:apm-orchestrate` agent
|
||||
are now actually discovered — they were not, before this fix.
|
||||
- The stale root-level `plugins/<name>/plugin.json` files (a near-duplicate of
|
||||
`.claude-plugin/plugin.json` that nothing read or wrote, flagged separately in issue #90's
|
||||
review) were deleted across all 6 plugins as part of the same cleanup.
|
||||
|
||||
## Considered options
|
||||
|
||||
**Patch `plugin.json`'s content-pointer fields to point directly at `.apm/` paths (rejected).**
|
||||
Claude Code's manifest schema documents these as legitimate override fields that accept custom
|
||||
paths — `plugins/kyberforge/docs/research/docs/claude-code-plugins/configuration.md` shows a real
|
||||
example (`"skills": "./custom/skills/"`, `"agents": ["./custom/agents/reviewer.md"]`), so the host
|
||||
side of this would work. Rejected because apm never emits such a pointer and would have to be
|
||||
worked around on every run to make it do so.
|
||||
|
||||
Be precise about the mechanism, because an earlier revision of this ADR overstated it. apm 0.28.0's
|
||||
`build_plugin_manifest` (`apm_cli/core/plugin_manifest.py`) does carry a strip loop, but its field
|
||||
list is `("agents", "skills", "commands", "instructions")` — `hooks` is **not** in it, and
|
||||
`instructions` **is**, which this ADR previously did not mention. More to the point, that loop can
|
||||
never fire: the manifest it operates on comes from `synthesize_plugin_json_from_apm_yml`
|
||||
(`apm_cli/deps/plugin_parser.py`), which only ever emits `name`, `version`, `description`,
|
||||
`author`, `license`, `homepage`, `repository` and `keywords`. The pointer fields are absent from
|
||||
apm's output because `apm.yml` has no schema for them, not because apm actively removes them — the
|
||||
`pop` loop is defensive dead code against a manifest shape apm does not produce.
|
||||
|
||||
The rejection is unaffected by that correction, only its framing. Honoring this option would still
|
||||
mean post-processing apm's compiled output on every `apm pack` run to add fields apm's schema has
|
||||
no way to express, rather than reusing `plugin_exporter.py`'s bundle-export mapping, which already
|
||||
does the right thing and only needed its output redirected to a path the installer reads. What it
|
||||
is *not* is a fight against a load-bearing apm code path — the honest statement is that apm has no
|
||||
input for these fields, and inventing one downstream is a workaround this ADR did not need.
|
||||
|
||||
**Point `marketplace.json`'s `source:` at `apm pack`'s `build/<name>-<version>/` output directly
|
||||
(rejected).** Would reuse the bundle exporter's correct mapping without adding a new script.
|
||||
Rejected: `build/` is a version-suffixed, regenerate-on-every-pack directory — pointing the
|
||||
marketplace at it would mean either committing a moving-target build artifact to version control
|
||||
(defeating the point of it being generated) or requiring every consumer's marketplace to run
|
||||
`apm pack` before install, a build step Claude Code's installer has no hook for — it clones/fetches
|
||||
source and scans directories; it does not execute a package manager's build command first.
|
||||
Copying the relevant subset back to the stable `plugins/<name>/` path — where `marketplace.json`
|
||||
already points — needed no change to the marketplace source model at all.
|
||||
|
||||
## Amendment (2026-08-13, revised 2026-08-14): Copilot's `plugin.json` gets an `mcpServers` *path*
|
||||
|
||||
PR #95's review (a follow-on to this same issue #90 workstream) found a second field apm's
|
||||
compiler drops for the Copilot ecosystem: `build_plugin_manifest` runs
|
||||
`manifest.pop("mcpServers", None)` on every Copilot-ecosystem `plugin.json`, its docstring stating
|
||||
the field is "not part of the Copilot plugin manifest schema." That claim is contradicted by this
|
||||
repo's own researched documentation —
|
||||
`plugins/kyberforge/docs/research/docs/github-copilot-plugins/configuration.md:49` documents
|
||||
`mcpServers` as a valid, optional `plugin.json` field, typed **"string or object — MCP server
|
||||
config path or inline definitions."**
|
||||
|
||||
This is not the same situation "Considered options" above rejected. There, apm emits no pointer
|
||||
because its schema has no input for one and the host auto-discovers the directories anyway, so
|
||||
nothing is missing. Here a field Copilot actually reads is actively removed on a premise that is
|
||||
wrong against documented Copilot behavior, and there is no auto-discovery mechanism that makes it
|
||||
redundant. Shipping the manifest as apm produces it would ship a manifest known to be incomplete.
|
||||
|
||||
`scripts/sync-plugin-content.sh`'s `reinject_mcp_servers()`, called from `sync_one()`, therefore
|
||||
sets `mcpServers` on `.github/plugin/plugin.json` after `apm pack` runs — to the **string
|
||||
`".mcp.json"`**, the path form of the documented type, not the resolved server objects. Only when
|
||||
the plugin's `.mcp.json` declares at least one server, matching apm's own Claude-ecosystem builder,
|
||||
which omits the field entirely rather than emitting `mcpServers: {}`.
|
||||
|
||||
**The payload is a path because an inlined object is a credential-leak path.** The original
|
||||
implementation copied `.mcp.json`'s resolved `mcpServers` object into the manifest with `jq`. That
|
||||
route bypasses apm's own `_sanitize_mcp_servers()` (`apm_cli/core/plugin_manifest.py`), which
|
||||
strips credential keys and redacts secret values out of `.mcp.json` precisely because — in its own
|
||||
words — "copying them verbatim into a committed `plugin.json` would exfiltrate them into the
|
||||
distributed artefact." Today's `.mcp.json` files here carry no `env` block, so nothing leaked; the
|
||||
first one that did would have written a live token into a tracked, published manifest, with the
|
||||
sanitizer sitting one code path away and never invoked. A path reference cannot carry a secret at
|
||||
all: the manifest names a file, and resolution happens in the host at load time. This also matches
|
||||
apm's documented posture for MCP secrets — `microsoft-apm/configuration.md:96-98` requires `${VAR}`
|
||||
indirection so secrets are "never committed to the manifest."
|
||||
|
||||
**Both modes re-inject**, not just real syncs: real mode writes into the plugin root directly,
|
||||
`--check` into its throwaway copy first, so the manifest diff compares against the same content a
|
||||
real sync would actually produce (see the script's own header). A check-mode re-injection is what
|
||||
keeps `--check` from reporting permanent phantom drift on every plugin that ships an `.mcp.json`.
|
||||
|
||||
This remains scoped to one field found to be incorrectly dropped. It does not reopen the
|
||||
content-pointer option rejected above: those fields stay absent because apm has no schema input for
|
||||
them and the host needs no pointer, which is a different situation from a documented field being
|
||||
actively removed.
|
||||
|
||||
Consequence: if a future apm release corrects the Copilot `mcpServers` omission, `reinject_mcp_servers()`
|
||||
and its call site become dead code and should be deleted — nothing else in this ADR depends on the
|
||||
reinjection existing beyond working around this specific upstream gap.
|
||||
|
||||
Line numbers are deliberately omitted above. An earlier revision of this amendment cited
|
||||
`reinject_mcp_servers()` at line 190 and its call site at line 269; both had already moved by the
|
||||
next review round of the same PR, and moved again with the edits recorded in the amendment below.
|
||||
A function name is stable enough to grep for; a line number in an ADR is stale by the next commit.
|
||||
|
||||
## Amendment (2026-08-14): the merged hooks file lands at `hooks/hooks.json`, not the plugin root
|
||||
|
||||
As originally executed, `sync-plugin-content.sh` wrote the merged hooks file to
|
||||
`plugins/<name>/hooks.json`. That path is scanned by nothing. Claude Code convention-scans
|
||||
`hooks/hooks.json`, and the "Plugin Directory Layout" table this ADR's own root-cause analysis
|
||||
quotes above says so:
|
||||
`plugins/kyberforge/docs/research/docs/claude-code-plugins/configuration.md:100` is the row naming
|
||||
`hooks/hooks.json`, six lines below the table's preamble at `:94` — "All content directories must
|
||||
be at the plugin root, not inside `.claude-plugin/`". The two are not the same line; an earlier
|
||||
revision of this amendment said they were. The implementation read the preamble's "at the plugin
|
||||
root" and dropped the file there, without reading the row that names the path. So this ADR shipped
|
||||
with the contract quoted correctly in its diagnosis and violated in its output — the flat mirror
|
||||
bridged skills and agents into discovery and left hooks exactly as undiscoverable as before the
|
||||
fix.
|
||||
|
||||
The merged file therefore moves to `plugins/<name>/hooks/hooks.json`. A root-level `hooks.json`
|
||||
left over from a prior sync is stale output: a real sync deletes it, `--check` reports it as
|
||||
drift. The real sync produced exactly these working-tree changes — `plugins/kyberforge/hooks.json`
|
||||
and `plugins/lint/hooks.json` deleted, `plugins/kyberforge/hooks/hooks.json` and
|
||||
`plugins/lint/hooks/hooks.json` created. Only those two plugins have an `.apm/hooks/` tree, so
|
||||
only those two grow a mirrored hooks file at all.
|
||||
|
||||
This does **not** reopen the "patch `plugin.json` pointer fields" option rejected above. The move
|
||||
needs no `hooks` pointer in `plugin.json`: `hooks/hooks.json` *is* the convention path, so the
|
||||
host finds it by auto-discovery, exactly as it finds `skills/` and `agents/`. The rejection stands
|
||||
for the reason it was made, once stated accurately — apm emits no pointer field for any of these,
|
||||
because `apm.yml` has no key that produces one, and none is needed when content sits at the
|
||||
convention path. (`hooks` was never in `build_plugin_manifest`'s strip list at all; see the
|
||||
corrected mechanism note under "Considered options".) Writing to the convention path is what makes
|
||||
the no-pointer premise true here rather than something to work around.
|
||||
|
||||
Read "the host finds it by auto-discovery" above as **Claude Code**, not both hosts. Copilot has no
|
||||
default for `hooks` and so discovers none — a real gap, examined and deliberately left open in the
|
||||
next amendment.
|
||||
|
||||
## Amendment (2026-08-14): no `hooks` pointer is re-injected for Copilot — the gap stays documented
|
||||
|
||||
PR #95's review found a third field, and it looks like the `mcpServers` amendment's exact twin:
|
||||
`plugins/kyberforge/docs/research/docs/github-copilot-plugins/configuration.md:47` types `hooks` as
|
||||
a `plugin.json` field, **"string or object"**, with **no default** — so Copilot has no convention
|
||||
path to scan — and `jq 'has("hooks")'` returns `false` for all six `plugins/*/.github/plugin/plugin.json`.
|
||||
Copilot therefore resolves **zero hooks from every plugin in this repo**. The facts are not in
|
||||
dispute; the remedy is.
|
||||
|
||||
State the mechanism correctly first, because it differs from `mcpServers` and the amendment above
|
||||
depends on that distinction. `mcpServers` is *actively removed* — `build_plugin_manifest` runs
|
||||
`manifest.pop("mcpServers", None)` on every Copilot manifest. `hooks` was **never in that strip
|
||||
list** (its field list is `("agents", "skills", "commands", "instructions")`, and the loop is dead
|
||||
code besides — see "Considered options"). This is an absence apm never fills, not a removal to
|
||||
reverse.
|
||||
|
||||
**Decision: do not re-inject. Document the gap.** The `mcpServers` exception was granted on three
|
||||
conditions, and `hooks` meets only two of them:
|
||||
|
||||
1. *A documented host schema field.* Met — `hooks` is in Copilot's own field table.
|
||||
2. *apm has no input that produces it.* Met — `apm.yml` has no key for it.
|
||||
3. *The payload is correct for the host regardless of content.* **Not met**, and this is the whole
|
||||
difference. `.mcp.json` is one host-agnostic format that both ecosystems read, so the string
|
||||
`".mcp.json"` is a true statement about the file no matter what is in it. Hooks have no such
|
||||
shared format: Claude Code reads
|
||||
`{"hooks": {"PreToolUse": [{"matcher": ..., "hooks": [...]}]}}` while Copilot requires
|
||||
`{"version": 1, "hooks": {"sessionStart": [{"type": "command", "bash": ..., "powershell": ...}]}}`
|
||||
— a mandatory `version`, lowercase and differently-named lifecycle events, and per-shell script
|
||||
keys. apm's exporter merges `.apm/hooks/*.json` into **exactly one** `hooks.json` with no
|
||||
per-target shaping (`_collect_hooks_from_apm`, `apm_cli/bundle/plugin_exporter.py`), and that one
|
||||
file also sits at Claude Code's convention path, where Claude Code will read it whatever it
|
||||
contains. So there is exactly one file and two incompatible readers of it.
|
||||
|
||||
A `hooks` pointer would therefore assert that a Claude-shaped file is Copilot-shaped. That trades an
|
||||
*incomplete* manifest for a *wrong* one, which is the opposite of the `mcpServers` amendment's
|
||||
reasoning ("shipping the manifest as apm produces it would ship a manifest known to be incomplete").
|
||||
|
||||
The "it changes nothing today, so it is zero-risk and correct-by-construction for the first real
|
||||
hook" argument does not survive the same check, in both halves. It is not inert today: both
|
||||
`hooks/hooks.json` files are `{"hooks": {}}`, which lacks the `version: 1` Copilot's schema
|
||||
requires, so a pointer would name a file invalid against the schema it is being pointed at from —
|
||||
a change from "declares no hooks" to "declares hooks, at an invalid file". And it is not
|
||||
correct-by-construction later: whoever writes the first real hook writes it in one of the two
|
||||
shapes, and the pointer is wrong in the Claude-shaped case (the case that actually happens, since
|
||||
Claude Code auto-discovers the same file and is what these hooks are authored against) while the
|
||||
Copilot-shaped case breaks Claude Code instead. No content makes both readers correct.
|
||||
|
||||
What would change this decision is upstream, not local: apm emitting a per-target hooks file (at
|
||||
which point a pointer names a file genuinely shaped for its reader), or the two hook schemas
|
||||
converging. Until then the honest artifact is a documented gap, recorded for authors in
|
||||
`plugins/kyberforge/docs/hooks.md` and pinned by a test asserting the Copilot manifest carries no
|
||||
`hooks` key — so that adding one is a deliberate act that has to confront the schema mismatch,
|
||||
rather than a plausible-looking one-liner nobody re-derives.
|
||||
|
||||
This does not weaken the `mcpServers` amendment. That exception was narrow on purpose, and this is
|
||||
what its third condition was for.
|
||||
|
||||
## Amendment (2026-08-14): symlinks under `.apm/` are dropped, and are now reported
|
||||
|
||||
apm's bundle exporter filters symlinks out of the bundle entirely — `f.is_file() and not
|
||||
f.is_symlink()` in `_collect_flat` and `_collect_recursive`, and the same test in
|
||||
`_collect_hooks_from_apm` (`apm_cli/bundle/plugin_exporter.py`). It emits no warning. A symlink
|
||||
placed under a plugin's `.apm/` therefore never reaches the mirror, and until now nothing said so.
|
||||
|
||||
This was **silent content loss, not drift**, and that distinction is why no existing gate caught it.
|
||||
Every other check in `sync-plugin-content.sh` compares the live mirror against a freshly synced
|
||||
copy — and both sides are built from that same bundle. The symlink is absent from both, they agree,
|
||||
and `--check` exits 0. There is no mismatch to detect, only an absence with nothing left to
|
||||
mismatch against. Reproduced on a fixture: `ln -s real.md link.md` under `.apm/skills/hello/`
|
||||
produced a mirror with no `link.md` and a `--check` at exit 0.
|
||||
|
||||
`check_apm_symlinks()` therefore reads the `.apm/` **source** tree directly — the only place the
|
||||
loss is visible — and reports each symlink in both modes, failing the run. It is reported rather
|
||||
than resolved: dereferencing and copying the target would make a real sync emit content the bundle
|
||||
does not contain, which is precisely the "reimplement apm's mapping outside apm" this ADR rejects.
|
||||
Telling the author is the in-contract half.
|
||||
|
||||
The scan covers only the `.apm/` directories apm's exporter actually reads
|
||||
(`agents`, `skills`, `prompts`, `commands`, `instructions`, `extensions`, `hooks`), and carves out
|
||||
`<category>/<name>/tests` to match the mirror's own exclusion — that subtree is not mirrored whether
|
||||
or not it holds a symlink, so nothing is lost there. The carve-out is depth-scoped for the same
|
||||
reason the `tests/` exclusion is: a symlink under `assets/templates/tests` sits in content the
|
||||
mirror does carry, and is reported.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Git now tracks real, visible duplication: `.apm/skills/<name>/SKILL.md` and
|
||||
`skills/<name>/SKILL.md` both exist and must match, likewise `.apm/agents/*.agent.md` vs.
|
||||
`agents/*.agent.md`, and `.apm/hooks/*.json` vs. the merged `hooks/hooks.json` (see the
|
||||
2026-08-14 amendment above for that path). This is an accepted
|
||||
tradeoff of bridging a gap apm itself doesn't close, not a bug — `.apm/` stays the single
|
||||
hand-edited source, and the drift gate (`check-plugin-content-sync`) is what keeps the mirror
|
||||
honest rather than trusting authors to remember to regenerate it by hand.
|
||||
- `scripts/check-manifests.sh`'s existing blind spot (flagged in the same issue #90 review round:
|
||||
it validated `plugin.json` fields that ADR-0015 already stopped populating, so a plugin shipping
|
||||
zero content could pass it silently) is fixed as part of the same workstream: those field checks
|
||||
are removed (nothing to check — the fields are correctly absent by design), and the
|
||||
content-presence question they were standing in for is now answered by
|
||||
`check-plugin-content-sync`, not re-implemented inside `check-manifests.sh`.
|
||||
- ADR-0015's "Status: executed" now carries a pointer to this ADR (see that ADR's Consequences)
|
||||
rather than being rewritten — the manifest-compilation half of its execution was correct and
|
||||
stands; this ADR fixes the second, previously-unverified half.
|
||||
- `CONTEXT.md`'s "Plugin" and "Plugin marketplace" glossary entries are updated to describe the
|
||||
flat mirror as a second compiled-output category, alongside the existing
|
||||
`.claude-plugin/plugin.json`/`marketplace.json` description. Superseded 2026-08-17: CONTEXT.md was
|
||||
cut back to one-line definitions and no longer describes either compiled-output category;
|
||||
`docs/spec/architecture.md` is where the mirror is documented.
|
||||
- A future apm release that ships a native `.apm/`-aware plugin.json compiler (closing this gap
|
||||
upstream) would let `sync-plugin-content.sh` and its drift gate be deleted outright — nothing in
|
||||
this ADR's decision depends on the flat mirror existing beyond satisfying the current installer's
|
||||
convention-scan contract.
|
||||
- **Reproduction note (2026-08-13):** the live behavioral test cited in "Decision" above
|
||||
(`claude --plugin-dir plugins/kyberforge -p "list your skills and agents"`) is only a clean
|
||||
kyberforge-only signal when run from a working directory outside this repo. Run literally as
|
||||
written, from this repo's root, this repo's own project-level `.claude/settings.json` sets
|
||||
`enabledPlugins` to true for all 6 holocron plugins (kyberforge, git, gitea, core, lint, bin), so
|
||||
Claude Code loads all 6 plugins' skills/agents, not just kyberforge's — conflating kyberforge's
|
||||
discoverability with the other 5 plugins' already-enabled content. To isolate the signal, run
|
||||
from a neutral cwd outside `/root/ai-development` with an absolute `--plugin-dir` path, e.g.
|
||||
`cd /some/neutral/dir && claude --plugin-dir /root/ai-development/plugins/kyberforge -p "list your skills and agents"`.
|
||||
- Reference: issue #90 (https://git.dev.rkdr.net/Defame1297/holocron/issues/90).
|
||||
172
docs/adr/0018-repo-consumes-its-own-plugins-through-apm.md
Normal file
172
docs/adr/0018-repo-consumes-its-own-plugins-through-apm.md
Normal file
@@ -0,0 +1,172 @@
|
||||
# This repo installs its own plugins through apm, not Claude Code's native plugin install
|
||||
|
||||
ADR-0015 moved plugin **authoring** to apm; ADR-0017 added the flat content mirror that keeps the
|
||||
authored `.apm/` tree discoverable by hosts that install natively. Both are about producing the
|
||||
marketplace. This ADR is about consuming it: how the plugins get onto the machine this repo is
|
||||
worked on.
|
||||
|
||||
**Correction (2026-09-14): the flat content mirror named above no longer exists.** ADR-0017 is
|
||||
superseded by ADR-0024, and commit `718c79a` deleted the mirror
|
||||
(`plugins/<name>/{skills,agents,hooks}/`) along with its generator, its test suite and its pre-push
|
||||
gate; native `claude plugin install` is no longer a supported path, so there are no longer "hosts
|
||||
that install natively" for it to serve. Nothing this ADR decides depends on the mirror — it appears
|
||||
here only as the other half of "producing the marketplace", and once more under "Install output is
|
||||
gitignored" below, where the two copies of plugin content ADR-0017 governed are now one, `.apm/`
|
||||
itself, and committing the deployed skills would make a second rather than a third. Read both
|
||||
mentions as historical.
|
||||
|
||||
**Status: executed (2026-08-14).** All six packages are installed into `/root/ai-development` by
|
||||
`apm install`; the six native project-scope installs (`claude plugin uninstall <name>@holocron
|
||||
--scope project`) are gone and `.claude/settings.json`'s `enabledPlugins` block is empty.
|
||||
|
||||
## Context
|
||||
|
||||
Until now the repo consumed its own output the same way any user would: `claude plugin install
|
||||
<name>@holocron`, six plugins enabled per-project in `.claude/settings.json`, the `holocron`
|
||||
marketplace registered in `~/.claude/plugins/known_marketplaces.json` with `autoUpdate: true`.
|
||||
That worked. It also meant the repo's dogfooding stopped one layer short of the tooling it
|
||||
publishes: `kyberforge` ships `apm-workflow` and `apm-install` skills describing an install path
|
||||
the repo itself did not take.
|
||||
|
||||
apm supports both scopes. `apm install --global` deploys to `~/.claude/`; plain `apm install`
|
||||
deploys to the project. Global was rejected deliberately — the switch should be provable in one
|
||||
repo before it changes how every other project on the machine resolves its skills.
|
||||
|
||||
## Decision
|
||||
|
||||
Root `apm.yml` declares all six packages under `dependencies.apm`, each as a `git:`/`path:` object
|
||||
against the holocron remote:
|
||||
|
||||
```yaml
|
||||
dependencies:
|
||||
apm:
|
||||
- git: git@git.dev.rkdr.net:Defame1297/holocron.git
|
||||
path: plugins/core
|
||||
```
|
||||
|
||||
`apm install` deploys them to `.claude/skills/<name>/` and `.claude/agents/<name>.md`.
|
||||
|
||||
Three sub-decisions inside that:
|
||||
|
||||
- **Object form over the `<name>@holocron` marketplace alias.** The alias is shorter and apm
|
||||
resolves it correctly (verified end-to-end against this remote), but it first requires
|
||||
`apm marketplace add`, which writes to `~/.apm/marketplaces.json` — user scope, outside the
|
||||
repo, and absent on a fresh clone. The object form needs nothing beyond the committed manifest.
|
||||
- **Remote source over local path.** apm accepts `path: /root/ai-development/plugins/<name>` as a
|
||||
local dependency, which would make the working tree live instantly. Rejected: it erases the
|
||||
distinction between editing a skill and shipping one, which is the entire point of having a
|
||||
marketplace. The remote form keeps the repo running the same released content every other
|
||||
consumer gets.
|
||||
- **Unpinned against the default branch.** Parity with the `autoUpdate: true` the native install
|
||||
had. apm warns on every install (`6 dependencies unpinned`); accepted knowingly. Pinning is a
|
||||
per-entry `ref:` away once the repo tags releases per package — today `git tag` lists one tag
|
||||
total, so there is nothing meaningful to pin to.
|
||||
|
||||
## Consequences
|
||||
|
||||
**Skills gain an unnamespaced name.** apm deploys plain project skills, so `git:git-commits` also
|
||||
answers to `git-commits` and `kyberforge:skill-audit` to `skill-audit`. This is not configurable —
|
||||
a project skill has no plugin to prefix. `AGENTS.md` and `CONTEXT.md` are updated to name the bare
|
||||
form, which is what apm deploys and the only form a repo consuming holocron through apm gets.
|
||||
|
||||
**Correction (2026-08-14): the namespaced form did not stop resolving.** *Superseded by the
|
||||
2026-08-17 correction below: the machine state this cites is no longer present. Both are kept
|
||||
because the pair is the finding — read neither as current.* An earlier revision of
|
||||
this consequence said every `<plugin>:<skill>` reference "was stale the moment the switch landed",
|
||||
and `AGENTS.md`/`CONTEXT.md` were written to match. That contradicts the "User scope is untouched,
|
||||
deliberately" consequence below, and the contradiction resolves against it: `~/.claude.json` still
|
||||
enables `core`, `git`, `gitea`, `kyberforge` and `lint` at user scope, so both names are live at
|
||||
once and a working `gitea:gitea-prs` is the user-scope copy answering. That doubling is the same
|
||||
"present twice under two names" outcome the "Keeping both install paths" alternative was rejected
|
||||
for — reached by leaving user scope alone rather than by adopting it, which is why it is a
|
||||
consequence to record rather than a decision to revisit. Prefer the bare name regardless: it
|
||||
survives those user-scope installs eventually being converted, and the namespaced form still
|
||||
resolves for anyone installing holocron natively, so skill bodies written for both audiences
|
||||
should name the bare skill.
|
||||
|
||||
**Correction (2026-08-17): the evidence under the correction above is gone, and the claim goes with
|
||||
it — not to its opposite.** Observed on this machine: `~/.claude/plugins/installed_plugins.json` is
|
||||
`{"version": 2, "plugins": {}}`; there is no `enabledPlugins` key anywhere in `~/.claude.json`
|
||||
(`grep -c enabledPlugins` returns 0); `~/.apm/marketplaces.json` is `{"marketplaces": []}`. The
|
||||
`holocron` entry in `~/.claude/plugins/known_marketplaces.json` survives, but a registered
|
||||
marketplace is not an installed plugin. So the user-scope installs the 2026-08-14 correction cited
|
||||
are not there, and neither is the state the *original* consequence described before it. The claim
|
||||
about the namespaced form has now been written twice off two different observations of the same
|
||||
machine, and this ADR has already reversed itself once on it. That is the finding: the fact is
|
||||
machine state, not a property of this decision, and it changes without any commit. No instruction
|
||||
file — `AGENTS.md`, `CONTEXT.md`, or a skill body — should assert either way whether
|
||||
`<plugin>:<skill>` resolves. The rule that survives every observation is the one that was always the
|
||||
actionable half: write the bare name, because it is the only form `apm install` produces.
|
||||
|
||||
**apm owns `.claude/settings.json`.** (ADR-0019 supersedes the "exactly `{"hooks": {}}`" claim
|
||||
below — once a package ships a hook, apm merges it into that file and the merged entry is apm's own
|
||||
output. The rule that nothing repo-authored goes in the file is unchanged.) `apm audit --ci` replays the install into a scratch tree and
|
||||
diffs it against the worktree. apm's hook integrator writes that file, so the replay expects
|
||||
exactly what apm would have written — `{"hooks": {}}` — and any repo-owned key in it is permanent
|
||||
drift that fails the `apm-audit-ci` pre-push hook. Verified both directions: with the pre-existing
|
||||
`enabledPlugins` block present, `1 of 10 check(s) failed`; reduced to `{"hooks": {}}`,
|
||||
`All 10 check(s) passed`. Nothing was lost in that reduction — `enabledPlugins` was empty after the
|
||||
native uninstall and the only `hooks` entry was an empty `PreToolUse: []` — but it does mean the
|
||||
file is no longer available for repo-owned settings. Machine-specific settings go in the gitignored
|
||||
`.claude/settings.local.json`, which apm does not deploy; shared enforcement belongs in
|
||||
`.pre-commit-config.yaml`, where this repo already keeps it.
|
||||
|
||||
**`apm_modules/` breaks naive tree walks.** apm materializes a full copy of every dependency there
|
||||
— including this repo's own plugins, `.bats` files and all. The dependency copies resolve their
|
||||
bats helpers relative to their own root, not this repo's, so `tests/run-tests.sh` went from 167
|
||||
tests passing to `334 tests, 167 failures` on the first install. Both discovery walks
|
||||
(`tests/run-bats.sh`, `tests/run-tests.sh`) now exclude `apm_modules/`, on the find side and on the
|
||||
`git ls-files` side that derives the expected set. Any future script that walks the repo tree needs
|
||||
the same exclusion.
|
||||
|
||||
**Install output is gitignored; the lockfile is not.** `.claude/skills/`, `.claude/agents/`, and
|
||||
`apm_modules/` are regenerated by `apm install`. Committing the deployed skills would add a third
|
||||
mirror of content ADR-0017 already governs two copies of. `apm.lock.yaml` is committed — it is what
|
||||
makes the install reproducible, and `apm audit --ci` checks it.
|
||||
|
||||
**MCP survived the switch; hooks were never at risk.** apm read `plugins/bin/.mcp.json` as a
|
||||
self-defined direct-dependency MCP server and configured `obsidian` into the repo's `.mcp.json`
|
||||
unprompted. The `gitea` and `context7` servers were never plugin-provided — they live in
|
||||
`~/.claude.json` and are untouched. Every plugin's `.apm/hooks/hooks.json` is `{"hooks": {}}`, so
|
||||
apm's "contributed no entries to claude settings; skipped" warning on `kyberforge` and `lint` is
|
||||
accurate and harmless.
|
||||
|
||||
**Correction (2026-09-14): the MCP propagation above stopped operating, and the files it read are
|
||||
deleted.** It ran on one code path only — `apm_cli/deps/plugin_parser.py` maps a plugin-root
|
||||
`.mcp.json` onto `.apm/.mcp.json` for packages apm treats as *marketplace plugins*. Commit
|
||||
`718c79a` (ADR-0024) deleted every `plugins/*/.claude-plugin/plugin.json` and
|
||||
`plugins/*/.github/plugin/plugin.json`, so each package is now a plain apm package and that path no
|
||||
longer runs. The supported declaration was never in use either: `plugins/bin/apm.yml` has
|
||||
`dependencies.mcp: []`. That left the six plugin-root `.mcp.json` files dead config — five of them
|
||||
empty stubs, only `plugins/bin`'s carrying the `obsidian` server — and all six are now deleted along
|
||||
with the server itself, which is not wanted. The repo-root `.mcp.json` was apm's own generated
|
||||
output that happened to be tracked; it is deleted and gitignored, on the same reasoning as
|
||||
`.claude/skills/`. The rest of this paragraph is unaffected: `gitea` and `context7` were never
|
||||
plugin-provided, and the hooks claim never depended on any of this.
|
||||
|
||||
**A `.apm/` edit now needs a round trip.** The dependency resolves from the remote, so an edit is
|
||||
invisible to the running session until it is pushed and the install is refreshed. Under the native
|
||||
install with `autoUpdate` the shape was the same; it was more noticeable here at first because the
|
||||
refresh is a manual step where marketplace auto-update was not — ADR-0019 automates it at
|
||||
`SessionStart`.
|
||||
|
||||
**Correction (2026-08-14): the refresh command is `apm update`, not `apm install`.** An earlier
|
||||
revision of this paragraph named `apm install`, which is wrong: `apm install` deploys from the
|
||||
pinned `resolved_commit` in `apm.lock.yaml` and does not re-resolve refs (`apm install --force`
|
||||
documents this explicitly — "does NOT refresh refs; use 'apm update' for that"). Running it after a
|
||||
merge redeploys the same content and reports success.
|
||||
|
||||
**User scope is untouched, deliberately.** This decision changed project scope only; whatever is
|
||||
natively installed at user scope was left alone, and converting it is a separate decision with a
|
||||
blast radius beyond this repo. The specific inventory this paragraph used to name
|
||||
(`bin@holocron`, `gitea@holocron`, a stale `hello-world@holocron`) is machine state and is stale —
|
||||
see the 2026-08-17 correction above. The decision recorded here is unaffected by what that state is.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **`apm install --global`.** Verified working in an isolated `HOME`: user-scope deploys land in
|
||||
`~/.claude/skills/` and `~/.claude/agents/`, and it is the only scope where a plugin's `bin/`
|
||||
executables deploy (moot here — every `bin/` in this repo is empty but for a README). Deferred,
|
||||
not rejected: it changes skill resolution for every project on the machine at once.
|
||||
- **Keeping both install paths.** Rejected: the same skill would be present twice under two names,
|
||||
and `.claude/settings.json` cannot hold `enabledPlugins` without failing `apm audit --ci`.
|
||||
@@ -0,0 +1,176 @@
|
||||
# A SessionStart hook keeps the apm install current, replacing a git hook that never ran
|
||||
|
||||
ADR-0018 switched this repo to consuming its own plugins through `apm install`, with the six
|
||||
packages declared as unpinned git refs against the holocron remote's default branch. That decision
|
||||
left a hole it named but did not fill: the deployed content goes stale the moment anyone merges,
|
||||
and nothing detects it.
|
||||
|
||||
**Status: accepted (2026-08-14).**
|
||||
|
||||
## Context
|
||||
|
||||
The pre-existing answer was `scripts/git-hooks/post-push`, which pulled the marketplace clone and
|
||||
ran `claude plugin update kyberforge`. Issue #78 filed it as a bug — the hook updated `kyberforge`
|
||||
but not `gitea`, so gitea skills stayed pinned at a pre-refactor version after #67 merged.
|
||||
|
||||
The issue's premise was wrong in a way nobody had noticed for six weeks. **Git has no client-side
|
||||
`post-push` hook.** `githooks(5)` does not list one, and git 2.39.5 does not invoke one.
|
||||
`scripts/install.sh` copies every file in `scripts/git-hooks/` into `.git/hooks/`, so
|
||||
`.git/hooks/post-push` existed on disk and looked installed. It had never fired. The hook did not
|
||||
skip `gitea`; it skipped everything. Both tests that appeared to cover it — `test-post-push.sh` and
|
||||
`test-git-hooks-install.sh` — asserted only that the script behaved correctly when invoked directly
|
||||
and that install.sh copied the file. Neither asserted that git ever runs it.
|
||||
|
||||
That also makes the original framing wrong. Refreshing on push assumes the person who pushes is the
|
||||
person who goes stale, which is backwards: your install goes stale when *someone else* merges, and a
|
||||
push of your own is neither necessary nor sufficient for it to have happened.
|
||||
|
||||
## Decision
|
||||
|
||||
A `SessionStart` hook, shipped in `plugins/kyberforge/.apm/hooks/`, checks whether the install is
|
||||
behind and refreshes it in place.
|
||||
|
||||
`SessionStart` is the correct trigger because the thing that goes stale is the skill content a
|
||||
*session* loads, and that is the moment the staleness does damage. It also enables two things a git
|
||||
hook structurally cannot do: `additionalContext` puts the notice into the agent's context rather
|
||||
than terminal scrollback nobody reads, and `reloadSkills: true` makes the host re-scan the skill
|
||||
directories after the hook returns, so a refresh lands in the running session without a restart.
|
||||
|
||||
apm's own lifecycle events (`pre-/post-install`, `pre-/post-update`, `pre-/post-uninstall`) were
|
||||
rejected: they fire around apm operations already chosen, so they can announce a refresh but never
|
||||
detect that one is needed.
|
||||
|
||||
Three sub-decisions:
|
||||
|
||||
- **Refresh automatically rather than report.** The hook runs `apm update --yes` and asks for a skill
|
||||
reload. The rejected alternative was to report and let a human run it. Auto-refresh costs a
|
||||
rewritten `apm.lock.yaml` — a committed file — appearing as an unexplained modification in the
|
||||
working tree, on any branch, at any time. The emitted notice says so explicitly for that reason.
|
||||
- **`plugins/kyberforge/.apm/hooks/`, not `.claude/settings.json`.** ADR-0018 established that apm
|
||||
owns `.claude/settings.json` and that any repo-authored key in it is permanent `apm audit --ci`
|
||||
drift. A hook shipped in a package is written into that file by apm itself, so it is apm's output
|
||||
and does not drift. `.claude/settings.local.json` also works but is gitignored and machine-local,
|
||||
which fails the requirement that this travel with the repo.
|
||||
- **`startup` matcher only.** `resume`, `clear`, `compact` and `fork` would re-run the check on every
|
||||
compaction, and a compaction is not an event after which the remote can have moved.
|
||||
|
||||
The executable-trust gate is switched on at the same time. Root `apm.yml` gains an `executables:`
|
||||
block allowing kyberforge's hooks and bin.
|
||||
|
||||
## Consequences
|
||||
|
||||
**The gate is off until something turns it on, and this repo had it off.** `apm approve --list`
|
||||
reports `Executable-trust gate disabled -- all executables deploy` until an `executables:` block
|
||||
exists in `apm.yml`. Any hook, bin, or MCP primitive a dependency shipped would have deployed with
|
||||
no prompt and no record. The block added here closes that for this repo; every other apm project on
|
||||
this machine still has it open.
|
||||
|
||||
**The allow key is version-pinned, and that is a live failure mode.** apm writes
|
||||
`kyberforge#1.5.0`, not `kyberforge` — and the release that ships this hook proved the point
|
||||
immediately, since bumping kyberforge to 1.5.0 required editing the key in the same commit. A
|
||||
kyberforge version bump makes the entry stop matching, the
|
||||
gate blocks the hook, and the install silently stops refreshing — the exact failure this ADR exists
|
||||
to end, reintroduced through the mechanism meant to secure it.
|
||||
|
||||
Matching is an exact dictionary lookup on the composed `name#version` string
|
||||
(`apm_cli/security/executables.py`, `is_package_approved`), so there is no wildcard or
|
||||
version-less key that would sidestep this — the key has to be edited on every bump, and the
|
||||
question is only what catches a missed edit. A comment in the `executables:` block is not enough:
|
||||
this repo gates generated-content drift, marketplace mirror drift and vale style drift
|
||||
deterministically, and a silent-staleness failure is strictly worse than any of them. So
|
||||
`scripts/check-executables-allow-sync.sh` runs at pre-push, parsing `version:` out of
|
||||
`plugins/kyberforge/apm.yml` and asserting root `apm.yml` carries the matching
|
||||
`kyberforge#<version>` key. The comment stays as the human-facing pointer; the hook is what
|
||||
actually holds. It parses with PyYAML where importable and falls back to a two-shape scan
|
||||
otherwise, so a missing pip package cannot become the thing that blocks every push.
|
||||
|
||||
**Trust is keyed on the version, not on the content.** `kyberforge#1.5.0` approves whatever
|
||||
`check-apm-current.sh` contains at the moment it is fetched, not the bytes that were reviewed when
|
||||
the key was written. Because the dependency is unpinned against the default branch and the hook
|
||||
runs `apm update --yes` unattended, an edit to that script landing on `main` deploys and executes
|
||||
on every contributor's machine at their next session start, with no second approval prompt and no
|
||||
diff shown. The trust gate constrains *which package* may ship an executable; it does not constrain
|
||||
what that executable does between version bumps. That is an accepted property of this design rather
|
||||
than an oversight — the remote is self-hosted, push access to `main` is already sufficient to
|
||||
change any skill body an agent will follow — but it is the reason the gate should not be read as a
|
||||
supply-chain control. Pinning each dependency to a `ref:` is what would make it one, and ADR-0018
|
||||
defers that until per-package release tags exist.
|
||||
|
||||
**A referenced hook script must be addressed at its `.apm/` path.** apm resolves
|
||||
`${CLAUDE_PLUGIN_ROOT}/...` against the installed package root, and `apm pack` keeps only `*.json`
|
||||
from `.apm/hooks/` when it builds the flat mirror. So `${CLAUDE_PLUGIN_ROOT}/hooks/check-apm-current.sh`
|
||||
resolves to the mirror, where the script does not exist — verified, apm reports
|
||||
`Hook script not found` and deploys a hook pointing at nothing. The working reference is
|
||||
`${CLAUDE_PLUGIN_ROOT}/.apm/hooks/check-apm-current.sh`. The script cannot simply be placed in
|
||||
`plugins/kyberforge/hooks/` either: that directory is `rm -rf`'d by every content sync (ADR-0017).
|
||||
A test pins the reference.
|
||||
|
||||
> **Amendment (2026-09-14) — the mirror half of that reasoning is gone; the conclusion is not.**
|
||||
> ADR-0024 deleted the flat mirror and `scripts/sync-plugin-content.sh`, so neither the `apm pack`
|
||||
> mirror-filtering behaviour nor the `rm -rf` content sync described above still happens. The
|
||||
> reference must stay exactly as written, for the reason that survives independently: `.apm/` is the
|
||||
> sole hand-edited authoring source (ADR-0015), it is what ships in the installed package root, and
|
||||
> `plugins/kyberforge/hooks/` no longer exists at all — so `${CLAUDE_PLUGIN_ROOT}/hooks/...` still
|
||||
> names a path with nothing at it, now because the directory is gone rather than because a sync
|
||||
> emptied it. `tests/test-apm-current-hook.sh` still pins the literal string.
|
||||
|
||||
**Session startup gets slower when the install is stale.** Measured: ~0.7 s for the `apm outdated`
|
||||
check when everything is current, ~10.4 s when six packages are behind and the refresh runs. The
|
||||
hook declares `timeout: 380` to cover a cold multi-package fetch. That number is not free-standing:
|
||||
the script imposes its own `timeout 60` on `apm outdated` and `timeout 300` on `apm update`, so the
|
||||
host-side timeout has to exceed their sum or the host kills the hook mid-update and leaves
|
||||
`.claude/skills/` half-deployed with no notice emitted. An earlier revision declared `320`, which
|
||||
was below the 360 s the script can legitimately take. A test asserts the invariant rather than the
|
||||
literal — it parses every `timeout N` out of the script, sums them, and requires the `hooks.json`
|
||||
value to be larger — so changing either side without the other fails the suite.
|
||||
|
||||
**Reading a human-readable CLI for a control decision cost a silent failure, again.** `apm outdated`
|
||||
has no `--json` or other machine-readable flag (confirmed against 0.28.0), so the hook must match
|
||||
its prose. The first attempt matched `outdated dependencies found` — plural only. apm emits
|
||||
`1 outdated dependency found` in the singular when exactly one package is behind
|
||||
(`apm_cli/commands/outdated.py`), so a single stale package was invisible: the hook exited 0
|
||||
silently and no refresh ran. With six packages merging independently, one-behind is the ordinary
|
||||
case rather than an edge, which means the mechanism failed most often in exactly the situation it
|
||||
exists for. The match is now `outdated dependenc(y|ies) found`.
|
||||
|
||||
The deeper lesson is the one `post-push` already taught and this repeated: every assertion about the
|
||||
hook mocked `apm`, so the suite was green while the hook could not detect the common case. Mocks
|
||||
verify the code against its author's belief about the interface, never the interface. The suite now
|
||||
carries one probe that stages a genuinely outdated dependency against a local git remote — offline,
|
||||
via `url.<path>.insteadOf`, so the twelve-hooks-pass-under-`unshare -rn` property survives — runs
|
||||
the real `apm outdated`, and replays its genuine output through the real hook. Reverting the grep
|
||||
to plural-only fails it.
|
||||
|
||||
**The hook cannot install itself.** Dependencies resolve from the remote, so the hook does not
|
||||
deploy until this change is merged and `apm update` has run once against the new default branch.
|
||||
Until then the repo has the mechanism in source and not in effect.
|
||||
|
||||
**`.claude/settings.json` stops being `{"hooks": {}}`.** apm merges the hook into it and tracks
|
||||
ownership in a `.claude/apm-hooks.json` sidecar, with the script copied to
|
||||
`.claude/hooks/<pkg>/`. The sidecar and the script directory are gitignored install output; the
|
||||
settings file remains committed, now with apm-generated content in it. ADR-0018's statement that the
|
||||
committed content is exactly `{"hooks": {}}` is superseded on that point only — the rule it was
|
||||
protecting, that nothing repo-authored goes in that file, is unchanged.
|
||||
|
||||
**Native consumers are protected by a guard, not by the gate.** A host installing holocron through
|
||||
`claude plugin install` auto-discovers `hooks/hooks.json` and does not consult apm's trust gate at
|
||||
all. The script therefore exits silently when there is no `apm.lock.yaml` in the working directory,
|
||||
which is what makes it inert in a repo that does not consume packages through apm. Copilot CLI sees
|
||||
no hook at all, for the reasons already documented in `plugins/kyberforge/docs/hooks.md`.
|
||||
|
||||
**`scripts/git-hooks/` is now empty.** `post-push` and `test-post-push.sh` are deleted.
|
||||
`install.sh`'s copy block is generic and is kept; `test-git-hooks-install.sh` now synthesizes its
|
||||
own fixture hook instead of depending on a real one existing, so the mechanism stays tested and can
|
||||
be used again if a hook git actually invokes is ever wanted.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **A `post-merge` git hook.** Real, unlike `post-push`, and verified to fire on both a
|
||||
fast-forward `git pull` and a `git pull --rebase`. Rejected as the primary mechanism because a
|
||||
pull is the wrong signal, and because it cannot reload skills in a running session. It remains
|
||||
the only option for a project that consumes apm packages without a Claude-family host.
|
||||
- **Reporting instead of refreshing.** See the sub-decision above.
|
||||
- **A seventh plugin holding only this hook**, to avoid shipping it to external kyberforge
|
||||
consumers. Rejected as disproportionate: the `apm.lock.yaml` guard already makes the hook inert
|
||||
for anyone not consuming through apm, and a package exists to be maintained, versioned, and
|
||||
registered in the marketplace.
|
||||
513
docs/adr/0020-skill-description-and-body-context-contract.md
Normal file
513
docs/adr/0020-skill-description-and-body-context-contract.md
Normal file
@@ -0,0 +1,513 @@
|
||||
# Skills and agents are authored against a context budget, not a spec ceiling
|
||||
|
||||
Every installed skill's `name` and `description` sits in every agent's context from the first token
|
||||
of every session, whether or not the skill is ever invoked. Across this repo's 39 skills that is
|
||||
23,427 characters — roughly 5,900 tokens — and the authoring rules that produced it optimised for
|
||||
triggering reliability with no counter-pressure on size. This ADR sets the budget, the shape, and the
|
||||
gates that hold them.
|
||||
|
||||
**Status: accepted (2026-08-14).**
|
||||
|
||||
## Context
|
||||
|
||||
Every `file:line` citation in this ADR is against the base commit the decision was taken on,
|
||||
`f9b919d7e3bd5e6b51fbdf88b32ace0438b313e0`, not against current `HEAD`. The change that carries this
|
||||
ADR rewrites several of the cited files, so a citation resolved against the worktree will land on
|
||||
unrelated text. Use `git show f9b919d:<path>` to follow one.
|
||||
|
||||
Measured before any change, at that commit. Method, so the figures are reproducible: sum
|
||||
`len(name) + len(description)` over the frontmatter of every `plugins/*/.apm/skills/*/SKILL.md`,
|
||||
folding `>` block scalars to the value the host actually loads (most descriptions here are folded
|
||||
scalars, so counting raw lines measures indentation instead); tokens at the standard
|
||||
~4-characters-per-token approximation `scripts/skill-size-check.sh` uses. Word counts are
|
||||
whitespace-separated tokens, and are stated as **body-only** or **whole-file** every time, never bare.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| 39 skill `name` + `description` | 23,427 chars, ~5,900 tokens, **preloaded every session** |
|
||||
| 4 agent `name` + `description` | 1,325 chars, ~330 tokens, preloaded every session |
|
||||
| skill bodies (body-only words) | median 684, mean 815, p90 1,349 |
|
||||
| skill files (whole-file words) | median 816, mean 927, p90 1,526 |
|
||||
| `MAX_WORDS` gate (`skill-audit/scripts/validate.sh:147`) | **2,770** whole-file — a density proxy, not a percentile |
|
||||
|
||||
That last row is worth stating plainly, because it is the first thing this ADR is about. 2,770 is not
|
||||
derived from the corpus distribution at all: per the derivation comment in
|
||||
`scripts/skill-size-check.sh`, it is 2,770 words at the densest observed 7.22 chars/word ≈ 20,000
|
||||
chars ≈ the agentskills.io ~5,000-token ceiling. Neither percentile reaches it — 2× the body-only p90
|
||||
is 2,698 and 2× the whole-file p90 is 3,052 — and reading it as "2× p90" would pair a whole-file gate
|
||||
against a body-only distribution, which is exactly the conflation this ADR exists to stop.
|
||||
|
||||
Three findings drove this, none of which is "the descriptions drifted".
|
||||
|
||||
**The rules mandate the bloat.** `skill-author/SKILL.md:104` requires indirect triggers ("even if the
|
||||
user doesn't mention X explicitly") and `skill-audit/references/description-quality.md:21` requires
|
||||
authors to "err toward being pushy". Both are enforced. The one rule that would delete the waste —
|
||||
`skill-author/SKILL.md:102`, "not the skill's internal mechanics" — is judgment-only and is absent
|
||||
from the FAIL conditions at `description-quality.md:45-50`. The enforced rules inflate; the deflating
|
||||
rule does not bite. The result is measurable: `gitea-files` spends 147 chars listing six verbs, then
|
||||
301 chars re-quoting the same six as user phrasings, in the same order. `apm-workflow` does the same
|
||||
with six capability clusters. Across the twelve longest descriptions, 30.7% is capability
|
||||
enumeration and 11.6% is composition or implementation detail that cannot affect a routing decision.
|
||||
|
||||
**Capability enumeration in a description is a correctness hazard, not only a token cost.**
|
||||
`plugins/kyberforge/docs/research/examples/skill-write/writing-skills/SKILL.md:154-158` reports a
|
||||
measured failure: "when
|
||||
a description summarizes the skill's workflow, an agent may follow the description instead of reading
|
||||
the full skill content. A description saying 'code review between tasks' caused an agent to do ONE
|
||||
review, even though the skill's flowchart clearly showed TWO reviews." `git-commits` is exactly that
|
||||
shape — 74% of its description is capability enumeration, including a rules table (`header max 100
|
||||
chars, lowercase subject, no trailing periods, 11 standard types`) an agent can act on without ever
|
||||
loading the body.
|
||||
|
||||
**The upstream sources cannot settle this.** The four skill-writing references under
|
||||
`plugins/kyberforge/docs/research/examples/skill-write/` disagree on what a description contains —
|
||||
when-only (`writing-skills/SKILL.md:99`), what-and-when (`skill-creator/SKILL.md:67`,
|
||||
`writing-skills/anthropic-best-practices.md:187`), triggers-only
|
||||
(`writing-great-skills/SKILL.md:28`), and what-plus-when-plus-negative
|
||||
(`write-skill/SKILL-TEMPLATE.md:5-6`). Those four paths are relative to that directory.
|
||||
`writing-skills` and the Anthropic document it bundles contradict each other inside one skill
|
||||
directory. They also disagree on whether
|
||||
500 lines is binding, on the inline-versus-bundle threshold, and on the TOC threshold (>100 lines vs
|
||||
>300 lines). "Grounded in the research" is therefore not available as a tiebreaker; a house choice is
|
||||
required and this is it.
|
||||
|
||||
A fourth observation shaped the body half. The best progressive-disclosure ratio in the repo belongs
|
||||
to `apm-workflow` — a 421-word body dispatching to 3,006 words of references — and the worst two
|
||||
belong to the skills that define the house standard: `skill-author` (2,623-word body / 1,247 words of
|
||||
references) and `agent-author` (2,582 / 1,664). Measured the other way, whole-file, those two are
|
||||
2,760 and 2,758 words — ten and twelve words under the 2,770 gate their own plugin enforces. A
|
||||
ceiling that nothing approaches is not a constraint; a ceiling that two files have grown into is a
|
||||
target. The two numbers for one file are the point: 2,623 and 2,760 describe the same `skill-author`,
|
||||
and only one of them is what either gate measures.
|
||||
|
||||
## Decision
|
||||
|
||||
### Descriptions
|
||||
|
||||
A description carries three things and nothing else: a **trigger clause**, at most one **capability
|
||||
clause**, and a **boundary clause**. Capability enumeration, output-format detail, composition notes
|
||||
("composes X rather than duplicating Y"), and implementation detail move to the body or to
|
||||
`README.md`.
|
||||
|
||||
- **250 characters SUGGESTION, 400 FAIL.** The agentskills.io 1,024-character limit remains as an
|
||||
unchanged spec backstop. The SUGGESTION tier is what moves the average; the FAIL tier only stops
|
||||
outliers.
|
||||
- **A missing, valueless or `null` `description:` is a hard FAIL** in all three validators. That
|
||||
reads as a trivial precondition and is not: a `description:` line with no value followed by
|
||||
`model: sonnet` let a line regex capture the *next* key, which looked non-empty, so the "missing or
|
||||
empty" branch never fired and every gate below it then early-returned on the genuinely empty folded
|
||||
value — exit 0, zero output, on a blocking pre-push gate. Presence is decided on the YAML-folded
|
||||
value and nowhere else. The field this contract is entirely about is the one field a gate must
|
||||
never fail to notice is absent.
|
||||
- **Boundary clauses compress** to `Not <thing> → <skill-name>.` and must name a target that
|
||||
resolves to a real skill or agent. Resolution walks up **from the file being checked** to an
|
||||
*authoring root* — the nearest ancestor holding `plugins/*/.apm/skills` or `plugins/*/.apm/agents`,
|
||||
falling back to the nearest ancestor holding `.git`. Two passes rather than one interleaved walk,
|
||||
so a nested `.git` (a submodule, a sub-package worktree) cannot beat a real monorepo root further
|
||||
up. When an authoring root is found the universe is every skill and agent under
|
||||
`<root>/plugins/*/`, plus the target's own apm package and the packages that package declares in
|
||||
its own `apm.yml` `dependencies.apm`. Sibling plugins resolve against each other, which is what a
|
||||
monorepo means. Deployed `.claude/`/`.agents/` trees are consulted **only** when the walk found no
|
||||
plugin monorepo root — whether it landed on a bare `.git` ancestor or on nothing at all. That is
|
||||
the consumer case, where there is no monorepo to read. The condition is which of the two passes
|
||||
matched, never a name-count delta: a single-plugin monorepo re-collects its own package and adds
|
||||
no new name, so a delta test reads zero there and would pull the deployed trees back in. What the
|
||||
resolver must never do is
|
||||
derive the universe from its own location: a `${BASH_SOURCE}`-relative repo root leaked this repo's
|
||||
39-skill universe into every consumer repo running the hook through pre-commit, so a consumer skill
|
||||
routing to `skill-audit` resolved against a plugin it had never installed. Checked
|
||||
deterministically. A description carrying **no** boundary clause at all is a SUGGESTION, for skills
|
||||
and agents alike: most descriptions want one, some genuinely have no near-miss sibling to exclude,
|
||||
and that judgment is not a script's to make.
|
||||
- **The verdict must not depend on whether `apm install` has been run.** Deployed trees are
|
||||
gitignored install output, present only on a machine that has run it. Four cross-plugin targets
|
||||
here (`gitea-branches` → `git-branches`, `gitea-branches` → `git-history`, `gitea-issues` →
|
||||
`git-branches`, `gitea-workflow` → `git-workflow`) once resolved through `.claude/skills/` alone,
|
||||
so the same commit measured 2 dangling targets on a developer machine and 6 on a fresh clone. A
|
||||
gate shipping hot with no baseline cannot give two answers. Under the walk-up those four resolve
|
||||
because sibling plugins are in the universe — no plugin here declares a cross-plugin apm
|
||||
dependency, and none needs to. Verified: a tree holding only `plugins/` and the root `apm.yml`,
|
||||
with no `.claude/` or `.agents/` anywhere, produced findings identical to the working tree. The
|
||||
figures that reproduction recorded — 26 description FAILs, 9 body FAILs, 2 dangling targets, 0
|
||||
missing references, 58 SUGGESTIONs — are the **pre-retrofit** corpus as it stood when the
|
||||
experiment ran, kept here as the evidence for the install-independence claim, not as a current
|
||||
reading. *Amended 2026-09-01: the #99 retrofit took the first three to zero. Measured at that
|
||||
date over the same install-free tree: 0 description FAILs, 0 body FAILs, 0 dangling targets, 0
|
||||
missing references, 29 SUGGESTIONs.* What the experiment establishes is that the two trees agree,
|
||||
not what either measured; re-derive rather than quote —
|
||||
`bash scripts/skill-size-check.sh plugins/*/.apm/skills/*/SKILL.md`.
|
||||
- **The universe is the apm marketplace, and nothing else.** A routing target resolves to a skill or
|
||||
an agent, or it does not resolve. Host built-ins are deliberately outside it: `/compact`, `/clear`
|
||||
and `/init` are Claude Code slash commands with no counterpart in Copilot CLI or Codex, so a
|
||||
vendor-neutral `.apm/` description routing to one is a portability defect and the hard FAIL is a
|
||||
true positive, not a false one. An allowlist of known built-ins was **rejected**: it answers a
|
||||
different question ("does this exist on *some* host?"), it cannot answer that portably from a
|
||||
single source file, and it goes stale the next time a host ships a command — reintroducing the
|
||||
same-commit-two-verdicts failure the bullet above exists to close. An author who needs to mention
|
||||
one writes it un-slashed (``the `compact` built-in``), which is not route notation and makes no
|
||||
routing claim.
|
||||
- **Blocking is scoped to a sentence, which makes sentence boundaries load-bearing.** A prose-form
|
||||
target earns a hard error only when its own sentence names another target that *resolves*; route
|
||||
notation (`/name`, `→ name`) is exempt and always blocks. So the splitter is part of the contract,
|
||||
not a detail of it. `e.g. "…"` is not a sentence end, and a sentence opening with a code span or a
|
||||
lowercase skill name is a start; getting either wrong moves targets between the two tiers in
|
||||
opposite directions — a stranded corroborator silently demotes a real finding to SUGGESTION, and a
|
||||
missed boundary lets one sentence vouch for a target it never stood beside, producing a hard FAIL
|
||||
with no escape hatch.
|
||||
- **The blanket pushiness rules are deleted.** `skill-author/SKILL.md:104` and
|
||||
`description-quality.md:21` are replaced by a conditional: add an indirect trigger only where the
|
||||
user's natural phrasing genuinely omits the domain word — true for the `gitea-*` family, false for
|
||||
`git-commits`. Stating the same trigger twice in two registers is a FAIL.
|
||||
|
||||
### Bodies
|
||||
|
||||
The body carries the **decision procedure only**: ordered steps, decision branches, gates, and which
|
||||
reference to load when. Lookup tables, spec restatements, output schemas, templates, and rationale
|
||||
prose move to `references/` behind an explicit "read X when Y" trigger.
|
||||
|
||||
- **600 words SUGGESTION, 900 FAIL, counted body-only** — everything after the closing `---` of the
|
||||
frontmatter. The 2,770-word / 500-line spec backstop is unchanged, keeps its existing meaning
|
||||
(conformance, not quality), and keeps counting the **whole file including frontmatter**. These are
|
||||
two different gates measuring two different things, and conflating them is what produced the
|
||||
current state.
|
||||
- **Dispatch is mandatory at two or more mutually exclusive flows.** The body carries the dispatch
|
||||
table and the gates that apply to every branch; each flow lives in its own self-contained
|
||||
`references/` file. This is `apm-workflow/SKILL.md:33-41` promoted from accident to rule. "Two
|
||||
mutually exclusive flows" is not decidable from file text, so this rule is auditor judgment — see
|
||||
Enforcement below for what that means and does not mean.
|
||||
- **Every `references/<file>.md` a body names must exist.** A dispatch table pointing at a file that
|
||||
was never written is a silently dead branch. Checked deterministically.
|
||||
- **Gotchas are constrained.** A Gotcha must state a fact that contradicts a reasonable default —
|
||||
something the agent gets wrong by acting sensibly. More than five entries is a SUGGESTION, as is a
|
||||
Gotchas section exceeding 25% of the body; both are countable and both are checked
|
||||
deterministically. A Gotcha that paraphrases a step in the body below it is a FAIL, but a FAIL an
|
||||
auditor issues, not a script — semantic equivalence is not pattern-matchable.
|
||||
|
||||
### Agents
|
||||
|
||||
Agents take the same description gates — they are preloaded identically — and **no body word gate**.
|
||||
A skill body is loaded into the caller's context, competing with the live conversation; an agent body
|
||||
becomes the system prompt of a fresh context. The rationale for the 900-word FAIL does not transfer.
|
||||
|
||||
That exemption is expressed in `agent-audit/scripts/validate.sh`, which has no body constant, and in
|
||||
the `files:` pattern of the `skill-size-check` pre-commit hook, which is `SKILL.md`-only. It is *not*
|
||||
expressed in `scripts/skill-size-check.sh` itself, which measures whatever path it is handed —
|
||||
running it directly over `plugins/*/.apm/agents/*.agent.md` exits 1 with 900-word body FAILs on
|
||||
`git-orchestrate` and `gitea-orchestrate`. *Amended 2026-09-01: this sentence named a third agent,
|
||||
`apm-orchestrate`, at 1,080 words. It is 876 today — a SUGGESTION, not a FAIL. Counts are
|
||||
deliberately no longer pinned here: agent bodies are edited like any other file and a figure in this
|
||||
paragraph goes stale the moment one is trimmed. Run the command.* Agents escape by
|
||||
file pattern, not by the script knowing the difference. Anyone widening that pattern to cover agents
|
||||
would silently enforce a gate this ADR declines to set.
|
||||
|
||||
A plugin-scope agent is a single file with no sibling `references/` directory, so it cannot disclose
|
||||
to itself — it can only delegate to skills. `agent-audit` therefore gains a **delegation check**: an
|
||||
agent body that restates a procedure owned by a skill it can invoke is a FAIL, with the fix being
|
||||
"invoke `<skill>` instead". Length falls out of delegation rather than being gated directly.
|
||||
|
||||
### Invocation as a design axis
|
||||
|
||||
`skill-author` asks whether a skill is model-invoked or hand-invoked before writing a description. A
|
||||
hand-invoked skill sets `disable-model-invocation: true` and carries one plain human-facing sentence
|
||||
with no trigger list.
|
||||
|
||||
Verified end-to-end rather than assumed: `plugins/bin/.apm/skills/zoom-out/SKILL.md:4` carries the
|
||||
flag, apm passes it through verbatim to both `.claude/skills/zoom-out/SKILL.md:4` and the flat mirror
|
||||
at `plugins/bin/skills/zoom-out/SKILL.md:4`, and `zoom-out` was — at the time of that check, when it
|
||||
was the only carrier — the one installed skill absent from the model-visible skill listing in a live
|
||||
session. It remains invocable as `/zoom-out`. `caveman` has since taken the flag as well, so the
|
||||
corpus now has **two** carriers. Do not read a carrier list off this page; re-derive it:
|
||||
|
||||
```
|
||||
grep -l '^disable-model-invocation: true' plugins/*/.apm/skills/*/SKILL.md
|
||||
```
|
||||
|
||||
### Merging siblings
|
||||
|
||||
Two skills that share substantial content, name each other as near-misses, and differ only in the
|
||||
type of input they take should be **one skill with a dispatch table**. This catches `skill-audit` +
|
||||
`agent-audit` and is scoped to them; the author pair is explicitly excluded, because
|
||||
`skill-author` and `agent-author` emit genuinely different artifacts (a skill directory versus a
|
||||
one-or-two-file agent pair, per ADR-0005 and ADR-0016) and their overlap is in the improve flow
|
||||
rather than the core job.
|
||||
|
||||
**DEFERRED — not implemented in the change that carries this ADR. Tracked as issue #101.** Both
|
||||
skills still exist separately, and this change made the split deeper rather than shallower: retrofit
|
||||
to the dispatch pattern took `skill-audit` from 3 reference files to 7 and `agent-audit` from 4 to 8,
|
||||
and their two same-named `references/description-quality.md` files now differ on 100 of ~120 lines
|
||||
after normalising `skill`/`agent`, where before they were closer. It has kept deepening since: the
|
||||
#99 retrofit added `finding-criteria.md` to `skill-audit`, drawing it level with `agent-audit`. Both
|
||||
figures move with the next retrofit, so measure rather than quote —
|
||||
`ls plugins/kyberforge/.apm/skills/<name>/references/ | grep -c '\.md$'`. The merge stays the
|
||||
decision; it reopens ADR-0008 (agent-audit's single-file invocation contract) and touches every call
|
||||
site in `skill-author`, `agent-author` and `forge`, which is why it is its own change and not a rider
|
||||
on this one. Recorded here rather than dropped, so the gap between the rule and the tree is deliberate
|
||||
and dated instead of discovered later.
|
||||
|
||||
### Enforcement and rollout
|
||||
|
||||
Gates land where the existing gates already live — no new layer. The table below is exhaustive about
|
||||
which tier each rule is in, because the failure this ADR is most exposed to is a rule filed under
|
||||
"Enforcement" that no validator implements:
|
||||
|
||||
| Check | Applies to | Tier | Home |
|
||||
|---|---|---|---|
|
||||
| description characters (250 SUGGESTION † / 400 FAIL) | skills, agents | deterministic | `scripts/skill-size-check.sh`; constants mirrored in `skill-audit/scripts/validate.sh` and `agent-audit/scripts/validate.sh` |
|
||||
| body-only words (600 SUGGESTION / 900 FAIL) | skills | deterministic | `skill-size-check.sh`, `skill-audit/scripts/validate.sh` |
|
||||
| description present and non-empty (ERROR) | skills, agents | deterministic | same |
|
||||
| boundary target resolves to a real skill or agent — **three** verdicts, not two (ERROR when written in route notation — `/name`, or any arrow form; or when a *terminal* bare name's own sentence names another target that resolves. SUGGESTION otherwise. INFO, "DID NOT RUN", exit 0, when no skill universe could be determined for the path at all — no authoring root above it, no apm package root, no declared apm dependencies, no deployed `.claude/` or `.agents/` tree: the targets are named and left unchecked) | skills, agents | deterministic | same |
|
||||
| boundary clause absent — `absent` (SUGGESTION) † | skills, agents | deterministic | same |
|
||||
| an arrow clause is present but no target can be read out of it — `unparsed` (SUGGESTION) † | skills, agents | deterministic | same |
|
||||
| one arrow clause naming two or more targets, of which only the first is resolved (SUGGESTION, issue #107) † | skills, agents | deterministic | same |
|
||||
| Gotchas entry count over five (SUGGESTION) | skills | deterministic | same |
|
||||
| Gotchas over 25% of the body (SUGGESTION) | skills | deterministic | same |
|
||||
| every `references/<file>.md` a body names exists (ERROR) | skills | deterministic | same |
|
||||
| description opener, composition notes in a description | skills, agents | prose pattern | `plugins/kyberforge/.apm/skills/*/assets/vale/styles/Kyberforge/` |
|
||||
| a Gotcha paraphrasing a body step | skills | **auditor judgment** | `references/body-discipline.md` |
|
||||
| dispatch at two or more mutually exclusive flows | skills | **auditor judgment** | `references/body-discipline.md` |
|
||||
| delegation: an agent body restating a skill's procedure | agents | **auditor judgment** | `agent-audit` |
|
||||
| capability enumeration, restatement, trigger quality | skills, agents | **auditor judgment** | `references/description-quality.md` |
|
||||
|
||||
The rows in bold are stated as FAILs in the Decision above and are FAILs an *auditor* issues. None of
|
||||
them is countable: "does this Gotcha paraphrase step 4", "are these two flows mutually exclusive" and
|
||||
"does this agent body restate what `git-commits` already owns" are semantic questions, and a script
|
||||
that guessed at them would be a worse gate than no gate, because it would be believed. They are not
|
||||
enforced, they are reviewed, and this table exists so that distinction is written down rather than
|
||||
inferred from whether a validator happens to have been written yet.
|
||||
|
||||
**† These four, and only these four, are lifted for a hand-invoked file** — one whose frontmatter
|
||||
carries `disable-model-invocation: true`, read as a boolean by `hand_invoked()` in all three scripts.
|
||||
No validator knew the field existed (issue **#108**), so every routing SUGGESTION above fired on
|
||||
exactly the shape the *Invocation as a design axis* section mandates, and the boundary-clause
|
||||
remedy — "so the router knows where NOT to send this skill" — was addressed to a router that cannot
|
||||
see the skill at all. An author who took the advice made the file worse.
|
||||
|
||||
What does **not** lift is the point of the carve-out. Both body word tiers stand: the body is still
|
||||
loaded on invocation and still competes with the caller's live conversation. The 400-character
|
||||
description FAIL stands: that description is not preloaded, but it is the one line a user reads when
|
||||
choosing from the `/` menu, and the ceiling is an outlier stop rather than a routing-quality budget —
|
||||
which is exactly why the 250-character *target* is the tier that lifts. And a target the description
|
||||
does happen to name is still resolved and can still dangle as a blocking ERROR. Mechanics, and the
|
||||
reason the field is read as a boolean rather than as a mention of the key: `docs/spec/gates.md`.
|
||||
|
||||
Two of the deterministic rows are tuned for **false positives over recall**, and what they decline to
|
||||
see is part of the contract. On target extraction: a bare hyphenated name counts only inside a
|
||||
boundary sentence, and a single-word name is never matchable bare — `research`, `triage`, `forge`,
|
||||
`prototype` and `tdd` are all real skill names *and* ordinary English, so it must be written
|
||||
`` `forge` `` or `/forge` to be seen at all. Grammar then decides whether a recognised target may
|
||||
raise an error: one followed by an ordinary lowercase noun is a compound **modifier**, not a route
|
||||
("use pre-commit hooks instead of ad-hoc scripts", "invoke the pull-request template"), so it is
|
||||
confirm-only — it still resolves and still counts as a route when the name exists, but it can never
|
||||
dangle. Only a *terminal* target can. The compressed arrow form `→ <name>` is exempt from that
|
||||
follower test and is always error-eligible, because nothing reads as a compound modifier after an
|
||||
arrow; a `/slash` target reached through a route verb is **not** exempt and takes the same test.
|
||||
*Amended 2026-08-31 — the `/slash` half is reversed: it is exempt too. See the amendment below.* The
|
||||
simpler rule — "only marked targets may dangle" — was available and would have been wrong here: both
|
||||
live true positives are bare, `research`'s "(use neuledge-context)" and the `gitea-labels-` /
|
||||
`milestones` fold. On the body-shape checks: a `## Gotchas` heading must *end* in "gotchas", not
|
||||
merely contain the word, so `## Gotcha handling` and `## Why gotchas matter` are prose sections and
|
||||
are skipped; fenced code blocks are masked out of heading detection and entry counting, so a fenced
|
||||
example list is not mistaken for the section; and a `references/` pointer named on a line
|
||||
that also says the file is gone ("removed", "deprecated", "no longer") is read as a historical
|
||||
mention rather than a dead dispatch entry. Note the 25% fraction is deliberately *not* fence-masked
|
||||
on either side — fenced lines are real body words, and the fraction is measured against the whole
|
||||
body.
|
||||
|
||||
**The deterministic tier blocks immediately, with no baseline file.**
|
||||
|
||||
Three pre-existing contradictions are fixed in the same change, because they are the contract:
|
||||
|
||||
- `skill-audit/SKILL.md:58` asks whether the description opens with an action verb ("Audits…",
|
||||
"Reviews…"), while `:56` defers the same question to `Kyberforge.DescriptionOpener` and
|
||||
`skill-author/SKILL.md:101` requires an imperative "Use when…" opener. The criterion is
|
||||
unsatisfiable against the house's own skills, both of which open with "Use when".
|
||||
- `DescriptionOpener.yml` is anchored to `^This (skill|agent)\b`, which misses a plain `This …`
|
||||
opener; it is widened here to `^This\b`. The anchor itself stays. Composition prose that sits
|
||||
*mid*-description — `gitea-workflow`'s "This is the human-facing entry point…" at character 377,
|
||||
`gitea-labels-milestones`'s "This is a cross-cutting shared skill…" at character 300 — was never in
|
||||
the opener rule's scope and correctly is not: under `scope: text.frontmatter.description` the `^`
|
||||
anchors to the start of the whole folded value, and un-anchoring to reach mid-description text was
|
||||
measured at 5 hits and 5 false positives and rejected (`LESSONS.md`, 2026-08-14). The real gap is
|
||||
that no rule covered that text at all, which a new token-list rule, `Kyberforge.CompositionNote`,
|
||||
closes: 10 alerts across four `gitea-*` skills, 0 false positives.
|
||||
- `description-quality.md:45-50` has no FAIL condition for internal-mechanics content, which is why
|
||||
`skill-author/SKILL.md:102` never bit.
|
||||
|
||||
## Amendment (2026-08-31): route notation short-circuits the follower test, `/name` included
|
||||
|
||||
The Enforcement section above exempts the arrow form from the follower test and then withholds the
|
||||
same exemption from `/name`: "a `/slash` target reached through a route verb is **not** exempt and
|
||||
takes the same test." That half is reversed. **Both spellings of route notation are exempt, and the
|
||||
exemption is decided before the follower test rather than weighed against it.**
|
||||
|
||||
Three things make the original call wrong rather than merely strict.
|
||||
|
||||
**It contradicted the promise the same paragraph makes.** Route notation is offered to an author as
|
||||
the way to get a target checked unconditionally — the SUGGESTION text on an unpromoted target says
|
||||
so in as many words: "write it as `/name` or `-> name` and it will be checked properly." Under the
|
||||
original rule that was true of one of the two spellings. `-> name` reached `_add()` with
|
||||
`strict=True` from both its call sites; `/name` did not, so it fell through to `_terminal()` and any
|
||||
follower outside `FOLLOWER_OK` demoted it. `Do not use for Y — use /no-such-skill afterwards.` exited
|
||||
0 — and, before the companion visibility fix, in total silence.
|
||||
|
||||
**The follower test's own justification does not reach `/name`.** That test exists for *prose*: a
|
||||
bare hyphenated token followed by an ordinary lowercase noun is a compound modifier, "pre-commit
|
||||
hooks" and "pull-request template". A leading slash is Claude Code's invocation syntax and occurs in
|
||||
no English compound, so there is no attributive reading to protect. The exemption was withheld from
|
||||
the one shape the rule it protects against cannot describe.
|
||||
|
||||
**`FOLLOWER_OK` is a closed whitelist of roughly eighty words, and a closed list is the wrong thing
|
||||
to hang a blocking gate on.** Leaving `/name` under it made *whether a commit is blocked* depend on
|
||||
whether someone had thought to enumerate the next word — the gate failing open on its own
|
||||
unfamiliarity. The bare-target path keeps the follower test precisely because it needs a brake it can
|
||||
justify; the notation path asked for one and was given the same brake by accident.
|
||||
|
||||
What is unchanged: the **corroboration** branch. A *bare* terminal name still earns its blocking
|
||||
ERROR only from a resolving sibling in the same sentence, and a compound modifier still cannot
|
||||
dangle at all. The conservative tuning that decision rests on is untouched — this amendment moves one
|
||||
explicitly-marked spelling out from under it, not the prose path.
|
||||
|
||||
Verified on fixtures inside a synthetic plugin tree: `… Do not use for Y — use /no-such-skill
|
||||
afterwards.` exits 1, while the same sentence with the bare `no-such-skill` exits 0 at SUGGESTION,
|
||||
and rises to a blocking ERROR the moment a resolving sibling joins it. The reasoning is recorded at
|
||||
the point of enforcement in `_add()`'s docstring in `scripts/skill-size-check.sh` and its two
|
||||
mirrored copies, and the verdict table in `docs/spec/gates.md` states the corrected shape.
|
||||
|
||||
## Consequences
|
||||
|
||||
**Editing any non-compliant skill now requires retrofitting it first.** At decision time, 30 of 39
|
||||
descriptions exceeded 400 characters and 13 of 39 bodies exceeded 900 words — the latter counted
|
||||
body-only, which is what the new gate measures; the pre-existing 2,770-word gate counts the whole
|
||||
file including frontmatter, and the two must not be conflated. The change that carries this ADR also
|
||||
retrofits kyberforge's own four author/audit skills, so the figures on landing are **26 and 9**.
|
||||
With the gate hot and no baseline, a one-line
|
||||
fix to `gitea-prs` cannot be committed until that skill meets the contract. This is deliberate — it
|
||||
guarantees convergence and avoids a half-state — but it means the retrofit is lazy and *mandatory*
|
||||
rather than deferred. Issue #99 tracks it and should be prioritised accordingly, and the risk it
|
||||
carries is the ordinary one for hot gates: a gate expensive enough to be inconvenient gets bypassed
|
||||
with `SKIP=` and loses its authority.
|
||||
|
||||
**A second hot gate ships alongside it, and it is easy to miss.** `Kyberforge.CompositionNote` is
|
||||
`level: error` like every other rule in that style, so at decision time `pre-commit run --all-files`
|
||||
was red on 10 alerts across `gitea-issues`, `gitea-labels-milestones`, `gitea-prs` and
|
||||
`gitea-workflow` independently of anything `skill-size-check` reports. Someone scoping the #99
|
||||
retrofit off the size findings alone would have fixed those and still been blocked. The two gates
|
||||
wanted fixing together, and were. *Amended 2026-09-01: that figure is historical. The Vale prefilter
|
||||
over the same 39 files now reports 0 errors, 0 warnings and 0 suggestions, so
|
||||
`Kyberforge.CompositionNote` fires nowhere in the corpus today. The rule is still hot and still
|
||||
independent of `skill-size-check`, so a new description can reintroduce it; `skill-size-check` does
|
||||
not cover the Vale half, and no `references/` file is linted by anything (`docs/spec/gates.md` has
|
||||
both causes, issue #117 tracks them). Re-derive rather than quote —*
|
||||
`bash plugins/kyberforge/.apm/skills/skill-audit/scripts/vale-wrap.sh plugins/*/.apm/skills/*/SKILL.md`.
|
||||
|
||||
**A ceiling does not produce an average.** If every author writes to the 400-character FAIL, the
|
||||
preload lands at 39 × 400 = 15,600 chars — a 33% cut off 23,427, not the ~50% intended. Writing to
|
||||
the 250-character SUGGESTION instead lands at 9,750, a 58% cut. The halving depends entirely on the
|
||||
250-character SUGGESTION tier being visible and respected. That tier works here in a way it does not
|
||||
elsewhere in this repo: `skill-audit` already reports `PASS (N suggestions)` as a first-class
|
||||
outcome. This is explicitly **not** the failure ADR-0013 records — Vale warnings are invisible
|
||||
because vale's exit code keys on `error` alone, but these gates live in `validate.sh` and
|
||||
`skill-audit`, where a SUGGESTION reaches the report. Realistic landing is somewhere in that 33-58%
|
||||
band, not a guaranteed 50%.
|
||||
|
||||
**A word gate cannot detect the defect it is standing in for.** `git-commits` carries twelve Gotchas
|
||||
of which four restate steps in its own Workflow (`:32` ≡ step 9, `:33` ≡ step 9, `:36` ≡ step 2,
|
||||
`:31` ≡ the description). Its body is 1,102 words and its whole file 1,217, so it does fail the
|
||||
900-word body FAIL — but for its length, not for the restatement. The four duplicated Gotchas are 114
|
||||
words between them; delete every one and the file still fails, while a skill 250 words shorter with
|
||||
the identical defect passes clean. The two properties are uncorrelated, which is why the counts are a
|
||||
backstop to the dispatch rule and the Gotchas constraint — both of which are auditor judgment for the
|
||||
semantic half, per the Enforcement table — and not a substitute for them. Reading the word gate as
|
||||
the mechanism is the specific mistake this paragraph exists to prevent.
|
||||
|
||||
**Some skills legitimately need more description budget than others.** A tiered limit keyed to
|
||||
sibling density was considered and rejected as too clever; the flat 250/400 pair means the `gitea-*`
|
||||
and `git-*` families — where every sibling shares a keyword and boundary clauses do real routing work
|
||||
— are the ones most likely to sit at the FAIL tier permanently. If the retrofit shows that family
|
||||
routing degrades, the tier is the first thing to revisit.
|
||||
|
||||
**Four broken routing targets were found; two were fixed here and two shortly after.** Tracked as
|
||||
issue #100.
|
||||
|
||||
- `skill-audit` routed to `/skill-improve` twice in its description plus `README.md:10`, and no such
|
||||
skill exists — the real target is `skill-author`. **Fixed here**, as a side effect of retrofitting
|
||||
kyberforge's own skills.
|
||||
- `agent-author` said "Do not use for read-only review — examine agent files manually", routing away
|
||||
from `agent-audit`, the correct sibling. **Fixed here**, same way. Note this one was never
|
||||
detectable by the resolvable-target check and never will be: "examine agent files manually" names
|
||||
no target, and a check that resolves names cannot see a name that is absent. A misroute to nowhere
|
||||
is a review finding, not a gate finding.
|
||||
- `research` routes to `neuledge-context`, which exists only inside that string. Was **live**;
|
||||
**fixed under #99** — the retrofitted description names no such target.
|
||||
- `gitea-issues` carries the literal string `gitea-labels- milestones` in its folded description, a
|
||||
stray space introduced by YAML wrapping mid-token, breaking the skill name in preloaded text. Was
|
||||
**live**, reported as a dangling `gitea-labels`; **fixed under #99** — the name now folds intact.
|
||||
|
||||
So the check fired on 3 of the 4 against the base commit and on 2 at the tip of the change that
|
||||
carried this ADR. **The corpus dangling set is now empty**, and that is asserted rather than
|
||||
observed: `tests/test-adr0020-targets.sh` pins the set as empty, so a new boundary clause naming a
|
||||
non-existent skill fails the suite instead of joining a backlog. `tests/test-skill-size-check.sh`
|
||||
probed the three original names rather than asserting a count; as each was retrofitted its probe was
|
||||
**removed, not skipped**, because a `pass "SKIP: …"` branch is an assertion-free result counted in
|
||||
the totals and makes the suite look one test stronger than it is. That file's commentary survives the
|
||||
probes and states the rule. Re-derive the current set — never read it off this page:
|
||||
|
||||
```
|
||||
bash scripts/skill-size-check.sh plugins/*/.apm/skills/*/SKILL.md | grep 'does not resolve'
|
||||
```
|
||||
|
||||
**Duplication between `skill-author` and `agent-author` survives un-gated.** The merge rule
|
||||
deliberately excludes the author pair, so the commit-verification argument in four near-copies, the
|
||||
root-cause grouping rule in four copies, and the wholesale clone of the "Improving an existing X"
|
||||
flow all remain. Cache isolation makes them structurally unavoidable
|
||||
(`skill-audit/SKILL.md:95` forbids cross-skill references; `LESSONS.md:107` records why), so the
|
||||
options are a sync gate or continued drift. This is an input to issue #101, which carries both halves
|
||||
of the kyberforge duplication problem — the deferred audit-pair merge and this — not a solved
|
||||
problem.
|
||||
|
||||
**Provenance frontmatter is explicitly out of scope.** `LESSONS.md:63` asserts that non-routing
|
||||
frontmatter (`source_keys`, `category`, `version`) is loaded at agent startup, which would make the
|
||||
7,352 characters of it across the corpus a third again on top of the description tax. Measured
|
||||
against a live session on this Claude Code version, it is not: the model-visible skill listing
|
||||
contains only `name` and `description`. That is host-observed rather than spec-guaranteed and says
|
||||
nothing about Copilot CLI, but it is sufficient to establish that cutting `source_keys` would break
|
||||
the ADR-0009 provenance machinery for no runtime gain. The metadata was added deliberately and stays.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
Upstream citations below are relative to
|
||||
`plugins/kyberforge/docs/research/examples/skill-write/`, as in Context above.
|
||||
|
||||
- **Keep pushiness, raise the budget to ~500 chars.** Undertriggering is the worse failure mode — a
|
||||
skill that never fires is worth nothing regardless of cost — and `skill-creator/SKILL.md:67`
|
||||
explicitly recommends being "pushy" against an observed undertriggering tendency. Rejected because
|
||||
that claim is an unmeasured assertion about an older model, and because the correctness hazard in
|
||||
`writing-skills/SKILL.md:154-158` cuts the other way: a fat description is not merely expensive, it is a
|
||||
shortcut agents take instead of reading the body. Would have landed a 35% cut.
|
||||
- **A trigger-eval loop to set lengths empirically.** `skill-creator/SKILL.md:337-404` specifies 20
|
||||
queries per skill, 8-10 positive and 8-10 near-miss, with a 60/40 train/test split selecting on
|
||||
test score. This is the rigorous answer and the repo has deliberately never built it. Rejected
|
||||
because it blocks the context cut behind a substantial new subsystem.
|
||||
- **A repo-level aggregate preload budget** (≤12,000 chars across all skills, checked at pre-push).
|
||||
The only option that measures the actual goal rather than a proxy. Rejected because it makes one
|
||||
skill's edit fail on account of another skill's growth, and because it is meaningless for an
|
||||
external consumer installing a subset of the plugins.
|
||||
- **500-word body FAIL, matching `writing-skills/SKILL.md:217-221`.** Best-grounded in upstream and
|
||||
would align this repo with the tightest source. Rejected because it fails 28 of 39 skills body-only
|
||||
(35 of 39 measured whole-file), and a blunt gate gets satisfied by deleting content rather than
|
||||
relocating it.
|
||||
- **A shrinking baseline file** recording each non-compliant skill's current numbers, failing only on
|
||||
growth. Would have made the retrofit a visible burn-down instead of a wall. Rejected in favour of
|
||||
hot gates.
|
||||
- **A sync gate over the duplicated spans** instead of a merge rule — generalising
|
||||
`scripts/check-vale-style-sync.sh` to cover shared prose so duplication persists but drift cannot.
|
||||
Rejected for the audit pair in favour of merging, which removes the duplication rather than
|
||||
policing it, and removes a mutually-excluding near-miss pair from the router at the same time. It
|
||||
remains the only available answer for the author pair.
|
||||
- **Merging `skill-author` + `agent-author` as well**, taking kyberforge from seven skills to five.
|
||||
Largest cut available. Rejected because it reopens ADR-0005, ADR-0008 and ADR-0016 together, and a
|
||||
merged author skill would carry both the skill-directory scaffold and the dual-provider agent
|
||||
scaffold behind one dispatch.
|
||||
- **Demoting Gotchas** to the end of the body or into `references/gotchas.md`, removing its
|
||||
position-based exemption from the dispatch rule. Maximum saving on the largest body construct
|
||||
(6,830 words, 21% of all body text). Rejected because a gotcha read after the mistake is worthless.
|
||||
232
docs/adr/0021-plugin-descriptions-state-a-domain-boundary.md
Normal file
232
docs/adr/0021-plugin-descriptions-state-a-domain-boundary.md
Normal file
@@ -0,0 +1,232 @@
|
||||
# A plugin's published description states its domain boundary and never enumerates its skills
|
||||
|
||||
Three of this repo's six plugins publish a `description` that lists the skills they ship. That style
|
||||
has now failed three times in four days, the third time inside the correction for the second. It is
|
||||
enforced by nothing, it obliges a marketplace release on every skill addition, and it was never
|
||||
applied to the other three plugins. This ADR retires it: a published description says what the
|
||||
plugin is *for*, and the inventory lives where an inventory can be read off the tree.
|
||||
|
||||
**Status: accepted (2026-08-17).**
|
||||
|
||||
## Context
|
||||
|
||||
A plugin's published description is one string authored twice — in `plugins/<name>/apm.yml` and in
|
||||
the matching `marketplace.packages[]` entry of the root `apm.yml` — and compiled into four generated
|
||||
files per plugin edit: the plugin's `.claude-plugin/plugin.json` and `.github/plugin/plugin.json`,
|
||||
plus the repo-wide `.claude-plugin/marketplace.json` and its `.github/plugin/marketplace.json`
|
||||
mirror. (`.agents/plugins/marketplace.json`, apm's codex profile, carried no per-package
|
||||
`description` or `version` at all and was unaffected — that file and the profile producing it were
|
||||
removed 2026-09-13; see the note below.) It is the only text a consumer sees in a marketplace listing before
|
||||
installing. It is **not** a SKILL.md `description`: it is never preloaded into an agent's context and
|
||||
routes nothing at runtime. ADR-0020 governs that other artifact; this one governs this one. The
|
||||
overlap is a finding, not a scope: ADR-0020 established that capability enumeration in a description
|
||||
is "a correctness hazard, not only a token cost". The hazard at this layer is different — staleness
|
||||
in published metadata rather than an agent shortcutting the body — but the enumeration is the same
|
||||
construct and it fails the same way.
|
||||
|
||||
Measured at `de84d1b`, the branch tip before this change. Each figure is reproducible from the tree:
|
||||
skill counts are `ls plugins/<name>/.apm/skills/ | wc -l`, description text is
|
||||
`plugins/<name>/apm.yml`.
|
||||
|
||||
| Plugin | Style | Skills | Items enumerated | Skills named | Unnamed |
|
||||
|---|---|---|---|---|---|
|
||||
| `bin` | enumeration | 11 | 8 | 9 | `caveman`, `zoom-out` |
|
||||
| `git` | enumeration | 9 | 8 | 8 | `git-workflow` |
|
||||
| `gitea` | enumeration | 7 | 7 | 6 | `gitea-workflow` |
|
||||
| `core` | boundary | 3 | — | — | — |
|
||||
| `kyberforge` | boundary | 7 | — | — | — |
|
||||
| `lint` | boundary | 2 | — | — | — |
|
||||
|
||||
Three failures, in order.
|
||||
|
||||
**`bb9158d` (2026-08-14) — `core`'s description described `bin`.** The text it deleted read
|
||||
"Cross-cutting utility skills for everyday AI-assisted coding — triage, diagnosis, architecture
|
||||
review, and session navigation." All four items are real skills and not one of them is `core`'s:
|
||||
they are `bin`'s `triage`, `diagnose`, `improve-codebase-architecture` and `zoom-out`. `core` ships
|
||||
`agentsmd-author`, `agentsmd-audit` and `provider-adapter-author`, and the published description
|
||||
named none of them.
|
||||
|
||||
This is the failure the whole style was later adopted against, and it is worth being exact about
|
||||
what it was, because the record has been read the other way twice since. It was **wrong content**,
|
||||
not an incomplete list. The description was a syntactically perfect, complete, four-item enumeration
|
||||
of a real skill set; it just belonged to a different plugin. Enumerating harder could not have caught
|
||||
it, and a gate that asked "does every enumerated item exist as a skill?" would have passed it — all
|
||||
four did exist. `bb9158d`'s own fix went the other direction: it replaced the enumeration with a
|
||||
domain boundary, and `core` has needed no correction since. The precedent set by that commit was
|
||||
therefore *boundary*, and the two commits below cite it while doing the opposite.
|
||||
|
||||
**`65bac15` (2026-08-17) — `git` advertised `gitea`'s domain, `gitea` advertised a skill that does
|
||||
not exist.** `git` read "conventional commits, branch management, pull requests, and feature flow";
|
||||
pull requests reach the forge over HTTP and are `gitea`'s, which is the exact boundary
|
||||
`docs/spec/architecture.md` draws between the two plugins. `gitea` read "issues, pull requests,
|
||||
milestones, releases, and wikis"; `grep -ri wiki plugins/gitea/.apm/` returns nothing and no wiki
|
||||
skill has ever existed. Both were repaired by re-enumerating.
|
||||
|
||||
**`de84d1b` (2026-08-17) — the re-enumeration was itself incomplete.** `bin`'s "A place for things to
|
||||
be binned" was replaced with an eight-item list over eleven skills; `caveman` and `zoom-out` are
|
||||
absent. `zoom-out` is the same skill `bb9158d` had called "session navigation" three days earlier
|
||||
while deleting it from the wrong plugin's description — named when it was in the wrong place,
|
||||
unnamed once it was in the right one. And the miss is not confined to `bin`: `git-workflow` is
|
||||
unnamed in `git`'s corrected description, though `65bac15`'s own commit message states it was added
|
||||
("omitting pc-author/pc-run, git-submodules and git-workflow"), and `gitea-workflow` is unnamed in
|
||||
`gitea`'s. Across the three plugins, 23 of 27 skills are named at the third attempt.
|
||||
|
||||
**Nothing checks any of this.** `scripts/check-manifests.sh` does not contain the string
|
||||
`description`. The three ADR-0020 validators (`scripts/skill-size-check.sh` and skill-audit's and
|
||||
agent-audit's `validate.sh`) gate on SKILL.md and agent frontmatter; they do open `apm.yml`, but only
|
||||
to read `dependencies.apm` when resolving the boundary-target universe — none of them reads the
|
||||
`description:` key, and their hook globs match `SKILL.md` and `*.agent.md` only. `apm audit --ci`,
|
||||
`apm pack --check-clean` and `scripts/sync-plugin-content.sh --check --all` all compare compiled
|
||||
output against `apm.yml`, so their entire job is to propagate whatever the description says into
|
||||
those four files byte-for-byte and confirm they match. The `wiki` claim passed every one of the fourteen pre-push hooks, every day it
|
||||
was published.
|
||||
|
||||
**And the obligation is unbounded.** Under enumeration, adding one skill to `bin`, `git` or `gitea`
|
||||
means editing two copies of a prose string on top of the version bumps and regeneration any skill
|
||||
addition already owes under this repo's release policy
|
||||
(`plugins/kyberforge/.apm/skills/apm-workflow/references/marketplace.md`). The bumps are not the
|
||||
marginal cost — the prose edit is, and it is the half nothing checks. A skill *rename* triggers the
|
||||
same, for a string no consumer can tell went stale. 27 of the repo's 39 skills sat behind
|
||||
a description carrying that obligation; the other 12 did not, and their three plugins have generated
|
||||
no defect of this class.
|
||||
|
||||
### Scope
|
||||
|
||||
This decision covers the six plugins this repo authors. The root marketplace also lists
|
||||
`mattpocock-skills`, a third-party package whose description is not this repo's to write; its entry
|
||||
is out of scope and is left as published upstream.
|
||||
|
||||
*(Note, 2026-09-13: `mattpocock-skills` has since been removed from the root marketplace. This
|
||||
section's scope statement is retained as the reasoning behind the boundary; the entry it describes
|
||||
no longer exists.)*
|
||||
|
||||
## Decision
|
||||
|
||||
**A plugin's published `description` states the plugin's domain boundary. It does not enumerate the
|
||||
skills the plugin ships, by name or by paraphrase.**
|
||||
|
||||
- The boundary answers "what kind of work belongs to this plugin, and where is its edge against its
|
||||
nearest sibling" — the question a consumer deciding whether to install is actually asking. It is
|
||||
stable under skill addition, rename and removal, which is the entire point: an artifact that does
|
||||
not change when the tree changes cannot go stale against it.
|
||||
- **The boundary must cover everything the plugin actually ships.** A boundary drawn narrower than
|
||||
the contents is the same defect as an incomplete enumeration, one level up, and it is the specific
|
||||
risk in this change. `git` carries `pc-author` and `pc-run`, which are not git operations at all;
|
||||
"Skills for working with Git" silently drops them, so the boundary names the pre-commit hooks
|
||||
explicitly rather than trusting a reader to file them under Git.
|
||||
- The two copies — package `apm.yml` and the root `marketplace.packages[]` entry — stay identical.
|
||||
This is already the rule in practice and both prior corrections state why: the root entry is what
|
||||
reaches the compiled marketplace, so fixing only the package manifest leaves it half-propagated.
|
||||
- The three descriptions, rewritten here, with `core`/`kyberforge`/`lint` shown for register:
|
||||
|
||||
| Plugin | Published description | Chars |
|
||||
|---|---|---|
|
||||
| `bin` | Skills for everyday AI-assisted development work that is not tied to a single tool, forge or language, and has not yet been split into a focused plugin. | 152 |
|
||||
| `git` | Skills and agents for working with a local Git clone over the git wire protocol, and for authoring and running the pre-commit hooks that guard it. | 146 |
|
||||
| `gitea` | Skills and agents for working with a Gitea forge through its HTTP API — the forge's own objects, as distinct from the local git clone. | 134 |
|
||||
| `core` | *(unchanged)* Skills for authoring and auditing a repo's AGENTS.md and the provider adapter files that defer to it. | 101 |
|
||||
| `kyberforge` | *(unchanged)* Skills and agents for creating, maintaining, and managing a Claude Code / Copilot CLI plugin marketplace. | 105 |
|
||||
| `lint` | *(unchanged)* Skills and agents for configuring and running linters. | 54 |
|
||||
|
||||
- **No gate is added.** This is a deliberate omission and the reasoning is below, not an item left
|
||||
for later.
|
||||
|
||||
### Why no gate
|
||||
|
||||
The check enumeration would need — "every skill directory appears in the description" — was writable
|
||||
in principle and was never written, including by the two commits that corrected an enumeration by
|
||||
enumerating again and had every reason to. It is also only half a check: it
|
||||
catches a skill missing from the list, and it cannot catch `wiki`, because "this noun does not name
|
||||
any skill" requires a vocabulary of permissible non-skill nouns that no one is going to maintain.
|
||||
Under a boundary there is no correspondence left to check, which is the property being bought.
|
||||
|
||||
What survives un-gated is `bb9158d`'s actual failure: a boundary that is simply wrong about its
|
||||
plugin. That was never machine-checkable in either style — the text was a well-formed description of
|
||||
a real plugin — and it is caught by the same review that has to happen when a published,
|
||||
consumer-facing string is edited at all. A gate that would catch it needs a declared per-plugin
|
||||
skill-to-boundary mapping for the description to be checked against, which is a second artifact
|
||||
requiring exactly the per-skill maintenance this ADR exists to delete, relocated one file over.
|
||||
|
||||
Two cheap partial gates were considered and rejected in the same breath. Forbidding a comma-separated
|
||||
run of three or more noun phrases is a prose heuristic that fires on `lint`'s perfectly good
|
||||
"configuring and running linters" class of sentence. Forbidding any string matching a skill directory
|
||||
name under `plugins/<name>/.apm/skills/` bans legitimate boundary vocabulary — `git-branches` exists,
|
||||
and a `git` boundary has every right to say "branches". Both would be believed, and both would be
|
||||
wrong, which ADR-0020 already records as worse than no gate.
|
||||
|
||||
## Considered options
|
||||
|
||||
**Keep enumeration and gate it.** The only option that makes the current style safe. Rejected on the
|
||||
three grounds above: the check is one-directional, it cannot see an invented capability, and it makes
|
||||
a marketplace release the consequence of adding a directory. It also hard-couples published consumer
|
||||
copy to internal directory names, so a skill rename becomes a version bump on the plugin and on the
|
||||
marketplace.
|
||||
|
||||
**Enumerate consistently across all six plugins**, on the grounds that the real defect is the split
|
||||
style. Rejected: it takes an obligation that has produced three failures on three plugins and applies
|
||||
it to six. The measured outcome of the most recent attempt to enumerate carefully, with the defect
|
||||
fresh and two prior commits as precedent, is four skills unnamed.
|
||||
|
||||
**Cap the description length**, mirroring ADR-0020's 250/400-character tiers, on the theory that a
|
||||
short description has no room to enumerate. Rejected because length does not measure correspondence:
|
||||
`gitea`'s failing description was 96 characters and asserted a skill that has never existed, while
|
||||
`bin`'s 176-character enumeration is under the same cap. All six descriptions here, before and after,
|
||||
sit inside ADR-0020's tiers; the tier would have been silent through all three failures.
|
||||
|
||||
**Delete the description to a bare name.** Rejected: apm's Claude marketplace mapper emits
|
||||
`description` into `marketplace.json`, and it is the only prose a consumer sees before installing.
|
||||
|
||||
**Point the description at the plugin's `README.md`.** Rejected: a marketplace listing renders a
|
||||
string, not a link — and the README's own plugin list carries the same enumeration with the same
|
||||
staleness, so this relocates the defect rather than fixing it.
|
||||
|
||||
## Consequences
|
||||
|
||||
**Three descriptions are rewritten and the compiled output regenerated.** Eight generated files
|
||||
change: `plugins/{bin,git,gitea}/.claude-plugin/plugin.json`,
|
||||
`plugins/{bin,git,gitea}/.github/plugin/plugin.json`, `.claude-plugin/marketplace.json` and its
|
||||
byte-identical `.github/plugin/marketplace.json` mirror. `.agents/plugins/marketplace.json` (the
|
||||
codex profile) is unchanged and correctly so — it carries no per-package `description` or `version`
|
||||
field at all, only `name`, `source`, `policy` and `category`.
|
||||
|
||||
**Version bumps, all PATCH under the `per_package` strategy:** `bin` 1.1.4 → 1.1.5, `git` 1.3.4 →
|
||||
1.3.5, `gitea` 1.3.5 → 1.3.6, `marketplace.version` 0.4.4 → 0.4.5.
|
||||
|
||||
**The root `apm.yml` top-level `version:` is restored to lockstep with `marketplace.version`,
|
||||
0.4.2 → 0.4.5.** These two fields have moved together in every commit that has ever touched root
|
||||
`apm.yml` — 0.3.2, 0.3.3, 0.3.4, 0.4.0, 0.4.1, 0.4.2 in both — until `65bac15` and
|
||||
`de84d1b` on this branch bumped `marketplace.version` to 0.4.3 and then 0.4.4 while leaving the
|
||||
top-level field at 0.4.2. Lockstep is not folklore: it is stated at
|
||||
`plugins/kyberforge/.apm/skills/apm-workflow/references/marketplace.md`. This is a defect, not a
|
||||
style: `apm.yml`'s comment inside the marketplace block records that the top-level `version:` is not inherited into the compiled output
|
||||
"despite being used elsewhere (e.g. by `apm audit`)", so the field is live and was silently two
|
||||
releases behind what the marketplace published. Closed here rather than tracked, because the
|
||||
correction is one line and the drift is three days old.
|
||||
|
||||
**`docs/spec/architecture.md`'s plugin table is unchanged and stays a routing table.** It answers
|
||||
"where does a new skill go" for someone working *inside* this repo; the published description answers
|
||||
"should I install this" for someone outside it. The two now read similarly, and that is not
|
||||
duplication to collapse — they have different readers and different lifecycles, and the table already
|
||||
says so in its own preamble ("These are routing boundaries, not inventories"). One caveat for whoever
|
||||
next edits that page: its closing sentence sends a reader to the published description "for what a
|
||||
consumer actually gets", which was true against an enumeration and is now a pointer to a second
|
||||
boundary statement. Neither artifact carries an inventory after this change, so that sentence was
|
||||
rewritten in the same branch to point at `plugins/<name>/.apm/skills/` and `README.md` instead.
|
||||
|
||||
**`README.md`'s plugin bullet list becomes the only place an inventory lives, and it still
|
||||
enumerates.** That is deliberate, but it makes the list load-bearing in a way it was not before, so
|
||||
its `bin`, `git` and `gitea` bullets were completed in the same branch to name every skill those
|
||||
plugins ship. This ADR does not otherwise extend to it: a README is a hand-read document where a
|
||||
list of what you get is the useful thing, it is not compiled into four files, and a stale line in it
|
||||
costs a reader a moment rather than misrepresenting a published package. The tradeoff that makes
|
||||
enumeration wrong in a marketplace manifest is precisely the one that makes it fine there.
|
||||
|
||||
**Nothing in the ADR-0020 gate set changes.** Its character and word tiers, its Vale rules and its
|
||||
three validators all read `SKILL.md` and `*.agent.md` frontmatter; none of them opens an `apm.yml`.
|
||||
The two contracts are adjacent and independent, and a future author retrofitting a skill under
|
||||
issue #99 is not touched by this ADR.
|
||||
|
||||
**The failure mode this leaves open is a wrong boundary, and it is un-gated by design.** If a fourth
|
||||
failure of this class occurs it will be a description that describes the wrong plugin — `bb9158d`'s
|
||||
shape, the one enumeration never addressed. That is the trigger to revisit, and the thing to build
|
||||
then is a declared skill-to-boundary mapping, not a return to enumeration.
|
||||
78
docs/adr/0022-skill-metadata-version-is-mandatory.md
Normal file
78
docs/adr/0022-skill-metadata-version-is-mandatory.md
Normal file
@@ -0,0 +1,78 @@
|
||||
# Every skill's `metadata.version` is mandatory, not a per-plugin option
|
||||
|
||||
**Status: accepted (2026-09-07).**
|
||||
|
||||
## Context
|
||||
|
||||
`metadata.version` is optional SKILL.md frontmatter (`create.md`'s "Optional frontmatter" list:
|
||||
"uncomment and fill in, or remove entirely"). `skill-author`'s own bump logic was written
|
||||
conditionally — "with `metadata.version` present, bump the minor version on create... and the
|
||||
patch version on improve" — which only makes sense if presence is a real per-skill choice.
|
||||
|
||||
Adoption never followed a rule; it followed the plugin. Of 39 skills, 12 carry a version:
|
||||
|
||||
| Plugin | Has it | Total |
|
||||
|---|---|---|
|
||||
| `core` | 3 | 3 |
|
||||
| `gitea` | 6 | 7 |
|
||||
| `lint` | 2 | 2 |
|
||||
| `git` | 1 | 9 |
|
||||
| `bin` | 0 | 11 |
|
||||
| `kyberforge` | 0 | 7 |
|
||||
|
||||
`core`, `gitea` and `lint` are consistent adopters (`gitea-files` the one gap); `bin` and
|
||||
`kyberforge` are consistent non-adopters; `git` has one outlier (`git-commits`, versioned for no
|
||||
plugin-specific reason found on inspection — no comment, no cross-reference, nothing distinguishing
|
||||
it from its eight siblings). Issue #127 raised this as an undocumented split: two house norms
|
||||
coexisting with no stated rule for which applies where, the same class of defect as an unstated
|
||||
`rtk`/bare-`git` convention (#113) found in the same audit pass.
|
||||
|
||||
## Decision
|
||||
|
||||
**Every skill's frontmatter carries `metadata.version`.** It is no longer optional, and no longer a
|
||||
per-plugin choice.
|
||||
|
||||
- **The 27 skills that never carried one are seeded at `1.0.0`**, not `0.1.0`. `0.1.0` is
|
||||
`skill-author`'s existing new-skill starting point, chosen for a skill with no revision history to
|
||||
its name yet. These 27 have all been through the ADR-0020 retrofit and repeated audit passes
|
||||
without ever tracking a version; crediting them with `0.1.0` would understate that, and there is
|
||||
no real history to justify seeding higher than a first stable release. `1.0.0` marks "versioned as
|
||||
of this retrofit," `0.1.0` keeps meaning "created and never yet revised."
|
||||
- **New skills still start at `0.1.0`.** `skill-author`'s create/improve bump convention is
|
||||
unchanged; only the presence of the field stops being conditional.
|
||||
- **The one outlier in the other direction, `git-commits`, keeps its existing value** (`0.1.3`) —
|
||||
it already had real tracked history under the old conditional rule, and this decision does not
|
||||
reset skills that were already compliant.
|
||||
- **`bin/write-docs`'s top-level `version:` moves into `metadata:`, normalized to `1.0.0`.** It is
|
||||
the one skill that carried a version outside the `metadata:` block, which is why the table above
|
||||
counts `bin` as 0 — a top-level `version:` is not `metadata.version`, and nothing reads it. #127
|
||||
raised it alongside the split because "does a skill carry a version" and "where does it live" are
|
||||
the same question. Its value (`1.0`) is not semver and carries no more real history than the 27
|
||||
unversioned skills, so it is relocated and reset to the same `1.0.0` seed rather than preserved
|
||||
like `git-commits`'s tracked `0.1.3`.
|
||||
- **`skill-frontmatter`'s pre-commit hook gains the check.** It already fails a SKILL.md missing
|
||||
`name:` or `description:`; a missing `metadata.version` is now the same class of failure, not a
|
||||
style nit an audit might or might not catch.
|
||||
|
||||
## Considered options
|
||||
|
||||
**Leave it per-plugin, document the split.** This was the initial framing of #127 and is coherent —
|
||||
`core`/`gitea`/`lint` keep it, `bin`/`kyberforge` don't, two outliers get normalized to match their
|
||||
plugin. Rejected on reconsideration: a rule that says "some plugins track this and some don't" is
|
||||
strictly harder to state, audit and onboard against than "every skill does," for a field whose entire
|
||||
job is answering "did this change since I last read it" — a question with the same shape everywhere
|
||||
it's asked, not one that varies by plugin domain.
|
||||
|
||||
**Drop the field corpus-wide.** Rejected: `skill-author` already depends on it to decide whether a
|
||||
create/improve pass owes a bump, so the 12 skills carrying it are not tracking dead weight — removing
|
||||
it discards real revision signal for no gain.
|
||||
|
||||
## Consequences
|
||||
|
||||
27 SKILL.md files gain `metadata.version: "1.0.0"`, and a 28th — `bin/write-docs` — reaches the same
|
||||
value by relocating its top-level `version: "1.0"` into `metadata:`. `skill-author`'s `create.md`
|
||||
moves the field from "Optional frontmatter" to the required list, citing this ADR. `skill-author`'s
|
||||
own SKILL.md drops the "with `metadata.version` present" conditional in its bump-rule line, since
|
||||
presence is no longer in question. `.pre-commit-config.yaml`'s `skill-frontmatter` hook is extended
|
||||
to require the field, closing the gap #113 and #118 both named in the same audit pass: a stated rule
|
||||
with nothing enforcing it drifts the same way an unstated one does.
|
||||
168
docs/adr/0023-rtk-prefix-marks-executable-commands-only.md
Normal file
168
docs/adr/0023-rtk-prefix-marks-executable-commands-only.md
Normal file
@@ -0,0 +1,168 @@
|
||||
# The `rtk` prefix marks executable commands only, and is repo-wide
|
||||
|
||||
**Status: accepted (2026-09-08).**
|
||||
|
||||
## Context
|
||||
|
||||
`CLAUDE.md` states the org convention as a golden rule: "Always prefix commands with `rtk`. If RTK
|
||||
has a dedicated filter, it uses it. If not, it passes through unchanged. This means RTK is always
|
||||
safe to use." Issue #113 observed that the rule had never been written down for skill *prose*, where
|
||||
a `git <subcommand>` mention can be either an instruction to execute or a reference to the concept,
|
||||
and that the corpus had drifted into carrying both spellings with no stated rule. PR #130 swept the
|
||||
`git` plugin and recorded a two-way split in `plugins/git/README.md`.
|
||||
|
||||
Review found two defects in that sweep, and both are in the premise rather than the execution.
|
||||
|
||||
**RTK is not output-transparent.** `rtk git --help` enumerates twelve filtered subcommands — `diff`,
|
||||
`log`, `status`, `show`, `add`, `commit`, `push`, `pull`, `branch`, `fetch`, `stash`, `worktree`.
|
||||
Everything else is a true passthrough. Inside that set the filter is not a formatting preference; it
|
||||
changes what the command *reports*. Measured against rtk 0.42.4:
|
||||
|
||||
| Command | What rtk does to it |
|
||||
|---|---|
|
||||
| `worktree list --porcelain -z` | discards both flags; no NUL separators, no `locked`/`lock_reason` field at all |
|
||||
| `worktree list -v` | abbreviates `/root/…` to `~/…`, collapses column alignment |
|
||||
| `branch --list <name>` | emits a phantom `* ` line even on no match |
|
||||
| `diff --name-only` / `--name-status` | appends a blank line and a `Changes:` trailer |
|
||||
| `diff --word-diff[=color\|=porcelain]` | emits no `[-removed-] {+added+}` markers; substitutes a diffstat |
|
||||
| `log -L` | truncates each diff body line at ~72 characters with an ellipsis |
|
||||
| `stash pop` (on conflict) | prints only `FAILED: git stash pop`, swallowing `CONFLICT`, `Unmerged paths` and the retained-entry notice |
|
||||
| `stash list` (empty) | prints `No stashes` where git prints nothing |
|
||||
|
||||
Every one of those falsified a skill that was written against the bare output. `git-worktrees`'s
|
||||
Step 2 required `locked` and `lock_reason` from a command whose rtk rendering has never carried
|
||||
them; `git-log-format.md` documented `[-removed-] {+added+}` markers beside a command that no longer
|
||||
produces them. The two-way split could not see any of this, because both halves of it are about what
|
||||
a *sentence* is doing and none of it is about what the *command* does.
|
||||
|
||||
**The rule is not `git`-plugin-scoped.** `plugins/git/README.md` claimed the `gitea-*` skills
|
||||
"contain no `git`/`rtk` mentions at all". Five `gitea-*` SKILL.md files run `git remote get-url
|
||||
origin` in a fenced ```bash Step block — the README's own canonical example of "executable,
|
||||
instructed" — plus `git branch --show-current` in a reference file and three `git remote -v` in
|
||||
`gitea-orchestrate.agent.md`. A convention stated inside one plugin's README is invisible from the
|
||||
plugin next door, which is how those eight sites stayed bare through the sweep that existed to find
|
||||
them.
|
||||
|
||||
## Decision
|
||||
|
||||
**One rule, three clauses, repo-wide** — every `plugins/*/.apm/skills/**` and
|
||||
`plugins/*/.apm/agents/**` file, not the `git` plugin alone.
|
||||
|
||||
1. **Executable and instructed → `rtk git`.** Anything telling the agent to run a command now: an
|
||||
imperative step, a dispatch-table "Run" cell, a fenced code-block procedure.
|
||||
`rtk git push -u origin <branch>`.
|
||||
2. **Illustrative or referential → bare `git`.** Naming a flag's behaviour, quoting a doc's own
|
||||
heading, describing a command in the abstract, warning against an anti-pattern. "`git switch`
|
||||
refuses rather than clobbering conflicting local edits."
|
||||
3. **Machine-parsed or interactive → bare `git`, and say why inline.** A command whose output the
|
||||
skill parses, where rtk is in the filtered set above; or a command that hands control to an
|
||||
interactive child process.
|
||||
|
||||
Clause 3 is the new one and it looks arbitrary without the table in Context, which is why the
|
||||
measurements are recorded here rather than left in a PR thread. It is applied per subcommand and per
|
||||
flag, not per skill: `tag --list` stays prefixed because rtk passes it through byte-identically,
|
||||
while `branch --list` two words away goes bare because it does not. `git remote get-url origin`,
|
||||
`git remote -v`, `git branch --show-current`, `git log --oneline -1` and `git add -u` were all
|
||||
re-measured as byte-identical passthroughs and are therefore prefixed, parsing notwithstanding.
|
||||
|
||||
Two consequences of that per-subcommand basis are worth stating, because both are load-bearing and
|
||||
neither is comfortable:
|
||||
|
||||
- **rtk's filtered set is a moving target.** `git rebase` and `git mergetool` are passthroughs on
|
||||
0.42.4 — verified under `script(1)`, both inherit a real TTY, contradicting an earlier report that
|
||||
they did not. They stay bare anyway, on the interactive limb: a token filter has nothing to offer a
|
||||
command that hands control to an editor, and the prefix would only buy exposure to whatever a later
|
||||
rtk version decides to do with those subcommands. The same reasoning makes the *inner* call in
|
||||
`` `rtk git remote add origin-push $(git config remote.origin.url)` `` bare while the outer stays
|
||||
prefixed — `config` passes through cleanly today, but its stdout becomes a remote URL that is then
|
||||
force-pushed to, and that is not a blast radius to lend to a future filter change.
|
||||
- **`branch --show-current` sits on the sharp edge.** It is in the filtered set, it is parsed, and it
|
||||
is prefixed — on a measurement, in a subcommand whose sibling `--list` is exactly the defect clause
|
||||
3 exists for. If rtk's `branch` filter is ever extended, that is the first site to break. It is
|
||||
called out rather than hedged, because a rule whose exceptions are unrecorded is the state this ADR
|
||||
is replacing.
|
||||
|
||||
**A clause-3 site says so inline, in a few words.** "bare, not `rtk`: rtk prints a phantom `* ` line
|
||||
even on no match". Without it the next sweep re-prefixes the command, which is how #113 recurs.
|
||||
|
||||
**The rule lives here, and `docs/spec/gates.md` carries the gate.** `plugins/git/README.md` is
|
||||
reduced to a pointer. It had also cited `git-workflow/references/hard-rules.md` as a place the rule
|
||||
was written down; that file contains no occurrence of "rtk", and the citation is removed rather than
|
||||
repaired.
|
||||
|
||||
**Clause 1 is enforced by a `check-rtk-prefix` pre-commit hook; clauses 2 and 3 are not enforceable
|
||||
and are not gated.** The hook checks the two places a `git` mention is unambiguously an instruction —
|
||||
a line in a shell-tagged code fence, and the opening backticked span of a "Run" column cell — and a
|
||||
deliberately-bare command opts out with the literal string `ADR-0023` on its own line. Its coverage
|
||||
limits are recorded in `docs/spec/gates.md`, not smoothed over.
|
||||
|
||||
## Considered options
|
||||
|
||||
**Add `compatibility:` frontmatter to every skill.** These six plugins are installable by third
|
||||
parties, and a consumer who installs `git` from the marketplace has no `rtk` on their PATH. Every
|
||||
prefixed command in the corpus is a plain `git` invocation with a word in front of it, so the prefix
|
||||
is *droppable*: delete `rtk ` and the command is correct. A `compatibility:` line per skill would
|
||||
state that in a machine-readable field. Rejected on cost. It is 39 lines of frontmatter restating one
|
||||
sentence, it is preloaded into every agent's context every session under ADR-0020's budget — the
|
||||
field is not free the way a line in a doc is — and it has no consumer: nothing reads
|
||||
`compatibility:`, so the field would be a comment with a colon in it. The consumer situation is
|
||||
documented here and in `plugins/git/README.md` instead, which is where a human installing a plugin
|
||||
actually looks. The same two-line note is owed to the other five plugin READMEs and is not yet
|
||||
written.
|
||||
|
||||
**Move rtk to the execution layer entirely.** Skills instruct bare `git` throughout; `CLAUDE.md`'s
|
||||
session rule handles prefixing at the point of execution. This is the strongest rejected option and
|
||||
it deserves the space: it closes the consumer gap and all eight output defects at once, because the
|
||||
executing agent knows what it is about to parse and the skill does not have to predict it. It also
|
||||
removes clause 3 entirely — there is nothing to except. Rejected because the prefix is lost wherever
|
||||
an agent copies a command literally, which is the common case for a fenced procedure block and the
|
||||
whole reason dispatch tables exist. The org convention's value is that the prefix is *already there*
|
||||
in the text the agent lifts; a rule that relies on the agent remembering to add it is the rule that
|
||||
produced the drift in the first place. Worth revisiting if rtk ever ships a shell shim, which would
|
||||
make the execution layer transparent and this trade different.
|
||||
|
||||
**Keep the two-way split and fix the eight sites by hand.** Rejected: the split has no vocabulary for
|
||||
"this command is executable, instructed, and must still be bare", so the eight sites would be
|
||||
unexplained exceptions and the next sweep re-prefixes them. That is the failure this ADR exists to
|
||||
stop, not a smaller version of it.
|
||||
|
||||
**Gate clauses 2 and 3 as well.** Rejected as undecidable. "Run `git switch <branch>`" and "`git
|
||||
switch` refuses rather than clobbering local edits" are the same token sequence; separating them is a
|
||||
judgement about what a sentence is doing. A gate that guessed would fire on correct content, and a
|
||||
gate that fires on correct content gets added to `SKIP`, which disarms clause 1 along with it.
|
||||
|
||||
## The boundary the rule does not decide
|
||||
|
||||
Two shapes in the corpus resisted the two-way split. The three-clause rule resolves one and does not
|
||||
resolve the other; both are recorded so an author meeting a third one knows which kind it is.
|
||||
|
||||
**`git-worktrees/SKILL.md`'s tracking row carries both spellings in one Run cell** — `rtk git
|
||||
worktree add --track -b <branch> <path> <remote>/<branch>` — always correct. `git worktree add
|
||||
<path> <branch>` expands to exactly this. **Resolved: the clauses apply per mention, not per row,
|
||||
per cell or per file.** The first is the instruction (clause 1), the second names what the first
|
||||
expands to (clause 2), and one table cell can hold one of each. The rule needed no change; the
|
||||
*gate* did, and it checks only a Run cell's opening span for exactly this reason.
|
||||
|
||||
**`git-submodules/references/setup-and-update.md:80` has a git command inside a quoted argument to
|
||||
another command** — `rtk git submodule foreach 'git pull origin main || :'`. **Not resolved: all
|
||||
three clauses describe a command the reading agent executes, and the inner `git pull` is not one.**
|
||||
It is the literal text of an argument that `git submodule foreach` hands to a subshell running inside
|
||||
each submodule's own working tree, where the local convention does not reach. The file already gets
|
||||
this right and already justifies it in prose two lines below ("the git calls in it are the
|
||||
submodule's own — that is the one place a bare `git` is correct"). **An author meeting this shape
|
||||
should do the same: leave the inner command bare and justify it inline.** It is deliberately not
|
||||
promoted to a fourth clause on one instance. The gate does not decide it either — it happens to pass
|
||||
this line, because the segment containing the inner command begins with `rtk`, and that is an
|
||||
accident of the split rather than an understanding of quoting.
|
||||
|
||||
## Consequences
|
||||
|
||||
Eleven sites in `plugins/git/.apm/skills/**` revert to bare `git` under clause 3, each carrying a
|
||||
short inline reason. Eight sites across `plugins/gitea/.apm/skills/**` and
|
||||
`plugins/gitea/.apm/agents/gitea-orchestrate.agent.md` gain the prefix under clause 1, and one in
|
||||
`pc-run/SKILL.md` that the #130 sweep's grep missed because the backtick opens with `SKIP=` rather
|
||||
than `git `. `plugins/git/README.md`'s Conventions section becomes a pointer here, minus a paragraph
|
||||
that was false about the `gitea-*` skills and a citation to a file that does not carry the rule.
|
||||
A `check-rtk-prefix` pre-commit hook and `tests/test-check-rtk-prefix.sh` land with it; the test runs
|
||||
the gate against the pre-sweep corpus on `main` and asserts it fails there, because a gate that only
|
||||
passes on the fixed tree proves nothing about the drift it was written for.
|
||||
259
docs/adr/0024-apm-is-the-only-supported-install-path.md
Normal file
259
docs/adr/0024-apm-is-the-only-supported-install-path.md
Normal file
@@ -0,0 +1,259 @@
|
||||
# apm is the only supported install path; the flat content mirror is deleted
|
||||
|
||||
**Supersedes ADR-0017** (plugin roots gain a compiled flat-directory mirror of `.apm/` content so
|
||||
Claude Code can discover it). ADR-0017's diagnosis was correct and is not in dispute: Claude Code's
|
||||
native installer convention-scans flat `skills/`/`agents/`/`hooks/` directories at the plugin root
|
||||
and has no model of `.apm/` at all, so without a mirror a natively-installed holocron plugin reports
|
||||
`Skills (0) Agents (0) Hooks (0)`. What changes here is not the mechanism but the premise — that the
|
||||
native install path is worth supporting. It is not, because nobody uses it.
|
||||
|
||||
**Status: accepted (2026-09-14).** The mirror, its generator, its test suite, its helper library and
|
||||
its pre-push gate are removed. `.apm/` remains the sole hand-edited authoring source, unchanged from
|
||||
ADR-0015. The root `marketplace:` block in `apm.yml` and the compiled
|
||||
`.claude-plugin/marketplace.json` it produces are **kept** — see "Also delete the marketplace
|
||||
catalogue" under considered options.
|
||||
|
||||
## Context
|
||||
|
||||
ADR-0018 moved this repo's own consumption of its own plugins onto `apm install`. From that point
|
||||
the flat mirror had no consumer inside this repo: it existed entirely for a hypothetical third party
|
||||
running `claude plugin install <name>@holocron`. No such consumer has ever been observed. The
|
||||
marketplace is on a private Gitea instance, and the repo has no telemetry, no issue traffic and no
|
||||
external clone record suggesting otherwise. The honest statement is that the native path has been
|
||||
maintained for an audience of zero.
|
||||
|
||||
What that audience costs is measurable:
|
||||
|
||||
| Artifact | Size |
|
||||
|---|---|
|
||||
| Tracked mirror files under `plugins/*/{skills,agents,hooks}/` | 213 files, ~20,000 lines |
|
||||
| `scripts/sync-plugin-content.sh` | 813 lines |
|
||||
| `tests/test-sync-plugin-content.sh` | 1,289 lines, 92 cases, ~83 s |
|
||||
| `scripts/lib/marketplace-plugins.sh` | helper, used only by the above |
|
||||
| `check-plugin-content-sync` pre-push hook | ~4.5 s per push |
|
||||
| `validate-plugins` pre-push hook | ~4.9 s per push (6 × `claude plugin validate --strict`) |
|
||||
|
||||
Roughly 22,000 lines of tracked content and tooling, and about 92 seconds on every push (83 + 4.5 +
|
||||
4.9; the two hook timings are the 2026-09-10 baseline measurements recorded in
|
||||
`SIMPLIFICATION-AUDIT.md`, not re-measured here). The test alone is close to 30% of `run-tests`'
|
||||
wall time — the single largest item in it.
|
||||
|
||||
**The native path's automated gate does not gate anything.** ADR-0017 cites
|
||||
`claude plugin validate --strict` passing on all six plugins as one of two verifications. That
|
||||
verification was re-run this session against a plugin directory with **every content directory
|
||||
deleted**, and it passed. `validate` reads the manifest; it never inspects content. It therefore
|
||||
cannot detect the exact `Skills (0) Agents (0) Hooks (0)` defect ADR-0017 was written to fix. The
|
||||
other half of ADR-0017's verification — the live behavioral test
|
||||
(`claude --plugin-dir plugins/kyberforge -p "list your skills and agents"`) — is a manual step,
|
||||
run by hand once in August 2026 and never since. So native-install correctness has been unguarded
|
||||
for a month, and the drift gate that runs on every push guards only that the mirror matches `.apm/`,
|
||||
not that the mirror works.
|
||||
|
||||
**Dropping native install does not reduce host coverage.** This is the fact that makes the decision
|
||||
cheap rather than a trade. apm's skills convergence deploys skills to `.agents/skills/<name>/SKILL.md`,
|
||||
the shared path read by Copilot, Cursor, Codex, Gemini, OpenCode and Windsurf, with Claude Code as
|
||||
the special case at `.claude/skills/`. The mirror served two hosts (Claude Code and Copilot, the
|
||||
latter only ever partially — see ADR-0017's own `hooks` amendment). apm serves eleven through those
|
||||
two roots: ten targets resolve to `.agents/skills/` — the six named above plus `agent-skills`,
|
||||
`antigravity`, `hermes` and `openclaw`, which are rooted there natively rather than by an explicit
|
||||
`deploy_root` — and `claude` is the eleventh at `.claude/skills/`. Four further targets (`kiro`,
|
||||
`grok-build`, `grok-cloud`, `copilot-cowork`) deploy skills under roots of their own, for fifteen in
|
||||
total; counted from `apm_cli/integration/targets.py` this session. A consumer who installs holocron
|
||||
through apm gets strictly more than one who installed it natively.
|
||||
|
||||
**Verified empirically, not reasoned about.** In a scratch clone with the mirror and the six
|
||||
per-plugin manifest pairs deleted:
|
||||
|
||||
- `apm marketplace add` still registers all 6 packages. It detects
|
||||
`.claude-plugin/marketplace.json` and reads the catalogue from there.
|
||||
- `apm install` still deploys 40 `SKILL.md` files across 39 skill directories, 4 agents, and
|
||||
kyberforge's `SessionStart` hook — identical to the baseline install from the unmodified tree.
|
||||
- `apm pack --check-versions --check-clean --dry-run` exits 0 ("Version alignment OK",
|
||||
"Marketplace working tree clean"), because it governs only the **root** `.claude-plugin/` outputs.
|
||||
The six per-plugin `plugin.json` pairs were never apm-pack-governed: they were generated by
|
||||
`apm pack --format plugin` invoked from inside `sync-plugin-content.sh`, so deleting the script
|
||||
deletes their producer and nothing is left asserting they should exist.
|
||||
|
||||
**Source-level proof the per-plugin manifests are droppable.**
|
||||
`apm_cli/deps/github_downloader_validation.py` probes package markers in a fixed order —
|
||||
`apm.yml`, then `SKILL.md`, then `plugin.json`, then `.github/plugin/plugin.json`, then
|
||||
`.claude-plugin/plugin.json` — and returns on the first hit. Every plugin here keeps its `apm.yml`,
|
||||
which is the first probe, so no `plugin.json` path is ever reached. The per-plugin manifests are not
|
||||
load-bearing for apm resolution; they were load-bearing only for the native installer.
|
||||
|
||||
## Decision
|
||||
|
||||
**apm is the only supported install path.** Concretely:
|
||||
|
||||
- Delete the flat mirror at every plugin root (`plugins/<name>/skills/`, `agents/`, `hooks/`) and
|
||||
the six per-plugin manifest pairs (`.claude-plugin/plugin.json`, `.github/plugin/plugin.json`).
|
||||
- Delete `scripts/sync-plugin-content.sh`, `tests/test-sync-plugin-content.sh`,
|
||||
`scripts/lib/marketplace-plugins.sh`, and the `check-plugin-content-sync` pre-push hook.
|
||||
- Delete the `validate-plugins` pre-push hook (`claude plugin validate --strict` over every plugin
|
||||
directory). It is removed on the evidence above — it reads manifests only, so with the per-plugin
|
||||
manifest pairs gone it has nothing left to read, and even while they existed it could not detect
|
||||
the empty-content defect. `validate-marketplace`, which validates the one manifest this repo still
|
||||
ships, is **kept**.
|
||||
- **Keep** the root `marketplace:` block in `apm.yml` and the compiled
|
||||
`.claude-plugin/marketplace.json`. apm's own marketplace consumers read that same file; removing
|
||||
it would stop holocron being an apm marketplace at all.
|
||||
|
||||
The asymmetry between those last two bullets is the whole subtlety of this ADR, and it exists
|
||||
because apm deliberately reuses Claude Code's catalogue format rather than inventing one. The
|
||||
catalogue is shared between the two ecosystems; the per-plugin content contract is not. Deleting the
|
||||
content is what ends native support; keeping the catalogue is what preserves apm support.
|
||||
|
||||
## Considered options
|
||||
|
||||
**Status quo — keep mirroring on every branch (rejected).** Pays ~22,000 tracked lines and ~92
|
||||
seconds per push for a path with no users and no working gate. It is not free in author attention
|
||||
either: ADR-0017 accrued four amendments in two days, every one of them about a detail of the
|
||||
mirroring mechanism rather than about the content being mirrored.
|
||||
|
||||
**Generate the mirror only at release, from a tag or a release branch (rejected).** Technically
|
||||
supported, and it is worth recording *why* it was rejected rather than leaving it to look like an
|
||||
oversight. Claude Code marketplace entries accept ref-pinned git sources, and apm already emits that
|
||||
exact shape: the `mattpocock-skills` entry removed from root `apm.yml` on 2026-09-13 compiled to
|
||||
`{"source": "github", "repo": ..., "ref": "v1.2.3", "sha": ..., "tag_pattern": "v{version}"}` — a
|
||||
`git-subdir` source pinned to a ref. So a release-only mirror would install correctly.
|
||||
|
||||
Rejected on three grounds, stacking:
|
||||
|
||||
1. It keeps the 813-line script and the 1,289-line test alive in full. It reduces how often they
|
||||
run, not how much there is to maintain — and the maintenance, not the runtime, is what ADR-0017's
|
||||
amendment history shows to be the real cost.
|
||||
2. It requires per-package tagging discipline this repo does not practise. `git tag` lists
|
||||
repo-level tags (`v1.0.0`, `v2.0.0`, `v2.0.1`) matching no package version under the
|
||||
`per_package` versioning mode the six plugins use. The tagging convention that would make
|
||||
ref-pinning meaningful would have to be invented first.
|
||||
3. **There is no CI in this repo at all.** Every gate here is a git hook on a developer's machine.
|
||||
A release-time regeneration step would therefore depend on a human remembering to run it, and its
|
||||
failure mode is silent: a release tag whose tree contains a stale or absent mirror installs
|
||||
natively and reports zero skills, which is precisely the ADR-0017 defect, reintroduced on the
|
||||
release path where it is hardest to notice.
|
||||
|
||||
**Also delete the marketplace catalogue (proposed, then rejected on evidence).** The initial shape
|
||||
of this decision deleted `.claude-plugin/marketplace.json` along with everything else, on the
|
||||
reasoning that it is a Claude Code artifact. That is wrong. `apm marketplace add` probes
|
||||
`_MARKETPLACE_PATHS` in `apm_cli/marketplace/client.py` — `marketplace.json`, then
|
||||
`.github/plugin/marketplace.json`, then `.claude-plugin/marketplace.json`, first hit wins — and that
|
||||
read is what makes holocron an apm marketplace and what gives consumers the `<name>@holocron`
|
||||
short-name form. Deleting it would have broken apm consumers in order to remove a file whose format
|
||||
Claude Code merely happens to share.
|
||||
|
||||
Note the probe order: `.claude-plugin/marketplace.json` is the **last** resort, not the first, and
|
||||
the `.github/plugin/marketplace.json` this repo deleted earlier outranked it. That deletion was
|
||||
still inconsequential, but for a reason that has to be established rather than assumed — dropping a
|
||||
higher-priority candidate only demotes resolution to the next one, and a reviewer reproduced
|
||||
`apm marketplace add` against the post-deletion tree: it registered, found all 6 plugins, and
|
||||
resolved via `.claude-plugin/marketplace.json`. What would break is deleting the last candidate,
|
||||
which is exactly what this option proposed.
|
||||
|
||||
**Declare holocron an apm marketplace as a new step (moot).** Considered as a follow-on to the
|
||||
above, and found to be already done: the `marketplace:` block in root `apm.yml` *is* the
|
||||
declaration, and `apm marketplace init` produces exactly that block. There is nothing to add.
|
||||
|
||||
## Consequences
|
||||
|
||||
**1. Native `claude plugin install` no longer works, and the failure is silent.** This is accepted,
|
||||
not overlooked. Because apm reuses Claude Code's catalogue format by design — an APM-based
|
||||
marketplace stays consumable by Claude Code's existing marketplace mechanism — a Claude Code user
|
||||
can still register holocron natively, and will then install six plugins containing zero skills,
|
||||
zero agents and zero hooks. No error is raised at any point; the manifests are valid and the
|
||||
directories are simply empty. There is no schema change available that would prevent this, because
|
||||
the catalogue format cannot express "this marketplace is not for you" — the compatibility is
|
||||
structural, and it is the same compatibility that makes keeping the catalogue correct for apm. A
|
||||
README note is the only available mitigation, and a note is not a gate.
|
||||
|
||||
**2. Consumers now receive dev-fixture files.** apm installs from `.apm/`, and `.apm/` contains the
|
||||
per-skill `tests/` directories the mirror explicitly stripped (ADR-0017's depth-scoped
|
||||
`<category>/<name>/tests` exclusion). 10 `.bats` files across 6 skills therefore now deploy into
|
||||
every consumer's skill directories. Suppressing them would mean switching all six `apm.yml` files
|
||||
from `includes: auto` to explicit include lists — and an explicit list that is wrong silently drops
|
||||
content, which is the same failure class ADR-0017 was written to fix. Trading a cosmetic problem for
|
||||
a correctness problem is a bad trade, so this is deferred deliberately rather than fixed in passing.
|
||||
|
||||
**3. `tests/run-bats.sh` must exclude `.claude/skills/`** from both its `find` walk and the
|
||||
`git ls-files` set-equality check that derives the expected test list. Deployed `.bats` files are
|
||||
now discoverable in the install output and would otherwise be found and double-run against a root
|
||||
they do not belong to — exactly the `apm_modules/` problem ADR-0018 recorded, arriving by a second
|
||||
route. Any future script that walks this repo's tree needs both exclusions.
|
||||
|
||||
**4. No version bumps.** There is no standing rule that would require one. The patch-bump-on-content-
|
||||
change convention this repo once followed was ADR-0006's, and ADR-0015 explicitly retired it as a
|
||||
dual-manifest artifact: "Conventions that existed only because of hand-authored dual manifests
|
||||
(ADR-0006's version-parity/patch-bump rule …) are obsolete under `apm.yml`'s single-manifest model
|
||||
and were deliberately dropped." ADR-0015 also records that apm "has no native version-bump
|
||||
automation at all", so nothing mechanical demands one either. What remains is the substantive test,
|
||||
and it is satisfied independently: nothing under `.apm/` is touched here, only compiled artifacts are
|
||||
removed, so the content every apm consumer receives is byte-identical before and after. This also
|
||||
avoids triggering the `executables.allow`
|
||||
`kyberforge#<version>` pin cascade ADR-0019 describes, which would otherwise turn a cleanup into a
|
||||
multi-file coordinated edit for no functional gain.
|
||||
|
||||
**5. Reintroduction recipe.** This is the insurance that made the decision acceptable, so it is
|
||||
stated concretely rather than left as "it's in git". To restore native install support: recover
|
||||
`scripts/sync-plugin-content.sh` from git history (`git log --diff-filter=D -- scripts/sync-plugin-content.sh`
|
||||
finds the deleting commit; `git show <sha>^:scripts/sync-plugin-content.sh` recovers it) and re-run
|
||||
it with `--all`; it regenerates both the mirror and the per-plugin manifest pairs, because
|
||||
`apm pack --format plugin` produces them together. Separately, apm resolves a marketplace at a git
|
||||
ref — default `main`, with `--ref` pinning — so a consumer who pins an older ref still gets a tree
|
||||
containing the mirror and is unaffected until they move forward.
|
||||
|
||||
**6. A negative result, pinned so it is not re-litigated: this does not relax the self-containment
|
||||
constraint.** The natural next thought is that with the native installer gone, the no-cross-skill-
|
||||
file-sharing rule (the rule that forced ADR-0014's Vale config duplication) could be relaxed,
|
||||
because that rule was read as a property of Claude Code's plugin cache-install. It is not.
|
||||
`plugins/kyberforge/.apm/skills/skill-author/references/deployment-modes.md`, sourced from the
|
||||
agentskills.io spec, states the constraint independently for **APM package mode**: file references
|
||||
inside `.apm/skills/<name>/` must not reach outside that skill's own directory, and the spec defines
|
||||
no cross-skill sharing mechanism. So cross-skill file sharing remains impossible under the only
|
||||
install path that survives, and ADR-0014's duplication rationale stands unchanged.
|
||||
|
||||
**7. Three of ADR-0017's four amendments become moot, and one loses its enforcement.** Recorded
|
||||
because each was a decision someone spent real effort on:
|
||||
|
||||
- The `mcpServers` re-injection amendment (2026-08-14) is moot. Its target was
|
||||
`.github/plugin/plugin.json`, which no longer exists; `reinject_mcp_servers()` dies with the
|
||||
script that called it. Its reasoning — a path string, never an inlined object, because inlining
|
||||
bypasses apm's credential sanitizer — is worth carrying forward as a general rule if per-plugin
|
||||
Copilot manifests ever return.
|
||||
- The `hooks`-pointer amendment (2026-08-14) is moot in the same way, and its outcome was to change
|
||||
nothing, so nothing is lost.
|
||||
- The `hooks/hooks.json` path-correction amendment (2026-08-14) is moot: there is no mirrored hooks
|
||||
file to place.
|
||||
- The symlink amendment (2026-08-14) is **not** moot, and this is the one real regression.
|
||||
`check_apm_symlinks()` read the `.apm/` source tree directly to report symlinks, because apm
|
||||
filters them out silently and the resulting content loss is invisible to any mirror-versus-mirror
|
||||
diff. That check dies with the script.
|
||||
|
||||
A previous revision of this ADR left open whether `apm install`'s own copy path drops symlinks the
|
||||
way the bundle exporter does. It does, and the mechanism is now confirmed by reading the installed
|
||||
apm source. Deployment filters them: `ignore_non_content()` in `apm_cli/security/gate.py` is a
|
||||
`shutil.copytree` ignore callback whose docstring states "Excludes symlinks (security)", and
|
||||
`apm_cli/integration/skill_integrator.py` passes it — or the equivalent `_build_copy_ignore()` —
|
||||
to every `copytree` that materialises a skill (`:424`, `:791`, `:1152`), with further per-file
|
||||
`is_symlink()` drops for `bin/` entries and the plugin manifest at `:1671` and `:1696`. None of
|
||||
these log the skip. Materialization into `apm_modules/`, by contrast, **dereferences**: for git
|
||||
sources `deps/github_downloader.py` copies the checkout out with `robust_copytree`/`robust_copy2`
|
||||
and no symlink filter, and for local sources `install/phases/local_content.py`'s
|
||||
`_copy_tree_dereferencing_validated()` resolves each in-package symlink and copies its content as
|
||||
a real file (erroring on dangling, escaping or circular ones). So a symlink under `.apm/` survives
|
||||
into `apm_modules/` as real content and is then silently dropped on deploy — the failure is at the
|
||||
deploy step, not the fetch, which is why it would not show up in a cache inspection.
|
||||
|
||||
**This is accepted, and no replacement guard is added.** No symlink exists under any `.apm/`
|
||||
today, and the mitigations available are worse than the exposure: a standalone repo-side checker
|
||||
would be a new script to maintain for a condition that has never occurred, and it could only warn,
|
||||
since the drop happens inside apm. The operative rule is therefore a convention rather than a
|
||||
gate: do not introduce a symlink under any `plugins/*/.apm/` tree. If one is ever needed, the
|
||||
filtering above is the reason it will not reach consumers, and closing the gap properly means an
|
||||
upstream change in apm, not a local one.
|
||||
|
||||
**8. ADR-0017's own exit condition was different from this one, and that is worth noting.** Its
|
||||
final consequence anticipated deletion, but conditioned it on an upstream fix: "a future apm release
|
||||
that ships a native `.apm/`-aware plugin.json compiler ... would let `sync-plugin-content.sh` and
|
||||
its drift gate be deleted outright." That release has not happened. The mirror is being deleted
|
||||
because the path it bridges has no users, not because apm closed the gap — the gap is still open,
|
||||
and a consumer who installs natively still hits it. ADR-0017's anticipated exit remains available
|
||||
and unclaimed; this ADR takes a different one.
|
||||
@@ -53,8 +53,7 @@ All skills — new and rebuilt — must follow this standard:
|
||||
- `name:` — matches directory name
|
||||
- `description:` — trigger-tested before writing the body (explicit, implicit, negative cases)
|
||||
- `metadata: category:` — from the category table above
|
||||
|
||||
`version:`, `updated:`, `when:`, `source:`, and `references:` are provenance/audit fields — they live in `META.md` alongside the SKILL.md (not in frontmatter). See `META-TEMPLATE.md` in `.agents/skills/write-skill/` for the META.md schema.
|
||||
- `metadata: version:` — mandatory for every skill (ADR-0022)
|
||||
|
||||
**Body required sections:**
|
||||
- Constraints (highest-ROI element — prevents overengineering)
|
||||
|
||||
@@ -103,21 +103,21 @@ Do not write the SKILL.md until the human has confirmed every section. The synth
|
||||
**c. SKILL.md** (sub-agent)
|
||||
Once all sections are confirmed, spawn a write agent to produce the SKILL.md using `write-skill` (or hand-write for bootstrap skills). The agent receives: trigger description, per-section decisions from step b, upstream content to incorporate, authoring standard (see below).
|
||||
|
||||
**c. META.md — `source:` and `references:` fields**
|
||||
Populate `META.md` after upstream review. Two distinct fields:
|
||||
- `source:` — upstream provenance tracking (repo slug, commit SHA, files adopted with inline comments, updated date). Present only if content was adopted. Absence = self-authored.
|
||||
- `references:` — general citations (research papers, documentation, standard specifications). Present only if the skill cites external research.
|
||||
**d. Provenance — source and reference records**
|
||||
Record provenance after upstream review. Two distinct kinds:
|
||||
- Upstream provenance (repo slug, commit SHA, files adopted with inline comments, updated date). Present only if content was adopted. Absence = self-authored.
|
||||
- General citations (research papers, documentation, standard specifications). Present only if the skill cites external research.
|
||||
|
||||
Both fields live in `META.md` alongside the SKILL.md — not in frontmatter. See `META-TEMPLATE.md` in `.agents/skills/write-skill/` for the full schema.
|
||||
Both are recorded in the skill's own `references/sources.md`, keyed by the `source_keys:` its SKILL.md and reference files declare. `validate-provenance.sh` checks that chain.
|
||||
|
||||
**d. eval.yaml** (sub-agent)
|
||||
**e. eval.yaml** (sub-agent)
|
||||
Invoke `write-eval` in two steps to preserve its confirmation gate:
|
||||
1. Sub-agent proposes test cases and returns the plan to the main conversation.
|
||||
2. Human confirms the plan; then sub-agent writes the file.
|
||||
|
||||
Do not pass pre-designed test cases directly to a write agent — that collapses the plan-then-confirm gate into a single step, bypassing write-eval's own constraint. Co-located at `.agents/evals/<category>/<skill-name>/eval.yaml`. Must contain all five required test types (see Eval schema below).
|
||||
|
||||
**e. HITL behavioral test**
|
||||
**f. HITL behavioral test**
|
||||
Human opens a fresh Claude session, invokes the skill with its trigger phrase, and verifies output. Do not batch more than 2–3 skills before running behavioral tests — output volume must stay within genuine human review capacity. An approval that cannot be meaningfully evaluated is not an approval.
|
||||
|
||||
### Step 6 — Session handoff
|
||||
@@ -157,12 +157,11 @@ name: skill-name
|
||||
description: <trigger description — routing only; written and tested first; max 1024 chars>
|
||||
metadata:
|
||||
category: <design|factory|implement|test|review|deploy|operate|cross-cutting|iac>
|
||||
version: <semver — mandatory for every skill; see ADR-0022>
|
||||
# allowed-tools: <add only when the skill has a narrow, well-defined tool surface; omit otherwise>
|
||||
---
|
||||
```
|
||||
|
||||
Frontmatter contains only these fields. `version`, `updated`, `when`, `source`, and `references` are provenance/audit fields — they are not used for routing or runtime execution. They live in `META.md` alongside the SKILL.md, loaded only when needed. See `META-TEMPLATE.md` in `.agents/skills/write-skill/` for the META.md schema.
|
||||
|
||||
### Body sections
|
||||
|
||||
Use `.agents/skills/write-skill/SKILL-TEMPLATE.md` as the authoritative structure reference. The template defines the required sections, correct order, XML grouping, and placeholder comments for each section.
|
||||
@@ -230,6 +229,6 @@ Upstream review happens per-skill during step 2, not once at chunk start.
|
||||
|
||||
## Open decisions carried forward
|
||||
|
||||
- **Bidirectional reference convention** — Chunk 4 (reference scanner tooling; reverse map "what files point to X?"). The `when:` field itself is resolved — it lives in `META.md` alongside every skill.
|
||||
- **Bidirectional reference convention** — Chunk 4 (reference scanner tooling; reverse map "what files point to X?").
|
||||
- **PRD/issue template scope** — refined during `write-prd` (0020) and `write-issue-spec` (0019) implementation
|
||||
- **Merging `zoom-out` into architect role** — revisit at Chunk 5 grill
|
||||
|
||||
@@ -19,32 +19,65 @@ project repo (local overrides)
|
||||
- **Executables** (`DEPLOY_EXECUTABLES`): `providers/claude-code/statusline-command.sh` → `~/.claude/statusline-command.sh` (with `+x`)
|
||||
- **Directories** (`DEPLOY_DIRS`): `core/` → `~/.claude/core/` (destination fully replaced on each deploy)
|
||||
|
||||
Skills are **not** deployed by `install.sh`. They are distributed as plugins and installed separately via `claude plugin install <name>@holocron`.
|
||||
Skills are **not** deployed by `install.sh`. They are distributed as plugins and installed separately — in this repo by `apm install` against the `dependencies.apm` entries in the root `apm.yml`, which lands them in `.claude/skills/` and `.claude/agents/` (ADR-0018). A consuming repo installs them the same way — apm is the only supported install path.
|
||||
|
||||
`~/.claude/CLAUDE.md` is a thin adapter, not a content source. It imports `~/.agents/AGENTS.md` (always-on rules) and `governance.md` (always-on governance), then lists the content index. All always-on content lives in `AGENTS.md` files so other providers can import the same source without duplication.
|
||||
`~/.claude/CLAUDE.md` is a thin adapter, not a content source. It imports `~/.agents/AGENTS.md` (always-on rules) and `governance.md` (always-on governance) and carries nothing else — the content index of on-demand instruction files sits in `core/AGENTS.md`, deployed to `~/.agents/AGENTS.md` and imported by it. All always-on content lives in `AGENTS.md` files so other providers can import the same source without duplication.
|
||||
|
||||
## Plugin model
|
||||
|
||||
Skills, agents, MCP servers, and hooks are distributed as self-contained plugin units under `plugins/`. Each plugin has a `plugin.json` manifest and is installed independently via `claude plugin install`.
|
||||
Skills, agents, MCP servers, and hooks are distributed as self-contained plugin units under `plugins/`, installed independently via `apm install`, here and in any consuming repo (ADR-0018). Self-contained is a hard constraint, not a description: a plugin is copied to a cache on install, so nothing inside it may reference a file outside its own directory. That is why the Vale styles are duplicated across two skills rather than shared (ADR-0014), and why ADR-0020's constants are copied into three validators rather than sourced from one. Each plugin is an **apm package**: `plugins/<name>/apm.yml` plus a hand-authored `plugins/<name>/.apm/{skills,agents,hooks,commands,instructions,extensions}/` tree (ADR-0015). There is no per-plugin `plugin.json` at all — apm reads `apm.yml`, and the repo's one generated manifest, `.claude-plugin/marketplace.json`, is compiled from that source.
|
||||
|
||||
Which plugin a new skill belongs in follows from what each one is scoped to. The boundary that matters most in practice is `core` vs `kyberforge`: `core` is the home for cross-cutting, repo-agnostic utility skills that a consumer would want against *their* repo, while `kyberforge` is meta-tooling for the holocron marketplace itself. A skill that authors a target repo's `AGENTS.md` is `core`; a skill that audits a `SKILL.md` against this marketplace's contract is `kyberforge`.
|
||||
|
||||
The second boundary worth stating is `git` vs `gitea`, because both own things called branches and both touch pull requests: `git` is whatever works over the git wire protocol against a local clone, `gitea` is whatever goes through the forge's HTTP API. That is why `git-branches` and `gitea-branches` both exist and are not duplicates.
|
||||
|
||||
These are routing boundaries, not inventories — they answer "where does a new skill go", so they deliberately do not enumerate what each plugin ships today. The plugin's published `description` in its `apm.yml` states the same boundary for a consumer deciding whether to install (ADR-0021); neither carries an inventory. For what a plugin ships today, read `plugins/<name>/.apm/skills/` or the plugin list in `README.md`.
|
||||
|
||||
| Plugin | Scope |
|
||||
|---|---|
|
||||
| `core` | Authoring and auditing a repo's `AGENTS.md` and the provider adapter files that defer to it |
|
||||
| `git` | Git operations and git hook tooling — anything driven over the git wire protocol against a local clone, plus the pre-commit hooks that guard it |
|
||||
| `gitea` | Anything reached through the Gitea HTTP API rather than the git wire protocol — the forge's own objects |
|
||||
| `kyberforge` | Creating and maintaining a Claude Code / Copilot CLI plugin marketplace — this repo's own meta-tooling |
|
||||
| `lint` | Configuring and running linters against a target repo; repo-agnostic, first linter is Vale |
|
||||
| `bin` | Unsorted skills that have not earned a home yet |
|
||||
|
||||
One compiler produces the generated content in the tree:
|
||||
|
||||
- **`apm pack` compiles the marketplace manifest** (ADR-0015). Repo-wide, from the root `apm.yml`'s `marketplace:` block: `.claude-plugin/marketplace.json` (apm's `claude` output profile) — the only manifest this repo generates or ships. Copilot CLI checks for a marketplace manifest at several conventional paths, falling back through `.github/plugin/marketplace.json` to `.claude-plugin/marketplace.json` — since this repo already generates the latter, no dedicated Copilot-path mirror is maintained.
|
||||
|
||||
apm is the only supported install path. A flat `skills/`, `agents/`, `hooks/` mirror used to be compiled to each plugin root so Claude Code's installer could convention-scan it, alongside a per-plugin `.claude-plugin/plugin.json` and `.github/plugin/plugin.json`; both are gone, together with native `claude plugin install` support. apm reads `plugins/<name>/apm.yml` and deploys from `.apm/` directly, and never probed those manifests.
|
||||
|
||||
`.apm/` is the sole hand-edited authoring source for plugin content. Hand-authored material that is not an `.apm/` primitive — `README.md`, `docs/`, `bin/`, `sources.md`, and per-plugin extras such as `plugins/gitea/references/` and `plugins/bin/evals/` — lives at the plugin **root**. A hand-edit to the generated `.claude-plugin/marketplace.json` is reported as drift by the `apm-pack-check-clean` pre-push hook.
|
||||
|
||||
Plugin-root documentation belongs in `docs/`. That convention is older than the mirror's removal: a hand-written `README.md` placed inside a mirrored directory used to be destroyed by the next sync with no drift report, which cost the repo one document — `plugins/kyberforge/hooks/README.md`, since restored to `plugins/kyberforge/docs/hooks.md`.
|
||||
|
||||
## Governance layer
|
||||
|
||||
`core/instructions/governance.md` is the always-on governance instruction file. Unlike the on-demand instruction files in the content index, governance.md is loaded into every Claude session via `@import` in `providers/claude-code/CLAUDE.md`. This is a technical guarantee, not a behavioural instruction — `@import` causes Claude Code to expand and load the file at launch, before any interaction begins.
|
||||
|
||||
Those on-demand files are plain markdown — no frontmatter, no schema. The agent decides when to read each one from task context and the content index label alone. Frontmatter is deferred until there is evidence that agents are loading the wrong files in practice; it is a deliberate deferral, not an oversight to close.
|
||||
|
||||
The governance layer has two phases:
|
||||
- **Phase 1** (complete): instruction and documentation layer — `governance.md` loaded via `@import`; `docs/ai-constitution.md` and `docs/wiki/HUMANS.md` as human-facing reference; `CONTEXT.md` extended with governance domain language.
|
||||
- **Phase 2** (planned): deterministic enforcement layer — pre-commit hooks, CI gates, secret scanning, licence scanning. Specified in `docs/research/governance_principles/CONTROLS.md`.
|
||||
|
||||
## AGENTS.md pattern
|
||||
|
||||
This repo uses two `AGENTS.md` files as the provider-agnostic source of always-on rules (ADR-0012):
|
||||
This repo uses two `AGENTS.md` files as the provider-agnostic source of always-on rules (ADR-0003):
|
||||
|
||||
- **Repo-level `AGENTS.md`** — instructions for agents working inside this repo (structure, key rules). Imported by repo `CLAUDE.md` via `@AGENTS.md`.
|
||||
- **Global `core/AGENTS.md`** — Communication and Behavior rules that apply across all projects. Deployed to `~/.agents/AGENTS.md`; imported by `~/.claude/CLAUDE.md` via `@~/.agents/AGENTS.md`.
|
||||
|
||||
Both `CLAUDE.md` files are thin adapters: they import from their respective `AGENTS.md` and add only Claude Code-specific syntax (`@import`, content index paths). They carry no original always-on content.
|
||||
|
||||
This repo also has a `CLAUDE.md` at its root — the Claude Code entry point for working in this repo. It imports `AGENTS.md` and `CONTEXT.md`, nothing more. This is distinct from `providers/claude-code/CLAUDE.md`, which is the global config deployed to `~/.claude/`.
|
||||
This repo also has a `CLAUDE.md` at its root — the Claude Code entry point for working in this repo. It imports `AGENTS.md` and nothing else; there is no `@CONTEXT.md` import. It is not import-only either: below the import sits a fenced `<!-- rtk-instructions v2 -->` … `<!-- /rtk-instructions -->` block carrying the RTK command-prefix convention, which is tool-specific content with no `AGENTS.md` source. This is distinct from `providers/claude-code/CLAUDE.md`, which is the global config deployed to `~/.claude/`.
|
||||
|
||||
`CONTEXT.md` is therefore **not** always-loaded. `AGENTS.md` instructs agents to read it at session start, which is a behavioural instruction, not an `@import` guarantee — `LESSONS.md`'s 2026-05-17 entry proposed adding the import and it was never applied. Treat that entry as open work rather than a record of a landed change.
|
||||
|
||||
## Reference conventions
|
||||
|
||||
The stated convention is that files referencing other files declare those references explicitly: the referencing file carries the forward reference (the content index in `core/AGENTS.md`, `references:` in frontmatter), the referenced file carries a `when:` field describing when it is loaded, and divergence between the two signals staleness. It is aspirational, not a description of the repo today — no file under `core/instructions/` carries frontmatter at all, `when:` appears in exactly one of the 39 `SKILL.md` sources under `plugins/*/.apm/skills/`, and the reference scanner script meant to derive the reverse map ("what files reference this file?") does not exist; `docs/notes/skill-implementation-workflow.md` still lists it as unbuilt work. Treat it as intent for instruction files, skills, and workflow documents, not as a rule the repo enforces.
|
||||
|
||||
## Provider model
|
||||
|
||||
@@ -52,4 +85,4 @@ This repo also has a `CLAUDE.md` at its root — the Claude Code entry point for
|
||||
|
||||
## Architectural decisions
|
||||
|
||||
Key hard-to-reverse decisions are recorded as ADRs in `docs/adr/`. See the index there for rationale on choices like the pull distribution model, copy-not-symlink coupling, and the two-tier CLAUDE.md structure.
|
||||
Key hard-to-reverse decisions are recorded as ADRs in `docs/adr/`. There is no index file — the directory holds numbered ADRs whose filenames state their decision, so `ls docs/adr/` is the index. Read a superseding ADR before the one it supersedes: ADR-0015 (apm as the authoring source of truth) supersedes ADR-0001 and moots ADR-0006, ADR-0017 corrects ADR-0015's host-discovery gap, and ADR-0019 supersedes one claim in ADR-0018 (that `.claude/settings.json`'s committed content is exactly `{"hooks": {}}`) while keeping the rule behind it. Entry points for the structure described on this page: ADR-0002 (two-tier CLAUDE.md), ADR-0003 (AGENTS.md as the provider-agnostic entry point), ADR-0015 and ADR-0017 (the two compilers behind the plugin roots).
|
||||
|
||||
908
docs/spec/gates.md
Normal file
908
docs/spec/gates.md
Normal file
@@ -0,0 +1,908 @@
|
||||
# Enforcement gates
|
||||
|
||||
Reference for this repo's pre-commit and pre-push hooks: what each one guards, what its numbers
|
||||
mean, and which shapes were tried and rejected. Read it when a gate fails, before changing anything
|
||||
in `.pre-commit-config.yaml`, or before "fixing" something that looks like an inconsistency — several
|
||||
of the oddities documented here are load-bearing and have already been re-litigated once.
|
||||
|
||||
`AGENTS.md` carries only the operative rules an agent needs in the moment. The reasoning lives here.
|
||||
|
||||
---
|
||||
|
||||
## Running the gates
|
||||
|
||||
| Command | Scope |
|
||||
|---|---|
|
||||
| `pre-commit run --all-files` | the commit-stage hooks |
|
||||
| `pre-commit run --hook-stage pre-push --all-files` | the push gate, one command — with one caveat below |
|
||||
| `pre-commit run skill-size-check --all-files` | just the ADR-0020 size/context gates |
|
||||
|
||||
Install hooks via `pc-run`, wiring **all three stages**. This repo's `.pre-commit-config.yaml` has no
|
||||
`default_install_hook_types`, so a plain install silently skips `commit-msg` (Conventional Commits)
|
||||
and `pre-push` (everything below).
|
||||
|
||||
The pre-push command reports **11** hooks, not 9. The extra two are pre-commit's own `meta` hooks,
|
||||
`check-hooks-apply` and `check-useless-excludes`: they declare no `stages:`, so they run at every
|
||||
stage including this one. Both are declared in this repo's `.pre-commit-config.yaml` like everything
|
||||
else — what separates them is `repo: meta` (pre-commit's own built-ins) from `repo: local`. Nine
|
||||
is the count of hooks this repo authors itself.
|
||||
|
||||
**The caveat: one of those 9 is a silent no-op under that invocation.**
|
||||
`check-release-needed` exits 0 immediately unless `PRE_COMMIT_REMOTE_BRANCH` equals
|
||||
`refs/heads/main`, and pre-commit exports that variable only from the real pre-push git hook during
|
||||
an actual `git push`. Running the stage by hand — or from a CI runner — therefore reports it
|
||||
`Passed` having checked nothing. That is by design for feature branches — pushing WIP must not be
|
||||
blocked on cutting a premature tag — but it means `--hook-stage pre-push --all-files` is a full
|
||||
rehearsal of 8 hooks and a skip of the ninth. The script's own header records the same gap for
|
||||
a PR merged through Gitea's merge button, where no local push happens at all.
|
||||
|
||||
## The pre-push gate
|
||||
|
||||
Nine hooks, grouped below by what they guard rather than by the order `.pre-commit-config.yaml` declares them in.
|
||||
|
||||
**Core checks**
|
||||
|
||||
| Hook | Guards |
|
||||
|---|---|
|
||||
| `run-tests` | `bash tests/run-tests.sh --strict` — the whole suite, skips fatal (see [Tests](#tests)) |
|
||||
|
||||
**Generated-content drift gates**
|
||||
|
||||
| Hook | Guards |
|
||||
|---|---|
|
||||
| `check-vale-style-sync` | skill-audit's Vale copy matches agent-audit's canonical copy, plus six glob-coverage probes (see [Vale](#vale)) |
|
||||
| `check-scope-walkup-sync` | `validate.sh`, `validate-provenance.sh`, `new-agent.sh` and `new-skill.sh`'s four independent `$HOME`/`.git`/`apm.yml` walk-up ports still agree behaviorally |
|
||||
| `check-executables-allow-sync` | root `apm.yml`'s `executables.allow` key names kyberforge's actual version (see [apm gates](#apm-gates)) |
|
||||
|
||||
`check-executables-allow-sync` is the odd one in this group: it guards a *silent failure* rather than
|
||||
drift in generated text.
|
||||
|
||||
**Artifact validators**
|
||||
|
||||
| Hook | Guards |
|
||||
|---|---|
|
||||
| `check-apm-agents-valid` | runs agent-audit's `validate.sh` over every real `plugins/*/.apm/agents/*.agent.md` (see [Agent files](#agent-files-take-the-description-gates-not-the-body-gate)) |
|
||||
|
||||
**apm's own gates**
|
||||
|
||||
| Hook | Guards |
|
||||
|---|---|
|
||||
| `apm-audit-ci` | `apm audit --ci` once per manifest — root plus each of the six plugin packages |
|
||||
| `apm-pack-check-clean` | `apm pack --check-versions --check-clean --dry-run` — the compiled marketplace still matches what `apm.yml` + `.apm/` would generate, and per-package versions agree with the `per_package` strategy |
|
||||
|
||||
**Host validators** (needs the `claude` CLI on PATH)
|
||||
|
||||
| Hook | Guards |
|
||||
|---|---|
|
||||
| `validate-marketplace` | `claude plugin validate --strict` on the root marketplace manifest |
|
||||
|
||||
**Release**
|
||||
|
||||
| Hook | Guards |
|
||||
|---|---|
|
||||
| `check-release-needed` | on a real `git push` to `main` only — fails if files exposed via `.pre-commit-hooks.yaml` changed since the last tag. A no-op everywhere else, including under `pre-commit run --hook-stage pre-push` (see [the caveat above](#running-the-gates)) |
|
||||
|
||||
Two of these shell out to `apm`: `apm-audit-ci` and `apm-pack-check-clean`. The second is a bare
|
||||
`apm …` entry and the first is a `bash -c` loop calling `apm` once per package, so without the CLI
|
||||
the push dies with an unhelpful "command not found". Install with `apm-install`, or
|
||||
`curl -sSL https://aka.ms/apm-unix | sh`; verify with `apm --version`.
|
||||
|
||||
## Skill and agent context gates (ADR-0020)
|
||||
|
||||
The `skill-size-check` pre-commit hook, scoped to `^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$`,
|
||||
runs `scripts/skill-size-check.sh`. It is also shipped to external repos as
|
||||
`kyberforge-skill-size-check` (see
|
||||
[External consumers](#external-consumers-the-root-pre-commit-hooksyaml)). Besides the ADR-0020
|
||||
gates below, it also asserts required frontmatter is present: `name`, a non-empty `description`, and
|
||||
a `metadata.version` matching three-part semver (`1.0.0`) — folded in from a formerly standalone
|
||||
`skill-frontmatter` hook that parsed the same fields with a shell script.
|
||||
|
||||
**Two things fall outside that scope, both deliberately.** The `[^/]+/SKILL\.md$` tail admits only a
|
||||
`SKILL.md` sitting directly in a skill directory under `.apm/skills/`:
|
||||
|
||||
- the `plugins/kyberforge/docs/research/examples/` reference skills, which are vendored upstream
|
||||
corpus and not this repo's to gate;
|
||||
- `plugins/kyberforge/.apm/skills/skill-author/assets/templates/SKILL.md` — inside `.apm/skills/`,
|
||||
but two directories deeper. It is the `FILL IN:` scaffold `skill-author` copies, so its
|
||||
`description: >` is a comment block rather than a description and every ADR-0020 measurement over
|
||||
it would be meaningless. A reader adjusting the pattern needs to know it is there.
|
||||
|
||||
Everything else it matches exactly, with nothing over- or under-caught. Re-derive both halves:
|
||||
|
||||
```
|
||||
git ls-files | grep -cE '^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$' # the real skills
|
||||
git ls-files | grep -E '^plugins/[^/]+/\.apm/skills/.*SKILL\.md$' \
|
||||
| grep -vE '^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$' # the scaffold only
|
||||
```
|
||||
|
||||
The first count equals the number of skill directories (`ls -d plugins/*/.apm/skills/*/ | wc -l`);
|
||||
the second returns exactly the template. Every other tracked `SKILL.md` in the tree is one of the
|
||||
four vendored `plugins/kyberforge/docs/research/examples/skill-write/` corpus files, excluded by the
|
||||
`.apm/skills/` segment — the first of the two deliberate exclusions above.
|
||||
|
||||
### Two independent gate families, neither replaced the other
|
||||
|
||||
**Family 1 — agentskills.io spec backstop** (unchanged, conformance not quality):
|
||||
|
||||
| Constant | Value | Measured over |
|
||||
|---|---|---|
|
||||
| `MAX_LINES` | 500 | whole file, **frontmatter included** |
|
||||
| `MAX_WORDS` | 2,770 | whole file, **frontmatter included** |
|
||||
|
||||
**Family 2 — ADR-0020 context budget** (measured differently, on purpose):
|
||||
|
||||
| Check | SUGGESTION | FAIL | Measured over |
|
||||
|---|---|---|---|
|
||||
| `description` characters | 250 | 400 | the YAML-**folded** value |
|
||||
| body words | 600 | 900 | **body only** — everything after the frontmatter's closing `---` |
|
||||
|
||||
Plus two hard FAILs with no suggestion tier:
|
||||
|
||||
- **A missing, valueless or `null` `description:`.** Not a skip. The description is the one field
|
||||
preloaded into every session, so a gate that declines to measure it reports green. (This is not
|
||||
hypothetical: `description:` with no value followed by `model: sonnet` let a line regex capture the
|
||||
*next* key, which looked non-empty, so the "missing or empty" branch never fired and every gate
|
||||
below early-returned on the genuinely empty folded value — exit 0, zero output, on a blocking gate.)
|
||||
- **Every `references/<file>.md` a body names must exist** on disk. A dispatch table pointing at a
|
||||
file that was never written is a silently dead branch, and nothing else in the gate/audit/vale
|
||||
stack notices it.
|
||||
|
||||
A file can sit well inside one family and fail the other. 2,770 whole-file words is a conformance
|
||||
backstop; 900 body-only words is a quality gate. Conflating them is what produced the current state.
|
||||
|
||||
### An unresolved routing target is not automatically a FAIL
|
||||
|
||||
A boundary-clause target that resolves to no skill or agent has **three** possible verdicts, not one
|
||||
(`unresolved_targets()` in `scripts/skill-size-check.sh`):
|
||||
|
||||
| Verdict | When |
|
||||
|---|---|
|
||||
| **SUGGESTION** — the default | the target does not resolve and neither promotion condition below holds |
|
||||
| **blocking ERROR** | the target is written in **route notation** — `/name` for any name, or any arrow form (a bare `-> name` only when the name is hyphenated, a backticked `` -> `name` `` for any — see the gap below); **or** it is a bare **terminal** name (not a compound modifier) **corroborated** by another target in the same sentence that *does* resolve |
|
||||
| **INFO, "DID NOT RUN"** | no skill universe could be determined for the path at all — the targets are named and left unchecked, exit 0 |
|
||||
|
||||
The default is deliberately soft because a hyphenated word in a boundary clause is as likely to be a
|
||||
tool, a file format or an English compound as a route: "pre-commit hooks" is prose about a tool and
|
||||
never reaches the check at all, being a compound modifier rather than a terminal name. The
|
||||
SUGGESTION text says how to opt in — write it as `/name` or `-> name` and it gets checked properly.
|
||||
|
||||
**The two promotion conditions are not symmetric, and the order matters.** `_add()` decides
|
||||
**notation first**: when the name is written `/name`, or reached through any arrow form, the target
|
||||
is marked error-eligible there and the terminal test is never run. Terminality gates only the *bare*
|
||||
path — a name in prose earns its error from corroboration, and a compound modifier can never dangle.
|
||||
Reading the row as "terminal AND (notation OR corroborated)" gets the notation half backwards: it
|
||||
predicts that `` … Do not use for Y — use /no-such-skill afterwards. `` is a SUGGESTION, because
|
||||
`afterwards` is a follower outside `FOLLOWER_OK`. It exits 1. That was the defect — `-> name` reached
|
||||
`_add()` with `strict=True` from both its call sites and `/name` did not, so the one spelling
|
||||
ADR-0020 offers an author who wants a route checked unconditionally was the one spelling a stray
|
||||
follower could silence.
|
||||
|
||||
**Known gap: a BARE arrow target must be hyphenated.** Target extraction is built on `NAME_HYPH` in
|
||||
`scripts/skill-size-check.sh`, which requires at least one hyphen, and `ARROW_BOUNDARY` inherits
|
||||
that. So `Not X -> gitea-prs` is extracted and checked, while `Not X -> triage` yields no target.
|
||||
The exclusion is deliberate, not an oversight: `research`, `triage`, `forge`, `prototype` and `tdd`
|
||||
are all real skill names *and* ordinary English, so a bare single-word rule would flag most of the
|
||||
corpus. The marked spellings carry no such restriction — `` `triage` `` and `/triage` are both
|
||||
extracted — and are the forms to prefer. **Both arrow spellings are recognised:** `ARROW_MARKED`,
|
||||
`ARROW_BOUNDARY` and `BOUNDARY_ARROW` are each built from `(?:->|→)`, so the unicode arrow `→`
|
||||
behaves exactly like `->` in every case below. Cite these constants by symbol name, never by line
|
||||
number: the script moves often enough that a pinned line lands a reader in an unrelated comment
|
||||
block and reads as plausible.
|
||||
|
||||
**The gap is no longer silent.** It used to be exactly that — no ERROR, no SUGGESTION, exit 0 — which
|
||||
made the dangling-target SUGGESTION's own advice unsafe for a single-word skill: taking it silenced
|
||||
the finding instead of checking it. `boundary_clause_status()` now separates the case out and
|
||||
reports it as `unparsed` (see below), naming the parse failure and the two spellings that fix it.
|
||||
The target is still not *resolved*; the author is now told so rather than left with a green gate.
|
||||
`tests/test-adr0020-targets.sh` covers both directions (`arrow-single-word-target` and the silent
|
||||
control `arrow-single-word-marked`).
|
||||
|
||||
Corroboration is what makes the soft default safe: a sentence whose *other* target resolves is
|
||||
demonstrably a routing sentence, so a sibling that does not resolve is a typo rather than a noun, and
|
||||
gets promoted.
|
||||
|
||||
### Target resolution walk
|
||||
|
||||
Resolution walks up **from the file being checked** — never from the script's own location. Deriving
|
||||
it from `${BASH_SOURCE}` leaked holocron's 39-skill universe into every consumer repo running the
|
||||
hook through pre-commit, so a consumer skill routing to `skill-audit` resolved against a plugin it
|
||||
had never installed.
|
||||
|
||||
The walk finds an **authoring root**: the nearest ancestor holding `plugins/*/.apm/skills` or
|
||||
`plugins/*/.apm/agents`, falling back to the nearest ancestor holding `.git`. **Two passes, not one
|
||||
interleaved walk**, so a nested `.git` (a submodule, a sub-package worktree) cannot beat a real
|
||||
monorepo root further up.
|
||||
|
||||
The universe is then:
|
||||
|
||||
1. every skill and agent under `<root>/plugins/*/` — sibling plugins resolve, which is what a
|
||||
monorepo means;
|
||||
2. the checked file's own apm package;
|
||||
3. the packages that package declares in **its own** `apm.yml` `dependencies.apm`.
|
||||
|
||||
The **root** manifest's `dependencies:` block is not read, and no plugin here declares a cross-plugin
|
||||
apm dependency — none needs to.
|
||||
|
||||
Deployed `.claude/` / `.agents/` trees are consulted **only** when the walk found no plugin monorepo
|
||||
root, whether it landed on a bare `.git` ancestor or on nothing at all. That is the consumer case.
|
||||
|
||||
**The gate keys on which of the two passes matched, never on whether the root contributed a new
|
||||
name.** A name-count delta looks equivalent and is not: `_collect_authoring_root()` re-collects the
|
||||
checked file's own plugin, whose names the earlier steps already added, so a single-plugin monorepo
|
||||
shows a delta of zero and would wrongly reach for the deployed trees — including the user's global
|
||||
`~/.claude/skills`, making the verdict depend on what happens to be installed.
|
||||
|
||||
Why it matters: those trees are gitignored `apm install` output, present only on a machine that has
|
||||
run it. Four cross-plugin targets here (`gitea-branches` → `git-branches`, `gitea-branches` →
|
||||
`git-history`, `gitea-issues` → `git-branches`, `gitea-workflow` → `git-workflow`) once resolved
|
||||
through `.claude/skills/` alone, so **the same commit measured 2 dangling targets on a developer
|
||||
machine and 6 on a fresh clone**. A gate shipping hot with no baseline cannot give two answers.
|
||||
|
||||
Verified: running the hook over a tree holding only `plugins/` and the root `apm.yml`, with no
|
||||
`.claude/` or `.agents/` anywhere, produces findings identical to the working tree — confirming the
|
||||
two trees agree on the current corpus, independent of what happens to be installed locally.
|
||||
|
||||
### Boundary-clause detection: three outcomes, not two
|
||||
|
||||
`boundary_clause_status()` returns one of three values, and the two findings get separate messages:
|
||||
|
||||
| Status | When | Reported as |
|
||||
|---|---|---|
|
||||
| `present` | a prose marker (`do not`, `instead`, `rather than`, `not for`) or an arrow clause was found | nothing |
|
||||
| `absent` | neither was found | SUGGESTION: add a boundary clause, in either form |
|
||||
| `unparsed` | an arrow clause was found and **no target could be read out of it** | SUGGESTION: the clause is present — this is a *parse* failure, not a missing clause |
|
||||
|
||||
The third had to be split out. Collapsing it into `absent` is a **wrong** finding, not a strict one:
|
||||
it sends the author to add a clause that is already there. Three of them instead reworded a correct
|
||||
clause until the regex accepted it, one stripping the very filename that discriminates the skill
|
||||
from its neighbour (**#110**).
|
||||
|
||||
`unparsed` is narrow and certain on purpose. It fires only on the arrow form, which *always* names a
|
||||
target, so zero targets means the name is written in a shape the extractor cannot see — in practice
|
||||
a bare single-word target, per the known gap above, and the message says to write it `` `name` `` or
|
||||
`/name`. A **prose** clause yielding no target is not reported at all: "Do not use for anything else"
|
||||
is a complete and legitimate boundary clause that names nowhere to go.
|
||||
|
||||
**One arrow, one target.** An arrow clause naming two or more targets draws its own SUGGESTION,
|
||||
quoting both names and asking for a split, because only the first is ever resolved: the conjunction
|
||||
continuation (`CONT_MARKED` / `CONT_ANY`) is wired to the prose route verbs and never to arrows. So
|
||||
`Not X -> a or b` resolved `a`, left `b` resolved by nothing and reported by nothing, and then let
|
||||
the audit print "1 of 1 boundary target(s) resolve" on a clause naming two — a gate under-reporting
|
||||
its own coverage, which is the one failure mode ADR-0020 says a gate must not have (**#107**). The
|
||||
clause is **rejected rather than the arrow scan extended**: extending it would widen the resolver's
|
||||
deliberately conservative false-positive tuning across every arrow in the corpus, where splitting
|
||||
costs the author one full stop. The convention is one arrow per target — `Not X -> a. Not Y -> b.` —
|
||||
already what every retrofitted `gitea-*` skill does in practice, now stated in
|
||||
`skill-author`'s `references/contract.md` instead of being folklore.
|
||||
|
||||
**Dotted filenames in a boundary clause now parse.** `CLAUSE_BODY` — what may sit between `Not` and
|
||||
the arrow — used to be `[^.;]`, a class that cannot cross a `.`, so every clause naming a dotted
|
||||
filename between the two (`AGENTS.md`, `.vale.ini`, `.pre-commit-config.yaml`) was invisible to both
|
||||
`BOUNDARY_ARROW` and `ARROW_BOUNDARY`. The two resulting failures were different sizes (**#110**):
|
||||
|
||||
- with a **backticked** target the clause was *misdiagnosed*. The backtick sweep still extracted the
|
||||
target, so the route was checked, but the gate reported "no boundary clause" on a clause that was
|
||||
present and working. That is the misdiagnosis the three rewordings above came from.
|
||||
- with a **bare** target the clause was *unchecked*. `ARROW_BOUNDARY` is the only extractor for a
|
||||
bare arrow target, so `Not AGENTS.md -> no-such-skill` produced no target, no dangling report and
|
||||
no missing-clause SUGGESTION. Silence, not noise — the worse of the two.
|
||||
|
||||
`CLAUSE_BODY` is now `(?:[^.;]|\.(?=\S))`: a dot inside a filename is followed by a non-space, a
|
||||
sentence-ending dot by whitespace or end of string, so the class crosses `AGENTS.md` and still stops
|
||||
at a real sentence end. **Read the second bullet forward as well as back:** a bare target sitting
|
||||
after a dotted filename is now extracted, resolved, and a blocking ERROR when it dangles, where the
|
||||
same clause used to pass unchecked in silence.
|
||||
|
||||
### SUGGESTION-only checks
|
||||
|
||||
Deterministic to measure, judgment to act on:
|
||||
|
||||
- a description with **no boundary clause at all** (`absent`);
|
||||
- an **arrow clause whose target could not be read** (`unparsed`);
|
||||
- an **arrow clause naming more than one target**;
|
||||
- a `## Gotchas` section with **more than five entries**;
|
||||
- a `## Gotchas` section over **25% of the body**.
|
||||
|
||||
### Hand-invoked skills are exempt from the routing rules, and only those
|
||||
|
||||
A skill or agent whose frontmatter carries `disable-model-invocation: true` skips three checks:
|
||||
|
||||
- the boundary-clause check, `absent` and `unparsed` alike;
|
||||
- the multi-target arrow check;
|
||||
- the 250-character description **target** (`hand_invoked()` in `scripts/skill-size-check.sh`).
|
||||
|
||||
It keeps the 400-character description FAIL and **both** body word tiers, and if its description
|
||||
does happen to name a target, that target is still resolved and can still dangle.
|
||||
|
||||
Why the exemption is right: `disable-model-invocation: true` removes the skill from the
|
||||
model-visible listing entirely — it is not preloaded, and the Skill tool refuses to call it — so its
|
||||
description is never matched against user intent. ADR-0020 and `skill-author`'s contract therefore
|
||||
give such a skill **one plain human-facing sentence**: no trigger list, no boundary clause. No
|
||||
validator knew the field existed (**#108**), so the boundary-clause SUGGESTION fired on exactly the
|
||||
shape the contract mandates, and its remedy — "so the router knows where NOT to send this skill" —
|
||||
was addressed to a router that cannot see the skill at all. An author who followed the advice made
|
||||
the file worse. There is no router to inform.
|
||||
|
||||
The half that does **not** lift is the point. The body is still loaded on invocation and still
|
||||
competes with the caller's live conversation, so neither body tier moves. The 400-character ceiling
|
||||
stands too: a hand-invoked description is not preloaded, but it is still the one line the user reads
|
||||
when choosing from the `/` menu, and that ceiling is an outlier stop rather than a routing-quality
|
||||
budget — which is precisely why the 250-character target is the tier that lifts.
|
||||
|
||||
The field is read as a **boolean**, not as a mention of the key. PyYAML already resolves the
|
||||
unquoted YAML 1.1 booleans, so the extra handling catches a quoted `"true"`, which a host reads as
|
||||
truthy; `disable-model-invocation: false` is the model-invoked case written out longhand and buys
|
||||
nothing. A frontmatter parse failure returns false rather than raising — the flag is a *modifier* on
|
||||
other checks, and `description_value()` on the same text already reports the broken frontmatter, so
|
||||
raising here would diagnose one file twice two different ways.
|
||||
|
||||
`caveman` and `zoom-out` are the two carriers here. `tests/test-skill-size-check.sh` pins both
|
||||
halves — what the carve-out lifts, each with a flag-removed control, and what it must not.
|
||||
|
||||
### `verbose: true` is load-bearing
|
||||
|
||||
The hook is declared `verbose: true` so the SUGGESTION tier is audible. pre-commit prints nothing at
|
||||
all for a passing hook, and a SUGGESTION deliberately does not fail — without verbose every
|
||||
suggestion is swallowed, which is exactly the invisibility ADR-0013 records for Vale warnings.
|
||||
ADR-0020's preload arithmetic depends on it: writing to the 400-char FAIL delivers roughly half the
|
||||
cut that writing to the 250-char SUGGESTION does, so the intended saving depends entirely on that
|
||||
tier being visible. The numbers, and the measurement method behind them, are not restated here —
|
||||
they live in ADR-0020's Consequences section, under "A ceiling does not produce an average", whose
|
||||
figures are pinned to the base commit the decision was taken on (`f9b919d`). Quoting them here would
|
||||
just create a second copy to go stale. It costs nothing on a clean file — the script prints only
|
||||
findings.
|
||||
|
||||
### Duplicated constants
|
||||
|
||||
`skill-audit`'s `validate.sh` holds a second copy of the four ADR-0020 constants
|
||||
(`DESC_SUGGEST_CHARS` / `DESC_MAX_CHARS` / `BODY_SUGGEST_WORDS` / `BODY_MAX_WORDS`), and
|
||||
`agent-audit`'s `validate.sh` holds a third copy of the two description constants. They are copied
|
||||
rather than imported because a cache-installed plugin's scripts cannot read files outside their own
|
||||
plugin directory. `tests/test-skill-size-check.sh` asserts the copies agree, so drift fails CI rather
|
||||
than silently letting an audit bless a skill the commit hook then rejects. The shared boundary
|
||||
resolver block is embedded verbatim in all three scripts between `BEGIN`/`END ADR-0020 SHARED
|
||||
BOUNDARY RESOLVER` markers and must stay byte-identical.
|
||||
|
||||
### `python3` and PyYAML are hard requirements
|
||||
|
||||
Both, and neither is a best-effort accelerator.
|
||||
|
||||
`python3` because the script measures the **folded** `description` value. Most descriptions here are
|
||||
`>`-block scalars, so a regex over the raw lines measures indentation and newlines instead of the
|
||||
value. Missing it fails the hook with an install pointer rather than skipping the ADR-0020 checks,
|
||||
which would be a vacuous green. In practice it is already present — pre-commit is itself a Python
|
||||
application.
|
||||
|
||||
**PyYAML** because the hand-rolled fallback frontmatter reader has been **removed deliberately**. It
|
||||
disagreed with a real parser across the FAIL boundary — one corpus description measured 270
|
||||
characters parsed and 412 unparsed — and a quoted `"description"` key or an explicit
|
||||
`description: null` returned empty from it, silently skipping the description *and* routing checks. A
|
||||
reader that mis-parses an unfamiliar scalar shape reports a clean pass on a file it never measured,
|
||||
which is the exact vacuous-green failure the `python3` check exists to avoid. `pip install pyyaml`
|
||||
(or `python3 -m pip install PyYAML`, or the distro's `python3-yaml`) if the hook reports it missing.
|
||||
|
||||
**Neither requirement generalises to every hook in this repo.** `check-rtk-prefix` needs `python3`
|
||||
but **not** PyYAML: it reads the markdown body and never touches frontmatter, so it has no scalar to
|
||||
fold.
|
||||
|
||||
## Agent files take the description gates, not the body gate
|
||||
|
||||
`check-apm-agents-valid` runs agent-audit's `validate.sh` over every real
|
||||
`plugins/*/.apm/agents/*.agent.md`. It derives its expected file set from `git ls-files` — the pattern
|
||||
`tests/run-bats.sh` established — so an agent file deleted from the worktree but still tracked fails
|
||||
the run, and **discovering zero agent files is an error, not a pass**. An untracked *new* agent file
|
||||
is still validated: the derivation is one-directional on purpose, so uncommitted work is not blocked
|
||||
but also cannot bypass the gate.
|
||||
|
||||
The hook exists because `validate.sh` was previously exercised only by `check-scope-walkup-sync`,
|
||||
against synthetic `mktemp` fixtures — it had never run against the agent files it governs. That is
|
||||
how ADR-0016 could be amended to bless a `disallowedTools` frontmatter field while `validate.sh`'s
|
||||
allowlist still rejected it: spec and enforcer disagreed and every gate stayed green.
|
||||
|
||||
Agents take the ADR-0020 **description** gates (agent-audit's `validate.sh` holds its own copy of
|
||||
those two constants) and, deliberately, **no body word gate**. A skill body is loaded into the
|
||||
caller's context and competes with the live conversation; an agent body becomes the system prompt of
|
||||
a *fresh* context. The rationale for the 900-word FAIL does not transfer. A bats test pins that
|
||||
absence in agent-audit's validator — adding a body gate there contradicts the ADR rather than fixing
|
||||
an inconsistency.
|
||||
|
||||
**Be precise about the scope of that guarantee: it holds for the *validator*, not for the shared
|
||||
script.** `scripts/skill-size-check.sh` applies its body gate to whatever path it is handed, and
|
||||
|
||||
```
|
||||
bash scripts/skill-size-check.sh plugins/*/.apm/agents/*.agent.md
|
||||
```
|
||||
|
||||
exits 1 today with 900-word body FAILs on `git-orchestrate` and `gitea-orchestrate`. (Counts are
|
||||
deliberately not pinned here — agent bodies are edited like any other file, and a figure in this
|
||||
paragraph goes stale the moment one is trimmed. Run the command.) Agent files escape only because
|
||||
the hook definitions filter on `SKILL.md`
|
||||
— a file-pattern accident that happens to implement the design, not the design itself. **Do not
|
||||
"extend" that hook's `files:` pattern to cover agents** on the assumption that the script already
|
||||
knows the difference; doing so silently enforces a gate ADR-0020 declines to set.
|
||||
|
||||
## Current retrofit status
|
||||
|
||||
The ADR-0020 gates ship hot, with no baseline file — a shrinking baseline was considered and
|
||||
rejected. The corpus is currently clean on both: 0 of 39 descriptions/bodies exceed their FAIL tier,
|
||||
0 dangling targets, 0 `Kyberforge.CompositionNote` (Vale) errors. History: issue #99.
|
||||
|
||||
Nothing is grandfathered — a new skill, or an edit that crosses a FAIL tier, is blocked on first
|
||||
commit. SUGGESTION counts are not pinned here; they move with every edit. Measure and check both
|
||||
gates before starting work on a skill:
|
||||
|
||||
```
|
||||
bash scripts/skill-size-check.sh plugins/*/.apm/skills/*/SKILL.md | grep -c '^SUGGESTION'
|
||||
pre-commit run --all-files # size AND Vale — skill-size-check alone can pass while Vale still blocks
|
||||
```
|
||||
|
||||
## The `rtk` prefix gate (ADR-0023)
|
||||
|
||||
`check-rtk-prefix` is a `repo: local` pre-commit hook running `scripts/check-rtk-prefix.sh` over
|
||||
`^plugins/[^/]+/\.apm/(skills/.*\.md|agents/.*\.agent\.md)$`, with `README.md` excluded. It enforces
|
||||
**ADR-0023 clause 1 and nothing else**: an executable, instructed local git command in plugin skill
|
||||
or agent content is written `rtk git`.
|
||||
|
||||
It is wider in file scope than the ADR-0020 hooks — every markdown file under a plugin's
|
||||
`.apm/skills/` and `.apm/agents/`, not `SKILL.md` alone — because the rule it enforces is about
|
||||
commands an agent runs, and most of those live in `references/`, which the ADR-0020 gates do not
|
||||
reach ([the `references/` blind spot](#the-blind-spot-references-is-unlinted-for-two-independent-reasons)).
|
||||
|
||||
### What it can decide, and what it declines to
|
||||
|
||||
ADR-0023 has three clauses and only the first is a pattern:
|
||||
|
||||
| Clause | Rule | Gated |
|
||||
|---|---|---|
|
||||
| 1 | executable + instructed → `rtk git` | yes |
|
||||
| 2 | illustrative / referential → bare `git` | no — undecidable |
|
||||
| 3 | machine-parsed or interactive → bare `git` | no — opt-out marker |
|
||||
|
||||
Clause 2 is a judgement about what a sentence is *doing*. "Run `git switch <branch>`" and "`git
|
||||
switch` refuses rather than clobbering local edits" are the same token sequence. A gate that guessed
|
||||
would fire on correct prose, and **a gate that fires on correct content gets added to `SKIP`** —
|
||||
which disarms clause 1 along with it. So the hook looks only at the two contexts where a `git`
|
||||
mention is unambiguously an instruction to execute:
|
||||
|
||||
- a line inside a fenced code block whose info string names a shell — `bash`, `sh`, `shell`, `zsh`,
|
||||
`console`, `shell-session`. Fences tagged `text`, `yaml`, `json`, or tagged with nothing, are **not**
|
||||
checked;
|
||||
- the **opening** backticked span of a "Run" column cell in a markdown dispatch table, and only the
|
||||
opening span.
|
||||
|
||||
That last narrowing is not fussiness. A Run cell routinely carries a command followed by prose about
|
||||
it, and the prose is clause 2. `git-worktrees/SKILL.md` has both shapes on adjacent rows — one cell
|
||||
reading `` `rtk git worktree add --track …` `` — always correct. `` `git worktree add <path>
|
||||
<branch>` `` expands to exactly this (instruction, then reference), and a `**Never** …` row whose Run
|
||||
cell is entirely explanation containing a bare `git push`. Checking every backticked span flags both;
|
||||
checking only a leading span flags neither, and still catches the ordinary
|
||||
`` | List | `git worktree list -v` | `` case the gate exists for.
|
||||
|
||||
### The clause-3 opt-out
|
||||
|
||||
A command that is deliberately bare — because rtk rewrites the output the skill parses, or because
|
||||
the command is interactive — is exempted by putting the literal string `ADR-0023` **on the same
|
||||
line**: in a shell comment for a code line, in the cell text for a table row.
|
||||
|
||||
Per line, never per block. A fenced procedure routinely mixes `rtk git` steps with one deliberately
|
||||
bare command (`git-remotes/references/push.md` does exactly that), and a block-level marker would
|
||||
silently disarm every checked line around the marked one. The cost is a repeated `# bare per
|
||||
ADR-0023` in the three blocks of `git-log-format.md` where every line is deliberately bare; that
|
||||
repetition is the price of the marked line being the only line the marker speaks for.
|
||||
|
||||
The marker is a plain substring match, so a line that mentions `ADR-0023` for an unrelated reason is
|
||||
also exempt. Accepted deliberately: the marker records an author's opt-out, it is not a security
|
||||
boundary, and a stricter form would only move the same trust to a different string.
|
||||
|
||||
### What it deliberately does not cover
|
||||
|
||||
- **Clause 2.** Nothing checks that an illustrative mention stayed bare. A sweep that re-prefixes a
|
||||
referential `git` passes this gate. The inline reasons ADR-0023 requires on clause-3 sites are the
|
||||
only defence, and they are prose.
|
||||
- **Prose bullets.** Most of `branch-operations.md`, `merging.md` and `rewrite-history.md` instruct
|
||||
in list items, not fences. Those are clause-1 sites the gate cannot see, because it cannot
|
||||
distinguish them from clause-2 mentions in the same list.
|
||||
- **`README.md`, excluded by pattern.** A skill-directory README is consumer-facing prose no agent
|
||||
loads, and the `git clone https://github.com/bats-core/…` lines in the seven `tests/README.md`
|
||||
files are setup instructions for a third party who has no `rtk`. Prefixing those would be actively
|
||||
wrong, not merely noisy — see ADR-0023's consumer section.
|
||||
- **Quoting.** The line splitter breaks on `;`, `|`, `&&`, `||`, `$(` and backticks without tracking
|
||||
quotes, so a git command inside a quoted argument is decided by accident.
|
||||
`rtk git submodule foreach 'git pull origin main || :'` passes because the segment holding the
|
||||
inner command begins with `rtk` — the right answer for the wrong reason. Write
|
||||
`foreach 'git a; git b'` and the second inner command is a false positive needing the marker.
|
||||
ADR-0023 records this shape as one the rule itself does not decide.
|
||||
- **Non-git commands.** Only `git` is checked. `rtk` fronts `gh`, `docker`, `kubectl` and others; no
|
||||
gate covers those, and the corpus does not currently instruct them.
|
||||
|
||||
`tests/test-check-rtk-prefix.sh` pins all of it, including the false-positive cases. Its first case
|
||||
reconstructs the plugin corpus as it stood on `main` before the #113 sweep and asserts the gate
|
||||
fails there with at least 20 findings, one of them the `gitea-*` `git remote get-url origin` drift
|
||||
the sweep missed — a gate that only passes on the already-fixed tree proves nothing about the drift
|
||||
it was written for.
|
||||
|
||||
## Vale
|
||||
|
||||
Install the `vale` binary — `brew install vale` (macOS), `snap install vale` (Linux),
|
||||
`choco install vale` (Windows), or see <https://vale.sh/docs/vale-cli/installation/>. No `vale sync`
|
||||
is needed: the `Kyberforge` styles are **committed** under
|
||||
`plugins/kyberforge/.apm/skills/{skill-audit,agent-audit}/assets/vale/styles/`, not downloaded
|
||||
packages (ADR-0014).
|
||||
|
||||
### Two copies, one canonical
|
||||
|
||||
Wiring Vale as a deterministic prefilter for `skill-audit`/`agent-audit`'s Description dimension
|
||||
(motivation: issue #84) is repo-specific, not part of the generic `lint` plugin, so it does not live
|
||||
in `plugins/lint/` — and per ADR-0014 it no longer lives at the repo root either. It lives **twice**,
|
||||
one copy per skill, both under `plugins/kyberforge/.apm/skills/`:
|
||||
|
||||
| Copy | Styles | `.vale.ini` sections |
|
||||
|---|---|---|
|
||||
| `agent-audit/assets/vale/` — **canonical** | `Kyberforge`, `KyberforgeCopilot` | `[**/agents/*.md]`, `[**/*.agent.md]` |
|
||||
| `skill-audit/assets/vale/` — smaller duplicate | `Kyberforge` | `[**/SKILL.md]` |
|
||||
|
||||
Duplicated rather than shared because a plugin's cache-install copies only each skill's own files —
|
||||
there is no cross-skill sharing to point at. `check-vale-style-sync` at pre-push is what keeps them
|
||||
from drifting; `KyberforgeCopilot` is the one deliberate inequality, being scoped only to `.agent.md`
|
||||
files for the Copilot-only "`Use proactively` has no effect" check.
|
||||
|
||||
### What Vale owns, and what stays LLM judgment
|
||||
|
||||
Eleven rule files across the two copies, six distinct rules:
|
||||
|
||||
| Rule | Vale scope | Bans | From |
|
||||
|---|---|---|---|
|
||||
| `Kyberforge.DescriptionOpener` | `text.frontmatter.description` | non-imperative openers ("This skill/agent…") | issue #84 |
|
||||
| `Kyberforge.VagueWording` | `text.frontmatter.description` | vague capability wording ("helps with", "utilize", …) | issue #84 |
|
||||
| `Kyberforge.PaddingPhrase` | `text` | generic "see `references/` for details" padding | issue #84 |
|
||||
| `KyberforgeCopilot.ProactivePhrase` | `text.frontmatter.description` | `Use proactively` (no effect in Copilot) | issue #84 |
|
||||
| `Kyberforge.SentenceOpenerThereIs` | `sentence` | "There is/are" sentence openers | ADR-0013 |
|
||||
| `Kyberforge.CompositionNote` | `text.frontmatter.description` | architecture and composition prose in a description | ADR-0020 |
|
||||
|
||||
Vale covers the **pattern-matchable** sub-checks named in issue #84 plus, per ADR-0013, one
|
||||
cherry-picked body-wide prose-pattern rule. Everything else stays LLM judgment: defaults-vs-menus,
|
||||
why-rationale, the non-pattern-matchable body-discipline calls, near-miss exclusion strength, and
|
||||
control calibration. New rules land directly in `styles/Kyberforge` and block immediately — there is
|
||||
no trial tier.
|
||||
|
||||
The cherry-pick record, so it is not re-litigated:
|
||||
|
||||
- `Kyberforge.SentenceOpenerThereIs` **landed** — 22 held-out hits, both in-corpus hits clean
|
||||
rewrites, zero suppressions needed.
|
||||
- `Kyberforge.VagueQualifier` was cherry-picked and then **deleted**. 2 hits across the corpus as it
|
||||
stood on 2026-08-08 (before the `.apm/` restructure): one marginal, and one unfixable false
|
||||
positive — `caveman/SKILL.md` quotes `of course` as an example of filler, a mention rather than a
|
||||
use — which forced the repo's only Vale suppression comments.
|
||||
- `governance.md` and `CONTROLS.md` were evaluated as rule sources and **excluded**: nothing
|
||||
prose-pattern-matchable to mine.
|
||||
|
||||
### Why every rule is `level: error`
|
||||
|
||||
Every alert is a FAIL, with no ignorable tier — same all-or-nothing model as shellcheck, the test
|
||||
suite, and conventional-pre-commit. Graded severities do not work here: **Vale's exit code keys on
|
||||
`error` alerts alone**, so a `warning` or `suggestion` rule exits 0, and pre-commit swallows a
|
||||
passing hook's output. Such a rule would be invisible and would block nothing.
|
||||
|
||||
`MinAlertLevel` and `--minAlertLevel` are correspondingly **absent** from both `.vale.ini` files and
|
||||
from the hook definitions. Under this model they are no-ops; adding one is not a missing knob.
|
||||
|
||||
The `verbose: true` escape hatch that makes `skill-size-check`'s SUGGESTION tier audible has no
|
||||
analogue here — Vale has no tier to make audible.
|
||||
|
||||
### External consumers: the root `.pre-commit-hooks.yaml`
|
||||
|
||||
The root `.pre-commit-hooks.yaml` exposes both Vale copies (`kyberforge-vale-audit-skill`,
|
||||
`kyberforge-vale-audit-agent`) plus `kyberforge-skill-size-check`, so any external repo can enforce
|
||||
the same rules with `repo: <this-repo-url>, rev: <tag>` in its own `.pre-commit-config.yaml`.
|
||||
pre-commit clones the pinned rev into its own cache, independent of whether Claude Code or the
|
||||
`kyberforge` plugin is installed at all; the same mechanism covers CI via `pre-commit run
|
||||
--all-files`. `skill-size-check` has no external asset dependency, so it needed no relocation under
|
||||
ADR-0014 — only exposure.
|
||||
|
||||
This repo's own `vale-audit-prefilter-skill` / `-agent` hooks consume the **identical**
|
||||
plugin-bundled copies via `repo: local`. Deliberately not a third root copy, and deliberately **not a
|
||||
pinned self-reference** — a pinned self-reference would lint working-tree edits against the last
|
||||
tagged release rather than against the change being made.
|
||||
|
||||
### Pre-commit
|
||||
|
||||
Two prefilter hooks, with `.apm/`-scoped `files:` patterns:
|
||||
|
||||
| Hook | Pattern |
|
||||
|---|---|
|
||||
| `vale-audit-prefilter-skill` | `^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$` |
|
||||
| `vale-audit-prefilter-agent` | `^plugins/[^/]+/\.apm/agents/[^/]+\.agent\.md$` |
|
||||
|
||||
Only the **authoring source** triggers them. The `skill-author` scaffold template and the four
|
||||
vendored research-corpus `SKILL.md` files match neither pattern (see
|
||||
[Skill and agent context gates](#skill-and-agent-context-gates-adr-0020), which uses the same scope
|
||||
for `skill-size-check`), so prose findings surface only when you edit a file this repo actually
|
||||
authors. Without the binary the hooks fail with a bare "command not found" and no install pointer.
|
||||
|
||||
**Two hooks, not one combined hook.** Both manifests split the prefilter in two precisely because a
|
||||
single hook can point at only one copy, and that copy would silently 0-file-skip the other file
|
||||
shape (see [A 0-file Vale run is NOT RUN](#a-0-file-vale-run-is-not-run)).
|
||||
|
||||
### The `.vale.ini` globs do no scoping
|
||||
|
||||
Each `.vale.ini`'s section globs are **path-agnostic** — `[**/SKILL.md]` for skill-audit's copy,
|
||||
`[**/agents/*.md]` and `[**/*.agent.md]` for agent-audit's — and constrain filename *shape*, not
|
||||
location: Vale's `*` crosses `/`. A `SKILL.md` outside `plugins/` (a project-scope
|
||||
`.claude/skills/foo/SKILL.md`, say) still matches `[**/SKILL.md]` and gets linted normally.
|
||||
|
||||
All scoping therefore comes from the pre-commit hook's own `files:` regex and from the audit skills
|
||||
passing one explicit file per invocation. The two manifests scope **differently on purpose**:
|
||||
|
||||
| Manifest | `-skill` | `-agent` |
|
||||
|---|---|---|
|
||||
| `.pre-commit-config.yaml` (pins this repo's layout) | `^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$` | `^plugins/[^/]+/\.apm/agents/[^/]+\.agent\.md$` |
|
||||
| `.pre-commit-hooks.yaml` (layout-agnostic for consumers) | `(^\|/)SKILL\.md$` | `(^\|/)agents/[^/]+\.md$\|\.agent\.md$` |
|
||||
|
||||
Narrowing a `.vale.ini` glob to a `plugins/`-shaped path to "tighten" it breaks the consumer case,
|
||||
and `check-vale-style-sync`'s probe set is built to catch exactly that.
|
||||
|
||||
### The blind spot: `references/` is unlinted, for two independent reasons
|
||||
|
||||
Every `references/*.md` file in the corpus is outside the prose gate. Count them with
|
||||
`git ls-files | grep -cE '^plugins/[^/]+/\.apm/skills/[^/]+/references/.*\.md$'` rather than reading
|
||||
a figure here; it moves with every retrofit. This is the gap that matters most, because the context
|
||||
contract's own remedy for an over-long body is to move prose **into** `references/` — the gate pushes
|
||||
text across its own boundary and then stops watching it.
|
||||
|
||||
**Closing either cause alone changes nothing.** There are two, and they are independent:
|
||||
|
||||
| Cause | Where | Effect on a `references/` file |
|
||||
|---|---|---|
|
||||
| the `Kyberforge` style is scoped `[**/SKILL.md]` | `skill-audit/assets/vale/.vale.ini` | matches no section, so Vale lints 0 files and exits 0 |
|
||||
| the hook's `files:` regex is `^plugins/[^/]+/\.apm/skills/[^/]+/SKILL\.md$` | `vale-audit-prefilter-skill` in `.pre-commit-config.yaml` | the file is never handed to Vale at all |
|
||||
|
||||
Verified both ways. Handing skill-audit's `vale-wrap.sh` a reference file directly — bypassing
|
||||
pre-commit entirely, so only the style scope is in play — prints `0 errors … in 0 files` and exits 0,
|
||||
where the same wrapper on a `SKILL.md` reports `in 1 file`. And the hook's `files:` regex, applied to
|
||||
`git ls-files`, selects only the skill-directory `SKILL.md` files scoped at the top of this page, so
|
||||
pre-commit never hands Vale a reference file to begin with. Widening the glob to `[**/*.md]` would
|
||||
still lint nothing through the hook; widening the hook's `files:` alone would hand Vale files its own
|
||||
config declines to match, which is the [0-file NOT RUN](#a-0-file-vale-run-is-not-run) shape — a
|
||||
green run that measured nothing. **Issue #117** records the style-scope half; the hook half has to
|
||||
land in the same change or the fix is cosmetic.
|
||||
|
||||
The consumer manifest is a third axis and does not rescue this either: `.pre-commit-hooks.yaml`'s
|
||||
`(^|/)SKILL\.md$` is layout-agnostic but still filename-shaped, so an external repo running
|
||||
`kyberforge-vale-audit-skill` has the same gap.
|
||||
|
||||
### `vale-wrap.sh`, never bare `vale`
|
||||
|
||||
Both audit skills' Step 1 and both pre-commit hooks call **each copy's own**
|
||||
`scripts/vale-wrap.sh`, not `vale`. It works around a confirmed **Vale 3.15.2** limitation:
|
||||
`text.frontmatter.description` silently stops matching on most — not all — multi-line descriptions.
|
||||
|
||||
Verified by reproduction on a deliberately-bad fixture, not assumed:
|
||||
|
||||
| Description scalar spanning 2+ lines | Vale's behaviour |
|
||||
|---|---|
|
||||
| `>` folded block | 0 alerts, exit 0 — **broken** |
|
||||
| plain (unquoted) continuation lines | 0 alerts, exit 0 — **broken** |
|
||||
| single- or double-quoted, wrapped | 0 alerts, exit 0 — **broken** |
|
||||
| `\|` literal block | alerts fire, exit 1 — lints normally |
|
||||
|
||||
The wrapper flattens the three broken forms to a single-line scalar in a scratch copy — or, for the
|
||||
rare value no inline scalar can spell verbatim, a `|-` block with one content line — padding with
|
||||
blank lines so **every other line number is unchanged**. `|` literal blocks and single-line
|
||||
descriptions pass through untouched. Most descriptions in this repo are `>` blocks, so before the
|
||||
wrapper a bad description in any of the three broken forms sailed straight through the prefilter.
|
||||
|
||||
### The `--config` argv defect
|
||||
|
||||
Handed **no `--config` at all**, the wrapper falls back to its own sibling `assets/vale/.vale.ini`,
|
||||
located from `${BASH_SOURCE[0]}` rather than from the cwd. That is why both manifests' `entry:` is
|
||||
now the bare script path with **no argument after it**.
|
||||
|
||||
pre-commit prefixes only `entry[0]` with the hook-repo clone path (`cmd = (prefix.path(cmd[0]),
|
||||
*cmd[1:])`), so every later argument resolves against the **consuming** repo's root. A `--config` in
|
||||
`.pre-commit-hooks.yaml` therefore pointed at a path no consumer has and hard-failed every external
|
||||
run with `E100 [--config] Runtime error`.
|
||||
|
||||
`.pre-commit-config.yaml` drops the argument too, deliberately keeping the two entries identical.
|
||||
The local `repo: local` hook resolved its `--config` correctly only because the consuming repo *was*
|
||||
this repo — and that divergence is why three review rounds exercised a path no external consumer
|
||||
takes and missed the defect. **Do not reintroduce a `--config` to either manifest to make the local
|
||||
run "explicit".**
|
||||
|
||||
An explicit `--config` from any other caller still wins, in all three argv forms (`--config X`,
|
||||
`--config=/abs`, `--config=rel`), and a relative one resolves against the caller's cwd — matching
|
||||
bare `vale`, not the repo root.
|
||||
|
||||
Both audit skills' Step 1 passes no `--config` either. Step 1 resolves the script relative to the
|
||||
skill's own directory so the call works from an installed plugin cache; a relative `--config`
|
||||
alongside it would resolve against the cwd instead, yielding `E100 Runtime error … does not exist`
|
||||
and exit 2 — which both skills' fallback misreads as "vale unavailable" and silently downgrades to
|
||||
full LLM judgment.
|
||||
|
||||
`tests/test-vale-wrap.sh` regression-tests this against **skill-audit's** copy specifically: its
|
||||
fixtures are all `SKILL.md`-shaped, and only skill-audit's `.vale.ini` carries that glob section.
|
||||
|
||||
### A 0-file Vale run is NOT RUN
|
||||
|
||||
Vale reports 0 files only when the path it is handed matches **no glob section at all** — a
|
||||
differently-named file, or a directory argument holding nothing that matches. That run prints
|
||||
|
||||
```
|
||||
✔ 0 errors ... in 0 files.
|
||||
```
|
||||
|
||||
and exits 0, indistinguishable from a clean pass. Both audits therefore treat a 0-file Vale run as
|
||||
**NOT RUN** and fall back to full LLM judgment rather than reporting the Description dimension
|
||||
clean.
|
||||
|
||||
### Pre-push
|
||||
|
||||
`vale` is a **pre-push** dependency too, not only pre-commit. `check-vale-style-sync` runs **six
|
||||
glob-coverage probes** by invoking `vale --config` — one representative path per file shape the
|
||||
prefilter is supposed to cover. They are the only assertions in the script that catch a `.vale.ini`
|
||||
glob typo (`[**/SKILL.md]` → `[**/SKILLS.md]`), the failure mode where every text-level check stays
|
||||
clean while vale lints zero files. As a warning this self-disabled on exactly that mutation and
|
||||
exited 0, and since pre-commit swallows a passing hook's output the stderr line was never seen — the
|
||||
hook reported `Passed`. Missing `vale` is therefore a hard failure here.
|
||||
|
||||
The opt-out is `CHECK_VALE_STYLE_SYNC_ALLOW_MISSING_VALE=1`, and **it is not `SKIP=`**: the hook
|
||||
still runs and still asserts everything verifiable from file text, but the six probes do not, and its
|
||||
summary says so explicitly —
|
||||
|
||||
```
|
||||
Vale style sync check passed (text-level only, vale unavailable): … 0 glob probe(s) verified.
|
||||
```
|
||||
|
||||
Use it only on a machine that genuinely cannot install `vale`, and read that line as "the glob axis
|
||||
was not checked", not as a pass. The hook is `verbose: true` for exactly that reason — its clean
|
||||
output is a single line, so it costs one line per push.
|
||||
|
||||
### Mentioning banned phrasing without tripping the rule
|
||||
|
||||
House convention: banned phrasing that must be **mentioned** rather than used goes in backticks or a
|
||||
fenced code block. Vale skips code spans and fences, so no suppression is needed — which is why this
|
||||
document quotes `Use proactively` and "There is/are" the way it does.
|
||||
|
||||
Inline `<!-- vale Rule = NO -->` is the fallback **only** where backticking is impossible. Use the
|
||||
HTML-comment form; the MDX `{/* */}` form does not work in plain Markdown. The one time a rule forced
|
||||
suppression comments, the rule was deleted instead (see the `VagueQualifier` entry above).
|
||||
|
||||
## Tests
|
||||
|
||||
```
|
||||
bash tests/run-tests.sh # every test-*.sh plus the bats suite
|
||||
bash tests/run-tests.sh --bats-only # just bats
|
||||
```
|
||||
|
||||
First run auto-initializes the bats submodules; no manual `git submodule update` needed.
|
||||
|
||||
**Exit 77 = SKIPPED.** A suite that skips because a dependency is missing does **not** fail an ad-hoc
|
||||
run. The pre-push hook invokes the same script as `--strict` (`RUN_TESTS_STRICT=1` is equivalent),
|
||||
where a skip **does** fail the push: at pre-push a skip means one of the documented dependencies is
|
||||
absent on this machine, so the gate would otherwise report success having run fewer suites than it
|
||||
appears to. Without `--strict` the gate once went green having verified 15 of 17 suites on a
|
||||
vale-less PATH, with the skip list swallowed. Without vale, three suites skip —
|
||||
`test-check-vale-style-sync.sh`, `test-vale-hooks-consumer.sh`, `test-vale-wrap.sh` — and the strict
|
||||
failure names each one and what to install.
|
||||
|
||||
`tests/run-bats.sh` derives the set of `.bats` files it expects from `git ls-files`, so a `.bats`
|
||||
file deleted from the worktree but still tracked in the index fails the run rather than silently
|
||||
shrinking the suite. Remove one with `git rm` (or stage the deletion) when intentional; an untracked
|
||||
new `.bats` file is picked up and needs no ceremony.
|
||||
|
||||
Both discovery walks (`tests/run-bats.sh` and `tests/run-tests.sh`) exclude `apm_modules/`:
|
||||
`apm install` materializes a full copy of every plugin there, and running a dependency's copy of a
|
||||
`.bats` file breaks its relative path to the bats helpers — **167 spurious failures** before the
|
||||
exclusion landed.
|
||||
|
||||
## apm gates
|
||||
|
||||
### `apm-audit-ci`
|
||||
|
||||
Runs `apm audit --ci` **once per manifest** — the root one and each of the six plugin packages —
|
||||
because the root-only invocation audits the marketplace manifest and **nothing else**, and
|
||||
`apm-pack-check-clean` does not parse plugin `dependencies:` blocks either. Verified: a malformed
|
||||
dependency entry passes `apm pack --check-versions --check-clean --dry-run` and fails
|
||||
`apm audit --ci` in that package's directory. Costs ~0.5s per package.
|
||||
|
||||
It verifies **exactly two things** per manifest and claims no more:
|
||||
|
||||
- **manifest-parse** — each `apm.yml` parses as a valid APM manifest. Unconditional; verified to fire
|
||||
on a dependency entry missing its `git`/`path`/`registry` field (`Cannot parse apm.yml`).
|
||||
- **lockfile-exists** — any package declaring dependencies has a consistent `apm.lock.yaml`.
|
||||
Conditional, and vacuous while every plugin `apm.yml` declares `dependencies: {apm: [], mcp: []}`;
|
||||
it arms itself the moment one does not (verified by adding a git dependency to
|
||||
`plugins/lint/apm.yml`).
|
||||
|
||||
It does **not** enforce an org policy. apm discovers one from the git remote and only understands
|
||||
github.com and Azure DevOps, so against this repo's self-hosted Gitea remote it prints:
|
||||
|
||||
```
|
||||
No org policy found at unknown; enforcement skipped
|
||||
```
|
||||
|
||||
**Do not "fix" that with `policy.fetch_failure_default: block` in `apm.yml`.** apm's own message
|
||||
suggests it; it was tried on a scratch copy and **rejected**. With no reachable policy source it does
|
||||
not make the check meaningful, it makes it permanently red — `apm audit --ci` exits 1 with
|
||||
`No org policy found at unknown (policy.fetch_failure_default=block)` on every push, forever. A gate
|
||||
that can never go green is not a gate. Revisit only if this repo gains a policy source apm can reach.
|
||||
|
||||
It also does not scan for hidden Unicode: that scan is plain `apm audit`, a different mode (`--ci`
|
||||
refuses to combine with `--file`/`--strip`/`--dry-run`/`PACKAGE`), and plain `apm audit` here reports
|
||||
`No apm.lock.yaml found -- nothing to scan` and exits 0. Adding it would buy a second vacuous check.
|
||||
|
||||
### `check-executables-allow-sync`
|
||||
|
||||
apm gates a package's `hooks/` and `bin/` on an **exact `<package>#<version>` dictionary lookup** in
|
||||
root `apm.yml`'s `executables.allow` (`apm_cli/security/executables.py`, `is_package_approved`).
|
||||
There is no wildcard and no version-less form.
|
||||
|
||||
So bumping `plugins/kyberforge/apm.yml`'s `version:` without bumping the key **errors nowhere**: the
|
||||
entry simply stops matching, the gate blocks the hook, kyberforge's `SessionStart` hook stops
|
||||
deploying, and the apm install goes quietly stale — the exact failure ADR-0019 exists to end,
|
||||
reintroduced through the mechanism meant to secure it. ADR-0019 records this as a live failure mode;
|
||||
the release that shipped the hook hit it immediately.
|
||||
|
||||
`scripts/check-executables-allow-sync.sh` parses `version:` out of `plugins/kyberforge/apm.yml` and
|
||||
asserts root `apm.yml` carries the matching `kyberforge#<version>` key. A comment in the
|
||||
`executables:` block stays as the human-facing pointer; the hook is what actually holds. It parses
|
||||
with PyYAML where importable and falls back to a two-shape scan otherwise, so a missing pip package
|
||||
cannot become the thing that blocks every push.
|
||||
|
||||
## `.claude/settings.json`
|
||||
|
||||
**apm owns this file. Nothing repo-authored goes in it.**
|
||||
|
||||
`apm audit --ci` replays the install into a scratch tree and diffs the result byte-for-byte, so
|
||||
anything apm would not have written there — an `enabledPlugins` block, a real `hooks` entry — is
|
||||
permanent drift that fails `apm-audit-ci`. A hook you want in this repo is authored in
|
||||
`plugins/<name>/.apm/hooks/` and deployed by apm, never hand-written here.
|
||||
|
||||
Its committed content is whatever apm last wrote, which today is the merged `SessionStart` entry for
|
||||
kyberforge's `check-apm-current.sh`. That is apm's own output and it belongs in the commit (ADR-0019;
|
||||
ADR-0018's statement that the committed content is exactly `{"hooks": {}}` is superseded on that
|
||||
point only). Machine-specific settings go in the gitignored `.claude/settings.local.json`, which apm
|
||||
does not deploy and the replay does not compare; shared enforcement belongs in
|
||||
`.pre-commit-config.yaml`.
|
||||
|
||||
### Why it is excluded from `pretty-format-json`
|
||||
|
||||
It is the **second and last alternation** in that hook's `exclude:` pattern, and the only one there
|
||||
for a reason other than "generated manifest". Mind which number you are quoting: **two alternations,
|
||||
expanding to two real files** — `.claude-plugin/marketplace.json`, plus this one.
|
||||
|
||||
`pretty-format-json --autofix` sorts object keys unless `--no-sort-keys` is passed, while apm's hook
|
||||
integrator emits insertion order (`matcher` before `hooks`, `type` before `command`). Leaving the
|
||||
file in that hook's scope therefore rewrites apm's output into a form apm would never produce on the
|
||||
way into **every** commit, and `apm-audit-ci` then reports permanent drift on a file with an empty
|
||||
`git diff` — exactly what happened when the `SessionStart` hook first landed in `2e395a4`. Re-running
|
||||
`apm install` fixes the file; leaving it in scope would re-break it on the very commit carrying the
|
||||
fix.
|
||||
|
||||
**Load-bearing. Do not tidy it out of that list** (see `LESSONS.md`, 2026-08-14).
|
||||
|
||||
## Pushing without a network
|
||||
|
||||
No pre-push hook needs the network. Every entry in root `apm.yml`'s `marketplace.packages[]`
|
||||
resolves from a local `./plugins/<name>` path, so `apm-pack-check-clean` never calls `git ls-remote`.
|
||||
|
||||
`apm-audit-ci` calls `apm` too but was always local: its org-policy discovery resolves nothing on
|
||||
this remote before any network call.
|
||||
|
||||
---
|
||||
|
||||
## See also
|
||||
|
||||
- `docs/adr/0020-skill-description-and-body-context-contract.md` — the context contract, its
|
||||
enforcement table (deterministic vs. auditor judgment), and every rejected alternative
|
||||
- `docs/adr/0019-session-start-hook-keeps-the-apm-install-current.md` — the `SessionStart` hook, the
|
||||
executable-trust gate, and the version-pinned allow key
|
||||
- `docs/adr/0024-apm-is-the-only-supported-install-path.md` — apm as the sole install path, and the
|
||||
deletion of the flat content mirror and its `check-plugin-content-sync` gate. It supersedes
|
||||
`docs/adr/0017-plugin-content-mirror-bridges-apm-to-host-discovery.md` (**superseded** — plugin
|
||||
content sync, kept as the historical record)
|
||||
- `docs/adr/0015-apm-replaces-plugin-marketplace-authoring.md`,
|
||||
`docs/adr/0014-vale-prefilter-ships-from-the-plugin.md` — apm-generated manifests, committed Vale
|
||||
styles
|
||||
- `docs/spec/architecture.md` — directory structure, install pipeline, what is generated and what is
|
||||
hand-authored
|
||||
- `.pre-commit-config.yaml` — the hooks themselves, with inline rationale comments
|
||||
@@ -1,17 +1,18 @@
|
||||
---
|
||||
name: caveman
|
||||
disable-model-invocation: true
|
||||
description: >
|
||||
Ultra-compressed communication mode. Cuts token usage ~75% by dropping
|
||||
filler, articles, and pleasantries while keeping full technical accuracy.
|
||||
Use when user says "caveman mode", "talk like caveman", "use caveman",
|
||||
"less tokens", "be brief", or invokes /caveman.
|
||||
Ultra-compressed output mode that drops articles, filler and pleasantries while
|
||||
keeping technical substance exact, cutting token usage by roughly 75%.
|
||||
metadata:
|
||||
version: "1.0.0"
|
||||
---
|
||||
|
||||
Respond terse like smart caveman. All technical substance stay. Only fluff die.
|
||||
|
||||
## Persistence
|
||||
|
||||
ACTIVE EVERY RESPONSE once triggered. No revert after many turns. No filler drift. Still active if unsure. Off only when user says "stop caveman" or "normal mode".
|
||||
ACTIVE EVERY RESPONSE once user type `/caveman`. No revert after many turns. No filler drift. Still active if unsure. Off only when user says "stop caveman" or "normal mode".
|
||||
|
||||
## Rules
|
||||
|
||||
91
plugins/bin/.apm/skills/diagnose/SKILL.md
Normal file
91
plugins/bin/.apm/skills/diagnose/SKILL.md
Normal file
@@ -0,0 +1,91 @@
|
||||
---
|
||||
name: diagnose
|
||||
description: >
|
||||
Use when the user says "diagnose this" or "debug this", reports something
|
||||
broken, throwing, or failing, or says something got slow. Not filing or
|
||||
triaging a reported bug -> `triage`. Not test-first feature work -> `tdd`.
|
||||
metadata:
|
||||
version: "1.0.1"
|
||||
---
|
||||
|
||||
# Diagnose
|
||||
|
||||
A discipline for hard bugs. Skip phases only when explicitly justified.
|
||||
|
||||
When exploring the codebase, use the domain glossary for a clear mental model of the relevant modules, and check ADRs in the area.
|
||||
|
||||
## Phase 1 — Build a feedback loop
|
||||
|
||||
**This is the skill.** Everything else is mechanical. If you have a fast, deterministic, agent-runnable pass/fail signal for the bug, you will find the cause — bisection, hypothesis-testing, and instrumentation all just consume that signal. If you don't have one, no amount of staring at code will save you.
|
||||
|
||||
Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**
|
||||
|
||||
**If you do not yet have such a signal, read `references/feedback-loops.md`** — ten ways to build one ordered by cost, and what to ask the user for when the bug resists reproduction entirely.
|
||||
|
||||
**If you do have one, it is probably not sharp enough yet.** Make it faster and more deterministic, and make it assert on the exact symptom rather than "didn't crash" — a 30-second flaky loop is barely better than no loop. If it stays slow or intermittent after that, read that file's "Iterate on the loop itself" and "Intermittent bugs" sections.
|
||||
|
||||
Do not proceed to Phase 2 until you have a loop you believe in. If you cannot build one, stop and say so explicitly, listing what you tried — never hypothesise without a signal.
|
||||
|
||||
## Phase 2 — Reproduce
|
||||
|
||||
Run the loop. Watch the bug appear.
|
||||
|
||||
Confirm:
|
||||
|
||||
- [ ] The loop produces the failure mode the **user** described — not a different failure that happens to be nearby. Wrong bug = wrong fix.
|
||||
- [ ] The failure is reproducible across multiple runs. If it is intermittent, `references/feedback-loops.md` defines the rate high enough to debug against — go back to Phase 1 and raise it.
|
||||
- [ ] You have captured the exact symptom (error message, wrong output, slow timing) so later phases can verify the fix actually addresses it.
|
||||
|
||||
Do not proceed until you reproduce the bug.
|
||||
|
||||
## Phase 3 — Hypothesise
|
||||
|
||||
Generate **3–5 ranked hypotheses** before testing any of them. Single-hypothesis generation anchors on the first plausible idea.
|
||||
|
||||
Each hypothesis must be **falsifiable**: state the prediction it makes.
|
||||
|
||||
> Format: "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."
|
||||
|
||||
If you cannot state the prediction, the hypothesis is a vibe — discard or sharpen it.
|
||||
|
||||
**Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly ("we just deployed a change to #3"), or know hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't block on it — proceed with your ranking if the user is AFK.
|
||||
|
||||
## Phase 4 — Instrument
|
||||
|
||||
Each probe must map to a specific prediction from Phase 3. **Change one variable at a time.**
|
||||
|
||||
Tool preference:
|
||||
|
||||
1. **Debugger / REPL inspection** if the env supports it. One breakpoint beats ten logs.
|
||||
2. **Targeted logs** at the boundaries that distinguish hypotheses.
|
||||
3. Never "log everything and grep".
|
||||
|
||||
**Tag every debug log** with a unique prefix, e.g. `[DEBUG-a4f2]`. Cleanup at the end becomes a single grep. Untagged logs survive; tagged logs die.
|
||||
|
||||
**Perf branch.** For performance regressions, logs are usually wrong. Instead: establish a baseline measurement (timing harness, `performance.now()`, profiler, query plan), then bisect. Measure first, fix second.
|
||||
|
||||
## Phase 5 — Fix + regression test
|
||||
|
||||
Write the regression test **before the fix** — but only at a **correct seam**: one where the test exercises the real bug pattern as it occurs at the call site. If the available seam looks too shallow, or you cannot tell whether it is, read `references/regression-seams.md`.
|
||||
|
||||
**If no correct seam exists, that itself is the finding.** Note it and carry it into Phase 6 — the architecture is preventing the bug from being locked down.
|
||||
|
||||
At a correct seam:
|
||||
|
||||
1. Turn the Phase 1 loop into a failing test at that seam, narrowed to the symptom captured in Phase 2.
|
||||
2. Watch it fail.
|
||||
3. Apply the fix.
|
||||
4. Watch it pass.
|
||||
5. Re-run the Phase 1 feedback loop against the original, un-narrowed scenario.
|
||||
|
||||
## Phase 6 — Cleanup + post-mortem
|
||||
|
||||
Required before declaring done:
|
||||
|
||||
- [ ] Original repro no longer reproduces (re-run the Phase 1 loop)
|
||||
- [ ] Regression test passes (or absence of seam is documented)
|
||||
- [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix)
|
||||
- [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug location)
|
||||
- [ ] The hypothesis that turned out correct is stated in the commit / PR message — so the next debugger learns
|
||||
|
||||
**Then ask: what would have prevented this bug?** If the answer involves architectural change (no good test seam, tangled callers, hidden coupling) hand off to the `/improve-codebase-architecture` skill with the specifics. Make the recommendation **after** the fix is in, not before — you have more information now than when you started.
|
||||
@@ -0,0 +1,40 @@
|
||||
# Constructing and sharpening a feedback loop
|
||||
|
||||
A feedback loop is a fast, deterministic, agent-runnable pass/fail signal for the bug. Build the right one and the bug is 90% fixed. This file covers the whole arc: building a loop, sharpening one you already have, and escalating when the bug resists reproduction.
|
||||
|
||||
## Ways to construct one — try them in roughly this order
|
||||
|
||||
1. **Failing test** at whatever seam reaches the bug — unit, integration, e2e.
|
||||
2. **Curl / HTTP script** against a running dev server.
|
||||
3. **CLI invocation** with a fixture input, diffing stdout against a known-good snapshot.
|
||||
4. **Headless browser script** (Playwright / Puppeteer) — drives the UI, asserts on DOM/console/network.
|
||||
5. **Replay a captured trace.** Save a real network request / payload / event log to disk; replay it through the code path in isolation.
|
||||
6. **Throwaway harness.** Spin up a minimal subset of the system (one service, mocked deps) that exercises the bug code path with a single function call.
|
||||
7. **Property / fuzz loop.** If the bug is "sometimes wrong output", run 1000 random inputs and look for the failure mode.
|
||||
8. **Bisection harness.** If the bug appeared between two known states (commit, dataset, version), automate "boot at state X, check, repeat" so you can `git bisect run` it.
|
||||
9. **Differential loop.** Run the same input through old-version vs new-version (or two configs) and diff outputs.
|
||||
10. **HITL bash script.** Last resort. If a human must click, drive _them_ with `assets/hitl-loop.template.sh` so the loop is still structured. Captured output feeds back to you.
|
||||
|
||||
## Iterate on the loop itself
|
||||
|
||||
Treat the loop as a product. Once you have _a_ loop, ask:
|
||||
|
||||
- Can I make it faster? (Cache setup, skip unrelated init, narrow the test scope.)
|
||||
- Can I make the signal sharper? (Assert on the specific symptom, not "didn't crash".)
|
||||
- Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, freeze network.)
|
||||
|
||||
A 30-second flaky loop is barely better than no loop. A 2-second deterministic loop is a debugging superpower.
|
||||
|
||||
## Intermittent bugs — raise the reproduction rate
|
||||
|
||||
If the loop only sometimes fails, the goal is not a clean repro but a **higher reproduction rate**. Loop the trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not — keep raising the rate until it's debuggable.
|
||||
|
||||
## When you genuinely cannot build a loop
|
||||
|
||||
Stop and say so explicitly. List what you tried. Ask the user for:
|
||||
|
||||
- access to whatever environment reproduces it,
|
||||
- a captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or
|
||||
- permission to add temporary production instrumentation.
|
||||
|
||||
Do **not** proceed to hypothesise without a loop. A hypothesis you cannot falsify against a signal is a guess, and the fix that follows it is unverifiable.
|
||||
@@ -0,0 +1,24 @@
|
||||
# Judging a regression-test seam
|
||||
|
||||
Read this when Phase 5 leaves you unsure whether the seam available for the regression test is the correct one — either because the obvious seam looks shallow, or because there appears to be no seam at all.
|
||||
|
||||
## What makes a seam correct
|
||||
|
||||
A correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site: the same entry point, the same participants, the same ordering, and the same state the real caller holds when it goes wrong.
|
||||
|
||||
## Seams that are too shallow
|
||||
|
||||
- A single-caller test when the bug only appears with multiple callers.
|
||||
- A unit test that cannot replicate the chain of calls that triggered the bug.
|
||||
- A test that reproduces the symptom by construction — asserting on a value the test itself set — rather than by driving the code path that produces it.
|
||||
- A test that mocks out the collaborator the bug actually lives in.
|
||||
|
||||
A regression test at a shallow seam gives false confidence. It passes forever, including after a change reintroduces the bug at the real call site, and it will be read by the next maintainer as proof the bug is locked down.
|
||||
|
||||
## When there is no correct seam
|
||||
|
||||
Do not force one, and do not settle for a shallow seam to have something green. Instead:
|
||||
|
||||
1. Apply the fix and verify it against the Phase 1 loop directly.
|
||||
2. Write down which seams you considered and why each was too shallow.
|
||||
3. Carry that into Phase 6's "what would have prevented this bug" question. A missing seam is an architecture finding — tangled callers, hidden coupling, or a module with no testable boundary — and the handoff is the `improve-codebase-architecture` skill, with those specifics attached.
|
||||
@@ -1,6 +1,12 @@
|
||||
---
|
||||
name: grill-me
|
||||
description: Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me".
|
||||
description: >
|
||||
Use when the user says "grill me" or wants a plan or design stress-tested by
|
||||
relentless interview — one question at a time, down each branch of the
|
||||
decision tree. Not a plan to challenge against `CONTEXT.md` and ADRs ->
|
||||
`grill-with-docs`.
|
||||
metadata:
|
||||
version: "1.0.0"
|
||||
---
|
||||
|
||||
Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer.
|
||||
@@ -1,6 +1,11 @@
|
||||
---
|
||||
name: grill-with-docs
|
||||
description: Grilling session that challenges your plan against the existing domain model, sharpens terminology, and updates documentation (CONTEXT.md, ADRs) inline as decisions crystallise. Use when user wants to stress-test a plan against their project's language and documented decisions.
|
||||
description: >
|
||||
Use when a plan should be stress-tested against the project's domain model —
|
||||
the interview challenges terms against `CONTEXT.md` and writes decisions into
|
||||
it and into ADRs as they land. Not a plain interview -> `grill-me`.
|
||||
metadata:
|
||||
version: "1.0.0"
|
||||
---
|
||||
|
||||
<what-to-do>
|
||||
@@ -71,7 +76,7 @@ When the user states how something works, check whether the code agrees. If you
|
||||
|
||||
### Update CONTEXT.md inline
|
||||
|
||||
When a term is resolved, update `CONTEXT.md` right there. Don't batch these up — capture them as they happen. Use the format in [CONTEXT-FORMAT.md](./CONTEXT-FORMAT.md).
|
||||
When a term is resolved, update `CONTEXT.md` right there. Don't batch these up — capture them as they happen. Use the format in [context-format.md](references/context-format.md).
|
||||
|
||||
Don't couple `CONTEXT.md` to implementation details. Only include terms that are meaningful to domain experts.
|
||||
|
||||
@@ -83,6 +88,6 @@ Only offer to create an ADR when all three are true:
|
||||
2. **Surprising without context** — a future reader will wonder "why did they do it this way?"
|
||||
3. **The result of a real trade-off** — there were genuine alternatives and you picked one for specific reasons
|
||||
|
||||
If any of the three is missing, skip the ADR. Use the format in [ADR-FORMAT.md](./ADR-FORMAT.md).
|
||||
If any of the three is missing, skip the ADR. Use the format in [adr-format.md](references/adr-format.md).
|
||||
|
||||
</supporting-info>
|
||||
@@ -1,6 +1,13 @@
|
||||
---
|
||||
name: improve-codebase-architecture
|
||||
description: Find deepening opportunities in a codebase, informed by the domain language in CONTEXT.md and the decisions in docs/adr/. Use when the user wants to improve architecture, find refactoring opportunities, consolidate tightly-coupled modules, or make a codebase more testable and AI-navigable.
|
||||
description: >
|
||||
Use when the user wants to improve architecture, find refactoring
|
||||
opportunities, consolidate tightly-coupled modules, or make a codebase more
|
||||
testable and AI-navigable — deepening opportunities that turn shallow modules
|
||||
into deep ones, informed by `CONTEXT.md` and `docs/adr/`. Not debugging a
|
||||
failure -> `diagnose`.
|
||||
metadata:
|
||||
version: "1.0.1"
|
||||
---
|
||||
|
||||
# Improve Codebase Architecture
|
||||
@@ -9,7 +16,7 @@ Surface architectural friction and propose **deepening opportunities** — refac
|
||||
|
||||
## Glossary
|
||||
|
||||
Use these terms exactly in every suggestion. Consistent language is the point — don't drift into "component," "service," "API," or "boundary." Full definitions in [LANGUAGE.md](LANGUAGE.md).
|
||||
Use these terms exactly in every suggestion. Consistent language is the point — don't drift into "component," "service," "API," or "boundary."
|
||||
|
||||
- **Module** — anything with an interface and an implementation (function, class, package, slice).
|
||||
- **Interface** — everything a caller must know to use the module: types, invariants, error modes, ordering, config. Not just the type signature.
|
||||
@@ -20,19 +27,21 @@ Use these terms exactly in every suggestion. Consistent language is the point
|
||||
- **Leverage** — what callers get from depth.
|
||||
- **Locality** — what maintainers get from depth: change, bugs, knowledge concentrated in one place.
|
||||
|
||||
Key principles (see [LANGUAGE.md](LANGUAGE.md) for the full list):
|
||||
Key principles:
|
||||
|
||||
- **Deletion test**: imagine deleting the module. If complexity vanishes, it was a pass-through. If complexity reappears across N callers, it was earning its keep.
|
||||
- **The interface is the test surface.**
|
||||
- **One adapter = hypothetical seam. Two adapters = real seam.**
|
||||
|
||||
If a term or principle above is ambiguous in the case in front of you, or you need the definitions and the principles the two lists leave out, read `references/language.md`.
|
||||
|
||||
This skill is _informed_ by the project's domain model. The domain language gives names to good seams; ADRs record decisions the skill should not re-litigate.
|
||||
|
||||
## Process
|
||||
|
||||
### 1. Explore
|
||||
|
||||
Read the project's domain glossary and any ADRs in the area you're touching first.
|
||||
Read the domain glossary and any ADRs in the area first.
|
||||
|
||||
Then use the Agent tool with `subagent_type=Explore` to walk the codebase. Don't follow rigid heuristics — explore organically and note where you experience friction:
|
||||
|
||||
@@ -53,7 +62,7 @@ Present a numbered list of deepening opportunities. For each candidate:
|
||||
- **Solution** — plain English description of what would change
|
||||
- **Benefits** — explained in terms of locality and leverage, and also in how tests would improve
|
||||
|
||||
**Use CONTEXT.md vocabulary for the domain, and [LANGUAGE.md](LANGUAGE.md) vocabulary for the architecture.** If `CONTEXT.md` defines "Order," talk about "the Order intake module" — not "the FooBarHandler," and not "the Order service."
|
||||
**Use CONTEXT.md vocabulary for the domain, and the architecture glossary above for the architecture.** If `CONTEXT.md` defines "Order," talk about "the Order intake module" — not "the FooBarHandler," and not "the Order service."
|
||||
|
||||
**ADR conflicts**: if a candidate contradicts an existing ADR, only surface it when the friction is real enough to warrant revisiting the ADR. Mark it clearly (e.g. _"contradicts ADR-0007 — but worth reopening because…"_). Don't list every theoretical refactor an ADR forbids.
|
||||
|
||||
@@ -65,7 +74,7 @@ Once the user picks a candidate, drop into a grilling conversation. Walk the des
|
||||
|
||||
Side effects happen inline as decisions crystallize:
|
||||
|
||||
- **Naming a deepened module after a concept not in `CONTEXT.md`?** Add the term to `CONTEXT.md` — same discipline as `/grill-with-docs` (see [CONTEXT-FORMAT.md](../grill-with-docs/CONTEXT-FORMAT.md)). Create the file lazily if it doesn't exist.
|
||||
- **Naming a deepened module after a concept not in `CONTEXT.md`?** Add the term to `CONTEXT.md` — same discipline as `grill-with-docs`, in the format `grill-with-docs`'s `references/context-format.md` defines. Create the file lazily if it doesn't exist.
|
||||
- **Sharpening a fuzzy term during the conversation?** Update `CONTEXT.md` right there.
|
||||
- **User rejects the candidate with a load-bearing reason?** Offer an ADR, framed as: _"Want me to record this as an ADR so future architecture reviews don't re-suggest it?"_ Only offer when the reason would actually be needed by a future explorer to avoid re-suggesting the same thing — skip ephemeral reasons ("not worth it right now") and self-evident ones. See [ADR-FORMAT.md](../grill-with-docs/ADR-FORMAT.md).
|
||||
- **Want to explore alternative interfaces for the deepened module?** See [INTERFACE-DESIGN.md](INTERFACE-DESIGN.md).
|
||||
- **User rejects the candidate with a load-bearing reason?** Offer an ADR, framed as: _"Want me to record this as an ADR so future architecture reviews don't re-suggest it?"_ Only offer when the reason would actually be needed by a future explorer to avoid re-suggesting the same thing — skip ephemeral reasons ("not worth it right now") and self-evident ones. See `grill-with-docs`'s `references/adr-format.md`.
|
||||
- **Want to explore alternative interfaces for the deepened module?** Read `references/interface-design.md`.
|
||||
@@ -1,6 +1,6 @@
|
||||
# Deepening
|
||||
|
||||
How to deepen a cluster of shallow modules safely, given its dependencies. Assumes the vocabulary in [LANGUAGE.md](LANGUAGE.md) — **module**, **interface**, **seam**, **adapter**.
|
||||
How to deepen a cluster of shallow modules safely, given its dependencies. Assumes the vocabulary in [language.md](language.md) — **module**, **interface**, **seam**, **adapter**.
|
||||
|
||||
## Dependency categories
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
When the user wants to explore alternative interfaces for a chosen deepening candidate, use this parallel sub-agent pattern. Based on "Design It Twice" (Ousterhout) — your first idea is unlikely to be the best.
|
||||
|
||||
Uses the vocabulary in [LANGUAGE.md](LANGUAGE.md) — **module**, **interface**, **seam**, **adapter**, **leverage**.
|
||||
Uses the vocabulary in [language.md](language.md) — **module**, **interface**, **seam**, **adapter**, **leverage**.
|
||||
|
||||
## Process
|
||||
|
||||
@@ -11,7 +11,7 @@ Uses the vocabulary in [LANGUAGE.md](LANGUAGE.md) — **module**, **interface**,
|
||||
Before spawning sub-agents, write a user-facing explanation of the problem space for the chosen candidate:
|
||||
|
||||
- The constraints any new interface would need to satisfy
|
||||
- The dependencies it would rely on, and which category they fall into (see [DEEPENING.md](DEEPENING.md))
|
||||
- The dependencies it would rely on, and which category they fall into (see [deepening.md](deepening.md))
|
||||
- A rough illustrative code sketch to ground the constraints — not a proposal, just a way to make the constraints concrete
|
||||
|
||||
Show this to the user, then immediately proceed to Step 2. The user reads and thinks while the sub-agents work in parallel.
|
||||
@@ -20,21 +20,21 @@ Show this to the user, then immediately proceed to Step 2. The user reads and th
|
||||
|
||||
Spawn 3+ sub-agents in parallel using the Agent tool. Each must produce a **radically different** interface for the deepened module.
|
||||
|
||||
Prompt each sub-agent with a separate technical brief (file paths, coupling details, dependency category from [DEEPENING.md](DEEPENING.md), what sits behind the seam). The brief is independent of the user-facing problem-space explanation in Step 1. Give each agent a different design constraint:
|
||||
Prompt each sub-agent with a separate technical brief (file paths, coupling details, dependency category from [deepening.md](deepening.md), what sits behind the seam). The brief is independent of the user-facing problem-space explanation in Step 1. Give each agent a different design constraint:
|
||||
|
||||
- Agent 1: "Minimize the interface — aim for 1–3 entry points max. Maximise leverage per entry point."
|
||||
- Agent 2: "Maximise flexibility — support many use cases and extension."
|
||||
- Agent 3: "Optimise for the most common caller — make the default case trivial."
|
||||
- Agent 4 (if applicable): "Design around ports & adapters for cross-seam dependencies."
|
||||
|
||||
Include both [LANGUAGE.md](LANGUAGE.md) vocabulary and CONTEXT.md vocabulary in the brief so each sub-agent names things consistently with the architecture language and the project's domain language.
|
||||
Include both [language.md](language.md) vocabulary and CONTEXT.md vocabulary in the brief so each sub-agent names things consistently with the architecture language and the project's domain language.
|
||||
|
||||
Each sub-agent outputs:
|
||||
|
||||
1. Interface (types, methods, params — plus invariants, ordering, error modes)
|
||||
2. Usage example showing how callers use it
|
||||
3. What the implementation hides behind the seam
|
||||
4. Dependency strategy and adapters (see [DEEPENING.md](DEEPENING.md))
|
||||
4. Dependency strategy and adapters (see [deepening.md](deepening.md))
|
||||
5. Trade-offs — where leverage is high, where it's thin
|
||||
|
||||
### 3. Present and compare
|
||||
@@ -1,6 +1,12 @@
|
||||
---
|
||||
name: prototype
|
||||
description: Build a throwaway prototype to flush out a design before committing to it. Routes between two branches — a runnable terminal app for state/business-logic questions, or several radically different UI variations toggleable from one route. Use when the user wants to prototype, sanity-check a data model or state machine, mock up a UI, explore design options, or says "prototype this", "let me play with it", "try a few designs".
|
||||
description: >
|
||||
Use when the user wants a throwaway prototype to answer a design question about
|
||||
a data model, state machine or business logic, or to mock up a UI in several
|
||||
variations. Not production code -> `tdd`. Not talking a design through ->
|
||||
`grill-me`.
|
||||
metadata:
|
||||
version: "1.0.0"
|
||||
---
|
||||
|
||||
# Prototype
|
||||
@@ -9,10 +15,12 @@ A prototype is **throwaway code that answers a question**. The question decides
|
||||
|
||||
## Pick a branch
|
||||
|
||||
Identify which question is being answered — from the user's prompt, the surrounding code, or by asking if the user is around:
|
||||
| Question being answered | Build | Reference |
|
||||
|---|---|---|
|
||||
| "Does this logic / state model feel right?" | A tiny interactive terminal app that pushes the state machine through cases that are hard to reason about on paper | `references/logic.md` |
|
||||
| "What should this look like?" | Several radically different UI variations on one route, switchable via a URL search param and a floating bottom bar | `references/ui.md` |
|
||||
|
||||
- **"Does this logic / state model feel right?"** → [LOGIC.md](LOGIC.md). Build a tiny interactive terminal app that pushes the state machine through cases that are hard to reason about on paper.
|
||||
- **"What should this look like?"** → [UI.md](UI.md). Generate several radically different UI variations on a single route, switchable via a URL search param and a floating bottom bar.
|
||||
Resolve the row from the user's prompt, the surrounding code, or by asking if the user is around, then read only that reference — each is self-contained.
|
||||
|
||||
The two branches produce fundamentally different artifacts — getting this wrong wastes the whole prototype. If the question is genuinely ambiguous and the user isn't reachable, default to whichever branch better matches the surrounding code (a backend module → logic; a page or component → UI) and state the assumption at the top of the prototype.
|
||||
|
||||
@@ -9,7 +9,7 @@ A tiny interactive terminal app that lets the user drive a state model by hand.
|
||||
- "I want to feel out what the API should look like before writing it."
|
||||
- Anything where the user wants to **press buttons and watch state change**.
|
||||
|
||||
If the question is "what should this look like" — wrong branch. Use [UI.md](UI.md).
|
||||
If the question is "what should this look like" — wrong branch. Read `references/ui.md`.
|
||||
|
||||
## Process
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
Generate **several radically different UI variations** on a single route, switchable from a floating bottom bar. The user flips between variants in the browser, picks one (or steals bits from each), then throws the rest away.
|
||||
|
||||
If the question is about logic/state rather than what something looks like — wrong branch. Use [LOGIC.md](LOGIC.md).
|
||||
If the question is about logic/state rather than what something looks like — wrong branch. Read `references/logic.md`.
|
||||
|
||||
## When this is the right shape
|
||||
|
||||
79
plugins/bin/.apm/skills/research/SKILL.md
Normal file
79
plugins/bin/.apm/skills/research/SKILL.md
Normal file
@@ -0,0 +1,79 @@
|
||||
---
|
||||
name: research
|
||||
description: >-
|
||||
Use when the user wants a tool, library, or API researched from canonical
|
||||
documentation into structured per-topic reference markdown files. Not
|
||||
documentation written from existing code or specs -> `write-docs`. Not a bug
|
||||
or incident -> `diagnose`.
|
||||
metadata:
|
||||
version: "1.0.0"
|
||||
category: research
|
||||
allowed-tools:
|
||||
- Grep
|
||||
- Glob
|
||||
- Read
|
||||
- Write
|
||||
- WebSearch
|
||||
- WebFetch
|
||||
- mcp__context7__resolve-library-id
|
||||
- mcp__context7__query-docs
|
||||
model: sonnet
|
||||
---
|
||||
|
||||
## Gotchas
|
||||
|
||||
- Never infer the output path. A run writes a directory's worth of files, and a guessed destination scatters them through someone's source tree. If the user named no path, stop and ask.
|
||||
- Write nothing outside the given output path. A file placed beside the agreed directory is one the user never asked for and will not think to look for.
|
||||
- Never write an empty topic file. A stub `troubleshooting.md` reads downstream as researched and closed.
|
||||
- A Context7 response that is a "no results" message, a redirect notice, or header-only boilerplate is not coverage. A topic area counts as covered only when the response carries at least one substantive paragraph.
|
||||
|
||||
## Step 1 — Scope against the working directory
|
||||
|
||||
Search for existing use of the topic — imports, config files, version pins, reference files already written — and narrow the research to what is missing: the version actually in use, the topics not yet documented.
|
||||
|
||||
The default topic areas are `overview`, `installation`, `configuration`, `cli-reference`,
|
||||
`api-reference`, `examples` and `troubleshooting` — one file each, and only where content exists.
|
||||
If what belongs in one of them is unclear, or the topic needs a file outside that set, read
|
||||
`references/topics.md` for the per-topic coverage table and the custom-topic naming rule.
|
||||
|
||||
## Step 2 — Resolve against Context7
|
||||
|
||||
If the topic is a library, framework, or API and the user gave no starting URLs, call `resolve-library-id` with the topic name and the user's full question — match quality depends on the question, not the bare name — then `query-docs` once per default topic area. Record each response as a source with slug `context7-<library-slug>`, and mark which topic areas it covered — those skip the web reads at step 4.
|
||||
|
||||
If the library does not resolve, or the user gave starting URLs, go to step 3. Explicit URLs are a source choice; do not second-guess them with a resolution attempt.
|
||||
|
||||
## Step 3 — Discover sources
|
||||
|
||||
If the user gave starting URLs, skip discovery: those URLs are the source list and go straight to step 4.
|
||||
|
||||
Otherwise, for every topic area Context7 did not cover, websearch for canonical documentation — `llms.txt`, official developer docs, and API references ahead of tutorials or blog posts. Collect three to five candidate URLs before reading any of them.
|
||||
|
||||
If nothing usable comes back, stop and report what was searched, then ask for starting URLs rather than settling for tutorials.
|
||||
|
||||
## Step 4 — Read the sources
|
||||
|
||||
`WebFetch` each URL in turn. No subagent tool is granted here, so the reads are serial and every fetched page lands in this context: reduce each page to notes by topic area, plus the links worth deepening, before fetching the next one.
|
||||
|
||||
## Step 5 — Deepen
|
||||
|
||||
`WebFetch` the links worth following, still one at a time and still reducing each page to notes. Stop a branch once its content turns repetitive or leaves the topic, and cap the whole step at roughly ten additional pages — serial reads make that cap a real budget, not a formality.
|
||||
|
||||
## Step 6 — Write
|
||||
|
||||
Merge every set of notes, Context7 and web alike, by topic area, then write, in the output path:
|
||||
|
||||
- `<topic>.md` for each topic area that has content, default or custom. Frontmatter carries `topic:` (the filename without `.md`) and `source_keys:` (kebab-case slugs matching `sources.md`); the body is prose in `##` sections, with no inline URLs.
|
||||
- `sources.md`, always, one `##` section per source — including sources that yielded nothing — with exactly these four fields:
|
||||
|
||||
```markdown
|
||||
- **URL:** <full URL>
|
||||
- **Description:** <one-line summary>
|
||||
- **Contributing files:** <topic files this source contributed to>
|
||||
- **Status:** `extracted` | `no content extracted`
|
||||
```
|
||||
|
||||
Spell those four field names exactly as given. The downstream provenance validator matches them literally; prose in their place parses as nothing, and the check passes having verified nothing.
|
||||
|
||||
Read `references/file-format.md` when the four fields above do not settle the case: what a slug should be, the `context7-<library-slug>` slug and `context7:<library-id>` URL convention for a Context7 source, or what belongs in a topic body versus a verbatim copy of the source.
|
||||
|
||||
If no topic area has content, write nothing at all, `sources.md` included, and report what was searched.
|
||||
@@ -1,6 +1,11 @@
|
||||
---
|
||||
name: tdd
|
||||
description: Test-driven development with red-green-refactor loop. Use when user wants to build features or fix bugs using TDD, mentions "red-green-refactor", wants integration tests, or asks for test-first development.
|
||||
description: >
|
||||
Use when the user wants a feature built or a bug fixed test-first, in a strict
|
||||
red-green-refactor loop, one behaviour at a time. Not diagnosing an existing
|
||||
bug -> `diagnose`. Not throwaway exploratory code -> `prototype`.
|
||||
metadata:
|
||||
version: "1.0.1"
|
||||
---
|
||||
|
||||
# Test-Driven Development
|
||||
@@ -13,7 +18,7 @@ description: Test-driven development with red-green-refactor loop. Use when user
|
||||
|
||||
**Bad tests** are coupled to implementation. They mock internal collaborators, test private methods, or verify through external means (like querying a database directly instead of using the interface). The warning sign: your test breaks when you refactor, but behavior hasn't changed. If you rename an internal function and tests fail, those tests were testing implementation, not behavior.
|
||||
|
||||
See [tests.md](tests.md) for examples and [mocking.md](mocking.md) for mocking guidelines.
|
||||
If you need worked examples of the difference — a behaviour-level test beside the implementation-coupled version of the same check — read `references/tests.md`. If a test needs a collaborator faked, read `references/mocking.md` before reaching for a mock.
|
||||
|
||||
## Anti-Pattern: Horizontal Slices
|
||||
|
||||
@@ -44,14 +49,14 @@ RIGHT (vertical):
|
||||
|
||||
### 1. Planning
|
||||
|
||||
When exploring the codebase, use the project's domain glossary so that test names and interface vocabulary match the project's language, and respect ADRs in the area you're touching.
|
||||
When exploring the codebase, use the domain glossary so test names and interface vocabulary match the project's language, and respect ADRs in the area.
|
||||
|
||||
Before writing any code:
|
||||
|
||||
- [ ] Confirm with user what interface changes are needed
|
||||
- [ ] Confirm with user which behaviors to test (prioritize)
|
||||
- [ ] Identify opportunities for [deep modules](deep-modules.md) (small interface, deep implementation)
|
||||
- [ ] Design interfaces for [testability](interface-design.md)
|
||||
- [ ] Identify opportunities for [deep modules](references/deep-modules.md) (small interface, deep implementation)
|
||||
- [ ] Design interfaces for [testability](references/interface-design.md)
|
||||
- [ ] List the behaviors to test (not implementation steps)
|
||||
- [ ] Get user approval on the plan
|
||||
|
||||
@@ -88,7 +93,7 @@ Rules:
|
||||
|
||||
### 4. Refactor
|
||||
|
||||
After all tests pass, look for [refactor candidates](refactoring.md):
|
||||
After all tests pass, look for [refactor candidates](references/refactoring.md):
|
||||
|
||||
- [ ] Extract duplication
|
||||
- [ ] Deepen modules (move complexity behind simple interfaces)
|
||||
@@ -1,6 +1,11 @@
|
||||
---
|
||||
name: triage
|
||||
description: Triage issues through a state machine driven by triage roles. Use when user wants to create an issue, triage issues, review incoming bugs or feature requests, prepare issues for an AFK agent, or manage issue workflow.
|
||||
description: >
|
||||
Use when the user wants an issue created, triaged, or moved through the
|
||||
tracker's triage states, or an issue prepared for an AFK agent. Not debugging
|
||||
the bug itself -> `diagnose`. Not fleshing out a design -> `grill-with-docs`.
|
||||
metadata:
|
||||
version: "1.0.1"
|
||||
---
|
||||
|
||||
# Triage
|
||||
@@ -15,8 +20,8 @@ Every comment or issue posted to the issue tracker during triage **must** start
|
||||
|
||||
## Reference docs
|
||||
|
||||
- [AGENT-BRIEF.md](AGENT-BRIEF.md) — how to write durable agent briefs
|
||||
- [OUT-OF-SCOPE.md](OUT-OF-SCOPE.md) — how the `.out-of-scope/` knowledge base works
|
||||
- [agent-brief.md](references/agent-brief.md) — how to write durable agent briefs
|
||||
- [out-of-scope.md](references/out-of-scope.md) — how the `.out-of-scope/` knowledge base works
|
||||
|
||||
## Roles
|
||||
|
||||
@@ -35,7 +40,7 @@ Five **state** roles:
|
||||
|
||||
Every triaged issue should carry exactly one category role and one state role. If state roles conflict, flag it and ask the maintainer before doing anything else.
|
||||
|
||||
These are canonical role names — the actual label strings used in the issue tracker may differ. The mapping should have been provided to you - run `/setup-matt-pocock-skills` if not.
|
||||
These are canonical role names — the actual label strings used in the issue tracker may differ. Resolve each canonical name against the tracker's live label set before applying it, using whichever tracker skill this install provides. If a name has no counterpart there, report the gap and ask the maintainer for the mapping — never substitute a guess.
|
||||
|
||||
State transitions: an unlabeled issue normally goes to `needs-triage` first; from there it moves to `needs-info`, `ready-for-agent`, `ready-for-human`, or `wontfix`. `needs-info` returns to `needs-triage` once the reporter replies. The maintainer can override at any time — flag transitions that look unusual and ask before proceeding.
|
||||
|
||||
@@ -60,7 +65,7 @@ Show counts and a one-line summary per issue. Let the maintainer pick.
|
||||
|
||||
## Triage a specific issue
|
||||
|
||||
1. **Gather context.** Read the full issue (body, comments, labels, reporter, dates). Parse any prior triage notes so you don't re-ask resolved questions. Explore the codebase using the project's domain glossary, respecting ADRs in the area. Read `.out-of-scope/*.md` and surface any prior rejection that resembles this issue.
|
||||
1. **Gather context.** Read the full issue (body, comments, labels, reporter, dates). Parse any prior triage notes so you don't re-ask resolved questions. Explore the codebase using the domain glossary, respecting ADRs in the area. Read `.out-of-scope/*.md` and surface any prior rejection that resembles this issue.
|
||||
|
||||
2. **Recommend.** Tell the maintainer your category and state recommendation with reasoning, plus a brief codebase summary relevant to the issue. Wait for direction.
|
||||
|
||||
@@ -69,11 +74,11 @@ Show counts and a one-line summary per issue. Let the maintainer pick.
|
||||
4. **Grill (if needed).** If the issue needs fleshing out, run a `/grill-with-docs` session.
|
||||
|
||||
5. **Apply the outcome:**
|
||||
- `ready-for-agent` — post an agent brief comment ([AGENT-BRIEF.md](AGENT-BRIEF.md)).
|
||||
- `ready-for-agent` — post an agent brief comment ([agent-brief.md](references/agent-brief.md)).
|
||||
- `ready-for-human` — same structure as an agent brief, but note why it can't be delegated (judgment calls, external access, design decisions, manual testing).
|
||||
- `needs-info` — post triage notes (template below).
|
||||
- `wontfix` (bug) — polite explanation, then close.
|
||||
- `wontfix` (enhancement) — write to `.out-of-scope/`, link to it from a comment, then close ([OUT-OF-SCOPE.md](OUT-OF-SCOPE.md)).
|
||||
- `wontfix` (enhancement) — write to `.out-of-scope/`, link to it from a comment, then close ([out-of-scope.md](references/out-of-scope.md)).
|
||||
- `needs-triage` — apply the role. Optional comment if there's partial progress.
|
||||
|
||||
## Quick state override
|
||||
@@ -1,6 +1,6 @@
|
||||
# Writing Agent Briefs
|
||||
|
||||
An agent brief is a structured comment posted on a GitHub issue when it moves to `ready-for-agent`. It is the authoritative specification that an AFK agent will work from. The original issue body and discussion are context — the agent brief is the contract.
|
||||
An agent brief is a structured comment posted on an issue in the issue tracker when it moves to `ready-for-agent`. It is the authoritative specification that an AFK agent will work from. The original issue body and discussion are context — the agent brief is the contract.
|
||||
|
||||
## Principles
|
||||
|
||||
@@ -27,7 +27,7 @@ Describe **what** the system should do, not **how** to implement it. The agent w
|
||||
|
||||
The agent needs to know when it's done. Every agent brief must have concrete, testable acceptance criteria. Each criterion should be independently verifiable.
|
||||
|
||||
- **Good:** "Running `gh issue list --label needs-triage` returns issues that have been through initial classification"
|
||||
- **Good:** "Querying the issue tracker for the `needs-triage` label returns issues that have been through initial classification"
|
||||
- **Bad:** "Triage should work correctly"
|
||||
|
||||
### Explicit scope boundaries
|
||||
@@ -1,10 +1,15 @@
|
||||
---
|
||||
name: write-docs
|
||||
description: Write documentation for X, document this module, create docs for this feature. Use when the user wants to produce or update technical documentation derived from code, spec, or existing artifacts. Do NOT use when the user wants a PRD, ADR, decision doc, or skill file — those have dedicated skills.
|
||||
version: "1.0"
|
||||
description: >
|
||||
Use when the user wants technical documentation produced or updated from code
|
||||
or spec, every claim traced to a source — "write docs for X", "document this
|
||||
module", "create docs for this feature", "write a README for this". Not an ADR
|
||||
or other decision record -> `grill-with-docs`. Not an external tool researched
|
||||
from its docs -> `research`.
|
||||
updated: 2026-05-17
|
||||
when: invoked by explicit trigger ("write docs for X", "document this module", "create docs for this feature") or implicit request to produce technical documentation from code or spec
|
||||
metadata:
|
||||
version: "1.0.0"
|
||||
category: implement
|
||||
source:
|
||||
- repo: anthropics/skills
|
||||
@@ -35,7 +40,8 @@ You are a technical writer that produces documentation by reading code and spec
|
||||
- User says "write docs for X", "document this", "create docs for this feature", "write a README for this"
|
||||
|
||||
**Do not use when:**
|
||||
- User wants a PRD, decision doc, or architecture proposal → `to-prd` or `grill-me`
|
||||
- User wants an ADR, decision doc, or architecture proposal → `grill-with-docs`, which writes ADRs
|
||||
- User wants a PRD → no skill in this set produces one; say so rather than redirecting
|
||||
- User wants to document a skill file (skill files are self-describing)
|
||||
- User wants marketing or blog copy
|
||||
- Documentation requires tacit organisational knowledge that cannot be read from code or spec
|
||||
@@ -88,7 +94,7 @@ You are a technical writer that produces documentation by reading code and spec
|
||||
- Stage skipped without a logged reason → flag and require the one-sentence log before continuing
|
||||
- Code behaviour is undocumentable (internal implementation detail, no public spec) → note as out-of-scope in the doc; do not invent an explanation
|
||||
- Reader Testing sub-agent fails on multiple questions → surface the failures, return to step 4; do not mark complete
|
||||
- Requested output is a PRD, decision doc, or architecture proposal → redirect to `to-prd`, `grill-me`, or `grill-with-docs`
|
||||
- Requested output is an ADR, decision doc, or architecture proposal → redirect to `grill-with-docs`; for a PRD, say no skill here produces one instead of redirecting
|
||||
|
||||
## Self-check
|
||||
|
||||
@@ -2,6 +2,8 @@
|
||||
name: zoom-out
|
||||
description: Tell the agent to zoom out and give broader context or a higher-level perspective. Use when you're unfamiliar with a section of code or need to understand how it fits into the bigger picture.
|
||||
disable-model-invocation: true
|
||||
metadata:
|
||||
version: "1.0.1"
|
||||
---
|
||||
|
||||
I don't know this area of code well. Go up a layer of abstraction. Give me a map of all the relevant modules and callers, using the project's domain glossary vocabulary.
|
||||
I don't know this area of code well. Go up a layer of abstraction. Give me a map of all the relevant modules and callers, using the project's domain glossary.
|
||||
@@ -1,12 +0,0 @@
|
||||
{
|
||||
"author": {
|
||||
"name": "Defame1297",
|
||||
"url": "https://git.dev.rkdr.net/Defame1297/"
|
||||
},
|
||||
"description": "A place for things to be binned",
|
||||
"displayName": "bin",
|
||||
"keywords": [],
|
||||
"license": "MIT",
|
||||
"name": "bin",
|
||||
"version": "1.1.1"
|
||||
}
|
||||
@@ -1,12 +0,0 @@
|
||||
{
|
||||
"mcpServers": {
|
||||
"obsidian": {
|
||||
"args": [
|
||||
"@bitbonsai/mcpvault@latest",
|
||||
"docs/"
|
||||
],
|
||||
"command": "npx",
|
||||
"type": "stdio"
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -4,36 +4,32 @@ A place for things to be binned
|
||||
|
||||
## Install
|
||||
|
||||
**Claude Code:**
|
||||
apm is the only supported install path (ADR-0024). Declare this package in the consuming project's `apm.yml`:
|
||||
|
||||
```bash
|
||||
claude plugin marketplace add <owner>/<repo>
|
||||
claude plugin install bin@<marketplace-name>
|
||||
```yaml
|
||||
dependencies:
|
||||
apm:
|
||||
- git: git@git.dev.rkdr.net:Defame1297/holocron.git
|
||||
path: plugins/bin
|
||||
```
|
||||
|
||||
**GitHub Copilot CLI:**
|
||||
Then:
|
||||
|
||||
```bash
|
||||
copilot plugin marketplace add <owner>/<repo>
|
||||
copilot plugin install bin
|
||||
apm install
|
||||
```
|
||||
|
||||
**Local (development):**
|
||||
The entry above is unpinned and tracks the remote's default branch — add `ref: <tag>` to pin a release. Registering the catalogue instead (`apm marketplace add git@git.dev.rkdr.net:Defame1297/holocron.git --name holocron`) gets you the `bin@holocron` short name, but writes to `~/.apm/marketplaces.json` at user scope; the git+path object needs nothing beyond the manifest.
|
||||
|
||||
```bash
|
||||
# Claude Code
|
||||
claude --plugin-dir ./plugins/bin
|
||||
|
||||
# GitHub Copilot CLI
|
||||
copilot plugin install ./plugins/bin
|
||||
```
|
||||
**Native plugin installs do not work.** This package ships no per-plugin manifest and no flat content directories, so a host that installs it natively gets zero skills — and Claude Code raises no error while doing it (ADR-0024).
|
||||
|
||||
## Contents
|
||||
|
||||
| Component | Path | Description |
|
||||
| -------------| ------------------------------------------------------| ---------------------------------------------------------------|
|
||||
| Skills | `skills/` | Slash commands available after install |
|
||||
| Agents | `agents/` | Role-based agents (`.md` for Claude, `.agent.md` for Copilot) |
|
||||
| Component | Path | Description |
|
||||
|---|---|---|
|
||||
| Skills | `.apm/skills/` | Slash commands available after install |
|
||||
|
||||
`.apm/` is the authoring source and the only thing apm deploys (ADR-0024). This plugin ships no agents.
|
||||
|
||||
## Author
|
||||
|
||||
|
||||
35
plugins/bin/apm.yml
Normal file
35
plugins/bin/apm.yml
Normal file
@@ -0,0 +1,35 @@
|
||||
name: bin
|
||||
version: 1.1.7
|
||||
description: Skills for everyday AI-assisted development work that is not tied to a single tool, forge or language, and has not yet been split into a focused plugin.
|
||||
author:
|
||||
name: Defame1297
|
||||
email: defame1297@rkdr.net
|
||||
url: https://git.dev.rkdr.net/Defame1297/
|
||||
license: MIT
|
||||
homepage: https://git.dev.rkdr.net/Defame1297/holocron/src/branch/main/plugins/bin
|
||||
repository: https://git.dev.rkdr.net/Defame1297/holocron/src/branch/main/plugins/bin
|
||||
keywords:
|
||||
- utility
|
||||
- diagnostics
|
||||
- prototyping
|
||||
- tdd
|
||||
- research
|
||||
|
||||
# Constrains what .apm/ may contain: instructions, skill, hybrid, or prompts
|
||||
type: skill
|
||||
|
||||
# Which agent platforms to deploy to.
|
||||
# Resolution order: --target flag > this field > auto-detect from filesystem.
|
||||
# Accepted values: agent-skills, antigravity, claude, codex, copilot, cursor, gemini, grok-build, kiro, opencode, windsurf
|
||||
targets:
|
||||
- claude
|
||||
- copilot
|
||||
- codex
|
||||
|
||||
dependencies:
|
||||
apm: []
|
||||
mcp: []
|
||||
includes: auto
|
||||
devDependencies:
|
||||
apm: []
|
||||
scripts: {}
|
||||
@@ -26,11 +26,6 @@ trigger_tests:
|
||||
query: "Research why these integration tests are failing"
|
||||
should_trigger: false
|
||||
|
||||
- id: negative-neuledge
|
||||
name: "Negative — MCP server setup goes to neuledge-context"
|
||||
query: "Install the neuledge context server and set it up"
|
||||
should_trigger: false
|
||||
|
||||
- id: negative-context7-direct-question
|
||||
name: "Negative — direct doc question goes to context7-mcp, not research"
|
||||
query: "What are the Next.js middleware options?"
|
||||
|
||||
@@ -1,15 +0,0 @@
|
||||
{
|
||||
"author": {
|
||||
"email": "defame1297@rkdr.net",
|
||||
"name": "Defame1297"
|
||||
},
|
||||
"description": "A place for things to be binned",
|
||||
"keywords": [],
|
||||
"license": "MIT",
|
||||
"mcpServers": ".mcp.json",
|
||||
"name": "bin",
|
||||
"skills": [
|
||||
"skills/"
|
||||
],
|
||||
"version": "1.1.1"
|
||||
}
|
||||
@@ -1,117 +0,0 @@
|
||||
---
|
||||
name: diagnose
|
||||
description: Disciplined diagnosis loop for hard bugs and performance regressions. Reproduce → minimise → hypothesise → instrument → fix → regression-test. Use when user says "diagnose this" / "debug this", reports a bug, says something is broken/throwing/failing, or describes a performance regression.
|
||||
---
|
||||
|
||||
# Diagnose
|
||||
|
||||
A discipline for hard bugs. Skip phases only when explicitly justified.
|
||||
|
||||
When exploring the codebase, use the project's domain glossary to get a clear mental model of the relevant modules, and check ADRs in the area you're touching.
|
||||
|
||||
## Phase 1 — Build a feedback loop
|
||||
|
||||
**This is the skill.** Everything else is mechanical. If you have a fast, deterministic, agent-runnable pass/fail signal for the bug, you will find the cause — bisection, hypothesis-testing, and instrumentation all just consume that signal. If you don't have one, no amount of staring at code will save you.
|
||||
|
||||
Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**
|
||||
|
||||
### Ways to construct one — try them in roughly this order
|
||||
|
||||
1. **Failing test** at whatever seam reaches the bug — unit, integration, e2e.
|
||||
2. **Curl / HTTP script** against a running dev server.
|
||||
3. **CLI invocation** with a fixture input, diffing stdout against a known-good snapshot.
|
||||
4. **Headless browser script** (Playwright / Puppeteer) — drives the UI, asserts on DOM/console/network.
|
||||
5. **Replay a captured trace.** Save a real network request / payload / event log to disk; replay it through the code path in isolation.
|
||||
6. **Throwaway harness.** Spin up a minimal subset of the system (one service, mocked deps) that exercises the bug code path with a single function call.
|
||||
7. **Property / fuzz loop.** If the bug is "sometimes wrong output", run 1000 random inputs and look for the failure mode.
|
||||
8. **Bisection harness.** If the bug appeared between two known states (commit, dataset, version), automate "boot at state X, check, repeat" so you can `git bisect run` it.
|
||||
9. **Differential loop.** Run the same input through old-version vs new-version (or two configs) and diff outputs.
|
||||
10. **HITL bash script.** Last resort. If a human must click, drive _them_ with `scripts/hitl-loop.template.sh` so the loop is still structured. Captured output feeds back to you.
|
||||
|
||||
Build the right feedback loop, and the bug is 90% fixed.
|
||||
|
||||
### Iterate on the loop itself
|
||||
|
||||
Treat the loop as a product. Once you have _a_ loop, ask:
|
||||
|
||||
- Can I make it faster? (Cache setup, skip unrelated init, narrow the test scope.)
|
||||
- Can I make the signal sharper? (Assert on the specific symptom, not "didn't crash".)
|
||||
- Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, freeze network.)
|
||||
|
||||
A 30-second flaky loop is barely better than no loop. A 2-second deterministic loop is a debugging superpower.
|
||||
|
||||
### Non-deterministic bugs
|
||||
|
||||
The goal is not a clean repro but a **higher reproduction rate**. Loop the trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not — keep raising the rate until it's debuggable.
|
||||
|
||||
### When you genuinely cannot build a loop
|
||||
|
||||
Stop and say so explicitly. List what you tried. Ask the user for: (a) access to whatever environment reproduces it, (b) a captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. Do **not** proceed to hypothesise without a loop.
|
||||
|
||||
Do not proceed to Phase 2 until you have a loop you believe in.
|
||||
|
||||
## Phase 2 — Reproduce
|
||||
|
||||
Run the loop. Watch the bug appear.
|
||||
|
||||
Confirm:
|
||||
|
||||
- [ ] The loop produces the failure mode the **user** described — not a different failure that happens to be nearby. Wrong bug = wrong fix.
|
||||
- [ ] The failure is reproducible across multiple runs (or, for non-deterministic bugs, reproducible at a high enough rate to debug against).
|
||||
- [ ] You have captured the exact symptom (error message, wrong output, slow timing) so later phases can verify the fix actually addresses it.
|
||||
|
||||
Do not proceed until you reproduce the bug.
|
||||
|
||||
## Phase 3 — Hypothesise
|
||||
|
||||
Generate **3–5 ranked hypotheses** before testing any of them. Single-hypothesis generation anchors on the first plausible idea.
|
||||
|
||||
Each hypothesis must be **falsifiable**: state the prediction it makes.
|
||||
|
||||
> Format: "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."
|
||||
|
||||
If you cannot state the prediction, the hypothesis is a vibe — discard or sharpen it.
|
||||
|
||||
**Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly ("we just deployed a change to #3"), or know hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't block on it — proceed with your ranking if the user is AFK.
|
||||
|
||||
## Phase 4 — Instrument
|
||||
|
||||
Each probe must map to a specific prediction from Phase 3. **Change one variable at a time.**
|
||||
|
||||
Tool preference:
|
||||
|
||||
1. **Debugger / REPL inspection** if the env supports it. One breakpoint beats ten logs.
|
||||
2. **Targeted logs** at the boundaries that distinguish hypotheses.
|
||||
3. Never "log everything and grep".
|
||||
|
||||
**Tag every debug log** with a unique prefix, e.g. `[DEBUG-a4f2]`. Cleanup at the end becomes a single grep. Untagged logs survive; tagged logs die.
|
||||
|
||||
**Perf branch.** For performance regressions, logs are usually wrong. Instead: establish a baseline measurement (timing harness, `performance.now()`, profiler, query plan), then bisect. Measure first, fix second.
|
||||
|
||||
## Phase 5 — Fix + regression test
|
||||
|
||||
Write the regression test **before the fix** — but only if there is a **correct seam** for it.
|
||||
|
||||
A correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site. If the only available seam is too shallow (single-caller test when the bug needs multiple callers, unit test that can't replicate the chain that triggered the bug), a regression test there gives false confidence.
|
||||
|
||||
**If no correct seam exists, that itself is the finding.** Note it. The codebase architecture is preventing the bug from being locked down. Flag this for the next phase.
|
||||
|
||||
If a correct seam exists:
|
||||
|
||||
1. Turn the minimised repro into a failing test at that seam.
|
||||
2. Watch it fail.
|
||||
3. Apply the fix.
|
||||
4. Watch it pass.
|
||||
5. Re-run the Phase 1 feedback loop against the original (un-minimised) scenario.
|
||||
|
||||
## Phase 6 — Cleanup + post-mortem
|
||||
|
||||
Required before declaring done:
|
||||
|
||||
- [ ] Original repro no longer reproduces (re-run the Phase 1 loop)
|
||||
- [ ] Regression test passes (or absence of seam is documented)
|
||||
- [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix)
|
||||
- [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug location)
|
||||
- [ ] The hypothesis that turned out correct is stated in the commit / PR message — so the next debugger learns
|
||||
|
||||
**Then ask: what would have prevented this bug?** If the answer involves architectural change (no good test seam, tangled callers, hidden coupling) hand off to the `/improve-codebase-architecture` skill with the specifics. Make the recommendation **after** the fix is in, not before — you have more information now than when you started.
|
||||
@@ -1,15 +0,0 @@
|
||||
```yaml
|
||||
version: "1.1"
|
||||
updated: 2026-06-21
|
||||
|
||||
when: >-
|
||||
Invoked when the user wants to gather structured reference documentation for a
|
||||
tool, library, or API from MCP documentation indexes or web sources. Typically
|
||||
run before writing a new skill that wraps an external tool, or any time
|
||||
reference files are needed for a topic. Triggered explicitly
|
||||
("/research <topic> <path>") or implicitly when the user asks to look up,
|
||||
gather, or pull docs for a topic before implementing something.
|
||||
|
||||
references:
|
||||
- .agents/skills/context7-mcp/SKILL.md # context7-mcp — MCP source channel integrated at step 2
|
||||
```
|
||||
@@ -1,97 +0,0 @@
|
||||
---
|
||||
name: research
|
||||
description: >-
|
||||
Use when the user wants to research a topic and generate structured reference
|
||||
markdown files. Handles: finding canonical docs for a tool/library/API via
|
||||
Context7 MCP or web sources, reading and deepening into linked pages,
|
||||
organizing extracted content into topic files (overview, installation,
|
||||
configuration, cli-reference, api-reference, examples, troubleshooting). Do
|
||||
NOT use when the user wants to write documentation from existing code or specs
|
||||
(use write-docs), install or manage the neuledge-context MCP server (use
|
||||
neuledge-context), or research a bug/incident (use diagnose).
|
||||
metadata:
|
||||
category: research
|
||||
allowed-tools:
|
||||
- WebSearch
|
||||
- WebFetch
|
||||
- Read
|
||||
- Write
|
||||
- mcp__context7__resolve-library-id
|
||||
- mcp__context7__query-docs
|
||||
model: sonnet
|
||||
---
|
||||
|
||||
<requirements>
|
||||
|
||||
## Required inputs
|
||||
|
||||
- **Topic** — the subject to research (tool, library, API, concept); inferred from user description if clear, ask if ambiguous
|
||||
- **Output path** — directory where reference files will be written; must be provided explicitly — do not infer or default
|
||||
- **Starting URLs** — optional; if provided, skip discovery websearch and read these first
|
||||
|
||||
## Constraints
|
||||
|
||||
- Never write files outside the explicitly provided output path
|
||||
- Skip any default topic file if no relevant content is found for it — do not create empty files
|
||||
- Create additional topic files beyond the default list when content warrants it (e.g. `webhooks.md`, `rate-limits.md`)
|
||||
- Subagents handle parallel source reading and link deepening — the orchestrator writes all files; subagents return summaries only, never write directly
|
||||
- Context7 MCP calls (`resolve-library-id`, `query-docs`) are made only by the orchestrator at step 2 — subagents must not call them
|
||||
- `sources.md` is always written, even if only one source was read
|
||||
- Each topic file must have frontmatter with `topic` and `source_keys`; body is prose only — no inline URLs
|
||||
- Source keys in `sources.md` must be kebab-case slugs: derived from the source domain or page title for web sources; for Context7 sources use `context7-<library-slug>` (e.g. `context7-vercel-next-js`)
|
||||
- Default topic list and file format spec live in `references/` sub-files — read them at step 1
|
||||
|
||||
</requirements>
|
||||
|
||||
<steps>
|
||||
|
||||
## Process
|
||||
|
||||
1. **Scan codebase.** Search the working directory for existing usage of the topic — imports, config files, version pins, existing reference files. Use findings to narrow research scope (e.g. target the version already in use, skip topics already documented). Read `references/topics.md` for the default topic list and `references/file-format.md` for the output file format spec.
|
||||
|
||||
2. **Try Context7.** If the topic is a library, framework, or API and no starting URLs were provided, call `resolve-library-id` with the topic name and the user's question. If a match resolves, call `query-docs` once per default topic area (see `references/topics.md`). Treat each response as a source summary with slug `context7-<library-slug>` (e.g. `context7-vercel-next-js`). A topic area has sufficient content when the Context7 response contains at least one substantive paragraph — not a "no results" message, redirect notice, or header-only boilerplate. Mark covered topic areas — skip their subagent web reads in step 4. If the library does not resolve, or starting URLs were provided (explicit source choice by the user), skip this step entirely.
|
||||
|
||||
3. **Discover sources.** For topics not covered by Context7 (or when no starting URLs were provided and Context7 did not resolve), websearch for canonical documentation (prefer `llms.txt`, developer docs, official API references over tutorials or blog posts). Collect 3–5 candidate URLs before reading any.
|
||||
|
||||
4. **Read sources in parallel.** Spawn one subagent per source URL. Each subagent fetches the page, extracts relevant content, identifies links worth deepening, and returns a structured summary (content by topic area + links to follow). Subagents do not write files.
|
||||
|
||||
5. **Deepen.** For each subagent that returned links worth following, spawn child subagents per branch. Continue until content becomes repetitive or out of scope. Cap at ~10 additional pages total across all branches.
|
||||
|
||||
6. **Consolidate.** Merge all subagent summaries (Context7 and web) by topic area. Identify which default topics have sufficient content and which custom topics emerged.
|
||||
|
||||
7. **Write topic files.** For each topic with content, write `<output-path>/<topic>.md` using the format in `references/file-format.md`. Orchestrator writes all files — never delegate file writing to a subagent.
|
||||
|
||||
8. **Write `sources.md`.** Write `<output-path>/sources.md` mapping each source slug to its URL (use `context7:<library-id>` as the URL for Context7 sources), description, and list of topic files it contributed to. Include sources that yielded no content, marked `no content extracted`.
|
||||
|
||||
## Output format
|
||||
|
||||
- `<output-path>/<topic>.md` per topic with content — formatted per `references/file-format.md`
|
||||
- `<output-path>/sources.md` — always produced; maps slug → URL, description, contributing files
|
||||
|
||||
</steps>
|
||||
|
||||
<checks>
|
||||
|
||||
## Failure handling
|
||||
|
||||
- Output path not provided — stop and ask; do not infer or default
|
||||
- No sources found after websearch — report what was searched, ask user to provide starting URLs
|
||||
- Subagent returns no usable content — skip that source, log in `sources.md` as `no content extracted`
|
||||
- All topic files would be empty — stop, report what was searched, do not write any files
|
||||
|
||||
## Self-check
|
||||
|
||||
- [ ] Codebase scanned before any websearch was performed
|
||||
- [ ] Output path was explicitly provided — not inferred
|
||||
- [ ] `references/topics.md` and `references/file-format.md` read at step 1
|
||||
- [ ] Context7 resolution attempted before websearch when topic is a library/framework/API
|
||||
- [ ] Context7 calls made only at orchestrator step 2 — no subagent called `resolve-library-id` or `query-docs`
|
||||
- [ ] Context7 sources recorded in `sources.md` with `context7:<library-id>` as URL
|
||||
- [ ] No topic file written without content
|
||||
- [ ] `sources.md` written with all sources read (including those with no content extracted)
|
||||
- [ ] All file writes performed by the orchestrator, not subagents
|
||||
- [ ] Each topic file has `topic` and `source_keys` frontmatter fields
|
||||
- [ ] All source keys in topic files have a matching entry in `sources.md`
|
||||
- [ ] No files written outside the provided output path
|
||||
|
||||
</checks>
|
||||
46
plugins/core/.apm/skills/agentsmd-audit/SKILL.md
Normal file
46
plugins/core/.apm/skills/agentsmd-audit/SKILL.md
Normal file
@@ -0,0 +1,46 @@
|
||||
---
|
||||
name: agentsmd-audit
|
||||
description: >
|
||||
Use when the user wants a repo's AGENTS.md audited for secrets, structure
|
||||
and drift — "is this AGENTS.md safe to commit" — or after a hand-edit
|
||||
outside `agentsmd-author`.
|
||||
Not converting a provider file -> `provider-adapter-author`.
|
||||
Not writing AGENTS.md -> `agentsmd-author`.
|
||||
allowed-tools: Bash Read
|
||||
metadata:
|
||||
category: docs
|
||||
source_keys:
|
||||
- agents-md-official
|
||||
- context7-websites-agents-md
|
||||
- context7-agentsmd-agents-md
|
||||
- governance-secrets-hard-prohibition
|
||||
version: "0.1.2"
|
||||
---
|
||||
|
||||
## Gotchas
|
||||
|
||||
- Always run all three checks — this skill does a single combined pass, not staged/gated passes. Don't skip structure or drift checks just because a secrets FAIL was found.
|
||||
- Never inspect or mention provider-specific adapter files (`CLAUDE.md`, `.cursor/rules/*.mdc`, `copilot-instructions.md`, etc.) — that's out of scope.
|
||||
- Gather findings internally; don't narrate PASS/FAIL per check as you go — surface them only in the final report.
|
||||
|
||||
## Step 1 — Run the validators
|
||||
|
||||
```bash
|
||||
bash scripts/validate-secrets.sh <repo-root>
|
||||
bash scripts/validate-structure.sh <repo-root>
|
||||
bash scripts/validate-drift.sh <repo-root>
|
||||
```
|
||||
|
||||
Each script walks the repo for every `AGENTS.md` file (root and nested, excluding `.git`, `node_modules`, `vendor`, and similar) and prints `FAIL` lines, plus `INFO`/`SUGGESTION` where applicable, with `Why`/`Fix` (or `Note`) per finding. A nonzero exit means at least one FAIL was found in that dimension. If a script cannot execute (`python3` unavailable, Bash denied), fall back to manual review: scan for real-looking credentials, check common sections are present, and spot-check a few referenced commands/paths by hand. Grade a manual finding the way the scripts grade theirs: a missing common section (e.g. no "Security" heading) is informational, not a failure — not every repo needs every section from the checklist. Only flag a FAIL when the file is empty, entirely unfilled placeholder text, or contains a real embedded secret/stale reference.
|
||||
|
||||
## Step 2 — Report
|
||||
|
||||
Open with a coverage line:
|
||||
|
||||
```text
|
||||
Checked: secrets · structure · drift
|
||||
```
|
||||
|
||||
Then output only findings that were found, in this order within a repo: `### Secrets`, `### Structure`, `### Drift`. Omit a dimension heading entirely if it produced nothing — its absence confirms it passed. Report each finding verbatim as emitted by the scripts (they already carry file:line, Why/Fix or Note).
|
||||
|
||||
Close with a `## Result` block holding one line: `PASS`, `PASS (N suggestions)`, or `FAIL (N fails · M suggestions)`, each optionally followed by ` · P info`. Omit the suggestion count when there are none, and omit `· P info` when there are none. INFO and SUGGESTION findings are observational — they never flip PASS to FAIL. Do not fix anything — this skill reports and proposes only. Point the user to `agentsmd-author` to apply fixes.
|
||||
@@ -87,14 +87,14 @@ for fpath in find_agents_md(repo_root):
|
||||
with open(fpath, encoding="utf-8", errors="replace") as f:
|
||||
lines = f.readlines()
|
||||
for i, line in enumerate(lines, start=1):
|
||||
if PLACEHOLDER_RE.search(line):
|
||||
continue
|
||||
for label, pattern in PATTERNS:
|
||||
m = pattern.search(line)
|
||||
if not m:
|
||||
continue
|
||||
# Re-check placeholder allowlist against just the matched value, in case
|
||||
# the placeholder marker sits outside the regex's own match span.
|
||||
# Scope the placeholder allowlist to the matched secret-candidate
|
||||
# substring only. Checking the whole line would let an unrelated
|
||||
# placeholder-looking token elsewhere on the line (e.g. in a
|
||||
# trailing comment) suppress detection of a real credential.
|
||||
value = m.group(0)
|
||||
if PLACEHOLDER_RE.search(value):
|
||||
continue
|
||||
@@ -18,7 +18,7 @@ git clone https://github.com/bats-core/bats-assert tests/test_helper/bats-assert
|
||||
Run all tests for this skill (from the repo root):
|
||||
|
||||
```bash
|
||||
bats plugins/core/skills/agentsmd-audit/tests/
|
||||
bats plugins/core/.apm/skills/agentsmd-audit/tests/
|
||||
```
|
||||
|
||||
## Files
|
||||
@@ -1,7 +1,7 @@
|
||||
#!/usr/bin/env bats
|
||||
|
||||
setup() {
|
||||
REPO_ROOT="$(cd "$BATS_TEST_DIRNAME/../../../../../" && pwd)"
|
||||
REPO_ROOT="$(cd "$BATS_TEST_DIRNAME/../../../../../../" && pwd)"
|
||||
load "$REPO_ROOT/tests/test_helper/bats-support/load"
|
||||
load "$REPO_ROOT/tests/test_helper/bats-assert/load"
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
#!/usr/bin/env bats
|
||||
|
||||
setup() {
|
||||
REPO_ROOT="$(cd "$BATS_TEST_DIRNAME/../../../../../" && pwd)"
|
||||
REPO_ROOT="$(cd "$BATS_TEST_DIRNAME/../../../../../../" && pwd)"
|
||||
load "$REPO_ROOT/tests/test_helper/bats-support/load"
|
||||
load "$REPO_ROOT/tests/test_helper/bats-assert/load"
|
||||
|
||||
@@ -52,6 +52,19 @@ EOF
|
||||
assert_output --partial "connection string"
|
||||
}
|
||||
|
||||
@test "still catches a real secret when a placeholder token sits elsewhere on the same line" {
|
||||
cat > "$TMPDIR/AGENTS.md" <<'EOF'
|
||||
# AGENTS.md
|
||||
|
||||
## Setup
|
||||
- AWS_ACCESS_KEY_ID=AKIAABCDEFGHIJKLMNOP # see your-token-here for an example, gitleaks:allow (synthetic fixture — this test verifies the placeholder allowlist is scoped to the matched value, not the whole line)
|
||||
EOF
|
||||
run bash "$SCRIPT" "$TMPDIR"
|
||||
assert_failure
|
||||
assert_output --partial "AWS access key ID"
|
||||
assert_output --partial "AGENTS.md:4"
|
||||
}
|
||||
|
||||
@test "detects secrets in a nested AGENTS.md, not just root" {
|
||||
mkdir -p "$TMPDIR/packages/api"
|
||||
cat > "$TMPDIR/AGENTS.md" <<'EOF'
|
||||
@@ -1,7 +1,7 @@
|
||||
#!/usr/bin/env bats
|
||||
|
||||
setup() {
|
||||
REPO_ROOT="$(cd "$BATS_TEST_DIRNAME/../../../../../" && pwd)"
|
||||
REPO_ROOT="$(cd "$BATS_TEST_DIRNAME/../../../../../../" && pwd)"
|
||||
load "$REPO_ROOT/tests/test_helper/bats-support/load"
|
||||
load "$REPO_ROOT/tests/test_helper/bats-assert/load"
|
||||
|
||||
43
plugins/core/.apm/skills/agentsmd-author/SKILL.md
Normal file
43
plugins/core/.apm/skills/agentsmd-author/SKILL.md
Normal file
@@ -0,0 +1,43 @@
|
||||
---
|
||||
name: agentsmd-author
|
||||
description: >
|
||||
Use when the user wants a repo's AGENTS.md written or updated, root or
|
||||
nested, including "document this for AI coding tools". Writes only verified
|
||||
conventions. Not review-only -> `agentsmd-audit`. Not converting CLAUDE.md ->
|
||||
`provider-adapter-author`.
|
||||
allowed-tools: Bash Read Write Edit
|
||||
metadata:
|
||||
category: docs
|
||||
source_keys:
|
||||
- agents-md-official
|
||||
- context7-websites-agents-md
|
||||
- context7-agentsmd-agents-md
|
||||
version: "0.1.2"
|
||||
---
|
||||
|
||||
## Gotchas
|
||||
|
||||
- Never invent a command. Every line under a setup/test/build section must come from something you actually found in the repo (`package.json` scripts, a `Makefile` target, a CI workflow step, a README). If you can't verify a command, don't include it.
|
||||
- Never write to a provider file yourself, in any circumstance: `CLAUDE.md`, `.cursor/rules/*.mdc`, `.github/copilot-instructions.md` and their equivalents are `provider-adapter-author`'s to own. That holds even when the user asks for one in the same breath as AGENTS.md, and even when the file is merely stale or missing a pointer rather than duplicating anything. Detect it and hand off.
|
||||
|
||||
## Step 1 — Explore the target repo
|
||||
|
||||
Before writing anything, gather real facts: package manager and scripts (`package.json`, `pyproject.toml`, `Cargo.toml`, etc.), a `Makefile` or task runner, CI config (`.github/workflows/`, etc.) for the commands it actually runs, linter/formatter config files, and any existing docs (`README.md`, existing `AGENTS.md`) describing conventions. Note whether any subdirectory looks like its own package with a different stack.
|
||||
|
||||
## Step 2 — Decide placement
|
||||
|
||||
- No `AGENTS.md` at the repo root yet → create one there first, covering whole-repo conventions.
|
||||
- A subdirectory has materially different build/test tooling or conventions than the root → create or update a nested `AGENTS.md` there, scoped to what's different. Convenience is not a reason to create one — without a distinct stack you are duplicating content the root already covers. **Don't repeat root-level content** in a nested file: the nearest-file-wins rule means it is read alone, never merged back with the root.
|
||||
- Otherwise → update the existing file(s) in place.
|
||||
|
||||
## Step 3 — Write or update
|
||||
|
||||
AGENTS.md has no required schema. Use only sections that reflect something real about the repo — never fill in every common-sections-checklist heading just because it exists, because a thin accurate file beats a padded generic one. Read `references/content-guide.md` for section-by-section guidance, a worked example, and what separates useful content from generic padding, before writing.
|
||||
|
||||
## Step 4 — Check for an existing provider file
|
||||
|
||||
Look for `CLAUDE.md`, `.cursor/rules/*.mdc`, `.github/copilot-instructions.md`, or similar in the target repo. If one exists, invoke the `provider-adapter-author` skill on it to reconcile — whether it duplicates content the AGENTS.md you just wrote/updated now owns, or is merely stale or missing a pointer to it. Never edit it yourself in either case.
|
||||
|
||||
## Step 5 — Audit and report
|
||||
|
||||
Invoke the `agentsmd-audit` skill on the target repo root — its validators take a `<repo-root>` and walk the tree for every AGENTS.md themselves; there is no per-file entry point. This closeout is mandatory, not optional, even when the change looks trivial — never sign the work off on your own judgment. Resolve any FAIL findings before considering the work done — re-invoke this skill's own writing steps to fix them, then re-run the audit, same as any other close-the-loop check. Report what was created/changed, whether a provider file was reconciled, and the audit's final result.
|
||||
52
plugins/core/.apm/skills/provider-adapter-author/SKILL.md
Normal file
52
plugins/core/.apm/skills/provider-adapter-author/SKILL.md
Normal file
@@ -0,0 +1,52 @@
|
||||
---
|
||||
name: provider-adapter-author
|
||||
description: >
|
||||
Use when a provider file (CLAUDE.md, .cursor rules, copilot-instructions)
|
||||
duplicating the repo's AGENTS.md should be cut to a thin adapter — "make
|
||||
CLAUDE.md just import AGENTS.md".
|
||||
Not writing the AGENTS file -> `agentsmd-author`.
|
||||
Not auditing the AGENTS file -> `agentsmd-audit`.
|
||||
allowed-tools: Bash Read Edit Write
|
||||
metadata:
|
||||
category: docs
|
||||
source_keys:
|
||||
- adr-0002-0003-two-tier-claude-md
|
||||
version: "0.1.1"
|
||||
---
|
||||
|
||||
## Gotchas
|
||||
|
||||
- Assume a provider has no cross-file import mechanism until you have confirmed it has one. Claude Code is the exception, not the rule: a `CLAUDE.md` may consist of nothing but `@path` lines, while the same `@AGENTS.md` line in a Cursor rule or a Copilot instructions file is inert text no tool resolves. Pass `--no-import-syntax` to `scripts/validate-adapter.sh` for those providers.
|
||||
|
||||
- Works standalone or composed-into by `agentsmd-author` — behave identically either way; do not assume a caller skill exists. Detect the provider file, confirm `AGENTS.md`, and run the closeout validator yourself in both cases (`references/provider-matrix.md`).
|
||||
|
||||
## Step 1 — Detect
|
||||
|
||||
Find the provider instruction file to convert. Before searching, read `references/provider-matrix.md` — skip it only when the target is already a known root `CLAUDE.md`, which is the common case.
|
||||
|
||||
Then confirm `AGENTS.md` exists at the repo root. If it does not, stop and tell the user to run `agentsmd-author` first — there is nothing to adapt to.
|
||||
|
||||
## Step 2 — Diff and rewrite
|
||||
|
||||
Read the provider file and `AGENTS.md` side by side. Separate the provider file's content into two buckets: lines that restate what `AGENTS.md` already owns (universal rules, conventions, project overview) versus lines that are genuinely provider-specific (tool syntax, IDE behavior, model-specific instructions). Rewrite the provider file:
|
||||
|
||||
- **Providers with import syntax** (Claude Code): replace the redundant bucket with an `@AGENTS.md` (or correct relative path) import on a line of its own, keep the provider-specific bucket below it. An import folded into a sentence is not the thin-adapter shape and `scripts/validate-adapter.sh` will not credit it — nor one inside a code fence, an indented block, or an HTML comment, nor one whose path does not resolve to a real, non-empty file on disk.
|
||||
- **Providers without import syntax** (Cursor, Copilot, etc.): replace the redundant bucket with a short sentence pointing at `AGENTS.md` ("See AGENTS.md at the repo root for ..."), keep the provider-specific bucket. A bare or negated mention is not a pointer and will not be credited.
|
||||
|
||||
The provider file is the only file this skill ever writes. Never create or edit `AGENTS.md` — not in this step, not in any step, whatever the payoff looks like.
|
||||
|
||||
Strip only what is genuinely redundant. Provider-specific material stays even when it is short — the goal is thin, not empty.
|
||||
|
||||
## Step 3 — Self-validate
|
||||
|
||||
Run the bundled check before finishing — this is the skill's own closeout gate; there is no separate paired audit skill for this concern:
|
||||
|
||||
```bash
|
||||
bash scripts/validate-adapter.sh [--no-import-syntax] [--max-lines N] <adapter-file> <agents-md-file>
|
||||
```
|
||||
|
||||
Fix any `FAIL` by editing the provider file, and re-run until it exits `0`. Exits `2` and `3` are not `FAIL`s and nothing was graded under either, so neither is a reason to touch the adapter: `2` means the invocation or the input is wrong (a bad, missing, or extra argument, an unknown option, or a file that is not UTF-8), and `3` means a named file exists but could not be read.
|
||||
|
||||
## Step 4 — Report
|
||||
|
||||
State which file was converted, what was removed versus kept, and the validator's final result.
|
||||
@@ -0,0 +1,32 @@
|
||||
---
|
||||
source_keys:
|
||||
- adr-0002-0003-two-tier-claude-md
|
||||
---
|
||||
|
||||
# Known provider instruction files
|
||||
|
||||
Which files to look for when detecting a provider-specific instruction file, whether each provider
|
||||
resolves a cross-file import, and what a thin adapter therefore looks like for it.
|
||||
|
||||
| Provider | File(s) | Import syntax | Thin adapter shape | Validator flag |
|
||||
|---|---|---|---|---|
|
||||
| Claude Code | `CLAUDE.md` at the repo root, plus any deployed copies | Yes — `@path` lines, e.g. `@AGENTS.md` | One or more `@` import lines; no other content is required | none |
|
||||
| Cursor | `.cursor/rules/*.mdc` | No | A short sentence pointing at `AGENTS.md`, plus the rule's own frontmatter and provider-specific body | `--no-import-syntax` |
|
||||
| GitHub Copilot | `.github/copilot-instructions.md` | No | A short sentence pointing at `AGENTS.md`, plus Copilot-only instructions | `--no-import-syntax` |
|
||||
| Anything else | tool-specific instruction file at whatever path the tool documents | Assume no | Text pointer, as above | `--no-import-syntax` |
|
||||
|
||||
A provider not listed here is not evidence it has an import mechanism. Confirm against that tool's
|
||||
own documentation before emitting an `@`-style line; an unresolved import reads as literal text and
|
||||
silently drops every rule the adapter was supposed to defer to.
|
||||
|
||||
Detection is a search, not a lookup: a repo may hold more than one of these, and each one converts
|
||||
independently against the same `AGENTS.md`.
|
||||
|
||||
## Standalone and composed runs behave identically
|
||||
|
||||
This skill is reached two ways: invoked directly by a user, and composed into by `agentsmd-author`
|
||||
once it has written or updated the repo's `AGENTS.md`. Behave identically either way — do not
|
||||
assume a caller skill exists. Detect the provider file yourself, confirm `AGENTS.md` yourself, and
|
||||
run the closeout validator yourself, rather than treating any step as already done by the caller or
|
||||
as something the caller will do afterwards. No handshake exists to rely on, and no state is
|
||||
passed in beyond the file paths.
|
||||
@@ -5,5 +5,5 @@
|
||||
- **URL:** (in-repo precedent — not an external source or plugin research corpus entry)
|
||||
- **Description:** This repo's own two-tier CLAUDE.md/AGENTS.md pattern: AGENTS.md is the provider-agnostic source of always-on rules; provider-specific files (CLAUDE.md) become thin adapters that import it (`@AGENTS.md` plus provider-specific additions). Grounds this skill's entire adapter-conversion design — the "thin adapter" shape, the `@`-import convention, and the size/duplication expectations enforced by `scripts/validate-adapter.sh`.
|
||||
- **Research doc:** docs/adr/0002-two-tier-claude-md.md, docs/adr/0003-agents-md-provider-agnostic-entry-point.md, providers/claude-code/CLAUDE.md (in-repo ADRs and a live example, not a plugin research corpus entry; referenced here since this skill's design is modeled directly on an existing implementation rather than external research)
|
||||
- **Contributing files:** SKILL.md
|
||||
- **Contributing files:** SKILL.md, references/provider-matrix.md
|
||||
- **Status:** `extracted`
|
||||
@@ -0,0 +1,28 @@
|
||||
# scripts/
|
||||
|
||||
Deterministic self-check this skill shells out to instead of relying on LLM judgment for a mechanical check.
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `validate-adapter.sh` | Checks a rewritten provider file (CLAUDE.md, etc.) has a working reference to AGENTS.md, doesn't duplicate its content, and stays under a thin-file line threshold |
|
||||
|
||||
Takes exactly `<adapter-file> <agents-md-file>`, with optional `--no-import-syntax` and `--max-lines N` flags (also accepted as `--max-lines=N`). A third positional argument or an unknown option is an error, not something quietly ignored.
|
||||
|
||||
## What counts as a reference to AGENTS.md
|
||||
|
||||
Both modes require the named path to be a real path segment ending in `AGENTS.md` — `AGENTS.md` or `…/AGENTS.md`, not `NOTAGENTS.md` — that resolves on disk, relative to the adapter file, to a non-empty file. An adapter deferring to a path that is not there defers to nothing, so the check has to touch the disk rather than pattern-match the line.
|
||||
|
||||
A reference only counts where something would actually resolve it. A line inside a fenced code block, an indented code block, or an HTML comment is not credited in either mode: Claude Code resolves an import in none of those, so a fenced `@AGENTS.md` is the silent-drop failure this gate exists to catch, not a pass.
|
||||
|
||||
Default mode wants a real import: `@AGENTS.md` alone on its own line, indented no more than three spaces. `--no-import-syntax` wants a prose pointer that reads as one — the sentence naming `AGENTS.md` must carry a deference cue (see, read, refer to, documented in, conventions, …) and must not be negated. `Do NOT read AGENTS.md; it is obsolete.` and `We deleted AGENTS.md last year.` name the file while pointing the reader away from it, and neither is a pointer.
|
||||
|
||||
## Exit codes
|
||||
|
||||
The distinction matters because the skill's closeout tells the agent to fix any non-zero exit by editing the provider file. That is right for exactly one of these.
|
||||
|
||||
| Code | Meaning | What to do |
|
||||
|------|---------|------------|
|
||||
| `0` | Passes every check | Nothing |
|
||||
| `1` | One or more `FAIL` findings printed to stdout — empty adapter, no working reference to AGENTS.md, excessive duplication, or not thin | Edit the provider file |
|
||||
| `2` | Usage or input error: a bad, missing, or extra argument, an unknown option, a path that is not a file, or a file that is not UTF-8. Nothing was graded, so there is no `FAIL` line | Fix the invocation or the file's encoding — do not edit the adapter |
|
||||
| `3` | A named input file exists but could not be read (permissions, I/O error). Nothing was graded and the adapter's contents are unknown | Fix the file's readability — do not edit the adapter |
|
||||
496
plugins/core/.apm/skills/provider-adapter-author/scripts/validate-adapter.sh
Executable file
496
plugins/core/.apm/skills/provider-adapter-author/scripts/validate-adapter.sh
Executable file
@@ -0,0 +1,496 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
usage() {
|
||||
cat <<EOF
|
||||
Usage: validate-adapter.sh [--no-import-syntax] [--max-lines N] <adapter-file> <agents-md-file>
|
||||
|
||||
Self-check gate for provider-adapter-author. Checks that a rewritten
|
||||
provider-specific instruction file (CLAUDE.md, .cursor/rules/*.mdc,
|
||||
copilot-instructions.md, etc.) is actually a thin adapter over AGENTS.md,
|
||||
not a duplicate copy of it.
|
||||
|
||||
Arguments:
|
||||
adapter-file Path to the provider-specific file to check.
|
||||
agents-md-file Path to the AGENTS.md file it should defer to.
|
||||
|
||||
Exactly two positional arguments are accepted. Extra ones are rejected
|
||||
rather than ignored: a third path silently graded nothing but the first
|
||||
two, so a typo'd invocation passed against the wrong file.
|
||||
|
||||
Options:
|
||||
--no-import-syntax The target provider has no native cross-file import
|
||||
mechanism. Require a plain-text pointer line naming
|
||||
"AGENTS.md" instead of an @import-style line; an
|
||||
@AGENTS.md line alone does not satisfy it, because
|
||||
such a provider never resolves it. Without this flag
|
||||
an actual @import line is required, and naming
|
||||
AGENTS.md in prose alone does not satisfy it.
|
||||
--max-lines N Max non-blank lines allowed in the adapter file before
|
||||
it's considered no longer "thin". Must be a
|
||||
non-negative integer. Default: 60.
|
||||
--help, -h Show this help and exit 0.
|
||||
-- End of options; every later argument is positional.
|
||||
|
||||
Both flags also accept the --flag=value form (--max-lines=40). An unknown
|
||||
option is reported as an unknown option, not as a missing file.
|
||||
|
||||
What counts as a reference:
|
||||
|
||||
In both modes the named path must be a real path segment ending in
|
||||
AGENTS.md ("AGENTS.md" or ".../AGENTS.md" — not NOTAGENTS.md), and it must
|
||||
resolve on disk, relative to the adapter file, to a non-empty file. An
|
||||
adapter deferring to a path that is not there defers to nothing.
|
||||
|
||||
A mention inside a fenced code block, an indented code block, or an HTML
|
||||
comment is not credited in either mode. Nothing resolves those, so an
|
||||
adapter whose only "import" is fenced silently defers to nothing.
|
||||
|
||||
With --no-import-syntax the pointer must read as a pointer: the sentence
|
||||
naming AGENTS.md has to carry a deference cue (see, read, refer to,
|
||||
documented in, conventions, ...) and must not be a negation ("do not read
|
||||
AGENTS.md", "we deleted AGENTS.md"). A bare mention is not a pointer.
|
||||
|
||||
Exit codes:
|
||||
0 Adapter file passes all checks
|
||||
1 One or more checks failed (empty file, no reference to AGENTS.md,
|
||||
excessive duplication, or file too long)
|
||||
2 Usage or input error — a bad, missing, or extra argument, an unknown
|
||||
option, a path that is not a file, or a file that is not UTF-8. Nothing
|
||||
was graded, so there is no FAIL line and no adapter edit to make: fix
|
||||
the invocation or the file's encoding and re-run. Kept distinct from 1
|
||||
because the skill's own closeout tells the agent to fix every non-zero
|
||||
exit by editing the provider file, which for a mistyped flag edits the
|
||||
wrong file forever.
|
||||
3 A named input file exists but could not be read (permissions, a
|
||||
directory swapped in mid-run, I/O error). Also not a FAIL: nothing was
|
||||
graded and the adapter's contents are unknown, so editing it is
|
||||
guesswork. Fix the file's readability and re-run.
|
||||
EOF
|
||||
}
|
||||
|
||||
NO_IMPORT_SYNTAX=0
|
||||
MAX_LINES=60
|
||||
ARGS=()
|
||||
END_OF_OPTS=0
|
||||
|
||||
require_int() {
|
||||
# $1 = the value to validate
|
||||
if [[ ! "$1" =~ ^[0-9]+$ ]]; then
|
||||
echo "Error: --max-lines expects a non-negative integer, got '$1'." >&2
|
||||
exit 2
|
||||
fi
|
||||
}
|
||||
|
||||
while [[ $# -gt 0 ]]; do
|
||||
if [[ $END_OF_OPTS -eq 1 ]]; then
|
||||
ARGS+=("$1")
|
||||
shift
|
||||
continue
|
||||
fi
|
||||
case "$1" in
|
||||
--)
|
||||
END_OF_OPTS=1
|
||||
shift
|
||||
;;
|
||||
--help|-h)
|
||||
usage
|
||||
exit 0
|
||||
;;
|
||||
--no-import-syntax)
|
||||
NO_IMPORT_SYNTAX=1
|
||||
shift
|
||||
;;
|
||||
--no-import-syntax=*)
|
||||
echo "Error: --no-import-syntax is a flag and takes no value (got '$1')." >&2
|
||||
exit 2
|
||||
;;
|
||||
--max-lines)
|
||||
if [[ $# -lt 2 ]]; then
|
||||
echo "Error: --max-lines requires a value (a non-negative integer)." >&2
|
||||
exit 2
|
||||
fi
|
||||
MAX_LINES="$2"
|
||||
require_int "$MAX_LINES"
|
||||
shift 2
|
||||
;;
|
||||
--max-lines=*)
|
||||
MAX_LINES="${1#--max-lines=}"
|
||||
if [[ -z "$MAX_LINES" ]]; then
|
||||
echo "Error: --max-lines requires a value (a non-negative integer)." >&2
|
||||
exit 2
|
||||
fi
|
||||
require_int "$MAX_LINES"
|
||||
shift
|
||||
;;
|
||||
-*)
|
||||
# Reported as an unknown option rather than falling through to the
|
||||
# positional bucket, where it used to surface as "'--bogus' is not a
|
||||
# file" — the right exit code attached to a diagnostic that sends the
|
||||
# reader looking for a path they never typed.
|
||||
echo "Error: unknown option '$1'." >&2
|
||||
echo "" >&2
|
||||
usage >&2
|
||||
exit 2
|
||||
;;
|
||||
*)
|
||||
ARGS+=("$1")
|
||||
shift
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
if [[ ${#ARGS[@]} -lt 2 ]]; then
|
||||
echo "Error: adapter-file and agents-md-file are required." >&2
|
||||
echo "" >&2
|
||||
usage >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
if [[ ${#ARGS[@]} -gt 2 ]]; then
|
||||
echo "Error: expected exactly 2 positional arguments (adapter-file and agents-md-file), got ${#ARGS[@]}: ${ARGS[*]}." >&2
|
||||
echo "" >&2
|
||||
usage >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
python3 -u - "${ARGS[0]}" "${ARGS[1]}" "$NO_IMPORT_SYNTAX" "$MAX_LINES" <<'PYTHON'
|
||||
import sys
|
||||
import os
|
||||
import re
|
||||
|
||||
adapter_path, agents_md_path, no_import_syntax, max_lines = sys.argv[1:5]
|
||||
no_import_syntax = no_import_syntax == "1"
|
||||
max_lines = int(max_lines)
|
||||
|
||||
EXIT_FAIL = 1
|
||||
EXIT_USAGE = 2
|
||||
EXIT_UNREADABLE = 3
|
||||
|
||||
if not os.path.isfile(adapter_path):
|
||||
print(f"Error: '{adapter_path}' is not a file.", file=sys.stderr)
|
||||
sys.exit(EXIT_USAGE)
|
||||
if not os.path.isfile(agents_md_path):
|
||||
print(f"Error: '{agents_md_path}' is not a file.", file=sys.stderr)
|
||||
sys.exit(EXIT_USAGE)
|
||||
|
||||
|
||||
def read_text(path):
|
||||
r"""File contents as text, UTF-8, every BOM stripped.
|
||||
|
||||
The BOM strip is not cosmetic. IMPORT_RE anchors on `^ {0,3}@`, and a BOM
|
||||
is not whitespace in Python, so a CLAUDE.md saved by an editor that emits
|
||||
one had its first line — the `@AGENTS.md` import, which is the whole
|
||||
adapter — silently treated as prose. The check then said "no reference to
|
||||
AGENTS.md" and told the author to add the line already sitting in front of
|
||||
them. Same class of silent BOM miss recorded in scripts/skill-size-check.sh;
|
||||
strip it at the reader so no later check has to know about it.
|
||||
|
||||
Every U+FEFF goes, not just one at offset 0. Stripping exactly the first
|
||||
one left the mirror-image false FAIL for a doubled BOM (two concatenated
|
||||
files, or a tool that re-adds one) and for a BOM mid-file at the head of
|
||||
the import line. U+FEFF has no meaning as a character in a markdown
|
||||
instruction file, so removing all of them cannot lose signal.
|
||||
|
||||
Decoding is strict, not errors="replace". Replacement mangles the file and
|
||||
the checks then grade the mangling: a UTF-16 adapter whose first line is
|
||||
`@AGENTS.md` decoded to interleaved NULs and failed as "no reference",
|
||||
which is a true FAIL for a false reason and points the fix at the wrong
|
||||
thing. But strict UTF-8 alone does not catch it — BOM-less UTF-16LE/BE and
|
||||
UTF-32LE are *valid* UTF-8, because NUL is a legal code point, so they
|
||||
decoded clean and produced exactly that false diagnosis anyway. The NUL
|
||||
byte is the complete signal and is checked first: no plausible markdown
|
||||
adapter contains one, and every UTF-16/32 encoding of ASCII is full of
|
||||
them. A file this gate cannot read gets an encoding diagnostic and exit 2,
|
||||
the same policy the ADR-0020 validators' read_text() uses.
|
||||
|
||||
A file that exists but cannot be read at all is neither a pass nor a FAIL —
|
||||
nothing was graded — so it exits 3 rather than 1. Exit 1 sends the skill's
|
||||
closeout into "fix the FAIL by editing the provider file", which for a file
|
||||
it cannot open is an instruction to edit blind.
|
||||
"""
|
||||
try:
|
||||
with open(path, "rb") as fh:
|
||||
raw = fh.read()
|
||||
except OSError as exc:
|
||||
print(f"Error: '{path}' exists but could not be read ({exc.strerror}). "
|
||||
"Nothing was checked — fix whatever is blocking the read "
|
||||
"(permissions, ownership, the underlying device) and re-run; do "
|
||||
"not edit the adapter on the strength of this.", file=sys.stderr)
|
||||
sys.exit(EXIT_UNREADABLE)
|
||||
if b"\x00" in raw:
|
||||
print(f"Error: '{path}' is not valid UTF-8 — it contains NUL bytes, so "
|
||||
"it is almost certainly UTF-16 or UTF-32 (with or without a BOM). "
|
||||
"Re-save it as UTF-8; this check does not guess at other "
|
||||
"encodings.", file=sys.stderr)
|
||||
sys.exit(EXIT_USAGE)
|
||||
try:
|
||||
text = raw.decode("utf-8")
|
||||
except UnicodeDecodeError as exc:
|
||||
print(f"Error: '{path}' is not valid UTF-8 ({exc.reason} at byte "
|
||||
f"{exc.start}) — re-save it as UTF-8; this check does not guess "
|
||||
"at other encodings.", file=sys.stderr)
|
||||
sys.exit(EXIT_USAGE)
|
||||
return text.replace("\ufeff", "")
|
||||
|
||||
|
||||
adapter_content = read_text(adapter_path)
|
||||
agents_md_content = read_text(agents_md_path)
|
||||
adapter_dir = os.path.dirname(os.path.abspath(adapter_path))
|
||||
|
||||
has_fail = False
|
||||
|
||||
if not adapter_content.strip():
|
||||
print(f"FAIL Adapter file is empty — {adapter_path}")
|
||||
print(" Why: An empty adapter carries no reference to AGENTS.md and no provider-specific content.")
|
||||
print(" Fix: Add at least an import (or text pointer) to AGENTS.md.")
|
||||
print()
|
||||
sys.exit(EXIT_FAIL)
|
||||
|
||||
|
||||
# --- Inert regions -----------------------------------------------------------
|
||||
#
|
||||
# A reference only counts where something would actually resolve it. Fenced
|
||||
# code blocks, indented code blocks and HTML comments are shown to the reader
|
||||
# (or hidden from them) as literal text; Claude Code resolves an @import in
|
||||
# none of them. Without this, a ```-fenced `@AGENTS.md` — the exact
|
||||
# copy-the-example-into-the-file mistake this gate exists to catch — exited 0
|
||||
# with the adapter deferring to nothing.
|
||||
#
|
||||
# Indented code blocks are handled by IMPORT_RE's `^ {0,3}` instead of by the
|
||||
# mask: four leading spaces is what opens an indented code block in CommonMark,
|
||||
# so an import has to sit within three. The mask deliberately does not apply
|
||||
# that rule to prose pointers, where four-space indentation is ordinary list
|
||||
# continuation rather than code.
|
||||
FENCE_RE = re.compile(r'^( {0,3})(`{3,}|~{3,})(.*)$')
|
||||
COMMENT_RE = re.compile(r'<!--.*?(?:-->|\Z)', re.DOTALL)
|
||||
|
||||
|
||||
def line_offsets(text):
|
||||
"""[(char offset, line without its terminator)] over `text`."""
|
||||
out = []
|
||||
off = 0
|
||||
for raw in text.splitlines(keepends=True):
|
||||
out.append((off, raw.rstrip("\r\n")))
|
||||
off += len(raw)
|
||||
return out
|
||||
|
||||
|
||||
def build_inert_mask(text, offsets):
|
||||
"""Per-character flags: 1 where a reference would never be resolved."""
|
||||
mask = bytearray(len(text))
|
||||
fence = None # (fence char, opening run length)
|
||||
for start, line in offsets:
|
||||
m = FENCE_RE.match(line)
|
||||
if fence is None:
|
||||
if m:
|
||||
fence = (m.group(2)[0], len(m.group(2)))
|
||||
for i in range(start, start + len(line)):
|
||||
mask[i] = 1
|
||||
continue
|
||||
for i in range(start, start + len(line)):
|
||||
mask[i] = 1
|
||||
if (m and m.group(2)[0] == fence[0]
|
||||
and len(m.group(2)) >= fence[1]
|
||||
and not m.group(3).strip()):
|
||||
fence = None
|
||||
for m in COMMENT_RE.finditer(text):
|
||||
if m.start() < len(mask) and mask[m.start()]:
|
||||
continue # a literal "<!--" printed inside a fence opens nothing
|
||||
for i in range(m.start(), min(m.end(), len(mask))):
|
||||
mask[i] = 1
|
||||
return mask
|
||||
|
||||
|
||||
# --- Reference shapes --------------------------------------------------------
|
||||
#
|
||||
# `\S*AGENTS\.md` had no path-separator boundary, so `@NOTAGENTS.md` and
|
||||
# `@zzzAGENTS.md` counted as imports of AGENTS.md. The matched path must end in
|
||||
# AGENTS.md as a whole segment.
|
||||
IMPORT_RE = re.compile(r'^ {0,3}@(?P<path>\S+?)\s*$')
|
||||
# A mention in prose: an optional relative path, then AGENTS.md, with no
|
||||
# identifier character glued to the front (so NOTAGENTS.md does not match) and
|
||||
# nothing glued to the back.
|
||||
MENTION_RE = re.compile(r'(?<![0-9A-Za-z_.\-/])((?:[\w.\-~]+/)*AGENTS\.md)(?![0-9A-Za-z])')
|
||||
|
||||
# A pointer has to read as a pointer. `"AGENTS.md" in ln` passed
|
||||
# "Do NOT read AGENTS.md; it is obsolete." and "We deleted AGENTS.md last
|
||||
# year." — both of which point the reader away from the file. Require a
|
||||
# deference cue in the naming sentence, and reject a negated one.
|
||||
DIRECTIVE_RE = re.compile(
|
||||
r'\b(see|read|refer|refers|referring|consult|consults|follow|follows|'
|
||||
r'defer|defers|deferring|described|documented|documents|covered|covers|'
|
||||
r'found|listed|specified|defined|governed|per|use|uses|using|apply|obey|'
|
||||
r'start|check|live|lives|contains|holds|carries|inherit|inherits|import|'
|
||||
r'imports|conventions|instructions|guidelines|guidance|rules|standards|'
|
||||
r'reference|setup)\b', re.I)
|
||||
NEGATION_RE = re.compile(
|
||||
r"(\bnot\b|n't\b|\bnever\b|\bno longer\b|\bdeleted\b|\bremoved\b|"
|
||||
r"\bobsolete\b|\bdeprecated\b|\bignore\b|\bignores\b|\bignoring\b|"
|
||||
r"\bdisregard\b|\bsuperseded\b|\bgone\b|\bunused\b|\bstale\b)", re.I)
|
||||
SENTENCE_SPLIT_RE = re.compile(r'(?<=[.;:!?])\s+')
|
||||
|
||||
|
||||
def sentence_around(line, index):
|
||||
"""(sentence of `line` containing character `index`, its start offset)."""
|
||||
bounds = [0]
|
||||
for m in SENTENCE_SPLIT_RE.finditer(line):
|
||||
bounds.append(m.end())
|
||||
bounds.append(len(line) + 1)
|
||||
for i in range(len(bounds) - 1):
|
||||
if bounds[i] <= index < bounds[i + 1]:
|
||||
return line[bounds[i]:bounds[i + 1]], bounds[i]
|
||||
return line, 0
|
||||
|
||||
|
||||
def reads_as_pointer(line, match):
|
||||
"""Does the sentence naming AGENTS.md actually point the reader at it?
|
||||
|
||||
The matched path is blanked out before the cues are applied. It is a
|
||||
filename, not prose, and leaving it in let its own characters vote: the
|
||||
perfectly ordinary `docs/does/not/exist/AGENTS.md` tripped the negation
|
||||
cue on the `not` path segment, so a pointer got rejected for the wrong
|
||||
reason and the near-miss line then reported the wrong diagnosis.
|
||||
"""
|
||||
sentence, sentence_start = sentence_around(line, match.start())
|
||||
rel_start = match.start() - sentence_start
|
||||
rel_end = match.end() - sentence_start
|
||||
probe = sentence[:rel_start] + " AGENTS.md " + sentence[rel_end:]
|
||||
if NEGATION_RE.search(probe):
|
||||
return False
|
||||
return bool(DIRECTIVE_RE.search(probe))
|
||||
|
||||
|
||||
def resolve(raw_path):
|
||||
"""An import/pointer path resolved the way the provider would resolve it."""
|
||||
p = os.path.expanduser(raw_path)
|
||||
if not os.path.isabs(p):
|
||||
p = os.path.join(adapter_dir, p)
|
||||
return os.path.normpath(p)
|
||||
|
||||
|
||||
def target_problem(raw_path):
|
||||
"""None if `raw_path` names a real, non-empty file; else why not."""
|
||||
resolved = resolve(raw_path)
|
||||
if not os.path.isfile(resolved):
|
||||
return f"'{raw_path}' resolves to {resolved}, which does not exist"
|
||||
try:
|
||||
if os.path.getsize(resolved) == 0:
|
||||
return f"'{raw_path}' resolves to {resolved}, which is empty"
|
||||
with open(resolved, "rb") as fh:
|
||||
if not fh.read().strip():
|
||||
return f"'{raw_path}' resolves to {resolved}, which is blank"
|
||||
except OSError as exc:
|
||||
return f"'{raw_path}' resolves to {resolved}, which cannot be read ({exc.strerror})"
|
||||
return None
|
||||
|
||||
|
||||
def names_agents_md(path):
|
||||
return path == "AGENTS.md" or path.endswith("/AGENTS.md")
|
||||
|
||||
|
||||
offsets = line_offsets(adapter_content)
|
||||
lines = [line for _, line in offsets]
|
||||
mask = build_inert_mask(adapter_content, offsets)
|
||||
|
||||
|
||||
def is_inert(abs_index):
|
||||
return abs_index < len(mask) and bool(mask[abs_index])
|
||||
|
||||
|
||||
# Lines shaped like an @AGENTS.md import, whether or not the target resolves.
|
||||
# Used to exclude them from the duplication denominator and from the prose
|
||||
# pointer scan, both of which only care about the shape.
|
||||
import_shaped_lines = set()
|
||||
# (line, raw path) for every import whose target actually resolves.
|
||||
live_imports = []
|
||||
# Diagnostics for imports that are the right shape but resolve to nothing.
|
||||
dead_imports = []
|
||||
# Imports that exist only inside a fence or an HTML comment.
|
||||
inert_imports = []
|
||||
|
||||
for start, line in offsets:
|
||||
m = IMPORT_RE.match(line)
|
||||
if not m or not names_agents_md(m.group("path")):
|
||||
continue
|
||||
at_index = start + line.index("@")
|
||||
if is_inert(at_index):
|
||||
inert_imports.append(line.strip())
|
||||
continue
|
||||
import_shaped_lines.add(line)
|
||||
problem = target_problem(m.group("path"))
|
||||
if problem:
|
||||
dead_imports.append(problem)
|
||||
else:
|
||||
live_imports.append(line)
|
||||
|
||||
live_pointers = []
|
||||
dead_pointers = []
|
||||
inert_pointers = []
|
||||
mention_only = []
|
||||
|
||||
for start, line in offsets:
|
||||
if line in import_shaped_lines:
|
||||
continue
|
||||
for m in MENTION_RE.finditer(line):
|
||||
if is_inert(start + m.start()):
|
||||
inert_pointers.append(line.strip())
|
||||
continue
|
||||
if not reads_as_pointer(line, m):
|
||||
mention_only.append(sentence_around(line, m.start())[0].strip())
|
||||
continue
|
||||
problem = target_problem(m.group(1))
|
||||
if problem:
|
||||
dead_pointers.append(problem)
|
||||
else:
|
||||
live_pointers.append(line)
|
||||
|
||||
if no_import_syntax:
|
||||
has_reference = bool(live_pointers)
|
||||
near_misses = dead_pointers + [f"{d} (inside a code fence or HTML comment)" for d in inert_pointers]
|
||||
near_misses += [f"'{s}' names AGENTS.md but does not point at it" for s in mention_only]
|
||||
else:
|
||||
has_reference = bool(live_imports)
|
||||
near_misses = dead_imports + [f"'{d}' is inside a code fence or HTML comment, where no import is resolved" for d in inert_imports]
|
||||
|
||||
if not has_reference:
|
||||
has_fail = True
|
||||
print(f"FAIL Adapter has no reference to AGENTS.md — {adapter_path}")
|
||||
if no_import_syntax:
|
||||
print(" Why: This provider resolves no cross-file import, so the adapter must point at AGENTS.md in prose; an `@AGENTS.md` line here is inert text. The pointer has to read as a pointer and name a file that is really there — a bare or negated mention (\"we deleted AGENTS.md\") defers nothing, and neither does a mention buried in a code fence or an HTML comment.")
|
||||
print(" Fix: Add a sentence like \"See AGENTS.md at the repo root for shared conventions.\", outside any fence, naming a path that exists relative to this file.")
|
||||
else:
|
||||
print(" Why: A thin adapter must import AGENTS.md with an `@AGENTS.md` line of its own, indented no more than three spaces, and the path must resolve to a real non-empty file. Naming the file mid-sentence or inside backticks is prose this check will not credit; putting the line inside a ``` fence, an indented code block, or an HTML comment is worse, because nothing resolves it and it looks right.")
|
||||
print(" Fix: Put `@AGENTS.md` (or the equivalent relative path) alone on its own line at the top level of the file, or pass --no-import-syntax if this provider resolves no imports.")
|
||||
for miss in near_misses:
|
||||
print(f" Near miss: {miss}")
|
||||
print()
|
||||
|
||||
# --- Duplication check ---
|
||||
non_import_lines = [ln for ln in lines if ln not in import_shaped_lines]
|
||||
adapter_lines = [ln.strip() for ln in non_import_lines if ln.strip()]
|
||||
agents_lines = {ln.strip() for ln in agents_md_content.splitlines() if ln.strip()}
|
||||
|
||||
if adapter_lines:
|
||||
overlap = sum(1 for ln in adapter_lines if ln in agents_lines)
|
||||
ratio = overlap / len(adapter_lines)
|
||||
if ratio > 0.3:
|
||||
has_fail = True
|
||||
print(f"FAIL Adapter duplicates AGENTS.md content — {adapter_path}")
|
||||
print(f" Why: {ratio:.0%} of the adapter's non-import lines already appear verbatim in AGENTS.md. A thin adapter should import shared content, not restate it.")
|
||||
print(" Fix: Remove the duplicated lines and rely on the AGENTS.md import (or pointer) instead.")
|
||||
print()
|
||||
|
||||
# --- Size check ---
|
||||
non_blank_count = len([ln for ln in lines if ln.strip()])
|
||||
if non_blank_count > max_lines:
|
||||
has_fail = True
|
||||
print(f"FAIL Adapter is not thin — {adapter_path}")
|
||||
print(f" Why: {non_blank_count} non-blank lines exceeds the {max_lines}-line threshold for a thin adapter.")
|
||||
print(" Fix: Delete the lines already covered by AGENTS.md; keep only genuinely provider-specific additions here.")
|
||||
print()
|
||||
|
||||
if has_fail:
|
||||
sys.exit(EXIT_FAIL)
|
||||
sys.exit(0)
|
||||
PYTHON
|
||||
@@ -18,7 +18,7 @@ git clone https://github.com/bats-core/bats-assert tests/test_helper/bats-assert
|
||||
Run all tests for this skill (from the repo root):
|
||||
|
||||
```bash
|
||||
bats plugins/core/skills/provider-adapter-author/tests/
|
||||
bats plugins/core/.apm/skills/provider-adapter-author/tests/
|
||||
```
|
||||
|
||||
## Files
|
||||
@@ -0,0 +1,519 @@
|
||||
#!/usr/bin/env bats
|
||||
|
||||
setup() {
|
||||
REPO_ROOT="$(cd "$BATS_TEST_DIRNAME/../../../../../../" && pwd)"
|
||||
load "$REPO_ROOT/tests/test_helper/bats-support/load"
|
||||
load "$REPO_ROOT/tests/test_helper/bats-assert/load"
|
||||
|
||||
SCRIPT="$(cd "$BATS_TEST_DIRNAME/../scripts" && pwd)/validate-adapter.sh"
|
||||
TMPDIR="$(mktemp -d)"
|
||||
AGENTS_MD="$TMPDIR/AGENTS.md"
|
||||
cat > "$AGENTS_MD" <<'EOF'
|
||||
# AGENTS.md
|
||||
|
||||
## Setup commands
|
||||
- Install deps: `pnpm install`
|
||||
- Run tests: `pnpm test`
|
||||
|
||||
## Code style
|
||||
- TypeScript strict mode, single quotes, no semicolons.
|
||||
EOF
|
||||
}
|
||||
|
||||
teardown() {
|
||||
rm -rf "$TMPDIR"
|
||||
}
|
||||
|
||||
@test "fails when the adapter file is empty" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
: > "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure
|
||||
assert_output --partial "empty"
|
||||
}
|
||||
|
||||
@test "fails when the adapter has no reference to AGENTS.md" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
# Claude-specific notes
|
||||
Use the internal linter before committing.
|
||||
EOF
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "passes a thin adapter with an @import line and provider-specific additions" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
@AGENTS.md
|
||||
@core/instructions/governance.md
|
||||
EOF
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_success
|
||||
}
|
||||
|
||||
@test "fails when the adapter duplicates most of AGENTS.md's content" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
@AGENTS.md
|
||||
|
||||
## Setup commands
|
||||
- Install deps: `pnpm install`
|
||||
- Run tests: `pnpm test`
|
||||
|
||||
## Code style
|
||||
- TypeScript strict mode, single quotes, no semicolons.
|
||||
EOF
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure
|
||||
assert_output --partial "duplicat"
|
||||
}
|
||||
|
||||
@test "fails when the adapter exceeds the max line threshold" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
{
|
||||
echo "@AGENTS.md"
|
||||
for i in $(seq 1 80); do echo "Provider-specific line $i unrelated to AGENTS.md content."; done
|
||||
} > "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure
|
||||
assert_output --partial "thin"
|
||||
}
|
||||
|
||||
@test "allows a custom --max-lines threshold" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
{
|
||||
echo "@AGENTS.md"
|
||||
for i in $(seq 1 10); do echo "Provider-specific line $i unrelated to AGENTS.md content."; done
|
||||
} > "$ADAPTER"
|
||||
run bash "$SCRIPT" --max-lines 5 "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure
|
||||
assert_output --partial "thin"
|
||||
}
|
||||
|
||||
@test "with --no-import-syntax, a text pointer to AGENTS.md is accepted instead of an @import line" {
|
||||
ADAPTER="$TMPDIR/copilot-instructions.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
See AGENTS.md at the repo root for setup, style, and testing conventions.
|
||||
|
||||
## Copilot-specific
|
||||
Prefer inline suggestions over chat for one-line edits.
|
||||
EOF
|
||||
run bash "$SCRIPT" --no-import-syntax "$ADAPTER" "$AGENTS_MD"
|
||||
assert_success
|
||||
}
|
||||
|
||||
@test "with --no-import-syntax, still fails if there is no mention of AGENTS.md at all" {
|
||||
ADAPTER="$TMPDIR/copilot-instructions.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
## Copilot-specific
|
||||
Prefer inline suggestions over chat for one-line edits.
|
||||
EOF
|
||||
run bash "$SCRIPT" --no-import-syntax "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "the two --no-import-syntax branches disagree: a text-pointer-only adapter fails in default mode" {
|
||||
ADAPTER="$TMPDIR/copilot-instructions.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
See AGENTS.md at the repo root for setup, style, and testing conventions.
|
||||
|
||||
## Copilot-specific
|
||||
Prefer inline suggestions over chat for one-line edits.
|
||||
EOF
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure
|
||||
assert_output --partial "no reference"
|
||||
|
||||
run bash "$SCRIPT" --no-import-syntax "$ADAPTER" "$AGENTS_MD"
|
||||
assert_success
|
||||
}
|
||||
|
||||
@test "with --no-import-syntax, an inert @AGENTS.md line alone is not a prose pointer" {
|
||||
ADAPTER="$TMPDIR/copilot-instructions.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
@AGENTS.md
|
||||
|
||||
## Copilot-specific
|
||||
Prefer inline suggestions over chat for one-line edits.
|
||||
EOF
|
||||
run bash "$SCRIPT" --no-import-syntax "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "--max-lines as the final argument reports a real error instead of failing silently" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
echo "@AGENTS.md" > "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD" --max-lines
|
||||
assert_failure 2
|
||||
assert_output --partial "--max-lines requires a value"
|
||||
}
|
||||
|
||||
@test "--max-lines rejects a non-numeric value with a real error" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
echo "@AGENTS.md" > "$ADAPTER"
|
||||
run bash "$SCRIPT" --max-lines abc "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 2
|
||||
assert_output --partial "non-negative integer"
|
||||
}
|
||||
|
||||
@test "--help exits 0 and documents usage" {
|
||||
run bash "$SCRIPT" --help
|
||||
assert_success
|
||||
assert_output --partial "Usage:"
|
||||
}
|
||||
|
||||
@test "fails with a clear error when the adapter file argument is missing" {
|
||||
run bash "$SCRIPT"
|
||||
assert_failure 2
|
||||
assert_output --partial "required"
|
||||
}
|
||||
|
||||
@test "a UTF-8 BOM before the @import line does not hide it" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
python3 -c "import sys; open(sys.argv[1], 'w', encoding='utf-8-sig').write('@AGENTS.md\n')" "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_success
|
||||
}
|
||||
|
||||
@test "a usage error exits 2, a genuine finding exits 1" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
echo "@AGENTS.md" > "$ADAPTER"
|
||||
|
||||
run bash "$SCRIPT" --max-lines -3 "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 2
|
||||
refute_output --partial "FAIL"
|
||||
|
||||
: > "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "FAIL"
|
||||
}
|
||||
|
||||
@test "a non-UTF-8 adapter is reported as an encoding error, not as a missing reference" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
python3 -c "import sys; open(sys.argv[1], 'wb').write('@AGENTS.md\n'.encode('utf-16'))" "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 2
|
||||
assert_output --partial "not valid UTF-8"
|
||||
refute_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "an @AGENTS.md folded into a sentence fails, and the message says the import needs its own line" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
See @AGENTS.md for shared conventions.
|
||||
EOF
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
assert_output --partial "line of its own"
|
||||
}
|
||||
|
||||
# --- Q1: a reference only counts where something would resolve it -------------
|
||||
|
||||
@test "an @AGENTS.md inside a backtick code fence is not credited as an import" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
# Claude notes
|
||||
|
||||
Put this at the top of the file:
|
||||
|
||||
```
|
||||
@AGENTS.md
|
||||
```
|
||||
EOF
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "an @AGENTS.md inside a tilde code fence is not credited as an import" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
~~~
|
||||
@AGENTS.md
|
||||
~~~
|
||||
EOF
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "an @AGENTS.md in a four-space indented code block is not credited as an import" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
# Claude notes
|
||||
|
||||
@AGENTS.md
|
||||
EOF
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "an @AGENTS.md inside a multi-line HTML comment is not credited as an import" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
# Claude notes
|
||||
|
||||
<!--
|
||||
@AGENTS.md
|
||||
-->
|
||||
EOF
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "an @AGENTS.md indented up to three spaces is still credited" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
printf ' @AGENTS.md\n' > "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_success
|
||||
}
|
||||
|
||||
# --- Q2: encodings that decode as valid UTF-8 but are not UTF-8 ---------------
|
||||
|
||||
@test "a BOM-less UTF-16LE adapter is an encoding error, not a missing reference" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
python3 -c "import sys; open(sys.argv[1], 'wb').write('@AGENTS.md\n'.encode('utf-16-le'))" "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 2
|
||||
assert_output --partial "not valid UTF-8"
|
||||
refute_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "a BOM-less UTF-32LE adapter is an encoding error, not a missing reference" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
python3 -c "import sys; open(sys.argv[1], 'wb').write('@AGENTS.md\n'.encode('utf-32-le'))" "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 2
|
||||
assert_output --partial "not valid UTF-8"
|
||||
refute_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "a doubled UTF-8 BOM does not hide the @import line" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
python3 -c "import sys; open(sys.argv[1], 'wb').write(('@AGENTS.md\n').encode('utf-8'))" "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_success
|
||||
}
|
||||
|
||||
@test "a BOM in front of a mid-file @import line does not hide it" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
python3 -c "import sys; open(sys.argv[1], 'wb').write(('# Claude notes\n\n@AGENTS.md\n').encode('utf-8'))" "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_success
|
||||
}
|
||||
|
||||
# --- Q3: a file that exists but cannot be read is not a FAIL -----------------
|
||||
|
||||
@test "an adapter that exists but cannot be read exits 3 with a diagnostic and no FAIL" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
echo "@AGENTS.md" > "$ADAPTER"
|
||||
chmod 000 "$ADAPTER"
|
||||
# chmod is not enough under a uid that bypasses it (root in CI containers).
|
||||
# /proc/self/mem is a regular file whose read returns EIO for every uid, so
|
||||
# it exercises the same branch where chmod cannot.
|
||||
if cat "$ADAPTER" >/dev/null 2>&1; then
|
||||
if [ -e /proc/self/mem ]; then
|
||||
ADAPTER=/proc/self/mem
|
||||
else
|
||||
chmod 644 "$TMPDIR/CLAUDE.md"
|
||||
skip "no way to make a readable-by-stat, unreadable-by-open file here"
|
||||
fi
|
||||
fi
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
chmod 644 "$TMPDIR/CLAUDE.md"
|
||||
assert_failure 3
|
||||
assert_output --partial "could not be read"
|
||||
refute_output --partial "FAIL"
|
||||
}
|
||||
|
||||
# --- Q4: the reference has to name, and resolve to, a real AGENTS.md ---------
|
||||
|
||||
@test "an @import naming a path that does not exist is not credited" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
echo "@docs/does/not/exist/AGENTS.md" > "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
assert_output --partial "does not exist"
|
||||
}
|
||||
|
||||
@test "@NOTAGENTS.md and @zzzAGENTS.md are not imports of AGENTS.md" {
|
||||
# The decoys are real files, so the on-disk resolution check cannot be what
|
||||
# rejects them. Only the path-segment boundary can — without the fixtures
|
||||
# this test passes against a substring match and proves nothing.
|
||||
cp "$AGENTS_MD" "$TMPDIR/NOTAGENTS.md"
|
||||
cp "$AGENTS_MD" "$TMPDIR/zzzAGENTS.md"
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
echo "@NOTAGENTS.md" > "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
|
||||
echo "@zzzAGENTS.md" > "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "an @import resolving to a zero-byte AGENTS.md is not credited" {
|
||||
SUB="$TMPDIR/empty"
|
||||
mkdir -p "$SUB"
|
||||
: > "$SUB/AGENTS.md"
|
||||
ADAPTER="$SUB/CLAUDE.md"
|
||||
echo "@AGENTS.md" > "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
assert_output --partial "empty"
|
||||
}
|
||||
|
||||
@test "an @import naming a real relative path to AGENTS.md is credited" {
|
||||
mkdir -p "$TMPDIR/docs"
|
||||
cp "$AGENTS_MD" "$TMPDIR/docs/AGENTS.md"
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
echo "@docs/AGENTS.md" > "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD"
|
||||
assert_success
|
||||
}
|
||||
|
||||
# --- Q5: --no-import-syntax needs a pointer, not a mention -------------------
|
||||
|
||||
@test "with --no-import-syntax, a negated mention of AGENTS.md is not a pointer" {
|
||||
ADAPTER="$TMPDIR/copilot-instructions.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
Do NOT read AGENTS.md; it is obsolete.
|
||||
EOF
|
||||
run bash "$SCRIPT" --no-import-syntax "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "with --no-import-syntax, a past-tense mention of a deleted AGENTS.md is not a pointer" {
|
||||
ADAPTER="$TMPDIR/copilot-instructions.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
We deleted AGENTS.md last year.
|
||||
EOF
|
||||
run bash "$SCRIPT" --no-import-syntax "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "with --no-import-syntax, a pointer inside a code fence is not credited" {
|
||||
ADAPTER="$TMPDIR/copilot-instructions.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
Example of what to write:
|
||||
|
||||
```
|
||||
See AGENTS.md at the repo root for shared conventions.
|
||||
```
|
||||
EOF
|
||||
run bash "$SCRIPT" --no-import-syntax "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "with --no-import-syntax, a pointer inside an HTML comment is not credited" {
|
||||
ADAPTER="$TMPDIR/copilot-instructions.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
# Copilot instructions
|
||||
|
||||
<!-- See AGENTS.md at the repo root for shared conventions. -->
|
||||
EOF
|
||||
run bash "$SCRIPT" --no-import-syntax "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "with --no-import-syntax, a name merely ending in AGENTS.md is not a pointer to it" {
|
||||
# Real decoy files, so the on-disk resolution check cannot be what rejects
|
||||
# these — only the token boundary in the mention pattern can. zzzAGENTS.md
|
||||
# is the load-bearing case: dropping the boundary from NOTAGENTS.md leaves
|
||||
# the fragment "NOT" behind, which the negation cue then rejects for an
|
||||
# unrelated reason, so that case alone would prove nothing.
|
||||
cp "$AGENTS_MD" "$TMPDIR/zzzAGENTS.md"
|
||||
cp "$AGENTS_MD" "$TMPDIR/NOTAGENTS.md"
|
||||
ADAPTER="$TMPDIR/copilot-instructions.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
See zzzAGENTS.md at the repo root for shared conventions.
|
||||
EOF
|
||||
run bash "$SCRIPT" --no-import-syntax "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
See NOTAGENTS.md at the repo root for shared conventions.
|
||||
EOF
|
||||
run bash "$SCRIPT" --no-import-syntax "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
|
||||
@test "with --no-import-syntax, a pointer naming a path that does not exist is not credited" {
|
||||
ADAPTER="$TMPDIR/copilot-instructions.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
See docs/does/not/exist/AGENTS.md for shared conventions.
|
||||
EOF
|
||||
run bash "$SCRIPT" --no-import-syntax "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
assert_output --partial "does not exist"
|
||||
}
|
||||
|
||||
# --- argument handling -------------------------------------------------------
|
||||
|
||||
@test "a third positional argument is rejected instead of silently ignored" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
echo "@AGENTS.md" > "$ADAPTER"
|
||||
run bash "$SCRIPT" "$ADAPTER" "$AGENTS_MD" "$TMPDIR/also-not-graded.md"
|
||||
assert_failure 2
|
||||
assert_output --partial "exactly 2 positional arguments"
|
||||
# The usage text this prints mentions the word FAIL, so refute the shape of
|
||||
# a real finding line rather than the bare word.
|
||||
refute_output --partial "FAIL Adapter"
|
||||
}
|
||||
|
||||
@test "an unknown option is reported as an unknown option, not as a missing file" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
echo "@AGENTS.md" > "$ADAPTER"
|
||||
run bash "$SCRIPT" --bogus "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 2
|
||||
assert_output --partial "unknown option '--bogus'"
|
||||
refute_output --partial "'--bogus' is not a file"
|
||||
refute_output --partial "FAIL Adapter"
|
||||
}
|
||||
|
||||
@test "--max-lines=N is accepted in the equals form" {
|
||||
ADAPTER="$TMPDIR/CLAUDE.md"
|
||||
{
|
||||
echo "@AGENTS.md"
|
||||
for i in $(seq 1 10); do echo "Provider-specific line $i unrelated to AGENTS.md content."; done
|
||||
} > "$ADAPTER"
|
||||
run bash "$SCRIPT" --max-lines=5 "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "thin"
|
||||
|
||||
run bash "$SCRIPT" --max-lines=40 "$ADAPTER" "$AGENTS_MD"
|
||||
assert_success
|
||||
}
|
||||
|
||||
@test "with --no-import-syntax, a bare mention with no deference cue is not a pointer" {
|
||||
ADAPTER="$TMPDIR/copilot-instructions.md"
|
||||
cat > "$ADAPTER" <<'EOF'
|
||||
# Copilot instructions
|
||||
|
||||
This repo also has an AGENTS.md.
|
||||
|
||||
Prefer inline suggestions over chat for one-line edits.
|
||||
EOF
|
||||
run bash "$SCRIPT" --no-import-syntax "$ADAPTER" "$AGENTS_MD"
|
||||
assert_failure 1
|
||||
assert_output --partial "no reference"
|
||||
}
|
||||
@@ -1,18 +0,0 @@
|
||||
{
|
||||
"author": {
|
||||
"name": "Defame1297",
|
||||
"url": "https://git.dev.rkdr.net/Defame1297/"
|
||||
},
|
||||
"description": "Cross-cutting utility skills for everyday AI-assisted coding \u2014 triage, diagnosis, architecture review, and session navigation.",
|
||||
"displayName": "Core",
|
||||
"keywords": [
|
||||
"cross-cutting",
|
||||
"triage",
|
||||
"diagnose",
|
||||
"architecture",
|
||||
"debug"
|
||||
],
|
||||
"license": "MIT",
|
||||
"name": "core",
|
||||
"version": "1.1.0"
|
||||
}
|
||||
@@ -1,3 +0,0 @@
|
||||
{
|
||||
"mcpServers": {}
|
||||
}
|
||||
@@ -1,38 +1,35 @@
|
||||
# core
|
||||
|
||||
Cross-cutting utility skills for everyday AI-assisted coding — triage, diagnosis, architecture review, and session navigation.
|
||||
Skills for authoring and auditing a repo's AGENTS.md and the provider adapter files that defer to it.
|
||||
|
||||
## Install
|
||||
|
||||
**Claude Code:**
|
||||
apm is the only supported install path (ADR-0024). Declare this package in the consuming project's `apm.yml`:
|
||||
|
||||
```bash
|
||||
claude plugin marketplace add <owner>/<repo>
|
||||
claude plugin install core@<marketplace-name>
|
||||
```yaml
|
||||
dependencies:
|
||||
apm:
|
||||
- git: git@git.dev.rkdr.net:Defame1297/holocron.git
|
||||
path: plugins/core
|
||||
```
|
||||
|
||||
**GitHub Copilot CLI:**
|
||||
Then:
|
||||
|
||||
```bash
|
||||
copilot plugin marketplace add <owner>/<repo>
|
||||
copilot plugin install core
|
||||
apm install
|
||||
```
|
||||
|
||||
**Local (development):**
|
||||
The entry above is unpinned and tracks the remote's default branch — add `ref: <tag>` to pin a release. Registering the catalogue instead (`apm marketplace add git@git.dev.rkdr.net:Defame1297/holocron.git --name holocron`) gets you the `core@holocron` short name, but writes to `~/.apm/marketplaces.json` at user scope; the git+path object needs nothing beyond the manifest.
|
||||
|
||||
```bash
|
||||
# Claude Code
|
||||
claude --plugin-dir ./plugins/core
|
||||
|
||||
# GitHub Copilot CLI
|
||||
copilot plugin install ./plugins/core
|
||||
```
|
||||
**Native plugin installs do not work.** This package ships no per-plugin manifest and no flat content directories, so a host that installs it natively gets zero skills — and Claude Code raises no error while doing it (ADR-0024).
|
||||
|
||||
## Contents
|
||||
|
||||
| Component | Path | Description |
|
||||
|---|---|---|
|
||||
| Skills | `skills/` | Slash commands available after install |
|
||||
| Skills | `.apm/skills/` | Slash commands available after install |
|
||||
|
||||
`.apm/` is the authoring source and the only thing apm deploys (ADR-0024).
|
||||
|
||||
## Skills
|
||||
|
||||
|
||||
35
plugins/core/apm.yml
Normal file
35
plugins/core/apm.yml
Normal file
@@ -0,0 +1,35 @@
|
||||
name: core
|
||||
version: 1.1.2
|
||||
description: Skills for authoring and auditing a repo's AGENTS.md and the provider adapter files that defer to it.
|
||||
author:
|
||||
name: Defame1297
|
||||
email: defame1297@rkdr.net
|
||||
url: https://git.dev.rkdr.net/Defame1297/
|
||||
license: MIT
|
||||
homepage: https://git.dev.rkdr.net/Defame1297/holocron/src/branch/main/plugins/core
|
||||
repository: https://git.dev.rkdr.net/Defame1297/holocron/src/branch/main/plugins/core
|
||||
keywords:
|
||||
- agents-md
|
||||
- documentation
|
||||
- audit
|
||||
- provider-adapter
|
||||
- governance
|
||||
|
||||
# Constrains what .apm/ may contain: instructions, skill, hybrid, or prompts
|
||||
type: skill
|
||||
|
||||
targets:
|
||||
- claude
|
||||
- copilot
|
||||
- codex
|
||||
|
||||
# "auto" publishes the authoritative local source layout, or list explicit
|
||||
# repo paths to define the complete publication set.
|
||||
includes: auto
|
||||
|
||||
dependencies:
|
||||
apm: []
|
||||
mcp: []
|
||||
devDependencies:
|
||||
apm: []
|
||||
scripts: {}
|
||||
@@ -1,3 +0,0 @@
|
||||
{
|
||||
"hooks": {}
|
||||
}
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user