Compare commits

..

107 Commits

Author SHA1 Message Date
Jesse Vincent
537d649ab3 feat(brainstorming): tooling question in the design presentation for new projects
For a project with no configured tooling, the design presentation now
includes a tooling question — linting/auto-formatting, unit-test
infrastructure, e2e test infrastructure, fuzz/mutation testing — whose
answers land in the spec's Global Constraints. Without this, sessions
never ask and ship plans with zero tooling constraints (0/3 in the
quorum control cell); with it, the ask fires before any code in 3/3
and the user's answer is honored verbatim, including exclusions.

Claude-Session: https://claude.ai/code/session_0185AJr98gHx5EmwqNeft4Sy
2026-08-06 17:28:47 -07:00
Drew Ritter
5f8f500b1d fix(release): wire Hermes into version bumps
Register the Hermes YAML manifest alongside the existing JSON manifests. Route manifest reads and writes by extension through jq or Mike Farah yq v4, with field names and values passed as data.

Preflight every present manifest before the mutating bump loop so a deterministic YAML read failure cannot leave earlier JSON manifests partially updated. Cover check, audit, bump, registry wiring, and byte-for-byte no-partial-write behavior with one focused fixture test.
2026-08-06 16:21:12 -07:00
Drew Ritter
707b155a38 docs: plan Hermes version-bump wiring
Record Drew's approved reduced design after the second staff review. Limit preflight to the mutating bump path, cover audit's independent read path, and require byte-for-byte proof that deterministic YAML failures cannot partially update earlier JSON manifests.

Provide one TDD implementation task for the Hermes registry entry, jq/yq dispatch, focused preflight, and three behavioral checks. Explicitly defer rollback, audit-status changes, nested YAML, runtime changes, and broader release-tool refactoring.
2026-08-06 16:21:12 -07:00
Drew Ritter
3e1ecde38f docs: reduce Hermes version-bump design
Incorporate the adversarial design review without turning the Hermes wiring follow-up into a general release-script refactor. Keep the existing jq path, add Mike Farah yq v4 only for .yaml, and retain one read-only preflight to prevent deterministic partial bumps.\n\nReduce the test contract to three behavioral cases and explicitly defer .yml support, nested YAML, rollback machinery, audit/status redesign, exhaustive failure matrices, and the separately discovered JSON-expression issue. This follows Drew's direction to avoid ceremony and overengineering.
2026-08-06 16:21:12 -07:00
Drew Ritter
ffe22811bf docs: design Hermes version-bump wiring
Document the agreed follow-up to PR #2025 on a branch based on its merged dev commit. The design registers the Hermes YAML manifest, keeps jq for existing JSON files, and uses Mike Farah yq v4 for a narrow top-level YAML field rather than adding a Bash parser.\n\nDefine focused failure behavior and behavioral tests while explicitly excluding nested YAML, Hermes runtime changes, and unrelated release-script refactors. This captures Drew's request to keep the implementation small and avoid process or abstraction overhead.
2026-08-06 16:21:12 -07:00
Drew Ritter
cfb310c69a Merge pull request #2089 from obra/fix/x13-illegibility
fix(sdd): reviewers re-read illegible evidence instead of re-running to regenerate it
2026-08-06 16:10:58 -07:00
Drew Ritter
fdd1763d77 Merge pull request #2086 from obra/fix/spec-travels-with-plan
fix(planning): the spec travels with the plan
2026-08-06 15:13:54 -07:00
Drew Ritter
af4bebf762 Merge pull request #2024 from obra/fix/worktree-cleanup-untracked-checkin
fix(finishing): check in with human partner when worktree removal hits untracked files
2026-08-06 12:21:45 -07:00
Drew Ritter
17b42c8128 fix(finishing): name the actual files in the refusal prompt
`git status --porcelain` collapses a wholly-untracked directory to a single
`?? docs/` line. In the shape of the incident this step exists for (#2016 — an
uncommitted plan document under an untracked `docs/` tree), the file list we
show the human partner therefore names no file at all:

    $ git -C "$WORKTREE_PATH" status --porcelain
    ?? docs/
    $ git -C "$WORKTREE_PATH" status --porcelain -uall
    ?? docs/superpowers/plans/2026-08-04-csv-export-rollout.md

Both forms produce identical (empty) output on a clean worktree, so this adds
no over-trigger surface.

Found while running this PR's behavioral micro-tests. Every treatment agent
dug past `?? docs/` unprompted and named the document, so the step did work —
but on the agent's own initiative rather than because the text asked for it.
That initiative is not reliable one tier down: Claude Haiku 4.5 on the control
arm failed for exactly this shape, asking a question that never named the file
and then deciding for the human when they deferred. Nothing in the prior
wording stopped a treatment agent from relaying `?? docs/` verbatim and
satisfying the letter of the instruction.

Re-ran the treatment cells against this amended text — Opus pass (refusal
fired, named the file), Haiku 4.5 pass (named the file) — no regression.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 12:14:55 -07:00
Drew Ritter
1245282b05 Merge pull request #1805 from obra/fix/render-graphs-no-shell
fix(writing-skills): make render-graphs ESM-compatible and shell-free
2026-08-06 11:50:32 -07:00
Drew Ritter
6819b42d97 Merge pull request #2064 from obra/fix/docs-codex-efficiency-campaign
docs: codex-efficiency fix-cycle spec and plan (campaign record)
2026-08-05 23:17:21 -07:00
Drew Ritter
02654f93bf test(writing-skills): cover render-graphs execution 2026-08-05 22:38:40 -07:00
Jesse Vincent
dcd3661b7c fix(writing-skills): run graphviz without a shell in render-graphs.js
The `dot` availability check shelled out to `which dot`, which is not a
command on Windows, so render-graphs.js reported graphviz as missing on
Windows even when it was installed. Replace it with a direct `dot -V`
probe via execFileSync.

Also switch the SVG render call from execSync to execFileSync('dot',
['-Tsvg']). Behavior is identical on macOS/Linux — the diagram source
was already passed via stdin, never interpolated into the command — but
running the binary directly removes the shell entirely.
2026-08-05 22:38:40 -07:00
Drew Ritter
9be44ebf40 Merge pull request #2025 from obra/hermes-harness-rebase
feat(hermes): Hermes Agent harness support — eval-verified pre_llm_call bootstrap
2026-08-05 18:23:46 -07:00
Drew Ritter
695744056e chore(hermes): align plugin version with dev
Update the Hermes plugin manifest from 6.1.1 to 6.2.0 so PR #2025 matches the current release version at the tip of origin/dev.\n\nThis intentionally does not change the version bump tooling. The existing release script supports JSON manifests only; YAML support will be handled separately on its own branch.
2026-08-05 17:57:31 -07:00
Kattni
fb518edf7b Moves Community up, and adds ToC. 2026-08-04 20:16:46 -07:00
Jesse Vincent
80b82abd8d fix(sdd): task reviewers re-read illegible evidence instead of re-running to regenerate it
Interrogation of reviewers who bypassed test-evidence leases showed a
convergent driver: when the report or receipt looked truncated or
couldn't be located, re-running the suite felt cheaper than re-reading —
evidence got regenerated instead of read. This paragraph names that
moment: re-read at the stated path, report a genuine gap to the
controller, and never re-run to regenerate what wasn't read.

Battery: 0/31 reviewer re-runs across 4 treatment reps vs 7/~59
reviewers in 5/8 control reps on the same scenario and classifier.

Claude-Session: https://claude.ai/code/session_0185AJr98gHx5EmwqNeft4Sy
2026-08-04 18:24:10 -07:00
Drew Ritter
05c2393b82 Merge pull request #2078 from obra/fix/x6a-sdd-batch-small-tasks
fix(sdd): batch small same-shape tasks into one dispatch
2026-08-04 14:26:31 -07:00
Drew Ritter
78cc189244 fix(sdd): batch reviews check the diff against the brief's file list
Batching moves N edits under one review, which changes the review's
failure profile: an implementer that silently skips one file of twelve
produces a diff full of correct, uniform edits — nothing conspicuous is
missing, and no seat in the pipeline was assigned to notice. The single
combined review is the only net for a dropped edit, but the reviewer
template never told it to count.

The batch brief already lists every file with its change, so the reviewer
reconciles the diff against that list file by file; a listed file with no
hunk is a Missing finding regardless of how clean the rest of the batch
looks. Conditional on a multi-file brief, so single-task reviews are
unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-04 14:25:40 -07:00
Drew Ritter
be76350536 Merge pull request #2080 from obra/fix/x7a-sdd-evidence-bearing-preflight
fix(sdd): preflight emits its checks as a ledger table and rules on what it surfaces
2026-08-04 14:22:57 -07:00
Drew Ritter
419dec7755 Merge dev into fix/x7a: resolve preflight paragraph with the composed 2077+2080 text
Both PRs rewrote the same preflight paragraph. Resolution is the composed
text published in #2080's description — the configuration the 3/3+3/3
composed eval grades ran.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-04 14:21:58 -07:00
Drew Ritter
2b195749df Merge pull request #2077 from obra/fix/x9a-sdd-never-stall
fix(sdd): rule and continue — non-catastrophic conflicts get ledgered rulings, not blocking questions
2026-08-04 14:16:16 -07:00
Drew Ritter
7a01a0e83a fix(sdd): one Ruling: token everywhere, exhaustive finish roll-up
The breaker's two ledger formats wrote lowercase 'ruling' (parked
findings, load-bearing adjudications), so the Finish section's
collect-every-`Ruling:`-line step missed exactly the rulings made under
the most pressure. Field evidence from an independent eval rep: a
breaker-cap run adjudicated correctly, wrote everything to the
plan-scoped ledger, deleted the workspace at finish, and left no durable
trace of the adjudication.

Capitalize the two breaker formats to the canonical token, and make the
finish roll-up explicitly exhaustive across preflight, parked, and
breaker rulings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-04 14:13:42 -07:00
Drew Ritter
8acf8e5f24 Merge pull request #2059 from obra/fix/t1-sdd-no-worker-reviewers
fix(sdd): dispatched subagents never dispatch subagents
2026-08-04 14:05:55 -07:00
Drew Ritter
50a924b0c4 Merge pull request #2062 from obra/fix/t5-codex-spawn-routing
fix(codex): explicit model+effort on every spawn, with config backstop
2026-08-04 14:04:31 -07:00
Drew Ritter
2a977c7095 Merge pull request #2061 from obra/fix/t2-codex-event-waits
fix(sdd,codex): event-driven bounded waits — 65-78% wait timeouts to 0%
2026-08-04 14:03:30 -07:00
Drew Ritter
50f787ca5c Merge pull request #2060 from obra/fix/t3-codex-tools-corrections
fix(codex): correct multi-agent guidance against the Codex source (V2)
2026-08-04 14:00:31 -07:00
Jesse Vincent
538d65120b fix(planning): the spec travels with the plan — Spec: header pointer + SDD reads it at setup
In controlled evals, an identical seeded-incoherence plan yielded 0-1/5
correct conflict resolutions when executed specless (controllers ruled
the conflicts 'internally explained') and 4-5/5 with the spec merely
present and named — even with no other skill-text changes. Cross-task
coherence turns out to be adjudicable only against ground truth above
the plan; this change makes that ground truth travel with the plan.

Claude-Session: https://claude.ai/code/session_0185AJr98gHx5EmwqNeft4Sy
2026-08-04 11:17:05 -07:00
Jesse Vincent
61f669ebc9 fix(sdd): preflight emits its pairwise checks as a ledger table and rules on what it surfaces
The pre-Task-1 conflict scan currently permits 'the scan is clean' with
no evidence the scan happened — mined sessions show controllers skipping
straight to dispatch and plan conflicts surfacing mid-execution as
blocking questions. Requiring the scan to emit one row per task pair
sharing a file/interface and one row per task's self-consistency turns
the claim into an artifact; in controlled evals the table appeared 3/3
with conflicts surfaced pre-dispatch, and the mechanism held 3/3 when
composed with the never-stall ruling change (#2077).

Claude-Session: https://claude.ai/code/session_0185AJr98gHx5EmwqNeft4Sy
2026-08-03 09:03:05 -07:00
Jesse Vincent
e7a4285985 fix(sdd): batch small same-shape tasks into one dispatch
Plans sometimes enumerate many tiny, same-shape edits (one-line fixes,
constant changes, a field added across files) as separate tasks. The
current loop dispatches a fresh implementer plus review per task, so a
12-micro-task plan costs ~24 subagent seats for what one subagent could
do in a single pass. In controlled evals on a micro-task plan, batching
cut cost 73% and dispatches 87% with better completion than control; on
a 5-non-trivial-task plan the rule correctly never batched (dispatch
counts and completion identical to control).

Claude-Session: https://claude.ai/code/session_0185AJr98gHx5EmwqNeft4Sy
2026-08-02 19:36:55 -07:00
Jesse Vincent
39f9602432 fix(sdd): rule and continue — non-catastrophic conflicts get ledgered rulings, not blocking questions
A donated session sat dormant 8h48m waiting for a plan-conflict answer
that cost ~zero tokens to decide. Wrong-ruling rework is bounded;
stalls are not. This encodes the never-stall doctrine: plan conflicts,
ambiguities, and cap exceptions get a controller ruling recorded in
the ledger and work proceeds; only irreversible/destructive actions,
security-sensitive actions, out-of-worktree side effects (merge/push/
publish), and totally-broken plans remain hard stops. Rulings surface
in the Finish report instead of as mid-run questions.

Evals: 3/3 no-stall vs control 3/3 stall-at-preflight on a
seeded-conflict SDD plan; catastrophic guard 5/5 (every rep reaching a
seeded DROP TABLE step refused it); re-validated 3/3 after rebase onto
the current fix-PR text; composes cleanly with the evidence-bearing
preflight treatment.

Claude-Session: https://claude.ai/code/session_0185AJr98gHx5EmwqNeft4Sy
2026-08-02 11:14:01 -07:00
Jesse Vincent
3ff8d15f15 docs: codex-efficiency fix-cycle spec and plan (campaign record) 2026-07-31 10:52:05 -07:00
Jesse Vincent
e9686d5c09 fix(codex): explicit model+effort on every spawn, config backstop
Depth-2 child-issued spawns omitted model 2/2 at CLI 0.146; model
without reasoning_effort resets effort to the model default.
2026-07-31 10:27:51 -07:00
Jesse Vincent
d8189d1587 fix(sdd,codex): bounded wait stretches with reconciliation
Round 2 proved the long-wait mechanism (65.1%->0.0% timeouts) but
20-38 min silent waits starved graders and let 1/51 children vanish;
bounded 5-10 min stretches with a status line and list_agents
reconcile keep the efficiency and restore observability.
2026-07-31 10:27:51 -07:00
Jesse Vincent
db4538fcb8 fix(sdd): controllers wait long or not at all
Docs-only wait guidance in the platform reference changed nothing
(65.1% vs 67.1% baseline wait-timeout rate); the discipline now lives
in the controller loop the session actually re-reads.
2026-07-31 10:27:50 -07:00
Jesse Vincent
9b8b14fe12 fix(codex): event-driven waiting instead of short polls
60-78% of wait_agent calls timed out across every measured corpus;
waits are event subscriptions, so one long wait replaces dozens of
polls at identical wake latency.
2026-07-31 10:27:50 -07:00
Jesse Vincent
7c560e048b fix(sdd): reviewers never dispatch subagents either
The first fix-cycle battery moved the depth-2 leak from implementers
(9/9 baseline -> 0/6) to a final reviewer that spawned two
sub-reviewers; the contract now reaches every dispatched role.
2026-07-31 09:39:55 -07:00
Jesse Vincent
75756d2900 fix(codex): correct multi-agent guidance against Codex source
Five claims contradicted by the Codex CLI source (V2 has no
close_agent; followup_task always reaches a child; role files attach
via agent_type; full-history forks accept model/effort; V2 spawn
allowlist). Citations: superpowers-autoresearch
docs/2026-07-29-codex-multiagent-v2-capabilities.md.
2026-07-31 09:39:55 -07:00
Jesse Vincent
2e7d681591 fix(sdd): implementers never dispatch subagents
Depth-2 worker-spawned reviewers were 9/9 same-task duplicate reviews
across four corpora in the codex-efficiency eval campaign.
2026-07-31 09:39:55 -07:00
Jesse Vincent
bb2a34b2a0 docs: remove the "We're Hiring" section from the README
The community engineer role has a candidate on trial, so the posting no
longer needs to be at the top of the README.
2026-07-27 11:43:14 -07:00
Jesse Vincent
7b4dc4d7fd Merge main back into dev after the v6.2.0 rebase-merge
Converges dev with the rebased release SHAs on main immediately, while
the trees are identical, so the merge is conflict-free and the next
dev -> main release shows only genuinely-new work.
2026-07-23 17:28:00 -07:00
Jesse Vincent
5b73c0f63a Merge main back into dev: converge histories after the 6.1.1 rebase-merge
All 11 conflicts are residue of the v6.1.1 rebase-merge, which replayed
dev's commits onto main as new SHAs: version manifests (dev 6.2.0 vs
main 6.1.1), RELEASE-NOTES.md, the porting-guide table (main predates
the #1969 dead-reference fix), and add/add on the Codex package script
and test (dev carries the later portability fixes). Every conflict
resolves to dev's side; the merged tree is identical to dev's.
2026-07-23 16:18:07 -07:00
Jesse Vincent
d262bc400c Release v6.2.0: SDD plan-scoped workspace and resume-based fix loop, skills compression sweep, Windows SessionStart fix (#2026)
Release notes for everything on dev since v6.1.1, plus the version bump
to 6.2.0 across all seven declared manifest files (bump-version.sh,
audit clean). Tagging and marketplace publication happen after the
dev -> main merge.
2026-07-23 16:16:32 -07:00
Jesse Vincent
b6613057ae test(hermes): realign suite with the pre_llm_call mechanism; slim docs to the README section
The 20-test suite still exercised the dead on_session_start/inject_message
mechanism (17 failures against the rewritten plugin). Rewritten for the
real contract: pre_llm_call registration + first-turn-only context return,
register_skill receiving pathlib.Path (the conftest mock now raises on str,
mirroring hermes' AttributeError that silently disables a plugin), both
install layouts resolving skills, loud failure when skills are missing,
tool mapping sourced verbatim from hermes-tools.md, and a bootstrap-size
guard against hermes' 10k-char context spill threshold. 19 tests, passing.

Install docs collapse into the README section per maintainer direction:
docs/README.hermes.md and .hermes-plugin/INSTALL.md are gone; the README
carries the two-line install plus the compaction caveat. plugin.yaml
version aligned to 6.1.1.
2026-07-23 16:05:54 -07:00
Jesse Vincent
178528c03e fix(hermes): working bootstrap injection via pre_llm_call + native skill registration
Empirical findings from the quorum eval bring-up (superpowers-evals
docs/experiments/2026-07-23-hermes-target-bringup.md):

- ctx.inject_message exists but returns False when called from
  on_session_start — nothing reaches the model. The documented path,
  a pre_llm_call hook returning {"context": ...} on is_first_turn,
  verifiably delivers (probe model echoed an injected codeword).
- ctx.register_skill requires a pathlib.Path; passing a str raises
  AttributeError inside hermes, which silently disables the entire
  plugin (no log line anywhere). This also means any exception in
  register() is invisible — keep register() failure-proof.
- Registered skills are namespaced by plugin name: models invoke
  skill_view("superpowers:brainstorming") and receive the stock
  SKILL.md — verified live on GLM 5.2, both install layouts.

The plugin now: resolves skills/ for both the git-clone layout
(.hermes-plugin/ and skills/ as siblings) and a flattened install,
raising loudly when neither matches; registers every stock skill with
Hermes' native loader (no per-harness skill copies); injects the
using-superpowers bootstrap via pre_llm_call on the first turn; and
sources the tool mapping from references/hermes-tools.md instead of
duplicating it. Injected context is transient (API-call time only, never
persisted in the session export) — verification of injection must be
behavioral.
2026-07-23 15:17:55 -07:00
Jesse Vincent
7b177613c0 feat(hermes): Hermes Agent harness support, rebased to a Hermes-only diff
Rebase of PR #1922 onto current dev: the ~14 files of v6.1.0-era
codex/release drift are dropped, the porting-guide edits (stale against
the post-prune rewrite, no Hermes content) are dropped, and the Hermes
surface is kept intact: .hermes-plugin/ (on_session_start bootstrap
injection), tests/hermes/ (20 tests, passing), docs/README.hermes.md,
references/hermes-tools.md, the Platform Adaptation row, README section,
and Python ignores.

Known open items from review, unchanged by this rebase: the injection
mechanism uses ctx.inject_message from on_session_start, which the
official plugin guide does not document (pre_llm_call returning
{"context": ...} is the sanctioned path), skills are not registered via
ctx.register_skill, and the acceptance transcript predates the fix.

Co-authored-by: kumarabd <kumarabd@users.noreply.github.com>
2026-07-23 12:18:13 -07:00
Jesse Vincent
1f0e2ab912 fix(finishing): check in with human partner when worktree removal hits untracked files
git worktree remove refuses when the tree holds modified or untracked
files, and the skill gave no guidance for that refusal — the natural
agent response was --force, permanently destroying files that exist
nowhere else (uncommitted plans, notes, scratch work). Reported twice
from real sessions (#2016's plan loss, #1223's dirty-tree ambiguity).

Step 6 now treats the refusal as a stop-and-ask moment: show the
untracked files, offer commit / relocate / delete, and only remove the
worktree after the human partner chooses. Adds a matching rationalization
row so --force-as-cleanup is named as the failure it is.
2026-07-23 11:55:24 -07:00
Jesse Vincent
0146173544 fix(systematic-debugging): find-polluter accepts ./-prefixed patterns and matches top-level tests
Follow-up to #2011 (which fixed the ./-prefix mismatch for the documented
pattern form): strip a leading ./ from the caller's pattern instead of
double-prefixing it into a never-matching ././ form, and also match the
pattern with '**/' collapsed, since find -path cannot match '**/' against
zero directory levels and silently skipped files directly under the base
directory (src/top.test.ts vs src/**/*.test.ts).

Adds a deterministic test suite for the script with a stubbed npm.
2026-07-23 10:54:50 -07:00
dev_Hakaze
54d0efefd7 fix(systematic-debugging): match find -path ./ prefix in find-polluter.sh (#2011)
find . emits ./-prefixed paths, so -path "src/**/*.test.ts" matched
nothing; wc -l on empty stdin then lied as "Found 1". Fixes #2008.

Co-authored-by: arimu1 <19286898+arimu1@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-23 10:53:39 -07:00
Mark Rada
55d28ddf10 docs(using-superpowers): drop dangling subagent-support anchor (#2010)
The prune in e7ddc25e removed the `## Subagent support` section from
antigravity-tools.md but left the inline cross-reference to it in the
dispatch table, so `[Subagent support](#subagent-support)` resolves to
nothing. An agent following the pointer to learn the difference between
the `self` and `research` subagent types lands nowhere.

Drop the dangling parenthetical. The guidance it pointed at survives in
the same table cell -- `self` for full-capability work, `research` for
read-only -- so no content is lost and the row still answers the
question the removed section answered.

gemini-tools.md carries the same cross-reference but retains its
`## Subagent support` heading, so its link is valid and is left alone.
2026-07-23 10:47:44 -07:00
Jesse Vincent
cc690476fc feat(sdd): lifecycle restructure with resume-based fix loop, five-round breaker, and rationalization table 2026-07-19 12:36:33 -07:00
Jesse Vincent
7ce7620d44 feat(sdd): align templates and codex reference with resume-based fix rounds 2026-07-19 12:36:33 -07:00
Jesse Vincent
f428cba185 feat(sdd): add scoped re-review prompt template 2026-07-19 12:36:33 -07:00
Jesse Vincent
eb1ff1f11f docs(plans): SDD fix-loop redesign implementation plan
Eight tasks across two repos: new re-review template, template/reference
alignment, full SKILL.md lifecycle restructure with move map, two
seeded-ledger fixture helpers, three quorum scenarios, and the RED/GREEN/
regression live-run campaign.
2026-07-19 12:36:33 -07:00
Jesse Vincent
bea92dce1a docs(specs): SDD fix-loop redesign design spec
Review-fix loop gets resume-the-implementer semantics, scoped
re-reviews, a five-round circuit breaker, and controller adjudication
at trip. SKILL.md reorganizes by lifecycle; Red Flags converts to a
rationalization table. Brainstormed with Jesse 2026-07-15.
2026-07-19 12:36:33 -07:00
Jesse Vincent
0634449ca6 fix(tests): stop the SDD skill test flaking on timing and prose case
tests/claude-code/test-subagent-driven-development.sh failed
intermittently for two independent reasons:

- Budget mismatch: the file runs 9 prompts with a 90s timeout each
  (810s worst case) inside the runner's 600s per-file ceiling, so slow
  backend days produced spurious timeouts. Raise the runner default to
  900s and fix the help text, which claimed the default was 300.
- Case-sensitive prose matching: the assert helpers grepped free-form
  model output case-sensitively, but models capitalize the skill's own
  headings — observed failures include "Do Not Trust the Report"
  missing pattern "not trust" and a structured answer missing
  "First:.*spec.*compliance". Match case-insensitively in
  assert_contains/assert_not_contains/assert_count/assert_order, widen
  two Test 5 keyword patterns to phrasings observed in real runs, and
  make assert_order dump the output on failure the way assert_contains
  already does, so the next flake is diagnosable.

Observed 3 failures across 4 runs before the change (timeout, two
distinct pattern misses); 3/3 consecutive full runs pass after it.
2026-07-19 12:04:46 -07:00
Jesse Vincent
3fe3cb0530 fix(codex): make package script and its test portable beyond macOS/bsdtar
The packaging pipeline only worked on a Mac with default umask, for
three stacked reasons:

- The deterministic-metadata tar flags (--uid/--gid/--uname/--gname)
  are bsdtar spellings; GNU tar rejects them, so the tar.gz archive
  step died on Linux. Detect the tar flavor and use --owner=:0
  --group=:0 --numeric-owner on GNU tar, which writes byte-identical
  ustar headers (uid/gid 0, empty uname/gname).
- Staged file modes depended on two umasks canceling out: git archive
  masks entry modes with tar.umask (git default 0002 -> 775), and the
  unflagged tar extraction re-masked with the process umask (022 on
  macOS -> 755, but 002 elsewhere -> 775). Pin tar.umask=0022 on the
  archive call and extract with -p so staged modes are canonical
  755/644 on every machine.
- The test's timestamp assertion parsed bsdtar's -tv column layout and
  expected epoch 0 rendered in a US timezone ("Dec 31 1969"); GNU tar
  uses different columns and UTC hosts render "1970-01-01". Assert
  mtime == 0 via python3 tarfile instead, matching how the test
  already checks zip timestamps.

tests/codex/test-package-codex-plugin.sh now passes on Linux/GNU tar;
the bsdtar branch preserves the exact flags that passed on macOS.
2026-07-19 12:04:46 -07:00
Jesse Vincent
fe0b24390e docs(windows): document shell:bash hook dispatch and the PowerShell/CMD fallback hazards 2026-07-19 12:03:59 -07:00
Jesse Vincent
df78c6bfaf fix(hooks): dispatch the SessionStart hook via Git Bash on Windows
The SessionStart command string starts with a quoted path, which breaks
both Windows shells Claude Code may hand it to: PowerShell parses the
leading quoted string as an expression and dies on the next bareword
('Unexpected token session-start', #1751), and cmd.exe's /c quote rule
drops the outer quotes when the path contains a metacharacter, so a
profile dir like C:\Users\Name(External) truncates the command at the
'(' (#1918). Either way the bootstrap silently never loads.

Declare shell: "bash" on the hook. Claude Code >= 2.1.81 then resolves
Git for Windows and runs the polyglot's bash path directly — the same
route it already picks when it detects Git Bash — and when Git Bash is
missing it surfaces an actionable install prompt instead of a parser
error. Older versions ignore the unknown key and behave exactly as
before (verified live on 2.0.77 and 2.1.80).

Verified end-to-end with real claude sessions: Linux (hook fires,
bootstrap injected), Windows 11 + Git Bash under a path containing
'(' and a space (fires, 3276-char context), and Windows 11 without
Git Bash (actionable error replaces the #1751 ParserError, reproduced
verbatim as control).

Fixes #1751
Fixes #1918
2026-07-19 12:03:59 -07:00
Jesse Vincent
30ff376cb6 chore(sdd): consistency sweep for plan-scoped workspace signatures 2026-07-19 12:03:18 -07:00
Jesse Vincent
75f4e9414e eval(sdd): GREEN results — plan-scoped resolution replaces cross-plan forensics 2026-07-19 12:03:18 -07:00
Jesse Vincent
c15e041e03 feat(sdd): plan-scoped durable progress — ledger names its plan, workspace dies at plan end
The start-of-skill ledger check is now scoped to the plan's own
workspace and keyed to the ledger's first line. Baseline eval (25/25
reps) showed controllers already refuse foreign ledgers — at a cost of
6-13 tool calls of cross-plan forensics per resume; plan-scoping makes
the answer structural instead. The workspace is deleted once the final
review is clean — git history is the durable record.
2026-07-19 12:03:18 -07:00
Jesse Vincent
9816a9cee2 feat(sdd): plan-scoped workspace — one .superpowers/sdd/<plan> dir per plan
sdd-workspace now requires the plan file and resolves
.superpowers/sdd/<plan-basename>/; task-brief and review-package write
into their plan's directory (review-package gains PLAN_FILE as its first
argument). Follow-up plans in the same working tree can no longer collide
with a previous plan's briefs, reports, or ledger.
2026-07-19 12:03:18 -07:00
Jesse Vincent
9d9eae52f9 eval(sdd): RED baseline — 25/25 controllers refuse stale ledgers, at a forensic cost 2026-07-19 12:03:18 -07:00
Jesse Vincent
6ddb0bfcd9 docs(specs): record eval re-scope — blind adoption did not reproduce, claims narrowed
25/25 baseline reps refused the stale foreign ledger via git forensics;
the spec's evaluation section now states the honest claims: structural
fix + measured disambiguation-cost delta + same-plan-resume regression
gate, shipping with explicit maintainer sign-off in place of a failing
S1 baseline.
2026-07-19 12:03:18 -07:00
Jesse Vincent
194907435d docs(plans): re-scope eval per maintainer decision — RED compiled, GREEN measures cost
Three RED rounds (25 reps, three framings incl. faithful compaction
resume) never reproduced blind stale-ledger adoption: sonnet controllers
forensically refuse foreign ledgers, spending 6-13 tool calls per resume
doing it. Jesse approved shipping the full change with the eval re-scoped
to what is true: Task 1 compiles the existing RED evidence, Task 4 runs
GREEN on a truthful v3 fixture (real implementations, rotating authors)
with an S2 released-text control, measuring regression safety and the
disambiguation-cost delta instead of an error rate.
2026-07-19 12:03:18 -07:00
Jesse Vincent
c10431b14c docs(plans): fixture v2 — real cited commits, matched task counts
Fixture v1 tripped the Task 1 STOP gate for the right reason: its
ledgers cited fabricated hashes, so RED agents dismissed them via git
forensics (S1 passed for the wrong mechanism, the S2 resume control
failed 5/5). v2 executes plan A's tasks as real commits, gives both
plans five tasks so numbering is ambiguous, adds a symmetric
resume-uncertainty line to the scenario prompt, hard-stops if the S2
control fails twice, and drops rm -rf from cleanup (hook-gated here).
2026-07-19 12:03:18 -07:00
Jesse Vincent
0da87665c8 docs(plans): SDD plan-scoped workspace implementation plan
Five tasks: RED baseline eval (writing-skills Iron Law — before any
skill edit), plan-scoped scripts via TDD, SKILL.md durable-progress
rewrite with mismatch guard and end-of-plan cleanup, GREEN eval with
refinement loop, consistency sweep. Eval = 5 fresh sonnet subagents per
scenario per arm, hand-scored.
2026-07-19 12:03:18 -07:00
Jesse Vincent
20940deae8 docs(specs): SDD plan-scoped workspace design
The .superpowers/sdd workspace has no plan identity and no end-of-life:
follow-up plans in the same worktree read the previous plan's ledger as
their own progress, and artifacts leak into git (observed in serf, three
contamination rounds and ad-hoc progress-p2/p3 workarounds). Structural
fix: per-plan workspace subdirs, ledger names its plan, delete the
workspace when the final review is clean.
2026-07-19 12:03:18 -07:00
Jesse Vincent
fb7b07088e docs: fix dead references to pruned claude-code-tools.md/copilot-tools.md
e7ddc25 deleted claude-code-tools.md and copilot-tools.md but left
writing-skills and the porting guide's reference-integration table
pointing at them. State the current architecture instead: Claude Code's
personal-skills path inline, and "no adapter file needed" for the
harnesses that ride the Claude Code-compatible tool surface.

Reported by @rasibintang (#1969, with a fix proposed in #1970).

Fixes #1969
2026-07-15 19:15:16 +00:00
Gaurav Dubey
7a81eb7177 test(pi): scope mapping assertions to the table, not whole file
The pi tokens (subagent, pi-subagents, Task, TODO.md) also appear in the
surrounding prose, so matching the whole file passed even with the mapping
table deleted — the exact regression this test exists to catch. Filter to
table rows (lines starting with '|') so the assertion fails when the table
is gone and passes on dev.

Reported by @muunkky on #1987 (approach from #1983); verified failing-first
by stripping the table rows from pi-tools.md.
2026-07-15 11:10:55 -07:00
Gaurav Dubey
2b1c06a849 test: realign antigravity + pi mapping assertions with pruned references
Commit e7ddc25 ('Prune per-harness tool-mapping boilerplate') deliberately
removed the skill-loading explainers and generic action->tool tables from
antigravity-tools.md and pi-tools.md, keeping only the harness-specific
notes (subagent dispatch, task tracking). It did not touch tests/, so two
content-assertion tests kept asserting the removed tokens and now fail on
both dev and main:

  - tests/antigravity/test-antigravity-tools.sh: asserted view_file,
    IsSkillFile, run_command, grep_search (all pruned)
  - tests/pi/test-pi-extension.mjs: asserted read/write/edit/bash (pruned)

Update both to assert only the surviving harness-specific mappings. No
reference or skill content is changed; only the stale test assertions.
2026-07-15 11:10:55 -07:00
Jesse Vincent
4562d18dcf refactor(skills): fold TDD Why Order Matters rebuttals into rationalization table
The eval verdict on this cut: deleting Why Order Matters and trusting the
compressed one-line table rows measurably degrades test-first behavior under
the exact pressure the section rebutted ("just write it, tests after") —
control 8/10 → treatment 5/10 at n=10, corroborated on both Claude and Codex.
Normal TDD triggering did not move (PPPPP → PPPPP both arms); the damage is
purely the pressure case.

So instead of trusting the compressed rows, fold the section's five prose
rebuttals into their Common Rationalizations rows so each row carries the
argument, not just the excuse label:

- "I'll test after" — passing immediately proves nothing (wrong thing /
  implementation-not-behavior / missed edge; you never saw it fail).
- "Already manually tested" — ad-hoc, no record, can't re-run, forgotten
  under pressure.
- "Deleting X hours is wasteful" — sunk cost; rewrite-high-confidence vs
  bolt-tests-on-after-low-confidence.
- "TDD will slow me down" — TDD is the pragmatic path; shortcuts mean
  debugging in production.
- "Tests after achieve same goals (spirit not ritual)" — what-does vs
  what-should; biased by the code you wrote; coverage without proof.

Still removes the 50-line section (~200 words / 45 lines net); the
arguments survive where an agent hits them mid-rationalization. Revalidate
with the tdd-holds-under-tests-later-pressure probe before merge.
2026-07-14 15:02:16 -07:00
Jesse Vincent
14603727c8 refactor(skills): drop The Bottom Line recap from receiving-code-review
Restates the evaluate-don't-obey frame, verification rule, and
no-performative-agreement rule, each detailed earlier at point of use.
The Common Mistakes table stays: it is the skill's one compact guard
table, the class this cleanup standardizes toward rather than deletes.
2026-07-14 15:02:16 -07:00
Jesse Vincent
019e79cc46 refactor(skills): drop The Bottom Line recap from writing-skills
Restates the Iron Law, the RED-GREEN-REFACTOR mapping, and the
TDD-for-docs framing, all stated in full earlier in the file.
2026-07-14 15:02:16 -07:00
Jesse Vincent
d74653cf74 refactor(skills): drop Remember recap from writing-plans
All four lines restate the Overview (DRY/YAGNI/TDD/frequent commits),
Task Structure (exact paths, commands with expected output), and No
Placeholders (complete code in every step).
2026-07-14 15:02:16 -07:00
Jesse Vincent
3550dd05cd refactor(skills): fold brainstorming Key Principles into points of use
Five of six principles restated the Checklist and Process sections
verbatim-in-spirit. The sixth, YAGNI, appeared nowhere else — it moves to
the Exploring approaches list where designs get shaped; the recap section
goes.
2026-07-14 15:02:16 -07:00
Jesse Vincent
8489d22016 refactor(skills): convert using-git-worktrees guard sections to rationalization table
Common Mistakes and Red Flags restated Steps 0-3 wholesale; both fold
into one Common Rationalizations table (house Excuse/Reality form) whose
five rows carry the tempting-thought version of each rule, including the
#1-mistake emphasis on bypassing native tools. Quick Reference stays as
the compact decision aid.
2026-07-14 15:02:16 -07:00
Jesse Vincent
22d65cf8f0 refactor(skills): trim requesting-code-review, keep review guards as a table
Integration with Workflows restated the When to Request Review triggers
grouped by caller (each-task / before-merge / when-stuck all appear at
point of use) — detritus, so it goes.

The intro's crafted-context sentence guarded two things at once, so keep
both as Common Rationalizations rows (house Excuse/Reality form) rather
than deleting the sentence. The skill's reader is the coordinator, not
the code's author:

- Don't review the diff inline — that burns the coordinator's context
  window; dispatch a subagent so the diff and evaluation live in its
  context and only findings return. ("preserves your own context for
  continued work")
- Don't hand the reviewer your session history — crafted context keeps it
  on the work product, not your thought process.
2026-07-14 15:02:16 -07:00
Jesse Vincent
9d941bec3b refactor(skills): drop Advantages section from subagent-driven-development
Five blocks of benefits and cost/benefit selling aimed at a reader who
has already invoked the skill; the vs-Executing-Plans comparison also
duplicates the one under When to Use. Integration section untouched
(PR #1932 owns it).
2026-07-14 15:02:16 -07:00
Jesse Vincent
9da6fec633 refactor(skills): trim quality claim from executing-plans subagent note
The tell-your-partner directive and the prefer-SDD instruction stay; the
significantly-higher-quality sentence restated them as a claim.
Integration section untouched (PR #1932 owns it).
2026-07-14 15:02:16 -07:00
Jesse Vincent
43d87baeed refactor(skills): drop persuasion sections from verification-before-completion
Why This Matters (failure-memory testimonials), the dishonesty reframing
in the Overview, and The Bottom Line recap all restate stakes the Iron
Law, gate function, and rationalization table already enforce. This is
the eval-gated class: the bet is that discipline holds without the
persuasion prose — evals on this branch decide.
2026-07-14 15:02:16 -07:00
Jesse Vincent
c81f29fc6b refactor(skills): drop social proof from systematic-debugging
Real-World Impact was statistics; the Overview opener restated the core
principle as motivation. The 95%-of-no-root-cause line stays: it guards
the bail-out point, which is rationalization control, not social proof.
Supporting Techniques/Related skills untouched (PR #1932 owns that).
2026-07-14 15:02:16 -07:00
Jesse Vincent
5e046b3db2 refactor(skills): drop social proof from dispatching-parallel-agents
Real-World Impact restated the Real Example from Session as statistics;
Key Benefits and the time-saved line sold the skill to a reader already
executing it. Instructions unchanged.
2026-07-14 15:02:16 -07:00
Jesse Vincent
92164e2d1a experiment: ground-up two-principle rewrite of writing-good-tests
Re-derived from scratch: every rule becomes a corollary of two principles
(every test names the break it catches; every test exercises the real
thing), one consolidated gate per principle, four example pairs kept, the
rest carried by prose. Scratch branch for comparison against the accreted
eight-rule version.
2026-07-13 14:25:55 -07:00
Jesse Vincent
5431cf3b1d refactor(skills): compress writing-good-tests additions; doc changes earn no tests
Prose additions from the last two passes tightened to the terse guard
form: change-detector rule, string-presence trap, and Rule 7's release
valve each drop to a few sentences. Rule 7 now settles the jurisdiction
question outright: trivial code and human prose earn no test; skills and
prompts are pressure-tested per writing-skills when edits change
behavior, never text-asserted. Micro-tested: a subject with a README
rewrite plus a skill typo fix, under tests-with-every-PR pressure,
shipped zero tests — declining the string assertions and the ceremonial
subagent pressure-test alike.
2026-07-13 14:25:55 -07:00
Jesse Vincent
cb830c74fb fix(skills): close the change-detector hole in writing-good-tests
Fresh-eyes review found falsifiable-but-worthless tests passed every
rule: a constant assertion can fail, uses a literal, mocks nothing — and
protects nothing, firing on intentional decisions while sleeping through
bugs. Rule 1 gains the what-break-would-this-catch question (absorbed
from the source skill's quality gate, missed in the first pass) with a
gate stop for change detectors; Rule 6's trivial-code list regains
constants; Rule 7 gains the release valve that trivial-only changes earn
no ceremonial test; the coverage-theater and change-detector smells join
Warning Signs; the Rule 6 example stops modeling exact-copy brittleness.
Micro-tested: under a tests-with-every-PR norm, a subject rejected both
draft constant tests citing the new gate and replaced them with a test of
the retry behavior the constant controls.
2026-07-13 14:25:55 -07:00
Jesse Vincent
6a2d0c211f feat(skills): absorb falsifiability discipline into writing-good-tests
Generalized from agentsview's testing-without-tautologies skill: a new
Iron Law and lead rule (name the production change that would fail the
test, derive expectations independently of the code under test), a
test-your-code-not-the-framework rule with the characterization-test
exception and the trivial-code guidance, branch-specific doubles folded
into Mock at the Right Level, a closing Mutation Check, and six new
warning-sign smells. Rule 1 carries the string-presence trap by name:
grep-style tests on scripts, skills, and prompts counterfeit
falsifiability — the observable is the artifact's behavior, never its
text — with a hard stop in the gate function. Repo-specific content
(testify, backend parity, test-level ladder) stays in the source skill.
Micro-tested: 3/3 tautology verdicts with correct rule citations and the
mutation check named unprompted; a RED-pressure subject refused the
10-second grep test and wrote a behavioral one citing the trap.
2026-07-13 14:25:55 -07:00
Jesse Vincent
6a8869c7d2 fix(skills): broaden writing-good-tests trigger to any test writing
The pointer fired only on adding mocks or test utilities; the doc's own
load-when line already says writing or changing tests. The narrow trigger
would skip the rules exactly when an agent thinks no mocks are involved.
2026-07-13 14:25:55 -07:00
Jesse Vincent
40b2f3aaca refactor(skills): reframe testing-anti-patterns as writing-good-tests
The disclosure doc becomes a catalog of what to do: six positively named
rules (assert on real behavior, cleanup in test utilities, mock at the
right level, mirror real data, tests ship with implementation, prefer
real components), each leading with the GOOD example and keeping the
violation as contrast. Iron Laws, gate functions, human-partner lines,
and warning signs all survive; The Bottom Line recap and the
TDD-prevents-these section fold into one Overview sentence. SKILL.md's
pointer moves into the Good Tests section it belongs with. Micro-tested
2/2: a mock-existence assertion got rewritten to a real-behavior
assertion citing Rule 1, and a test-only teardown method plus a
to-be-safe mock were both rejected citing Rules 2 and 3.
2026-07-13 14:25:55 -07:00
Jesse Vincent
f68c94334d fix(skills): capture worktree path before Step 5 changes directory
Step 6 recomputed WORKTREE_PATH after Option 1 and discard had already
cd'd to the main repo root, so --show-toplevel returned the main root:
the provenance check could never match, cleanup silently no-oped, and the
branch delete failed with the worktree still attached. A test subject had
to deviate from the literal skill to produce a working sequence. The
capture moves to Step 2 (still inside the workspace); Step 6 consumes
Step 2's values and drops its redundant recompute and MAIN_ROOT
derivation. Also: Option 2 gains the detached-HEAD push variant its menu
advertises, and the stale-green rationalization row states what a green
run proves instead of asserting the tree changed. Re-verified: merge-flow
and discard-flow subjects both walk the literal skill to correct cleanup
with concrete paths and no deviations.
2026-07-13 14:25:32 -07:00
Jesse Vincent
df93818856 refactor(skills): compress finishing-a-development-branch, adopt rationalization table
Red Flags and Common Mistakes fold into one Common Rationalizations table
(house Excuse/Reality form); every prior entry maps to a table row or an
inline sentence in the step it guards. Instructions rephrase positively —
what to do rather than what to avoid — with negations remaining only in
statements of fact. Workflow prose tightens throughout; menus, detection
mechanics, cleanup provenance, and the typed-discard ritual are unchanged.
Re-verified 4/4 after the rewrite: both menus verbatim, the lukewarm-human
pressure arm cited the rationalizations table when declining to offer
discard, and a prose discard request still required the literal typed
word.
2026-07-13 14:25:32 -07:00
Jesse Vincent
6f81c378ac refactor(skills): make PR creation forge-agnostic in finishing-a-development-branch
Naming gh and glab implicitly blessed two forges; Gitea, Forgejo,
Bitbucket and others are equally valid. Point at the forge's CLI or the
creation URL printed on push instead of naming tools.
2026-07-13 14:25:32 -07:00
Jesse Vincent
a0487b028f refactor(skills): stop offering to discard work in finishing-a-development-branch
The completion menu dates from when throwing away branches was routine;
offering 'Discard this work' beside 'Merge' on every completion advertised
destroying finished, passing work. The menu is now 3 options (2 detached
HEAD); discard survives as an explicit-request-only path with the same
typed-confirmation ritual and cleanup mechanics. Fresh-eyes fixes in the
same pass: Option 2 actually creates the pull/merge request
(platform-neutral tooling) and reports the URL; Step 3's base-branch
detection drops a command that printed a SHA instead of choosing a branch
(ask when not known); Option 1 gains a failure branch (merged-result test
failures stop cleanup); description trimmed to trigger-only. Micro-tested
4/4: both menus verbatim with no discard, no discard offer even when the
human sounded lukewarm about the feature, and a prose 'throw it all away'
still required the typed confirmation before any deletion.
2026-07-13 14:25:32 -07:00
Jesse Vincent
5ce5a40703 refactor(skills): fold systematic-debugging Related-skills block into Phase 4
Same treatment as subagent-driven-development and executing-plans: the
test-driven-development entry duplicated the reference already at Phase 4
Step 1, and the verification-before-completion entry was a sole carrier —
it moves to its point of use in Phase 4 Step 3 (Verify Fix). Micro-tested
2/2: subjects at the just-implemented-a-fix point invoke
verification-before-completion before any success claim, including under
ship-pressure.
2026-07-13 14:25:08 -07:00
Jesse Vincent
ab4fa6b09f refactor(skills): fold Integration skill lists into points of use
The list-style Integration sections in subagent-driven-development and
executing-plans duplicated references that already exist where the flow
uses them (process digraph, When to Use, prompt templates, Step 3), so
they added maintenance cost without carrying behavior. The one entry not
duplicated anywhere — the using-git-worktrees isolated-workspace
requirement — moves to its point of use: SDD's Pre-Flight Plan Review and
executing-plans' Step 1. Micro-tested 5/5: controllers at skill start
establish or verify the worktree before reading the plan or dispatching
Task 1, including under skip-the-ceremony pressure. The prose Integration
sections in requesting-code-review and other skills are unchanged — they
carry placement content, not an index.
2026-07-13 14:25:08 -07:00
Ada Sen
096e15aa73 Revert "Remove Gemini CLI support"
This reverts commit 711d895ce7.
2026-07-10 11:58:08 -04:00
Jesse Vincent
c809093a2a Release v6.1.1: fix Codex SessionStart hook re-registration, add Codex portal packaging 2026-07-02 14:53:00 -07:00
Drew Ritter
97506cefd7 Preserve hooks in Codex package manifest 2026-07-02 14:53:00 -07:00
Drew Ritter
4ecbbcd0b4 Strip hooks from Codex portal package 2026-07-02 14:53:00 -07:00
Drew Ritter
53106e6536 docs: re-anchor Shape A examples away from Codex 2026-07-02 14:53:00 -07:00
Drew Ritter
89338e5113 chore(codex): remove orphaned session-start-codex hook + refresh hook docs
hooks/session-start-codex has had no caller since "Remove Codex hooks"
(#1845) deleted hooks-codex.json and its manifest registration; the
Codex manifest now declares an empty hooks object so Codex registers no
session-start hook at all. The script is Codex-specific dead code —
nothing executes it on Codex or any other harness.

- Delete hooks/session-start-codex.
- tests/hooks/test-session-start.sh: drop the two Codex cases that are
  redundant with the generic session-start tests (nested-format and the
  legacy-warning omission are already covered by the Claude Code cases).
  Re-point the "wrapper dispatches" case to the live `session-start`
  script so run-hook.cmd dispatch coverage — used by Claude Code and
  Cursor in production — is preserved rather than lost.
- docs/porting-to-a-new-harness.md: Codex is no longer a Shape A
  (shell-hook) harness, so re-anchor that worked example to Cursor (a
  live shell-hook harness that demonstrates the same per-harness field,
  schema, and matcher variance) and mark Codex as native skill discovery
  with no session-start hook. Clears the references to the deleted
  hooks-codex.json.
- docs/windows/polyglot-hooks.md: the "check hooks-codex.json" pointer
  referenced a file deleted in #1845; re-point to hooks-cursor.json.

RELEASE-NOTES.md keeps its historical mention of hooks-codex.json (it
accurately records what that release did). The tests/codex-plugin-sync
fixtures build their own synthetic session-start-codex and test the sync
mechanism generically, so they are intentionally left as-is.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 14:53:00 -07:00
Drew Ritter
c842f8871a Fix Codex plugin category 2026-07-02 14:53:00 -07:00
Drew Ritter
6752471ad9 Default Codex portal package to zip 2026-07-02 14:53:00 -07:00
Drew Ritter
371a26cf99 Harden Codex package script checks 2026-07-02 14:53:00 -07:00
Drew Ritter
3bb0a3faa3 Add Codex portal package script 2026-07-02 14:53:00 -07:00
Drew Ritter
2d05b63edc fix(codex): suppress SessionStart hook auto-discovery with empty hooks object
Codex auto-discovers a plugin's hooks/hooks.json whenever the Codex
manifest has no `hooks` field: load_plugin_hooks falls back to a
hardcoded DEFAULT_HOOKS_CONFIG_FILE = "hooks/hooks.json" and registers
it. hooks/hooks.json is the Claude Code SessionStart hook, it is tracked
in this repo, and the Codex marketplace installs the whole repo root
(source url "./"), so the fallback re-registered the SessionStart hook
and its install-time trust prompt on Codex.

Removing the Codex hook file and the manifest `hooks` pointer (commit
"Remove Codex hooks") did not disable the hook on Codex — it removed the
explicit declaration that was overriding the fallback, so the fallback
took over and found the Claude hooks/hooks.json.

Declare an empty inline hooks object ({}) in .codex-plugin/plugin.json.
It parses as an empty inline hook set and stops Codex reaching the
auto-discovery fallback. An absent field, an empty array ([]), and an
empty inline list all collapse back to the fallback, so the value must
be exactly {}.

Update the test to assert the manifest declares hooks: {} (and that
hooks/hooks.json exists, which is what makes the declaration necessary),
replacing the prior assertion that the field was absent — which passed
while the hook was still being auto-discovered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 14:53:00 -07:00
28 changed files with 2612 additions and 44 deletions

6
.gitignore vendored
View File

@@ -11,3 +11,9 @@ triage/
# development (see CLAUDE.md / README.md). It is not part of the published
# plugin, so the whole directory is ignored here.
evals/
# Python
__pycache__/
*.pyc
*.pyo
.pytest_cache/

104
.hermes-plugin/__init__.py Normal file
View File

@@ -0,0 +1,104 @@
import os
import re
from pathlib import Path
BOOTSTRAP_MARKER = "superpowers:using-superpowers bootstrap for hermes"
def _skills_dir() -> str:
"""Locate the stock skills/ tree for either supported install layout.
- git-clone install (`hermes plugins install obra/superpowers`): the plugin
dir is the repo root, so `.hermes-plugin/` and `skills/` are siblings and
this module resolves `../skills`.
- flattened install (plugin files copied to the plugin dir root): `skills/`
sits next to this module.
Raises loudly when neither matches — a bootstrap that silently skips is how
a broken install masquerades as a working one.
"""
here = os.path.dirname(os.path.realpath(__file__))
candidates = (
os.path.realpath(os.path.join(here, "..", "skills")),
os.path.realpath(os.path.join(here, "skills")),
)
for cand in candidates:
if os.path.isfile(os.path.join(cand, "using-superpowers", "SKILL.md")):
return cand
raise RuntimeError(
"superpowers plugin: cannot find the skills/ tree "
f"(looked at {candidates}). Reinstall with "
"`hermes plugins install obra/superpowers`."
)
def _strip_frontmatter(content: str) -> str:
match = re.match(r"^---\n[\s\S]*?\n---\n([\s\S]*)$", content)
return (match.group(1) if match else content).strip()
def _build_bootstrap(skills_dir: str) -> str:
with open(
os.path.join(skills_dir, "using-superpowers", "SKILL.md"),
encoding="utf-8",
) as f:
body = _strip_frontmatter(f.read())
tools_path = os.path.join(
skills_dir, "using-superpowers", "references", "hermes-tools.md"
)
with open(tools_path, encoding="utf-8") as f:
tool_mapping = f.read().strip()
return (
f"<EXTREMELY_IMPORTANT>\n"
f"{BOOTSTRAP_MARKER}\n\n"
f"You have superpowers.\n\n"
f"The using-superpowers skill content is included below and is already "
f"loaded for this Hermes session. Follow it now. "
f"Do not try to load using-superpowers again.\n\n"
f"{body}\n\n"
f"## Loading Superpowers Skills on Hermes\n\n"
f"Superpowers skills are registered with Hermes' native skill loader: "
f'invoke one with `skill_view("superpowers:skill-name")` '
f'(for example `skill_view("superpowers:brainstorming")`). '
f"If a namespaced lookup returns 'not found', read the skill file "
f"directly instead:\n"
f'`read_file("{skills_dir}/skill-name/SKILL.md")`\n\n'
f"The superpowers skills directory is: `{skills_dir}`\n\n"
f"{tool_mapping}\n"
f"</EXTREMELY_IMPORTANT>"
)
def register(ctx):
skills_dir = _skills_dir()
bootstrap = _build_bootstrap(skills_dir)
# Register every stock skill with Hermes' native loader so skill_view can
# load them on demand. Standard markdown; no conversion (plugin guide).
# register_skill requires a pathlib.Path — a str raises AttributeError and
# hermes silently disables the whole plugin (verified 2026-07-23).
for name in sorted(os.listdir(skills_dir)):
skill_md = os.path.join(skills_dir, name, "SKILL.md")
if os.path.isfile(skill_md):
ctx.register_skill(name, Path(skill_md))
# pre_llm_call returning {"context": ...} is the documented injection path
# (on_session_start return values are ignored, and ctx.inject_message
# refuses from that hook — verified empirically 2026-07-23). The context is
# appended to the first turn's user message.
def pre_llm_call(
session_id=None,
user_message=None,
conversation_history=None,
is_first_turn=None,
model=None,
platform=None,
**kwargs,
):
if is_first_turn:
return {"context": bootstrap}
return None
ctx.register_hook("pre_llm_call", pre_llm_call)

View File

@@ -0,0 +1,6 @@
name: superpowers
version: 6.2.0
description: Superpowers skills and workflow bootstrap for Hermes Agent
author: obra
provides_hooks:
- pre_llm_call

View File

@@ -1,6 +1,7 @@
{
"files": [
{ "path": "package.json", "field": "version" },
{ "path": ".hermes-plugin/plugin.yaml", "field": "version" },
{ "path": ".claude-plugin/plugin.json", "field": "version" },
{ "path": ".cursor-plugin/plugin.json", "field": "version" },
{ "path": ".codex-plugin/plugin.json", "field": "version" },

View File

@@ -2,10 +2,35 @@
Superpowers is a complete software development methodology for your coding agents, built on top of a set of composable skills and some initial instructions that make sure your agent uses them.
## Table of Contents
- [Quickstart](#quickstart)
- [How it works](#how-it-works)
- [Commercial Services](#commercial-services)
- [Installation](#installation)
- [Claude Code](#claude-code)
- [Antigravity](#antigravity)
- [Codex App](#codex-app)
- [Codex CLI](#codex-cli)
- [Cursor](#cursor)
- [Factory Droid](#factory-droid)
- [Gemini CLI](#gemini-cli)
- [GitHub Copilot CLI](#github-copilot-cli)
- [Kimi Code](#kimi-code)
- [OpenCode](#opencode)
- [Pi](#pi)
- [The Basic Workflow](#the-basic-workflow)
- [Community](#community)
- [What's Inside](#whats-inside)
- [Philosophy](#philosophy)
- [Contributing](#contributing)
- [Updating](#updating)
- [License](#license)
- [Visual companion telemetry](#visual-companion-telemetry)
## Quickstart
Give your agent Superpowers: [Claude Code](#claude-code), [Antigravity](#antigravity), [Codex App](#codex-app), [Codex CLI](#codex-cli), [Cursor](#cursor), [Factory Droid](#factory-droid), [Gemini CLI](#gemini-cli), [GitHub Copilot CLI](#github-copilot-cli), [Kimi Code](#kimi-code), [OpenCode](#opencode), [Pi](#pi).
Give your agent Superpowers: [Claude Code](#claude-code), [Antigravity](#antigravity), [Codex App](#codex-app), [Codex CLI](#codex-cli), [Cursor](#cursor), [Factory Droid](#factory-droid), [Gemini CLI](#gemini-cli), [GitHub Copilot CLI](#github-copilot-cli), [Hermes Agent](#hermes-agent), [Kimi Code](#kimi-code), [OpenCode](#opencode), [Pi](#pi).
## How it works
@@ -193,6 +218,18 @@ pi -e /path/to/superpowers
The Pi package loads the Superpowers skills and a small extension that injects the `using-superpowers` bootstrap at session startup and again after compaction. Pi has native skills, so no compatibility `Skill` tool is required. Subagent and task-list tools remain optional Pi companion packages.
### Hermes Agent
Install Superpowers as a Hermes plugin from this repository:
```bash
hermes plugins install obra/superpowers --enable
```
Restart any active Hermes sessions after installing. Note: Hermes has no
post-compaction hook, so a very long session that compacts over its first
turn loses the bootstrap — start a fresh session if skills stop triggering.
## The Basic Workflow
1. **brainstorming** - Activates before writing code. Refines rough ideas through questions, explores alternatives, presents design in sections for validation. Saves design document.
@@ -211,6 +248,14 @@ The Pi package loads the Superpowers skills and a small extension that injects t
**The agent checks for relevant skills before any task.** Mandatory workflows, not suggestions.
## Community
Superpowers is built by [Jesse Vincent](https://blog.fsck.com) and the rest of the folks at [Prime Radiant](https://primeradiant.com).
- **Discord**: [Join us](https://discord.gg/35wsABTejz) for community support, questions, and sharing what you're building with Superpowers
- **Issues**: https://github.com/obra/superpowers/issues
- **Release announcements**: [Sign up](https://primeradiant.com/superpowers/) to get notified about new versions
## What's Inside
### Skills Library
@@ -271,11 +316,3 @@ MIT License - see LICENSE file for details
## Visual companion telemetry
Because skills and plugins don't provide any feedback to creators, we have no idea how many of you are using Superpowers. By default, the Prime Radiant logo on brainstorming's optional visual companion feature is loaded from our website. It includes the version of Superpowers in use. It does not include any details about your project, prompt, or coding agent. We don't see your clicks or anything about what you're building. This helps us have a rough idea of how many folks are using Superpowers and which version of Superpowers they're using. It's 100% optional. To disable this, set the environment variable `SUPERPOWERS_DISABLE_TELEMETRY` to any true value. Superpowers also honors Claude Code's `DISABLE_TELEMETRY` and `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC` opt-outs.
## Community
Superpowers is built by [Jesse Vincent](https://blog.fsck.com) and the rest of the folks at [Prime Radiant](https://primeradiant.com).
- **Discord**: [Join us](https://discord.gg/35wsABTejz) for community support, questions, and sharing what you're building with Superpowers
- **Issues**: https://github.com/obra/superpowers/issues
- **Release announcements**: [Sign up](https://primeradiant.com/superpowers/) to get notified about new versions

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,304 @@
# Hermes Version-Bump Wiring Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Keep the Hermes YAML manifest version synchronized with every other declared release manifest.
**Spec:** `docs/superpowers/specs/2026-08-05-hermes-version-bump-wiring-design.md`
**Architecture:** Extend the existing release script with a small extension-based dispatcher: JSON continues through `jq`, while `.yaml` uses Mike Farah `yq` v4. Before the mutating bump loop, read every present manifest through that dispatcher so deterministic format or field failures occur before the first write.
**Tech Stack:** Bash 3.2-compatible shell, `jq`, Mike Farah `yq` v4, existing shell-lint tooling.
## Global Constraints
- Support only `.json` and `.yaml`; `.yml` and other extensions remain unsupported.
- YAML fields are present top-level strings; nested YAML fields are out of scope.
- Pass the YAML field and new value through environment data, never interpolate either into a `yq` expression.
- Keep `yq` confined to maintainer release tooling; do not add a plugin runtime dependency.
- Preserve the existing missing-file behavior: `--check` reports missing files and a bump skips them.
- Preflight only the mutating bump path; do not add rollback or transactional writes.
- Do not change audit status behavior, version validation, or the existing JSON field-expression implementation.
---
## File Map
- Create: `tests/version-bump/test-bump-version.sh`
- Exercise the real script in temporary JSON/YAML fixtures and check the real registry.
- Modify: `scripts/bump-version.sh`
- Add YAML read/write helpers, format dispatch, and bump-only read preflight.
- Modify: `.version-bump.json`
- Register `.hermes-plugin/plugin.yaml` at top-level field `version`.
### Task 1: Wire Hermes Into The Existing Version-Bump Script
**Files:**
- Create: `tests/version-bump/test-bump-version.sh`
- Modify: `scripts/bump-version.sh`
- Modify: `.version-bump.json`
**Interfaces:**
- Consumes: `.version-bump.json` records shaped as `{ "path": string, "field": string }`.
- Produces: `read_manifest_field FILE FIELD`, `write_manifest_field FILE FIELD VALUE`, and `preflight_manifests` Bash helpers.
- [ ] **Step 1: Fetch the current development base**
Run:
```bash
git fetch origin dev
```
Expected: command exits 0 and refreshes `origin/dev`.
- [ ] **Step 2: Rebase the task branch**
Run:
```bash
git rebase origin/dev
```
Expected: command exits 0, and `git status --short --branch` no longer reports the branch behind `origin/dev`.
- [ ] **Step 3: Add the initial failing behavioral test**
Create `tests/version-bump/test-bump-version.sh` with the happy-path fixture and real registry assertion:
```bash
#!/usr/bin/env bash
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
REPO_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)"
SCRIPT_SOURCE="$REPO_ROOT/scripts/bump-version.sh"
TEST_ROOT="$(mktemp -d)"
cleanup() {
rm -rf "$TEST_ROOT"
}
trap cleanup EXIT
fail() {
echo "FAIL: $*" >&2
exit 1
}
make_fixture() {
local repo="$1"
local yaml_body="$2"
mkdir -p "$repo/scripts" "$repo/.hermes-plugin"
cp "$SCRIPT_SOURCE" "$repo/scripts/bump-version.sh"
cat >"$repo/.version-bump.json" <<'JSON'
{
"files": [
{ "path": "package.json", "field": "version" },
{ "path": ".hermes-plugin/plugin.yaml", "field": "version" }
],
"audit": { "exclude": [] }
}
JSON
cat >"$repo/package.json" <<'JSON'
{
"name": "fixture",
"version": "1.2.3"
}
JSON
printf '%s\n' "$yaml_body" >"$repo/.hermes-plugin/plugin.yaml"
}
happy_repo="$TEST_ROOT/happy"
make_fixture "$happy_repo" $'name: superpowers\nversion: 1.2.3'
/bin/bash "$happy_repo/scripts/bump-version.sh" --check >"$TEST_ROOT/check.out"
/bin/bash "$happy_repo/scripts/bump-version.sh" --audit >"$TEST_ROOT/audit.out"
/bin/bash "$happy_repo/scripts/bump-version.sh" 2.3.4 >"$TEST_ROOT/bump.out"
[[ "$(jq -r '.version' "$happy_repo/package.json")" == "2.3.4" ]] \
|| fail "JSON manifest was not bumped"
[[ "$(yq -r '.version' "$happy_repo/.hermes-plugin/plugin.yaml")" == "2.3.4" ]] \
|| fail "YAML manifest was not bumped"
jq -e '
any(.files[];
.path == ".hermes-plugin/plugin.yaml" and .field == "version")
' "$REPO_ROOT/.version-bump.json" >/dev/null \
|| fail "Hermes manifest is not registered"
echo "Version-bump tests passed"
```
- [ ] **Step 4: Run the test to verify RED**
Run:
```bash
/bin/bash tests/version-bump/test-bump-version.sh
```
Expected: FAIL before `Version-bump tests passed`; the current JSON-only reader cannot process the YAML fixture.
- [ ] **Step 5: Add minimal YAML dispatch and register Hermes**
In `scripts/bump-version.sh`, add these helpers after `write_json_field`:
```bash
require_tool() {
command -v "$1" >/dev/null 2>&1 || {
echo "error: required tool '$1' is not on PATH" >&2
return 1
}
}
read_yaml_field() {
local file="$1" field="$2"
require_tool yq || return 1
FIELD="$field" yq -er '.[strenv(FIELD)] | select(tag == "!!str")' "$file"
}
write_yaml_field() {
local file="$1" field="$2" value="$3"
FIELD="$field" VALUE="$value" \
yq -i '.[strenv(FIELD)] = strenv(VALUE)' "$file"
}
read_manifest_field() {
local file="$1"
case "$file" in
*.json) read_json_field "$@" ;;
*.yaml) read_yaml_field "$@" ;;
*)
echo "error: unsupported manifest format: $file" >&2
return 1
;;
esac
}
write_manifest_field() {
local file="$1"
case "$file" in
*.json) write_json_field "$@" ;;
*.yaml) write_yaml_field "$@" ;;
*)
echo "error: unsupported manifest format: $file" >&2
return 1
;;
esac
}
```
Replace the three command-path calls to `read_json_field` with `read_manifest_field`, and replace the bump-path call to `write_json_field` with `write_manifest_field`.
Add this exact entry to `.version-bump.json` immediately after `package.json`:
```json
{ "path": ".hermes-plugin/plugin.yaml", "field": "version" },
```
- [ ] **Step 6: Run the initial test to verify GREEN**
Run:
```bash
/bin/bash tests/version-bump/test-bump-version.sh
```
Expected: PASS with `Version-bump tests passed`.
- [ ] **Step 7: Add the failing no-partial-write regression**
Insert this block before the final success message in `tests/version-bump/test-bump-version.sh`:
```bash
invalid_repo="$TEST_ROOT/invalid"
make_fixture "$invalid_repo" $'name: superpowers\nversion: 123'
cp "$invalid_repo/package.json" "$TEST_ROOT/package.before"
cp "$invalid_repo/.hermes-plugin/plugin.yaml" "$TEST_ROOT/plugin.before"
if /bin/bash "$invalid_repo/scripts/bump-version.sh" 2.3.4 \
>"$TEST_ROOT/invalid.out" 2>&1; then
fail "bump accepted a non-string YAML version"
fi
cmp -s "$TEST_ROOT/package.before" "$invalid_repo/package.json" \
|| fail "JSON manifest changed before YAML validation failed"
cmp -s "$TEST_ROOT/plugin.before" "$invalid_repo/.hermes-plugin/plugin.yaml" \
|| fail "invalid YAML manifest changed"
```
- [ ] **Step 8: Run the regression to verify RED**
Run:
```bash
/bin/bash tests/version-bump/test-bump-version.sh
```
Expected: FAIL with `JSON manifest changed before YAML validation failed`; without preflight, the JSON manifest is written before the later YAML reader rejects its non-string version.
- [ ] **Step 9: Add the bump-only preflight**
Add this helper after `declared_files` in `scripts/bump-version.sh`:
```bash
preflight_manifests() {
local path field fullpath
require_tool jq || return 1
while IFS=$'\t' read -r path field; do
fullpath="$REPO_ROOT/$path"
[[ -f "$fullpath" ]] || continue
if ! read_manifest_field "$fullpath" "$field" >/dev/null; then
echo "error: cannot read declared manifest: $path ($field)" >&2
return 1
fi
done < <(declared_files)
}
```
Call it in `cmd_bump` after version-format validation and before the first bump output or write:
```bash
preflight_manifests
echo "Bumping all declared files to $new_version..."
```
- [ ] **Step 10: Run focused verification**
Run:
```bash
/bin/bash tests/version-bump/test-bump-version.sh
scripts/lint-shell.sh scripts/bump-version.sh tests/version-bump/test-bump-version.sh
scripts/bump-version.sh --check
git diff --check
```
Expected:
- The behavioral test prints `Version-bump tests passed`.
- Shell lint reports both scripts with no errors.
- `--check` lists eight declared manifests, including `.hermes-plugin/plugin.yaml`, all at `6.2.0`.
- `git diff --check` prints nothing.
- [ ] **Step 11: Review and commit the implementation**
Run:
```bash
git status --short
git diff -- .version-bump.json scripts/bump-version.sh tests/version-bump/test-bump-version.sh
git add .version-bump.json scripts/bump-version.sh tests/version-bump/test-bump-version.sh
git commit \
-m "fix(release): wire Hermes into version bumps" \
-m "Register the Hermes YAML manifest alongside the existing JSON manifests. Route manifest reads and writes by extension through jq or Mike Farah yq v4, with field names and values passed as data." \
-m "Preflight every present manifest before the mutating bump loop so a deterministic YAML read failure cannot leave earlier JSON manifests partially updated. Cover check, audit, bump, registry wiring, and byte-for-byte no-partial-write behavior with one focused fixture test."
```
Expected: the commit succeeds with only the three implementation paths staged.

View File

@@ -0,0 +1,252 @@
# Codex Efficiency Fixes — Design
Date: 2026-07-30
Status: approved by Jesse (in-session)
Branch: `codex-efficiency-fixes` off `dev`
## Sources
- Eval campaign closeout: `superpowers-autoresearch/reports/2026-07-codex-efficiency-campaign.md`
(treatment table §4; every treatment below has a scorer and a measured
`dev` baseline).
- Codex source recon: `superpowers-autoresearch/docs/2026-07-29-codex-multiagent-v2-capabilities.md`
(file:line citations against the Codex CLI source; grounds T2, T3, T5).
- Published experiment write-ups: `superpowers-evals/docs/experiments/`.
- Drew's spinout stack (PRs #2036, #2035) is **evidence, not adopted text**:
Jesse wants to dig into those fixes in more detail before adopting any
of them; they inform the problem statements only.
## Goal
Ship the five evidence-strong treatments from the codex-efficiency eval
campaign as superpowers skill/doc changes, each graded against its
pre-registered criterion by the campaign's scorers before its PR is cut.
Phase 2 (everything else in the closeout treatment table) follows, each
item gated on new baseline work first.
## Scope decisions (settled with Jesse)
- **Phase 1 = the evidence-strong five** (T1T5 below). Phase 2 items
each need a failing baseline before any fix ships (discrimination
rule: inconclusive-by-zero is a stop).
- **One branch, PR per treatment.** Development and batteries happen on
`codex-efficiency-fixes`; when a treatment beats its criterion, it is
cut into its own PR against `dev` with its eval evidence. No merge
without Jesse's per-PR approval.
- **T4 ships cross-harness with a global regression battery** (Claude
Code, Codex, Gemini), variant C shape: ceremony scales, approval never
does.
## The five treatments
### T1. SDD worker-review prohibition
**Evidence:** 9/9 depth-2 spawns across 4 corpora were implementer-issued
reviewers; all 9 were same-task duplicates of the review the controller
dispatches anyway. The dispatch contract never says review is not the
worker's job; "self-review" in the implementer prompt gets reified into a
reviewer subagent on harnesses where children can spawn (Codex).
**Changes:**
- `skills/subagent-driven-development/implementer-prompt.md`: an explicit
"You do not dispatch subagents" clause — self-review means reading your
own diff; the controller owns all review dispatch; a reviewer you spawn
duplicates a review the process already provides.
- `skills/subagent-driven-development/SKILL.md`: one dispatch-contract
line in the task loop, plus a Red Flags row: "An independent review
would strengthen my report" → review is the controller's next step;
your reviewer is a duplicate seat.
- Harness-agnostic wording (no-op where children cannot spawn).
**Graded by:** `score_e6.py` (depth-2 spawns by spawner role, duplicate
review families); `score_e5.py` for the same-scope variant.
**Baseline:** 9/9 worker-issued, 0 counter-examples.
**Criterion:** 0 worker-issued depth-2 spawns AND review coverage
preserved (every task still gets exactly one controller-dispatched task
review).
### T2. Event-driven waiting
**Evidence:** 6078% of `wait_agent` calls time out in every corpus
(dev 67.1%, spinout 60.2%). Source recon: V2 waits are event
subscriptions, not polls — one long wait has the same wake latency as a
10s poll at ~1/90th the calls; a completed child's FINAL_ANSWER is pushed
into the parent's mailbox and drained into the next model request with no
wait at all.
**Changes** (`skills/using-superpowers/references/codex-tools.md`):
- Never short-timeout poll.
- While local work remains, do not wait — child results arrive with your
next turn via the mailbox.
- When genuinely idle, issue ONE `wait_agent` with a long `timeout_ms`
(900000+; harness max 3600000).
- V2 caveat stated: completion mail carries `trigger_turn=false` and will
not wake an idle controller — that is the one job `wait_agent` has.
**Graded by:** `score_e7.py` (timeout rate, inter-poll cadence,
cache-rebill estimate — the rebill figure stays labeled as an estimate).
**Baseline:** dev 67.1% timeout rate.
**Criterion:** timeout rate < 25% with no loss of task completion.
### T3. codex-tools.md corrections
**Evidence:** five claims in the current guidance are contradicted by the
Codex source (all file:line-cited in the capabilities doc):
1. `close_agent` does not exist in multi-agent V2 (V1-only). V2 LRU-evicts
finished children automatically; not closing costs nothing;
`followup_task` transparently reloads an evicted child.
2. Fix rounds can always resume the implementer via `followup_task`
dev's "if your harness cannot send another message to a spawned agent,
dispatch each fix round as a fresh implementer" branch is dead on V2.
3. Role files (`~/.codex/agents/**.toml`) DO attach to spawns via
`agent_type` on isolated forks (0.145+).
4. Full-history forks accept `model`/`reasoning_effort` overrides; only
`agent_type` is refused. (Isolated forks remain the SDD guidance for
context-hygiene reasons, stated accurately.)
5. Dispatch guidance must never name non-V2 model presets — the V2 spawn
allowlist is v2 presets only; others hard-error.
**Changes:** rewrite the multi-agent paragraph of
`skills/using-superpowers/references/codex-tools.md` to be
version-honest (V1 vs V2 behavior labeled where they differ).
**Graded by:** source citation (already verified); no scorer regressions
on the shared battery. `score_e8.py` is retained as a V1/V2 schema
detector, not a hygiene grader — no `close_agent` checklist ships.
### T4. Brainstorming three-path router (variant C: approval always)
**Evidence:** micro — the current HARD-GATE text pushes a bounded task to
FULL ceremony 5/5, while Z-null (no guidance) and a three-path router
both differentiate 5/5: the absolute wording suppresses discrimination
the model draws natively. FULL battery — ceremony volume scales
moderately (16.7 vs 24.0 tool calls, bounded vs arch), but the
two-document ritual (spec file → plan file) ran unconditionally in every
rep. The measured waste is the unconditional artifact ritual, not the
approval gate.
**Design (variant C):** three paths scale the ARTIFACT; every path keeps
human approval before implementation:
- **Spike** (feasibility question, explicitly throwaway): present the
question and the intended probe in 23 sentences, get a nod, go. No
docs. Findings return as a recommendation; anything built stays labeled
throwaway.
- **Bounded** (well-scoped change to an existing, understood flow):
present a short design in chat, get approval, implement. No spec file,
no writing-plans invocation.
- **Architectural** (restructures components, new subsystem, public
interface change): the full current flow — spec doc, review,
writing-plans.
**Guards (all ship with the router):**
- Classification is said out loud ("this looks bounded, so I'll present a
short design here rather than write a spec") so the human can override.
- When in doubt between two paths, take the heavier one.
- One-way ratchet: hidden complexity discovered mid-path upgrades the
path; never downgrade mid-task.
- New Red Flags rows targeting classification-as-escape-hatch ("I'll call
it bounded to skip the doc").
**Changes** (`skills/brainstorming/SKILL.md`): HARD-GATE keeps "no
implementation before approval" and drops "regardless of perceived
simplicity" as the ceremony driver; anti-pattern section reframed (the
sin is skipping approval, not skipping documents); checklist steps 69
become the architectural path; process-flow graph gains the router; Red
Flags rows added. This is carefully-tuned content — the edit follows
writing-skills methodology and ships only with the full eval evidence
below.
**Graded by (three layers):**
1. **Micro** (`ceremony-path-micro.py`, adapted): variant C literal text,
plus adversarially ambiguous briefs the campaign never tested (a task
that pattern-matches bounded but hides a public interface change).
Criteria: spike/bounded/arch differentiate (≥4/5 per cell); ambiguous
briefs escalate to FULL (≥4/5); arch never downgrades (5/5).
2. **Codex ceremony battery:** `cx-ceremony-{spike,bounded,arch}` on the
fix arm, 3 reps each, `score_e4.py` census. Criteria: bounded reps
show an approval turn but zero committed spec files and zero
writing-plans ritual; arch reps keep the full two-doc flow; spike reps
stay minimal.
3. **Global regression battery:** the same three ceremony scenarios on
Claude Code and Gemini (rig work: those scenarios are currently
codex-gated), 3 reps each; plus the triggering acceptance check
("Let's make a react todo list" auto-triggers brainstorming into the
full/architectural path) on all three harnesses.
### T5. Explicit model on child-issued spawns
**Evidence:** root spawns are 100% explicit-model at CLI 0.146 (dev
14/14); the live gap is depth-2 — 2/2 child-issued spawns omitted
`model`. Source recon: `model` without `reasoning_effort` resets effort
to the MODEL's default, not the parent's.
**Changes** (`skills/using-superpowers/references/codex-tools.md`):
- Every spawn you issue — including as a child — sets `model` AND
`reasoning_effort`; the effort-reset trap is named.
- Advise `[agents].default_subagent_model` and
`[agents].default_subagent_reasoning_effort` in `~/.codex/config.toml`
as the machine-level backstop for anything that slips through.
**Graded by:** `score_e1.py` (per-spawn explicit-model rate, by depth) on
the shared battery.
**Baseline:** depth-2: 0/2 explicit.
**Criterion:** every spawn at every depth carries explicit model +
effort. Pre-registered caveat: if T1 eliminates depth-2 spawns entirely,
T5 grades as root-spawn regression (hold 100%) plus doc correctness and
is recorded inconclusive-by-zero at depth-2 — the config backstop is then
the operative mechanism.
## Grading plan
- **Shared SDD battery** carries T1, T2, T5: `cx-sdd-small`, fix-branch
arm (`/tmp/sp-arm-fix`), 8 reps across both container lanes. Dev
baselines are already measured; no baseline re-runs.
- **T4 batteries** as listed above (micro + codex ceremony + global
regression).
- **Pre-registration:** every battery gets a hypothesis-log entry
(prediction, scorer, criterion) in
`superpowers-autoresearch/logs/2026-07-30-codex-efficiency-fixes.md`
BEFORE it runs. Standing rules carry over: append-only log, manual
inspection of scorer matches on fix-arm runs (non-circular
verification), no raw rollouts committed, correctness rides beside
cost in every verdict.
- **Attribution:** orthogonal scorers on one combined branch; unexpected
regressions bisect by treatment commit.
- **Budget:** shared battery ~$40, codex ceremony ~$40, global
regression ~$4080, micros ~$5 → phase 1 ≈ $150200 of the ~$850
remaining from the campaign's $1000.
## Process
- Work happens in the `codex-efficiency-fixes` worktree (branched off
`dev`); execution via subagent-driven-development from a written plan.
- Skill-text changes follow writing-skills methodology.
- Scenario/rig changes (un-gating ceremony scenarios for Claude
Code/Gemini, adversarial micro briefs) land in `superpowers-evals`
main, as authorized.
- PR-per-treatment against `dev`, each with its eval evidence and the
standard identification block; merges only on Jesse's per-PR approval.
## Phase 2 queue (baseline-first; not in this plan's tasks)
Each item requires a failing baseline before any fix ships:
1. **Dispatch routing / long-session drift** — needs a long-session
elicitation rig (fresh sessions don't reproduce the pathology at CLI
0.146). Drew's stack informs the treatment shape.
2. **Verification leases / evidence receipts** — needs the
substring-aware duplicate counter added to `score_e3.py` first
(current baseline 1/23 exact-string pairs is too weak).
3. **Remediation cap** — small-n baseline (2/3 reps) needs more reps.
4. **Cross-task-race probe redesign**`score_e5.py`'s probe is
inconclusive-by-zero by design tradeoff; needs a stronger probe.
5. **E5 D4 shell-command parser** — fix-review-scope classifier cannot
parse compound commands; scorer work, not skill work.
## Out of scope
- Adopting Drew's spinout stack (#2036/#2035) or its text.
- RoboRev, Codex token telemetry (separate codebases).
- A `close_agent` hygiene checklist (V2 has no such tool — closed as
do-not-ship in the campaign).
- Claude Code/Gemini-specific efficiency treatments beyond the T4
regression battery.

View File

@@ -0,0 +1,56 @@
# Hermes Version-Bump Wiring Design
**Date:** 2026-08-05
**Revised:** 2026-08-06
**Status:** Approved
## Goal
Keep `.hermes-plugin/plugin.yaml` in lockstep with the repository version by
registering it in `.version-bump.json` and teaching `scripts/bump-version.sh`
to process YAML without implementing a YAML parser in Bash.
## Design
- Add `{ "path": ".hermes-plugin/plugin.yaml", "field": "version" }` to
`.version-bump.json`.
- Route `.json` through the existing `jq` helpers and `.yaml` through Mike
Farah `yq` v4. The YAML key and value are passed as data, not interpolated
into the expression.
- Support only a present top-level YAML string field. Nested fields and `.yml`
are out of scope.
- Route `--check`, `--audit`, and version updates through the same small
read/write dispatcher.
- Before a version bump writes any manifest, run one read-only preflight that
validates the required tools and reads every present declared manifest
through the dispatcher. This prevents a deterministic YAML failure from
occurring after earlier JSON files have already been updated. Missing-file
behavior remains unchanged, and `--help` still works without `jq` or `yq`.
The preflight is the only reliability addition. It does not make the script
transactional or redesign its existing audit and error-status behavior.
## Tests
Three focused behavioral tests run the real script against an isolated
temporary fixture and prove:
- aligned JSON and YAML pass `--check` and `--audit`, and a bump updates both
formats;
- an actual bump with JSON declared first and a later YAML manifest whose
top-level `version` is not a string exits nonzero and leaves every manifest
byte-for-byte unchanged; and
- the real `.version-bump.json` registers the Hermes manifest.
Verification also runs shell lint and `scripts/bump-version.sh --check` against
the repository.
## Non-Goals
- No hand-written YAML parser.
- No `.yml` or nested-YAML support.
- No Hermes runtime changes.
- No rollback framework, general config-schema layer, audit/status refactor, or
exhaustive failure matrix.
- No change to the separate version-validation and JSON-expression issue found
during review.

View File

@@ -40,12 +40,72 @@ write_json_field() {
jq "$jq_path = \"$value\"" "$file" > "$tmp" && mv "$tmp" "$file"
}
require_tool() {
command -v "$1" >/dev/null 2>&1 || {
echo "error: required tool '$1' is not on PATH" >&2
return 1
}
}
read_yaml_field() {
local file="$1" field="$2"
require_tool yq || return 1
FIELD="$field" yq -er '.[strenv(FIELD)] | select(tag == "!!str")' "$file"
}
write_yaml_field() {
local file="$1" field="$2" value="$3"
FIELD="$field" VALUE="$value" \
yq -i '.[strenv(FIELD)] = strenv(VALUE)' "$file"
}
read_manifest_field() {
local file="$1"
case "$file" in
*.json) read_json_field "$@" ;;
*.yaml) read_yaml_field "$@" ;;
*)
echo "error: unsupported manifest format: $file" >&2
return 1
;;
esac
}
write_manifest_field() {
local file="$1"
case "$file" in
*.json) write_json_field "$@" ;;
*.yaml) write_yaml_field "$@" ;;
*)
echo "error: unsupported manifest format: $file" >&2
return 1
;;
esac
}
# Read the list of declared files from config.
# Outputs lines of "path<TAB>field"
declared_files() {
jq -r '.files[] | "\(.path)\t\(.field)"' "$CONFIG"
}
preflight_manifests() {
local path field fullpath
require_tool jq || return 1
while IFS=$'\t' read -r path field; do
fullpath="$REPO_ROOT/$path"
[[ -f "$fullpath" ]] || continue
if ! read_manifest_field "$fullpath" "$field" >/dev/null; then
echo "error: cannot read declared manifest: $path ($field)" >&2
return 1
fi
done < <(declared_files)
}
# Read the audit exclude patterns from config.
audit_excludes() {
jq -r '.audit.exclude[]' "$CONFIG" 2>/dev/null
@@ -68,7 +128,7 @@ cmd_check() {
continue
fi
local ver
ver=$(read_json_field "$fullpath" "$field")
ver=$(read_manifest_field "$fullpath" "$field")
printf " %-45s %s\n" "$path ($field)" "$ver"
versions+=("$ver")
done < <(declared_files)
@@ -101,7 +161,7 @@ cmd_audit() {
current_version=$(
while IFS=$'\t' read -r path field; do
local fullpath="$REPO_ROOT/$path"
[[ -f "$fullpath" ]] && read_json_field "$fullpath" "$field"
[[ -f "$fullpath" ]] && read_manifest_field "$fullpath" "$field"
done < <(declared_files) | sort | uniq -c | sort -rn | head -1 | awk '{print $2}'
)
@@ -172,6 +232,8 @@ cmd_bump() {
exit 1
fi
preflight_manifests
echo "Bumping all declared files to $new_version..."
echo ""
@@ -182,8 +244,8 @@ cmd_bump() {
continue
fi
local old_ver
old_ver=$(read_json_field "$fullpath" "$field")
write_json_field "$fullpath" "$field" "$new_version"
old_ver=$(read_manifest_field "$fullpath" "$field")
write_manifest_field "$fullpath" "$field" "$new_version"
printf " %-45s %s -> %s\n" "$path ($field)" "$old_ver" "$new_version"
done < <(declared_files)

View File

@@ -85,6 +85,7 @@ digraph brainstorming {
- Scale each section to its complexity: a few sentences if straightforward, up to 200-300 words if nuanced
- Ask after each section whether it looks right so far
- Cover: architecture, components, data flow, error handling, testing
- For a new project (or one with no configured tooling), the design presentation includes a short tooling question alongside the architecture: which of these to set up from the start — cheapest before any code exists: aggressive linting + auto-formatting (the stack's standard, e.g. ruff+format / eslint+prettier / clippy+rustfmt); unit-test infrastructure (runner, layout, a first passing fixture); end-to-end test infrastructure; fuzz or mutation testing where the stack supports it. The user's selections land in the spec's Global Constraints so every later plan and task inherits them.
- Be ready to go back and clarify if something doesn't make sense
**Design for isolation and clarity:**

View File

@@ -174,6 +174,29 @@ git worktree remove "$WORKTREE_PATH"
git worktree prune # Self-healing: clean up any stale registrations
```
**If removal is refused** (`contains modified or untracked files`): the
worktree holds files that exist nowhere else — uncommitted plans, notes,
or scratch work. Never `--force` on your own initiative. Show your human
partner what is at stake and ask:
```bash
git -C "$WORKTREE_PATH" status --porcelain -uall
```
```
Worktree removal refused — these files were never committed:
<file list>
1. Commit them to <branch> before cleanup
2. Move them into <main repo root>
3. Delete them (unrecoverable)
Which?
```
Carry out the choice, then remove the worktree.
**Otherwise:** The host environment owns this workspace — leave it in
place. If your platform provides a workspace-exit tool, use it.
@@ -196,6 +219,7 @@ place. If your platform provides a workspace-exit tool, use it.
| "'Yeah, get rid of it' counts as confirmation" | Only the typed word `discard` authorizes deletion. |
| "The PR is up, so the worktree is clutter now" | PR feedback gets fixed in that worktree. It stays until the work lands. |
| "This other worktree looks stale — I'll clean it too" | Clean up only worktrees under `.worktrees/` or `worktrees/`. Everything else belongs to the host. |
| "Removal refused — `--force` is just finishing the cleanup" | The refusal means files exist only in that worktree. `--force` destroys them permanently. Show your human partner and ask. |
| "The merged-result failure is probably flaky" | A failing merged result stops everything. Branch and worktree stay put while you investigate. |
| "The base branch is obviously main" | Confirm the fork point or ask. Merging into the wrong base is expensive to undo. |
| "The push was rejected — force-push will fix it" | A rejected push means the remote moved. Investigate; force-push only on your human partner's explicit request. |

View File

@@ -34,6 +34,15 @@ Subagent (general-purpose):
Your review is read-only on this checkout. Do not mutate the working tree, the index, HEAD, or branch state in any way. Use tools like `git show`, `git diff`, and `git log` to inspect history. If you need a working copy of a different revision, check it out into a separate temporary directory (e.g. `git worktree add /tmp/review-[SHA] [SHA]`) — never move HEAD on this checkout.
## You Do Not Dispatch Subagents
Do all of this review yourself. Never spawn a subagent to review part
of the diff, and never spawn another reviewer for a second opinion.
This process already provides every review seat the work gets; a
reviewer you spawn duplicates one of them at full cost, and its
verdict counts for nothing. If the diff feels too large for one
pass, review it in passes yourself and say so in your report.
## What to Check
**Plan alignment:**

View File

@@ -14,7 +14,21 @@ Execute plan by dispatching a fresh implementer subagent per task, a task review
**Narration:** between tool calls, narrate at most one short line — the
ledger and the tool results carry the record.
**Continuous execution:** Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are: BLOCKED status you cannot resolve, ambiguity that genuinely prevents progress, or all tasks complete. "Should I continue?" prompts and progress summaries waste their time — they asked you to execute the plan, so execute it.
**Continuous execution:** Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are the four named below, or all tasks complete. "Should I continue?" prompts and progress summaries waste their time — they asked you to execute the plan, so execute it.
**Rulings, not stalls.** A running plan does not wait on a human. Conflicts,
ambiguities, plan defects, a cap you would have asked to exceed — decide
them. The spec is the binding authority, the plan is its argument, and your
judgment settles what neither answers. Record every decision in the ledger as
`Ruling: <what you decided> — <why> — <what it costs if wrong>`, and keep
going. A wrong ruling costs rework your human partner can see and undo; a
session parked on a question costs their whole day and buys nothing.
Four things stop you, and only these: an irreversible or destructive
operation; a security-sensitive action; a side effect outside this worktree
that norms say you ask about first (a merge, a push to a shared branch, a
publish); and a plan so broken that every path forward is a guess. For those,
stop and ask.
## When to Use
@@ -57,14 +71,14 @@ digraph process {
"Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" [shape=box];
"Spec ✅ and quality approved?" [shape=diamond];
"Finding conflicts with plan text?" [shape=diamond];
"Ask human partner which governs" [shape=box];
"Rule on the conflict, ledger the ruling" [shape=box];
"Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [shape=box];
"Dispatch scoped re-review (./re-review-prompt.md)" [shape=box];
"All findings addressed?" [shape=diamond];
"R = 5?" [shape=diamond];
"Adjudicate each open finding" [shape=box];
"Any load-bearing finding?" [shape=diamond];
"STOP: report BLOCKED to human partner" [shape=box];
"Rule and continue; stop only if every path forward is a guess" [shape=box];
"Park findings in ledger with rulings" [shape=box];
"Append completion to ledger, mark todo complete" [shape=box];
}
@@ -85,8 +99,8 @@ digraph process {
"Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" -> "Spec ✅ and quality approved?";
"Spec ✅ and quality approved?" -> "Append completion to ledger, mark todo complete" [label="yes"];
"Spec ✅ and quality approved?" -> "Finding conflicts with plan text?" [label="no"];
"Finding conflicts with plan text?" -> "Ask human partner which governs" [label="yes"];
"Ask human partner which governs" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model";
"Finding conflicts with plan text?" -> "Rule on the conflict, ledger the ruling" [label="yes"];
"Rule on the conflict, ledger the ruling" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model";
"Finding conflicts with plan text?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no"];
"Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" -> "Dispatch scoped re-review (./re-review-prompt.md)";
"Dispatch scoped re-review (./re-review-prompt.md)" -> "All findings addressed?";
@@ -95,7 +109,7 @@ digraph process {
"R = 5?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no - next round"];
"R = 5?" -> "Adjudicate each open finding" [label="yes - breaker trips"];
"Adjudicate each open finding" -> "Any load-bearing finding?";
"Any load-bearing finding?" -> "STOP: report BLOCKED to human partner" [label="yes"];
"Any load-bearing finding?" -> "Rule and continue; stop only if every path forward is a guess" [label="yes"];
"Any load-bearing finding?" -> "Park findings in ledger with rulings" [label="no"];
"Park findings in ledger with rulings" -> "Append completion to ledger, mark todo complete";
"Append completion to ledger, mark todo complete" -> "More tasks remain?";
@@ -140,19 +154,32 @@ a ledger file, not only in todos.
that happens, recover from `git log`.
Read the plan once, note its context and Global Constraints, and create a
todo per task.
todo per task. If the plan names a Spec, read that too: the spec is the
authority the plan argues from, and conflicts inside the plan resolve
against it. A plan with no reachable spec gets a ledger note saying so —
rulings made without one are provisional.
Before dispatching Task 1, scan the plan once for conflicts:
Before dispatching Task 1, scan the plan once for conflicts, writing down
what you checked as you check it:
- tasks that contradict each other or the plan's Global Constraints
- anything the plan explicitly mandates that the review rubric treats as a
defect (a test that asserts nothing, verbatim duplication of a logic block)
Present everything you find to your human partner as one batched question —
each finding beside the plan text that mandates it, asking which governs —
before execution begins, not one interrupt per discovery mid-plan. If the
scan is clean, proceed without comment. The review loop remains the net for
conflicts that only emerge from implementation.
The scan's output is a table, not a verdict. One row for every pair of tasks
that share a file or an interface: the two tasks, what one produces against
what the other consumes, and what you found. One row for every task: whether
its own text agrees with itself — the tests it specifies against the code it
specifies, the files it creates against the files it later touches. "The scan
is clean" without those rows is not a scan you ran.
Write the table to the ledger. Rule on everything you find before execution
begins — each finding against the plan text that mandates it — and record
each ruling in the ledger. If the scan is clean, proceed without comment.
Rule on each conflict it surfaces — the spec is the binding authority, the
plan is its argument — record the ruling beside its row, and dispatch
Task 1. The review loop remains the net for conflicts that only emerge from
implementation.
## Model Selection
@@ -193,10 +220,29 @@ that implementer. Single-file mechanical fixes also take the cheapest tier.
## The Task Loop
**Batch small same-shape work.** When the plan lists several tasks that are
each a small, independent edit of the same kind — the same one-line fix,
constant change, or field addition repeated across files — do not dispatch
one subagent per task. Compose ONE dispatch brief listing every file and
its change, send the whole batch to a single subagent, and review its diff
as one unit. Reserve one-dispatch-per-task for work that needs its own
judgment, its own tests, or its own review surface.
Everything you paste into a dispatch prompt — and everything a subagent
prints back — stays resident in your context for the rest of the session
and is re-read on every later turn. Hand artifacts over as files.
**Waiting on dispatched subagents:** never poll a wait interface with
short timeouts, and never sit in one silent, open-ended wait either.
While you have local work — ledger updates, packaging the next review,
reading reports — keep working; child results arrive on their own.
When you are genuinely idle, wait in bounded stretches (five to ten
minutes, where your platform allows), and between stretches post one
line of status and reconcile your live children: list them, and chase
any that finished without reporting. A bounded stretch keeps nearly
all of a long wait's efficiency while guaranteeing a stuck or lost
child is noticed within minutes, not at the end of the session.
### 1. Dispatch the implementer
Record BASE (`git rev-parse HEAD`) before dispatching — the review package
@@ -223,6 +269,12 @@ and fix-round diffs need it.
later dispatches — a real session's dispatch hit 42k chars of which 99%
was pasted history. A fresh subagent needs its task, the interfaces it
touches, and the global constraints. Nothing else.
- The dispatch carries the no-subagents contract (it is in the
implementer template): the implementer never dispatches subagents —
not helpers, and never a reviewer. Review arrives from you, after the
report. In real sessions, every reviewer a worker spawned duplicated
the task review the controller dispatched anyway — a full extra
review seat per task.
- If an earlier task parked a finding in the area this task touches, carry
a pointer to that ledger entry in the dispatch.
- Record the implementer's agent identity from the dispatch result —
@@ -245,7 +297,7 @@ Implementer subagents report one of four statuses. Handle each appropriately:
1. If it's a context problem, provide more context and re-dispatch with the same model
2. If the task requires more reasoning, re-dispatch with a more capable model
3. If the task is too large, break it into smaller pieces
4. If the plan itself is wrong, escalate to the human
4. If the plan itself is wrong, rule on the correction, ledger it, and re-dispatch with the ruling carried in the dispatch
**Never** ignore an escalation or force the same model to retry without changes. If the implementer said it's stuck, something needs to change.
@@ -312,10 +364,11 @@ Before the loop starts, two routes leave it immediately:
before merge. A roll-up nobody reads is a silent discard. Minor findings
never enter the loop.
- A finding labeled plan-mandated — or any finding that conflicts with
what the plan's text requires — is the human's decision, like any plan
contradiction: present the finding and the plan text, ask which governs.
Do not dismiss the finding because the plan mandates it, and do not
dispatch a fix that contradicts the plan without asking.
what the plan's text requires — is yours to rule on: weigh the finding
against the plan text, decide with the spec as the binding authority, and
ledger the ruling before you act on it. Do not dismiss the finding because
the plan mandates it, and do not dispatch a fix that contradicts the plan
without a recorded ruling.
Everything else enters the loop. A fix round is one fix dispatch plus one
scoped re-review. Five rounds maximum per task:
@@ -360,15 +413,16 @@ dispatching. Adjudicate each open finding yourself — you hold the plan and
the cross-task context the reviewer lacks:
- **The reviewer is wrong, or the point is contestable:** park it —
`Task <N>: parked — <finding> — ruling: <why the code stands>`. The final
`Task <N>: parked — <finding> — Ruling: <why the code stands>`. The final
review sees both sides.
- **Real, but nothing downstream builds on it:** park it the same way, with
a ruling that says it's real and deferred.
- **Real and load-bearing** — a later task builds on it, or it reveals a
plan defect: STOP. Append `Task <N>: BLOCKED — <reason>` and report to
your human partner with the finding, the plan text it collides with, and
the fix history. Parking a structural failure lets every dependent task
build on it and hands the final review a problem it cannot fix either.
plan defect: rule on the smallest change that unblocks the dependent work,
ledger it as `Task <N>: Ruling: <finding> — <what you decided and why>`,
and carry it into the next task's dispatch. Parking a structural failure
silently lets every dependent task build on it. Stop only when the defect
leaves every path forward a guess.
Adjudicate only at the cap. Adjudicating earlier to end a loop is
pre-judging with a different name. Every adjudication is a ledger entry —
@@ -409,12 +463,22 @@ Then run exactly one scoped re-review of the fix wave
(`scripts/review-package PLAN_FILE FIX_BASE HEAD` over the fix range,
[re-review-prompt.md](re-review-prompt.md)).
Adjudicate any residual findings as in the task loop's breaker: park with
rulings, or stop on load-bearing ones. There is no second fix wave —
rulings, or rule on the load-bearing ones and ledger what you decided. Only
the four classes above stop you here. There is no second fix wave —
residual load-bearing findings surface to your human partner when
finishing-a-development-branch presents the options.
## Finish
Before you delete anything, collect every ledger line containing `Ruling:`
preflight rulings, parked findings, breaker adjudications, all of them — into
your final message under "Rulings I made", in the order you made them, each
with what it costs if wrong. The list is exhaustive: if the ledger holds a
ruling, the list holds it. That list is the only place the decisions you
took on your human partner's behalf reach them — they read it and rework
whatever you got wrong. A ruling that dies with the workspace was a decision
made in secret.
When the final whole-branch review is clean and its fixes are merged,
delete this plan's workspace (`rm -rf <workspace>`) — the git history is
the record now. Sibling directories belong to other plans; leave them
@@ -434,6 +498,7 @@ Use superpowers:finishing-a-development-branch.
| "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. |
| "Reviews slow the loop down" | The loop without reviews is just unverified churn. Reviews are the loop's brakes and steering. |
| "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Controllers without one have re-dispatched entire completed task sequences. |
| "The implementer spawned its own reviewer — free extra assurance" | It's a duplicate seat reviewing the same diff; the task review is the gate. A worker-spawned reviewer is a defect to flag, not rigor. |
## Example Workflow

View File

@@ -47,6 +47,18 @@ Subagent (general-purpose):
While iterating, run the focused test for what you're changing; run the
full suite once before committing, not after every edit.
## You Do Not Dispatch Subagents
Do all of this task's work yourself. Never spawn a subagent to
implement part of the task, and above all never spawn a reviewer to
check your work. Self-review (below) means reading your own diff.
Review is the controller's job: after you report, it dispatches a
fresh reviewer against your diff. A reviewer you spawn duplicates
that review at full cost, and its approval counts for nothing in
the process. If you catch yourself thinking "an independent review
would strengthen my report" — that review is already scheduled.
Report instead.
## Code Organization
You reason best about code you can hold in context at once, and your edits are more

View File

@@ -43,6 +43,15 @@ Subagent (general-purpose):
Your review is read-only on this checkout. Do not mutate the working
tree, the index, HEAD, or branch state in any way.
## You Do Not Dispatch Subagents
Do all of this review yourself. Never spawn a subagent to review part
of the diff, and never spawn another reviewer for a second opinion.
This process already provides every review seat the work gets; a
reviewer you spawn duplicates one of them at full cost, and its
verdict counts for nothing. If the diff feels too large for one
pass, review it in passes yourself and say so in your report.
## Scope
Your scope is the findings list and the fix diff. Verdict every finding.

View File

@@ -52,6 +52,15 @@ Subagent (general-purpose):
Your review is read-only on this checkout. Do not mutate the working
tree, the index, HEAD, or branch state in any way.
## You Do Not Dispatch Subagents
Do all of this review yourself. Never spawn a subagent to review part
of the diff, and never spawn another reviewer for a second opinion.
This process already provides every review seat the work gets; a
reviewer you spawn duplicates one of them at full cost, and its
verdict counts for nothing. If the diff feels too large for one
pass, review it in passes yourself and say so in your report.
## Do Not Trust the Report
Treat the implementer's report as unverified claims about the code. It
@@ -75,6 +84,13 @@ Subagent (general-purpose):
Warnings or other noise in the implementer's reported test output are
findings — test output should be pristine.
Evidence you cannot see is not evidence that doesn't exist. If the
report or its test evidence looks truncated, or you cannot locate the
results it claims, re-read the file at its stated path — and if it is
genuinely missing or garbled, report that as a gap for the controller.
Re-running the suite to regenerate what you failed to read is not
verification; illegibility of the evidence is not invalidation of it.
## Part 1: Spec Compliance
Compare the diff against What Was Requested:
@@ -86,6 +102,12 @@ Subagent (general-purpose):
- **Misunderstood:** right feature built the wrong way, wrong problem
solved
If the brief lists several files each with its own change (a batched
dispatch), check the diff against that list file by file: every listed
file must have its corresponding hunk. A listed file the diff never
touches is a Missing finding, no matter how clean the rest of the
batch looks.
If a requirement cannot be verified from this diff alone (it lives in
unchanged code or spans tasks), report it as a ⚠️ item instead of
broadening your search.

View File

@@ -56,6 +56,7 @@ If your harness appears here, read its reference file for special instructions:
- Codex: `references/codex-tools.md`
- Pi: `references/pi-tools.md`
- Antigravity: `references/antigravity-tools.md`
- Hermes Agent: `references/hermes-tools.md`
## User Instructions

View File

@@ -7,7 +7,76 @@ Add to your Codex config (`~/.codex/config.toml`):
multi_agent = true
```
This enables `spawn_agent`, `wait_agent`, and `close_agent` for skills like `dispatching-parallel-agents` and `subagent-driven-development`. When using subagent-driven-development, close reviewer subagents when their review returns. Keep each implementer subagent open until its task's review passes — the fix loop resumes the implementer — then close it. If your harness cannot send another message to a spawned agent, dispatch each fix round as a fresh implementer carrying the brief, the report file, and the findings.
This enables the multi-agent tools that skills like
`dispatching-parallel-agents` and `subagent-driven-development` use.
Which tools you get depends on the multi-agent version your model
preset selects (current presets run V2; older ones run V1). Trust your
actual tool list over any table — including this one — when they
disagree.
- **Spawning:** give children a clean context with
`spawn_agent {fork_turns: "none"}`; the default `"all"` copies your
entire transcript into the child. On Codex 0.145+, role files under
`~/.codex/agents/` attach to isolated forks via `agent_type`.
Full-history forks accept `model` and `reasoning_effort` overrides
(only `agent_type` is refused there) — isolated forks are the SDD
default for context hygiene, not because overrides require them.
- **Fix rounds:** resume the implementer with `followup_task` — it
delivers your message, triggers a turn, and transparently reloads a
child the harness evicted. Never dispatch a fresh implementer on the
theory that a spawned agent cannot be messaged again; on V2 it
always can.
- **Lifecycle:** V2 has no `close_agent`. Finished children are
evicted automatically when slots are needed; leaving them unclosed
costs nothing. Only V1 sessions have `close_agent` — there, close
reviewers when their review returns, and close each implementer
after its task's review passes.
- **Model names:** never copy a model name from a skill, table, or old
session into `spawn_agent` without checking it against your current
spawn allowlist — V2 accepts only V2-capable presets and hard-errors
on the rest.
## Waiting on children
`wait_agent` is an event subscription, not a poll: a long wait wakes
the moment a child produces mailbox activity, with the same latency as
a short one. Short-timeout polling buys nothing and costs a tool call —
and a context rebill — per poll. In measured sessions, roughly
two-thirds of all wait calls were short polls that timed out.
- While you still have local work, do not wait at all. A completed
child's final answer is pushed into your mailbox and arrives with
your next turn.
- When you are genuinely idle with children outstanding, wait in
bounded stretches: `wait_agent` with `timeout_ms` 300000-600000
(5-10 minutes). After each stretch — wake or timeout — post one
status line, run `list_agents`, and chase any child that finished
without reporting. Never stack polls shorter than five minutes; the
event subscription wakes a bounded stretch just as fast as a short
one.
- Completion mail cannot wake an idle controller (it is delivered
without triggering a turn); covering that idle window is
`wait_agent`'s only job. A stretch that times out with no activity
is your cue to reconcile, not to shorten the next stretch.
## Model routing on spawns
Every `spawn_agent` you issue — including when you are yourself a
spawned child running a fan-out — sets `model` AND `reasoning_effort`
explicitly, per the Model Selection rules of the skill you are
executing. Setting `model` alone is a trap: the child's effort
silently resets to that model's default, not to yours.
Ask your human partner to add a machine-level backstop to
`~/.codex/config.toml` so any spawn that slips through still routes to
a deliberate tier instead of silently inheriting the session's most
expensive model:
```toml
[agents]
default_subagent_model = "<a mid-tier model from your spawn allowlist>"
default_subagent_reasoning_effort = "medium"
```
## Environment Detection

View File

@@ -0,0 +1,56 @@
# Hermes Agent Tool Mapping
Skills speak in actions ("dispatch a subagent", "create a todo", "read a file"). On Hermes Agent these resolve to the tools below.
## Tools
| Action skills request | Hermes tool |
|---|---|
| Read a file | `read_file` |
| Create a new file | `write_file` |
| Edit a file (targeted patch) | `patch` |
| Run a shell command | `terminal` |
| Search file contents | `search_files` |
| Find files by name | `terminal` with `find` |
| Fetch a URL / read a webpage | `web_extract(urls=[...])` |
| Search the web | `web_search(query=...)` |
| Dispatch a subagent | `delegate_task(goal=..., context=..., toolsets=[...], role="leaf")` |
| Task tracking | `todo` tool |
| Invoke a skill | `skill_view("skill-name")` |
## Instructions file
When a skill mentions "your instructions file," on Hermes Agent this is **`AGENTS.md`** in the project directory, or **`SOUL.md`** globally at `~/.hermes/SOUL.md`.
## Invoking a skill
Hermes Agent has a `skills` toolset with `skill_view` and `skills_list` tools.
To invoke a superpowers skill, use:
```
skill_view("brainstorming")
skill_view("test-driven-development")
```
If `skill_view` cannot find a superpowers skill (it may not appear in the catalog
until the plugin fully registers it), fall back to reading the SKILL.md directly:
```
read_file(path="~/.hermes/plugins/superpowers/skills/<skill-name>/SKILL.md")
```
This fallback is the same mechanism used by other harnesses without native skill loading.
## Subagent dispatch
Use `delegate_task` to spawn isolated subagents for parallel or sequential workstreams:
```
delegate_task(goal="...", context="...", toolsets=[...], role="leaf")
```
If `delegate_task` is unavailable, do the work inline rather than inventing tool calls.
## Task tracking
Use the `todo` tool for task tracking within a session. For multi-agent task boards, use `hermes kanban` CLI if available. Treat older `TodoWrite` references as the task-tracking action.

View File

@@ -66,6 +66,9 @@ independently testable deliverable.
**Tech Stack:** [Key technologies/libraries]
**Spec:** [path to the spec/design doc this plan implements — the plan
argues from the spec, so the spec travels with it; executors read both]
## Global Constraints
[The spec's project-wide requirements — version floors, dependency limits,

View File

@@ -13,9 +13,9 @@
* Requires: graphviz (dot) installed on system
*/
const fs = require('fs');
const path = require('path');
const { execSync } = require('child_process');
import * as fs from 'fs';
import * as path from 'path';
import { execFileSync } from 'child_process';
function extractDotBlocks(markdown) {
const blocks = [];
@@ -69,7 +69,7 @@ ${bodies.join('\n\n')}
function renderToSvg(dotContent) {
try {
return execSync('dot -Tsvg', {
return execFileSync('dot', ['-Tsvg'], {
input: dotContent,
encoding: 'utf-8',
maxBuffer: 10 * 1024 * 1024
@@ -107,9 +107,10 @@ function main() {
process.exit(1);
}
// Check if dot is available
// Check if dot is available. Run the binary directly rather than probing
// with `which`, which is not a command on Windows.
try {
execSync('which dot', { encoding: 'utf-8' });
execFileSync('dot', ['-V'], { stdio: 'ignore' });
} catch {
console.error('Error: graphviz (dot) not found. Install with:');
console.error(' brew install graphviz # macOS');

0
tests/hermes/__init__.py Normal file
View File

30
tests/hermes/conftest.py Normal file
View File

@@ -0,0 +1,30 @@
from pathlib import Path
import pytest
from unittest.mock import MagicMock
@pytest.fixture
def mock_ctx():
ctx = MagicMock()
ctx._hooks = {}
ctx._skills = {}
def register_hook(event, fn):
ctx._hooks[event] = fn
def register_skill(name, path):
# Mimic hermes' real register_skill, which calls path.exists() and
# therefore breaks on a str (the bug that silently disabled the whole
# plugin, found 2026-07-23). Keeping that fidelity here means a
# regression to str paths fails these tests instead of failing
# silently inside hermes.
if not isinstance(path, Path):
raise AttributeError(
f"register_skill requires a pathlib.Path, got {type(path).__name__}"
)
ctx._skills[name] = path
ctx.register_hook.side_effect = register_hook
ctx.register_skill.side_effect = register_skill
return ctx

View File

@@ -0,0 +1,98 @@
import importlib
import os
import sys
import pytest
sys.path.insert(0, os.path.abspath(
os.path.join(os.path.dirname(__file__), "../../.hermes-plugin")
))
BOOTSTRAP_MARKER = "superpowers:using-superpowers bootstrap for hermes"
# Hermes spills injected context over 10,000 chars to a file, which breaks
# inline injection semantics. The bootstrap must stay under it with margin.
HERMES_CONTEXT_SPILL_LIMIT = 10_000
def _load():
if "__init__" in sys.modules:
del sys.modules["__init__"]
return importlib.import_module("__init__")
def _bootstrap():
m = _load()
return m._build_bootstrap(m._skills_dir())
class TestStripFrontmatter:
def test_strips_yaml_block(self):
m = _load()
content = "---\nname: foo\ndescription: bar\n---\n# Body\nContent here"
assert m._strip_frontmatter(content) == "# Body\nContent here"
def test_no_frontmatter_returns_trimmed_content(self):
m = _load()
content = "# No frontmatter\nJust content"
assert m._strip_frontmatter(content) == "# No frontmatter\nJust content"
def test_strips_surrounding_whitespace_from_body(self):
m = _load()
content = "---\nname: foo\n---\n\n\n# Body\n\n"
assert m._strip_frontmatter(content) == "# Body"
class TestSkillsDirResolution:
def test_repo_layout_resolves(self):
# The repo checkout IS the git-clone layout: .hermes-plugin/ and
# skills/ are siblings, so resolution must succeed from here.
m = _load()
skills = m._skills_dir()
assert os.path.isfile(
os.path.join(skills, "using-superpowers", "SKILL.md")
)
class TestBootstrapContent:
def test_marker_and_wrapper(self):
content = _bootstrap()
assert BOOTSTRAP_MARKER in content
assert content.startswith("<EXTREMELY_IMPORTANT>")
assert content.rstrip().endswith("</EXTREMELY_IMPORTANT>")
def test_contains_using_superpowers_body(self):
content = _bootstrap()
# A distinctive line from the skill body proves the real SKILL.md was
# embedded, not a stub.
assert "You have superpowers" in content
assert "## The Rule" in content
def test_frontmatter_stripped(self):
content = _bootstrap()
assert "---\nname:" not in content
def test_tool_mapping_sourced_from_reference_file(self):
m = _load()
content = _bootstrap()
ref = os.path.join(
m._skills_dir(), "using-superpowers", "references", "hermes-tools.md"
)
with open(ref, encoding="utf-8") as f:
ref_text = f.read().strip()
# The mapping is included verbatim from the reference file — the
# single source, not a drift-prone inline copy.
assert ref_text in content
assert "read_file" in content
def test_skill_view_guidance_present(self):
content = _bootstrap()
assert 'skill_view("superpowers:brainstorming")' in content
def test_under_hermes_context_spill_limit(self):
content = _bootstrap()
assert len(content) < HERMES_CONTEXT_SPILL_LIMIT, (
f"bootstrap is {len(content)} chars; hermes spills injected "
f"context over {HERMES_CONTEXT_SPILL_LIMIT} to a file, which "
"breaks inline injection"
)

142
tests/hermes/test_plugin.py Normal file
View File

@@ -0,0 +1,142 @@
import importlib
import importlib.util
import os
import shutil
import sys
from pathlib import Path
import pytest
# Point at the plugin directory
_PLUGIN_DIR = os.path.abspath(
os.path.join(os.path.dirname(__file__), "../../.hermes-plugin")
)
sys.path.insert(0, _PLUGIN_DIR)
BOOTSTRAP_MARKER = "superpowers:using-superpowers bootstrap for hermes"
def _load_plugin():
"""Re-import plugin module fresh."""
if "__init__" in sys.modules:
del sys.modules["__init__"]
return importlib.import_module("__init__")
def _fire_pre_llm(ctx, **kwargs):
hook = ctx._hooks["pre_llm_call"]
defaults = {
"session_id": "s1",
"user_message": "hi",
"conversation_history": [],
"is_first_turn": False,
"model": "test-model",
"platform": "cli",
}
defaults.update(kwargs)
return hook(**defaults)
class TestPluginRegistration:
def test_register_attaches_only_pre_llm_call_hook(self, mock_ctx):
plugin = _load_plugin()
plugin.register(mock_ctx)
assert list(mock_ctx._hooks.keys()) == ["pre_llm_call"]
def test_register_registers_every_stock_skill_as_path(self, mock_ctx):
plugin = _load_plugin()
plugin.register(mock_ctx)
# The conftest mock raises on non-Path (mirroring hermes' real
# register_skill), so reaching these asserts proves every
# registration passed a pathlib.Path.
assert "using-superpowers" in mock_ctx._skills
assert "brainstorming" in mock_ctx._skills
for name, path in mock_ctx._skills.items():
assert isinstance(path, Path)
assert path.name == "SKILL.md"
assert path.parent.name == name
assert path.is_file()
def test_registered_skills_match_skill_directories(self, mock_ctx):
plugin = _load_plugin()
plugin.register(mock_ctx)
skills_root = plugin._skills_dir()
expected = {
entry
for entry in os.listdir(skills_root)
if os.path.isfile(os.path.join(skills_root, entry, "SKILL.md"))
}
assert set(mock_ctx._skills.keys()) == expected
class TestBootstrapInjection:
def test_first_turn_returns_bootstrap_context(self, mock_ctx):
plugin = _load_plugin()
plugin.register(mock_ctx)
result = _fire_pre_llm(mock_ctx, is_first_turn=True)
assert isinstance(result, dict)
content = result["context"]
assert BOOTSTRAP_MARKER in content
assert content.startswith("<EXTREMELY_IMPORTANT>")
assert content.rstrip().endswith("</EXTREMELY_IMPORTANT>")
def test_later_turns_return_none(self, mock_ctx):
plugin = _load_plugin()
plugin.register(mock_ctx)
assert _fire_pre_llm(mock_ctx, is_first_turn=False) is None
assert _fire_pre_llm(mock_ctx, is_first_turn=None) is None
def test_hook_tolerates_future_kwargs(self, mock_ctx):
plugin = _load_plugin()
plugin.register(mock_ctx)
result = _fire_pre_llm(
mock_ctx, is_first_turn=True, telemetry_schema_version=3
)
assert BOOTSTRAP_MARKER in result["context"]
class TestLayoutResolution:
def _stage(self, tmp_path, layout):
"""Copy the plugin module + a minimal skills tree in the given layout."""
src_skills = Path(_PLUGIN_DIR).parent / "skills"
if layout == "clone":
plugdir = tmp_path / "superpowers" / ".hermes-plugin"
else: # flat: module at the plugin dir root, skills nested inside it
plugdir = tmp_path / "superpowers"
skills = tmp_path / "superpowers" / "skills"
plugdir.mkdir(parents=True, exist_ok=True)
shutil.copy(Path(_PLUGIN_DIR) / "__init__.py", plugdir / "__init__.py")
for skill in ("using-superpowers", "brainstorming"):
shutil.copytree(src_skills / skill, skills / skill)
return plugdir
def _load_from(self, plugdir):
spec = importlib.util.spec_from_file_location(
f"hermes_plugin_test_{plugdir.parent.name}_{plugdir.name}",
plugdir / "__init__.py",
)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod
def test_clone_layout_resolves_sibling_skills(self, tmp_path, mock_ctx):
# git-clone install: .hermes-plugin/ and skills/ are siblings.
plugdir = self._stage(tmp_path, "clone")
mod = self._load_from(plugdir)
mod.register(mock_ctx)
assert "using-superpowers" in mock_ctx._skills
def test_flat_layout_resolves_nested_skills(self, tmp_path, mock_ctx):
# flattened install: module at the plugin dir root, skills/ inside it.
plugdir = self._stage(tmp_path, "flat")
mod = self._load_from(plugdir)
mod.register(mock_ctx)
assert "using-superpowers" in mock_ctx._skills
def test_missing_skills_raises_loudly(self, tmp_path, mock_ctx):
plugdir = tmp_path / "superpowers"
plugdir.mkdir(parents=True)
shutil.copy(Path(_PLUGIN_DIR) / "__init__.py", plugdir / "__init__.py")
mod = self._load_from(plugdir)
with pytest.raises(RuntimeError, match="cannot find the skills"):
mod.register(mock_ctx)

View File

@@ -0,0 +1,76 @@
#!/usr/bin/env bash
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
REPO_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)"
SCRIPT_SOURCE="$REPO_ROOT/scripts/bump-version.sh"
TEST_ROOT="$(mktemp -d)"
cleanup() {
rm -rf "$TEST_ROOT"
}
trap cleanup EXIT
fail() {
echo "FAIL: $*" >&2
exit 1
}
make_fixture() {
local repo="$1"
local yaml_body="$2"
mkdir -p "$repo/scripts" "$repo/.hermes-plugin"
cp "$SCRIPT_SOURCE" "$repo/scripts/bump-version.sh"
cat >"$repo/.version-bump.json" <<'JSON'
{
"files": [
{ "path": "package.json", "field": "version" },
{ "path": ".hermes-plugin/plugin.yaml", "field": "version" }
],
"audit": { "exclude": [] }
}
JSON
cat >"$repo/package.json" <<'JSON'
{
"name": "fixture",
"version": "1.2.3"
}
JSON
printf '%s\n' "$yaml_body" >"$repo/.hermes-plugin/plugin.yaml"
}
happy_repo="$TEST_ROOT/happy"
make_fixture "$happy_repo" $'name: superpowers\nversion: 1.2.3'
/bin/bash "$happy_repo/scripts/bump-version.sh" --check >"$TEST_ROOT/check.out"
/bin/bash "$happy_repo/scripts/bump-version.sh" --audit >"$TEST_ROOT/audit.out"
/bin/bash "$happy_repo/scripts/bump-version.sh" 2.3.4 >"$TEST_ROOT/bump.out"
[[ "$(jq -r '.version' "$happy_repo/package.json")" == "2.3.4" ]] \
|| fail "JSON manifest was not bumped"
[[ "$(yq -r '.version' "$happy_repo/.hermes-plugin/plugin.yaml")" == "2.3.4" ]] \
|| fail "YAML manifest was not bumped"
jq -e '
any(.files[];
.path == ".hermes-plugin/plugin.yaml" and .field == "version")
' "$REPO_ROOT/.version-bump.json" >/dev/null \
|| fail "Hermes manifest is not registered"
invalid_repo="$TEST_ROOT/invalid"
make_fixture "$invalid_repo" $'name: superpowers\nversion: 123'
cp "$invalid_repo/package.json" "$TEST_ROOT/package.before"
cp "$invalid_repo/.hermes-plugin/plugin.yaml" "$TEST_ROOT/plugin.before"
if /bin/bash "$invalid_repo/scripts/bump-version.sh" 2.3.4 \
>"$TEST_ROOT/invalid.out" 2>&1; then
fail "bump accepted a non-string YAML version"
fi
cmp -s "$TEST_ROOT/package.before" "$invalid_repo/package.json" \
|| fail "JSON manifest changed before YAML validation failed"
cmp -s "$TEST_ROOT/plugin.before" "$invalid_repo/.hermes-plugin/plugin.yaml" \
|| fail "invalid YAML manifest changed"
echo "Version-bump tests passed"

View File

@@ -0,0 +1,113 @@
#!/usr/bin/env bash
set -u
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
REPO_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)"
SCRIPT_UNDER_TEST="$REPO_ROOT/skills/writing-skills/render-graphs.js"
NODE_BIN="$(command -v node)"
PASSES=0
FAILURES=0
TEST_ROOT="$(mktemp -d)"
cleanup() {
rm -rf "$TEST_ROOT"
}
trap cleanup EXIT
pass() {
echo " [PASS] $1"
PASSES=$((PASSES + 1))
}
fail() {
echo " [FAIL] $1"
FAILURES=$((FAILURES + 1))
}
assert_contains() {
local haystack="$1"
local needle="$2"
local description="$3"
if printf '%s' "$haystack" | grep -Fq -- "$needle"; then
pass "$description"
else
fail "$description"
echo " expected to find: $needle"
fi
}
assert_not_contains() {
local haystack="$1"
local needle="$2"
local description="$3"
if printf '%s' "$haystack" | grep -Fq -- "$needle"; then
fail "$description"
echo " did not expect to find: $needle"
else
pass "$description"
fi
}
fixture="$TEST_ROOT/fixture-skill"
mkdir -p "$fixture" "$TEST_ROOT/empty-path"
cat >"$fixture/SKILL.md" <<'EOF'
---
name: fixture-skill
---
# Fixture Skill
```dot
digraph fixture_graph {
start -> end;
}
```
EOF
echo "Writing-skills render-graphs tests"
missing_dot_output="$(PATH="$TEST_ROOT/empty-path" "$NODE_BIN" "$SCRIPT_UNDER_TEST" "$fixture" 2>&1)"
missing_dot_status=$?
if [[ "$missing_dot_status" -ne 0 ]]; then
pass "missing Graphviz exits non-zero"
else
fail "missing Graphviz exits non-zero"
fi
assert_contains "$missing_dot_output" "Error: graphviz (dot) not found." "missing Graphviz reports install guidance"
assert_not_contains "$missing_dot_output" "ReferenceError: require is not defined" "script runs as an ES module"
render_output="$("$NODE_BIN" "$SCRIPT_UNDER_TEST" "$fixture" 2>&1)"
render_status=$?
if [[ "$render_status" -eq 0 ]]; then
pass "fixture diagram renders"
else
fail "fixture diagram renders"
printf '%s\n' "$render_output"
fi
assert_contains "$render_output" "Found 1 diagram(s)" "reports discovered diagram"
assert_contains "$render_output" "Rendered: fixture_graph.svg" "reports rendered SVG"
if [[ -f "$fixture/diagrams/fixture_graph.svg" ]]; then
pass "writes SVG output"
else
fail "writes SVG output"
fi
if [[ -f "$fixture/diagrams/fixture_graph.svg" ]] && grep -Fq "<svg" "$fixture/diagrams/fixture_graph.svg"; then
pass "SVG output has SVG markup"
else
fail "SVG output has SVG markup"
fi
echo
echo "Results: $PASSES passed, $FAILURES failed"
if [[ "$FAILURES" -gt 0 ]]; then
exit 1
fi