mirror of
https://github.com/obra/superpowers.git
synced 2026-07-24 03:04:02 +08:00
Compare commits
4 Commits
exp/loop-e
...
fix/worktr
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
1f0e2ab912 | ||
|
|
0146173544 | ||
|
|
54d0efefd7 | ||
|
|
55d28ddf10 |
@@ -174,6 +174,29 @@ git worktree remove "$WORKTREE_PATH"
|
||||
git worktree prune # Self-healing: clean up any stale registrations
|
||||
```
|
||||
|
||||
**If removal is refused** (`contains modified or untracked files`): the
|
||||
worktree holds files that exist nowhere else — uncommitted plans, notes,
|
||||
or scratch work. Never `--force` on your own initiative. Show your human
|
||||
partner what is at stake and ask:
|
||||
|
||||
```bash
|
||||
git -C "$WORKTREE_PATH" status --porcelain
|
||||
```
|
||||
|
||||
```
|
||||
Worktree removal refused — these files were never committed:
|
||||
|
||||
<file list>
|
||||
|
||||
1. Commit them to <branch> before cleanup
|
||||
2. Move them into <main repo root>
|
||||
3. Delete them (unrecoverable)
|
||||
|
||||
Which?
|
||||
```
|
||||
|
||||
Carry out the choice, then remove the worktree.
|
||||
|
||||
**Otherwise:** The host environment owns this workspace — leave it in
|
||||
place. If your platform provides a workspace-exit tool, use it.
|
||||
|
||||
@@ -196,6 +219,7 @@ place. If your platform provides a workspace-exit tool, use it.
|
||||
| "'Yeah, get rid of it' counts as confirmation" | Only the typed word `discard` authorizes deletion. |
|
||||
| "The PR is up, so the worktree is clutter now" | PR feedback gets fixed in that worktree. It stays until the work lands. |
|
||||
| "This other worktree looks stale — I'll clean it too" | Clean up only worktrees under `.worktrees/` or `worktrees/`. Everything else belongs to the host. |
|
||||
| "Removal refused — `--force` is just finishing the cleanup" | The refusal means files exist only in that worktree. `--force` destroys them permanently. Show your human partner and ask. |
|
||||
| "The merged-result failure is probably flaky" | A failing merged result stops everything. Branch and worktree stay put while you investigate. |
|
||||
| "The base branch is obviously main" | Confirm the fork point or ask. Merging into the wrong base is expensive to undo. |
|
||||
| "The push was rejected — force-push will fix it" | A rejected push means the remote moved. Investigate; force-push only on your human partner's explicit request. |
|
||||
|
||||
@@ -211,15 +211,9 @@ and fix-round diffs need it.
|
||||
first — it is your requirements, with the exact values to use verbatim";
|
||||
(3) interfaces and decisions from earlier tasks that the brief cannot
|
||||
know; (4) your resolution of any ambiguity you noticed in the brief;
|
||||
(5) the report-file path and report contract; (6) the plan's binding
|
||||
constraint values — any Global Constraint that mandates an exact
|
||||
mechanical form (commit-message rules, naming rules, fixed literals) —
|
||||
pasted verbatim. Task-specific exact values (numbers, magic strings,
|
||||
signatures, test cases) appear only in the brief; binding constraint
|
||||
values are the one exception — they ride in EVERY dispatch, because a
|
||||
subagent that must recall a constraint from memory will reconstruct it
|
||||
from its own priors instead. Never make a subagent read the whole plan
|
||||
file.
|
||||
(5) the report-file path and report contract. Exact values (numbers,
|
||||
magic strings, signatures, test cases) appear only in the brief. Never
|
||||
make a subagent read the whole plan file.
|
||||
- **Report file:** name the implementer's report file after the brief
|
||||
(brief `…/task-N-brief.md` → report `…/task-N-report.md`) and put it in
|
||||
the dispatch prompt. The implementer writes the full report there and
|
||||
@@ -279,10 +273,6 @@ needed.
|
||||
- **Reviewer inputs:** the task reviewer gets three paths — the same brief
|
||||
file, the report file, and the review package — plus the global
|
||||
constraints that bind the task.
|
||||
- **Persist the verdict:** when the review returns, write its full text to
|
||||
`<workspace>/task-<N>-review.md` before acting on it. The file is the
|
||||
gate: completion checks for the artifact, not for your memory of a
|
||||
verdict.
|
||||
- The global-constraints block you hand the reviewer is its attention
|
||||
lens. Copy the binding requirements verbatim from the plan's Global
|
||||
Constraints section or the spec: exact values, exact formats, and the
|
||||
@@ -333,26 +323,22 @@ scoped re-review. Five rounds maximum per task:
|
||||
verbatim. Its context is intact: it knows the task, the code, and its own
|
||||
choices. If your harness cannot send another message to a live subagent,
|
||||
dispatch a fresh implementer carrying the brief path, the report-file path,
|
||||
the findings, and the binding constraint values verbatim — the report file
|
||||
is the persistent memory either way.
|
||||
and the findings — the report file is the persistent memory either way.
|
||||
|
||||
**Rounds 4-5 — dispatch a fresh implementer on a more capable model** (per
|
||||
Model Selection), with the brief path, the report-file path, the open
|
||||
findings, the binding constraint values verbatim, and this framing: "A
|
||||
prior implementer attempted this task [N] times; you own it now. Read the
|
||||
report file for what was tried." A loop that survives three resumes usually
|
||||
means the implementer cannot see its own problem — fresh eyes and a
|
||||
capability bump in one move.
|
||||
findings, and this framing: "A prior implementer attempted this task
|
||||
[N] times; you own it now. Read the report file for what was tried." A loop
|
||||
that survives three resumes usually means the implementer cannot see its
|
||||
own problem — fresh eyes and a capability bump in one move.
|
||||
|
||||
**Every round, either way:** the implementer fixes, re-runs the tests
|
||||
covering the amended code, appends its fix report to the same report file,
|
||||
and returns the short contract. Before re-dispatching the reviewer, confirm
|
||||
the fix report contains the covering tests, the command run, and the output;
|
||||
dispatch the re-review once all three are present. Name the covering test
|
||||
files in the fix message — a one-line fix does not need the whole suite.
|
||||
Append the re-review's returned text to `<workspace>/task-<N>-review.md` as
|
||||
well — the artifact accumulates every verdict, and the file's last entry is
|
||||
the one completion relies on.
|
||||
the fix report contains the covering tests, the command run, and the
|
||||
output; dispatch the re-review once all three are present. Name the
|
||||
covering test files in the fix message — a one-line fix does not need the
|
||||
whole suite.
|
||||
|
||||
**The re-review is scoped.** Run `scripts/review-package PLAN_FILE FIX_BASE HEAD`
|
||||
where FIX_BASE is the head the previous review saw, and dispatch
|
||||
@@ -397,20 +383,10 @@ message as your other bookkeeping:
|
||||
- `Task <N>: complete (commits <base7>..<head7>, review clean)`
|
||||
- `Task <N>: complete (commits <base7>..<head7>, <K> parked)` after a
|
||||
tripped breaker
|
||||
- `Task <N>: complete (commits <base7>..<head7>, deviation parked: <rule>)`
|
||||
when a landed commit violates a mechanical constraint that nothing
|
||||
downstream builds on. Record the ruling in the ledger. A green build with
|
||||
a parked, recorded deviation is complete — do not fail the task, and do
|
||||
not rewrite landed history to chase cosmetics. A load-bearing violation
|
||||
is different: that is a BLOCKED, not a deviation. This valve is not the
|
||||
breaker: it needs no exhausted fix rounds — a mechanical deviation
|
||||
discovered at completion parks here directly, with its ruling in the ledger.
|
||||
|
||||
Then mark the todo complete and move on — but only once
|
||||
`<workspace>/task-<N>-review.md` exists; a completion line without its review
|
||||
artifact is invalid, whatever you remember about the review. Never move to
|
||||
the next task while the review has open Critical/Important issues that are
|
||||
neither fixed nor parked-with-ruling at the cap.
|
||||
Then mark the todo complete and move on. Never move to the next task while
|
||||
the review has open Critical/Important issues that are neither fixed nor
|
||||
parked-with-ruling at the cap.
|
||||
|
||||
## Final Review
|
||||
|
||||
@@ -423,22 +399,17 @@ on the most capable available model (see Model Selection), using
|
||||
superpowers:requesting-code-review's
|
||||
[code-reviewer.md](../requesting-code-review/code-reviewer.md). Point it at
|
||||
the ledger's deferred-minor and parked lines so it can triage which must be
|
||||
fixed before merge. Write the returned review to `<workspace>/final-review.md`
|
||||
before dispatching any fix wave — the merge decision cites the artifact, not a
|
||||
recollection.
|
||||
fixed before merge.
|
||||
|
||||
If the final whole-branch review returns findings, dispatch ONE fix subagent
|
||||
with the complete findings list and the binding constraint values
|
||||
verbatim — not one fixer per finding.
|
||||
with the complete findings list — not one fixer per finding.
|
||||
Per-finding fixers each rebuild context and re-run suites; a real
|
||||
session's final-review fix wave cost more than all its tasks combined.
|
||||
Then run exactly one scoped re-review of the fix wave
|
||||
(`scripts/review-package PLAN_FILE FIX_BASE HEAD` over the fix range,
|
||||
[re-review-prompt.md](re-review-prompt.md)).
|
||||
Adjudicate any residual findings as in the task loop's breaker: park with
|
||||
rulings, or stop on load-bearing ones. The same valve applies to
|
||||
mechanical-constraint misses discovered at the end: non-load-bearing means
|
||||
parked with a ruling, not a failed branch. There is no second fix wave —
|
||||
rulings, or stop on load-bearing ones. There is no second fix wave —
|
||||
residual load-bearing findings surface to your human partner when
|
||||
finishing-a-development-branch presents the options.
|
||||
|
||||
|
||||
@@ -15,12 +15,6 @@ Subagent (general-purpose):
|
||||
Read your task brief first: [BRIEF_FILE]
|
||||
It contains the full task text from the plan.
|
||||
|
||||
## Binding Constraint Values
|
||||
|
||||
[CONSTRAINT_VALUES — the plan's mechanical constraints, pasted verbatim
|
||||
by the dispatcher. If a rule here mandates an exact form (commit-message
|
||||
text, naming, fixed literals), reproduce it exactly — never from memory.]
|
||||
|
||||
## Context
|
||||
|
||||
[Scene-setting: where this fits, dependencies, architectural context]
|
||||
@@ -42,12 +36,8 @@ Subagent (general-purpose):
|
||||
2. Write tests (following TDD if task says to)
|
||||
3. Verify implementation works
|
||||
4. Commit your work
|
||||
5. Immediately after each commit, verify it against the Binding Constraint
|
||||
Values above (commit-message rules, naming rules, fixed literals) while
|
||||
history is still local. A miss is cheap now — amend or forward-fix at
|
||||
once — and expensive after your work is delivered.
|
||||
6. Self-review (see below)
|
||||
7. Report back
|
||||
5. Self-review (see below)
|
||||
6. Report back
|
||||
|
||||
Work from: [directory]
|
||||
|
||||
@@ -106,10 +96,6 @@ Subagent (general-purpose):
|
||||
- Did I only build what was requested?
|
||||
- Did I follow existing patterns in the codebase?
|
||||
|
||||
**Constraints:**
|
||||
- Does every commit message satisfy the Binding Constraint Values exactly?
|
||||
- Did I reproduce mandated literals from the constraint text, not from memory?
|
||||
|
||||
**Testing:**
|
||||
- Do tests actually verify behavior (not just mock behavior)?
|
||||
- Did I follow TDD if required?
|
||||
@@ -118,12 +104,6 @@ Subagent (general-purpose):
|
||||
|
||||
If you find issues during self-review, fix them now before reporting.
|
||||
|
||||
If a constraint violation is already in a landed commit you cannot safely
|
||||
amend, forward-fix it in a new commit when possible; when it is not, report
|
||||
the deviation explicitly with Status DONE_WITH_CONCERNS — never report the
|
||||
task incomplete solely for a cosmetic miss on an otherwise green build.
|
||||
Whether the deviation parks or blocks is the dispatching controller's call.
|
||||
|
||||
## After Review Findings
|
||||
|
||||
If the task review finds issues, you will be resumed with the findings.
|
||||
@@ -156,8 +136,7 @@ Subagent (general-purpose):
|
||||
If BLOCKED or NEEDS_CONTEXT, put the specifics in the final message
|
||||
itself — the controller acts on it directly.
|
||||
|
||||
Use DONE_WITH_CONCERNS if you completed the work but have doubts about correctness, or
|
||||
when you are reporting a known deviation you could not safely fix.
|
||||
Use DONE_WITH_CONCERNS if you completed the work but have doubts about correctness.
|
||||
Use BLOCKED if you cannot complete the task. Use NEEDS_CONTEXT if you need
|
||||
information that wasn't provided. Never silently produce work you're unsure about.
|
||||
```
|
||||
|
||||
@@ -18,9 +18,18 @@ echo "🔍 Searching for test that creates: $POLLUTION_CHECK"
|
||||
echo "Test pattern: $TEST_PATTERN"
|
||||
echo ""
|
||||
|
||||
# Get list of test files
|
||||
TEST_FILES=$(find . -path "$TEST_PATTERN" | sort)
|
||||
TOTAL=$(echo "$TEST_FILES" | wc -l | tr -d ' ')
|
||||
# Get list of test files (find . emits ./-prefixed paths, so accept the
|
||||
# pattern written with or without a leading ./)
|
||||
TEST_PATTERN="${TEST_PATTERN#./}"
|
||||
# find -path can't match '**/' against zero directory levels, so a pattern
|
||||
# like src/**/*.test.ts would skip src/top.test.ts; also try the pattern
|
||||
# with '**/' collapsed to cover files directly under the base directory.
|
||||
TEST_FILES=$(find . \( -path "./$TEST_PATTERN" -o -path "./${TEST_PATTERN//\*\*\//}" \) | sort -u)
|
||||
if [ -z "$TEST_FILES" ]; then
|
||||
TOTAL=0
|
||||
else
|
||||
TOTAL=$(printf '%s\n' "$TEST_FILES" | wc -l | tr -d ' ')
|
||||
fi
|
||||
|
||||
echo "Found $TOTAL test files"
|
||||
echo ""
|
||||
|
||||
@@ -4,7 +4,7 @@ Skills speak in actions ("dispatch a subagent", "create a todo", "read a file").
|
||||
|
||||
| Action skills request | Antigravity CLI equivalent |
|
||||
|----------------------|----------------------|
|
||||
| Dispatch a subagent (`Subagent (general-purpose):` template) | `invoke_subagent` with a built-in `TypeName` — `self` for full-capability work, `research` for read-only (see [Subagent support](#subagent-support)) |
|
||||
| Dispatch a subagent (`Subagent (general-purpose):` template) | `invoke_subagent` with a built-in `TypeName` — `self` for full-capability work, `research` for read-only |
|
||||
| Task tracking ("create a todo", "mark complete") | a **task artifact** — `write_to_file` with `IsArtifact: true` and `ArtifactType: "task"` (see [Task tracking](#task-tracking)). **Not** `manage_task`, which manages background processes. |
|
||||
|
||||
## Task tracking
|
||||
|
||||
@@ -71,11 +71,7 @@ independently testable deliverable.
|
||||
[The spec's project-wide requirements — version floors, dependency limits,
|
||||
naming and copy rules, platform requirements — one line each, with exact
|
||||
values copied verbatim from the spec. Every task's requirements implicitly
|
||||
include this section. Three things may never enter it: cosmetic absolutes
|
||||
on every commit (a fixed trailer or byline the work does not need), your
|
||||
own identity or model name promoted into a rule, and environment
|
||||
constraints (versions, platforms, paths) you have not verified against the
|
||||
environment the plan will execute in.]
|
||||
include this section.]
|
||||
|
||||
---
|
||||
```
|
||||
@@ -129,8 +125,6 @@ git commit -m "feat: add specific feature"
|
||||
```
|
||||
````
|
||||
|
||||
Commit messages describe the change. Never mandate session boilerplate — trailers, bylines, model names — as a per-commit rule; what your session stamps on its commits is not a requirement of the work.
|
||||
|
||||
## No Placeholders
|
||||
|
||||
Every step must contain the actual content an engineer needs. These are **plan failures** — never write them:
|
||||
@@ -151,8 +145,6 @@ After writing the complete plan, look at the spec with fresh eyes and check the
|
||||
|
||||
**3. Type consistency:** Do the types, method signatures, and property names you used in later tasks match what you defined in earlier tasks? A function called `clearLayers()` in Task 3 but `clearFullLayers()` in Task 7 is a bug.
|
||||
|
||||
**4. Constraint hygiene:** Does any Global Constraint mandate a per-commit cosmetic absolute, name the authoring model or session, or assert an environment fact (version floor, platform, path) you did not verify? Cut or verify it.
|
||||
|
||||
If you find issues, fix them inline. No need to re-review — just fix and move on. If you find a spec requirement with no task, add the task.
|
||||
|
||||
## Execution Handoff
|
||||
|
||||
90
tests/systematic-debugging/test-find-polluter.sh
Executable file
90
tests/systematic-debugging/test-find-polluter.sh
Executable file
@@ -0,0 +1,90 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
REPO_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)"
|
||||
SCRIPT_UNDER_TEST="$REPO_ROOT/skills/systematic-debugging/find-polluter.sh"
|
||||
|
||||
FAILURES=0
|
||||
TEST_ROOT="$(mktemp -d)"
|
||||
|
||||
cleanup() {
|
||||
rm -rf "$TEST_ROOT"
|
||||
}
|
||||
trap cleanup EXIT
|
||||
|
||||
pass() {
|
||||
echo " [PASS] $1"
|
||||
}
|
||||
|
||||
fail() {
|
||||
echo " [FAIL] $1"
|
||||
FAILURES=$((FAILURES + 1))
|
||||
}
|
||||
|
||||
assert_contains() {
|
||||
local haystack="$1"
|
||||
local needle="$2"
|
||||
local description="$3"
|
||||
|
||||
if printf '%s' "$haystack" | grep -Fq -- "$needle"; then
|
||||
pass "$description"
|
||||
else
|
||||
fail "$description (expected output to contain: $needle)"
|
||||
fi
|
||||
}
|
||||
|
||||
# Toy project: one top-level test, one nested test. A stubbed `npm` on PATH
|
||||
# creates the pollution marker whenever any test runs, so the first test file
|
||||
# executed is always identified as the polluter.
|
||||
setup_project() {
|
||||
PROJECT="$TEST_ROOT/project"
|
||||
rm -rf "$PROJECT"
|
||||
mkdir -p "$PROJECT/src/feature" "$PROJECT/bin"
|
||||
echo "test('top')" > "$PROJECT/src/top.test.ts"
|
||||
echo "test('nested')" > "$PROJECT/src/feature/nested.test.ts"
|
||||
cat > "$PROJECT/bin/npm" <<'EOF'
|
||||
#!/usr/bin/env bash
|
||||
touch pollution.marker
|
||||
EOF
|
||||
chmod +x "$PROJECT/bin/npm"
|
||||
}
|
||||
|
||||
# run_polluter <pattern> — runs the script in the toy project with the stub
|
||||
# npm first on PATH; captures combined output, never aborts on exit code.
|
||||
run_polluter() {
|
||||
local pattern="$1"
|
||||
rm -f "$PROJECT/pollution.marker"
|
||||
(
|
||||
cd "$PROJECT"
|
||||
PATH="$PROJECT/bin:$PATH" "$SCRIPT_UNDER_TEST" 'pollution.marker' "$pattern" 2>&1
|
||||
) || true
|
||||
}
|
||||
|
||||
echo "Test: documented pattern finds nested test files (issue #2008)"
|
||||
setup_project
|
||||
OUTPUT="$(run_polluter 'src/**/*.test.ts')"
|
||||
assert_contains "$OUTPUT" "FOUND POLLUTER" "documented pattern runs tests and detects pollution"
|
||||
|
||||
echo "Test: documented pattern also finds top-level test files"
|
||||
setup_project
|
||||
OUTPUT="$(run_polluter 'src/**/*.test.ts')"
|
||||
assert_contains "$OUTPUT" "Found 2 test files" "src/**/*.test.ts matches src/top.test.ts and src/feature/nested.test.ts"
|
||||
|
||||
echo "Test: ./-prefixed pattern matches the same files"
|
||||
setup_project
|
||||
OUTPUT="$(run_polluter './src/**/*.test.ts')"
|
||||
assert_contains "$OUTPUT" "Found 2 test files" "leading ./ on the pattern is accepted"
|
||||
|
||||
echo "Test: non-matching pattern reports an honest zero"
|
||||
setup_project
|
||||
OUTPUT="$(run_polluter 'nomatch/**/*.test.ts')"
|
||||
assert_contains "$OUTPUT" "Found 0 test files" "empty result counts as 0, not 1"
|
||||
assert_contains "$OUTPUT" "No polluter found" "empty result exits via the clean path"
|
||||
|
||||
echo ""
|
||||
if [ "$FAILURES" -gt 0 ]; then
|
||||
echo "$FAILURES test(s) failed"
|
||||
exit 1
|
||||
fi
|
||||
echo "All tests passed"
|
||||
Reference in New Issue
Block a user