Compare commits

..

4 Commits

Author SHA1 Message Date
Jesse Vincent
1f0e2ab912 fix(finishing): check in with human partner when worktree removal hits untracked files
git worktree remove refuses when the tree holds modified or untracked
files, and the skill gave no guidance for that refusal — the natural
agent response was --force, permanently destroying files that exist
nowhere else (uncommitted plans, notes, scratch work). Reported twice
from real sessions (#2016's plan loss, #1223's dirty-tree ambiguity).

Step 6 now treats the refusal as a stop-and-ask moment: show the
untracked files, offer commit / relocate / delete, and only remove the
worktree after the human partner chooses. Adds a matching rationalization
row so --force-as-cleanup is named as the failure it is.
2026-07-23 11:55:24 -07:00
Jesse Vincent
0146173544 fix(systematic-debugging): find-polluter accepts ./-prefixed patterns and matches top-level tests
Follow-up to #2011 (which fixed the ./-prefix mismatch for the documented
pattern form): strip a leading ./ from the caller's pattern instead of
double-prefixing it into a never-matching ././ form, and also match the
pattern with '**/' collapsed, since find -path cannot match '**/' against
zero directory levels and silently skipped files directly under the base
directory (src/top.test.ts vs src/**/*.test.ts).

Adds a deterministic test suite for the script with a stubbed npm.
2026-07-23 10:54:50 -07:00
dev_Hakaze
54d0efefd7 fix(systematic-debugging): match find -path ./ prefix in find-polluter.sh (#2011)
find . emits ./-prefixed paths, so -path "src/**/*.test.ts" matched
nothing; wc -l on empty stdin then lied as "Found 1". Fixes #2008.

Co-authored-by: arimu1 <19286898+arimu1@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-23 10:53:39 -07:00
Mark Rada
55d28ddf10 docs(using-superpowers): drop dangling subagent-support anchor (#2010)
The prune in e7ddc25e removed the `## Subagent support` section from
antigravity-tools.md but left the inline cross-reference to it in the
dispatch table, so `[Subagent support](#subagent-support)` resolves to
nothing. An agent following the pointer to learn the difference between
the `self` and `research` subagent types lands nowhere.

Drop the dangling parenthetical. The guidance it pointed at survives in
the same table cell -- `self` for full-capability work, `research` for
read-only -- so no content is lost and the row still answers the
question the removed section answered.

gemini-tools.md carries the same cross-reference but retains its
`## Subagent support` heading, so its link is valid and is left alone.
2026-07-23 10:47:44 -07:00
7 changed files with 151 additions and 105 deletions

View File

@@ -174,6 +174,29 @@ git worktree remove "$WORKTREE_PATH"
git worktree prune # Self-healing: clean up any stale registrations
```
**If removal is refused** (`contains modified or untracked files`): the
worktree holds files that exist nowhere else — uncommitted plans, notes,
or scratch work. Never `--force` on your own initiative. Show your human
partner what is at stake and ask:
```bash
git -C "$WORKTREE_PATH" status --porcelain
```
```
Worktree removal refused — these files were never committed:
<file list>
1. Commit them to <branch> before cleanup
2. Move them into <main repo root>
3. Delete them (unrecoverable)
Which?
```
Carry out the choice, then remove the worktree.
**Otherwise:** The host environment owns this workspace — leave it in
place. If your platform provides a workspace-exit tool, use it.
@@ -196,6 +219,7 @@ place. If your platform provides a workspace-exit tool, use it.
| "'Yeah, get rid of it' counts as confirmation" | Only the typed word `discard` authorizes deletion. |
| "The PR is up, so the worktree is clutter now" | PR feedback gets fixed in that worktree. It stays until the work lands. |
| "This other worktree looks stale — I'll clean it too" | Clean up only worktrees under `.worktrees/` or `worktrees/`. Everything else belongs to the host. |
| "Removal refused — `--force` is just finishing the cleanup" | The refusal means files exist only in that worktree. `--force` destroys them permanently. Show your human partner and ask. |
| "The merged-result failure is probably flaky" | A failing merged result stops everything. Branch and worktree stay put while you investigate. |
| "The base branch is obviously main" | Confirm the fork point or ask. Merging into the wrong base is expensive to undo. |
| "The push was rejected — force-push will fix it" | A rejected push means the remote moved. Investigate; force-push only on your human partner's explicit request. |

View File

@@ -133,15 +133,6 @@ a ledger file, not only in todos.
plan's progress: leave it in place and start your own, fresh.
- Create the ledger with its identity as the first line:
`# SDD ledger — plan: <plan file path>`.
- During that same Setup read, copy the plan's Global Constraints section
verbatim into `<workspace>/constraints.md`. Every dispatch's binding
constraint values paste from that file — the plan itself stays closed
after Setup, even across compaction.
- The workspace is never committed to the project repo. Do not `git add`
anything under it, and never write a commit whose purpose is to record,
correct, or tidy a workspace artifact — reports and ledgers are session
records, not deliverables. If the repo lacks a `.gitignore` entry for
`.superpowers/`, leave the directory untracked rather than committing it.
- The ledger is your recovery map: the commits it names exist in git even
when your context no longer remembers creating them. After compaction,
trust the ledger and `git log` over your own recollection.
@@ -149,11 +140,7 @@ a ledger file, not only in todos.
that happens, recover from `git log`.
Read the plan once, note its context and Global Constraints, and create a
todo per task. That is the plan's one full read for the whole session:
after Setup, the ledger and `scripts/task-brief` extracts are your working
memory — re-reading the plan or spec late in the run (to "double-check"
completion, to rebuild the final-review dispatch) re-buys context you
already paid for and is forbidden.
todo per task.
Before dispatching Task 1, scan the plan once for conflicts:
@@ -208,10 +195,7 @@ that implementer. Single-file mechanical fixes also take the cheapest tier.
Everything you paste into a dispatch prompt — and everything a subagent
prints back — stays resident in your context for the rest of the session
and is re-read on every later turn. Hand artifacts over as files. The same
tax applies to your own words: checkpoint in one short line, keep
bookkeeping in the ledger file, and never paste back into the conversation
what a file already holds.
and is re-read on every later turn. Hand artifacts over as files.
### 1. Dispatch the implementer
@@ -227,15 +211,9 @@ and fix-round diffs need it.
first — it is your requirements, with the exact values to use verbatim";
(3) interfaces and decisions from earlier tasks that the brief cannot
know; (4) your resolution of any ambiguity you noticed in the brief;
(5) the report-file path and report contract; (6) the plan's binding
constraint values — any Global Constraint that mandates an exact
mechanical form (commit-message rules, naming rules, fixed literals) —
pasted verbatim. Task-specific exact values (numbers, magic strings,
signatures, test cases) appear only in the brief; binding constraint
values are the one exception — they ride in EVERY dispatch, because a
subagent that must recall a constraint from memory will reconstruct it
from its own priors instead. Never make a subagent read the whole plan
file.
(5) the report-file path and report contract. Exact values (numbers,
magic strings, signatures, test cases) appear only in the brief. Never
make a subagent read the whole plan file.
- **Report file:** name the implementer's report file after the brief
(brief `…/task-N-brief.md` → report `…/task-N-report.md`) and put it in
the dispatch prompt. The implementer writes the full report there and
@@ -295,10 +273,6 @@ needed.
- **Reviewer inputs:** the task reviewer gets three paths — the same brief
file, the report file, and the review package — plus the global
constraints that bind the task.
- **Persist the verdict:** when the review returns, write its full text to
`<workspace>/task-<N>-review.md` before acting on it. The file is the
gate: completion checks for the artifact, not for your memory of a
verdict.
- The global-constraints block you hand the reviewer is its attention
lens. Copy the binding requirements verbatim from the plan's Global
Constraints section or the spec: exact values, exact formats, and the
@@ -349,26 +323,22 @@ scoped re-review. Five rounds maximum per task:
verbatim. Its context is intact: it knows the task, the code, and its own
choices. If your harness cannot send another message to a live subagent,
dispatch a fresh implementer carrying the brief path, the report-file path,
the findings, and the binding constraint values verbatim — the report file
is the persistent memory either way.
and the findings — the report file is the persistent memory either way.
**Rounds 4-5 — dispatch a fresh implementer on a more capable model** (per
Model Selection), with the brief path, the report-file path, the open
findings, the binding constraint values verbatim, and this framing: "A
prior implementer attempted this task [N] times; you own it now. Read the
report file for what was tried." A loop that survives three resumes usually
means the implementer cannot see its own problem — fresh eyes and a
capability bump in one move.
findings, and this framing: "A prior implementer attempted this task
[N] times; you own it now. Read the report file for what was tried." A loop
that survives three resumes usually means the implementer cannot see its
own problem — fresh eyes and a capability bump in one move.
**Every round, either way:** the implementer fixes, re-runs the tests
covering the amended code, appends its fix report to the same report file,
and returns the short contract. Before re-dispatching the reviewer, confirm
the fix report contains the covering tests, the command run, and the output;
dispatch the re-review once all three are present. Name the covering test
files in the fix message — a one-line fix does not need the whole suite.
Append the re-review's returned text to `<workspace>/task-<N>-review.md` as
well — the artifact accumulates every verdict, and the file's last entry is
the one completion relies on.
the fix report contains the covering tests, the command run, and the
output; dispatch the re-review once all three are present. Name the
covering test files in the fix message — a one-line fix does not need the
whole suite.
**The re-review is scoped.** Run `scripts/review-package PLAN_FILE FIX_BASE HEAD`
where FIX_BASE is the head the previous review saw, and dispatch
@@ -413,20 +383,10 @@ message as your other bookkeeping:
- `Task <N>: complete (commits <base7>..<head7>, review clean)`
- `Task <N>: complete (commits <base7>..<head7>, <K> parked)` after a
tripped breaker
- `Task <N>: complete (commits <base7>..<head7>, deviation parked: <rule>)`
when a landed commit violates a mechanical constraint that nothing
downstream builds on. Record the ruling in the ledger. A green build with
a parked, recorded deviation is complete — do not fail the task, and do
not rewrite landed history to chase cosmetics. A load-bearing violation
is different: that is a BLOCKED, not a deviation. This valve is not the
breaker: it needs no exhausted fix rounds — a mechanical deviation
discovered at completion parks here directly, with its ruling in the ledger.
Then mark the todo complete and move on — but only once
`<workspace>/task-<N>-review.md` exists; a completion line without its review
artifact is invalid, whatever you remember about the review. Never move to
the next task while the review has open Critical/Important issues that are
neither fixed nor parked-with-ruling at the cap.
Then mark the todo complete and move on. Never move to the next task while
the review has open Critical/Important issues that are neither fixed nor
parked-with-ruling at the cap.
## Final Review
@@ -439,25 +399,17 @@ on the most capable available model (see Model Selection), using
superpowers:requesting-code-review's
[code-reviewer.md](../requesting-code-review/code-reviewer.md). Point it at
the ledger's deferred-minor and parked lines so it can triage which must be
fixed before merge. Build that dispatch from the ledger alone — the
completion lines, parked rulings, and deferred minors are the whole-run
summary; do not re-read the plan, the spec, or per-task reports to
reconstruct what the ledger already states. Write the returned review to
`<workspace>/final-review.md` before dispatching any fix wave — the merge
decision cites the artifact, not a recollection.
fixed before merge.
If the final whole-branch review returns findings, dispatch ONE fix subagent
with the complete findings list and the binding constraint values
verbatim — not one fixer per finding.
with the complete findings list — not one fixer per finding.
Per-finding fixers each rebuild context and re-run suites; a real
session's final-review fix wave cost more than all its tasks combined.
Then run exactly one scoped re-review of the fix wave
(`scripts/review-package PLAN_FILE FIX_BASE HEAD` over the fix range,
[re-review-prompt.md](re-review-prompt.md)).
Adjudicate any residual findings as in the task loop's breaker: park with
rulings, or stop on load-bearing ones. The same valve applies to
mechanical-constraint misses discovered at the end: non-load-bearing means
parked with a ruling, not a failed branch. There is no second fix wave —
rulings, or stop on load-bearing ones. There is no second fix wave —
residual load-bearing findings surface to your human partner when
finishing-a-development-branch presents the options.

View File

@@ -15,12 +15,6 @@ Subagent (general-purpose):
Read your task brief first: [BRIEF_FILE]
It contains the full task text from the plan.
## Binding Constraint Values
[CONSTRAINT_VALUES — the plan's mechanical constraints, pasted verbatim
by the dispatcher. If a rule here mandates an exact form (commit-message
text, naming, fixed literals), reproduce it exactly — never from memory.]
## Context
[Scene-setting: where this fits, dependencies, architectural context]
@@ -42,12 +36,8 @@ Subagent (general-purpose):
2. Write tests (following TDD if task says to)
3. Verify implementation works
4. Commit your work
5. Immediately after each commit, verify it against the Binding Constraint
Values above (commit-message rules, naming rules, fixed literals) while
history is still local. A miss is cheap now — amend or forward-fix at
once — and expensive after your work is delivered.
6. Self-review (see below)
7. Report back
5. Self-review (see below)
6. Report back
Work from: [directory]
@@ -106,10 +96,6 @@ Subagent (general-purpose):
- Did I only build what was requested?
- Did I follow existing patterns in the codebase?
**Constraints:**
- Does every commit message satisfy the Binding Constraint Values exactly?
- Did I reproduce mandated literals from the constraint text, not from memory?
**Testing:**
- Do tests actually verify behavior (not just mock behavior)?
- Did I follow TDD if required?
@@ -118,12 +104,6 @@ Subagent (general-purpose):
If you find issues during self-review, fix them now before reporting.
If a constraint violation is already in a landed commit you cannot safely
amend, forward-fix it in a new commit when possible; when it is not, report
the deviation explicitly with Status DONE_WITH_CONCERNS — never report the
task incomplete solely for a cosmetic miss on an otherwise green build.
Whether the deviation parks or blocks is the dispatching controller's call.
## After Review Findings
If the task review finds issues, you will be resumed with the findings.
@@ -156,8 +136,7 @@ Subagent (general-purpose):
If BLOCKED or NEEDS_CONTEXT, put the specifics in the final message
itself — the controller acts on it directly.
Use DONE_WITH_CONCERNS if you completed the work but have doubts about correctness, or
when you are reporting a known deviation you could not safely fix.
Use DONE_WITH_CONCERNS if you completed the work but have doubts about correctness.
Use BLOCKED if you cannot complete the task. Use NEEDS_CONTEXT if you need
information that wasn't provided. Never silently produce work you're unsure about.
```

View File

@@ -18,9 +18,18 @@ echo "🔍 Searching for test that creates: $POLLUTION_CHECK"
echo "Test pattern: $TEST_PATTERN"
echo ""
# Get list of test files
TEST_FILES=$(find . -path "$TEST_PATTERN" | sort)
TOTAL=$(echo "$TEST_FILES" | wc -l | tr -d ' ')
# Get list of test files (find . emits ./-prefixed paths, so accept the
# pattern written with or without a leading ./)
TEST_PATTERN="${TEST_PATTERN#./}"
# find -path can't match '**/' against zero directory levels, so a pattern
# like src/**/*.test.ts would skip src/top.test.ts; also try the pattern
# with '**/' collapsed to cover files directly under the base directory.
TEST_FILES=$(find . \( -path "./$TEST_PATTERN" -o -path "./${TEST_PATTERN//\*\*\//}" \) | sort -u)
if [ -z "$TEST_FILES" ]; then
TOTAL=0
else
TOTAL=$(printf '%s\n' "$TEST_FILES" | wc -l | tr -d ' ')
fi
echo "Found $TOTAL test files"
echo ""

View File

@@ -4,7 +4,7 @@ Skills speak in actions ("dispatch a subagent", "create a todo", "read a file").
| Action skills request | Antigravity CLI equivalent |
|----------------------|----------------------|
| Dispatch a subagent (`Subagent (general-purpose):` template) | `invoke_subagent` with a built-in `TypeName``self` for full-capability work, `research` for read-only (see [Subagent support](#subagent-support)) |
| Dispatch a subagent (`Subagent (general-purpose):` template) | `invoke_subagent` with a built-in `TypeName``self` for full-capability work, `research` for read-only |
| Task tracking ("create a todo", "mark complete") | a **task artifact**`write_to_file` with `IsArtifact: true` and `ArtifactType: "task"` (see [Task tracking](#task-tracking)). **Not** `manage_task`, which manages background processes. |
## Task tracking

View File

@@ -71,11 +71,7 @@ independently testable deliverable.
[The spec's project-wide requirements — version floors, dependency limits,
naming and copy rules, platform requirements — one line each, with exact
values copied verbatim from the spec. Every task's requirements implicitly
include this section. Three things may never enter it: cosmetic absolutes
on every commit (a fixed trailer or byline the work does not need), your
own identity or model name promoted into a rule, and environment
constraints (versions, platforms, paths) you have not verified against the
environment the plan will execute in.]
include this section.]
---
```
@@ -129,8 +125,6 @@ git commit -m "feat: add specific feature"
```
````
Commit messages describe the change. Never mandate session boilerplate — trailers, bylines, model names — as a per-commit rule; what your session stamps on its commits is not a requirement of the work.
## No Placeholders
Every step must contain the actual content an engineer needs. These are **plan failures** — never write them:
@@ -151,8 +145,6 @@ After writing the complete plan, look at the spec with fresh eyes and check the
**3. Type consistency:** Do the types, method signatures, and property names you used in later tasks match what you defined in earlier tasks? A function called `clearLayers()` in Task 3 but `clearFullLayers()` in Task 7 is a bug.
**4. Constraint hygiene:** Does any Global Constraint mandate a per-commit cosmetic absolute, name the authoring model or session, or assert an environment fact (version floor, platform, path) you did not verify? Cut or verify it.
If you find issues, fix them inline. No need to re-review — just fix and move on. If you find a spec requirement with no task, add the task.
## Execution Handoff

View File

@@ -0,0 +1,90 @@
#!/usr/bin/env bash
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
REPO_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)"
SCRIPT_UNDER_TEST="$REPO_ROOT/skills/systematic-debugging/find-polluter.sh"
FAILURES=0
TEST_ROOT="$(mktemp -d)"
cleanup() {
rm -rf "$TEST_ROOT"
}
trap cleanup EXIT
pass() {
echo " [PASS] $1"
}
fail() {
echo " [FAIL] $1"
FAILURES=$((FAILURES + 1))
}
assert_contains() {
local haystack="$1"
local needle="$2"
local description="$3"
if printf '%s' "$haystack" | grep -Fq -- "$needle"; then
pass "$description"
else
fail "$description (expected output to contain: $needle)"
fi
}
# Toy project: one top-level test, one nested test. A stubbed `npm` on PATH
# creates the pollution marker whenever any test runs, so the first test file
# executed is always identified as the polluter.
setup_project() {
PROJECT="$TEST_ROOT/project"
rm -rf "$PROJECT"
mkdir -p "$PROJECT/src/feature" "$PROJECT/bin"
echo "test('top')" > "$PROJECT/src/top.test.ts"
echo "test('nested')" > "$PROJECT/src/feature/nested.test.ts"
cat > "$PROJECT/bin/npm" <<'EOF'
#!/usr/bin/env bash
touch pollution.marker
EOF
chmod +x "$PROJECT/bin/npm"
}
# run_polluter <pattern> — runs the script in the toy project with the stub
# npm first on PATH; captures combined output, never aborts on exit code.
run_polluter() {
local pattern="$1"
rm -f "$PROJECT/pollution.marker"
(
cd "$PROJECT"
PATH="$PROJECT/bin:$PATH" "$SCRIPT_UNDER_TEST" 'pollution.marker' "$pattern" 2>&1
) || true
}
echo "Test: documented pattern finds nested test files (issue #2008)"
setup_project
OUTPUT="$(run_polluter 'src/**/*.test.ts')"
assert_contains "$OUTPUT" "FOUND POLLUTER" "documented pattern runs tests and detects pollution"
echo "Test: documented pattern also finds top-level test files"
setup_project
OUTPUT="$(run_polluter 'src/**/*.test.ts')"
assert_contains "$OUTPUT" "Found 2 test files" "src/**/*.test.ts matches src/top.test.ts and src/feature/nested.test.ts"
echo "Test: ./-prefixed pattern matches the same files"
setup_project
OUTPUT="$(run_polluter './src/**/*.test.ts')"
assert_contains "$OUTPUT" "Found 2 test files" "leading ./ on the pattern is accepted"
echo "Test: non-matching pattern reports an honest zero"
setup_project
OUTPUT="$(run_polluter 'nomatch/**/*.test.ts')"
assert_contains "$OUTPUT" "Found 0 test files" "empty result counts as 0, not 1"
assert_contains "$OUTPUT" "No polluter found" "empty result exits via the clean path"
echo ""
if [ "$FAILURES" -gt 0 ]; then
echo "$FAILURES test(s) failed"
exit 1
fi
echo "All tests passed"