mirror of
https://github.com/obra/superpowers.git
synced 2026-07-23 10:44:01 +08:00
feat(sdd): evidence-locked gate claims in the implementer report
Responds to maintainer review asking for better implementer self-reflection. The report format already demands command+output for TDD evidence and fix-round covering tests, but the full-suite claim asked only for prose — and that is the claim that rotted in the field: three consecutive fix rounds shipped a "full suite passes" claim the fix reviewer found unreproducible, each time because the suite had run before later edits made the result stale. Two changes, both at the moment the report is written: every claimed gate needs its exact command and the tail of fresh output — fresh meaning after the final edit, otherwise rerun or report the gate as unverified — and a missing pasted output is itself a defect for the reviewer to flag, which puts enforcement at the consumption side the same way the dispatch hints do. Fix reports claiming a full-suite pass need a fresh run, not the pre-findings one. This is verification-before-completion's gate function relocated into the one prompt subagents actually receive. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -110,14 +110,21 @@ Subagent (general-purpose):
|
|||||||
Fix them, re-run the tests that cover the amended code, and append a fix
|
Fix them, re-run the tests that cover the amended code, and append a fix
|
||||||
report to your report file: what you changed, the covering tests you
|
report to your report file: what you changed, the covering tests you
|
||||||
ran, the command, and the output. Reviewers will not re-run tests for
|
ran, the command, and the output. Reviewers will not re-run tests for
|
||||||
you — your report is the test evidence. Then reply with the same short
|
you — your report is the test evidence. If your fix report claims a
|
||||||
status contract as your first report.
|
full-suite pass, that claim needs a fresh run after your last edit —
|
||||||
|
a suite run from before the findings arrived no longer counts. Then
|
||||||
|
reply with the same short status contract as your first report.
|
||||||
|
|
||||||
## Report Format
|
## Report Format
|
||||||
|
|
||||||
Write your full report to [REPORT_FILE]:
|
Write your full report to [REPORT_FILE]:
|
||||||
- What you implemented (or what you attempted, if blocked)
|
- What you implemented (or what you attempted, if blocked)
|
||||||
- What you tested and test results
|
- What you tested, and for every gate you claim — focused tests, full
|
||||||
|
suite, lint, build — the exact command and the tail of its fresh
|
||||||
|
output. Fresh means run after your final edit: if you edited anything
|
||||||
|
since your last full-suite run, that run is stale — rerun it or
|
||||||
|
report the suite as unverified. A gate claim without pasted fresh
|
||||||
|
output is itself a defect for the reviewer to flag.
|
||||||
- **TDD Evidence** (if TDD was required for this task):
|
- **TDD Evidence** (if TDD was required for this task):
|
||||||
- RED: command run, relevant failing output before implementation, and why the failure was expected
|
- RED: command run, relevant failing output before implementation, and why the failure was expected
|
||||||
- GREEN: command run and relevant passing output after implementation
|
- GREEN: command run and relevant passing output after implementation
|
||||||
|
|||||||
Reference in New Issue
Block a user