Compare commits

..

1 Commits

Author SHA1 Message Date
Jesse Vincent
61f669ebc9 fix(sdd): preflight emits its pairwise checks as a ledger table and rules on what it surfaces
The pre-Task-1 conflict scan currently permits 'the scan is clean' with
no evidence the scan happened — mined sessions show controllers skipping
straight to dispatch and plan conflicts surfacing mid-execution as
blocking questions. Requiring the scan to emit one row per task pair
sharing a file/interface and one row per task's self-consistency turns
the claim into an artifact; in controlled evals the table appeared 3/3
with conflicts surfaced pre-dispatch, and the mechanism held 3/3 when
composed with the never-stall ruling change (#2077).

Claude-Session: https://claude.ai/code/session_0185AJr98gHx5EmwqNeft4Sy
2026-08-03 09:03:05 -07:00
2 changed files with 14 additions and 34 deletions

View File

@@ -142,17 +142,24 @@ a ledger file, not only in todos.
Read the plan once, note its context and Global Constraints, and create a
todo per task.
Before dispatching Task 1, scan the plan once for conflicts:
Before dispatching Task 1, scan the plan once for conflicts, writing down
what you checked as you check it:
- tasks that contradict each other or the plan's Global Constraints
- anything the plan explicitly mandates that the review rubric treats as a
defect (a test that asserts nothing, verbatim duplication of a logic block)
Present everything you find to your human partner as one batched question —
each finding beside the plan text that mandates it, asking which governs —
before execution begins, not one interrupt per discovery mid-plan. If the
scan is clean, proceed without comment. The review loop remains the net for
conflicts that only emerge from implementation.
The scan's output is a table, not a verdict. One row for every pair of tasks
that share a file or an interface: the two tasks, what one produces against
what the other consumes, and what you found. One row for every task: whether
its own text agrees with itself — the tests it specifies against the code it
specifies, the files it creates against the files it later touches. "The scan
is clean" without those rows is not a scan you ran.
Write the table to the ledger. Rule on each conflict it surfaces — the spec
is the binding authority, the plan is its argument — record the ruling beside
its row, and dispatch Task 1. The review loop remains the net for conflicts
that only emerge from implementation.
## Model Selection

View File

@@ -7,34 +7,7 @@ Add to your Codex config (`~/.codex/config.toml`):
multi_agent = true
```
This enables the multi-agent tools that skills like
`dispatching-parallel-agents` and `subagent-driven-development` use.
Which tools you get depends on the multi-agent version your model
preset selects (current presets run V2; older ones run V1). Trust your
actual tool list over any table — including this one — when they
disagree.
- **Spawning:** give children a clean context with
`spawn_agent {fork_turns: "none"}`; the default `"all"` copies your
entire transcript into the child. On Codex 0.145+, role files under
`~/.codex/agents/` attach to isolated forks via `agent_type`.
Full-history forks accept `model` and `reasoning_effort` overrides
(only `agent_type` is refused there) — isolated forks are the SDD
default for context hygiene, not because overrides require them.
- **Fix rounds:** resume the implementer with `followup_task` — it
delivers your message, triggers a turn, and transparently reloads a
child the harness evicted. Never dispatch a fresh implementer on the
theory that a spawned agent cannot be messaged again; on V2 it
always can.
- **Lifecycle:** V2 has no `close_agent`. Finished children are
evicted automatically when slots are needed; leaving them unclosed
costs nothing. Only V1 sessions have `close_agent` — there, close
reviewers when their review returns, and close each implementer
after its task's review passes.
- **Model names:** never copy a model name from a skill, table, or old
session into `spawn_agent` without checking it against your current
spawn allowlist — V2 accepts only V2-capable presets and hard-errors
on the rest.
This enables `spawn_agent`, `wait_agent`, and `close_agent` for skills like `dispatching-parallel-agents` and `subagent-driven-development`. When using subagent-driven-development, close reviewer subagents when their review returns. Keep each implementer subagent open until its task's review passes — the fix loop resumes the implementer — then close it. If your harness cannot send another message to a spawned agent, dispatch each fix round as a fresh implementer carrying the brief, the report file, and the findings.
## Environment Detection