fix(sdd,codex): bounded wait stretches with reconciliation

Round 2 proved the long-wait mechanism (65.1%->0.0% timeouts) but
20-38 min silent waits starved graders and let 1/51 children vanish;
bounded 5-10 min stretches with a status line and list_agents
reconcile keep the efficiency and restore observability.
This commit is contained in:
Jesse Vincent
2026-07-30 15:57:32 -07:00
parent db4538fcb8
commit d8189d1587
2 changed files with 18 additions and 11 deletions

View File

@@ -198,12 +198,15 @@ prints back — stays resident in your context for the rest of the session
and is re-read on every later turn. Hand artifacts over as files.
**Waiting on dispatched subagents:** never poll a wait interface with
short timeouts. While you have local work — ledger updates, packaging
the next review, reading reports — keep working; child results arrive
on their own. Wait only when you are genuinely idle, and then issue one
long wait (fifteen minutes or more, where your platform allows it)
instead of many short ones: a long wait wakes just as fast and costs
one call instead of dozens.
short timeouts, and never sit in one silent, open-ended wait either.
While you have local work — ledger updates, packaging the next review,
reading reports — keep working; child results arrive on their own.
When you are genuinely idle, wait in bounded stretches (five to ten
minutes, where your platform allows), and between stretches post one
line of status and reconcile your live children: list them, and chase
any that finished without reporting. A bounded stretch keeps nearly
all of a long wait's efficiency while guaranteeing a stuck or lost
child is noticed within minutes, not at the end of the session.
### 1. Dispatch the implementer

View File

@@ -47,13 +47,17 @@ two-thirds of all wait calls were short polls that timed out.
- While you still have local work, do not wait at all. A completed
child's final answer is pushed into your mailbox and arrives with
your next turn.
- When you are genuinely idle with children outstanding, issue ONE
`wait_agent` with a long `timeout_ms` — 900000 (15 minutes) or more —
and let the event wake you.
- When you are genuinely idle with children outstanding, wait in
bounded stretches: `wait_agent` with `timeout_ms` 300000-600000
(5-10 minutes). After each stretch — wake or timeout — post one
status line, run `list_agents`, and chase any child that finished
without reporting. Never stack polls shorter than five minutes; the
event subscription wakes a bounded stretch just as fast as a short
one.
- Completion mail cannot wake an idle controller (it is delivered
without triggering a turn); covering that idle window is
`wait_agent`'s only job. If a long wait times out, check
`list_agents` for stuck children — do not fall back to short polls.
`wait_agent`'s only job. A stretch that times out with no activity
is your cue to reconcile, not to shorten the next stretch.
## Environment Detection