Repository navigation
Conversation
…watch wakes Subagents cannot watch pull requests; watches they inherited end without a wake. A host rate limit no longer counts toward giving up on a watch. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ApprovabilityVerdict: Approved at Macroscope's review found this PR approvable — This is a focused server bug fix that prevents subagents from creating or honoring pull-request watches that cannot deliver useful results, while preserving parent-thread watches and pausing cleanly through host rate limits. The production changes are narrowly scoped and covered by targeted tests, with no sensitive, schema, deployment, or static-analysis impact. You can add or adjust custom eligibility rules. Learn more. |
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
…lone Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Dismissing prior approval to re-evaluate 74f80f8
|
Note 🤖 Claude Opus 5.5 responding on behalf of Theo Closing as superseded by #16208, which shipped every behavior here: subagents cannot watch, watches they inherited end without a wake, and a rate limit never ends a watch. #16208 also shares one read per pull request across watchers and slows the cadence. After #16208 landed, I checked whether its subagent refusal text reaches the agent. It does: for a declared tool failure, Effect's MCP server sends |
Delegated subagents were burning huge numbers of tokens on pull request watch wakes nobody reads. Across the fleet since Oct 1, finished subagents ran about 4,000 extra times. On cup2 alone that was 3,635 extra runs and about 1.48B input tokens, against about 398M for the subagents' real work. One review subagent ran 151 times.
Two things combined:
Fix:
watch_pull_requestreturns a typed refusal whose text reaches the agent, telling it to put the pull request in its result so the parent can watch it. The orchestrator refuses the command too.Tests cover a rate-limited watch (stays on, no wake after 30 passes) and a subagent with an inherited watch (new watch refused, inherited one ended without a host read or wake). Both fail without the fix. The Stop test now gives its subagent an inherited watch instead of starting one, because subagents can no longer start one.
Not changed: review subagents still call
link_pull_requestfor the pull request they review, because the runtime instructions tell every agent to link pull requests it works on. Links do not wake anyone, so this costs a few tool calls per subagent.Reviewed with sol-loop: 2 rounds with GPT-6.1 Sol on high.
Created with Claude Opus 5.5 in Claude Code.