Skip to content

fix(server): queued runs land on the timeline where they started - #18094

Draft
juliusmarminge wants to merge 2 commits into
mainfrom
fix/run-timeline-position
Draft

juliusmarminge wants to merge 2 commits into
mainfrom
fix/run-timeline-position

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Oct 11, 2026 •

Copy link
Copy Markdown
Member

Server half of #16987. #17764 already fixed queued runs vanishing on the client.

Problem

Queued runs render above a steer that ran first. Every turn item's position sits in a band of its run's ordinal (ordinal * 1_000_000 + n), and the ordinal is fixed when the run is queued. Take a queued message B (run 2) and then a steer C (run 3). C starts first, but B's items still sort under run 2, so B shows above C even though it ran after it.

Fix

  • Bands follow start order. Each run gets a timeline band the first time it writes a turn item, which happens when it starts. The band is the thread's highest band plus one. Runless items stay in band 0. The bands live in a new orchestration_v2_run_timeline_bands table owned by TurnItemPositionStore. EventSink no longer passes run ordinals in.
  • Run ordinals are unchanged. Run ids, throughRunOrdinal and provider-thread ranges are all derived from them.
  • Existing threads keep their order. Migration 061 gives each run that already has items the band its positions were written in, which is its old ordinal. A run without items gets a band when it starts.
  • Rebuild. ProjectionMaintenance refills the bands from stored positions with the same query.
  • Payload ordinals stay as they are. The ordinal * 100 + n placeholders in turn-item payloads (queued and direct start, handoff, checkpoint, run signals) are left alone. Every write goes through normalize, which replaces them with the next position in the run's band. A queued run writes nothing until it starts, so its message lands in its start-order band.

Tests

  • runtimeLayer.test.ts: "places a queued run that starts late after the runs that started before it".
    • The test sends A, queues B, interrupts A with the queue held, starts C, then resumes B.
    • It checks the order is A, C, B, through the snapshot window, after a rebuild, and after appending D.
    • On main it fails with expected ['A','B','C'] to deeply equal ['A','C','B'].
  • 061_RunTimelineBands.test.ts: the backfill keeps stored bands, and a run started afterwards goes above all of them.
  • Hard-coded migration lists: 061 is added to both.
  • Suites run: migration and runtime suites (81), ThreadFork and ProjectionStore (35), all orchestrator replay fixtures (137), and the apps/server typecheck.

Not covered yet

Fork "through run" cuts still compare run ordinals:

  • which functions: visibleTurnItemsThroughRun, the timeline index, itemCountThroughRun, fork_boundary, and the portable-fork handoff list;
  • the effect: a fork through C still includes the late-starting B.

The fix compares bands at those sites. Those are the same functions #17763 changes, so it follows as a separate PR once #17763 lands.

🤖 Generated with Claude Code (Opus 5.5)

A run's turn items were banded by its run ordinal, which is fixed when the
run is queued. A message queued before a steer therefore rendered above the
steer even though the steer's run started first (#16987).

Each run now gets its own timeline band the first time it writes an item,
which is when it starts. TurnItemPositionStore owns the new
orchestration_v2_run_timeline_bands table. Migration 061 and the projection
rebuild both derive a run's band from the positions its items already hold,
so existing threads keep their order. The run ordinal itself is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Oct 11, 2026
@juliusmarminge juliusmarminge added the macroscope-review Opt PRs made by unvouched contributors in for Macroscope review. Vouched contributors auto-reviews label Oct 11, 2026
@github-actions

github-actions Bot commented Oct 11, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire 5.0 KiB 5.0 KiB 0 B (0.0%) 6.8 KiB ✅
Codex Thread snapshot wire 3.8 KiB 3.8 KiB 0 B (0.0%) 4.9 KiB ✅
Codex Live turn WebSocket wire 1.2 KiB 1.2 KiB 0 B (0.0%) 2.0 KiB ✅
Codex Live turn WebSocket decoded 20.9 KiB 20.9 KiB 0 B (0.0%) 29.3 KiB ✅
Codex Live turn messages 2 2 0 (0.0%) 8 ✅
Claude Total thread wire 5.0 KiB 5.0 KiB 0 B (0.0%) 6.8 KiB ✅
Claude Thread snapshot wire 3.8 KiB 3.8 KiB 0 B (0.0%) 4.9 KiB ✅
Claude Live turn WebSocket wire 1.2 KiB 1.2 KiB 0 B (0.0%) 2.0 KiB ✅
Claude Live turn WebSocket decoded 21.2 KiB 21.2 KiB 0 B (0.0%) 29.3 KiB ✅
Claude Live turn messages 2 2 0 (0.0%) 8 ✅

Baseline: a908a7a · PR result: bafb06d · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 108.5 KiB
  • Claude decoded thread snapshot: 108.8 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@voltcrash

voltcrash commented Oct 11, 2026 •

Copy link
Copy Markdown
Contributor

Note

GPT 6.1 Sol responding on behalf of voltcrash

@juliusmarminge would you consider merging #17773 for the queued-run ordering fix instead?

Both address the same transcript bug, while #17764 already handles queued runs disappearing on the client. #17773 moves a queued run's ordinal forward before it starts when newer work has already started. It keeps run IDs unchanged and preserves the execution-order model used by existing ordinal consumers, without adding a table or migration.

The separate timeline bands in this PR are a reasonable design, but the conversion is incomplete. As the description notes, the transcript can show A → C → B while a fork through C still includes B.

Checkpoint restore needs consideration too: CheckpointRollbackService selects later runs using run.ordinal > targetOrdinal. With execution order A(1) → C(3) → B(2), restoring C would not select B as a later run. My concern is that restore could leave B in the conversation despite its work occurring after C. This is a source-based concern; I haven't reproduced the restore flow.

I'd favor #17773 with focused held-queue → resume → fork/restore regression coverage before merging. The timeline-band approach could then be reconsidered once the remaining ordering consumers are reconciled. Is there a constraint requiring queued run ordinals to remain immutable that makes #17773 unsuitable?

@Mnigos

Mnigos commented Oct 11, 2026

Copy link
Copy Markdown
Contributor

Note

🤖 Claude Fable 5.1 on behalf of Mnigos

I hit this bug myself (a steer promoted from the queue into a held PR-watch run that started after an automatic continuation, the #16987 pattern), so I read this PR and #17773 (#17773) against that case. Both fix it for future runs; neither repairs already-recorded history, which is fine. Two things I found here, from source review plus an isolated SQLite check, not a runtime run:

  1. Rebuild can change the next allocation after an item changes runs. The allocator reserves a new band for the run before ON CONFLICT DO NOTHING preserves the item's old position (TurnItemPositionStore.ts ~L58), while ProjectionMaintenance derives the band from that old position (~L178). With an existing item at 1,000,002 and run C at 2,000,001, the next item lands at 3,000,001 live but at 1,000,003 after a rebuild, so it changes sides of C. A source-supported trigger is retrying an unaccepted mailbox steer with the same message id into another run, then rebuilding before a fresh item establishes the band.
  2. First write is not always start. A held run that already wrote an item before the newer run started keeps that early band, so the "band on first write" rule misses that variant; a regression with a prewritten held item would make the boundary explicit.

For what it's worth, #17773's dequeue-time ordinal reassignment is the smaller model, since fork cuts and ordinal-based "latest run" lookups follow it without a second source of order; it has the same prewritten-item gap and would need the same test.

@bitconym

Copy link
Copy Markdown

Additional real-thread evidence for the ordering defect tracked in #16987 (read-only inspection on macOS, 2026-10-11, Claude provider):

  • Run 105 was queued at 11:29:38.228 UTC and started at 12:43:01.937 UTC.
  • Runs 106–112 were requested later but all started before run 105. For example, run 106 was requested at 11:33:55.745 and started at 11:35:43.765 UTC.
  • The original run-105 user message has timeline position 105000001. A later status? message promoted into that run at 13:11:09.304 UTC has position 105000182. Both therefore sort before the intervening higher-numbered runs, even though run 105 started after them.
  • The user could no longer find the queued task in the main conversation after it left the queue and could not tell whether it had been picked up.

This matches the start-order problem this PR addresses. I checked current main's TurnItemPositionStore: it still allocates positions using runOrdinal * 1_000_000. #17764 is merged for the separate partial-timeline filtering problem; this comment does not establish whether the affected installed client includes that fix.

Separately, a queued run 115 was cancelled before starting: requested 13:12:56.565 UTC, cancelled 13:13:04.299 UTC, no startedAt. Its user message remains persisted, but it has no orchestration_v2_turn_item_positions row. That explains why that message is absent from the timeline, but the cause of cancellation is unknown and this PR does not claim to address it.

Run 105 was cancelled at 13:12:09.253 UTC and followed by continuation runs 114 and 116. These records alone do not prove whether Claude read status? or that restart recovery lost it. I am keeping that hypothesis separate from the confirmed ordering evidence.

Inspected with Codex (gpt-6.1-sol) in T3 Code; no live database writes or application changes.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

macroscope-review Opt PRs made by unvouched contributors in for Macroscope review. Vouched contributors auto-reviews size:L 100-499 changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants