Skip to content

fix(server): a failed rollback shows an error instead of hanging - #13887

Merged
juliusmarminge merged 2 commits into
t3code/codex-turn-mappingfrom
v2/rollback-failure-signal
Sep 27, 2026
Merged

juliusmarminge merged 2 commits into
t3code/codex-turn-mappingfrom
v2/rollback-failure-signal

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Sep 26, 2026 •

Copy link
Copy Markdown
Member

A provider rollback that fails ("Edit from here" / revert) left the web UI on "Thinking" for 2 minutes, then showed a generic timeout. The provider-thread.rollback effect retried five times and failed silently. The client had nothing to wait on except runs reaching rolled_back.

What changed

Server

  • On the last failed attempt of provider-thread.rollback, EffectWorker dispatches an internal checkpoint.rollback.fail command. Its decider emits thread.metadata-updated with a new optional thread.rollbackFailure: { requestId, message }. The marker goes through the normal command → event → projector path; the worker writes no projection state itself.
  • The next accepted checkpoint.rollback clears the marker, so a stale failure never rejects a later rollback.
  • The contract change is one optional field on OrchestrationV2AppThread (the JSON mirror derives from it, and there is no new export). Old clients and stored events still decode.
  • The catch-all unexpected-failure rollback message is now user-facing, not an internal id dump.

Web

  • onRevertToTurnCount allocates the rollback's command id and passes it to waitForRevertedMessage. That function rejects with the server's reason as soon as thread.rollbackFailure names that request. The existing setThreadError path shows it.
  • A rewind no longer counts toward isWorking, so the timeline stops showing the "Thinking" row. The composer keeps its existing disabled state ("Rewinding conversation", inert overlay). Compaction still stays disabled during a rewind. No components/ui changes.

Mobile: no change. It has no rewind flow; apps/mobile never dispatches checkpoint.rollback.

Verification

  • Server: new runtimeLayer.test.ts case. The real worker drains a rollback whose provider session cannot open, with the clock advanced through the retry backoff. It asserts that the effect ends failed, that thread.rollbackFailure is projected with the request id and message, and that the next rollback clears it.
    • Without the worker change: fails (expected undefined to deeply equal {…}).
    • With it: passes.
  • Web: new waitForRevertedMessage logic tests (no rendering). A projected failure for the request rejects with its message; a failure recorded for an earlier request is ignored.
    • Without the rejection branch: the first test hangs until timeout.
    • With it: passes.
  • vp test run, all passing:
    • server runtimeLayer, EffectWorker, CheckpointRollbackService: 68/68
    • contracts orchestrationV2: 27/27
    • web ChatView.logic, MessagesTimeline.logic, composerDraftStore: 405/405
  • tsc --noEmit is clean for packages/contracts, packages/client-runtime, apps/server, apps/web, and apps/mobile.
  • vp lint on the touched files reports only pre-existing warnings. knip's unused-export count is unchanged (785).
  • Not run: repo-wide checks, and a real-client pass in the browser.

Model: Claude Opus 5.5 (Claude Code)

🤖 Generated with Claude Code

The provider-thread.rollback effect retried five times and then failed
silently, so clients waited for runs to reach rolled_back until they timed
out. On the last failed attempt the worker now dispatches an internal
checkpoint.rollback.fail command, which sets thread.rollbackFailure with a
user-facing reason through the normal decider and projector. The next
checkpoint.rollback clears it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Sep 26, 2026
}),
/** Server-only: records that the provider rollback for `requestId` failed for good. */
Schema.Struct({
type: Schema.Literal("checkpoint.rollback.fail"),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium src/orchestrationV2.ts:2607

A late checkpoint.rollback.fail from an older rollback overwrites the thread's rollbackFailure after a newer rollback has already cleared it, so a successful newer rollback still leaves a stale failure visible. This command carries only the old requestId, and its handler records the failure without verifying that it is still the current rollback; add a current-request/generation check and ignore superseded failures.

Also found in 3 other location(s)

apps/server/src/orchestration-v2/EffectWorker.ts:362

The failure marker is dispatched after the rollback effect has already failed, without verifying that its rollback request is still the latest one. If a user starts a new rollback in the interval before this dispatch obtains the thread lock, that new command sees no existing marker and cannot clear it; the old attempt then writes its rollbackFailure after the newer rollback (even if the latter succeeds). This leaves a stale failure visible on the thread and contradicts the intended clearing behavior.

apps/server/src/orchestration-v2/Orchestrator.ts:7939

dispatchCheckpointRollbackFail unconditionally overwrites rollbackFailure for command.requestId. A prior rollback can be waiting for its retry while a later rollback is accepted and clears the marker; effects are only serialized while one is running, so the later rollback can settle before the older retry exhausts. When that older effect then fails, this line restores its stale failure marker, leaving the thread reporting a failed rollback after the newer request completed.

apps/server/src/orchestration-v2/Orchestrator.ts:8929

Dispatching checkpoint.rollback.fail unconditionally records the failure even when its original rollback has been superseded. For example, rollback A can still be retrying, rollback B can be accepted while no marker exists, and then A's final failed retry dispatches this command afterward; the handler invoked here writes A's rollbackFailure over B's state. The stale error remains until yet another rollback and can be displayed as a failure after the newer rollback has succeeded.

🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @packages/contracts/src/orchestrationV2.ts around line 2607:

A late `checkpoint.rollback.fail` from an older rollback overwrites the thread's `rollbackFailure` after a newer rollback has already cleared it, so a successful newer rollback still leaves a stale failure visible. This command carries only the old `requestId`, and its handler records the failure without verifying that it is still the current rollback; add a current-request/generation check and ignore superseded failures.

Also found in 3 other location(s):
- apps/server/src/orchestration-v2/EffectWorker.ts:362 -- The failure marker is dispatched after the rollback effect has already failed, without verifying that its rollback request is still the latest one. If a user starts a new rollback in the interval before this dispatch obtains the thread lock, that new command sees no existing marker and cannot clear it; the old attempt then writes its `rollbackFailure` after the newer rollback (even if the latter succeeds). This leaves a stale failure visible on the thread and contradicts the intended clearing behavior.
- apps/server/src/orchestration-v2/Orchestrator.ts:7939 -- `dispatchCheckpointRollbackFail` unconditionally overwrites `rollbackFailure` for `command.requestId`. A prior rollback can be waiting for its retry while a later rollback is accepted and clears the marker; effects are only serialized while one is `running`, so the later rollback can settle before the older retry exhausts. When that older effect then fails, this line restores its stale failure marker, leaving the thread reporting a failed rollback after the newer request completed.
- apps/server/src/orchestration-v2/Orchestrator.ts:8929 -- Dispatching `checkpoint.rollback.fail` unconditionally records the failure even when its original rollback has been superseded. For example, rollback A can still be retrying, rollback B can be accepted while no marker exists, and then A's final failed retry dispatches this command afterward; the handler invoked here writes A's `rollbackFailure` over B's state. The stale error remains until yet another rollback and can be displayed as a failure after the newer rollback has succeeded.

}),
/** Server-only: records that the provider rollback for `requestId` failed for good. */
Schema.Struct({
type: Schema.Literal("checkpoint.rollback.fail"),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium src/orchestrationV2.ts:2607

Any client with the normal operate scope can submit checkpoint.rollback.fail through orchestration.dispatchCommand and persist a fabricated rollbackFailure for an arbitrary requestId and message, including an active rollback, causing the UI to reject a real rollback. The Server-only comment does not enforce that restriction because this member is part of the publicly accepted OrchestrationV2Command union; remove it from the public schema or enforce server-only authorization before dispatch.

🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @packages/contracts/src/orchestrationV2.ts around line 2607:

Any client with the normal operate scope can submit `checkpoint.rollback.fail` through `orchestration.dispatchCommand` and persist a fabricated `rollbackFailure` for an arbitrary `requestId` and message, including an active rollback, causing the UI to reject a real rollback. The `Server-only` comment does not enforce that restriction because this member is part of the publicly accepted `OrchestrationV2Command` union; remove it from the public schema or enforce server-only authorization before dispatch.

} from "./EffectOutbox.ts";
import { CheckpointRollbackServiceV2 } from "./CheckpointRollbackService.ts";
import {
CheckpointRollbackServiceV2,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This changed import consumes CheckpointRollbackServiceV2 as an Effect service, so import the local module as a namespace and access its public members through that namespace (including the message constant). This preserves the service-module boundary convention. The fix also updates the service references below, so it is not confined to this diff hunk.

Posted via Macroscope — Effect Service Conventions

commandId: CommandId.make(`${effect.commandId}:rollback-failed`),
threadId: effect.threadId,
requestId: effect.commandId,
message: Option.match(Cause.findErrorOption(cause), {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Persisting Cause.findErrorOption(cause).message exposes arbitrary lower-level failure text to clients; it can contain provider payloads, URLs, or other sensitive details. Suggest persisting the bounded, safe rollback message instead and keeping the underlying cause in server-side diagnostics.

Suggested change
message: Option.match(Cause.findErrorOption(cause), {
message: ROLLBACK_FAILED_MESSAGE,

Posted via Macroscope — Effect Service Conventions

@github-actions

github-actions Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 4.9 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.4 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 4.9 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 20.7 KiB — 29.3 KiB ✅
Claude Live turn messages — 1 — 8 ✅

Baseline: unavailable · PR result: 8ab4665 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@macroscopeapp

macroscopeapp Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This production rollback fix adds a new failure-recording command and changes existing server and UI behavior. The command is publicly dispatchable despite being labeled server-only, late failures can leave stale rollback state, and raw provider errors may be exposed to clients.

Not approved because:

  • 2 blocking correctness issues found at or above your repo's Minimum Blocking Severity

Adjust the Minimum Blocking Severity for this repo — including turning it Off — in Settings. You can add or adjust custom eligibility rules. Learn more.

…ing row

waitForRevertedMessage now takes the rollback's command id and rejects with
the server's reason as soon as thread.rollbackFailure names that request,
instead of waiting up to 2 minutes. A rewind no longer counts as agent work,
so the timeline stops showing Thinking; the composer keeps its "Rewinding
conversation" disabled state. The rollback failure schema is inlined so the
contract adds no new export.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added size:L 100-499 changed lines (additions + deletions). and removed size:M 30-99 changed lines (additions + deletions). labels Sep 26, 2026
@juliusmarminge
juliusmarminge merged commit 9fcc9a4 into t3code/codex-turn-mapping Sep 27, 2026
24 of 25 checks passed
@juliusmarminge
juliusmarminge deleted the v2/rollback-failure-signal branch September 27, 2026 00:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L 100-499 changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant