Skip to content

Retry compose lifecycle commands on a panic, keep both ends of script output, bump compose to v5.5.1 - #40

Open
earakely-scale wants to merge 4 commits into
mainfrom
edgararakelyan/fd-3360-universe-load-dies-on-an-intermittent-docker-compose-panic
Open

earakely-scale wants to merge 4 commits into
mainfrom
edgararakelyan/fd-3360-universe-load-dies-on-an-intermittent-docker-compose-panic

Conversation

@earakely-scale

@earakely-scale earakely-scale commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Why

load_artifact of a universe onto a multi env on modal_vm intermittently fails with Script failed (exit 2) and the tail of a Go goroutine dump, raised from MultiEnv._load_from_snapshot at docker compose up -d --force-recreate … (one dump also shows the docker compose rm -sf -v servicedb && up -d servicedb path). Exit 2 is Go's unrecovered-panic exit; the binary is compose v2.29.7, and the panic family is compose's known data races during up --force-recreate / stop (docker/compose#10319, #10468, #12335, #12787, #12834). Six first attempts on one env hit it between 2026-09-30 and 2026-10-01; the activity-level retry hid all of them except a run started in-process, where the first crash was fatal. Diagnosis was blocked by exec_script keeping only the last 1500 chars of stderr, which for a Go crash never includes the panic: line.

What

  • VmSandbox.exec_script gains retry_exit_codes; _load_from_snapshot passes max_retries=2, retry_exit_codes=(2,) on its two compose lifecycle commands, both idempotent. The pg_isready probe keeps its own loop and is untouched (pg_isready itself exits 2 for "no response").
  • The error raised by exec_script now carries the first and last 1500 chars of stdout and stderr around an elision marker (clip_output), so the next crash is attributable.
  • _COMPOSE_VERSION v2.29.7 → v5.5.1, as a separate commit. The VM image's docker.io is Docker Engine 29.1.3 today; the plugin was from September 2024.

Verification

  • tst/unit: 5768 passed, 7 new (retry only on listed codes, fail fast otherwise, give up after max_retries, panic header and tail both survive, clipping exact).
  • Live chaos run, dev stage, branch installed editable: deploy the env, replace the VM's docker-compose plugin with a shim that prints a fake goroutine dump and exits 2 on its first up -d --force-recreate and execs the real binary afterwards, load the universe; expect one logged retry and a clean load. Then a plain deploy + load for regression on the new compose. Results in a comment below.

Linear: FD-3360

🤖 Generated with Claude Code

RetriggerConfidence Score: 4/5

The PR should meet the repository’s variable-placement rule before merging.

Fix All in CursorFindings

  1. P2 Retry counts lack names ▶
Fix with agent prompt
### Issue 1
src/agent_env/env/envs/multi_env.py:undefined-599
Both snapshot commands hardcode `max_retries=2`, and the retry log in `sandbox.py` uses an inline `200`-character limit. The repository requires magic numbers to be stored as descriptively named class or instance variables. Name these limits before merging so their purpose is clear and the two commands stay in sync.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

Multi-environment snapshot restores retry Compose exit code 2, and script errors keep both ends of long output. Modal VM images also move to Compose v5.5.1.

  • Retries the two Compose lifecycle commands on the selected exit code.
  • Keeps the start and end of long output in errors and retry logs.
  • Updates the Compose plugin version in Modal VM images.

Reviews (2) · Last reviewed commit: "Name the compose crash retry budget and ..."

earakely-scale and others added 2 commits October 1, 2026 11:23
…ipt output

A universe load onto a multi env dies when the docker compose binary itself
crashes during the snapshot restore (`rm -sf -v` + `up -d`, or
`up -d --force-recreate`): compose has data races that end in a Go runtime
panic, exit code 2, and the step failed on the first one even though both
commands are idempotent. exec_script grows an opt-in retry_exit_codes, and the
two compose calls in _load_from_snapshot use it with (2,).

The error raised by exec_script kept only the last 1500 chars of each stream,
which for a Go crash is the end of the goroutine list and never the panic
header. It now keeps the first and last 1500 chars around an elision marker.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The VM image installs Ubuntu 22.04's docker.io, now Docker Engine 29.x, next
to a compose plugin from September 2024. Move to the current release so the
engine and the plugin come from the same era.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@earakely-scale
earakely-scale requested a review from a team as a code owner October 1, 2026 21:23
The servicedb recreate merges stderr into stdout, so the retry warning showed
an empty stderr for exactly the crash it exists to surface.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Comment thread src/agent_env/env/envs/multi_env.py Outdated
await self._sandbox.exec_script(
f"cd {GATEWAY_APP_DIR} && docker compose rm -sf -v {DATABASE_SERVICE_NAME} && docker compose up -d {DATABASE_SERVICE_NAME} 2>&1"
f"cd {GATEWAY_APP_DIR} && docker compose rm -sf -v {DATABASE_SERVICE_NAME} && docker compose up -d {DATABASE_SERVICE_NAME} 2>&1",
max_retries=2, retry_exit_codes=(2,),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Retry counts lack names

Both snapshot commands hardcode max_retries=2, and the retry log in sandbox.py uses an inline 200-character limit. The repository requires magic numbers to be stored as descriptively named class or instance variables. Name these limits before merging so their purpose is clear and the two commands stay in sync.

Rule Used: Store magic numbers as class or instance variables with descriptive names rather than using them inline in the code. (source)

Learned From
scaleapi/scaleapi#126388

Prompt To Fix With AI
This is a comment left during a code review.
Path: src/agent_env/env/envs/multi_env.py
Line: 599

Comment:
**Retry counts lack names**

Both snapshot commands hardcode `max_retries=2`, and the retry log in `sandbox.py` uses an inline `200`-character limit. The repository requires magic numbers to be stored as descriptively named class or instance variables. Name these limits before merging so their purpose is clear and the two commands stay in sync.

**Rule Used:** Store magic numbers as class or instance variables with descriptive names rather than using them inline in the code. ([source](https://app.greptile.com/scale-ai/-/custom-context?memory=002e0051-41ad-46c1-9098-47433c580150))

**Learned From**
[scaleapi/scaleapi#126388](https://github.com/scaleapi/scaleapi/pull/126388)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Fix in Cursor Fix in Claude Code Fix in Codex

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in bbf6386: the two compose calls now share _COMPOSE_CRASH_RETRIES / _COMPOSE_CRASH_EXIT_CODES (module constants in multi_env.py), and the retry log uses _RETRY_LOG_CLIP_CHARS next to _OUTPUT_CLIP_CHARS in sandbox.py.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@earakely-scale

Copy link
Copy Markdown
Collaborator Author

Live verification

Branch installed editable; env hg4-with-gsuite-bundle-1174 v34 + universe hg4-2026-08-31 v35 (clean snapshot v8, the same records as the failing run); modal_vm 2 vCPU / 8 GB, gateway_mode=consistent; throwaway env instances, all torn down afterwards (tunnels no longer answer). The chaos shim replaces the VM's docker-compose CLI plugin with a wrapper that prints panic: concurrent map writes plus 300 goroutine frames and exits 2 on the first compose command matching a pattern, then execs the real binary.

run shim fires on retries logged load containers after gateway card MCP tools linear_list_issues wall time
chaos A up -d --force-recreate … 1 ok, snapshot restore 17 up (15 healthy + pgweb, db-mcp, which have no healthcheck) 200 357 ¹ rows returned 297 s
chaos B rm -sf -v servicedb && up -d servicedb 1 ok, snapshot restore 17 up 200 357 rows returned 419 s ²
regression none 0 ok, snapshot restore 17 up 200 357 rows returned 499 s

¹ run A's harness listed tools with a wrong filter and counted 0; fixed before B and the regression run. ² includes a ~90 s Mongo write timeout on the instance record from my machine, unrelated to the change.

Both chaos runs show the new warning with the panic header preserved:

exec_script exit 2 (retryable), retrying in 1s (attempt 1/3); output: 'panic: concurrent map writes (chaos shim)\n\ngoroutine 1 [running]:\n…\n... [16658 chars elided] ...\n…readLoop(0xc0004e66c0) frame 298…'

Inside every VM: Docker Compose version v5.5.1, Docker version 29.1.3, build 29.1.3-0ubuntu3~22.04.2, so Modal rebuilt the image from the new pin and the full restore (rm -sf -v, up -d, exec -T pg_isready, up -d --force-recreate) runs on it.

Unit suite: 5768 passed. Greptile: one P2 (unnamed constants), fixed in bbf6386.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant