Skip to content

perf(engine): wait for a missing composition timeline once per render, not once per browser - #5227

Open
miguel-heygen wants to merge 3 commits into
mainfrom
perf/timeline-wait-once-per-render
Open

miguel-heygen wants to merge 3 commits into
mainfrom
perf/timeline-wait-once-per-render

Conversation

@miguel-heygen

@miguel-heygen miguel-heygen commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

What

A composition whose animation is not a GSAP timeline (CSS keyframes, canvas, requestAnimationFrame) and whose host lacks data-no-timeline now pays the timeline wait once per render instead of once per browser session.

Why

The engine waits up to the player-ready timeout (45 s by default) for every [data-composition-id] host to register window.__timelines[id]. A host that never registers costs that full wait in every session of a render: the calibration session, then the capture workers, and each static-frame verification page. The outcome cannot change between sessions of the same render, so every wait after the first is spent on a result already known.

A 5-second CSS-only composition (the css-spinner-render-compat fixture with data-no-timeline removed) took 93.8 s to render on a Mac: 45.6 s in calibration and 47.5 s in capture.

Related work

Builds on the fail-fast for failed script loads (scriptLoadFailures) in the same poll. Does not change the timeout, the warning, or the opt-out.

How

  • SubTimelineWaitMemo (engine types.ts) is a render-scoped record of host ids whose wait already timed out. The producer creates one per render and passes it through the probe options and buildCaptureOptions, so every session of the render shares it. A session created without one gets its own.
  • waitForSubCompositionTimelines replaces the two identical wait blocks in initializeSession. It passes the memo's ids to pollSubCompositionTimelines, and records the pending ids after a real timeout (not after a VFX-failure stop).
  • pollSubCompositionTimelines skips known ids while polling, then lists hosts that are still unregistered. If a known host is still missing, the session reports timeout with the same pending ids and the same sub_timeline_readiness_timeout warning, and it skips the rebind, exactly like a session that waited. If it registered since, the session reports ready and rebinds.
  • The static-frame verification page reuses the session's memo, so it no longer waits a second time inside the same session. It also gets the session's failed-script list, so a 404'd scene script ends its wait after the 2 s grace like the main page, instead of the full timeout.
  • Only a clean timeout is remembered: the waiting session ran the full wait (no VFX stop) and saw no page errors and no failed script loads. A render whose composition throws or 404s a script keeps today's behaviour exactly, so sub_timeline_script_failure still reaches the render policy and still fails the render. A replaying session that already sees a failed script load reports script_failure, not timeout.
  • Distributed chunks create one memo per chunk, shared by the chunk's sessions.
  • pollSubCompositionTimelines takes an options object instead of eight positional parameters; existing callers and tests are updated.

Test plan

  • Unit tests added: frameCapture-subTimelineMemo.test.ts runs the real page-side check scripts against a fake DOM on fake timers:
    • a second session sharing the memo settles with no timer advance and reports timeout plus the pending id;
    • a host the memo does not name is still waited for the full timeout;
    • a memo host that has registered since reports ready and rebinds.
  • frameCapture-staticDedupVerificationPage.test.ts: the verification page settles with no timer advance for a memo host, and settles between 1.999 s and 2.5 s when a script failed to load.
  • Also in frameCapture-subTimelineMemo.test.ts: no memo after a page error, no memo after a VFX stop, and a replaying session with a failed script load reports script_failure.
  • Seven mutations, each red on its target test: drop the memo write; drop the skip in the page check; drop the verification page's failed-script list; drop the verification page's memo; memoise despite a page error; memoise despite a VFX stop; drop the replay's script-failure classification.
  • The four touched test files, 3 runs in a row: 26 passed, exit 0 each.

Known residual differences (accepted)

  • A composition that did not register in 45 s in the first session but would register within the first poll of a later session is now reported as a timeout there, without the rebind. That needs a timeline that misses a 45 s window and then appears in milliseconds.
  • The static-frame verification page now stops waiting 2 s after a failed script load, the same rule the main page already applies. A composition that registers after that window is verified without its rebind, exactly as the main page captures it.
  • Before/After on the same Linux box (numbers below); output frames compared via ffmpeg -f framemd5.

Before / After

Linux, 8 cores, software GPU (screenshot capture), Chrome headless shell 152. Median of 5 runs (min-max), same box, load1 under 8 throughout; many-cuts is 3 runs. Output compared by decoded frame hash (ffmpeg -map 0:v -f framemd5) and audio hash (-f md5).

Composition Before After Output
CSS-only clip, no timeline (css-spinner-render-compat without data-no-timeline), 5 s 100.1 s (97.7-106.4) 51.1 s (50.9-51.4) identical, all 10 runs
GSAP timeline on the root, nested host with a CSS spinner and no timeline, 4 s 92.7 s (92.7-93.2) 47.7 s (47.7-47.9) identical, all 10 runs
many-cuts (GSAP, every host registers) 11.45 s (11.23-12.10) 11.30 s (11.10-11.83) identical, all 6 runs

Re-run at this head (2 runs per composition): video and audio hashes are identical to origin/main for all three.

Where it went: capture setup (the first session's wait) is unchanged at ~45.7 s. Frame capture drops from 46.8-49.8 s to 1.8-3.1 s because the workers no longer wait again.

The first 45 s wait per render is unchanged on purpose. Skipping it would mean guessing from the page that no timeline will ever arrive, which changes the documented data-no-timeline contract.

Independent review

An adversarial review at this head found no blocking or major issues. An earlier round found one blocker, now fixed: a replaying session could downgrade a script failure to a plain timeout. Minor items accepted:

  • m1: covered under "Known residual differences".
  • m2: covered under "Known residual differences".
  • m4: a failure that hits only a later session, after a first session that ran clean for the full timeout, is reported as a timeout there. Same class as m1: it needs sessions of one render to diverge.
  • m5: the "no failed script loads" condition on writing the memo has no dedicated test. It only matters for a failure that lands in the final poll interval, and that first session already reports a plain timeout, so replaying it stays consistent.

…, not once per browser

A composition host that never registers a timeline (CSS, canvas or rAF animation
without data-no-timeline) made every capture session of a render wait the full
player-ready timeout: calibration, then the workers, then each static-frame check
page. The first session that times out now records the host ids in a memo shared
by the render's sessions; later sessions report the same timeout warning without
waiting again.
@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Edit accuracy: accurate 2059 (base branch 2059), smooth 1601 of those

The gate passes.
Smoothness is reported in the artifact, not gated. A case fails only if it fails 2 of 3 runs.

Quarantined, measured but not gated (0)

…d chunks share one memo

The verification page now gets the session's failed-script list, so it bails after
the short grace like the main page instead of waiting the full timeout. Distributed
chunks share one timeline-wait memo across their sessions. The poll takes an
options object instead of eight positional parameters.
… still fails the render

A session that skipped the wait classified it at once, before a late page error
could arrive, so a render that used to stop on a script failure could ship. The
memo is now written only by a wait that ran its full length with no page errors
and no failed script loads; a replaying session that already sees a failed script
load reports a script failure.
@miguel-heygen
miguel-heygen marked this pull request as ready for review October 8, 2026 16:05

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant