Skip to content

fix(compositor): restore the 2.0 export speed on release/v2.0.0 - #984

Merged
EtienneLescot merged 8 commits into
release/v2.0.0from
cherry/v2.0.0-export-perf
Oct 2, 2026
Merged

EtienneLescot merged 8 commits into
release/v2.0.0from
cherry/v2.0.0-export-perf

Conversation

@EtienneLescot

@EtienneLescot EtienneLescot commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Cherry-picks the code of #982 onto the frozen release/v2.0.0: the Linux export of 2.0.0-rc.12 runs at 2.43× the ffmpeg floor against 1.29× for 1.11.0-rc.1 on the same machine and session, and this restores it (1.29×, byte-identical MP4). Full write-up in #982.

The eight code commits, git cherry-pick -x in order; the two docs commits of #982 are left on main, as the release contract asks:

  • fix(compositor): keep the 3D models out of the Linux layer pipeline
  • perf(compositor): evaluate one corner distance per pixel on Linux
  • perf(compositor): scissor the Linux cursor trail composite to the trail
  • perf(compositor): copy a static background instead of redrawing it on Linux
  • fix(compositor): pick the Linux layer pipeline from the layer's mode
  • fix(compositor): keep the 3D models out of the D3D11 and Metal layer shaders
  • perf(compositor): evaluate one corner distance per pixel on Windows and macOS
  • perf(compositor): compare the background key without copying the background

crates/ here is identical to #982's tip (git diff 73de470e -- crates/ is empty), so #982's CI results carry over. #983 (the same optimizations ported to Windows and macOS) is not included: it is not a regression fix.

After merge: rerun prerelease.yml with rc_number 13 to tag the branch tip.

Type of change

  • Bug fix
  • Performance

Release impact

  • Patch

Desktop impact

  • Windows
  • macOS
  • Linux

Testing

See #982: Linux GPU tests and benchmark on a Radeon 610M; Windows and macOS compiled and tested in CI only (the macOS runner renders the 3D-model tests on Metal), not measured on hardware.

🤖 Generated with Claude Code

Every layer is drawn by one WGSL pixel shader that branches on its mode, and
ACO allocates registers for the heaviest branch. With the ray-traced cursor
model, its click impact and the device frame (modes 15-17) in it, the shader
needed 168 VGPRs and ran 6 waves per SIMD instead of 32: the wallpaper, the
shadow and the screen paid that occupancy on every pixel. On a Radeon 610M the
benchmark export went from 31 s (1.11.0-rc.1) to 57 s (2.0.0-rc.12).

Compile layer.wgsl twice, with and without LAYER_MODELS, and draw only the
three modelled layers through the second pipeline. The flat one needs 48 VGPRs
(20 waves). Export 57 s -> 36.5 s, same pixels.

(cherry picked from commit 1bd3e4d)
`select` evaluates both of its operands, so the screen's corner mask computed
three continuous-corner SDFs per pixel (sd_round_rect, and sd_screen_under_bar
which calls it twice) where one was used: ~0.6 ms per 1080p frame on the two
compute units of a Radeon 610M. Branch on the uniform instead, in the screen
tail and in the tilted mode 8.

sd_round_rect also returns the plain box distance away from the corners, where
the exponent and the extent cancel out, so exp2 and the divisions run only near
them. Screen draw 1.53 -> 0.81 ms; the export is unchanged to the byte.

(cherry picked from commit 9c7f661)
The cursor trail accumulates in its own target and was then copied back over
the whole output: 0.5 ms per 1080p frame on a Radeon 610M, on every frame where
the cursor moves, for a sprite a few dozen pixels wide. Only the trail's quads
are ever painted into `accum`, so the copy is scissored to their union.

(cherry picked from commit 1204aca)
… Linux

Pass 1 sampled the 3840 px wallpaper without mips onto the whole output every
frame: 1.65 ms on a Radeon 610M, more than the screen itself, for a picture
that does not change. Keep the background as pass 1 and its blur left it,
keyed on everything that decides it (clear colour, layer, blur amount, render
size), and copy it into the render target while the key holds. An animated
background is never kept, nor one whose image failed to load.

Benchmark export on that machine: 34.1 s -> 30.9 s, against 31.3 s for
1.11.0-rc.1 in the same session and 57 s for 2.0.0-rc.12. The MP4 is
byte-identical to rc.12's.

(cherry picked from commit c00e5b7)
The models pipeline was chosen at the three call sites known to draw a 3D
model: the cursor model, its click impact and the device frame. The device's
shadow is a mode-17 model too (device_shadow_cb builds on device_frame_cb), and
it went through the flat pipeline, which compiles no model and drew nothing.

LayerCB::needs_models now says which layers need the models variant, shared by
the three backends, and every Linux layer is drawn by draw_layer, which reads
it from the LayerBind that make_bind returns. A call site can no longer get it
wrong: forcing the flag off fails the device shadow, device frame, cursor trail
and click impact tests.

(cherry picked from commit 294b0eb)
…shaders

Same split as on Linux: ps_main is compiled once without LAYER_MODELS for
every layer, and once with it for the cursor model, its click impact, the
device and its shadow, chosen per draw from LayerCB::needs_models. D3D11 gets
ps_main_models from build.rs (a D3D_SHADER_MACRO on the same source); Metal
compiles a second library with the macro prefixed, and draw_layer switches to
the models variant of the pass's pipeline for the draw.

The flat ps_main compiles to 40 KB of DXIL against 137 KB with the models
(dxc 1.9, ps_6_0). Neither backend can be built from Linux: compile_hlsl was
type-checked against x86_64-pc-windows-msvc, the HLSL compiled with dxc, and
every_shader_entry_point_compiles now builds the Metal variant too.

(cherry picked from commit 48c2778)
…nd macOS

The HLSL and Metal ports of 9c7f661. `?:` with two SDF calls computed three
continuous-corner SDFs per screen pixel where one is used (HLSL evaluates both
operands), and sd_round_rect paid for its exponent away from the corners,
where it cancels out. A uniform branch, and the box distance off the corners.

(cherry picked from commit 527f934)
…ground

The Linux background cache built its key every frame from the image path's
bytes, so a wallpaper stored as a data URI was copied in full, several
megabytes, on every frame. The key is now a BackgroundKey kept with the cached
copy (the scene's background, blur, fallback colour, render size) and
compared field by field against the requested one; it is only built, and the
background cloned, when a new background is kept. Shared in frame_geometry so
the other backends can use the same rule.

(cherry picked from commit 73de470)
@coderabbitai

coderabbitai Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 7ffb3745-c466-4c61-a840-6da7c2423230

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@EtienneLescot
EtienneLescot merged commit d227e05 into release/v2.0.0 Oct 2, 2026
18 of 19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant