Skip to content

perf(layout): share a flat layout's decode across the splits reading it - #10272

Open
joseph-isaacs wants to merge 2 commits into
developfrom
ji/flat-shared-decode
Open

joseph-isaacs wants to merge 2 commits into
developfrom
ji/flat-shared-decode

Conversation

@joseph-isaacs

Copy link
Copy Markdown
Contributor

Summary

FlatReader::array_future built a new decode future on every call, and a scan calls it once per split, for filter and projection alike. A flat layout that spans several splits was therefore fetched, deserialized and UTF-8 validated once per split. For example, a 1M-row utf8 chunk read in 8,192-row splits was decoded 123 times, taking 537 ms where one decode takes 15 ms.

Changes

  • FlatReader holds a WeakShared handle to the in-flight decode. A split that asks for the array while another split still holds the decode reuses it.
  • The handle is weak, so the decoded array is freed as soon as the last split using it finishes. A scan does not keep every chunk of a file alive.
  • The segment request is now registered only when a new decode starts. A reused decode already holds a request for the same segment, so a hit no longer registers a read and then cancels it.
  • Tests:
    • splits_share_one_decode_until_dropped: two outstanding splits of one layout make a single segment request, and once both are done the next split requests the segment again. The test fails (2 requests instead of 1) when the cache lookup is disabled.
    • test_flat_chunk_scan_with_row_count_splits (vortex-file): split scans over a flat chunk with unaligned, oversized and many-small splits return the right rows.

Local measurements

Decode reuse, counted with temporary counters that are not committed, one run per query:

query reused new decodes
TPC-H SF=1 lineitem, decode all columns 10,533 1,199
TPC-H Q6 299 164
FineWeb sample, dump = ... 1,158 664
ClickBench 10 partitions, decode all columns 267,851 7,947
ClickBench Q12 196 144

Peak RSS and median time against develop. Peak RSS is the max of 3 processes; time is the median of 9 warm rounds.

query threads develop RSS this PR RSS develop ms this PR ms
TPC-H decode all 1 198 MB 196 MB 362 349
TPC-H decode all 4 246 MB 229 MB 139 134
TPC-H Q6 4 159 MB 163 MB 21.0 19.8
FineWeb dump = ... 4 413 MB 385 MB 61.8 54.6
FineWeb text LIKE 4 697 MB 674 MB 541 502
ClickBench decode all 1 764 MB 721 MB 8,954 7,440
ClickBench decode all 4 956 MB 887 MB 2,903 2,330
ClickBench Q12 4 516 MB 524 MB 559 605

Peak memory is within noise of develop and usually lower. ClickBench Q12 at 4 threads varied between 582 and 715 ms in earlier runs, so its difference here is within noise.

Checks run on this branch:

  • cargo test -p vortex-layout --lib: 307 passed
  • cargo test -p vortex-file --lib: 150 passed
  • cargo clippy -p vortex-layout -p vortex-file --all-targets --all-features -- -D warnings
  • cargo +nightly-2026-09-10 fmt -p vortex-layout -p vortex-file --check

🤖 Generated with Claude Code

https://claude.ai/code/session_019q6jggoqVzaHb8FiJJtUDb


Generated by Claude Code

joseph-isaacs and others added 2 commits October 3, 2026 08:16
`FlatReader::array_future` built a fresh decode future on every call,
and a scan calls it once per split. A flat layout larger than the split
size was therefore deserialized, and its strings UTF-8 validated, once
per split: a 1M-row `utf8` chunk read in ten default splits decoded ten
times, and in 8,192-row splits 123 times (537 ms instead of 15 ms).

Keep a weak handle to the in-flight decode so that splits of the same
layout share it. Holding it weakly frees the decoded array as soon as
the last split using it is done, so a scan still does not keep every
chunk of a file alive.

Single-threaded canonical scans of Vorticity's per-encoding corpus
(1M rows, vortex 0.86.1 files), ns/row:

  varbinview 145.3 -> 17.6, struct 130.3 -> 15.2, map 127.3 -> 16.6,
  table_mixed 303.1 -> 48.3, varbin 151.5 -> 26.3, listview 15.9 -> 3.5

Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019q6jggoqVzaHb8FiJJtUDb
`array_future` registered a segment request before checking for a
decode still shared by another split, so every reuse registered a read
and cancelled it on drop. Check first; a shared decode already holds a
request for the same segment.

Test that two outstanding splits of one flat layout make a single
segment request, and that once both are done the decode is released and
the next split requests the segment again.

Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019q6jggoqVzaHb8FiJJtUDb
@joseph-isaacs joseph-isaacs added the changelog/performance A performance improvement label Oct 3, 2026 — with Claude
@codspeed

codspeed Bot commented Oct 3, 2026

Copy link
Copy Markdown

Merging this PR will improve performance by 51.99%

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 1 improved benchmark
✅ 2102 untouched benchmarks
⏩ 503 skipped benchmarks1

Performance Changes

Mode Benchmark BASE HEAD Efficiency
⚡ Simulation density_sweep_single_slice[0.9] 50.8 µs 33.4 µs +51.99%

Tip

Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.


Comparing ji/flat-shared-decode (382588e) with develop (9a08b82)

Open in CodSpeed

Footnotes

  1. 503 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/performance A performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant