You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Arrow string-view imports pin whole Parquet pages, so the writer never coalesces string columns #10403
String-view columns imported from Arrow keep every Parquet page their views point into, so nbytes() reports up to 15× the bytes the rows use. The writer's coalescing step (RepartitionStrategy, 1 MiB minimum) trusts nbytes(), so every 8,192-row string block looks like it already holds more than 1 MiB and never merges. TPC-H l_comment ends up with 733 chunks of 8,192 rows instead of about 244 chunks of about 24,576 rows.
Cause
arrow-rs's Parquet reader builds each batch's Utf8View/BinaryView views over the whole decompressed page. The TPC-H Parquet stores every string column as Utf8View.
The bench converter's VarBinViewBuilder (compaction_threshold = 0, no deduplication) keeps the buffers and adds a page once for every batch that references it.
VarBinView::slice keeps all data buffers, and nbytes() sums full buffer lengths.
Numeric columns are unaffected because a slice owns exactly len × width bytes.
Evidence (TPC-H SF1, bench converter path)
Column
nbytes() fed to writer (B/row)
Bytes the rows use (B/row)
Chunks written
lineitem.l_comment
637
41
733 × 8,192 rows
orders.o_comment
1,064
65
184 × 8,192 rows
customer.c_comment
1,064
89
19 × 8,192 rows
part.p_name
780
49
25 × 8,192 rows
SpatialBench customer string columns behave the same way. ClickBench and PolarSignals are not affected because their strings import as offset-based Utf8/Binary.
Fix in progress
Branch ji/repartition-view-nbytes, commit d49946a: from_arrow_byte_view now trims view buffers. It drops buffers no valid row references, slices buffers that are at least half used (no copy), and copies sparse values. Fully used buffers are still shared without copying.
Parquet was used as the control in every run and stayed flat. These numbers come from a 2-core VM with stable 1.97; DuckDB was not measured.
Remaining: what the writer measures
The import fix covers Arrow sources only. Any upstream slice of a VarBinView still makes the coalescer over-measure, because nbytes() reports retained memory.
Proposal: keep nbytes() as retained memory. Change the VarBinView arm of UncompressedSizeInBytes to count views plus referenced bytes, matching the ListView arm, which already rebuilds with MakeExact before measuring. Then have RepartitionStrategy measure blocks with that aggregate instead of nbytes().
Alternatives considered:
A second aggregate or a precision option: not needed, because nbytes() is already the cheap upper bound.
Canonical::compact() in the coalescer: also unpins memory and moves work the compressor already does, but should_compact's thresholds (50% utilisation, 128 B/row) leave up to 2× slack in the measurement.
The aggregate change should be coordinated with #10177 and #10235.
Summary
String-view columns imported from Arrow keep every Parquet page their views point into, so
nbytes()reports up to 15× the bytes the rows use. The writer's coalescing step (RepartitionStrategy, 1 MiB minimum) trustsnbytes(), so every 8,192-row string block looks like it already holds more than 1 MiB and never merges. TPC-Hl_commentends up with 733 chunks of 8,192 rows instead of about 244 chunks of about 24,576 rows.Cause
Utf8View/BinaryViewviews over the whole decompressed page. The TPC-H Parquet stores every string column asUtf8View.from_arrow_byte_viewimports those buffers whole, without copying. Trim child data on Arrow slice imports #10379 trimmed slicedUtf8/Binary/Listimports but not view types.VarBinViewBuilder(compaction_threshold = 0, no deduplication) keeps the buffers and adds a page once for every batch that references it.VarBinView::slicekeeps all data buffers, andnbytes()sums full buffer lengths.Numeric columns are unaffected because a slice owns exactly
len × widthbytes.Evidence (TPC-H SF1, bench converter path)
nbytes()fed to writer (B/row)lineitem.l_commentorders.o_commentcustomer.c_commentpart.p_nameSpatialBench
customerstring columns behave the same way. ClickBench and PolarSignals are not affected because their strings import as offset-basedUtf8/Binary.Fix in progress
Branch
ji/repartition-view-nbytes, commit d49946a:from_arrow_byte_viewnow trims view buffers. It drops buffers no valid row references, slices buffers that are at least half used (no copy), and copies sparse values. Fully used buffers are still shared without copying.Results with the fix:
l_commentnbytes()fed to writerl_commentchunks (median rows)compress-benchVortex compress,TPC-H l_comment(wholelineitem)compress-benchVortex decompress, chunkedo_comment),vortex-file-compressedvortex-compactstring-benchl_commentParquet was used as the control in every run and stayed flat. These numbers come from a 2-core VM with stable 1.97; DuckDB was not measured.
Remaining: what the writer measures
The import fix covers Arrow sources only. Any upstream slice of a
VarBinViewstill makes the coalescer over-measure, becausenbytes()reports retained memory.Proposal: keep
nbytes()as retained memory. Change theVarBinViewarm ofUncompressedSizeInBytesto count views plus referenced bytes, matching theListViewarm, which already rebuilds withMakeExactbefore measuring. Then haveRepartitionStrategymeasure blocks with that aggregate instead ofnbytes().Alternatives considered:
nbytes()is already the cheap upper bound.Canonical::compact()in the coalescer: also unpins memory and moves work the compressor already does, butshould_compact's thresholds (50% utilisation, 128 B/row) leave up to 2× slack in the measurement.The aggregate change should be coordinated with #10177 and #10235.
Generated by Claude Code