Skip to content

Slice blocked BitPacked without decoding - #10250

Draft
mhk197 wants to merge 1 commit into
mk/bitpacked-v2-schemefrom
mk/bitpacked-blocked-slice
Draft

mhk197 wants to merge 1 commit into
mk/bitpacked-v2-schemefrom
mk/bitpacked-blocked-slice

Conversation

@mhk197

@mhk197 mhk197 commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Tracking Issue: #10167

Stacked on #10245; this PR targets mk/bitpacked-v2-scheme.

Summary

Slicing a BitPacked array whose blocks have their own bit widths now runs as a kernel instead of decoding the whole array first. The result is still a blocked BitPacked array, holding only the blocks the slice overlaps.

Block offsets are relative to the first one, so slicing keeps the overlapped blocks' boundaries as they are and cuts the packed buffer between the first and last of them. For example, rows 1500..4100 of an array with block offsets [b0, b1, b2, b3, b4, b5] keep blocks 1 to 4:

  • block offsets: [b1, b2, b3, b4, b5], a zero-copy slice of the child;
  • packed buffer: bytes b1 - b0..b5 - b0, a zero-copy slice;
  • offset: 476, since row 1500 is 476 rows into block 1.

Only b0, b1 and b5 are read, with execute_scalar, so this is a SliceKernel like FastLanes RLE's, whose chunk offsets are also relative to the first. SliceReduce still declines blocked arrays, since it never reads buffers, and slice stays lazy until execution. Release builds don't validate boundaries on construction, so the kernel checks the three it reads against the packed length before slicing. Patches and validity are sliced as for a global width.

Benchmarks

Local divan benchmark, not committed: slice 1500 rows into 128Ki u32 values with block widths cycling 1 to 16 bits, then canonicalize. Apple M5 Max, median.

rows unblocked_kernel (global width 16) blocked_decode (no kernel) blocked_kernel (this PR)
1Ki 0.46 µs 10.7 µs 1.0 µs
16Ki 1.46 µs 10.7 µs 1.9 µs
64Ki 5.5 µs 10.6 µs 5.8 µs

The kernel is 5.6–10.7x faster than decoding first at 1Ki and 16Ki rows, and within 1.3x of a single width from 16Ki rows. The slice alone takes 250 ns, against 125 ns for a single width. Most of the remaining gap at 1Ki rows is the blocked decoder's fixed cost of executing and validating the block offsets.

Testing

  • Ranges inside one block, across blocks, block-aligned, to the end, and empty both at a block boundary and inside a block, with primitive and FoR-encoded block offsets.
  • The sliced array keeps only the overlapped blocks: its offset, packed length and block offsets.
  • Nulls and patches.
  • Slicing a sliced array.

Checks

  • cargo nextest run -p vortex-fastlanes (676 passed)
  • cargo clippy -p vortex-fastlanes --all-targets --all-features -- -D warnings
  • rustfmt +nightly-2026-09-10 on the changed file

Signed-off-by: Matt Katz <mhkatz97@gmail.com>
@mhk197 mhk197 added the changelog/performance A performance improvement label Oct 2, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/performance A performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant