Repository navigation
Blocked BitPacked: store block offsets as a child - #10004
Conversation
Merging this PR will degrade performance by 41.92%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | take[small_m/shuffled/primitive/nonnull/chunks=2048/indices=64] |
279.4 µs | 481 µs | -41.92% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing mk/bitpacked-stack-03-width-child (9c01ab9) with develop (9a17a1d)
Footnotes
-
476 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
1 benchmark was run, but is now archived. If it was deleted in another branch, consider rebasing to remove it from the report. Instead if it was added back, click here to restore it. ↩
6dc1734 to
9dc388a
Compare
9dc388a to
505f31f
Compare
BitPacked: store bitpacked block offsets as a child
BitPacked: store bitpacked block offsets as a childBitPacked: store block offsets as a child
|
|
||
| /// Check that each block between `boundaries` is a whole number of 128-byte rows of at most | ||
| /// `max_bit_width` bits, and that the boundaries span `packed_len` bytes. | ||
| fn validate_primitive_offsets<T: Copy + Display>( |
There was a problem hiding this comment.
this is very expensive. It is necessary to do a full pass of blocked offsets to make sure the bitpacked array is valid unfortunately, since it depends on array values now as well as metadata
There was a problem hiding this comment.
Additionally, I expect this check to rarely run since the block offsets will be compressed
0304db0 to
5f3f2c3
Compare
| /// Every block is packed at this bit width. | ||
| Global(u8), | ||
| /// Byte boundaries of the packed blocks, from which each block's bit width is derived. | ||
| Blocked(ArrayRef), |
There was a problem hiding this comment.
make it a view?
| Blocked(ArrayRef), | |
| Blocked(&'a ArrayRef), |
There was a problem hiding this comment.
or have a viewed version
There was a problem hiding this comment.
I'll make this into BitWidthsView enum and add an owned BitWidths enum for when we need to do into_parts
| fn bit_widths(&self) -> BitWidths { | ||
| match (self.global_bit_width, self.block_offsets()) { | ||
| (Some(bit_width), None) => BitWidths::Global(bit_width), | ||
| (None, Some(block_offsets)) => BitWidths::Blocked(block_offsets.clone()), |
There was a problem hiding this comment.
So we don't need this clone
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
5f3f2c3 to
ad8ba8c
Compare
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Tracking Issue: #10167
Summary
Add an optional
block_offsetschild toBitPackedArraythat stores the byte boundaries of independently sized packed blocks, and make the shared bit width inBitPackedDataoptional.An array has exactly one of the two: a global bit width, as today, or block offsets.
Design
Bit widths
BitPackedData::global_bit_widthisSome(width)when every block shares a width, andNonewhenblock_offsetsisSome. Validation requires exactly one of them.BitPackedArrayExt::bit_widths()replacesbit_width():Paths that need one width (decoding,
scalar_at, v1 serialization and CUDA) return an error for blocked arrays until the decoder follow-ups.There is a viewed and owned version of this.
Block offsets
block_offsetsis a non-nullable unsigned integer array withnum_blocks + 1entries, including a trailing end boundary.Boundaries are relative to the first entry. Block
ioccupies bytesblock_offsets[i] - block_offsets[0]toblock_offsets[i + 1] - block_offsets[0]of the packed buffer, at(block_offsets[i + 1] - block_offsets[i]) / 128bits, because a 1024-value block at widthwtakes128 * wbytes.For example, three blocks at widths 3, 0 and 7 have offsets
[0, 384, 384, 1280]. Offsets with equal steps are still blocked; nothing converts them to a global width.Validation
ptype.bit_width()(previously at most 64), and the packed buffer must holdnum_blocks * 128 * widthbytes, as before.num_blocks + 1entries. Host-resident primitive offsets are also checked boundary by boundary: they must be nondecreasing, each block must be a whole number of 128-byte rows at mostptype.bit_width()bits wide, and together they must span exactly the packed buffer. Other encodings and device-resident offsets only get the dtype and length checks.Scope
Nothing in this PR produces block offsets, but arrays built with
try_new_with_block_offsetscan be constructed and inspected. Kernels decline them, and decoding, scalar access, serialization and CUDA return an error.Wire format and compatibility
fastlanes.bitpackedis unchanged. A global-width array has no block offsets child, so it serializes exactly as before, and the reader builds a global-width array from the metadatabit_width. Serializing a blocked array with this plugin returns an error. Compression goldens are unchanged.API changes
BitPackedArrayExt::bit_width()andBitPackedData::bit_width()are replaced byBitPackedArrayExt::bit_widths() -> BitWidths.BitPackedDataParts::bit_width: u8is replaced bybit_widths: BitWidths.BitPacked::try_new_with_block_offsetsconstructs an array from packed data and block boundaries.BitPacked::try_newandBitPackedData::try_neware unchanged.BitUnpackedChunks::try_newand bothunpacked_chunksaccessors return an error for blocked arrays.bitpack_decompress::unpack_singlereturnsVortexResult<Scalar>.max_packed_value()accessor is removed.FastLanesBitPackedArray.bit_widthproperty returnsint | None.