Repository navigation
Conversation
Merging this PR will improve performance by 10.22%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | WallTime | words_gather_dispatch_neon[65536] |
2.3 µs | 2 µs | +10.22% |
Tip
Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.
Comparing mk/bitpacked-variable-encode-shared (bb0d98a) with mk/bitpacked-variable-encode (fbf3501)2
Footnotes
-
518 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
No successful run was found on
mk/bitpacked-variable-encode(f55497a) during the generation of this report, so f915a0d was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
de28cfc to
fbf3501
Compare
edacfb4 to
eb38aaa
Compare
Codecov Report❌ Patch coverage is
☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
fbf3501 to
f55497a
Compare
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
Signed-off-by: Matt Katz <mhkatz97@gmail.com>
eb38aaa to
bb0d98a
Compare
f55497a to
977cca1
Compare
Tracking Issue: #10167
Stacked on #10203; this PR targets
mk/bitpacked-variable-encode.Summary
Share the packing loop and patch gathering between the global-width and blocked encoders. They differed only in where each 1024-value block's bit width comes from.
bitpack_blocks(array, capacity, bit_width)is the existingbitpack_primitiveloop, taking the width as a closure.gather_patches_with(parray, bit_width, ..)is the existinggather_patchesbody, with the same change.Both live in
bitpack_compress/mod.rsnext toensure_non_negative_integers, the other helper the two encoders share.bitpack_primitiveandgather_patcheskeep their signatures and passmove |_| bit_width. The blocked encoder passes|block| bit_widths[block]and drops its own copies. The encoder is 63 lines shorter overall.The bit-width histogram stays duplicated. Sharing it left the loop out of line under
#[inline], which adds a call per array, and#[inline(always)]is denied by clippy.Codegen
Each caller gets its own copy of the shared code, compiled with its own width closure. The global copy sees a constant width, so it compiles like the original loop. Comparing release assembly (
codegen-units=1) before and after:gather_patches_implcopy is identical. The width closures aremove; capturing the width by reference added a load.bitpack_uncheckedis 97% the same once register names are normalized. The differences are instruction order and block layout around the per-chunkunchecked_packcalls.Results
Median walltime over 10 interleaved runs of the before and after bench binaries, 100 iterations per sample for
bitpacked_compress_u32and 20 forbitpack_blocked_compress:single_encoding_throughput::bitpacked_compress_u32bitpack_blocked::bitpack_blocked_compressalp_compress::compress_rdf64, 6 cases(2000, 0.1)(2000, 0.1)has measured −5.4%, −3.8% and +3.9% across runs.bitpacked_compress_u32andcompress_rdare the benches CodSpeed flagged on an earlier attempt at this sharing. That attempt zeroed a padding block on every call and passed the width as&dyn Fn.