Repository navigation
[codex] Experiment: Fearless SIMD byte compression regression - #10321
joseph-isaacs wants to merge 1 commit into
Conversation
Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
Merging this PR will degrade performance by 0.93%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | compare[1] |
114.3 µs | 135.1 µs | -15.4% |
| ⚡ | Simulation | compare[1] |
155.6 µs | 134.1 µs | +16.02% |
| 🆕 | Simulation | allocated[u8, 0.6] |
N/A | 170.9 µs | N/A |
| 🆕 | Simulation | allocated[u8, 0.75] |
N/A | 179.2 µs | N/A |
| 🆕 | Simulation | in_place[u8, 0.6] |
N/A | 137.9 µs | N/A |
| 🆕 | Simulation | in_place[u8, 0.75] |
N/A | 136.1 µs | N/A |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ji/fearless-simd-byte-compression (f96c310) with develop (c6e51ba)
Footnotes
-
518 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
|
Closing: this experiment's own measurements show a regression (allocated u8 compress takes 13–15% more time on ARM64), so the native NEON kernel stays. The result is recorded here for reference. Generated by Claude Code |
Summary
Replace the ARM64 one-byte filter-compression kernel with a Fearless SIMD byte shuffle. It preserves the original eight-element chunks and density threshold, pads each load to the library's 16-byte vector width, and writes only eight bytes. The 16-, 32-, and 64-bit NEON kernels and all x86 paths remain unchanged.
Changes
This is an isolated performance-loss experiment. Fresh measurements show allocated output taking 13–15% more time; in-place output ranges from about 1% less to 4% more time, so the loss is not uniform across output modes. Add a direct byte-compression benchmark for allocated and in-place output, with scalar-output comparisons over eight bitmap offsets and six boundary/tail lengths before timing. No performance claim is made for wider element widths.
Benchmark results
ARM64 macOS, Rust 1.98.0, standard optimized Cargo bench profile. Median of three per-run medians, alternating baseline/candidate order, with no concurrent builds from this task. Baseline:
c6e51ba4de887d5795e2d1335fc30025dd4a1189. Ratios below 1 mean less time; above 1 mean a regression.allocated/u8/0.6allocated/u8/0.75in_place/u8/0.6in_place/u8/0.75For added benchmarks, copy the same benchmark source and Cargo bench entry onto the baseline before building.
Validation
vortex-arraynextest execution was not completed locally: compilation of the large test executable was canceled after prolonged builds; the narrower kernel tests above completed.cargo +nightly-2026-09-10 fmt -p vortex-arrayandgit diff --check.develop.Related experiments