Repository navigation
[codex] Experiment: Fearless SIMD boolean packing - #10320
joseph-isaacs wants to merge 1 commit into
Conversation
Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
Merging this PR will degrade performance by 30.93%
|
|
Closing: the benchmarks here are neutral to slightly worse (0–4% more time on the gather and bool-slice paths), so there's no win to justify the change. The result is recorded here for reference. Generated by Claude Code |
Summary
Replace default and runtime-dispatched boolean packing with Fearless SIMD comparison masks and
to_bitmask(). The baseline backend stays inline with arbitrary predicates, while the cheap-predicate path uses runtime dispatch. Miri retains scalar packing; explicit architecture helpers remain available for direct comparisons.Changes
This isolates the boolean-packing migration originally classified as a loss. The earlier combined experiment measured a 48–51% regression against custom NEON, but the fresh standalone comparison below does not reproduce that magnitude: gather packing is roughly neutral (0–3% more time), and the 65,536-element public bool-slice conversion takes about 4% more time. Keep this as an investigation draft, not evidence of a repeatable large regression. Public-entry-point runs varied substantially, so small differences should not be treated as established wins or losses. Bitmap counting, rank-select, and filtering are unchanged.
Benchmark results
ARM64 macOS, Rust 1.98.0, standard optimized Cargo bench profile. Median of three per-run medians, alternating baseline/candidate order, with no concurrent builds from this task. Baseline:
c6e51ba4de887d5795e2d1335fc30025dd4a1189. Ratios below 1 mean less time; above 1 mean a regression.words_gather_dispatch/1048576words_gather_dispatch/65536words_gather_neon/1048576words_gather_neon/65536collect_bool_u32_gt/1024collect_bool_u32_gt/65536from_bool_slice/1024from_bool_slice/65536cargo bench -p vortex-buffer --bench collect_bool -- 'words_gather_(dispatch|neon)|collect_bool_u32_gt|from_bool_slice' --sample-count 300 --min-time 0.2For added benchmarks, copy the same benchmark source and Cargo bench entry onto the baseline before building.
Validation
cargo nextest run --offline --release -p vortex-buffer: 881 passed.cargo check --offline -p vortex-buffer --target x86_64-unknown-linux-gnu: passed (cross-compilation only).cargo +nightly-2026-09-10 fmt -p vortex-bufferandgit diff --check.develop; the explicit NEON helper is also measured in the candidate binary.Related experiments