Skip to content

[codex] Experiment: Fearless SIMD boolean packing - #10320

Closed
joseph-isaacs wants to merge 1 commit into
developfrom
ji/fearless-simd-bool-packing
Closed

joseph-isaacs wants to merge 1 commit into
developfrom
ji/fearless-simd-bool-packing

Conversation

@joseph-isaacs

@joseph-isaacs joseph-isaacs commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Replace default and runtime-dispatched boolean packing with Fearless SIMD comparison masks and to_bitmask(). The baseline backend stays inline with arbitrary predicates, while the cheap-predicate path uses runtime dispatch. Miri retains scalar packing; explicit architecture helpers remain available for direct comparisons.

Changes

This isolates the boolean-packing migration originally classified as a loss. The earlier combined experiment measured a 48–51% regression against custom NEON, but the fresh standalone comparison below does not reproduce that magnitude: gather packing is roughly neutral (0–3% more time), and the 65,536-element public bool-slice conversion takes about 4% more time. Keep this as an investigation draft, not evidence of a repeatable large regression. Public-entry-point runs varied substantially, so small differences should not be treated as established wins or losses. Bitmap counting, rank-select, and filtering are unchanged.

Benchmark results

ARM64 macOS, Rust 1.98.0, standard optimized Cargo bench profile. Median of three per-run medians, alternating baseline/candidate order, with no concurrent builds from this task. Baseline: c6e51ba4de887d5795e2d1335fc30025dd4a1189. Ratios below 1 mean less time; above 1 mean a regression.

Benchmark Current develop This PR Time ratio
words_gather_dispatch/1048576 33660.00 ns 33700.00 ns 1.001×
words_gather_dispatch/65536 2041.00 ns 2103.00 ns 1.030×
words_gather_neon/1048576 33580.00 ns 33660.00 ns 1.002×
words_gather_neon/65536 2103.00 ns 2103.00 ns 1.000×
collect_bool_u32_gt/1024 361.50 ns 361.10 ns 0.999×
collect_bool_u32_gt/65536 20490.00 ns 19740.00 ns 0.963×
from_bool_slice/1024 43.22 ns 42.88 ns 0.992×
from_bool_slice/65536 2041.00 ns 2124.00 ns 1.041×
cargo bench -p vortex-buffer --bench collect_bool -- 'words_gather_(dispatch|neon)|collect_bool_u32_gt|from_bool_slice' --sample-count 300 --min-time 0.2

For added benchmarks, copy the same benchmark source and Cargo bench entry onto the baseline before building.

Validation

  • cargo nextest run --offline --release -p vortex-buffer: 881 passed.
  • cargo check --offline -p vortex-buffer --target x86_64-unknown-linux-gnu: passed (cross-compilation only).
  • cargo +nightly-2026-09-10 fmt -p vortex-buffer and git diff --check.
  • Repeated packing benchmarks against current develop; the explicit NEON helper is also measured in the candidate binary.
  • x86 execution and clippy were not run.

Related experiments

Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
@codspeed

codspeed Bot commented Oct 5, 2026

Copy link
Copy Markdown

Merging this PR will degrade performance by 30.93%

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 9 improved benchmarks
❌ 17 regressed benchmarks
✅ 2075 untouched benchmarks
⏩ 518 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ WallTime deferred_bool_avx2[16384, (ConstantLhs, PartialAccepted)] 33.5 µs 231.9 µs -85.54%
❌ WallTime deferred_bool_avx512[16384, (ConstantLhs, PartialAccepted)] 38.2 µs 232.2 µs -83.57%
❌ WallTime deferred_bool_avx2[16384, (Columns, PartialAccepted)] 42.6 µs 231.5 µs -81.59%
❌ WallTime deferred_bool_avx512[16384, (Columns, PartialAccepted)] 43.6 µs 231.9 µs -81.19%
❌ WallTime deferred_bool_neon[16384, (ConstantLhs, PartialAccepted)] 54.8 µs 181.6 µs -69.84%
❌ WallTime deferred_bool_neon[16384, (Columns, PartialAccepted)] 68.8 µs 174.3 µs -60.53%
❌ WallTime deferred_bool_avx512[16384, (ConstantLhs, NullOnlyFailure)] 285.7 µs 498.5 µs -42.69%
❌ WallTime deferred_bool_avx2[16384, (Columns, NullOnlyFailure)] 278.4 µs 467.9 µs -40.49%
❌ WallTime deferred_bool_avx2[16384, (ConstantLhs, NullOnlyFailure)] 297.4 µs 498.9 µs -40.39%
❌ WallTime deferred_bool_avx512[16384, (Columns, NullOnlyFailure)] 302 µs 467.6 µs -35.42%
❌ Simulation from_bool_slice[1024] 3.4 µs 5.2 µs -34.34%
❌ WallTime words_gather_dispatch_neon[65536] 2.1 µs 2.8 µs -27.16%
❌ WallTime deferred_bool_neon[16384, (ConstantLhs, NullOnlyFailure)] 375.5 µs 496.1 µs -24.3%
❌ WallTime words_gather_dispatch_neon[1048576] 29.8 µs 38.5 µs -22.52%
❌ WallTime deferred_bool_neon[16384, (Columns, NullOnlyFailure)] 381.1 µs 478.3 µs -20.32%
❌ Simulation compare[1] 114.3 µs 135.1 µs -15.4%
❌ WallTime words_gather_dispatch_avx512[65536] 967 ns 1,075 ns -10.05%
⚡ WallTime infallible_bool_avx2[i64, PerRowPerRow] 10.4 µs 5.6 µs +85.76%
⚡ WallTime infallible_bool_constant_avx2[PerRowConstant] 10.3 µs 5.6 µs +85.28%
⚡ WallTime deferred_bool_constant_avx2[PerRowConstant] 10.8 µs 6 µs +80.25%
... ... ... ... ... ...

ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing ji/fearless-simd-bool-packing (428719b) with develop (c6e51ba)

Open in CodSpeed

Footnotes

  1. 518 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

Copy link
Copy Markdown
Contributor Author

Closing: the benchmarks here are neutral to slightly worse (0–4% more time on the gather and bool-slice paths), so there's no win to justify the change. The result is recorded here for reference.


Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant