Skip to content

perf(buffer): use fearless_simd for x86 popcount - #10323

Closed
joseph-isaacs wants to merge 1 commit into
developfrom
ji/fearless-simd-buffer
Closed

joseph-isaacs wants to merge 1 commit into
developfrom
ji/fearless-simd-buffer

Conversation

@joseph-isaacs

Copy link
Copy Markdown
Contributor

Summary

Part of the fearless_simd migration (see #10319, #10320, #10321). This is one of three per-crate PRs, alongside the vortex-array and vortex-layout PRs, and contains only switches that benchmarked equal or faster. It replaces the hand-written x86 count_ones kernels with a single fearless_simd kernel.

Changes

  • vortex-buffer/src/bit/count_ones.rs: the AVX2 nibble-lookup kernel and the AVX-512 VPOPCNTDQ kernel become one kernel that popcounts u64 lanes of 64-byte vectors. On Ice Lake-class AVX-512 this lowers to vpopcntq; on AVX2 it lowers to the same nibble lookup plus vpsadbw. The SIMD Level is detected once and cached in a LazyLock, because Level::new() probes every feature on each call.
  • Inputs under 32 bytes, non-x86 targets and Miri keep the scalar path, as before.
  • Adds fearless_simd = "1.0.0" to the workspace. This is the same line as the other fearless PRs, so they merge trivially.

Benchmarks

The host is a 4-vCPU Sapphire Rapids VM, so it is noisy; compare numbers within a run.

In-repo cargo bench -p vortex-buffer --bench vortex_bitbuffer -- true_count_vortex_buffer, alternating runs:

Bits develop (run 1 / run 2) This PR (run 1 / run 2)
16,384 14.85 / 15.06 ns 20.06 / 14.59 ns
65,536 53.86 / 51.95 ns 61.01 / 51.54 ns

The second run is level. The first run has this PR slower (35% at 16,384 bits, 13% at 65,536); I read that as VM noise because the next run reversed it. A third alternating run would settle it.

Standalone kernel comparison (medians):

Bytes Native VPOPCNTDQ Fearless at AVX-512 level Native AVX2 Fearless at AVX2 level*
8 KiB 91 ns 83 ns 396 ns 395 ns
128 KiB 2.29 µs 2.19 µs 6.72 µs 7.35 µs

* Forced with RUSTFLAGS="--cfg disable_dispatch_avx512", which emulates AVX2-only CPUs and Skylake-X. Skylake-X has no VPOPCNTDQ, so it already used the AVX2 kernel.

Testing

  • cargo check -p vortex-buffer --all-targets is clean. The existing count_ones tests cover offsets and lengths across the 32/64-byte boundaries.
  • In the standalone harness, results matched the scalar reference for all lengths 0–300, 4,096 and 65,537 on x86, and on aarch64 under qemu.
  • Clippy, fmt and the full test suite were not run locally.

🤖 Generated with Claude Code

https://claude.ai/code/session_01YZRAJuGZmP4JRrmAbQQqZx


Generated by Claude Code

Replace the hand-written AVX2 nibble-lookup and AVX-512 VPOPCNTDQ
`count_ones` kernels with one fearless_simd kernel that popcounts u64
lanes of 64-byte vectors. The SIMD level is detected once and cached.
Non-x86 targets and Miri keep the scalar path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YZRAJuGZmP4JRrmAbQQqZx
Signed-off-by: Claude <noreply@anthropic.com>
@codspeed

codspeed Bot commented Oct 6, 2026

Copy link
Copy Markdown

Merging this PR will regress 1 benchmark

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 11 improved benchmarks
❌ 1 regressed benchmark
✅ 2089 untouched benchmarks
⏩ 518 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ Simulation compare[1] 114.3 µs 135 µs -15.32%
⚡ WallTime mul_u64_nonnull_neon 41.1 µs 29.1 µs +41.36%
⚡ WallTime multiply_shapes_neon[(32768, PerRowPerRow)] 39.3 µs 33 µs +19.24%
⚡ WallTime mul_i64_nonnull_neon 39.2 µs 32.9 µs +19.13%
⚡ Simulation sum_v2_i64 224.5 µs 193.8 µs +15.81%
⚡ Simulation sum_i64 224.2 µs 193.8 µs +15.71%
⚡ Simulation slice_bounds[(4096, 0.1)] 3.8 µs 3.4 µs +13.21%
⚡ Simulation slice_bounds[(4096, 0.9)] 3.9 µs 3.4 µs +13%
⚡ Simulation rank_single[(1024, 0.1)] 2.6 µs 2.4 µs +11.56%
⚡ Simulation rank_single[(1024, 0.9)] 2.6 µs 2.4 µs +11.53%
⚡ Simulation rank_single[(16384, 0.9)] 3.4 µs 3.1 µs +10.68%
⚡ Simulation rank_single[(16384, 0.1)] 4.2 µs 3.8 µs +10.25%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing ji/fearless-simd-buffer (6b1366f) with develop (c6e51ba)

Open in CodSpeed

Footnotes

  1. 518 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

Copy link
Copy Markdown
Contributor Author

Closing: superseded by #10327, which merged this change with follow-ups (portable kernel without cfg gating, #[simd], no extra LazyLock, multi-chunk test).


Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants