Repository navigation
perf(buffer): use fearless_simd for x86 popcount - #10323
joseph-isaacs wants to merge 1 commit into
Conversation
Replace the hand-written AVX2 nibble-lookup and AVX-512 VPOPCNTDQ `count_ones` kernels with one fearless_simd kernel that popcounts u64 lanes of 64-byte vectors. The SIMD level is detected once and cached. Non-x86 targets and Miri keep the scalar path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YZRAJuGZmP4JRrmAbQQqZx Signed-off-by: Claude <noreply@anthropic.com>
Merging this PR will regress 1 benchmark
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | compare[1] |
114.3 µs | 135 µs | -15.32% |
| ⚡ | WallTime | mul_u64_nonnull_neon |
41.1 µs | 29.1 µs | +41.36% |
| ⚡ | WallTime | multiply_shapes_neon[(32768, PerRowPerRow)] |
39.3 µs | 33 µs | +19.24% |
| ⚡ | WallTime | mul_i64_nonnull_neon |
39.2 µs | 32.9 µs | +19.13% |
| ⚡ | Simulation | sum_v2_i64 |
224.5 µs | 193.8 µs | +15.81% |
| ⚡ | Simulation | sum_i64 |
224.2 µs | 193.8 µs | +15.71% |
| ⚡ | Simulation | slice_bounds[(4096, 0.1)] |
3.8 µs | 3.4 µs | +13.21% |
| ⚡ | Simulation | slice_bounds[(4096, 0.9)] |
3.9 µs | 3.4 µs | +13% |
| ⚡ | Simulation | rank_single[(1024, 0.1)] |
2.6 µs | 2.4 µs | +11.56% |
| ⚡ | Simulation | rank_single[(1024, 0.9)] |
2.6 µs | 2.4 µs | +11.53% |
| ⚡ | Simulation | rank_single[(16384, 0.9)] |
3.4 µs | 3.1 µs | +10.68% |
| ⚡ | Simulation | rank_single[(16384, 0.1)] |
4.2 µs | 3.8 µs | +10.25% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ji/fearless-simd-buffer (6b1366f) with develop (c6e51ba)
Footnotes
-
518 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
|
Closing: superseded by #10327, which merged this change with follow-ups (portable kernel without Generated by Claude Code |
Summary
Part of the fearless_simd migration (see #10319, #10320, #10321). This is one of three per-crate PRs, alongside the
vortex-arrayandvortex-layoutPRs, and contains only switches that benchmarked equal or faster. It replaces the hand-written x86count_oneskernels with a single fearless_simd kernel.Changes
vortex-buffer/src/bit/count_ones.rs: the AVX2 nibble-lookup kernel and the AVX-512 VPOPCNTDQ kernel become one kernel that popcountsu64lanes of 64-byte vectors. On Ice Lake-class AVX-512 this lowers tovpopcntq; on AVX2 it lowers to the same nibble lookup plusvpsadbw. The SIMDLevelis detected once and cached in aLazyLock, becauseLevel::new()probes every feature on each call.fearless_simd = "1.0.0"to the workspace. This is the same line as the other fearless PRs, so they merge trivially.Benchmarks
The host is a 4-vCPU Sapphire Rapids VM, so it is noisy; compare numbers within a run.
In-repo
cargo bench -p vortex-buffer --bench vortex_bitbuffer -- true_count_vortex_buffer, alternating runs:The second run is level. The first run has this PR slower (35% at 16,384 bits, 13% at 65,536); I read that as VM noise because the next run reversed it. A third alternating run would settle it.
Standalone kernel comparison (medians):
* Forced with
RUSTFLAGS="--cfg disable_dispatch_avx512", which emulates AVX2-only CPUs and Skylake-X. Skylake-X has no VPOPCNTDQ, so it already used the AVX2 kernel.Testing
cargo check -p vortex-buffer --all-targetsis clean. The existingcount_onestests cover offsets and lengths across the 32/64-byte boundaries.🤖 Generated with Claude Code
https://claude.ai/code/session_01YZRAJuGZmP4JRrmAbQQqZx
Generated by Claude Code