Skip to content

[codex] Experiment: Fearless SIMD byte compression regression - #10321

Closed
joseph-isaacs wants to merge 1 commit into
developfrom
ji/fearless-simd-byte-compression
Closed

joseph-isaacs wants to merge 1 commit into
developfrom
ji/fearless-simd-byte-compression

Conversation

@joseph-isaacs

@joseph-isaacs joseph-isaacs commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Replace the ARM64 one-byte filter-compression kernel with a Fearless SIMD byte shuffle. It preserves the original eight-element chunks and density threshold, pads each load to the library's 16-byte vector width, and writes only eight bytes. The 16-, 32-, and 64-bit NEON kernels and all x86 paths remain unchanged.

Changes

This is an isolated performance-loss experiment. Fresh measurements show allocated output taking 13–15% more time; in-place output ranges from about 1% less to 4% more time, so the loss is not uniform across output modes. Add a direct byte-compression benchmark for allocated and in-place output, with scalar-output comparisons over eight bitmap offsets and six boundary/tail lengths before timing. No performance claim is made for wider element widths.

Benchmark results

ARM64 macOS, Rust 1.98.0, standard optimized Cargo bench profile. Median of three per-run medians, alternating baseline/candidate order, with no concurrent builds from this task. Baseline: c6e51ba4de887d5795e2d1335fc30025dd4a1189. Ratios below 1 mean less time; above 1 mean a regression.

Benchmark Current develop This PR Time ratio
allocated/u8/0.6 8040.00 ns 9082.00 ns 1.130×
allocated/u8/0.75 7999.00 ns 9208.00 ns 1.151×
in_place/u8/0.6 8249.00 ns 8582.00 ns 1.040×
in_place/u8/0.75 8332.00 ns 8249.00 ns 0.990×
cargo bench -p vortex-array --bench filter_u8_simd -- --sample-count 300 --min-time 0.2

For added benchmarks, copy the same benchmark source and Cargo bench entry onto the baseline before building.

Validation

  • Six existing SIMD compression tests passed through a standalone Rust test harness importing the production modules.
  • Full vortex-array nextest execution was not completed locally: compilation of the large test executable was canceled after prolonged builds; the narrower kernel tests above completed.
  • Direct benchmark startup validated 96 scalar-versus-SIMD comparisons per invocation (8 offsets × 6 lengths × 2 output modes).
  • cargo +nightly-2026-09-10 fmt -p vortex-array and git diff --check.
  • Repeated direct byte-compression benchmarks against current develop.
  • x86 execution and clippy were not run.

Related experiments

Signed-off-by: Joe Isaacs <joe.isaacs@live.co.uk>
@codspeed

codspeed Bot commented Oct 5, 2026

Copy link
Copy Markdown

Merging this PR will degrade performance by 0.93%

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 1 improved benchmark
❌ 1 regressed benchmark
✅ 2099 untouched benchmarks
🆕 4 new benchmarks
⏩ 518 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ Simulation compare[1] 114.3 µs 135.1 µs -15.4%
⚡ Simulation compare[1] 155.6 µs 134.1 µs +16.02%
🆕 Simulation allocated[u8, 0.6] N/A 170.9 µs N/A
🆕 Simulation allocated[u8, 0.75] N/A 179.2 µs N/A
🆕 Simulation in_place[u8, 0.6] N/A 137.9 µs N/A
🆕 Simulation in_place[u8, 0.75] N/A 136.1 µs N/A

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing ji/fearless-simd-byte-compression (f96c310) with develop (c6e51ba)

Open in CodSpeed

Footnotes

  1. 518 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

Copy link
Copy Markdown
Contributor Author

Closing: this experiment's own measurements show a regression (allocated u8 compress takes 13–15% more time on ARM64), so the native NEON kernel stays. The result is recorded here for reference.


Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant