Repository navigation
[EXPERIMENTAL] perf: make checked lane chunking opt-in - #10304
connortsui20 wants to merge 3 commits into
Conversation
Merging this PR will improve performance by 38.95%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | WallTime | mul_i32_nonnull_avx2 |
32.7 µs | 18.4 µs | +78.1% |
| ⚡ | WallTime | mul_i32_nullable_avx2 |
33.9 µs | 19.5 µs | +73.65% |
| ⚡ | WallTime | mul_i16_nonnull_avx2 |
29.1 µs | 21.7 µs | +34.06% |
| ⚡ | WallTime | mul_i16_nonnull_avx512 |
16.5 µs | 14.7 µs | +12.55% |
| ⚡ | WallTime | mul_i16_nonnull_neon |
24 µs | 21.6 µs | +10.99% |
| Simulation | density_sweep_dense_runs[0.9] |
55.7 µs | < 1 ns | N/A | |
| Simulation | bench_compare_sliced_dict_primitive[(3333, 10000)] |
79.7 µs | < 1 ns | N/A |
Tip
Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ct/checked-lane-chunks-opt-in (b6ff4dc) with develop (76381d5)
Footnotes
-
518 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
2511a5e to
10a8626
Compare
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Summary
Alternative to #10299. Fixed-size failure reductions help LLVM combine signed multiplication overflow checks, but applying them to every checked operation can lose efficient unsigned multiply instructions and add repeated reductions.
Keep chunking opt-in so each operation can select the loop that suits its failure evidence. Enable it for
i16andi32multiplication, which improved across the x86 and ARM targets in the original investigation. Wide integer multiplication needs a separate policy because its results depend on the target ISA.Changes
Preserve the default checked visitor and keep its chunked counterpart separate. Sharing callback construction and dispatch between them makes LLVM lose the default ARM i64/u64 loop's two-lane form, even when chunking is disabled. Select chunking inside the primitive-type dispatcher to preserve codegen-unit placement. Nullable dense attempts support the same opt-in while retaining ordinary valid-row retries. Batch-constant inputs in the nullable dense attempt retain their existing loop.
API Changes
Adds
map_checked_chunked_into,RowVisitor::visit_deferred_chunked, and its prepared form. Chunked failure reduction requires an associative OR operation withDefaultas its identity. Existing entry points keep their behavior, and array, SQL, and Python APIs are unchanged.