perf(compute): tile checked lane kernels consistently - #10299
connortsui20 wants to merge 3 commits into
Conversation
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Merging this PR will regress 9 benchmarks
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | mul_u16_nonnull_avx2 |
11.2 µs | 20.8 µs | -45.9% |
| ❌ | WallTime | mul_u16_nonnull_avx512 |
11.3 µs | 19.5 µs | -42.22% |
| ❌ | WallTime | mul_u8_nonnull_avx512 |
22.5 µs | 29.7 µs | -24.41% |
| ❌ | WallTime | mul_u32_nonnull_avx2 |
12.8 µs | 16.9 µs | -24.07% |
| ❌ | WallTime | mul_u8_nonnull_avx2 |
20.5 µs | 26 µs | -20.94% |
| ❌ | WallTime | add_u32_nonnull_avx2 |
13 µs | 15.4 µs | -15.86% |
| ❌ | WallTime | mul_u32_nonnull_avx512 |
12 µs | 13.8 µs | -13.33% |
| ❌ | WallTime | add_u32_nonnull_neon |
13.5 µs | 15.5 µs | -12.93% |
| ❌ | WallTime | add_i32_nonnull_avx512 |
12 µs | 13.5 µs | -10.9% |
| ⚡ | WallTime | mul_u64_nonnull_avx512 |
52.7 µs | 29 µs | +81.43% |
| ⚡ | WallTime | mul_i32_nonnull_avx2 |
32.7 µs | 18.1 µs | +80.4% |
| ⚡ | WallTime | mul_i32_nullable_avx2 |
33.8 µs | 19.3 µs | +74.94% |
| ⚡ | WallTime | mul_u64_nonnull_neon |
39.4 µs | 24.7 µs | +59.27% |
| ⚡ | WallTime | mul_i16_nonnull_avx2 |
29.2 µs | 21.5 µs | +35.84% |
| ⚡ | WallTime | multiply_shapes_neon[(32768, PerRowPerRow)] |
38.9 µs | 30.2 µs | +29.02% |
| ⚡ | WallTime | mul_i64_nonnull_neon |
38.6 µs | 30.2 µs | +27.71% |
| ⚡ | WallTime | mul_i8_nonnull_avx2 |
29.9 µs | 24.7 µs | +20.89% |
| ⚡ | WallTime | add_constant_shapes_avx2[(32768, PerRowConstant)] |
12.3 µs | 10.9 µs | +12.67% |
| ⚡ | WallTime | add_constant_shapes_avx2[(32768, ConstantPerRow)] |
12.3 µs | 10.9 µs | +12.39% |
| ⚡ | WallTime | mul_i8_nonnull_neon |
25.6 µs | 22.8 µs | +12.19% |
| ... | ... | ... | ... | ... | ... |
ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ct/checked-lane-chunks (47db41a) with develop (80aaa32)2
Footnotes
-
503 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
No successful run was found on
develop(23a59e3) during the generation of this report, so 80aaa32 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
This reverts commit 09bc49a. Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Summary
Adding Narrow changed LLVM's unrolling of the existing checked 64-bit multiplication loop on Linux ARM, even though the loop's source was unchanged. Chunking
map_checked_intogives LLVM a fixed trip count for the checked loop, as the other unmasked lane kernels already do.Known limitation: chunking also regresses other benchmarks, particularly unsigned multiplication on AVX2 and AVX512. Making the chunk length const generic and removing forced inlining made the regressions worse, so that experiment was reverted.
Changes
Uses the same chunk size as the other unmasked lane kernels, with a runtime count and an
#[inline(always)]helper. Failure evidence combines across full chunks and the remainder.