Repository navigation
Add benchmarks for round tripping aggregate partials - #9916
Conversation
Adds a microbenchmark for the partial round-trip the zoned writer performs: accumulate one partial per zone, convert each to a scalar with `partial_scalar`, then fold those scalars back into a single accumulator with `combine_partials`. Each half is timed separately, plus the whole round-trip, over two partial shapes: a bloom filter, whose partial is a byte blob that grows with the filter size, and `Sum`, whose partial is a single value. Signed-off-by: Robert Kruszewski <robert@spiraldb.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YCjt1pTdVXr8dHar9ALtTY
`SumV2`'s partial is a `{sum, is_overflow, is_empty}` struct rather than a bare
scalar, so the round-trip covers struct scalar construction and field access
rather than a single primitive.
Signed-off-by: Robert Kruszewski <robert@spiraldb.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YCjt1pTdVXr8dHar9ALtTY
Drops the `partial_scalar` and round-trip benchmarks, leaving `combine_partials`, which reads a stored partial out of its scalar representation and merges it into an accumulator. Renames the target to match. Signed-off-by: Robert Kruszewski <robert@spiraldb.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YCjt1pTdVXr8dHar9ALtTY
Merging this PR will degrade performance by 12.82%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | arrow_checked_add_u32_neon[16384] |
12.9 µs | 20.4 µs | -36.95% |
| ⚡ | Simulation | allocate_drop_bytes[0] |
635.5 ns | 527.2 ns | +20.55% |
| 🆕 | Simulation | bloom[256] |
N/A | 149.8 µs | N/A |
| 🆕 | Simulation | sum_v2 |
N/A | 65.6 µs | N/A |
| 🆕 | Simulation | bloom[1024] |
N/A | 653.9 µs | N/A |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing claude/stoic-goodall-wprp78 (03cb988) with develop (f15c0b8)
Footnotes
-
218 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
Halves the zones per iteration and quarters the rows per zone, keeping every case well under a millisecond. Rows per zone only affect setup, so the merge itself is measured over the same partial shapes as before. Signed-off-by: Robert Kruszewski <robert@spiraldb.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YCjt1pTdVXr8dHar9ALtTY
At 8192 blocks the merge took 2.7ms under CodSpeed's instrumentation. A quarter of the blocks brings that under a millisecond, and a 64KiB filter still pushes the merge's working set out of L1. Signed-off-by: Robert Kruszewski <robert@spiraldb.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YCjt1pTdVXr8dHar9ALtTY
Doubles the zones merged per iteration and halves the large filter, keeping the slowest case where it was while giving the two cheap cases twice the work to measure. Rows per zone only change how densely the filters are populated, which a bitwise OR merge is indifferent to. Signed-off-by: Robert Kruszewski <robert@spiraldb.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YCjt1pTdVXr8dHar9ALtTY
This is missing in our benchmark coverage and is useful to check things like #9816