Repository navigation
Preserve overflow when combining file sums - #10228
connortsui20 wants to merge 4 commits into
Conversation
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Merging this PR will improve performance by 25.42%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | WallTime | mul_u64_nonnull_neon |
40.2 µs | 28.9 µs | +38.9% |
| ⚡ | WallTime | mul_i64_nonnull_neon |
39.1 µs | 32.7 µs | +19.51% |
| ⚡ | WallTime | multiply_shapes_neon[(32768, PerRowPerRow)] |
38.9 µs | 32.7 µs | +18.86% |
Tip
Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ct/stats-streaming-summaries (47e17ed) with develop (23a59e3)2
Footnotes
-
503 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
No successful run was found on
develop(bf82f23) during the generation of this report, so 23a59e3 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩
|
Closing in favor of #10328. It fixes the same bug with a small guard in Generated by Claude Code |
## Summary - Tracking Issue: #10177 The file writer computes the file `Stat::Sum` in two passes. `push_chunk` stores each chunk's sum in a nullable column, and `as_stats_set` sums that column again. Both passes use the legacy `Sum` aggregate, which returns zero for empty input and null only on overflow. The second pass skips nulls as if they were missing values, so the footer could report an exact sum of only the chunks that did not overflow. For chunks `[i64::MAX, 1]` and `[2]`, the footer reported `2`. This happens with real data when some chunks overflow and others don't. Examples: microsecond timestamps stored as plain `i64`, where full batches overflow but a short last batch does not, or `i64::MAX` / `u64::MAX` sentinel values that only appear in some chunks. This replaces #10228 with a much smaller fix. ## Changes - In `StatsAccumulator::as_stats_set`, leave `Sum` absent when any chunk sum is null. This follows the existing guard for truncated varlen `Max` right above it. - Adds an rstest unit test in `vortex-layout` (one chunk overflows, all chunks overflow, the total overflows, nullable values, an all-null chunk) and an end-to-end writer test in `vortex-file` that checks both the returned footer and the footer read back from the file. This is a stopgap for the legacy `Stat::Sum` path only. When file statistics move to aggregate function partials (for example `SumV2`, which tracks overflow in an explicit `is_overflow` field), chunk results are merged as partials instead of summed as values, and this guard can be deleted along with `Stat::Sum`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_014jTJihcDfBHCS9AQcQHLjR --------- Signed-off-by: Claude <noreply@anthropic.com> Co-authored-by: Claude <noreply@anthropic.com>
Tracking Issue: #10177
Summary
The file writer sums finalized chunk results through a nullable intermediate column. If one chunk overflows, its null sum is skipped, so a later chunk can turn an overflowed file sum into an incorrect exact total. For chunks
[i64::MAX, 1]and[2], the footer currently reports2.Changes
Merge each chunk's accumulator state before finalizing the file sum, preserving overflow and applying the original input's decimal precision throughout. Keep the existing floating-point chunk grouping, NaN-skipping policy, and chunk-cache publication. The legacy footer API represents an overflowed sum as absent. This fix can be reviewed and merged independently of the statistics API migration.
Related Work
The migration stack starts at #10226 and continues with #10227. This fix targets
developindependently.