Skip to content

Compute writer chunk hints through the aggregate cache - #10266

Closed
connortsui20 wants to merge 1 commit into
ct/stats-remaining-consumersfrom
ct/stats-writer-cache-computation
Closed

connortsui20 wants to merge 1 commit into
ct/stats-remaining-consumersfrom
ct/stats-writer-cache-computation

Conversation

@connortsui20

@connortsui20 connortsui20 commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Superseded by #10269, which groups this change with the related statistics work. The original commits and branch are preserved.

Original PR description

Tracking Issue: #10177

Summary

The file writer needs both a cached answer for each chunk and the partial state needed to combine chunks. A finalized answer cannot always supply that state: two individually sorted chunks can be out of order at their shared boundary. Publishing through a bound computation keeps both results tied to the input that produced them.

Changes

Reuse the writer's chunk accumulators through the aggregate cache API. Preserve fused extrema, cached Min/Max recovery, overflow state, and the existing omission of null chunk hints. For review, start with the bound computation contract, then its core accumulator dispatch, followed by the writer callers.

API Changes

Add AggregationsRef::compute_into. It resets and computes a matching core accumulator, caches an exact non-null final, and leaves the real partial state available for merging. compute_result continues to cache exact nulls.

Stack

Depends on #10265 in migration stack #10229.

Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
@connortsui20 connortsui20 added the changelog/feature A new feature label Oct 3, 2026
@connortsui20
connortsui20 added this pull request to stack #10229 October 3, 2026 02:24
@codspeed

codspeed Bot commented Oct 3, 2026

Copy link
Copy Markdown

Merging this PR will improve performance by 36.17%

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 3 improved benchmarks
✅ 2100 untouched benchmarks
⏩ 503 skipped benchmarks1

Performance Changes

Mode Benchmark BASE HEAD Efficiency
⚡ Simulation bench_compare_sliced_dict_primitive[(2500, 10000)] 83 µs 42.8 µs +94.17%
⚡ Simulation chunked_opt_bool_canonical_into[(1000, 10)] 69.9 µs 59.3 µs +17.81%
⚡ Simulation compress_fsst[(500, 4, 8)] 216.9 µs 196.5 µs +10.37%

Tip

Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.


Comparing ct/stats-writer-cache-computation (aef451a) with ct/stats-remaining-consumers (019c756)

Open in CodSpeed

Footnotes

  1. 503 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@connortsui20
connortsui20 removed this pull request from stack #10229 October 5, 2026 10:31
@connortsui20
connortsui20 deleted the ct/stats-writer-cache-computation branch October 5, 2026 10:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/feature A new feature

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant