Skip to content

Introduce aggregate results and array caching - #10235

Draft
connortsui20 wants to merge 8 commits into
developfrom
ct/stats-array-cache
Draft

connortsui20 wants to merge 8 commits into
developfrom
ct/stats-array-cache

Conversation

@connortsui20

@connortsui20 connortsui20 commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Tracking Issue: #10177

Summary

The fixed statistics slots duplicate the aggregate API and cannot distinguish functions with different options. This introduces one aggregate-keyed store in the existing array metadata slot, so callers can reuse computed results while the legacy API is migrated.

Cached answers and streaming accumulator states have different contracts. An exact sortedness answer, for example, does not contain the endpoints needed to combine batches. Result-to-state conversion is therefore explicit, and representation-dependent results are excluded when facts transfer to another representation.

Changes

  • Add immutable result snapshots and an input-bound cache keyed by the aggregate function and its options.
  • Keep missing results, exact nulls, and bounds distinct.
  • Define which aggregates can recover partial states and which results survive representation changes.
  • Preserve array clone sharing and transfer snapshots during execution without retaining the source array.

For review, start with stats/results, then stats/aggregations, then the aggregate contracts and array integration.

API Changes

Adds AggregateResults and AggregationsRef for lookup, explicit computation, and snapshots. The legacy facade remains temporarily available for callers migrated in the next PR.

Stack

  1. This PR: aggregate cache and contracts.
  2. Migrate statistics consumers and file summaries to aggregates #10265: migrate callers and persisted summaries.
  3. Preserve aggregate facts from writers and array producers #10269: populate aggregate metadata from writers and producers.

Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Keep generic finalized results and legacy statistics in one cache so callers can migrate without maintaining two mutable stores. Recover typed partials through each aggregate contract, preserving exact null results and full option identity.

Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Historical count fields retain their u64 wire type even when current aggregate kernels decline an input dtype. Preserve those hints through the temporary facade and node serialization.

Release the cache write lock before dropping removed functions so custom destructors can reenter without deadlocking.

Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
@codspeed

codspeed Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Merging this PR will degrade performance by 7.7%

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 12 improved benchmarks
❌ 29 regressed benchmarks
✅ 2062 untouched benchmarks
⏩ 503 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ Simulation null_count_run_end[(10000, 1024, 0.01)] 6.6 µs 11 µs -39.92%
❌ Simulation null_count_run_end[(32000, 256, 0.01)] 6.6 µs 11 µs -39.92%
❌ Simulation null_count_run_end[(10000, 256, 0.01)] 6.6 µs 10.9 µs -39.12%
❌ Simulation null_count_run_end[(32000, 1024, 0.01)] 6.6 µs 10.9 µs -39.12%
❌ Simulation sparse_null_count 62 µs 89.6 µs -30.77%
❌ Simulation decompress[u8, (4000, 256)] 33.8 µs 46.8 µs -27.78%
❌ Simulation slice_dict_tight_loop[10000] 699.2 µs 901.9 µs -22.48%
❌ Simulation bench_compare_sliced_dict_primitive[(2000, 10000)] 85.6 µs 106.4 µs -19.59%
❌ Simulation slice_primitive_tight_loop[10000] 427.9 µs 529.7 µs -19.22%
❌ Simulation whole_array_sum_partially_valid 65.1 µs 80.1 µs -18.78%
❌ Simulation search_index_above_max_chunked 526.9 µs 636.2 µs -17.18%
❌ Simulation search_index_in_range_chunked 528.2 µs 637.5 µs -17.14%
❌ Simulation bitpacked_compress_u32 36.2 µs 43.3 µs -16.44%
❌ Simulation take_varbin 332 µs 396.4 µs -16.24%
❌ Simulation sequence_compress_u32 42.8 µs 51 µs -16.04%
❌ Simulation canonicalize_sparse_list[(512, 7, 4)] 532.3 µs 625.1 µs -14.84%
❌ Simulation canonicalize_sparse_list[(1024, 17, 8)] 458.2 µs 535 µs -14.36%
❌ Simulation decompress[u16, (4000, 256)] 50.8 µs 59.1 µs -14.09%
❌ Simulation bitpacked_decompress_u32 51.4 µs 59.6 µs -13.71%
❌ Simulation bench_compare_varbinview[(10000, 512)] 158.6 µs 181.9 µs -12.82%
... ... ... ... ... ...

ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing ct/stats-array-cache (1643e97) with ct/stats-result-scope (e457d5f)

Open in CodSpeed

Footnotes

  1. 503 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/break A breaking API change

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant