Skip to content

perf(array): pack Boolean output during RowFn dense retry - #9986

Merged
connortsui20 merged 1 commit into
developfrom
ct/row-fn-bool-retry
Sep 23, 2026
Merged

connortsui20 merged 1 commit into
developfrom
ct/row-fn-bool-retry

Conversation

@connortsui20

@connortsui20 connortsui20 commented Sep 22, 2026 •

Copy link
Copy Markdown
Member

Stacked on #9988, which adds the CI benchmark without changing production code.

Summary

Routes deferred Boolean dense attempts through the packed collector, preserving the requested multiversioning flag. Keeps the source and failure accumulator in one borrowed state, and inlines the collector's scalar tail so LLVM can retain the alias information needed to vectorize the row loop. Decode errors remain terminal, and null-only failures still retry valid rows.

The earlier x86 regressions came from scalarized collector loops. Optimized Linux IR and the native CI executables confirm that these loops now vectorize.

Changes

Native CI medians in microseconds per eight 16,384-row batches, with partial validity and the multiversioned collector:

Target Columns before → after Constant LHS before → after
AVX2 111.0 → 44.53 217.1 → 34.58
AVX-512 42.81 → 45.07 206.9 → 37.25
NEON 68.32 → 71.01 173.3 → 57.62

Constant RHS also improves on every target. Small column-only movements remain uncertain after one paired run. All final cases are below 1 ms, with a maximum printed sample of 555.8 µs.

Passed 175 RowFn tests, 14 Boolean packing tests, and scoped Clippy. The CodSpeed check flags no Boolean retry regressions, but remains red on three other, unattributed benchmark alerts. This remains a draft.

Four-stage comparison, compiler evidence, and raw results. Evidence stays on the separate branch.

@codspeed

codspeed Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Merging this PR will regress 2 benchmarks

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ 3 benchmarks measured no execution time

Nothing ran under measurement, usually because the compiler removed the code under test. These results are not comparable, so they count as unchanged.

Preventing compiler optimizations

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 15 improved benchmarks
❌ 2 regressed benchmarks
✅ 2215 untouched benchmarks
⏩ 329 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ Simulation take_fsl_random[128, 10] 32.7 µs 58.4 µs -44.05%
❌ WallTime filtered_sink_i64_avx2[OneNullInEight] 22 µs 26 µs -15.53%
⚡ WallTime deferred_bool_avx2[16384, (ConstantLhs, PartialAccepted)] 206.4 µs 33.9 µs ×6.1
⚡ WallTime deferred_bool_avx512[16384, (ConstantLhs, PartialAccepted)] 206.3 µs 38.3 µs ×5.4
⚡ Simulation take_fsl_u32_random[64, 100] 96 µs 30.3 µs ×3.2
⚡ WallTime deferred_bool_neon[16384, (ConstantLhs, PartialAccepted)] 169.8 µs 54.4 µs ×3.1
⚡ WallTime deferred_bool_avx2[16384, (Columns, PartialAccepted)] 110.4 µs 43 µs ×2.6
⚡ Simulation take_fsl_nullable_random[256, 10] 105.5 µs 49.1 µs ×2.1
⚡ Simulation take_fsl_f16_random[256, 100] 141.2 µs 87.8 µs +60.89%
⚡ WallTime deferred_bool_avx2[16384, (ConstantLhs, NullOnlyFailure)] 452.8 µs 282.5 µs +60.27%
⚡ WallTime deferred_bool_avx512[16384, (ConstantLhs, NullOnlyFailure)] 452.8 µs 303.7 µs +49.12%
⚡ WallTime decode_avx512[8192, (Inline, OneNullInEight)] 111 µs 83.3 µs +33.27%
⚡ WallTime deferred_bool_neon[16384, (ConstantLhs, NullOnlyFailure)] 479.5 µs 378.2 µs +26.77%
⚡ WallTime deferred_bool_avx2[16384, (Columns, NullOnlyFailure)] 344.2 µs 279.3 µs +23.27%
⚡ WallTime decode_avx512[8192, (External, OneNullInEight)] 175.4 µs 152.1 µs +15.31%
⚡ Simulation take_fsl_random[64, 100] 143.9 µs 125.3 µs +14.85%
⚡ Simulation take_filter_list_nullable_random_mask_random_indices[256, 50] 155.4 µs 138.1 µs +12.52%
⚠️ Simulation fixed_16_advancing_ptr_safe[100] < 1 ns < 1 ns N/A
⚠️ Simulation preverify_advancing_ptr_unchecked[1000] < 1 ns < 1 ns N/A
⚠️ Simulation preverify_advancing_ptr_unchecked[10000] < 1 ns < 1 ns N/A

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing ct/row-fn-bool-retry (994b7dd) with develop (2c8c7ee)

Open in CodSpeed

Footnotes

  1. 329 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@connortsui20
connortsui20 changed the base branch from develop to ct/row-fn-bool-bench September 22, 2026 18:34
@connortsui20
connortsui20 added this pull request to stack #9990 September 22, 2026 18:34
@connortsui20 connortsui20 added the changelog/performance A performance improvement label Sep 22, 2026
@connortsui20
connortsui20 force-pushed the ct/row-fn-bool-retry branch 2 times, most recently from 363bac7 to 65c67e1 Compare September 22, 2026 20:54
@connortsui20
connortsui20 force-pushed the ct/row-fn-bool-retry branch 2 times, most recently from 967c158 to fcb0a8b Compare September 23, 2026 02:26
@connortsui20
connortsui20 marked this pull request as ready for review September 23, 2026 02:27
Base automatically changed from ct/row-fn-bool-bench to develop September 23, 2026 09:27
myrrc pushed a commit that referenced this pull request Sep 23, 2026
## Summary

Adds a deferred Boolean RowFn benchmark to establish the CI baseline for
#9986. It uses the multiversioned collector over partially valid input
and covers four cases: an accepted dense attempt and a null-only retry,
each with two columns and with a constant left operand against a column.

## Changes

Uses the existing native AVX2, AVX-512, and NEON jobs, so there are 12
benchmarks in total. Each timed iteration executes eight 16,384-row
batches to clear the wall-time floor. Fixture and execution-context
setup and result destruction stay outside timing. Backtraces follow the
CI setting. This layer changes no production code.

An earlier revision ran 72 cases (both collector flags, constants in
either position, all-valid input, and observable failures). All of them
completed in native CI, with medians from 32.31 to 490.6 µs and a
maximum printed sample of 712.1 µs.

[Comparison with #9986 and raw
results](https://github.com/vortex-data/vortex/blob/ct/row-fn-performance-evidence/research/row-fn-engine/performance/ci-bool-capture.md)
come from that earlier revision.

---------

Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
@myrrc
myrrc force-pushed the ct/row-fn-bool-retry branch from 2446896 to 7d6bfa2 Compare September 23, 2026 09:28
@connortsui20
connortsui20 marked this pull request as draft September 23, 2026 14:18
`ExecuteDenseWithRetry` inherited the `visit_prepared_deferred_bool`
default, which forwards to the generic owned-output path. A deferred
Boolean kernel under the retry policy therefore collected one byte per
row and packed the bytes afterwards.

The visitor now implements the method and calls the new
`execute_bool_dense_attempt`, which packs bits during evaluation and
still separates row failure from terminal decode errors. That executor
keeps its own copy of the row loop, for the reason
`execute_owned_dense_attempt` records: a shared helper changes the
optimized dense kernel even when it inlines.

Two source shapes are load-bearing for code generation. The collector
captures one mutable borrow of `DeferredBoolState` rather than capturing
its four parts separately. The word loop in `vortex-buffer` packs its
tail with a private always-inlined copy of `collect_bool_word_scalar`,
so the public entry point and the scalar benchmark baselines keep their
own inlining behavior. Without either shape the multiversioned loop
stays scalar.

The new tests cover sliced validity, batch-constant inputs, word
boundary lengths, and a decode error that must stay terminal.

Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
@connortsui20
connortsui20 merged commit 6c2eb22 into develop Sep 23, 2026
97 of 100 checks passed
@connortsui20
connortsui20 deleted the ct/row-fn-bool-retry branch September 23, 2026 21:40
connortsui20 added a commit that referenced this pull request Sep 24, 2026
Depends on #10016.

RowFn output payloads bypass the execution allocator. This change
allocates them directly through `ctx.allocator()` across owned,
deferred-retry, selected, filtered, constant, and sink execution, while
preserving zero-copy primitive publication and empty-output paths.

`OutputElement` chooses its collection storage through an associated
buffer type and an allocation hook. The executor writes through
`OutputBuffer` slots, and the buffer implementation constructs the
array. Vortex primitive and Boolean implementations use `BufferMut`,
while scalar and fixed-size-list sinks use the same storage contract.
UTF-8 descriptors, external bytes, and polygon payloads also use the
execution allocator. Physical sink parameters remain separate from
allocation resources.

Regressions check ownership of returned payloads using canonical inputs
prepared before allocation tracking, including a context override,
constant UTF-8 output, retry execution, and zero-copy reuse. A
zero-sized output with `Vec` storage exercises the owned execution paths
without requiring `BufferMut`. Boolean collector selection is preserved.

The Boolean dense-retry path from merged #9986 also uses the execution
allocator, with coverage in the existing packed-output allocator test.

Validation before the rebase: 168 focused comparison, mask, and RowFn
tests passed on the combined stack through #9979, along with `cargo
clippy -p vortex-array --all-targets --all-features -- -D warnings`.
Tests, formatting, and benchmarks were not rerun locally after the
rebase.

---------

Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/performance A performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants