perf(array): pack Boolean output during RowFn dense retry - #9986
Conversation
5cdf87a to
2047abf
Compare
Merging this PR will regress 2 benchmarks
|
2047abf to
1312f0a
Compare
1312f0a to
31c9d94
Compare
363bac7 to
65c67e1
Compare
967c158 to
fcb0a8b
Compare
fcb0a8b to
2446896
Compare
## Summary Adds a deferred Boolean RowFn benchmark to establish the CI baseline for #9986. It uses the multiversioned collector over partially valid input and covers four cases: an accepted dense attempt and a null-only retry, each with two columns and with a constant left operand against a column. ## Changes Uses the existing native AVX2, AVX-512, and NEON jobs, so there are 12 benchmarks in total. Each timed iteration executes eight 16,384-row batches to clear the wall-time floor. Fixture and execution-context setup and result destruction stay outside timing. Backtraces follow the CI setting. This layer changes no production code. An earlier revision ran 72 cases (both collector flags, constants in either position, all-valid input, and observable failures). All of them completed in native CI, with medians from 32.31 to 490.6 µs and a maximum printed sample of 712.1 µs. [Comparison with #9986 and raw results](https://github.com/vortex-data/vortex/blob/ct/row-fn-performance-evidence/research/row-fn-engine/performance/ci-bool-capture.md) come from that earlier revision. --------- Signed-off-by: Connor Tsui <connor.tsui20@gmail.com> Signed-off-by: Claude <noreply@anthropic.com> Co-authored-by: Claude <noreply@anthropic.com>
2446896 to
7d6bfa2
Compare
7d6bfa2 to
bf57c5a
Compare
bf57c5a to
7ce6b1b
Compare
`ExecuteDenseWithRetry` inherited the `visit_prepared_deferred_bool` default, which forwards to the generic owned-output path. A deferred Boolean kernel under the retry policy therefore collected one byte per row and packed the bytes afterwards. The visitor now implements the method and calls the new `execute_bool_dense_attempt`, which packs bits during evaluation and still separates row failure from terminal decode errors. That executor keeps its own copy of the row loop, for the reason `execute_owned_dense_attempt` records: a shared helper changes the optimized dense kernel even when it inlines. Two source shapes are load-bearing for code generation. The collector captures one mutable borrow of `DeferredBoolState` rather than capturing its four parts separately. The word loop in `vortex-buffer` packs its tail with a private always-inlined copy of `collect_bool_word_scalar`, so the public entry point and the scalar benchmark baselines keep their own inlining behavior. Without either shape the multiversioned loop stays scalar. The new tests cover sliced validity, batch-constant inputs, word boundary lengths, and a decode error that must stay terminal. Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
7ce6b1b to
994b7dd
Compare
Depends on #10016. RowFn output payloads bypass the execution allocator. This change allocates them directly through `ctx.allocator()` across owned, deferred-retry, selected, filtered, constant, and sink execution, while preserving zero-copy primitive publication and empty-output paths. `OutputElement` chooses its collection storage through an associated buffer type and an allocation hook. The executor writes through `OutputBuffer` slots, and the buffer implementation constructs the array. Vortex primitive and Boolean implementations use `BufferMut`, while scalar and fixed-size-list sinks use the same storage contract. UTF-8 descriptors, external bytes, and polygon payloads also use the execution allocator. Physical sink parameters remain separate from allocation resources. Regressions check ownership of returned payloads using canonical inputs prepared before allocation tracking, including a context override, constant UTF-8 output, retry execution, and zero-copy reuse. A zero-sized output with `Vec` storage exercises the owned execution paths without requiring `BufferMut`. Boolean collector selection is preserved. The Boolean dense-retry path from merged #9986 also uses the execution allocator, with coverage in the existing packed-output allocator test. Validation before the rebase: 168 focused comparison, mask, and RowFn tests passed on the combined stack through #9979, along with `cargo clippy -p vortex-array --all-targets --all-features -- -D warnings`. Tests, formatting, and benchmarks were not rerun locally after the rebase. --------- Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Stacked on #9988, which adds the CI benchmark without changing production code.
Summary
Routes deferred Boolean dense attempts through the packed collector, preserving the requested multiversioning flag. Keeps the source and failure accumulator in one borrowed state, and inlines the collector's scalar tail so LLVM can retain the alias information needed to vectorize the row loop. Decode errors remain terminal, and null-only failures still retry valid rows.
The earlier x86 regressions came from scalarized collector loops. Optimized Linux IR and the native CI executables confirm that these loops now vectorize.
Changes
Native CI medians in microseconds per eight 16,384-row batches, with partial validity and the multiversioned collector:
Constant RHS also improves on every target. Small column-only movements remain uncertain after one paired run. All final cases are below 1 ms, with a maximum printed sample of 555.8 µs.
Passed 175 RowFn tests, 14 Boolean packing tests, and scoped Clippy. The CodSpeed check flags no Boolean retry regressions, but remains red on three other, unattributed benchmark alerts. This remains a draft.
Four-stage comparison, compiler evidence, and raw results. Evidence stays on the separate branch.