Repository navigation
Read RowFn UTF-8 byte lengths from native metadata - #10263
Closed
connortsui20 wants to merge 1 commit into
Closed
connortsui20 wants to merge 1 commit into
connortsui20 wants to merge 1 commit into
Conversation
connortsui20
added this pull request to stack #10264
October 3, 2026 02:09
connortsui20
force-pushed
the
ct/row-fn-utf8-lengths
branch
from
October 6, 2026 11:38
a2c8a07 to
1fba759
Compare
Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
connortsui20
force-pushed
the
ct/row-fn-utf8-lengths
branch
from
October 6, 2026 11:50
1fba759 to
54057ac
Compare
Member
Author
|
Closing along with #10262, which this builds on. Its replacement is #10340. The measured gain was about 10 µs per 65,536 rows for a length comparison, and it needed two new public unsafe input elements plus layout selection before dispatch. Length predicates already have a good path: Generated by Claude Code |
connortsui20
added a commit
that referenced
this pull request
Oct 7, 2026
## Summary Replaces #10262 and #10263, whose offset elements needed a layout choice before `RowFn::dispatch`, which only sees dtypes. Here `Utf8Column` picks its storage per batch, so every existing RowFn skips building a 16-byte view per row for `VarBin` input. ## Changes Offsets that do not fit in `u32`, or that delimit invalid UTF-8 at any row (null rows included), fall back to string views. The row accessors are now `#[inline]` because they are not generic, so row loops in other crates made one call per row, which is why `VarBinView` input also got faster. On the new `row_fn_utf8_input` bench, a `starts_with` RowFn over 16,384 rows goes from 227.5 µs to 34.1 µs on `VarBin` with 8-byte values and from 79.7 µs to 65.3 µs on `VarBinView` with 8-byte values. <details> <summary>Benchmark methodology and full results</summary> Baseline is `develop` at `659c7f7`. Candidate is this commit before its rebase onto `d9ad4cf` (`a51871e`). The two binaries ran alternately 7 times on a 4-core cloud VM with the bench profile. Each figure is the median of the 7 per-run divan medians. A later rerun of the rebased commit with only doc and comment changes matched these numbers. | Input | `develop` | This PR | |---|---|---| | `VarBin`, 8-byte values | 227.5 µs | 34.1 µs | | `VarBin`, 64-byte values | 252.9 µs | 47.9 µs | | `VarBinView`, 8-byte values | 79.7 µs | 65.3 µs | | `VarBinView`, 64-byte values | 120.4 µs | 113.1 µs | </details> ## API Changes `Utf8Column` now yields `&str`, and `Utf8View` is removed. Callers that only dereference the value compile unchanged, but `x == value.as_ref()` becomes ambiguous because `str` implements `AsRef` for several targets, so compare against `value` directly. A null row still yields an unspecified string, which offset storage now takes from the stored bytes instead of an empty string. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_018FefTpTuDky8bGPr5R9Loq Signed-off-by: Claude <noreply@anthropic.com> Co-authored-by: Claude <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Depends on #10262 for the shared validated offset storage.
Length predicates currently prepare UTF-8 strings or materialize a
u64byte-length array before comparison. Explicit offset-length and view-length input elements let a RowFn compare native headers directly into Boolean output without acquiring payload buffers or constructing strings.Changes
The offset element shares the validated offset domain from the preceding draft. The view element owns only initialized string-view headers. Null lengths are unspecified, and the executor propagates input validity. Both elements reject unsupported physical layouts and report fallible decoding, which preserves the existing valid-only policy for partially valid inputs.
This remains a draft pending investigation of the research prototype's reproducible dense x86 offset regressions and measurements of the exact API port. Matching hot decode instructions support a code-placement hypothesis, but do not establish the cause. No consumer is switched and no query gain is claimed.
Research evidence and the unresolved regression
On develop
9a08b82dc683cd3fff8dcde834c382141f425f48, the corrected private view-length prototype measured 13,074 ns/batch against 22,870 ns for native byte length followed by comparison on 65,536 dense view rows. The median paired ratio was 0.572, with five ratios from 0.569 to 0.580. The native path already avoids UTF-8 validation, so this gain includes eliminating its intermediate 524,288-byte length buffer. Device-buffer controls establish that the metadata binding does not request payload host reads.Against the first private offset prototype, the corrected version regressed from 46,864 to 60,721 ns/batch under AVX2 release at 65,536 rows, with median paired ratio 1.296. The validation instruction sequences matched after relocation normalization, while linked placement changed. That regression remains unresolved and blocks qualification.
Measurements used c7i.4xlarge, Rust 1.98.0, mimalloc, AVX2, five alternating process pairs, and 30 samples per process. Benchmark used 16 codegen units and release used one, with LTO disabled in both. CPU affinity and an offline SMT sibling controlled execution. These figures measure private research bindings with infallible decode, not this API port with its different nullable policy.