Skip to content

fix: decode CAST(binary AS string) JVM-compatibly instead of unsafe reinterpret#4763

Merged
andygrove merged 3 commits into
apache:mainfrom
andygrove:fix-4488-cast-binary-string-ub
Jul 15, 2026
Merged

fix: decode CAST(binary AS string) JVM-compatibly instead of unsafe reinterpret#4763
andygrove merged 3 commits into
apache:mainfrom
andygrove:fix-4488-cast-binary-string-ub

Conversation

@andygrove

@andygrove andygrove commented Jun 29, 2026

Copy link
Copy Markdown
Member

Which issue does this PR close?

Closes #4488.

Rationale for this change

Native CAST(<binary> AS string) ran through cast_binary_formatter, which used unsafe { String::from_utf8_unchecked(value.to_vec()) } to turn non-UTF-8 bytes into a Rust String. String carries a documented invariant that its bytes are valid UTF-8, so constructing one from arbitrary bytes is undefined behaviour, and an Arrow string array that holds invalid UTF-8 is unsound for the downstream native string kernels that read it via StringArray::value.

Spark's StringType tolerates arbitrary bytes while Arrow's Utf8 type requires valid UTF-8, so Comet cannot store the raw bytes natively. Instead, decode them the same way the JVM does.

What changes are included in this PR?

  • Native (cast.rs): cast_binary_formatter now decodes with decode_utf8_spark_lossy, which replaces each ill-formed sequence with U+FFFD exactly as the JVM's new String(bytes, UTF_8) does — including the surrogate-range cases (ED A0..BF) where Rust's str::from_utf8_lossy emits a different number of U+FFFD. The output is always valid UTF-8 and matches Spark's rendered result byte-for-byte, so the cast stays Compatible (native by default, no opt-in or gating). The ToPrettyString formatting path is unchanged.
  • Shared decoder: decode_utf8_spark_lossy already existed in the native shuffle (native shuffle: get_string should not panic on non-UTF-8 bytes (use lossy decode) #4521). Move it into the shared spark-expr utils module so the native shuffle and CAST(binary AS string) use a single implementation; the shuffle now imports it from there.
  • Docs: add a "Strings with non-UTF-8 bytes" section to the compatibility guide, a "Producing strings from arbitrary bytes" note to the contributor guide (adding_a_new_expression.md), and update the cast expression-audit note.

The only behavior that differs from Spark is byte-preserving round-trips (e.g. CAST(CAST(X'FF' AS STRING) AS BINARY)): Spark keeps the original bytes, Comet has the U+FFFD encoding. This is consistent with the native shuffle and is documented.

How are these changes tested?

  • Native: decode_utf8_spark_lossy JVM-parity tests (moved alongside the function) and a cast_binary_to_string test that exercises the surrogate-range case; cargo fmt and clippy clean.
  • JVM: full CometCastSuite (161 passed, 9 ignored) and all CometSqlFileTestSuite cast files pass on Spark 4.1, with no changes to the existing tests (they match because the decoder matches the JVM).

@andygrove
andygrove marked this pull request as draft June 29, 2026 21:28
@andygrove
andygrove marked this pull request as ready for review June 29, 2026 21:29
@andygrove
andygrove marked this pull request as draft June 29, 2026 21:33
@andygrove andygrove changed the title fix: make CAST(binary AS string) safe and gate it behind a dedicated opt-in fix: make CAST(binary AS string) safe and gate it behind a dedicated opt-in [WIP] Jun 29, 2026
…einterpret

Native CAST(binary AS string) built a Rust String from non-UTF-8 bytes via
from_utf8_unchecked, which is undefined behaviour (apache#4488), and an Arrow string
array holding invalid UTF-8 is unsound for the downstream native string kernels
that read it via StringArray::value.

Decode the bytes with decode_utf8_spark_lossy, which replaces ill-formed
sequences with U+FFFD exactly as the JVM's new String(bytes, UTF_8) does,
including the surrogate-range cases where Rust's from_utf8_lossy diverges. The
output is always valid UTF-8 and matches Spark's rendered result byte-for-byte,
so the cast stays Compatible (native, no opt-in). It diverges from Spark only
under byte-level round-trips such as CAST(CAST(x AS string) AS binary).

This decoder already existed in the native shuffle (apache#4521); move it into the
shared spark-expr utils module so shuffle and cast use one implementation.
Document the policy in the compatibility guide and the contributor guide.
@andygrove
andygrove force-pushed the fix-4488-cast-binary-string-ub branch from 3790a09 to 8e2dcd2 Compare June 29, 2026 22:12
@andygrove andygrove changed the title fix: make CAST(binary AS string) safe and gate it behind a dedicated opt-in [WIP] fix: decode CAST(binary AS string) JVM-compatibly instead of unsafe reinterpret Jun 29, 2026
@andygrove
andygrove marked this pull request as ready for review June 29, 2026 22:16

@mbutrovich mbutrovich left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a nice fix, @andygrove. Replacing from_utf8_unchecked with a decode pass removes real undefined behaviour, and folding the surrogate-range handling so the output matches new String(bytes, UTF_8) byte-for-byte is the right call for Spark parity. Consolidating the decoder into spark-expr utils so the shuffle and the cast share one implementation is a good cleanup too, and the JVM-oracle tests that moved with it make the parity intent clear.

I confirmed the existing cast BinaryType to StringType test feeds 8 random bytes per row through castTest, so the invalid-byte path is already exercised end-to-end against Spark and CI is green on it. The compatibility doc is accurate about the one observable divergence (byte-preserving round-trips).

A few suggestions below. The one I would most encourage a look at is the sibling BinaryOutputStyle::Utf8 formatter, since it is the same operation one branch over and currently panics on the exact input class this PR is about. The rest are a small performance idea and two minor notes.

Thanks for tackling this.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks like the same UTF8String.fromBytes operation as the cast path you just fixed, one formatter over. As far as I can tell it is reachable in production: with spark.sql.binaryOutputStyle=UTF8 (Spark 4.0+), a ToPrettyString / show() over a binary column whose bytes are not valid UTF-8 would land here and hit the unwrap(), which panics the executor rather than rendering like Spark. Spark's ToStringBase UTF8 BinaryFormatter goes through the same fromBytes plus replacement, so I think it has the same U+FFFD semantics as the plain cast.

Since decode_utf8_spark_lossy is now shared, would it make sense to route this arm through it as well so both binary-to-string formatters behave the same?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in bdf9491. You are right that this arm was reachable and panicking: with spark.sql.binaryOutputStyle=UTF8, ToPrettyString/show() over a binary column with non-UTF-8 bytes lands here, and String::from_utf8(..).unwrap() panics the executor. Spark renders it via UTF8String.fromBytes(bytes).toString, i.e. new String(bytes, UTF_8), so this now routes through the shared decode_utf8_spark_lossy and yields the same U+FFFD semantics as the plain cast. Added a cast_binary_to_string test with binary_output_style = Utf8 over the same invalid-byte inputs (including the surrogate-range [ED A0 80] collapse) to lock the behavior in.

// kernels, and matches Spark's rendered output byte-for-byte. It diverges from Spark only under
// byte-level round-trips such as CAST(CAST(x AS string) AS binary), where Spark still has the
// original bytes and Comet has the U+FFFD replacements.
decode_utf8_spark_lossy(value).into_owned()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Small performance thought: in the common all-valid-UTF-8 case the decoder hands back a borrowed Cow, then into_owned() allocates a fresh String and copies the bytes in, and the collect::<GenericStringArray>() in cast_binary_to_string immediately copies those same bytes a second time into the Arrow value buffer before dropping the String. That is one malloc plus one free per non-null row that does not buy anything. It is not a regression, the old code allocated too, but since this function is already being reworked it might be a good moment to capture the win.

One option that matches what cast_array_to_string already does at lines 641 and 703 is to build with a GenericStringBuilder<O> and append the &str straight from the Cow, so the valid path copies once instead of twice. Roughly:

fn cast_binary_to_string<O: OffsetSizeTrait>(
    array: &dyn Array,
    spark_cast_options: &SparkCastOptions,
) -> Result<ArrayRef, ArrowError> {
    let input = array
        .as_any()
        .downcast_ref::<GenericByteArray<GenericBinaryType<O>>>()
        .unwrap();

    let mut builder = GenericStringBuilder::<O>::new();
    for value in input.iter() {
        match value {
            Some(bytes) => match spark_cast_options.binary_output_style {
                Some(s) => builder.append_value(spark_binary_formatter(bytes, s)),
                None => builder.append_value(decode_utf8_spark_lossy(bytes).as_ref()),
            },
            None => builder.append_null(),
        }
    }
    Ok(Arc::new(builder.finish()))
}

That keeps the ToPrettyString styles exactly as they are (they still build owned strings, which is unavoidable) and only removes the throwaway allocation on the hot cast path.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Applied in bdf9491. cast_binary_to_string now builds with a GenericStringBuilder<O> and appends the value straight from the decoder, matching the cast_array_to_string pattern you pointed to. To keep the doc comment on the default-cast path I left cast_binary_formatter in place but changed it to return Cow<'_, str>, so the valid path borrows and append_value copies once into the Arrow buffer instead of allocating a throwaway String first. The ToPrettyString styles still build owned strings, unchanged.

Comment thread native/spark-expr/src/utils.rs Outdated
/// per-class malformed lengths below (E0/ED overlong & surrogate handling, F0/F4 range checks)
/// match the observable replacement behavior of the JDK UTF-8 decoder; they were determined from
/// observed `new String(bytes, UTF_8)` output, not by reviewing the OpenJDK source.
pub fn decode_utf8_spark_lossy(bytes: &[u8]) -> Cow<'_, str> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: decode_utf8_spark_lossy is a pure byte-to-string primitive with no expression dependencies, and its sibling JVM-bytes helper bytes_to_i128 already lives in datafusion-comet-common, which both the shuffle and spark-expr already depend on. Putting the decoder there too would avoid the shuffle crate reaching into the expression crate just for this, and would keep the two "render arbitrary JVM bytes" helpers side by side.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in bdf9491. decode_utf8_spark_lossy and its JVM-oracle tests now live in datafusion-comet-common next to bytes_to_i128, and the shuffle crate imports it from there instead of reaching into datafusion-comet-spark-expr.

Comment thread native/spark-expr/src/utils.rs Outdated
let n = bytes.len();
let mut out = String::with_capacity(n);
let mut i = 0;
while i < n {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The decoder is correct and the branch guards all earn their place (they also double as the proof that the three char::from_u32(cp).unwrap() calls cannot fire). Two tiny idiomatic touches:

  • The bounds checks use i + 1 >= n followed by raw bytes[i + 1] indexing. Using bytes.get(i + 1) folds the bounds check into the access and reads a bit more like idiomatic Rust.
  • The char::from_u32(cp).unwrap() calls are provably safe given the guards above them. A one-line // guards above keep cp a valid scalar would save the next reader from re-deriving that, or char::from_u32_unchecked with a SAFETY note if you want to shave the check.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added the // guards above keep cp a valid scalar note on each of the three char::from_u32(cp).unwrap() calls in bdf9491.

On the bytes.get(i + 1) suggestion, I left the explicit i + 1 >= n style as is. Several of those bounds checks do double duty: on a truncated lead they also drive the i = n skip-to-EOF behavior, so folding them into bytes.get() would only apply cleanly to a subset of branches. Converting just those would make the decoder mix two bounds-checking idioms, and for a delicate byte-for-byte parity function I would rather keep it uniform. Happy to revisit if you feel strongly.

@andygrove andygrove added this to the 1.0.0 milestone Jul 3, 2026
Addresses review feedback on the CAST(binary AS string) UB fix:

- BinaryOutputStyle::Utf8 previously rendered via
  String::from_utf8(..).unwrap(), panicking the executor on non-UTF-8
  input. With spark.sql.binaryOutputStyle=UTF8 (Spark 4.0+), a
  ToPrettyString/show() over a binary column with invalid bytes hit
  this. Decode lossily instead so invalid bytes become U+FFFD, matching
  Spark's new String(bytes, UTF_8).
- Build cast_binary_to_string with a GenericStringBuilder and append the
  decoder's borrowed Cow, so the common valid-UTF-8 path copies once
  instead of allocating a throwaway String and copying twice.
- Move decode_utf8_spark_lossy into datafusion-comet-common next to
  bytes_to_i128 so the shuffle crate no longer reaches into spark-expr.
- Document why the char::from_u32 unwraps in the decoder are infallible.

@parthchandra parthchandra left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One request for a test @andygrove otherwise this looks good.

// byte-level round-trips such as CAST(CAST(x AS string) AS binary), where Spark still has the
// original bytes and Comet has the U+FFFD replacements. Returning `Cow` lets the valid path
// borrow so the caller appends without an intermediate allocation.
decode_utf8_spark_lossy(value)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we add a test to verify that invalid characters have been replaced after this call? (Any future changes in the behavior of decode_utf8_spark_lossy will then be caught by the test.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added in fb126fd. There are Rust unit tests for this in cast.rs, but nothing pinned the behavior at the Spark level, so CometCastSuite now has cast BinaryType to StringType - invalid UTF-8 bytes are replaced: it runs deterministic invalid inputs (two invalid bytes, the surrogate-range ED A0 80 that collapses to a single U+FFFD where Rust's from_utf8_lossy would emit three, a truncated two-byte lead, plus a valid and an empty value) through castTest for Spark parity, then asserts the collected values equal both new String(bytes, UTF_8) (the JVM oracle) and the literal expected strings. A change in decode_utf8_spark_lossy now fails an end-to-end test, not just the unit tests.

// byte-level round-trips such as CAST(CAST(x AS string) AS binary), where Spark still has the
// original bytes and Comet has the U+FFFD replacements. Returning `Cow` lets the valid path
// borrow so the caller appends without an intermediate allocation.
decode_utf8_spark_lossy(value)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This lossy conversion collapses distinct Spark strings before downstream native operations. For example, if column b contains X'FF':

CAST(b AS STRING) = CAST(X'EFBFBD' AS STRING)

Spark returns false because CAST(binary AS string) preserves the raw bytes and UTF8String equality compares them directly. Comet will decode both values to U+FFFD and return true. The same issue can affect joins, grouping, ordering, and byte-based string operations such as contains.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You are right, and the docs in this PR understated the divergence: they only called out byte-level round-trips. Distinct ill-formed sequences all decode to the same U+FFFD, so CAST(b AS STRING) = CAST(X'EFBFBD' AS STRING) is false in Spark (UTF8String compares the raw bytes) and true in Comet, and the same collapse can affect joins, grouping, ordering, and byte-based functions such as contains. fb126fd documents that class explicitly in the compatibility guide, the cast expression-audit note, and the comment on cast_binary_formatter.

I have kept the cast Compatible rather than downgrading it, for two reasons:

  1. Comet already behaves this way independently of the cast. The native columnar shuffle decodes with the same decode_utf8_spark_lossy (native shuffle: get_string should not panic on non-UTF-8 bytes (use lossy decode) #4521), so an invalid-UTF-8 string that goes through a Comet shuffle already collapses to U+FFFD regardless of how it was produced. Making just this one cast fall back would not restore byte-exact string identity in a Comet plan.

  2. Falling back would trade this divergence for the UB the PR removes. If the cast runs on the JVM, Spark produces a raw-byte UTF8String, and that value then enters the native pipeline through the unchecked Arrow FFI import (from_ffi does not validate UTF-8), where any downstream native string kernel reading it via StringArray::value materializes an invalid &str. That is Gap B in [EPIC] Consistent handling of invalid UTF-8 in native StringType (ingress policy) #4764.

The real fix is a single ingress policy (decode, do not reinterpret, at every boundary), tracked by the epic #4764, which the compatibility guide now links from this section. Happy to be argued out of this if you think the value-identity divergence is severe enough to warrant gating the cast in the meantime.

bytes is undefined behaviour, and an Arrow string array that holds invalid UTF-8 is unsound for the
downstream kernels that read it via `StringArray::value` (which assumes the UTF-8 invariant).

Use `crate::utils::decode_utf8_spark_lossy`. It replaces each ill-formed sequence with `U+FFFD`

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

decode_utf8_spark_lossy was moved to datafusion-comet-common and is not exported from this crate’s utils module.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in fb126fd: the reference now points at datafusion_comet_common::decode_utf8_spark_lossy (native/common/src/utils.rs), which is where it moved in the previous round.

… in a Spark test

Addresses further review feedback on the CAST(binary AS string) UB fix:

- Add an end-to-end CometCastSuite test over known invalid byte sequences
  (two invalid bytes, the surrogate-range collapse, a truncated lead) that
  checks Comet against Spark and pins the U+FFFD output, so a change in
  decode_utf8_spark_lossy is caught at the Spark level.
- Document that decoding is not byte-preserving in a second way: distinct
  ill-formed byte sequences all decode to U+FFFD, so strings that differ in
  Spark can compare equal in Comet, which can affect equality, joins,
  grouping, ordering, and byte-based functions such as contains. Previously
  only the round-trip divergence was called out.
- Fix the contributor-guide reference to the decoder, which now lives in
  datafusion-comet-common rather than the spark-expr utils module.

@parthchandra parthchandra left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@andygrove
andygrove merged commit faf9445 into apache:main Jul 15, 2026
70 checks passed
@andygrove

Copy link
Copy Markdown
Member Author

Merged. Thanks @parthchandra @manuzhang

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] CAST(BinaryType AS StringType) uses unsafe from_utf8_unchecked (undefined behaviour)

4 participants