You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
On that note, should we review the pyarrow Schema to Iceberg Schema type mappings within the repository and ensure that all types that are supported in the existing parquet type -> Spark data type -> Iceberg data type conversions are supported in parquet type -> PyArrow data type -> Iceberg data type conversions?
Add Arrow LargeString as an Iceberg data type. Map 1:1 with Arrow data type. The physical representation will still be backed by string.
Arrow LargeString is already converted to Iceberg String type in create_table by _convert_schema_if_needed (see Arrow: Support large-string #382). So when writing an Arrow table (in overwrite/append), convert the given Arrow table schema to the table's schema, after checking the two schemas are compatible.
Feature Request / Improvement
Currently,
large_stringdata type is converted tostring(link)This breaks the parquet writer when we're writing an Arrow table with a
large_stringcolumnSee pola-rs/polars#9795