Honor identity sort orders for Arrow table writes - #3830
Open
adonm wants to merge 3 commits into
Open
Conversation
Author
|
Synced with latest main (now 13 commits ahead of the original base, including the pyarrow 24→25 upgrade). pyarrow 25 deprecated the SortOptions-level Verified locally: unit test, docker-based REST catalog integration test ( |
adonm
force-pushed
the
feat/identity-sorted-writes
branch
from
August 25, 2026 04:59
eeef9cd to
3660095
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Related: #271 and #3848
Rationale for this change
PyIceberg accepts table sort metadata but currently writes unsorted files and hard-codes
sort_order_id=None. This prevents readers from safely using sort-order-aware pruning and leaves manifest metadata inconsistent with users' write intent.What changes?
Honors table sort orders for materialized
pyarrow.Tablewrites when every sort field uses an identity transform and one consistent null placement:WriteTaskDataFile.sort_order_idUnsupported transforms, nested/missing fields, mixed null placement, and streaming
RecordBatchReaderwrites preserve current behavior: data is not claimed as sorted and the file sort-order ID remains null. A warning explains why.Are these changes tested?
sort_order_idagainst a REST catalogPYTHONPATH=. uv run pytest tests/io/test_pyarrow.py -k "sort_table_for_identity_sort_order or write_sorted_data_files_per_partition" -q— passedruff checkon changed files — passedgit diff --check— passedAre there any user-facing changes?
Yes. Tables with an identity-transform sort order now write physically sorted data files carrying the truthful
sort_order_id; previously all files were written unsorted with a null sort-order ID. Unsupported sort orders keep the previous behavior. Changelog label requested.Tooling note: developed with assistance from DS v4 Pro. I reviewed and verified the changes.