Repository navigation
Core: Don't prune equality delete manifest entries by non-key column stats - #18338
Open
findinpath wants to merge 1 commit into
Open
findinpath wants to merge 1 commit into
findinpath wants to merge 1 commit into
Conversation
findinpath
force-pushed
the
findinpath/equality-deletes
branch
2 times, most recently
from
October 1, 2026 11:53
2ec8a70 to
adcf116
Compare
chenjian2664
reviewed
Oct 1, 2026
findinpath
force-pushed
the
findinpath/equality-deletes
branch
2 times, most recently
from
October 1, 2026 13:29
f2cc7ae to
efa28c2
Compare
pvary
reviewed
Oct 5, 2026
pvary
reviewed
Oct 5, 2026
pvary
reviewed
Oct 5, 2026
Contributor
|
I'm not entirely sure how much effort do we want to put into fixing this, as there is no current use-case for this. Especially considering that we would like to get rid of the equality deletes in the long run. |
…stats An equality delete's match condition depends only on its equality_ids columns. Per spec, the file may legitimately carry additional columns of the deleted row, but their stats describe values that play no part in the match condition. InclusiveMetricsEvaluator evaluated the scan's row filter against all of a delete file's column stats, so a predicate on a non-key column could incorrectly prune a delete manifest entry and leave a row that should have been deleted in the query result. Narrow stats to a file's equality_ids columns before metrics evaluation for equality delete files; pruning using the equality-key columns' own stats remains sound and is preserved.
findinpath
force-pushed
the
findinpath/equality-deletes
branch
from
October 7, 2026 04:36
efa28c2 to
0f942a8
Compare
Contributor
Author
@pvary This represents a silent correctness issue. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An equality delete's match condition depends only on its equality_ids
columns. Per spec, the file may legitimately carry additional columns of
the deleted row, but their stats describe values that play no part in the
match condition.
ManifestReader's stats-based pruning evaluated the scan'srow filter against all of a delete file's column stats, so a predicate on
a non-key column could incorrectly prune a delete manifest entry and leave
a row that should have been deleted in the query result.
Only skip metrics evaluation for an equality delete file when the row
filter references a column outside its equality_ids; pruning using the
equality-key columns' own stats remains sound and is preserved.
Additional context
https://iceberg.apache.org/spec/#equality-delete-files
Equality delete files identify deleted rows in a collection of data files by one or more column values, and may optionally contain additional columns of the deleted row.
Is there a specific writer that does this (spark? flink? something else?)
Oracle GoldenGate
Issue found through
trinodb/trinotrinodb/trino#31399