Skip to content

Spec: Add stats-only fields to table metadata - #18405

Open
rdblue wants to merge 1 commit into
apache:mainfrom
rdblue:spec-add-stats-only-fields
Open

rdblue wants to merge 1 commit into
apache:mainfrom
rdblue:spec-add-stats-only-fields

Conversation

@rdblue

@rdblue rdblue commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

This adds "stats-only fields" to table metadata. A stats-only field is used to track stats for a non-materialized derived value. For example, one could be used to track stats for to_lower_case(name).

This introduces two cases: partition fields for the result of bucket partitioning, and value expressions for use cases like the one above.

This was discussed on the dev list in the thread, "[DISCUSS] Field ID tracking for non-materialized columns in v4".

@github-actions github-actions Bot added the Specification Issues that may introduce spec changes. label Oct 6, 2026
Comment thread format/spec.md
| _optional_ | _optional_ | _optional_ | **`partition-statistics`** | A list (optional) of [partition statistics](#partition-statistics). |
| | | _required_ | **`next-row-id`** | A `long` higher than all assigned row IDs; the next snapshot’s `first-row-id`. See [Row Lineage](#row-lineage). |
| | | _optional_ | **`encryption-keys`** | A list (optional) of [encryption keys](#encryption-keys) used for table encryption. |
| | | _optional_ | **`stats-only-fields`** | A list (optional) of [stats-only fields](#stats-only-fields) used to track stats for derived values. |

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added this to the v3 table as well because this is currently a forward-compatible change that doesn't need to be tied to v4.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

An optional field on v3 is readable by implementations that ignore unknown metadata. But do we expect v3 writer to put the stats in the maps keyed by column ID (lower/upper_bounds, value_counts, etc.). I am wondering if we want to add new v3 writer behavior at this stage.

Appendix E's v3 section probably need to be updated if we decide to keep this.

Comment thread format/spec.md

Writers must preserve existing stats for stats-only fields listed in a table's `stats-only-fields`. Writers should produce stats when possible for stats-only fields. If an expression is not supported or produces a different output type when bound, a writer should produce no stats.

The data type of a stats-only field may only change according to the type promotion rules above.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I allowed type promotion because input values can be promoted and result in a promoted output type. This can happen with identity partitions as well as expressions.

Comment thread format/spec.md

The `expr-value` type stores stats for the result of a [value expression](https://iceberg.apache.org/expressions-spec#value-expressions), stored in the `expr` field. The output type of the value expression is stored in the `data-type` field and must be a primitive or variant.

Readers must not fail when an unsupported stats-only field `type` is found; stats for unsupported types must be ignored.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is intended to make new types forward compatible.

Comment thread format/spec.md

Readers must not fail when an unsupported stats-only field `type` is found; stats for unsupported types must be ignored.

Writers must preserve existing stats for stats-only fields listed in a table's `stats-only-fields`. Writers should produce stats when possible for stats-only fields. If an expression is not supported or produces a different output type when bound, a writer should produce no stats.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The only issue here is the type used to build the content stats schema.

Right now, data-type is not required for partition-value because the type may be promoted (int -> long) when the source column is widened and I didn't want to include metadata that could be stale.

However, by not having data-type for all cases, the requirement to preserve existing stats becomes harder because the output type for a new stats-only field type may not be known.

It might be a good idea to require data-type for all types except partition-value. But if we add a sort-value later we would want to do the same thing.

@stevenzwu stevenzwu Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we require data-type unless the stats-only field type defines how to derive its current result

A stats-only field stores its current result type in `data-type` 
unless its `type` defines how to derive that result type. 
`data-type` must be a primitive or variant. 
`partition-value` derives its result type from the partition field's 
result type and must omit `data-type`. The result type 
may change only according to the type promotion rules above.

Comment thread format/spec.md
|**`default-sort-order-id`**|`JSON int`|`0`|
|**`refs`**|`JSON map with string key and object value:`<br />`{`<br />&nbsp;&nbsp;`"<name>": {`<br />&nbsp;&nbsp;`"snapshot-id": <id>,`<br />&nbsp;&nbsp;`"type": <type>,`<br />&nbsp;&nbsp;`"max-ref-age-ms": <long>,`<br />&nbsp;&nbsp;`...`<br />&nbsp;&nbsp;`}`<br />&nbsp;&nbsp;`...`<br />`}`|`{`<br />&nbsp;&nbsp;`"test": {`<br />&nbsp;&nbsp;`"snapshot-id": 123456789000,`<br />&nbsp;&nbsp;`"type": "tag",`<br />&nbsp;&nbsp;`"max-ref-age-ms": 10000000`<br />&nbsp;&nbsp;`}`<br />`}`|
|**`encryption-keys`**|`JSON list of encryption key objects`|`[ {"key-id": "5f819b", "key-metadata": "aWNlYmVyZwo="} ]`|
|**`stats-only-fields`**|`JSON list of stats-only field objects`|`[ {"field-id": 102, "type": "partition-value", "partition-field-id": 1001} ]`|

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should we add an example for expression-value type?

Comment thread format/spec.md

Stats-only fields are used to track stats for derived values that are not part of the table schema and are not materialized. A stats-only field consists of:

* A **`field-id`** assigned by incrementing the table's `last-field-id`

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The table metadata field is last-column-id, not last-field-id.

Comment thread format/spec.md

Readers must not fail when an unsupported stats-only field `type` is found; stats for unsupported types must be ignored.

Writers must preserve existing stats for stats-only fields listed in a table's `stats-only-fields`. Writers should produce stats when possible for stats-only fields. If an expression is not supported or produces a different output type when bound, a writer should produce no stats.

@stevenzwu stevenzwu Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we require data-type unless the stats-only field type defines how to derive its current result

A stats-only field stores its current result type in `data-type` 
unless its `type` defines how to derive that result type. 
`data-type` must be a primitive or variant. 
`partition-value` derives its result type from the partition field's 
result type and must omit `data-type`. The result type 
may change only according to the type promotion rules above.

Comment thread format/spec.md
| _optional_ | _optional_ | _optional_ | **`partition-statistics`** | A list (optional) of [partition statistics](#partition-statistics). |
| | | _required_ | **`next-row-id`** | A `long` higher than all assigned row IDs; the next snapshot’s `first-row-id`. See [Row Lineage](#row-lineage). |
| | | _optional_ | **`encryption-keys`** | A list (optional) of [encryption keys](#encryption-keys) used for table encryption. |
| | | _optional_ | **`stats-only-fields`** | A list (optional) of [stats-only fields](#stats-only-fields) used to track stats for derived values. |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

An optional field on v3 is readable by implementations that ignore unknown metadata. But do we expect v3 writer to put the stats in the maps keyed by column ID (lower/upper_bounds, value_counts, etc.). I am wondering if we want to add new v3 writer behavior at this stage.

Appendix E's v3 section probably need to be updated if we decide to keep this.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Specification Issues that may introduce spec changes.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants