Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 24 additions & 0 deletions format/spec.md
Original file line number Diff line number Diff line change
Expand Up @@ -436,6 +436,27 @@ Two rows are the "same"---that is, the rows represent the same entity---if the i

Identifier fields may be nested in structs but cannot be nested within maps or lists. Float, double, and optional fields cannot be used as identifier fields and a nested field cannot be used as an identifier field if it is nested in an optional struct, to avoid null values in identifiers.

#### Stats-only Fields

Stats are tracked in manifests by field ID. Schema fields have assigned IDs, but additional field IDs may be assigned to track stats for derived values. For instance, lower and upper bounds for `to_lower_case(name)` are useful for case-insensitive file pruning.

Stats-only fields are used to track stats for derived values that are not part of the table schema and are not materialized. A stats-only field consists of:

* A **`field-id`** assigned by incrementing the table's `last-column-id`
* A **`type`** that can be `partition-value` or `expr-value`

@danielcweeks danielcweeks Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is the primary function of partition-value to tie the partition-field-id to a "stats" field-id? Would this only be used for bucket/multi-arg transforms?

I assume we wouldn't track identity columns this way.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's correct. This is to be able to store the output of a non-monotonic function. I would not expect anyone to use this for identity because the easier path is to store the identity value in the field's stats directly.

* An optional **`data-type`** that determines the bound type; must be a primitive or variant
* Type-specific fields that define how derived values are produced

The `partition-value` type stores stats for the output of a partition field, identified by a `partition-field-id` type-specific field. The bound type is the partition field's result type and `data-type` is omitted. This may be used in v4 to filter by bucket partition values.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This may be used in v4 to filter by bucket partition values.

do we need to mention "v4" here? the stats-only fields are allowed in v3. if v3 writer populates these partition stats fields, do we need to restrict v3 readers from using them (although they don't need to).

if we keep it, maybe v4 can be renamed to v4+ for future proof?


The `expr-value` type stores stats for the result of a [value expression](https://iceberg.apache.org/expressions-spec#value-expressions), stored in the `expr` field. Expressions must use only ID references. The output type of the value expression must be stored in the `data-type` field.

Readers must not fail when an unsupported stats-only field `type` is found; stats for unsupported types must be ignored.
Comment thread
pvary marked this conversation as resolved.

Writers must preserve existing stats for stats-only fields listed in a table's `stats-fields` if the `data-type` is known (either set or specified by a supported type). Writers should produce stats when possible for stats-only fields. A writer should produce no stats by setting the field stats struct to null when an expression is not supported, produces a different output type when bound, or binding fails.

@stevenzwu stevenzwu Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

produces a different output type when bound

nit: this reads a bit awkward to me. it is one of the 3 conditions for when. maybe sth like

binding produces a different output type


The data type of a stats-only field may only change according to the type promotion rules above.
Comment thread
pvary marked this conversation as resolved.

#### Reserved Field IDs

Iceberg tables must not use field ids greater than 2147483447 (`Integer.MAX_VALUE - 200`). This id range is reserved for metadata columns that can be used in user data schemas, like the `_file` column that holds the file path in which a row was stored.
Expand Down Expand Up @@ -1169,6 +1190,7 @@ Table metadata consists of the following fields:
| _optional_ | _optional_ | _optional_ | **`partition-statistics`** | A list (optional) of [partition statistics](#partition-statistics). |
| | | _required_ | **`next-row-id`** | A `long` higher than all assigned row IDs; the next snapshot’s `first-row-id`. See [Row Lineage](#row-lineage). |
| | | _optional_ | **`encryption-keys`** | A list (optional) of [encryption keys](#encryption-keys) used for table encryption. |
| | | _optional_ | **`stats-fields`** | A list (optional) of [stats-only fields](#stats-only-fields) used to track stats for derived values. |
=== "v4"
| v4 | Field | Description |
|------------|-----------------------------|-------------|
Expand Down Expand Up @@ -1197,6 +1219,7 @@ Table metadata consists of the following fields:
| _optional_ | **`partition-statistics`** | A list (optional) of [partition statistics](#partition-statistics). |
| _required_ | **`next-row-id`** | A `long` higher than all assigned row IDs; the next snapshot's `first-row-id`. See [Row Lineage](#row-lineage). |
| _optional_ | **`encryption-keys`** | A list (optional) of [encryption keys](#encryption-keys) used for table encryption. |
| _optional_ | **`stats-fields`** | A list (optional) of [stats-only fields](#stats-only-fields) used to track stats for derived values. |

For serialization details, see Appendix C.

Expand Down Expand Up @@ -1801,6 +1824,7 @@ A metadata JSON file may be compressed with [GZIP](https://datatracker.ietf.org/
|**`default-sort-order-id`**|`JSON int`|`0`|
|**`refs`**|`JSON map with string key and object value:`<br />`{`<br />&nbsp;&nbsp;`"<name>": {`<br />&nbsp;&nbsp;`"snapshot-id": <id>,`<br />&nbsp;&nbsp;`"type": <type>,`<br />&nbsp;&nbsp;`"max-ref-age-ms": <long>,`<br />&nbsp;&nbsp;`...`<br />&nbsp;&nbsp;`}`<br />&nbsp;&nbsp;`...`<br />`}`|`{`<br />&nbsp;&nbsp;`"test": {`<br />&nbsp;&nbsp;`"snapshot-id": 123456789000,`<br />&nbsp;&nbsp;`"type": "tag",`<br />&nbsp;&nbsp;`"max-ref-age-ms": 10000000`<br />&nbsp;&nbsp;`}`<br />`}`|
|**`encryption-keys`**|`JSON list of encryption key objects`|`[ {"key-id": "5f819b", "key-metadata": "aWNlYmVyZwo="} ]`|
|**`stats-only-fields`**|`JSON list of stats-only field objects`|`[ {"field-id": 102, "type": "partition-value", "partition-field-id": 1001} ]`|
Comment thread
stevenzwu marked this conversation as resolved.

### Name Mapping Serialization

Expand Down
Loading