Repository navigation
Conversation
| | _optional_ | _optional_ | _optional_ | **`partition-statistics`** | A list (optional) of [partition statistics](#partition-statistics). | | ||
| | | | _required_ | **`next-row-id`** | A `long` higher than all assigned row IDs; the next snapshot’s `first-row-id`. See [Row Lineage](#row-lineage). | | ||
| | | | _optional_ | **`encryption-keys`** | A list (optional) of [encryption keys](#encryption-keys) used for table encryption. | | ||
| | | | _optional_ | **`stats-only-fields`** | A list (optional) of [stats-only fields](#stats-only-fields) used to track stats for derived values. | |
There was a problem hiding this comment.
I added this to the v3 table as well because this is currently a forward-compatible change that doesn't need to be tied to v4.
There was a problem hiding this comment.
An optional field on v3 is readable by implementations that ignore unknown metadata. But do we expect v3 writer to put the stats in the maps keyed by column ID (lower/upper_bounds, value_counts, etc.). I am wondering if we want to add new v3 writer behavior at this stage.
Appendix E's v3 section probably need to be updated if we decide to keep this.
|
|
||
| Writers must preserve existing stats for stats-only fields listed in a table's `stats-only-fields`. Writers should produce stats when possible for stats-only fields. If an expression is not supported or produces a different output type when bound, a writer should produce no stats. | ||
|
|
||
| The data type of a stats-only field may only change according to the type promotion rules above. |
There was a problem hiding this comment.
I allowed type promotion because input values can be promoted and result in a promoted output type. This can happen with identity partitions as well as expressions.
|
|
||
| The `expr-value` type stores stats for the result of a [value expression](https://iceberg.apache.org/expressions-spec#value-expressions), stored in the `expr` field. The output type of the value expression is stored in the `data-type` field and must be a primitive or variant. | ||
|
|
||
| Readers must not fail when an unsupported stats-only field `type` is found; stats for unsupported types must be ignored. |
There was a problem hiding this comment.
This is intended to make new types forward compatible.
|
|
||
| Readers must not fail when an unsupported stats-only field `type` is found; stats for unsupported types must be ignored. | ||
|
|
||
| Writers must preserve existing stats for stats-only fields listed in a table's `stats-only-fields`. Writers should produce stats when possible for stats-only fields. If an expression is not supported or produces a different output type when bound, a writer should produce no stats. |
There was a problem hiding this comment.
The only issue here is the type used to build the content stats schema.
Right now, data-type is not required for partition-value because the type may be promoted (int -> long) when the source column is widened and I didn't want to include metadata that could be stale.
However, by not having data-type for all cases, the requirement to preserve existing stats becomes harder because the output type for a new stats-only field type may not be known.
It might be a good idea to require data-type for all types except partition-value. But if we add a sort-value later we would want to do the same thing.
There was a problem hiding this comment.
Could we require data-type unless the stats-only field type defines how to derive its current result
A stats-only field stores its current result type in `data-type`
unless its `type` defines how to derive that result type.
`data-type` must be a primitive or variant.
`partition-value` derives its result type from the partition field's
result type and must omit `data-type`. The result type
may change only according to the type promotion rules above.
| |**`default-sort-order-id`**|`JSON int`|`0`| | ||
| |**`refs`**|`JSON map with string key and object value:`<br />`{`<br /> `"<name>": {`<br /> `"snapshot-id": <id>,`<br /> `"type": <type>,`<br /> `"max-ref-age-ms": <long>,`<br /> `...`<br /> `}`<br /> `...`<br />`}`|`{`<br /> `"test": {`<br /> `"snapshot-id": 123456789000,`<br /> `"type": "tag",`<br /> `"max-ref-age-ms": 10000000`<br /> `}`<br />`}`| | ||
| |**`encryption-keys`**|`JSON list of encryption key objects`|`[ {"key-id": "5f819b", "key-metadata": "aWNlYmVyZwo="} ]`| | ||
| |**`stats-only-fields`**|`JSON list of stats-only field objects`|`[ {"field-id": 102, "type": "partition-value", "partition-field-id": 1001} ]`| |
There was a problem hiding this comment.
should we add an example for expression-value type?
|
|
||
| Stats-only fields are used to track stats for derived values that are not part of the table schema and are not materialized. A stats-only field consists of: | ||
|
|
||
| * A **`field-id`** assigned by incrementing the table's `last-field-id` |
There was a problem hiding this comment.
The table metadata field is last-column-id, not last-field-id.
|
|
||
| Readers must not fail when an unsupported stats-only field `type` is found; stats for unsupported types must be ignored. | ||
|
|
||
| Writers must preserve existing stats for stats-only fields listed in a table's `stats-only-fields`. Writers should produce stats when possible for stats-only fields. If an expression is not supported or produces a different output type when bound, a writer should produce no stats. |
There was a problem hiding this comment.
Could we require data-type unless the stats-only field type defines how to derive its current result
A stats-only field stores its current result type in `data-type`
unless its `type` defines how to derive that result type.
`data-type` must be a primitive or variant.
`partition-value` derives its result type from the partition field's
result type and must omit `data-type`. The result type
may change only according to the type promotion rules above.
| | _optional_ | _optional_ | _optional_ | **`partition-statistics`** | A list (optional) of [partition statistics](#partition-statistics). | | ||
| | | | _required_ | **`next-row-id`** | A `long` higher than all assigned row IDs; the next snapshot’s `first-row-id`. See [Row Lineage](#row-lineage). | | ||
| | | | _optional_ | **`encryption-keys`** | A list (optional) of [encryption keys](#encryption-keys) used for table encryption. | | ||
| | | | _optional_ | **`stats-only-fields`** | A list (optional) of [stats-only fields](#stats-only-fields) used to track stats for derived values. | |
There was a problem hiding this comment.
An optional field on v3 is readable by implementations that ignore unknown metadata. But do we expect v3 writer to put the stats in the maps keyed by column ID (lower/upper_bounds, value_counts, etc.). I am wondering if we want to add new v3 writer behavior at this stage.
Appendix E's v3 section probably need to be updated if we decide to keep this.
This adds "stats-only fields" to table metadata. A stats-only field is used to track stats for a non-materialized derived value. For example, one could be used to track stats for
to_lower_case(name).This introduces two cases: partition fields for the result of bucket partitioning, and value expressions for use cases like the one above.
This was discussed on the dev list in the thread, "[DISCUSS] Field ID tracking for non-materialized columns in v4".