All posts

Apache Iceberg 1.12.0: what's new, breaking changes, and upgrade guide

Apache Iceberg 1.12.0 ships geospatial types, Hilbert clustering, Flink equality-delete conversion, and format v4 groundwork. It also removes Spark 3.4 and position deletes with row data. Here is what changed and how to plan the upgrade.

13 min read Iceberg

TL;DR

  • Apache Iceberg 1.12.0 is a substantial release: 777 commits from 142 contributors.
  • Geospatial types (geometry and geography) arrive as first-class types, stored as WKB, with Spark 4.1 supporting them end to end.
  • Hilbert-curve clustering joins Z-order as a rewrite_data_files sort strategy; benchmark before adopting.
  • Flink gains a maintenance task that converts equality deletes to deletion vectors, cutting read cost for upsert pipelines.
  • Format version 4 plumbing is everywhere in the code, but V4 is not production-ready. Do not set format-version=4 on real tables.
  • Breaking: Spark 3.4 support is removed, position deletes with row data (PDWR) can no longer be written, and old S3 signer classes are gone.
  • Upgrade by table behavior, not by dependency file. The PDWR and Spark 3.4 changes need a plan before you bump.

1. What’s new in 1.12.0

1.1 Geospatial types arrive

Geospatial types

Iceberg now has first-class geometry and geography types, stored as WKB (Well-Known Binary). Both Avro and Parquet support reading and writing them with self-describing logical types. Spark 4.1 supports them end to end, including row-level DELETE, UPDATE, and MERGE.

Until now, spatial data in Iceberg meant a binary or string column plus a convention your team had to remember. Native types let engines and catalogs know what the column actually holds. This is a meaningful step for teams working with location data, mapping, or any spatial analytics workload.

Behavior change to watch: GeometryType and GeographyType now include the resolved CRS (Coordinate Reference System) and edge algorithm in their toString() output. A plain geometry now prints as geometry(OGC:CRS84), and geography prints as geography(OGC:CRS84, spherical). If you have tests or tooling that compare type strings, they may need updating. The community decided to keep this change because equal types printing differently looked like a bug, and function definition IDs built from parameter type strings would otherwise treat them as different overloads.

1.2 Hilbert clustering alongside Z-order

Hilbert vs Z-order

rewrite_data_files gains a second space-filling curve. Z-order is still there, and now you can also cluster on a Hilbert curve. Hilbert curves usually keep nearby values closer together than Z-order does, which can mean tighter file-level min/max statistics when you filter on several columns at once.

-- Existing: Z-order clustering
CALL my_catalog.system.rewrite_data_files(
  table => 'db.events',
  strategy => 'sort',
  sort_order => 'zorder(device_id, event_ts)'
);

-- New in 1.12: Hilbert curve clustering
CALL my_catalog.system.rewrite_data_files(
  table => 'db.events',
  strategy => 'sort',
  sort_order => 'hilbert(device_id, event_ts)'
);

Important details:

  • hilbert() has no options of its own. Each column contributes its full 8-byte primitive width to the index, so Z-order knobs like var-length-contribution and max-output-size do not apply.
  • You cannot combine hilbert() with zorder() in the same sort order.
  • Do not assume Hilbert always wins. Benchmark it on your real query patterns before switching.

Flink deletion vectors

This may be the most useful change for streaming teams. Flink upsert pipelines typically write equality deletes, which are cheap to write but expensive to read because every reader must join them against data files. The new ConvertEqualityDeletes maintenance task turns those into row-position deletion vectors (DVs), which readers can apply cheaply.

How it works: the Flink IcebergSink writes data files and equality deletes to a staging branch. The converter then resolves the deletes and commits the data files plus DVs to the target branch (usually main).

// Flink TableMaintenance, Iceberg 1.12
.add(ConvertEqualityDeletes.builder()
    .stagingBranch("staging")
    .equalityFieldColumns(ImmutableList.of("id"))
    .scheduleOnEqDeleteFileCount(10)
    .parallelism(4))

Requirements and caveats:

  • Requires table format version 3 or higher.
  • Does not support row lineage yet.
  • equalityFieldColumns must match the columns your writer uses.
  • The task keeps a primary-key index in Flink state; RocksDB is recommended.
  • Avoid running other continuous writers against the target branch, or conversion cycles can keep losing the commit race.

1.4 Format v4 groundwork

Format versions

Iceberg 1.12 contains a lot of code labeled v4, but format version 4 is not ready for production tables. This is implementation groundwork that shows where the format is heading.

What is present in the code:

  • Relative paths in the evolving specification, and utilities for converting absolute locations to relative ones. This can make table metadata less tied to one bucket or storage prefix.
  • Parquet or Avro manifests (instead of Avro-only), which can bring richer types and reader tooling to manifest data.
  • Content-statistics structures and filters, carrying structured index and average-size information alongside tracked files.
  • Tracked-file builders and adapters, plus a v4 manifest reader.
  • A draft Mumbling Bitmap specification and a read-only implementation.
  • EntryStatus.MODIFIED and TrackingBuilder status derivation, expressing changes that do not fit the older added/existing/deleted model cleanly.

What you should not do:

  • Do not set format-version=4 on production tables.
  • Do not build application code that parses V4 structures.
  • Do not build a roadmap around draft field names.

The useful action today is to keep table movement and metadata assumptions out of application code, and follow the specification through later releases.

1.5 Scan planning, metrics, and maintenance

The less visible changes in 1.12 may save more time than the headline features.

  • Partition-statistics scans can now project columns and apply filters, reducing the amount of metadata an optimizer must read.
  • Snapshot expiration correctly reads delete manifests in cases that previously went wrong.
  • A scan-based action removes dangling delete files.
  • RESTMetricsReporter.report() no longer blocks the calling thread, so emitting metrics cannot stall the work they describe. Still monitor queue size, delivery failures, and dropped reports.
  • Adaptive split-sizing options are documented, making it easier to balance planning overhead against parallelism.

Reader and file-format fixes cover Arrow integer-to-long promotion, dictionary-encoded INT96 timestamps with offsets, large decimals, decimal default values, end-of-file handling in object-store streams, Parquet row-group size tracking, and faster decimal handling that avoids unnecessary BigInteger conversions.

1.6 Kafka Connect stops hiding failures

Kafka Connect’s most important change is simple: commit failures are surfaced instead of silently swallowed.

  • Adds a metric for partial commit failures.
  • Bounded retries for transient commit exceptions.
  • Fixes rebalance scenarios that could commit files from a prior attempt.
  • Tracks control-topic offsets as a high-water mark.

Schema and conversion fixes cover null-valued records after a schema change, UUID conversion from Avro, decimal inference, MongoDB timestamp and date arrays, and Variant shredding. Removed deprecated members in TableReference and IcebergWriterResult may affect custom transforms or extensions at compile time.

1.7 Security and supply-chain work

  • The project published an Iceberg security model describing trust boundaries across catalogs, clients, storage, credentials, and metadata.
  • CI now makes vulnerability scanning blocking on relevant pull requests.
  • Jackson versions were aligned and updated to address reported vulnerabilities.
  • OAuth token refresh keeps optional parameters during non-exchange refreshes.
  • AWS REST signing can use assumed-role credentials.
  • RemoteSigningConfig implements the formalized remote signing spec, and old S3 signer classes are removed.

A library upgrade does not secure an Iceberg deployment by itself. Use the security model to review who can alter catalog metadata, who can request temporary storage credentials, who can delete files, where tokens are cached, and how audit events connect catalog actions to object-store actions.

2. Breaking changes

This is the section to share with whoever owns your Spark and Flink jobs. Several removals land in 1.12.0, and they break code, not just deprecate it.

Breaking changes

2.1 Spark 3.4 support removed

PR #14122 removes the Spark 3.4 runtime entirely. Spark 3.5, 4.0, and 4.1 remain, with version bumps to Spark 3.5.9, 4.0.4, and 4.1.3.

Spark 4.2 is deliberately excluded from the 1.12.0 source and binary artifacts. The community merged the Spark 4.2 module into main but excluded it from the release, with Anton Okolnychyi arguing that each new Spark release is a chance to fix DSv2 tech debt.

2.2 Position deletes with row data (PDWR) cannot be written

PR #17706 removes the feature the community deprecated in 1.11. The rewrite_position_delete_files and rewrite_table_path actions now fail with an explicit exception on files that still carry row data.

This was a deliberate community decision. At the August 26 sync, the group agreed that table owners should make explicit decisions about existing PDWR rather than having maintenance actions silently drop the row column.

2.3 Deprecated methods removed

  • Deprecated SparkTableUtil methods in Spark 4.0 and 4.1 are removed.
  • SparkFilters is deprecated in favor of SparkV2Filters.

2.4 Old S3 signer classes removed

PR #17627 drops deprecated AWS signer classes and properties. The new RemoteSigningConfig implementation follows the formalized remote signing spec.

3. Migrating from 1.11 to 1.12

Upgrade by table behavior, not by dependency file. The following sequence catches the failures most likely to escape a basic smoke test.

3.1 Pre-upgrade inventory

Build an engine and feature inventory before touching anything:

Item What to record
Engine clients Every Spark, Flink, Kafka Connect, Python, Rust, query-engine, and maintenance client that touches each catalog
Runtime versions Spark, Flink, and other engine versions in use
Iceberg library versions Current version on each client
Catalog client REST, Hive, Glue, JDBC, Nessie, or other
Authentication method OAuth2, SigV4, assumed-role, or other
File I/O implementation S3, GCS, ADLS, HDFS, or other
Table format versions V1, V2, V3 for each table
Delete-file strategies Copy-on-write vs merge-on-read; position vs equality deletes

This inventory tells you which breaking changes affect you and which do not.

3.2 Position deletes with row data (PDWR)

If you still run V2 tables with position deletes that include row data, your maintenance jobs will start failing after the upgrade. You have two ways out.

Option A: upgrade to V3 and rewrite as deletion vectors.

-- Upgrade table format version
ALTER TABLE prod.db.trips SET TBLPROPERTIES ('format-version' = '3');

-- Then rewrite position delete files (now allowed on V3)
CALL prod.system.rewrite_position_delete_files(table => 'db.trips');

Option B: stay on V2 and let data compaction fold those deletes into data files.

-- Run data compaction to rewrite position deletes into data files
CALL prod.system.rewrite_data_files(
  table => 'db.trips',
  strategy => 'sort'
);

Pick one before you upgrade, not after the first failed job.

How to find affected tables:

-- Query position deletes across tables to find PDWR
SELECT * FROM prod.db.trips.position_deletes LIMIT 10;
-- If the row column is populated, this table has PDWR

3.3 Spark 3.4 removal

If any job runs Spark 3.4, it must move to Spark 3.5, 4.0, or 4.1 before upgrading to Iceberg 1.12.0.

Recommended path:

  1. Upgrade Spark to 3.5.x first (3.5.9 is the 1.12.0 baseline).
  2. Run your test suite.
  3. Then upgrade Iceberg to 1.12.0.

Spark 4.2 users should stay on Iceberg 1.11.x until a later Iceberg release adds official 4.2 artifacts.

3.4 Deprecated API removals

Search your codebase for:

  • SparkTableUtil method calls (in Spark 4.0/4.1 projects)
  • SparkFilters usage (replace with SparkV2Filters)
  • Any code referencing removed S3 signer classes or properties

These fail at compile time, not runtime, so a full build catches them.

3.5 Testing checklist

After upgrading in staging, run these targeted tests:

Test area Why it matters
Arrow integer-to-long promotion Fixed in 1.12; verify with affected types
Dictionary-encoded INT96 timestamps with offsets Fixed in 1.12
Large decimals Fixed in 1.12
Decimal default values Fixed in 1.12
End-of-file handling in object-store streams Fixed in 1.12
Parquet row-group size tracking Fixed in 1.12
Kafka Connect commit failures Force a catalog commit failure and confirm it is reported
Kafka Connect rebalance Trigger a task rebalance during a pending commit
Kafka Connect restart Restart from a known control-topic offset
Geospatial types (if used) Verify toString() behavior change does not break tests

For Kafka Connect specifically, the release notes recommend three forced-failure drills: make a catalog commit fail, trigger a task rebalance during a pending commit, and restart from a known control-topic offset. Confirm the connector reports the failure, avoids duplicate table data, and resumes from the expected point.

4. Upgrade sequence in one page

The order that catches the failures most likely to escape a smoke test:

  1. Build an engine and feature inventory. List every Spark, Flink, Kafka Connect, Python, Rust, query-engine, and maintenance client that touches each catalog. Record its runtime version, Iceberg library version, catalog client, authentication method, and file I/O implementation.
  2. Inventory table format versions and delete-file strategies. Any V2 table writing PDWR needs a plan before upgrade.
  3. Test the specific changes that matter to your tables. Arrow integer-to-long promotion, dictionary-encoded INT96 timestamps with offsets, large decimals, decimal default values, end-of-file handling in object-store streams, Parquet row-group size tracking.
  4. Force failure drills on Kafka Connect. Make a catalog commit fail, trigger a task rebalance during a pending commit, restart from a known control-topic offset. Confirm the connector reports the failure, avoids duplicate table data, and resumes from the expected point.

5. Quick reference

Feature Status in 1.12.0 Action required
Geospatial types Available None unless using spatial data
Hilbert clustering Available Benchmark before adopting
Flink equality-delete conversion Available Requires V3 table; read docs first
Format V4 Not production-ready Do not use format-version=4
PDWR Removed Migrate V2 tables or compact
Spark 3.4 Removed Upgrade to Spark 3.5+
Spark 4.2 Excluded from release Stay on Iceberg 1.11.x
Old S3 signers Removed Update to RemoteSigningConfig
Kafka Connect commit failures Now surfaced Force failure drills after upgrade

Conclusion

Apache Iceberg 1.12.0 is a release that rewards planning. Geospatial types and Hilbert clustering are worth testing. Flink equality-delete conversion is worth adopting for streaming CDC pipelines. The breaking changes around Spark 3.4 and PDWR are worth planning for before you bump the version on Friday afternoon.

The V4 code is worth watching, not using. The specification work on relative paths, Parquet manifests, and content statistics shows where the format is heading, but setting format-version=4 on production tables today is not a supported path.

When you move off this release, re-check the format version range first, then the defaults in the properties table, then whether SparkProcedures has grown: procedures are the part of the surface that changes most between releases.

References

Trademarks

Apache Iceberg, Apache Spark, Apache Flink, Apache Hive, Apache Parquet, Apache Avro, Apache Hudi, Apache Polaris (incubating) and Apache are either registered trademarks or trademarks of The Apache Software Foundation in the United States and other countries. Delta Lake is a trademark of the Linux Foundation.

Found this useful?

These posts and tools are free. If one saved you an afternoon, you can buy me a coffee.

Buy me a coffee