All posts

Branching, tagging and write-audit-publish in Iceberg: CI for a table

A branch lets a job write to a table nobody is reading, and publishing is a pointer move. Here is that measured - branch isolation, what spark.wap.branch redirects that is easy to miss, fast-forward, tags with retention, and the divergence that makes fast-forward refuse.

10 min read Iceberg

TL;DR

  • A branch is an independent pointer into the same table. Writing four rows to a branch left main at 2 rows while the branch held 4 — same table, same files, no copy.
  • spark.wap.branch redirects reads as well as writes. With it set, SELECT count(*) on the table returned the branch’s 4 rows; unsetting it revealed main had 2 all along. A job that writes and then validates in the same session is validating the branch, which is what you want — and it means you cannot check main from inside that session.
  • Publishing is fast_forward, a pointer move rather than a merge. It took main from 2 rows to 4 and returned the snapshot ids on both sides.
  • Iceberg has no branch merge. Once main moved independently, fast_forward refused: Cannot fast-forward: main is not an ancestor of feature.
  • Discarding work is free. DROP BRANCH removed the ref and left main untouched — nothing to clean up, nothing to undo.
  • Tags are named snapshots with optional retention. CREATE TAG ... RETAIN 30 DAYS showed up in refs as max_reference_age_in_ms = 2592000000.

The pattern every data team reinvents is “check it before anyone sees it”. Usually that means writing to a staging table, running assertions, and then copying or swapping — which costs a full copy of the data and a rename that is not atomic anywhere object storage is involved.

Iceberg does it with a pointer. A branch is another named reference into the same metadata tree, so a write to a branch produces files that only readers of that branch can see, and publishing is updating which snapshot main names.

What a branch actually is

The three-tier model has the catalog holding one pointer to the current metadata file, and that file holding a tree of snapshots. A branch adds a named reference to a snapshot in that tree.

flowchart LR
    S1["snapshot 1"] --> S2["snapshot 2"]
    S2 --> S3["snapshot 3<br/>(main)"]
    S2 --> B1["snapshot 4<br/>(audit)"]

Nothing is copied. Both refs point into one set of data files, and a file referenced by either is kept. That is why creating a branch is instant whatever the table’s size, and why deleting one costs nothing.

ALTER TABLE ice.brdemo.events CREATE BRANCH audit;

Branches are genuinely isolated

A table with two rows, a branch, then a write to the branch only:

INSERT INTO ice.brdemo.events.branch_audit VALUES (3,'c'), (4,'d');
MAIN_ROWS   2
AUDIT_ROWS  4
REFS        [('audit','BRANCH'), ('main','BRANCH')]

Readers of the table see two rows throughout. The refs metadata table lists every named reference and its kind, which is the place to look when you want to know what exists:

SELECT name, type, snapshot_id, max_reference_age_in_ms FROM ice.brdemo.events.refs;

Note that main is itself a branch. There is nothing special about it beyond being the default.

Write-audit-publish, and the redirect that surprises people

Writing to table.branch_audit works, but it requires changing the job’s SQL — which is exactly what you do not want, because then the audited job is not the job that runs in production.

That is what the WAP configuration is for. Enable it on the table, set one session property, and an unmodified job writes to the branch:

CREATE TABLE ice.brdemo.wap (id INT, val STRING) USING iceberg
TBLPROPERTIES ('write.wap.enabled'='true');
spark.conf.set("spark.wap.branch", "staging")
spark.sql("INSERT INTO ice.brdemo.wap VALUES (3,'c'), (4,'d')")   # no branch named anywhere

Measured, with main holding 2 rows before:

with spark.wap.branch set:     main = 4      staging = 4
after unsetting it:            main = 2

Read the first line carefully, because it is the part that catches people. While spark.wap.branch is set, reads are redirected too. SELECT count(*) against the plain table name returned the branch’s four rows, not main’s two. Only after unsetting the property did the table show its real state.

That is the correct design — a job that writes and then asserts should be asserting on what it just wrote — but it has a consequence worth knowing: you cannot check whether main changed from inside the session that has WAP set. Every query in that session is pointed at the branch. Verify from another session, or unset first.

The whole workflow is then three steps and no data movement:

# 1. write - the production job, unmodified
spark.conf.set("spark.wap.branch", "staging")
run_the_job()

# 2. audit - against the branch
bad = spark.sql("SELECT count(*) FROM ice.brdemo.wap.branch_staging WHERE val IS NULL").first()[0]

# 3. publish - a pointer move
if bad == 0:
    spark.sql("CALL ice.system.fast_forward('brdemo.wap','main','staging')")
AUDIT_NULLS        0
AFTER_PUBLISH main = 4

If the audit fails, nothing is published and nothing needs undoing. Readers of the table never saw the candidate data at any point.

Publishing is a pointer move, not a merge

fast_forward returns exactly what it did:

Row(branch_updated='main',
    previous_ref=5583543293194953086,
    updated_ref=6588505601775582655)

Two snapshot ids: where main pointed, and where it points now. That is the whole operation.

Which brings the important limitation. Iceberg has no branch merge. A fast-forward is only valid when the target is an ancestor of the source — when main has not moved since the branch was created. Let both move and it refuses:

MAIN 2   FEATURE 2
Cannot fast-forward: main is not an ancestor of feature

main took an independent write, so the two references diverged, and there is no operation in Iceberg that reconciles them. The mental model to carry is git’s fast-forward without git’s merge: a branch is a candidate for becoming the current state, not something you combine.

The practical consequences:

  • Keep branches short-lived. The longer a branch exists, the more likely main moves and the candidate becomes unpublishable.
  • One writer to main while a branch is outstanding, or accept that you may have to redo the work on a fresh branch.
  • Re-branch rather than reconcile. If divergence happens, create a new branch from current main and re-run, which is usually cheap because the job is unmodified anyway.

For picking up a single snapshot rather than advancing a whole branch there is cherrypick_snapshot, which applies one snapshot’s changes onto the current state — a narrower tool for a narrower case.

Throwing work away is free

The failure path is the reason to like this pattern:

ALTER TABLE ice.brdemo.div DROP BRANCH feature;
AFTER_DROP_REFS   ['main']
MAIN_UNCHANGED    2

The ref is gone and main is untouched. The data files that only the branch referenced become unreferenced and are removed by the ordinary expiry and maintenance jobs — there is no special cleanup, and crucially no half-published state to repair.

Compare that with the staging-table version of this workflow, where a failed run leaves a populated staging table someone has to notice and truncate.

Tags: a snapshot with a name

A tag is a reference that is not meant to move. Where a branch is “work in progress”, a tag is “this exact state, remembered”:

ALTER TABLE ice.brdemo.wap CREATE TAG `q3-close` RETAIN 30 DAYS;
REFS  [('main','BRANCH',None), ('q3-close','TAG',2592000000), ('staging','BRANCH',None)]

2592000000 ms is exactly thirty days, and it is what stops snapshot expiry from removing the files that tag depends on. Without retention, a tag is only as durable as your expiry policy — which is the failure people hit when a “reproducible” tag stops resolving months later.

Reading through it is ordinary time travel by name:

SELECT count(*) FROM ice.brdemo.wap VERSION AS OF 'q3-close';
4

Which is the case tags earn their place in: a reported number that must still be reproducible after the underlying table has moved on. A tag named for the report is self-documenting in a way a raw snapshot id is not — see time travel and rollback for the snapshot-id form and what rollback leaves behind.

Where this fits in practice

Want Use
Validate before anyone reads it branch + spark.wap.branch + fast_forward
Remember a state for reproducibility tag, with RETAIN
Undo a bad write already published rollback, not branching
Combine two lines of work not available — re-branch and re-run
Long-lived parallel versions of a table not what branches are for; divergence makes them unpublishable

The sweet spot is short-lived, single-purpose branches in an automated pipeline: create, write, assert, publish or drop, all inside one job run. Branches held open for days drift into the divergence problem, and the cost of re-running is usually lower than the cost of discovering you cannot publish.

Common misconceptions

“Branching copies the table.” It adds a reference into the existing snapshot tree. Nothing is copied at any size.

“spark.wap.branch only affects writes.” It redirects reads in that session too — measured above, main appeared to have 4 rows until the property was unset.

“I can merge a branch into main.” There is no merge. fast_forward requires the target to be an ancestor, and refuses otherwise.

“A tag is permanent.” Only with RETAIN, or until expiry removes the snapshot it names.

“main is special.” It is a branch like any other; refs lists it as one.

“A failed branch needs cleaning up.” Drop the ref. Ordinary maintenance reclaims the files.

Frequently asked questions

Does WAP work outside Spark? The branch and tag mechanics are in the format and the table property is engine-neutral, but spark.wap.branch is a Spark session property. Other engines expose branch writes their own way, or not at all — worth checking rather than assuming, the same caution that applies to Iceberg views.

Can two jobs write to different branches at once? Yes, and that is one of the better arguments for the feature — they touch different refs and do not contend on main.

What happens to branch files if I never publish? They stay referenced while the branch exists, and become collectable once it is dropped. A forgotten branch is a storage leak, so clean up branches the way you clean up snapshots.

Is there a retention setting for branches? Yes — branches accept retention properties too, which is how you stop a forgotten branch pinning old snapshots forever.

How does this compare with Hudi? Hudi has savepoints for pinning a state and restore for going back, but no branch-and-publish workflow of this shape. The nearest equivalent to the isolation here is writing to a separate table and swapping.

Conclusion

The numbers are small and the point is structural: main at 2 rows while a branch held 4, a publish that moved a pointer rather than data, and a drop that cost nothing. That is the write-audit-publish pattern with the copy removed from it.

Two things are worth remembering past the syntax. spark.wap.branch redirects reads as well as writes, so validation inside that session is validating the branch — intended, and misleading if you forget. And there is no merge: the moment main moves independently, fast_forward refuses and the candidate has to be rebuilt. Branches here are a staging mechanism with a cheap discard, not a version-control system.

Used that way — short-lived, automated, dropped on failure — they replace the staging table and the copy that goes with it.

References

Trademarks

Apache Iceberg, Apache Spark, Apache Hudi, Apache and the Apache feather logo are either registered trademarks or trademarks of The Apache Software Foundation in the United States and other countries.

Found this useful?

These posts and tools are free. If one saved you an afternoon, you can buy me a coffee.

Buy me a coffee