Guides
The posts on this site in the order worth reading them, grouped by what you are trying to learn: how Spark executes a query, how the table formats store a table, and which tool to reach for.
- How does Spark actually execute a query?
- What changes once the query never ends?
- Which Spark reference should I keep open?
- How do the table formats store a table?
- How does a Hudi table work, end to end?
- How does an Iceberg table work, end to end?
- How do I run a table format on Spark?
- How do I run any of this locally?
- Which tool do I reach for?
- Everything, from basic to advanced
The category and archive pages list everything by topic and by date. This page is different: it puts the posts in a reading order, and says what each one gives you, so you can start somewhere sensible rather than at whatever was published most recently.
Each path is built so that every post assumes only what the ones above it already explained.
How does Spark actually execute a query?
Ten posts that build on each other, from the runtime up to the optimizer and back down to what fails in production. If you read one, read the first.
- Apache Spark architecture is the foundation: who schedules what, how a job becomes stages and tasks, where a shuffle is written, and which component to suspect when a job misbehaves. Everything else here assumes it.
- Inside Spark’s Catalyst optimizer takes one query apart rule by rule: the four trees, what a rule is in the source, and how to switch one off and watch the plan change.
- Adaptive query execution covers what the engine revises once it has measured the data, and, just as usefully, the parts of the plan it never revisits.
- Apache Spark joins in depth is the decision that dominates most jobs: five strategies, why there are five, and how to make each one appear on demand.
- Spark memory management accounts for every region of an executor’s memory and the arithmetic that sizes it, including the one hard floor you cannot tune below.
- Spark shuffle internals goes under the stage boundary the first post described: Spark has three shuffle writers and picks one per shuffle without ever naming the choice. The decision tree here is read from the source and confirmed by running it.
- Data skew in Apache Spark is the most common way all of the above goes wrong in production, with six fixes and the measurements that tell you which one you need.
- Spark performance tuning is the capstone: seventeen techniques, each with the code, the evidence the optimization actually fired, and the measurement that says whether it helped. Most tuning advice is a list of settings with no way to tell.
- Accelerating Spark with DataFusion Comet swaps those operators for native ones without changing the query, which is the clearest way to see which parts of a plan the JVM was costing you.
- Every Apache Spark release puts the whole thing on a timeline: what each version changed, which defaults moved underneath you, and what a given upgrade is actually worth.
What changes once the query never ends?
Streaming is the same engine with a different contract: the query outlives any one run, so correctness depends on what it wrote down before it stopped.
- Structured Streaming internals is the place to start: a streaming query is a loop that writes what it is about to do, does it, then writes that it finished. This is what that loop puts on disk, how the state store is laid out underneath it, and which parts of a query a restart can never change.
- Streaming into Iceberg is the same loop pointed at a table format, where every micro-batch commits a snapshot and the trigger you pick decides how fast your metadata grows. It appears again as step 10 of the Iceberg path below, which is the other way into it.
For changing the log level on a query you cannot restart, see Logging a Spark stream in the reference list above.
Which Spark reference should I keep open?
These are lookup material rather than a reading order. Open them beside your editor.
- Apache Spark: the complete cheat sheet for the config key, default and syntax you half remember.
- A Spark JVM playbook for custom logging, class loading, stack size and proxy problems, which are the failures that look like Spark bugs and are not.
- Logging a Spark stream for the same problem on a job you cannot restart, where log volume follows partition count and the level has to change while the query runs.
- Top Apache Spark interview questions as a self-test over everything above, with answers.
How do the table formats store a table?
Start by choosing, then go deep on the one you chose.
- Open table formats in practice is the comparison: what actually differs between Iceberg, Hudi and Delta Lake, rather than what the marketing pages differ on.
- Apache Hudi, if you chose Hudi: it has its own reading path below, seven posts from the first upsert to concurrent writers.
- Apache Iceberg architecture for what a commit actually swaps, and why the catalog is part of the table.
- Apache XTable incremental sync for keeping one table readable as more than one format without re-converting everything each time.
The matching reference for step 4 is the XTable cheat sheet.
How does a Hudi table work, end to end?
Seven posts, from what an upsert does on disk to what goes wrong when two writers commit at once. Each one assumes only the ones above it.
Keep the Hudi cheat sheet open while you read: table types, the timeline, writer properties, indexes, table services and procedures, with the config key and default for each.
- Apache Hudi: an introduction is the place to start.
- Apache Hudi architecture for what happens when you upsert a row, down to bytes on object storage.
- Key generators decide what a record’s identity is and where it lands — the one Hudi decision you cannot change after the first write.
- The Hudi index is the single biggest lever on upsert cost — it decides which file group a record belongs to.
- Hudi Merge-on-Read for where the write you skipped goes, and which read pays for it.
- Incremental processing, the capability Hudi is named for: what an incremental read actually returns, how CDC mode differs, and the timestamp change in 1.x that makes the obvious checkpoint loop re-read every batch.
- Concurrency control, where two of three configurations lost a whole commit without raising anything, including the one most people deploy.
For running Hudi from a thin client or making a stalled write explain itself, see the Spark Connect and logging posts under How do I run a table format on Spark?
How does an Iceberg table work, end to end?
The longest path on the site, and the one that rewards being read in order: eighteen posts from why the format exists to what you have to schedule in production. Each one assumes only the ones above it.
Keep the Iceberg cheat sheet
open while you read: the metadata tables, row-level operation modes, branches and
tags, the CALL procedures and the maintenance jobs, with the property and default
for each.
The shape of a table
- What Hive tables could not do is the motivation: a Hive table is a directory, and every limitation follows from that. Iceberg replaces the directory with a metadata tree that names every file.
- The three tiers — catalog, metadata, data — because knowing which layer a problem lives in is most of operating one.
- Iceberg catalogs answer one question, which metadata file is current, and how they answer it decides whether concurrent writers are safe.
How data moves through it
- Life of a write: three files and one pointer swap, and what ACID actually means here — measured with two writers hammering the same table.
- Life of a read: four prunes before a byte of data is touched, the last of which only works if your data is sorted.
- Format versions for what v1 to v4 change — chiefly whether a row can be deleted without rewriting a file.
Changing a table without rewriting it
- Schema evolution: why a rename costs nothing, and why field ids are the whole mechanism.
- Hidden partitioning and partition evolution: changing the layout without rewriting history.
Getting data in
- Appends, overwrites and row-level SQL, and the different mark each leaves in the snapshot log.
- Streaming ingestion, where the trigger you choose decides how fast your metadata grows.
- Copy-on-write or merge-on-read:
the same
DELETEproducing two completely different file layouts. - CDC into Iceberg, which works on day one and degrades quietly: three merges, three delete files, 960 dead records, and a row count that never moved.
Operating one
- Time travel and rollback: undoing a bad write in one statement.
- Migrating Hive tables — snapshot first, migrate second, and why that order matters.
- Sorting and clustering: what the per-file min/max statistics can and cannot skip, measured — the same query going from 48 files scanned to 1, and the z-order that made it worse.
- Running Iceberg in production: the settings fixed at create time, and the jobs nobody schedules.
- Iceberg views: a view definition that is versioned the way a table is — and a cross-engine promise that did not hold when it was tested between Spark and Trino.
- Branching, tagging and write-audit-publish: validating a write before anyone can read it, with the copy removed — plus the read redirect that catches people, and why there is no merge.
Two posts sit alongside this path rather than in it: table maintenance for Iceberg and Hudi, on why compaction reclaims nothing until a separate expiry job runs, and one table, three formats, on reading the same Parquet files as Iceberg, Delta and Hudi.
How do I run a table format on Spark?
The posts where the two halves of the site meet.
- Apache Iceberg through Spark Connect, including the session extension that breaks time travel.
- Apache Hudi through Spark Connect, running the full table surface from a thin client.
- Spark and Hudi logging, for making a write explain itself when it stalls.
- Iceberg on Apache Polaris, for a managed REST catalog instead of Hive, against two storage backends.
How do I run any of this locally?
Everything measured on this site was measured on a laptop. These are the environments that made that possible, in increasing order of how much they bring up.
- A local Iceberg playground is the smallest thing that works: Spark, MinIO and a REST catalog in one compose file, when Iceberg is all you need.
- Run every demo on this blog locally is the full stack — MinIO, a Hive Metastore, Spark, Kafka and Trino — plus the demo scripts that reproduce the numbers in these posts on your own machine.
Which tool do I reach for?
Small, self-contained utilities and command-line notes.
- Spark Configuration Generator
and Spark Submit Command Formatter
for building and reading long
spark-submitlines. - Spark Submit generator for Iceberg for getting the catalog flags right the first time.
- Inspecting Parquet files with parquet-cli for looking inside a file before blaming the engine.
- Kerberos setup in Linux for the authentication step that blocks everything else.
Everything, from basic to advanced
- Apache Iceberg 1.12.0: what’s new, breaking changes, and upgrade guide Oct 2026
- Data Modeling for Data Engineers: From Foundations to Lakehouse Design Oct 2026
- Spark on Kubernetes on your Mac: kind, spark-submit, the Spark Operator and Iceberg Sep 2026
- Apache Spark architecture: what actually happens when you call an action Sep 2026
- Apache Spark: the complete cheat sheet Sep 2026
- Open table formats in practice: what actually differs between Iceberg, Hudi and Delta Lake Sep 2026
- Run every demo on this blog locally: one compose stack for Hudi, Iceberg and Delta Sep 2026
- What Hive tables could not do, and why Iceberg exists Sep 2026
- Apache Hudi: an introduction to its architecture, features and what makes it different Sep 2026
- A local Iceberg playground: Spark, MinIO and a REST catalog in one compose file Sep 2026
- Inspecting Parquet files from the command line with parquet-cli Sep 2026
- Spark Submit Command Formatter tool Sep 2026
- Spark Configuration Generator tool Sep 2026
- The three tiers of an Iceberg table: catalog, metadata, data Sep 2026
- Apache Iceberg on Spark: the complete cheat sheet Sep 2026
- Inside Spark’s Catalyst optimizer: how a query gets rewritten, rule by rule Sep 2026
- Adaptive query execution in Apache Spark: what it re-plans, and what it leaves alone Sep 2026
- Apache Spark joins in depth: the five strategies and how one gets picked Aug 2026
- Spark memory management: every region, and the arithmetic that sizes it Aug 2026
- Apache Hudi architecture: what actually happens when you upsert a row Aug 2026
- The Hudi index: how one lookup decides the cost of every upsert Aug 2026
- Hudi Merge-on-Read: the write you skipped, and the read that pays for it Aug 2026
- Apache Hudi on Spark: the complete cheat sheet Aug 2026
- Apache Iceberg architecture: what actually happens when you commit Aug 2026
- Iceberg catalogs: the one pointer that makes a table a table Aug 2026
- Life of a write query in Iceberg: three files, one pointer swap, and what ACID means here Aug 2026
- Life of a read query in Iceberg: four prunes before a byte of data is read Aug 2026
- Iceberg format versions: what v1, v2, v3 and v4 actually change Aug 2026
- Schema evolution in Iceberg: why a rename costs nothing Aug 2026
- Hidden partitioning and partition evolution: changing layout without rewriting history Aug 2026
- Loading data into Iceberg: appends, overwrites and row-level SQL Aug 2026
- Apache XTable incremental sync: how to keep your conversions fast Jul 2026
- Apache XTable cheat sheet: one table, every format Jul 2026
- Spark shuffle internals: the three writers, and which one your job gets Jul 2026
- Data skew in Apache Spark: how to see it, and nine ways to fix it Jul 2026
- Spark performance tuning: seventeen techniques, and how to prove each one worked Jul 2026
- Native Spark execution: Apache Gluten and DataFusion Comet compared Jul 2026
- Every Apache Spark release, and the problem each one was built to solve Jul 2026
- Spark Connect: the driver becomes a server, and your application becomes a client Jul 2026
- Structured Streaming internals: the checkpoint, the state store, and what a restart cannot change Jul 2026
- Streaming into Iceberg: availableNow, checkpoints, and the snapshot count nobody watches Jul 2026
- Copy-on-write or merge-on-read: choosing how an Iceberg table pays for a delete Jul 2026
- CDC into Iceberg: the delete files nobody counted Jul 2026
- Time travel and rollback in Iceberg: undoing a bad write in one statement Jul 2026
- Apache Iceberg through Spark Connect: the full surface, and the extension that breaks time travel Jul 2026
- Apache Hudi through Spark Connect: running the full table surface from a thin client Jul 2026
- Running Iceberg on Apache Polaris: one compose file, two storage backends Jul 2026
- Hudi key generators: the one decision you cannot change later Jun 2026
- Spark Submit Command generator using Iceberg Catalog Jun 2026
- Branching, tagging and write-audit-publish in Iceberg: CI for a table Jun 2026
- Table maintenance for Iceberg and Hudi: compaction, expiry, and the space you did not reclaim Jun 2026
- Concurrency control in Hudi: the lock provider that protects nothing Jun 2026
- One table, three formats: Iceberg, Delta and Hudi interoperability Jun 2026
- Iceberg views: a versioned view definition, and the cross-engine promise tested Jun 2026
- Migrating Hive tables to Iceberg: snapshot first, migrate second Jun 2026
- Sorting and clustering in Iceberg: what min/max statistics can and cannot skip Jun 2026
- Running Iceberg in production: file sizes, object-storage paths, and the jobs you must schedule Jun 2026
- Incremental processing in Hudi: reading only what changed, and the timestamp that betrays you Jun 2026
- Spark and Hudi logging: making a write explain itself Jun 2026
- A Spark JVM playbook: custom logging, class loading, stack size and proxies Jun 2026
- Logging a Spark stream: the job you cannot restart to debug Jun 2026
- Kerberos Setup in Linux Jun 2026
- Top Apache Spark interview questions, with answers Jun 2026