Skip to content
RR
Ranga Reddy
Senior Data Engineer/Architect
Home
Guides
Archive
Categories
Links
About Me
Archive
61 posts across 1 years, ordered from basic to advanced.
2026
61
Sep 30
Spark on Kubernetes on your Mac: kind, spark-submit, the Spark Operator and Iceberg
Spark
#Spark
#Kubernetes
#Operator
#Iceberg
Sep 27
Apache Spark architecture: what actually happens when you call an action
Spark
#Spark
#Architecture
#Internals
#Shuffle
#Scheduling
Sep 25
Apache Spark: the complete cheat sheet
Spark
#Spark
#SQL
#Tuning
#Streaming
#Reference
Sep 23
Open table formats in practice: what actually differs between Iceberg, Hudi and Delta Lake
Lakehouse
#Iceberg
#Hudi
#Delta
#Lakehouse
#Spark
Sep 21
Run every demo on this blog locally: one compose stack for Hudi, Iceberg and Delta
Lakehouse
#Docker
#Iceberg
#Hudi
#Delta
Sep 19
What Hive tables could not do, and why Iceberg exists
Iceberg
#Iceberg
#Hive
#Metadata
#Lakehouse
Sep 17
Apache Hudi: an introduction to its architecture, features and what makes it different
Hudi
#Hudi
#Lakehouse
#Spark
#Indexing
#Streaming
Sep 15
A local Iceberg playground: Spark, MinIO and a REST catalog in one compose file
Iceberg
#Iceberg
#MinIO
#Docker
#Catalog
Sep 13
Inspecting Parquet files from the command line with parquet-cli
Tools
#Tools
#Parquet
Sep 11
Spark Submit Command Formatter tool
Tool
Spark
#Spark
#Utilities
Sep 9
Spark Configuration Generator tool
Tool
Spark
#Spark
#Utilities
#Generator
Sep 7
The three tiers of an Iceberg table: catalog, metadata, data
Iceberg
#Iceberg
#Metadata
#Manifests
#Architecture
Sep 5
Apache Iceberg on Spark: the complete cheat sheet
Iceberg
#Iceberg
#Lakehouse
#Spark
#SQL
#Reference
Sep 3
Inside Spark's Catalyst optimizer: how a query gets rewritten, rule by rule
Spark
#Spark
#Catalyst
#Optimizer
#Internals
#Codegen
Sep 1
Adaptive query execution in Apache Spark: what it re-plans, and what it leaves alone
Spark
#Spark
#AQE
#Performance
#Tuning
#Internals
Aug 30
Apache Spark joins in depth: the five strategies and how one gets picked
Spark
#Spark
#Joins
#Catalyst
#Performance
#Internals
Aug 28
Spark memory management: every region, and the arithmetic that sizes it
Spark
#Spark
#Memory
#Tuning
#Internals
#Performance
Aug 24
Apache Hudi architecture: what actually happens when you upsert a row
Hudi
#Hudi
#Architecture
#Internals
#Compaction
#Indexing
Aug 22
The Hudi index: how one lookup decides the cost of every upsert
Hudi
#Hudi
#Indexing
#Metadata
#Upserts
Aug 20
Hudi Merge-on-Read: the write you skipped, and the read that pays for it
Hudi
#Hudi
#Compaction
#Timeline
#Concurrency
Aug 18
Apache Hudi on Spark: the complete cheat sheet
Hudi
#Hudi
#Lakehouse
#Spark
#PySpark
#Reference
Aug 16
Apache Iceberg architecture: what actually happens when you commit
Iceberg
#Iceberg
#Architecture
#Internals
#Snapshots
#Catalog
Aug 14
Iceberg catalogs: the one pointer that makes a table a table
Iceberg
#Iceberg
#Catalog
#REST
#Governance
Aug 12
Life of a write query in Iceberg: three files, one pointer swap, and what ACID means here
Iceberg
#Iceberg
#Commits
#Concurrency
#Transactions
Aug 10
Life of a read query in Iceberg: four prunes before a byte of data is read
Iceberg
#Iceberg
#Pruning
#Metadata
#Query
Aug 8
Iceberg format versions: what v1, v2, v3 and v4 actually change
Iceberg
#Iceberg
#Deletes
#Lineage
#Compatibility
Aug 6
Schema evolution in Iceberg: why a rename costs nothing
Iceberg
#Iceberg
#Schema
#Evolution
#Compatibility
Aug 4
Hidden partitioning and partition evolution: changing layout without rewriting history
Iceberg
#Iceberg
#Partitioning
#Pruning
#Transforms
Aug 2
Loading data into Iceberg: appends, overwrites and row-level SQL
Iceberg
#Iceberg
#Ingestion
#Merge
#Overwrite
Jul 31
Apache XTable incremental sync: how to keep your conversions fast
Lakehouse
#XTable
#Hudi
#Iceberg
#Delta
#Lakehouse
Jul 29
Apache XTable cheat sheet: one table, every format
Lakehouse
#XTable
#Hudi
#Iceberg
#Delta
#Paimon
#Reference
Jul 27
Spark shuffle internals: the three writers, and which one your job gets
Spark
#Spark
#Shuffle
#Serialization
#Internals
Jul 25
Data skew in Apache Spark: how to see it, and nine ways to fix it
Spark
#Spark
#Skew
#Performance
#Joins
#Tuning
Jul 23
Spark performance tuning: seventeen techniques, and how to prove each one worked
Spark
#Spark
#Performance
#Joins
#AQE
Jul 21
Native Spark execution: Apache Gluten and DataFusion Comet compared
Spark
#Spark
#Gluten
#Comet
#Velox
Jul 19
Every Apache Spark release, and the problem each one was built to solve
Spark
#Spark
#Releases
#Migration
#Defaults
#Internals
Jul 17
Spark Connect: the driver becomes a server, and your application becomes a client
Spark
#SparkConnect
#gRPC
#Arrow
#PySpark
Jul 15
Structured Streaming internals: the checkpoint, the state store, and what a restart cannot change
Spark
#Spark
#Streaming
#State
#Watermarks
Jul 13
Streaming into Iceberg: availableNow, checkpoints, and the snapshot count nobody watches
Iceberg
#Iceberg
#Streaming
#Checkpoints
#Snapshots
Jul 11
Copy-on-write or merge-on-read: choosing how an Iceberg table pays for a delete
Iceberg
#Iceberg
#Deletes
#Compaction
#Streaming
Jul 9
CDC into Iceberg: the delete files nobody counted
Iceberg
#Iceberg
#CDC
#Deletes
#Compaction
Jul 7
Time travel and rollback in Iceberg: undoing a bad write in one statement
Iceberg
#Iceberg
#Snapshots
#Recovery
#History
Jul 5
Apache Iceberg through Spark Connect: the full surface, and the extension that breaks time travel
Iceberg
#Spark
#SparkConnect
#Iceberg
#Lakehouse
Jul 3
Apache Hudi through Spark Connect: running the full table surface from a thin client
Hudi
#Spark
#SparkConnect
#Hudi
#Lakehouse
Jul 1
Running Iceberg on Apache Polaris: one compose file, two storage backends
Iceberg
#Iceberg
#Polaris
#Catalog
#Spark
#MinIO
Jun 30
Hudi key generators: the one decision you cannot change later
Hudi
#Hudi
#KeyGenerator
#Partitioning
#Configuration
Jun 29
Spark Submit Command generator using Iceberg Catalog
Tool
Spark
#Spark
#Utilities
#Iceberg
Jun 28
Branching, tagging and write-audit-publish in Iceberg: CI for a table
Iceberg
#Iceberg
#Branching
#Tagging
#WAP
#Production
Jun 27
Table maintenance for Iceberg and Hudi: compaction, expiry, and the space you did not reclaim
Lakehouse
#Iceberg
#Hudi
#Compaction
#Cleaning
Jun 26
Concurrency control in Hudi: the lock provider that protects nothing
Hudi
#Hudi
#Concurrency
#Locking
#Timeline
#Production
Jun 25
One table, three formats: Iceberg, Delta and Hudi interoperability
Lakehouse
#Iceberg
#Delta
#Hudi
#XTable
Jun 24
Iceberg views: a versioned view definition, and the cross-engine promise tested
Iceberg
#Iceberg
#Views
#Catalogs
#Trino
#Interoperability
Jun 23
Migrating Hive tables to Iceberg: snapshot first, migrate second
Iceberg
#Iceberg
#Migration
#Hive
#Parquet
Jun 22
Sorting and clustering in Iceberg: what min/max statistics can and cannot skip
Iceberg
#Iceberg
#Performance
#Sorting
#Clustering
#Maintenance
Jun 21
Running Iceberg in production: file sizes, object-storage paths, and the jobs you must schedule
Iceberg
#Iceberg
#Production
#Tuning
#Storage
Jun 20
Incremental processing in Hudi: reading only what changed, and the timestamp that betrays you
Hudi
#Hudi
#Incremental
#CDC
#Timeline
#Streaming
Jun 19
Spark and Hudi logging: making a write explain itself
Hudi
#Hudi
#Spark
#Logging
#Log4j
#Troubleshoot
Jun 17
A Spark JVM playbook: custom logging, class loading, stack size and proxies
Spark
#Spark
#Troubleshoot
#Logging
Jun 15
Logging a Spark stream: the job you cannot restart to debug
Spark
#Spark
#Streaming
#Logging
#Log4j
Jun 13
Kerberos Setup in Linux
Linux
#Linux
#Kerberos
#Security
Jun 11
Top Apache Spark interview questions, with answers
Spark
#Spark
#Interview
#Reference
#Scenarios