Skip to content
RR
Ranga Reddy
Senior Data Engineer/Architect
Home
Guides
Archive
Categories
Links
About Me
Categories
6 categories covering 61 posts.
Hudi
10
Iceberg
23
Lakehouse
6
Linux
1
Spark
20
Tools
1
Hudi
10
2026-09-17
Apache Hudi: an introduction to its architecture, features and what makes it different
#Hudi
#Lakehouse
#Spark
#Indexing
#Streaming
2026-08-24
Apache Hudi architecture: what actually happens when you upsert a row
#Hudi
#Architecture
#Internals
#Compaction
#Indexing
2026-08-22
The Hudi index: how one lookup decides the cost of every upsert
#Hudi
#Indexing
#Metadata
#Upserts
2026-08-20
Hudi Merge-on-Read: the write you skipped, and the read that pays for it
#Hudi
#Compaction
#Timeline
#Concurrency
2026-08-18
Apache Hudi on Spark: the complete cheat sheet
#Hudi
#Lakehouse
#Spark
#PySpark
#Reference
2026-07-03
Apache Hudi through Spark Connect: running the full table surface from a thin client
#Spark
#SparkConnect
#Hudi
#Lakehouse
2026-06-30
Hudi key generators: the one decision you cannot change later
#Hudi
#KeyGenerator
#Partitioning
#Configuration
2026-06-26
Concurrency control in Hudi: the lock provider that protects nothing
#Hudi
#Concurrency
#Locking
#Timeline
#Production
2026-06-20
Incremental processing in Hudi: reading only what changed, and the timestamp that betrays you
#Hudi
#Incremental
#CDC
#Timeline
#Streaming
2026-06-19
Spark and Hudi logging: making a write explain itself
#Hudi
#Spark
#Logging
#Log4j
#Troubleshoot
Iceberg
23
2026-09-19
What Hive tables could not do, and why Iceberg exists
#Iceberg
#Hive
#Metadata
#Lakehouse
2026-09-15
A local Iceberg playground: Spark, MinIO and a REST catalog in one compose file
#Iceberg
#MinIO
#Docker
#Catalog
2026-09-07
The three tiers of an Iceberg table: catalog, metadata, data
#Iceberg
#Metadata
#Manifests
#Architecture
2026-09-05
Apache Iceberg on Spark: the complete cheat sheet
#Iceberg
#Lakehouse
#Spark
#SQL
#Reference
2026-08-16
Apache Iceberg architecture: what actually happens when you commit
#Iceberg
#Architecture
#Internals
#Snapshots
#Catalog
2026-08-14
Iceberg catalogs: the one pointer that makes a table a table
#Iceberg
#Catalog
#REST
#Governance
2026-08-12
Life of a write query in Iceberg: three files, one pointer swap, and what ACID means here
#Iceberg
#Commits
#Concurrency
#Transactions
2026-08-10
Life of a read query in Iceberg: four prunes before a byte of data is read
#Iceberg
#Pruning
#Metadata
#Query
2026-08-08
Iceberg format versions: what v1, v2, v3 and v4 actually change
#Iceberg
#Deletes
#Lineage
#Compatibility
2026-08-06
Schema evolution in Iceberg: why a rename costs nothing
#Iceberg
#Schema
#Evolution
#Compatibility
2026-08-04
Hidden partitioning and partition evolution: changing layout without rewriting history
#Iceberg
#Partitioning
#Pruning
#Transforms
2026-08-02
Loading data into Iceberg: appends, overwrites and row-level SQL
#Iceberg
#Ingestion
#Merge
#Overwrite
2026-07-13
Streaming into Iceberg: availableNow, checkpoints, and the snapshot count nobody watches
#Iceberg
#Streaming
#Checkpoints
#Snapshots
2026-07-11
Copy-on-write or merge-on-read: choosing how an Iceberg table pays for a delete
#Iceberg
#Deletes
#Compaction
#Streaming
2026-07-09
CDC into Iceberg: the delete files nobody counted
#Iceberg
#CDC
#Deletes
#Compaction
2026-07-07
Time travel and rollback in Iceberg: undoing a bad write in one statement
#Iceberg
#Snapshots
#Recovery
#History
2026-07-05
Apache Iceberg through Spark Connect: the full surface, and the extension that breaks time travel
#Spark
#SparkConnect
#Iceberg
#Lakehouse
2026-07-01
Running Iceberg on Apache Polaris: one compose file, two storage backends
#Iceberg
#Polaris
#Catalog
#Spark
#MinIO
2026-06-28
Branching, tagging and write-audit-publish in Iceberg: CI for a table
#Iceberg
#Branching
#Tagging
#WAP
#Production
2026-06-24
Iceberg views: a versioned view definition, and the cross-engine promise tested
#Iceberg
#Views
#Catalogs
#Trino
#Interoperability
2026-06-23
Migrating Hive tables to Iceberg: snapshot first, migrate second
#Iceberg
#Migration
#Hive
#Parquet
2026-06-22
Sorting and clustering in Iceberg: what min/max statistics can and cannot skip
#Iceberg
#Performance
#Sorting
#Clustering
#Maintenance
2026-06-21
Running Iceberg in production: file sizes, object-storage paths, and the jobs you must schedule
#Iceberg
#Production
#Tuning
#Storage
Lakehouse
6
2026-09-23
Open table formats in practice: what actually differs between Iceberg, Hudi and Delta Lake
#Iceberg
#Hudi
#Delta
#Lakehouse
#Spark
2026-09-21
Run every demo on this blog locally: one compose stack for Hudi, Iceberg and Delta
#Docker
#Iceberg
#Hudi
#Delta
2026-07-31
Apache XTable incremental sync: how to keep your conversions fast
#XTable
#Hudi
#Iceberg
#Delta
#Lakehouse
2026-07-29
Apache XTable cheat sheet: one table, every format
#XTable
#Hudi
#Iceberg
#Delta
#Paimon
#Reference
2026-06-27
Table maintenance for Iceberg and Hudi: compaction, expiry, and the space you did not reclaim
#Iceberg
#Hudi
#Compaction
#Cleaning
2026-06-25
One table, three formats: Iceberg, Delta and Hudi interoperability
#Iceberg
#Delta
#Hudi
#XTable
Linux
1
2026-06-13
Kerberos Setup in Linux
#Linux
#Kerberos
#Security
Spark
20
2026-09-30
Spark on Kubernetes on your Mac: kind, spark-submit, the Spark Operator and Iceberg
#Spark
#Kubernetes
#Operator
#Iceberg
2026-09-27
Apache Spark architecture: what actually happens when you call an action
#Spark
#Architecture
#Internals
#Shuffle
#Scheduling
2026-09-25
Apache Spark: the complete cheat sheet
#Spark
#SQL
#Tuning
#Streaming
#Reference
2026-09-11
Spark Submit Command Formatter tool
#Spark
#Utilities
2026-09-09
Spark Configuration Generator tool
#Spark
#Utilities
#Generator
2026-09-03
Inside Spark's Catalyst optimizer: how a query gets rewritten, rule by rule
#Spark
#Catalyst
#Optimizer
#Internals
#Codegen
2026-09-01
Adaptive query execution in Apache Spark: what it re-plans, and what it leaves alone
#Spark
#AQE
#Performance
#Tuning
#Internals
2026-08-30
Apache Spark joins in depth: the five strategies and how one gets picked
#Spark
#Joins
#Catalyst
#Performance
#Internals
2026-08-28
Spark memory management: every region, and the arithmetic that sizes it
#Spark
#Memory
#Tuning
#Internals
#Performance
2026-07-27
Spark shuffle internals: the three writers, and which one your job gets
#Spark
#Shuffle
#Serialization
#Internals
2026-07-25
Data skew in Apache Spark: how to see it, and nine ways to fix it
#Spark
#Skew
#Performance
#Joins
#Tuning
2026-07-23
Spark performance tuning: seventeen techniques, and how to prove each one worked
#Spark
#Performance
#Joins
#AQE
2026-07-21
Native Spark execution: Apache Gluten and DataFusion Comet compared
#Spark
#Gluten
#Comet
#Velox
2026-07-19
Every Apache Spark release, and the problem each one was built to solve
#Spark
#Releases
#Migration
#Defaults
#Internals
2026-07-17
Spark Connect: the driver becomes a server, and your application becomes a client
#SparkConnect
#gRPC
#Arrow
#PySpark
2026-07-15
Structured Streaming internals: the checkpoint, the state store, and what a restart cannot change
#Spark
#Streaming
#State
#Watermarks
2026-06-29
Spark Submit Command generator using Iceberg Catalog
#Spark
#Utilities
#Iceberg
2026-06-17
A Spark JVM playbook: custom logging, class loading, stack size and proxies
#Spark
#Troubleshoot
#Logging
2026-06-15
Logging a Spark stream: the job you cannot restart to debug
#Spark
#Streaming
#Logging
#Log4j
2026-06-11
Top Apache Spark interview questions, with answers
#Spark
#Interview
#Reference
#Scenarios
Tools
1
2026-09-13
Inspecting Parquet files from the command line with parquet-cli
#Tools
#Parquet