Skip to content
RR Ranga Reddy Senior Data Engineer
  • Home
  • Guides
  • Archive
  • Categories
  • Links
  • About Me

Posts, page 7

Newest first. 46 posts in total.

Hudi

Apache Hudi architecture: what actually happens when you upsert a row

Sep 12, 2026 21 min read
  • #Hudi
  • #Architecture
  • #Internals
  • #Compaction
  • #Indexing
Spark

Apache Spark architecture: what actually happens when you call an action

Sep 11, 2026 29 min read
  • #Spark
  • #Architecture
  • #Internals
  • #Shuffle
  • #Scheduling
Spark

Apache Spark: the complete cheat sheet

Sep 10, 2026 33 min read
  • #Spark
  • #SQL
  • #Tuning
  • #Streaming
  • #Reference
Hudi

The Hudi index: how one lookup decides the cost of every upsert

Sep 10, 2026 16 min read
  • #Hudi
  • #Indexing
  • #Metadata
  • #Upserts
Iceberg

Apache Iceberg on Spark: the complete cheat sheet

Sep 9, 2026 21 min read
  • #Iceberg
  • #Lakehouse
  • #Spark
  • #SQL
  • #Reference
Hudi

Apache Hudi on Spark: the complete cheat sheet

Sep 8, 2026 35 min read
  • #Hudi
  • #Lakehouse
  • #Spark
  • #PySpark
  • #Reference
First Previous Page 7 of 8 Next Last

Most read

  1. CDC Into Iceberg views
  2. Streaming Into Iceberg views
  3. Iceberg Playground views
  4. Iceberg Format Versions views
  5. Iceberg Write Path views
  6. Copy-on-Write vs Merge-on-Read views
  7. Iceberg In Production views
  8. Hive To Iceberg views
  9. Iceberg Read Path views
  10. Why Iceberg Exists views
  11. Table Maintenance views
  12. Structured Streaming views
  13. Spark Performance Tuning views
  14. Spark Shuffle Internals views
  15. DataFusion Comet views
  16. Iceberg on Polaris views
  17. Spark Submit for Iceberg views
  18. Shuffle Partition Generator views
  19. Parquet Command Line Tools views
  20. Spark Submit Formatter views
  21. Spark Configuration Generator views
  22. Kerberos Setup on Linux views
  23. Streaming Query Logging views
  24. Spark Release History views
  25. Spark Memory Management views
  26. Spark Adaptive Query Execution views
  27. Fixing Data Skew views
  28. Inside the Catalyst Optimizer views
  29. Spark Connect with Apache Iceberg views
  30. Spark Connect with Apache Hudi views
  31. Join Strategies in Apache Spark views
  32. Top Apache Spark Interview Questions views
  33. Hudi Merge-on-Read views
  34. Spark and Hudi Logging views
  35. Apache XTable Cheat Sheet views
  36. Inside Apache Iceberg Architecture views
  37. Inside Apache Hudi Architecture views
  38. Inside Apache Spark Architecture views
  39. Apache Spark Cheat Sheet views
  40. Hudi Index Types views
  41. Apache Iceberg Cheat Sheet views
  42. Apache Hudi Cheat Sheet views
  43. Getting Started with Apache Hudi views
  44. Hudi vs Iceberg vs Delta Lake views
  45. XTable Incremental Sync views
  46. Spark JVM Troubleshooting views

Categories

  • Hudi 6
  • Iceberg 13
  • Lakehouse 4
  • Linux 1
  • Spark 21
  • Tools 1

Tags

AQE2 Architecture3 Arrow1 CDC1 Catalog3 Catalyst2 Checkpoints1 Cleaning1 Codegen1 Comet1 Commits1 Compaction5 Compatibility1 Concurrency2 DataFusion1 Defaults1 Deletes3 Delta3 Docker1 Generator1 Hive2 Hudi11 Iceberg19 Indexing3 Internals9 Interview1 Joins3 Kerberos1 Lakehouse8 Lineage1 Linux1 Log4j2 Logging3 Memory1 Metadata3 Migration2 MinIO2 Optimizer1 Paimon1 Parquet2 Performance5 Polaris1 Production1 Pruning1 PySpark1 Query1 Reference5 Releases1 SQL2 Scenarios1 Scheduling1 Security1 Serialization1 Shuffle3 Skew1 Snapshots3 Spark27 SparkConnect2 State1 Storage1 Streaming6 Timeline1 Tools1 Troubleshoot2 Tuning6 Upserts1 Utilities4 Watermarks1 XTable2
RR Ranga Reddy

Hands-on notes, production troubleshooting and interactive tools for Apache Hudi, Iceberg, Spark, Kafka and the lakehouse stack underneath them.

Explore

  • Home
  • Guides
  • Archive
  • Categories
  • Links
  • About Me

You can contact me

  • GitHub
  • LinkedIn
  • X
  • Stack Overflow
  • Email

© 2026 Ranga Reddy. All rights reserved.

visits · unique visitors See all stats