All posts

Hudi key generators: the one decision you cannot change later

The key generator decides what a record's identity is and where it lands on disk. Six of them measured side by side, including the two that produce different keys for the same field, and what happens when you try to switch one on a table that already exists.

9 min read Hudi

TL;DR

  • The key generator turns a row into two things: the record key that decides identity for upserts, and the partition path that decides where it is written.
  • SimpleKeyGenerator and ComplexKeyGenerator produce different keys for the same field. Simple gives 1; Complex gives id:1 — the field name is part of the key.
  • That holds even with a single field, so choosing Complex “in case I add a field later” silently changes your key format from day one.
  • You cannot switch afterwards. Hudi records the generator in table config and refuses a mismatched write outright: Config conflict(key, current value, existing value): KeyGenerator. That is a good failure — the alternative would be every record becoming a new record.
  • hive_style_partitioning changes the path from IN to country=IN. It is a layout decision that affects every engine reading the table afterwards.
  • TimestampBasedKeyGenerator turns an epoch column into a date path — 1772323200000 becomes 2026/03/01 — and fails the write outright if you omit its output format rather than writing something useless.

Most Hudi configuration can be changed. Compaction can be retuned, indexes can be swapped, file sizes can be adjusted, and the table carries on. The key generator is the exception, and it is decided in the first write.

It is worth a few minutes up front because it defines the two things everything else depends on: what makes two rows “the same row”, and what the directory layout is.

What a key generator produces

Every Hudi record carries metadata columns, and two of them come from the key generator:

flowchart LR
    R["incoming row<br/>id=1, country=IN, city=mumbai"] --> KG["key generator"]
    KG --> K["_hoodie_record_key<br/>identity for upserts"]
    KG --> P["_hoodie_partition_path<br/>where it is written"]

The record key is what the index looks up to decide whether an incoming row is an insert or an update. The partition path is the directory. Get either wrong and the table is wrong in a way that is expensive to fix.

Six configurations, measured

The same three rows — (1, IN, mumbai), (2, IN, delhi), (3, US, austin) — written six ways, reading back the two metadata columns:

Configuration _hoodie_record_key _hoodie_partition_path
SimpleKeyGenerator on id, partition country 1 IN
ComplexKeyGenerator on id,country, partition country,city id:1,country:IN IN/mumbai
ComplexKeyGenerator on id alone id:1 IN
NonpartitionedKeyGenerator on id 1 `` (empty)
SimpleKeyGenerator + hive_style_partitioning=true 1 country=IN
TimestampBasedKeyGenerator on an epoch-millis column 1 2026/03/01

Rows one and three are the ones to look at twice.

Simple and Complex are not interchangeable

For the same single field id, the two generators write different keys:

SimpleKeyGenerator    ->  '1'
ComplexKeyGenerator   ->  'id:1'

Complex encodes the field name into the key. That is what lets it compose several fields unambiguously — id:1,country:IN cannot be confused with a different pairing — and it means the format is not a superset of Simple’s, it is a different format.

The practical consequence catches people who are being careful. “I will use Complex even though I only have one key field, so I can add another later without changing anything” produces id:1 from the first write onwards. Adding a second field later changes the key again, to id:1,country:IN, so the flexibility you thought you bought is not there either.

Pick Simple for a single key field. Pick Complex when you genuinely have a composite key.

You cannot change it later

Writing a table with Simple, then attempting an upsert of the same rows with Complex, does not corrupt anything — Hudi refuses:

org.apache.hudi.exception.HoodieException: Config conflict(key, current value, existing value):
KeyGenerator:  org.apache.hudi.keygen.ComplexKeyGenerator  org.apache.hudi.keygen.SimpleKeyGenerator

This is the right behaviour, and worth understanding rather than working around. The generator is stored in hoodie.properties as part of the table’s identity. If the write were allowed, every incoming record would compute a key that matches nothing in the index — so every update would become an insert, and the table would quietly double. Failing the write is much better than discovering that later.

The migration path when you genuinely need a different key is the expensive one: write a new table with the desired configuration and backfill it. There is no in-place conversion, which is precisely why this is worth getting right in the first write.

Partition path shapes

Hive-style partitioning changes the directory name from the bare value to a key=value pair:

hive_style_partitioning = false   ->   IN/
hive_style_partitioning = true    ->   country=IN/

The key=value form is what Hive, Spark and Trino discover automatically when reading a directory tree directly, so it is the friendlier layout for anything reading outside Hudi’s own catalog integration. Like the generator, it is a property of how the table was written; changing it later means new partitions in one shape and old ones in another.

Nonpartitioned gives an empty partition path and a flat table. That is the right choice more often than people expect — a table of a few million rows with no natural low-cardinality partition column is better flat than partitioned on something high-cardinality, which produces the small-files problem instead.

Multiple partition fields nest, in the order given: partitionpath.field = country,city yields IN/mumbai.

Timestamps into dates

The most common real requirement is partitioning by day from a timestamp column, which is what TimestampBasedKeyGenerator exists for:

"hoodie.datasource.write.partitionpath.field": "event_ms",
"hoodie.datasource.write.keygenerator.class":
    "org.apache.hudi.keygen.TimestampBasedKeyGenerator",
"hoodie.keygen.timebased.timestamp.type": "EPOCHMILLISECONDS",
"hoodie.keygen.timebased.output.dateformat": "yyyy/MM/dd",
"hoodie.keygen.timebased.timezone": "UTC",
1772323200000  ->  2026/03/01
1772409600000  ->  2026/03/02

Three things to get right, and the first two are the ones that bite:

  • timestamp.type must match the column. EPOCHMILLISECONDS, UNIX_TIMESTAMP (seconds), DATE_STRING, SCALAR and so on. Feeding seconds to a milliseconds configuration produces dates in 1970 with no error.
  • output.dateformat is required. Omitting it fails the write rather than producing something unusable, which is the better outcome but surprises people who expect a default.
  • Set timezone explicitly. Otherwise the partition a row lands in depends on the JVM’s default timezone, so the same data written from two clusters can land in different days. This is the one that produces the “why are there rows in yesterday’s partition” question months later.

Which one to use

Situation Generator
Single key field, single partition field SimpleKeyGenerator
Composite key, or multiple partition fields ComplexKeyGenerator
No sensible partition column NonpartitionedKeyGenerator
Partition by day from a timestamp TimestampBasedKeyGenerator
Partition transform that none of the above expresses CustomKeyGenerator, or your own

CustomKeyGenerator allows per-field types — some fields treated as simple values, others as timestamps — which is how you get country/2026/03/01 from one generator. Writing your own is supported and rarely necessary.

A note on global versus non-global indexes, because the two decisions interact. With a non-global index, uniqueness of the record key is only enforced within a partition, so the same key can exist in two partitions and an update that changes the partition value produces a duplicate rather than a move. A global index enforces uniqueness across the whole table and can handle the move, at the cost of a more expensive lookup. If your partition column can change for a given record, that pushes you toward a global index — and that is a key-generator decision as much as an index one.

Common misconceptions

“Simple and Complex produce the same key for one field.” Simple gives 1, Complex gives id:1. Measured above.

“I can switch generators if I need to.” Hudi refuses the write with a config conflict. The path is a new table and a backfill.

“Using Complex up front keeps my options open.” It fixes a different key format from the first write, and adding a field changes the key again anyway.

“hive_style_partitioning is cosmetic.” It changes the directory layout that every other engine reads.

“Not partitioning is a mistake.” Partitioning on a high-cardinality column is the more common and more expensive mistake. Flat is fine for many tables.

“The timestamp generator infers my format.” It requires output.dateformat and fails without it, and it will silently use the JVM timezone if you do not set one.

Frequently asked questions

Where is the generator recorded? In hoodie.properties under the table’s .hoodie directory, alongside the table type and version. That is the file the config-conflict check reads.

Can I change the partition field without changing the generator? That is also part of the recorded configuration and also refused. New partition layouts mean a new table.

Does the record key have to be unique? Within a partition for a non-global index, across the table for a global one. That choice belongs with the key design, not after it.

What if my key fields contain the separator characters? Complex keys use : and , as structure. Values containing them make keys ambiguous — worth avoiding, or normalising before the write.

Does any of this apply to Iceberg? No. Iceberg has no record key at all; identity is positional and partitioning is a hidden transform on a column that can be evolved later. The immovability described here is specific to Hudi’s upsert model.

Conclusion

The key generator is a small amount of configuration that fixes two large properties of a table: what counts as the same record, and what the directory tree looks like. Hudi will not let you change it afterwards — which is the correct behaviour, since the alternative is every update becoming an insert — so it is worth the few minutes before the first write.

The specific thing to check before that write: if you have one key field, use SimpleKeyGenerator. Reaching for ComplexKeyGenerator to keep options open gives you id:1 instead of 1 permanently and does not actually keep the options open.

References

Trademarks

Apache Hudi, Apache Spark, Apache Hive, Apache Iceberg, Apache and the Apache feather logo are either registered trademarks or trademarks of The Apache Software Foundation in the United States and other countries. Trino is a trademark of the Trino Software Foundation.

Found this useful?

These posts and tools are free. If one saved you an afternoon, you can buy me a coffee.

Buy me a coffee