Hudi key generators: the one decision you cannot change later
The key generator decides what a record's identity is and where it lands on disk. Six of them measured side by side, including the two that produce different keys for the same field, and what happens when you try to switch one on a table that already exists.
- What a key generator produces
- Six configurations, measured
- Simple and Complex are not interchangeable
- You cannot change it later
- Partition path shapes
- Timestamps into dates
- Which one to use
- Common misconceptions
- Frequently asked questions
- Conclusion
- References
- Trademarks
TL;DR
- The key generator turns a row into two things: the record key that decides identity for upserts, and the partition path that decides where it is written.
SimpleKeyGeneratorandComplexKeyGeneratorproduce different keys for the same field. Simple gives1; Complex givesid:1— the field name is part of the key.- That holds even with a single field, so choosing Complex “in case I add a field later” silently changes your key format from day one.
- You cannot switch afterwards. Hudi records the generator in table config and refuses a mismatched write outright:
Config conflict(key, current value, existing value): KeyGenerator. That is a good failure — the alternative would be every record becoming a new record.hive_style_partitioningchanges the path fromINtocountry=IN. It is a layout decision that affects every engine reading the table afterwards.TimestampBasedKeyGeneratorturns an epoch column into a date path —1772323200000becomes2026/03/01— and fails the write outright if you omit its output format rather than writing something useless.
Most Hudi configuration can be changed. Compaction can be retuned, indexes can be swapped, file sizes can be adjusted, and the table carries on. The key generator is the exception, and it is decided in the first write.
It is worth a few minutes up front because it defines the two things everything else depends on: what makes two rows “the same row”, and what the directory layout is.
What a key generator produces
Every Hudi record carries metadata columns, and two of them come from the key generator:
flowchart LR
R["incoming row<br/>id=1, country=IN, city=mumbai"] --> KG["key generator"]
KG --> K["_hoodie_record_key<br/>identity for upserts"]
KG --> P["_hoodie_partition_path<br/>where it is written"]
The record key is what the index looks up to decide whether an incoming row is an insert or an update. The partition path is the directory. Get either wrong and the table is wrong in a way that is expensive to fix.
Six configurations, measured
The same three rows — (1, IN, mumbai), (2, IN, delhi), (3, US, austin) —
written six ways, reading back the two metadata columns:
| Configuration | _hoodie_record_key |
_hoodie_partition_path |
|---|---|---|
SimpleKeyGenerator on id, partition country |
1 |
IN |
ComplexKeyGenerator on id,country, partition country,city |
id:1,country:IN |
IN/mumbai |
ComplexKeyGenerator on id alone |
id:1 |
IN |
NonpartitionedKeyGenerator on id |
1 |
`` (empty) |
SimpleKeyGenerator + hive_style_partitioning=true |
1 |
country=IN |
TimestampBasedKeyGenerator on an epoch-millis column |
1 |
2026/03/01 |
Rows one and three are the ones to look at twice.
Simple and Complex are not interchangeable
For the same single field id, the two generators write different keys:
SimpleKeyGenerator -> '1'
ComplexKeyGenerator -> 'id:1'
Complex encodes the field name into the key. That is what lets it compose
several fields unambiguously — id:1,country:IN cannot be confused with a
different pairing — and it means the format is not a superset of Simple’s, it is
a different format.
The practical consequence catches people who are being careful. “I will use
Complex even though I only have one key field, so I can add another later without
changing anything” produces id:1 from the first write onwards. Adding a second
field later changes the key again, to id:1,country:IN, so the flexibility you
thought you bought is not there either.
Pick Simple for a single key field. Pick Complex when you genuinely have a composite key.
You cannot change it later
Writing a table with Simple, then attempting an upsert of the same rows with Complex, does not corrupt anything — Hudi refuses:
org.apache.hudi.exception.HoodieException: Config conflict(key, current value, existing value):
KeyGenerator: org.apache.hudi.keygen.ComplexKeyGenerator org.apache.hudi.keygen.SimpleKeyGenerator
This is the right behaviour, and worth understanding rather than working around.
The generator is stored in hoodie.properties as part of the table’s identity. If
the write were allowed, every incoming record would compute a key that matches
nothing in the index — so every update would become an insert, and the table would
quietly double. Failing the write is much better than discovering that later.
The migration path when you genuinely need a different key is the expensive one: write a new table with the desired configuration and backfill it. There is no in-place conversion, which is precisely why this is worth getting right in the first write.
Partition path shapes
Hive-style partitioning changes the directory name from the bare value to a
key=value pair:
hive_style_partitioning = false -> IN/
hive_style_partitioning = true -> country=IN/
The key=value form is what Hive, Spark and Trino discover automatically when
reading a directory tree directly, so it is the friendlier layout for anything
reading outside Hudi’s own catalog integration. Like the generator, it is a
property of how the table was written; changing it later means new partitions in
one shape and old ones in another.
Nonpartitioned gives an empty partition path and a flat table. That is the right choice more often than people expect — a table of a few million rows with no natural low-cardinality partition column is better flat than partitioned on something high-cardinality, which produces the small-files problem instead.
Multiple partition fields nest, in the order given:
partitionpath.field = country,city yields IN/mumbai.
Timestamps into dates
The most common real requirement is partitioning by day from a timestamp column,
which is what TimestampBasedKeyGenerator exists for:
"hoodie.datasource.write.partitionpath.field": "event_ms",
"hoodie.datasource.write.keygenerator.class":
"org.apache.hudi.keygen.TimestampBasedKeyGenerator",
"hoodie.keygen.timebased.timestamp.type": "EPOCHMILLISECONDS",
"hoodie.keygen.timebased.output.dateformat": "yyyy/MM/dd",
"hoodie.keygen.timebased.timezone": "UTC",
1772323200000 -> 2026/03/01
1772409600000 -> 2026/03/02
Three things to get right, and the first two are the ones that bite:
timestamp.typemust match the column.EPOCHMILLISECONDS,UNIX_TIMESTAMP(seconds),DATE_STRING,SCALARand so on. Feeding seconds to a milliseconds configuration produces dates in 1970 with no error.output.dateformatis required. Omitting it fails the write rather than producing something unusable, which is the better outcome but surprises people who expect a default.- Set
timezoneexplicitly. Otherwise the partition a row lands in depends on the JVM’s default timezone, so the same data written from two clusters can land in different days. This is the one that produces the “why are there rows in yesterday’s partition” question months later.
Which one to use
| Situation | Generator |
|---|---|
| Single key field, single partition field | SimpleKeyGenerator |
| Composite key, or multiple partition fields | ComplexKeyGenerator |
| No sensible partition column | NonpartitionedKeyGenerator |
| Partition by day from a timestamp | TimestampBasedKeyGenerator |
| Partition transform that none of the above expresses | CustomKeyGenerator, or your own |
CustomKeyGenerator allows per-field types — some fields treated as simple
values, others as timestamps — which is how you get country/2026/03/01 from one
generator. Writing your own is supported and rarely necessary.
A note on global versus non-global indexes, because the two decisions interact. With a non-global index, uniqueness of the record key is only enforced within a partition, so the same key can exist in two partitions and an update that changes the partition value produces a duplicate rather than a move. A global index enforces uniqueness across the whole table and can handle the move, at the cost of a more expensive lookup. If your partition column can change for a given record, that pushes you toward a global index — and that is a key-generator decision as much as an index one.
Common misconceptions
“Simple and Complex produce the same key for one field.” Simple gives 1,
Complex gives id:1. Measured above.
“I can switch generators if I need to.” Hudi refuses the write with a config conflict. The path is a new table and a backfill.
“Using Complex up front keeps my options open.” It fixes a different key format from the first write, and adding a field changes the key again anyway.
“hive_style_partitioning is cosmetic.” It changes the directory layout that every other engine reads.
“Not partitioning is a mistake.” Partitioning on a high-cardinality column is the more common and more expensive mistake. Flat is fine for many tables.
“The timestamp generator infers my format.” It requires output.dateformat
and fails without it, and it will silently use the JVM timezone if you do not set
one.
Frequently asked questions
Where is the generator recorded?
In hoodie.properties under the table’s .hoodie directory, alongside the table
type and version. That is the file the config-conflict check reads.
Can I change the partition field without changing the generator? That is also part of the recorded configuration and also refused. New partition layouts mean a new table.
Does the record key have to be unique? Within a partition for a non-global index, across the table for a global one. That choice belongs with the key design, not after it.
What if my key fields contain the separator characters?
Complex keys use : and , as structure. Values containing them make keys
ambiguous — worth avoiding, or normalising before the write.
Does any of this apply to Iceberg? No. Iceberg has no record key at all; identity is positional and partitioning is a hidden transform on a column that can be evolved later. The immovability described here is specific to Hudi’s upsert model.
Conclusion
The key generator is a small amount of configuration that fixes two large properties of a table: what counts as the same record, and what the directory tree looks like. Hudi will not let you change it afterwards — which is the correct behaviour, since the alternative is every update becoming an insert — so it is worth the few minutes before the first write.
The specific thing to check before that write: if you have one key field, use
SimpleKeyGenerator. Reaching for ComplexKeyGenerator to keep options open
gives you id:1 instead of 1 permanently and does not actually keep the options
open.
References
- Hudi key generation for the generator classes and their configuration
- Hudi configurations for
hoodie.datasource.write.keygenerator.*and the timestamp options KeyGeneratorimplementations for what each one actually does- The Hudi index for how the record key is used once it exists, and global versus non-global lookup
- Apache Hudi architecture for where the key lands in a commit
Trademarks
Apache Hudi, Apache Spark, Apache Hive, Apache Iceberg, Apache and the Apache feather logo are either registered trademarks or trademarks of The Apache Software Foundation in the United States and other countries. Trino is a trademark of the Trino Software Foundation.
Found this useful?
These posts and tools are free. If one saved you an afternoon, you can buy me a coffee.