Integration / Apache Kafka Interview questions
How does Kafka Streams manage local state using state stores?
A state store is a local, embedded key-value store (RocksDB-backed by default, though an in-memory option exists) that a Kafka Streams application uses to hold running state for stateful operations — aggregations, counts, joins — that need to remember something across records rather than processing each one independently.
builder.stream("orders") .groupByKey() .count(Materialized.as("order-counts-store"));
The critical piece for fault tolerance is that every state store is continuously backed by an internal, compacted changelog topic — every update to the local store is also written to that topic, so if the application instance crashes or is rebalanced to a different machine, the new instance can rebuild the exact same state by replaying the changelog rather than losing that accumulated state entirely. Because the changelog uses compaction, replaying it to rebuild state reads only the latest value per key rather than the full historical update stream, keeping recovery time proportional to the number of distinct keys rather than the total number of updates ever made.
More Related questions...