Database / Apache Cassandra Intermediate and Advanced interview questions
What is Change Data Capture (CDC) in Cassandra?
Change Data Capture lets external systems consume a stream of the mutations written to a Cassandra table, without polling the table itself — useful for feeding data pipelines, search indexes, caches, or event-driven architectures.
ALTER TABLE orders WITH cdc = true;
- Once enabled on a table, commit log segments containing mutations for that table are moved to a dedicated CDC directory once they'd normally be recycled, instead of being discarded.
- An external agent or connector (commonly paired with a Kafka connector) reads and parses these commit log segments to extract the change events.
- Cassandra enforces a size limit (
cdc_total_space_in_mb) on the CDC directory; if consumers fall behind and it fills up, writes to CDC-enabled tables can be rejected until space frees up.
CDC is lower-level than a managed change-stream product: it hands you raw commit log data to parse, so most teams pair it with existing connector tooling rather than writing a parser from scratch. It's the standard mechanism for building real-time pipelines off of Cassandra without adding read load to the cluster itself.
More Related questions...