Database / Apache Cassandra Intermediate and Advanced interview questions
Why shouldn't Cassandra batches be used to improve write throughput?
It's a common misconception carried over from relational databases that batching writes always improves throughput. In Cassandra, batching across multiple partitions usually does the opposite.
- A multi-partition logged batch has to write to the batchlog first, then fan out each individual statement to its own set of replicas — that's more total work than sending the same statements independently, not less.
- The coordinator handling the batch becomes a bottleneck: instead of spreading writes evenly across the cluster, it has to orchestrate every statement in the batch itself.
- Large batches increase the risk of hitting
batch_size_warn_threshold/batch_size_fail_threshold, and can cause coordinator-side memory pressure.
In Cassandra, throughput comes from spreading independent writes across many coordinators and replicas in parallel — typically via async driver calls — not from bundling them together. Batches should be reserved for genuine atomicity needs on a small number of statements, ideally within a single partition, rather than treated as a bulk-loading or performance optimization.
More Related questions...