Database / ScyllaDB Interview questions
Why should you avoid large batch statements in ScyllaDB?
CQL's BATCH statement groups multiple writes together, but unless every statement in the batch shares the same partition key, ScyllaDB has to coordinate the batch as a distributed operation across multiple partitions, which adds a lot of overhead compared to sending the same writes individually.
A large multi-partition batch forces a single coordinator node to buffer and manage writes destined for many different replica sets at once, increasing memory pressure and the risk of timeouts, and it doesn't actually make those writes atomic in the way a relational transaction would - it just batches network round trips. If all statements in a batch do share one partition key, the batch is efficient and reasonable, since it's genuinely a single-partition operation. The general guidance is to keep batches small, single-partition, and use them for genuine logical grouping rather than as a shortcut to bulk-load unrelated writes.
More Related questions...