Prev Next

Database / Apache Cassandra Intermediate and Advanced interview questions

How do you model time-series data in Cassandra?

Good time-series modeling in Cassandra centers on two goals: keep partitions bounded in size, and keep the most commonly queried time range fast to retrieve.

CREATE TABLE sensor_readings (
  sensor_id text,
  day       text,        -- bucket, e.g. '2026-07-21'
  reading_time timestamp,
  value     double,
  PRIMARY KEY ((sensor_id, day), reading_time)
) WITH CLUSTERING ORDER BY (reading_time DESC);

  1. Bucket the partition key by a time unit (day, hour, or week depending on write volume) combined with the natural entity id, so no single partition grows without bound.
  2. Cluster by timestamp, typically descending, so the most recent readings — the ones usually queried — are at the front of the partition and returned fastest.
  3. Use TTLs to expire old data automatically if it doesn't need to be kept forever, pairing well with Time Window Compaction Strategy so whole SSTables age out together.
  4. Choose bucket size based on write rate: a high-frequency sensor might bucket hourly, while a low-frequency one might bucket monthly, so each bucket stays a reasonable, similar size.

This pattern — bucketed partition key, descending clustering key, TTL, TWCS — is the standard playbook for metrics, logs, events, and IoT data in Cassandra.

Why is the partition key bucketed by time unit in a time-series model?
Why is clustering order typically set to DESC on the timestamp column?

Invest now in Acorns!!! 🚀 Join Acorns and get your $5 bonus!
Acorns Logo

Invest now in Acorns!!! 🚀
Join Acorns and get your $5 bonus!

Earn passively and while sleeping

Acorns is a micro-investing app that automatically invests your "spare change" from daily purchases into diversified, expert-built portfolios of ETFs. It is designed for beginners, allowing you to start investing with as little as $5. The service automates saving and investing. Disclosure: I may receive a referral bonus.

Robinhood Logo

Invest now!!! Get Free equity stock (US, UK only)!

Use Robinhood app to invest in stocks. It is safe and secure. Use the Referral link to claim your free stock when you sign up!.

The Robinhood app makes it easy to trade stocks, crypto and more.


Webull Logo

Webull! Receive free stock by signing up using the link: Webull signup.

More Related questions...

What is the difference between a partition key and a clustering key? How does Cassandra achieve tunable consistency? What is the difference between consistency levels ONE, QUORUM, and ALL? Explain the write path in Cassandra? Explain the read path in Cassandra? What is the role of a coordinator node in Cassandra? What is the gossip protocol in Cassandra? How does Cassandra detect node failure? What are virtual nodes (vnodes) and why does Cassandra use them? What is consistent hashing and how does Cassandra use it? What is a partitioner in Cassandra? What is a snitch in Cassandra and what does it do? What is the difference between SimpleStrategy and NetworkTopologyStrategy? What is hinted handoff in Cassandra? What is read repair in Cassandra? What is the difference between hinted handoff and read repair? What is anti-entropy repair and why is it needed? What is the difference between full repair and incremental repair? What are tombstones in Cassandra? Why can excessive tombstones degrade Cassandra performance? What is gc_grace_seconds and why does it matter? What are the different compaction strategies available in Cassandra? When would you choose Leveled Compaction Strategy over Size-Tiered Compaction Strategy? What is Time Window Compaction Strategy used for? What are lightweight transactions (LWT) in Cassandra? Why are lightweight transactions expensive in Cassandra? What role does the Paxos protocol play in Cassandra's lightweight transactions? What are secondary indexes in Cassandra, and when should you avoid them? What is a materialized view in Cassandra? What is the difference between a secondary index and a materialized view? What is SASI (SSTable Attached Secondary Index) in Cassandra? What are User Defined Types (UDTs) in Cassandra? What are counter columns in Cassandra and what are their limitations? What is a wide partition in Cassandra and why is it a problem? How do you model time-series data in Cassandra? What is the ALLOW FILTERING clause and why is it risky? What is a batch statement in Cassandra, and what's the difference between logged and unlogged batches? Why shouldn't Cassandra batches be used to improve write throughput? What is Change Data Capture (CDC) in Cassandra? What is speculative retry in Cassandra? What is token awareness in Cassandra drivers? What is the role of Merkle trees in Cassandra's repair process? How do you add a new node to a Cassandra cluster? What is nodetool cleanup used for? How do you handle consistency across multiple datacenters in Cassandra?
Show more question and Answers...

Integration

Comments & Discussions