Cassandra Interview Questions (Top 100 Questions with Answers)
Master Apache Cassandra Interview Questions with production-oriented questions covering Cassandra Architecture, Data Model, Replication, Partitioning, Consistency Levels, CQL, Performance Tuning, Compaction, Repair, and real-world production scenarios.
Introduction
Apache Cassandra is a highly scalable distributed NoSQL database designed for handling massive amounts of data across multiple servers with no single point of failure.
It is widely used in
- Banking
- IoT
- Messaging Systems
- Telecommunication
- Recommendation Engines
- Fraud Detection
- Time-Series Applications
Major companies using Cassandra include
- Netflix
- Apple
- Uber
- Spotify
- Cisco
- eBay
- IBM
This guide contains the Top 100 Cassandra Interview Questions frequently asked in Java, Backend, Cloud, DevOps, and Solution Architect interviews.
Cassandra Interview Roadmap
Cassandra Basics
│
▼
Architecture
│
▼
Data Model
│
▼
Partitioning
│
▼
Replication
│
▼
Consistency
│
▼
Compaction
│
▼
Repair
│
▼
Performance
Cassandra Fundamentals
1. What is Apache Cassandra?
Apache Cassandra is a distributed NoSQL wide-column database designed for high availability and linear scalability.
2. Why Cassandra?
- No Single Point of Failure
- Linear Scalability
- High Availability
- Multi-Data Center Support
- Fault Tolerance
3. Cassandra vs Relational Database?
| Cassandra | SQL Database |
|---|---|
| NoSQL | Relational |
| Wide Column | Tables |
| Horizontal Scaling | Mostly Vertical Scaling |
| Eventual Consistency | Strong Consistency |
4. Cassandra vs MongoDB?
| Cassandra | MongoDB |
|---|---|
| Wide Column | Document |
| Better Write Performance | Flexible Schema |
| Tunable Consistency | Rich Queries |
5. Cassandra Use Cases?
- Time Series
- IoT
- Messaging
- Logging
- Analytics
- User Activity
- Recommendation Systems
Architecture
6. Explain Cassandra Architecture.
Main components
- Cluster
- Node
- Data Center
- Rack
- Gossip Protocol
7. What is a Node?
A single Cassandra server.
8. What is a Cluster?
Collection of Cassandra nodes.
9. What is a Data Center?
Logical grouping of nodes.
10. What is a Rack?
Physical grouping of nodes.
Peer-to-Peer Architecture
11. Does Cassandra have a Master node?
No.
Every node is equal.
12. Advantages of Peer-to-Peer Architecture?
- High Availability
- No Single Point of Failure
- Better Scalability
13. What is Gossip Protocol?
Nodes exchange cluster state information.
14. What is Snitch?
Determines network topology.
15. What is Seed Node?
Starting point for cluster discovery.
Data Model
16. What is a Keyspace?
Equivalent to a database.
17. What is a Table?
Stores rows inside a keyspace.
18. What is Partition Key?
Determines which node stores data.
19. What is Clustering Column?
Defines row ordering inside a partition.
20. Primary Key?
Partition Key
Clustering Columns.
Partitioning
21. Why Partitioning?
Distributes data evenly.
22. Partitioner?
Hashes partition keys.
23. Murmur3 Partitioner?
Default partitioner.
24. Token?
Hash value assigned to a partition.
25. Token Ring?
Logical ring containing all tokens.
26. Hot Partition?
One partition receiving excessive traffic.
27. Partition Size Best Practice?
Keep partitions reasonably small (typically under a few hundred MB).
28. Wide Rows?
Large number of clustering rows.
29. Skinny Rows?
Few clustering rows.
30. Partition Key Selection?
- High Cardinality
- Even Distribution
- Query Driven
Replication
31. What is Replication Factor?
Number of data copies.
32. Replication Strategies?
- SimpleStrategy
- NetworkTopologyStrategy
33. SimpleStrategy?
Single Data Center.
34. NetworkTopologyStrategy?
Multiple Data Centers.
35. Replica Node?
Stores replicated data.
36. Hinted Handoff?
Stores temporary hints for unavailable nodes.
37. Read Repair?
Synchronizes inconsistent replicas during reads.
38. Anti-Entropy Repair?
Synchronizes replicas using repair operations.
39. Repair Command?
nodetool repair
40. Incremental Repair?
Repairs only changed data.
Consistency
41. What is Tunable Consistency?
Consistency level can be selected per request.
42. Consistency Levels?
- ONE
- TWO
- THREE
- QUORUM
- LOCAL_QUORUM
- EACH_QUORUM
- ALL
43. QUORUM?
Majority of replicas respond.
44. ONE?
One replica acknowledgement.
45. ALL?
All replicas acknowledge.
46. Strong Consistency Formula?
R + W > RF
47. Eventual Consistency?
Replicas become consistent over time.
48. Read Consistency?
Controls read acknowledgements.
49. Write Consistency?
Controls write acknowledgements.
50. Best Production Consistency?
LOCAL_QUORUM is commonly used in multi-data-center deployments.
CQL
51. What is CQL?
Cassandra Query Language.
52. INSERT?
Adds data.
53. UPDATE?
Updates rows.
54. DELETE?
Deletes rows.
55. SELECT?
Retrieves data.
56. Does Cassandra support JOIN?
No.
57. Foreign Keys?
Not supported.
58. Transactions?
Lightweight Transactions (LWT) supported using Paxos.
59. Batch Statement?
Groups multiple operations.
60. Prepared Statements?
Improve performance and security.
Storage Engine
61. Memtable?
In-memory write buffer.
62. SSTable?
Immutable on-disk storage file.
63. Commit Log?
Ensures durability before Memtable flush.
64. Bloom Filter?
Avoids unnecessary disk reads.
65. Index Summary?
Speeds SSTable lookup.
Compaction
66. What is Compaction?
Merges SSTables.
67. Why Compaction?
- Remove Deleted Data
- Merge SSTables
- Improve Reads
68. Compaction Strategies?
- SizeTiered
- Leveled
- TimeWindow
69. Tombstone?
Marker indicating deleted data.
70. Tombstone Problems?
Slow queries
Large scans
Performance
71. Why is Cassandra Fast?
Sequential writes.
72. Write Path?
Commit Log
↓
Memtable
↓
SSTable
73. Read Path?
Memtable
↓
Bloom Filter
↓
SSTables
74. Common Performance Problems?
- Hot Partitions
- Tombstones
- Large Partitions
- Poor Data Model
75. Performance Tuning?
- Proper Partition Key
- Compaction
- Repair
- Caching
Monitoring
76. nodetool status?
Cluster health.
77. nodetool tpstats?
Thread pools.
78. nodetool cfstats?
Table statistics.
79. nodetool netstats?
Streaming information.
80. nodetool compactionstats?
Compaction progress.
Production Scenarios
81. Node Failure?
Replica serves requests.
82. Replica Down?
Hinted Handoff.
83. Cluster Expansion?
Add nodes.
Rebalance automatically.
84. High Latency?
Check
- Hot Partitions
- Compaction
- GC
85. Slow Reads?
Bloom Filters
Cache
Partition Size.
86. Large Tombstones?
Run compaction.
87. Uneven Data Distribution?
Review Partition Key.
88. High Disk Usage?
Review SSTables
Compaction
TTL.
89. Data Inconsistency?
Run Repair.
90. Multi Data Center Deployment?
Use
NetworkTopologyStrategy.
Senior-Level Questions
91. Cassandra vs HBase?
Peer-to-Peer
vs
Master-Worker.
92. Cassandra vs DynamoDB?
Self-managed
vs
AWS Managed.
93. Cassandra vs Redis?
Persistent Storage
vs
Memory.
94. Cassandra vs PostgreSQL?
Distributed NoSQL
vs
Relational.
95. CAP Theorem?
Cassandra prioritizes Availability and Partition Tolerance while offering tunable consistency.
96. Explain Gossip Internals.
Nodes exchange heartbeat and cluster state information periodically.
97. Explain Read Repair Internals.
Out-of-date replicas are updated during read operations.
98. Explain Compaction Internals.
Multiple SSTables are merged into fewer optimized SSTables while removing obsolete data.
99. Best Monitoring Tools?
- Prometheus
- Grafana
- DataStax OpsCenter
- nodetool
- JMX
100. Cassandra Performance Checklist?
- Proper Partition Key
- Correct Consistency Level
- Bloom Filters
- Compaction
- Repair
- Monitor GC
- Avoid Hot Partitions
- Optimize Data Model
Cassandra Write Workflow
Application
│
▼
Coordinator Node
│
▼
Commit Log
│
▼
Memtable
│
▼
SSTable
│
▼
Replicas
Cassandra Read Workflow
Application
│
▼
Coordinator Node
│
▼
Memtable
│
▼
Bloom Filter
│
▼
SSTables
│
▼
Return Result
Quick Revision
| Area | Focus |
|---|---|
| Architecture | Peer-to-Peer |
| Storage | Memtable, SSTable |
| Replication | RF, Hinted Handoff |
| Consistency | QUORUM, ONE |
| Partitioning | Token Ring |
| Performance | Bloom Filter |
| Repair | Read Repair |
| Compaction | SSTable Merge |
| Monitoring | nodetool |
| Scaling | Add Nodes |
Interview Tips
During Cassandra interviews
- Explain the peer-to-peer architecture before discussing replication.
- Understand the write path: Commit Log → Memtable → SSTable.
- Explain the read path using Bloom Filters and SSTables.
- Be comfortable discussing Replication Factor, Consistency Levels, and CAP Theorem.
- Know how Hinted Handoff, Read Repair, and Anti-Entropy Repair work.
- Understand Compaction Strategies and when to use them.
- Explain how to select an effective Partition Key to avoid hot partitions.
- Relate answers to production systems handling billions of writes.
Summary
Apache Cassandra is a highly available, distributed NoSQL database built for massive scalability and fault tolerance. Strong Cassandra interview performance requires understanding peer-to-peer architecture, partitioning, replication, consistency levels, Memtables, SSTables, Commit Logs, compaction, repair, and performance tuning.
Mastering these 100 Cassandra interview questions prepares you for Backend Developer, Java Developer, Big Data Engineer, Distributed Systems Engineer, Solution Architect, and Cassandra Administrator interviews.