Cassandra Interview Questions (Top 100 Questions with Answers)

Master Apache Cassandra Interview Questions with production-oriented questions covering Cassandra Architecture, Data Model, Replication, Partitioning, Consistency Levels, CQL, Performance Tuning, Compaction, Repair, and real-world production scenarios.

Introduction

Apache Cassandra is a highly scalable distributed NoSQL database designed for handling massive amounts of data across multiple servers with no single point of failure.

It is widely used in

  • Banking
  • IoT
  • Messaging Systems
  • Telecommunication
  • Recommendation Engines
  • Fraud Detection
  • Time-Series Applications

Major companies using Cassandra include

  • Netflix
  • Apple
  • Uber
  • Spotify
  • Cisco
  • Instagram
  • eBay
  • IBM

This guide contains the Top 100 Cassandra Interview Questions frequently asked in Java, Backend, Cloud, DevOps, and Solution Architect interviews.


Cassandra Interview Roadmap

Cassandra Basics
        │
        ▼
Architecture
        │
        ▼
Data Model
        │
        ▼
Partitioning
        │
        ▼
Replication
        │
        ▼
Consistency
        │
        ▼
Compaction
        │
        ▼
Repair
        │
        ▼
Performance

Cassandra Fundamentals

1. What is Apache Cassandra?

Apache Cassandra is a distributed NoSQL wide-column database designed for high availability and linear scalability.


2. Why Cassandra?

  • No Single Point of Failure
  • Linear Scalability
  • High Availability
  • Multi-Data Center Support
  • Fault Tolerance

3. Cassandra vs Relational Database?

Cassandra SQL Database
NoSQL Relational
Wide Column Tables
Horizontal Scaling Mostly Vertical Scaling
Eventual Consistency Strong Consistency

4. Cassandra vs MongoDB?

Cassandra MongoDB
Wide Column Document
Better Write Performance Flexible Schema
Tunable Consistency Rich Queries

5. Cassandra Use Cases?

  • Time Series
  • IoT
  • Messaging
  • Logging
  • Analytics
  • User Activity
  • Recommendation Systems

Architecture

6. Explain Cassandra Architecture.

Main components

  • Cluster
  • Node
  • Data Center
  • Rack
  • Gossip Protocol

7. What is a Node?

A single Cassandra server.


8. What is a Cluster?

Collection of Cassandra nodes.


9. What is a Data Center?

Logical grouping of nodes.


10. What is a Rack?

Physical grouping of nodes.


Peer-to-Peer Architecture

11. Does Cassandra have a Master node?

No.

Every node is equal.


12. Advantages of Peer-to-Peer Architecture?

  • High Availability
  • No Single Point of Failure
  • Better Scalability

13. What is Gossip Protocol?

Nodes exchange cluster state information.


14. What is Snitch?

Determines network topology.


15. What is Seed Node?

Starting point for cluster discovery.


Data Model

16. What is a Keyspace?

Equivalent to a database.


17. What is a Table?

Stores rows inside a keyspace.


18. What is Partition Key?

Determines which node stores data.


19. What is Clustering Column?

Defines row ordering inside a partition.


20. Primary Key?

Partition Key

Clustering Columns.


Partitioning

21. Why Partitioning?

Distributes data evenly.


22. Partitioner?

Hashes partition keys.


23. Murmur3 Partitioner?

Default partitioner.


24. Token?

Hash value assigned to a partition.


25. Token Ring?

Logical ring containing all tokens.


26. Hot Partition?

One partition receiving excessive traffic.


27. Partition Size Best Practice?

Keep partitions reasonably small (typically under a few hundred MB).


28. Wide Rows?

Large number of clustering rows.


29. Skinny Rows?

Few clustering rows.


30. Partition Key Selection?

  • High Cardinality
  • Even Distribution
  • Query Driven

Replication

31. What is Replication Factor?

Number of data copies.


32. Replication Strategies?

  • SimpleStrategy
  • NetworkTopologyStrategy

33. SimpleStrategy?

Single Data Center.


34. NetworkTopologyStrategy?

Multiple Data Centers.


35. Replica Node?

Stores replicated data.


36. Hinted Handoff?

Stores temporary hints for unavailable nodes.


37. Read Repair?

Synchronizes inconsistent replicas during reads.


38. Anti-Entropy Repair?

Synchronizes replicas using repair operations.


39. Repair Command?

nodetool repair

40. Incremental Repair?

Repairs only changed data.


Consistency

41. What is Tunable Consistency?

Consistency level can be selected per request.


42. Consistency Levels?

  • ONE
  • TWO
  • THREE
  • QUORUM
  • LOCAL_QUORUM
  • EACH_QUORUM
  • ALL

43. QUORUM?

Majority of replicas respond.


44. ONE?

One replica acknowledgement.


45. ALL?

All replicas acknowledge.


46. Strong Consistency Formula?

R + W > RF

47. Eventual Consistency?

Replicas become consistent over time.


48. Read Consistency?

Controls read acknowledgements.


49. Write Consistency?

Controls write acknowledgements.


50. Best Production Consistency?

LOCAL_QUORUM is commonly used in multi-data-center deployments.


CQL

51. What is CQL?

Cassandra Query Language.


52. INSERT?

Adds data.


53. UPDATE?

Updates rows.


54. DELETE?

Deletes rows.


55. SELECT?

Retrieves data.


56. Does Cassandra support JOIN?

No.


57. Foreign Keys?

Not supported.


58. Transactions?

Lightweight Transactions (LWT) supported using Paxos.


59. Batch Statement?

Groups multiple operations.


60. Prepared Statements?

Improve performance and security.


Storage Engine

61. Memtable?

In-memory write buffer.


62. SSTable?

Immutable on-disk storage file.


63. Commit Log?

Ensures durability before Memtable flush.


64. Bloom Filter?

Avoids unnecessary disk reads.


65. Index Summary?

Speeds SSTable lookup.


Compaction

66. What is Compaction?

Merges SSTables.


67. Why Compaction?

  • Remove Deleted Data
  • Merge SSTables
  • Improve Reads

68. Compaction Strategies?

  • SizeTiered
  • Leveled
  • TimeWindow

69. Tombstone?

Marker indicating deleted data.


70. Tombstone Problems?

Slow queries

Large scans


Performance

71. Why is Cassandra Fast?

Sequential writes.


72. Write Path?

Commit Log

Memtable

SSTable


73. Read Path?

Memtable

Bloom Filter

SSTables


74. Common Performance Problems?

  • Hot Partitions
  • Tombstones
  • Large Partitions
  • Poor Data Model

75. Performance Tuning?

  • Proper Partition Key
  • Compaction
  • Repair
  • Caching

Monitoring

76. nodetool status?

Cluster health.


77. nodetool tpstats?

Thread pools.


78. nodetool cfstats?

Table statistics.


79. nodetool netstats?

Streaming information.


80. nodetool compactionstats?

Compaction progress.


Production Scenarios

81. Node Failure?

Replica serves requests.


82. Replica Down?

Hinted Handoff.


83. Cluster Expansion?

Add nodes.

Rebalance automatically.


84. High Latency?

Check

  • Hot Partitions
  • Compaction
  • GC

85. Slow Reads?

Bloom Filters

Cache

Partition Size.


86. Large Tombstones?

Run compaction.


87. Uneven Data Distribution?

Review Partition Key.


88. High Disk Usage?

Review SSTables

Compaction

TTL.


89. Data Inconsistency?

Run Repair.


90. Multi Data Center Deployment?

Use

NetworkTopologyStrategy.


Senior-Level Questions

91. Cassandra vs HBase?

Peer-to-Peer

vs

Master-Worker.


92. Cassandra vs DynamoDB?

Self-managed

vs

AWS Managed.


93. Cassandra vs Redis?

Persistent Storage

vs

Memory.


94. Cassandra vs PostgreSQL?

Distributed NoSQL

vs

Relational.


95. CAP Theorem?

Cassandra prioritizes Availability and Partition Tolerance while offering tunable consistency.


96. Explain Gossip Internals.

Nodes exchange heartbeat and cluster state information periodically.


97. Explain Read Repair Internals.

Out-of-date replicas are updated during read operations.


98. Explain Compaction Internals.

Multiple SSTables are merged into fewer optimized SSTables while removing obsolete data.


99. Best Monitoring Tools?

  • Prometheus
  • Grafana
  • DataStax OpsCenter
  • nodetool
  • JMX

100. Cassandra Performance Checklist?

  • Proper Partition Key
  • Correct Consistency Level
  • Bloom Filters
  • Compaction
  • Repair
  • Monitor GC
  • Avoid Hot Partitions
  • Optimize Data Model

Cassandra Write Workflow

Application
      │
      ▼
Coordinator Node
      │
      ▼
Commit Log
      │
      ▼
Memtable
      │
      ▼
SSTable
      │
      ▼
Replicas

Cassandra Read Workflow

Application
      │
      ▼
Coordinator Node
      │
      ▼
Memtable
      │
      ▼
Bloom Filter
      │
      ▼
SSTables
      │
      ▼
Return Result

Quick Revision

Area Focus
Architecture Peer-to-Peer
Storage Memtable, SSTable
Replication RF, Hinted Handoff
Consistency QUORUM, ONE
Partitioning Token Ring
Performance Bloom Filter
Repair Read Repair
Compaction SSTable Merge
Monitoring nodetool
Scaling Add Nodes

Interview Tips

During Cassandra interviews

  • Explain the peer-to-peer architecture before discussing replication.
  • Understand the write path: Commit Log → Memtable → SSTable.
  • Explain the read path using Bloom Filters and SSTables.
  • Be comfortable discussing Replication Factor, Consistency Levels, and CAP Theorem.
  • Know how Hinted Handoff, Read Repair, and Anti-Entropy Repair work.
  • Understand Compaction Strategies and when to use them.
  • Explain how to select an effective Partition Key to avoid hot partitions.
  • Relate answers to production systems handling billions of writes.

Summary

Apache Cassandra is a highly available, distributed NoSQL database built for massive scalability and fault tolerance. Strong Cassandra interview performance requires understanding peer-to-peer architecture, partitioning, replication, consistency levels, Memtables, SSTables, Commit Logs, compaction, repair, and performance tuning.

Mastering these 100 Cassandra interview questions prepares you for Backend Developer, Java Developer, Big Data Engineer, Distributed Systems Engineer, Solution Architect, and Cassandra Administrator interviews.