Cassandra Basics Interview Questions
Master Apache Cassandra fundamentals with interview-focused questions and answers covering architecture, CAP theorem, BASE, CQL, keyspaces, tables, partition keys, clustering columns, TTL, tombstones, and production best practices.
Introduction
Apache Cassandra is one of the world's most popular distributed NoSQL databases. It was designed to handle massive amounts of data across multiple commodity servers without any single point of failure.
Today Cassandra powers applications at companies like Netflix, Apple, Uber, Instagram, Spotify, eBay, Discord, Reddit, Cisco, and many banking organizations where high availability, fault tolerance, and horizontal scalability are critical.
Unlike traditional relational databases, Cassandra follows a distributed, masterless, peer-to-peer architecture, making it an excellent choice for globally distributed applications.
This guide contains the most frequently asked Cassandra interview questions for Java Developers, Backend Engineers, Database Engineers, Solution Architects, and Principal Engineers.
Cassandra Overview
flowchart LR
Client --> CoordinatorNode
CoordinatorNode --> Node1
CoordinatorNode --> Node2
CoordinatorNode --> Node3
Node1 --> Storage
Node2 --> Storage
Node3 --> Storage
1. What is Apache Cassandra?
Answer
Apache Cassandra is an open-source distributed NoSQL wide-column database designed for
- High Availability
- Horizontal Scalability
- Fault Tolerance
- High Write Throughput
- Multi-Data Center Replication
It stores data across multiple nodes without requiring a master server.
2. Why was Cassandra created?
Answer
Traditional relational databases faced problems such as
- Single Point of Failure
- Limited Horizontal Scaling
- Downtime During Maintenance
- Poor Performance for Massive Writes
Cassandra was designed to solve these challenges by providing
- Masterless Architecture
- Automatic Replication
- Linear Scalability
- Always-On Availability
3. What type of database is Cassandra?
Cassandra is
- NoSQL Database
- Distributed Database
- Wide Column Store
- Column Family Database
4. What are the key features of Cassandra?
- Masterless Architecture
- Peer-to-Peer Communication
- Linear Scalability
- Automatic Replication
- Fault Tolerance
- High Write Performance
- Tunable Consistency
- Multi Data Center Support
5. Which companies use Cassandra?
Examples include
- Netflix
- Apple
- Uber
- Spotify
- Discord
- Cisco
6. What are common use cases?
- Time Series Data
- IoT
- Banking Transactions
- User Activity Tracking
- Messaging Systems
- Recommendation Systems
- Logging Systems
- Event Storage
7. Cassandra vs Traditional RDBMS?
| Cassandra | RDBMS |
|---|---|
| NoSQL | Relational |
| Schema Flexible | Fixed Schema |
| Horizontal Scaling | Vertical Scaling |
| Eventually Consistent | Strongly Consistent |
| High Availability | High Consistency |
| Denormalized Data | Normalized Data |
8. Cassandra vs MongoDB?
| Cassandra | MongoDB |
|---|---|
| Wide Column | Document |
| Excellent Writes | Flexible Documents |
| Time Series | JSON Storage |
| Multi-DC Native | Replica Set |
9. Cassandra vs HBase?
| Cassandra | HBase |
|---|---|
| Peer-to-Peer | Master-Slave |
| No HDFS Required | Uses HDFS |
| Better Availability | Better Hadoop Integration |
10. Explain CAP Theorem.
CAP states that a distributed database can guarantee only two of
- Consistency
- Availability
- Partition Tolerance
Cassandra chooses
- Availability
- Partition Tolerance
11. What is BASE?
BASE stands for
- Basically Available
- Soft State
- Eventual Consistency
Unlike relational databases, Cassandra follows BASE instead of ACID.
12. What is Eventual Consistency?
Data may not be immediately consistent across all nodes.
Eventually every replica reaches the same state.
13. What is a Keyspace?
A Keyspace is similar to a database in relational systems.
It contains
- Tables
- Replication Configuration
14. What is a Table?
A table stores rows and columns similar to relational databases.
However,
Rows are distributed across nodes.
15. What is a Row?
A row represents one record inside a table.
Example
Employee
ID
Name
Salary
16. What is a Column?
A column stores a single value.
Example
Name
=
Venugopal
17. What is a Primary Key?
Primary Key uniquely identifies a row.
Structure
Primary Key
=
Partition Key
+
Clustering Columns
18. What is a Partition Key?
Partition Key decides
Which node stores the data.
Example
CustomerID
19. What are Clustering Columns?
Clustering Columns determine
The sorting order inside a partition.
20. What is a Composite Primary Key?
Example
PRIMARY KEY ((customer_id), order_date)
Partition Key
customer_id
Clustering Column
order_date
21. Explain Cassandra Data Types.
Examples
- text
- int
- bigint
- uuid
- timeuuid
- boolean
- timestamp
- list
- set
- map
22. What is CQL?
CQL
Cassandra Query Language
Looks similar to SQL but is designed for distributed databases.
23. Example CQL Create Table
CREATE TABLE employee (
id UUID PRIMARY KEY,
name TEXT,
salary DECIMAL
);
24. Insert Data
INSERT INTO employee
(id,name,salary)
VALUES
(uuid(), 'Venugopal', 5000);
25. Read Data
SELECT *
FROM employee
WHERE id=?;
26. Update Data
UPDATE employee
SET salary=6000
WHERE id=?;
27. Delete Data
DELETE
FROM employee
WHERE id=?;
28. What is TTL?
TTL
Time To Live
Automatically deletes data after a specified time.
Example
USING TTL 3600
Data expires after one hour.
29. What are Tombstones?
Deleting data creates
Tombstones
instead of immediate removal.
Actual deletion occurs during
Compaction.
30. What is TimeUUID?
TimeUUID combines
- UUID
- Timestamp
Useful for
- Event Ordering
- Time-Series Data
31. What are Collections?
Supported Collections
- List
- Set
- Map
Used for storing multiple values in one column.
32. What are Counters?
Counters support atomic increment operations.
Example
Page Views
Likes
Downloads
33. What are Static Columns?
Shared across every row inside the same partition.
Useful for
Customer Information
shared across Orders.
34. What are Secondary Indexes?
Indexes created on non-primary key columns.
Useful only for low-cardinality queries.
Not recommended for heavy workloads.
35. What are Materialized Views?
Automatically maintain another query pattern.
Use carefully because they increase write overhead.
36. What are Lightweight Transactions (LWT)?
Support conditional updates.
Example
IF NOT EXISTS
Uses Paxos protocol.
37. What are Batch Operations?
Execute multiple statements together.
Best for
Atomic operations on the same partition.
Avoid very large batches.
38. Advantages of Cassandra
- Highly Available
- Fault Tolerant
- Linear Scalability
- Multi-Region Support
- Excellent Write Performance
- No Single Point of Failure
39. Limitations of Cassandra
- No JOINs
- No Foreign Keys
- Limited Aggregations
- Query-Driven Modeling
- Eventual Consistency
- Denormalization Required
40. Explain a Real Production Use Case.
A banking application stores transaction history.
Each transaction is written to Cassandra.
Partition Key
Account_ID
Clustering Column
Transaction_Time
Benefits
- Millions of Writes Per Second
- Multi-Region Replication
- No Downtime
- Fast Transaction History Retrieval
Enterprise Best Practices
- Design tables based on queries.
- Choose partition keys carefully.
- Keep partitions balanced.
- Avoid large partitions.
- Use UUID or TimeUUID appropriately.
- Minimize tombstones.
- Use TTL only when necessary.
- Avoid ALLOW FILTERING in production.
- Monitor compaction and repair jobs.
- Use NetworkTopologyStrategy for multi-data-center deployments.
Quick Revision
| Topic | Key Point |
|---|---|
| Database Type | NoSQL Wide Column |
| Architecture | Peer-to-Peer |
| Master | None |
| Scaling | Horizontal |
| CAP | AP |
| Model | BASE |
| Language | CQL |
| Keyspace | Database |
| Table | Collection of Rows |
| Partition Key | Data Distribution |
| Clustering Column | Sorting |
| TTL | Auto Expiration |
| Tombstone | Deleted Marker |
| TimeUUID | Ordered UUID |
| LWT | Conditional Update |
Interview Tips
Interviewers commonly ask follow-up questions such as
- Why Cassandra instead of PostgreSQL?
- How does Cassandra scale?
- What happens when a node fails?
- Why are joins not supported?
- Explain Partition Key with examples.
- Explain Eventual Consistency.
- What are Tombstones?
- Why does Cassandra perform well for writes?
- Explain CAP theorem in Cassandra.
- Explain BASE vs ACID.
Be prepared to answer with real production examples rather than only theoretical definitions.
Summary
Apache Cassandra is a highly available, distributed, masterless NoSQL database built for applications requiring continuous availability, fault tolerance, and massive horizontal scalability. Its peer-to-peer architecture, tunable consistency, and efficient write path make it ideal for time-series, event-driven, IoT, financial, and globally distributed systems.
Understanding Cassandra fundamentals—including CAP theorem, BASE, keyspaces, tables, partition keys, clustering columns, CQL, TTL, tombstones, and common production practices—provides the foundation for advanced topics such as architecture, data modeling, replication, consistency tuning, and performance optimization covered in the upcoming chapters.