Cassandra Basics Interview Questions

Master Apache Cassandra fundamentals with interview-focused questions and answers covering architecture, CAP theorem, BASE, CQL, keyspaces, tables, partition keys, clustering columns, TTL, tombstones, and production best practices.

Introduction

Apache Cassandra is one of the world's most popular distributed NoSQL databases. It was designed to handle massive amounts of data across multiple commodity servers without any single point of failure.

Today Cassandra powers applications at companies like Netflix, Apple, Uber, Instagram, Spotify, eBay, Discord, Reddit, Cisco, and many banking organizations where high availability, fault tolerance, and horizontal scalability are critical.

Unlike traditional relational databases, Cassandra follows a distributed, masterless, peer-to-peer architecture, making it an excellent choice for globally distributed applications.

This guide contains the most frequently asked Cassandra interview questions for Java Developers, Backend Engineers, Database Engineers, Solution Architects, and Principal Engineers.


Cassandra Overview

flowchart LR

Client --> CoordinatorNode

CoordinatorNode --> Node1

CoordinatorNode --> Node2

CoordinatorNode --> Node3

Node1 --> Storage

Node2 --> Storage

Node3 --> Storage

1. What is Apache Cassandra?

Answer

Apache Cassandra is an open-source distributed NoSQL wide-column database designed for

  • High Availability
  • Horizontal Scalability
  • Fault Tolerance
  • High Write Throughput
  • Multi-Data Center Replication

It stores data across multiple nodes without requiring a master server.


2. Why was Cassandra created?

Answer

Traditional relational databases faced problems such as

  • Single Point of Failure
  • Limited Horizontal Scaling
  • Downtime During Maintenance
  • Poor Performance for Massive Writes

Cassandra was designed to solve these challenges by providing

  • Masterless Architecture
  • Automatic Replication
  • Linear Scalability
  • Always-On Availability

3. What type of database is Cassandra?

Cassandra is

  • NoSQL Database
  • Distributed Database
  • Wide Column Store
  • Column Family Database

4. What are the key features of Cassandra?

  • Masterless Architecture
  • Peer-to-Peer Communication
  • Linear Scalability
  • Automatic Replication
  • Fault Tolerance
  • High Write Performance
  • Tunable Consistency
  • Multi Data Center Support

5. Which companies use Cassandra?

Examples include

  • Netflix
  • Apple
  • Uber
  • Spotify
  • Instagram
  • Discord
  • Reddit
  • Cisco

6. What are common use cases?

  • Time Series Data
  • IoT
  • Banking Transactions
  • User Activity Tracking
  • Messaging Systems
  • Recommendation Systems
  • Logging Systems
  • Event Storage

7. Cassandra vs Traditional RDBMS?

Cassandra RDBMS
NoSQL Relational
Schema Flexible Fixed Schema
Horizontal Scaling Vertical Scaling
Eventually Consistent Strongly Consistent
High Availability High Consistency
Denormalized Data Normalized Data

8. Cassandra vs MongoDB?

Cassandra MongoDB
Wide Column Document
Excellent Writes Flexible Documents
Time Series JSON Storage
Multi-DC Native Replica Set

9. Cassandra vs HBase?

Cassandra HBase
Peer-to-Peer Master-Slave
No HDFS Required Uses HDFS
Better Availability Better Hadoop Integration

10. Explain CAP Theorem.

CAP states that a distributed database can guarantee only two of

  • Consistency
  • Availability
  • Partition Tolerance

Cassandra chooses

  • Availability
  • Partition Tolerance

11. What is BASE?

BASE stands for

  • Basically Available
  • Soft State
  • Eventual Consistency

Unlike relational databases, Cassandra follows BASE instead of ACID.


12. What is Eventual Consistency?

Data may not be immediately consistent across all nodes.

Eventually every replica reaches the same state.


13. What is a Keyspace?

A Keyspace is similar to a database in relational systems.

It contains

  • Tables
  • Replication Configuration

14. What is a Table?

A table stores rows and columns similar to relational databases.

However,

Rows are distributed across nodes.


15. What is a Row?

A row represents one record inside a table.

Example

Employee

ID

Name

Salary

16. What is a Column?

A column stores a single value.

Example

Name

=

Venugopal

17. What is a Primary Key?

Primary Key uniquely identifies a row.

Structure

Primary Key

=

Partition Key

+

Clustering Columns

18. What is a Partition Key?

Partition Key decides

Which node stores the data.

Example

CustomerID

19. What are Clustering Columns?

Clustering Columns determine

The sorting order inside a partition.


20. What is a Composite Primary Key?

Example

PRIMARY KEY ((customer_id), order_date)

Partition Key

customer_id

Clustering Column

order_date

21. Explain Cassandra Data Types.

Examples

  • text
  • int
  • bigint
  • uuid
  • timeuuid
  • boolean
  • timestamp
  • list
  • set
  • map

22. What is CQL?

CQL

Cassandra Query Language

Looks similar to SQL but is designed for distributed databases.


23. Example CQL Create Table

CREATE TABLE employee (

id UUID PRIMARY KEY,

name TEXT,

salary DECIMAL

);

24. Insert Data

INSERT INTO employee

(id,name,salary)

VALUES

(uuid(), 'Venugopal', 5000);

25. Read Data

SELECT *

FROM employee

WHERE id=?;

26. Update Data

UPDATE employee

SET salary=6000

WHERE id=?;

27. Delete Data

DELETE

FROM employee

WHERE id=?;

28. What is TTL?

TTL

Time To Live

Automatically deletes data after a specified time.

Example

USING TTL 3600

Data expires after one hour.


29. What are Tombstones?

Deleting data creates

Tombstones

instead of immediate removal.

Actual deletion occurs during

Compaction.


30. What is TimeUUID?

TimeUUID combines

  • UUID
  • Timestamp

Useful for

  • Event Ordering
  • Time-Series Data

31. What are Collections?

Supported Collections

  • List
  • Set
  • Map

Used for storing multiple values in one column.


32. What are Counters?

Counters support atomic increment operations.

Example

Page Views

Likes

Downloads

33. What are Static Columns?

Shared across every row inside the same partition.

Useful for

Customer Information

shared across Orders.


34. What are Secondary Indexes?

Indexes created on non-primary key columns.

Useful only for low-cardinality queries.

Not recommended for heavy workloads.


35. What are Materialized Views?

Automatically maintain another query pattern.

Use carefully because they increase write overhead.


36. What are Lightweight Transactions (LWT)?

Support conditional updates.

Example

IF NOT EXISTS

Uses Paxos protocol.


37. What are Batch Operations?

Execute multiple statements together.

Best for

Atomic operations on the same partition.

Avoid very large batches.


38. Advantages of Cassandra

  • Highly Available
  • Fault Tolerant
  • Linear Scalability
  • Multi-Region Support
  • Excellent Write Performance
  • No Single Point of Failure

39. Limitations of Cassandra

  • No JOINs
  • No Foreign Keys
  • Limited Aggregations
  • Query-Driven Modeling
  • Eventual Consistency
  • Denormalization Required

40. Explain a Real Production Use Case.

A banking application stores transaction history.

Each transaction is written to Cassandra.

Partition Key

Account_ID

Clustering Column

Transaction_Time

Benefits

  • Millions of Writes Per Second
  • Multi-Region Replication
  • No Downtime
  • Fast Transaction History Retrieval

Enterprise Best Practices

  • Design tables based on queries.
  • Choose partition keys carefully.
  • Keep partitions balanced.
  • Avoid large partitions.
  • Use UUID or TimeUUID appropriately.
  • Minimize tombstones.
  • Use TTL only when necessary.
  • Avoid ALLOW FILTERING in production.
  • Monitor compaction and repair jobs.
  • Use NetworkTopologyStrategy for multi-data-center deployments.

Quick Revision

Topic Key Point
Database Type NoSQL Wide Column
Architecture Peer-to-Peer
Master None
Scaling Horizontal
CAP AP
Model BASE
Language CQL
Keyspace Database
Table Collection of Rows
Partition Key Data Distribution
Clustering Column Sorting
TTL Auto Expiration
Tombstone Deleted Marker
TimeUUID Ordered UUID
LWT Conditional Update

Interview Tips

Interviewers commonly ask follow-up questions such as

  • Why Cassandra instead of PostgreSQL?
  • How does Cassandra scale?
  • What happens when a node fails?
  • Why are joins not supported?
  • Explain Partition Key with examples.
  • Explain Eventual Consistency.
  • What are Tombstones?
  • Why does Cassandra perform well for writes?
  • Explain CAP theorem in Cassandra.
  • Explain BASE vs ACID.

Be prepared to answer with real production examples rather than only theoretical definitions.


Summary

Apache Cassandra is a highly available, distributed, masterless NoSQL database built for applications requiring continuous availability, fault tolerance, and massive horizontal scalability. Its peer-to-peer architecture, tunable consistency, and efficient write path make it ideal for time-series, event-driven, IoT, financial, and globally distributed systems.

Understanding Cassandra fundamentals—including CAP theorem, BASE, keyspaces, tables, partition keys, clustering columns, CQL, TTL, tombstones, and common production practices—provides the foundation for advanced topics such as architecture, data modeling, replication, consistency tuning, and performance optimization covered in the upcoming chapters.