Cassandra Replication Interview Questions
Master Apache Cassandra Replication with interview questions covering Replication Factor, Replication Strategies, Multi-Data Center Replication, Consistency, Hint Handoff, Read Repair, Anti-Entropy Repair, Replica Placement, and production best practices.
Introduction
Replication is one of Cassandra's biggest strengths.
Unlike traditional databases that rely on a single primary server, Cassandra automatically stores multiple copies of data across different nodes, racks, and data centers.
Replication provides
- High Availability
- Fault Tolerance
- Disaster Recovery
- Zero Single Point of Failure
- Multi-Region Support
Almost every enterprise Cassandra deployment uses replication to ensure applications continue running even if multiple servers fail.
This guide covers the most frequently asked Cassandra Replication interview questions.
Cassandra Replication Architecture
flowchart LR
Client --> Coordinator
Coordinator --> Replica1
Coordinator --> Replica2
Coordinator --> Replica3
1. What is Replication in Cassandra?
Answer
Replication is the process of storing multiple copies of the same data across different nodes.
Purpose
- High Availability
- Fault Tolerance
- Disaster Recovery
- Improved Read Performance
2. Why is Replication important?
Without replication
One Node Failure
↓
Data Loss
With replication
Node Failure
↓
Other Replicas Continue Serving Data
3. What is Replication Factor (RF)?
Replication Factor specifies the number of copies Cassandra stores.
Example
RF = 3
Means
Replica1
Replica2
Replica3
Replication Factor Example
flowchart LR
Partition --> NodeA
Partition --> NodeB
Partition --> NodeC
4. How is Replication Factor configured?
Configured while creating a keyspace.
Example
CREATE KEYSPACE banking
WITH replication = {
'class':'NetworkTopologyStrategy',
'US-East':3
};
5. What happens if RF = 1?
Only one copy exists.
Problems
- No High Availability
- Data Loss Possible
- Poor Fault Tolerance
Not recommended for production.
6. What is the recommended Replication Factor?
Production Recommendation
RF = 3
Benefits
- High Availability
- Good Performance
- Strong Consistency Support
7. What is Replica Placement?
Replica Placement determines where copies of data are stored.
Goals
- Even Distribution
- Rack Awareness
- Data Center Awareness
Replica Placement
flowchart LR
Partition --> Rack1
Partition --> Rack2
Partition --> Rack3
8. What is SimpleStrategy?
SimpleStrategy is used for
- Single Data Center
- Development
- Testing
Example
CREATE KEYSPACE demo
WITH replication = {
'class':'SimpleStrategy',
'replication_factor':3
};
9. What is NetworkTopologyStrategy?
Designed for production.
Supports
- Multiple Data Centers
- Rack Awareness
- Independent Replication
Example
CREATE KEYSPACE banking
WITH replication = {
'class':'NetworkTopologyStrategy',
'US-East':3,
'Europe':3
};
10. Difference between SimpleStrategy and NetworkTopologyStrategy?
| SimpleStrategy | NetworkTopologyStrategy |
|---|---|
| Single DC | Multi DC |
| No Rack Awareness | Rack Aware |
| Development | Production |
| Limited Availability | High Availability |
11. What is Multi-Data Center Replication?
Data is replicated across multiple geographical regions.
Example
US-East
↓
Europe
↓
Asia
Benefits
- Disaster Recovery
- Low Latency
- Business Continuity
Multi-DC Architecture
flowchart LR
US-East
<--> Europe
Europe
<--> Asia
12. What is Rack Awareness?
Cassandra stores replicas on different racks.
Benefits
- Hardware Failure Protection
- Better Availability
13. What is Coordinator Node?
Coordinator Node
- Receives Request
- Sends to Replicas
- Waits for Required Consistency
- Returns Response
Any node can become coordinator.
Write Flow
flowchart LR
Client --> Coordinator --> Replica1
Coordinator --> Replica2
Coordinator --> Replica3
14. How does Cassandra replicate writes?
Steps
- Client sends write.
- Coordinator identifies replicas.
- Writes sent to replicas.
- Required consistency acknowledgments received.
- Success returned.
15. How are reads handled?
Coordinator
↓
Queries replicas
↓
Consistency achieved
↓
Latest data returned
Read Flow
flowchart LR
Client --> Coordinator --> Replica1
Coordinator --> Replica2
Coordinator --> MergeResult --> Client
16. What is Hint Handoff?
If one replica is unavailable
↓
Coordinator stores
Hint
↓
Delivers later.
Hint Handoff Architecture
flowchart LR
ReplicaDown --> HintStored --> ReplicaRecovered --> ReplayHint
17. What is Read Repair?
During reads
↓
Coordinator detects inconsistent replicas
↓
Automatically repairs them.
Read Repair Flow
flowchart LR
ReadRequest --> ReplicaComparison --> RepairReplica --> UpdatedReplica
18. What is Anti-Entropy Repair?
Background synchronization between replicas.
Uses
Merkle Trees
Command
nodetool repair
19. Why is Repair necessary?
Replicas may become inconsistent because of
- Node Failure
- Network Partition
- Hint Expiration
- Temporary Outages
Repair synchronizes all replicas.
20. What are Merkle Trees?
Binary hash trees used to compare replica data efficiently.
Only changed data is synchronized.
21. What happens if a replica goes down?
Coordinator
↓
Stores Hint
↓
Other replicas continue serving requests
↓
Repair later synchronizes data
22. Can Cassandra survive node failures?
Yes.
If enough replicas remain available according to the configured consistency level.
23. Relationship between Replication and Consistency?
Replication
↓
Stores copies
Consistency
↓
Determines how many copies participate.
Replication + Consistency
flowchart LR
ReplicationFactor --> ReplicaCopies --> ConsistencyLevel --> SuccessfulReadWrite
24. Example
RF = 3
Write QUORUM
Read QUORUM
Formula
Read + Write > RF
2 + 2 > 3
Strong consistency.
25. What happens if RF = 3 and ONE is used?
Write succeeds after
1 Replica
Eventually remaining replicas synchronize.
26. What happens if RF = 3 and ALL is used?
All replicas must respond.
Advantages
- Highest Consistency
Disadvantages
- Higher Latency
- Lower Availability
27. What is Replica Synchronization?
Methods
- Read Repair
- Hint Handoff
- Anti-Entropy Repair
28. How does Cassandra support Disaster Recovery?
Multi-Data Center Replication
↓
Automatic Failover
↓
Replica Availability
Disaster Recovery
flowchart LR
PrimaryDC --> SecondaryDC --> ApplicationContinues
29. Real Banking Example
Requirement
Zero Downtime
Architecture
US-East
RF=3
Europe
RF=3
Consistency
LOCAL_QUORUM
Benefits
- Low Latency
- Strong Local Consistency
- Region Failure Protection
30. Real E-Commerce Example
Order Database
Replication Factor = 3
Read = ONE
Write = QUORUM
Benefits
- Fast Reads
- Reliable Writes
- High Availability
Enterprise Best Practices
- Use NetworkTopologyStrategy for production.
- Configure RF = 3 for most workloads.
- Distribute replicas across multiple racks.
- Schedule
nodetool repairregularly. - Monitor replication health continuously.
- Use
LOCAL_QUORUMfor multi-data-center deployments. - Avoid
SimpleStrategyin production. - Monitor hint backlog.
- Test node failures periodically.
- Keep all nodes time synchronized.
Quick Revision
| Topic | Key Point |
|---|---|
| Replication | Multiple Data Copies |
| Replication Factor | Number of Copies |
| RF Recommended | 3 |
| SimpleStrategy | Single DC |
| NetworkTopologyStrategy | Multi DC |
| Coordinator | Handles Client Requests |
| Hint Handoff | Stores Missed Writes |
| Read Repair | Fix During Reads |
| Anti-Entropy Repair | Background Synchronization |
| Merkle Tree | Compare Replicas |
| Multi-DC | Disaster Recovery |
Interview Tips
Interviewers frequently ask
- What is Replication Factor?
- Why is RF = 3 recommended?
- Explain SimpleStrategy vs NetworkTopologyStrategy.
- What happens when a node fails?
- Explain Hint Handoff.
- Explain Read Repair.
- Explain Anti-Entropy Repair.
- What are Merkle Trees?
- Explain Multi-Data Center Replication.
- How does replication work with consistency levels?
Always explain how replication, consistency, and coordinator nodes work together to provide high availability, fault tolerance, and scalability.
Summary
Replication is a fundamental capability of Apache Cassandra that enables high availability, fault tolerance, and disaster recovery. By storing multiple copies of data across nodes, racks, and data centers, Cassandra continues serving requests even during hardware failures or network partitions.
Understanding replication factors, replication strategies, coordinator nodes, replica placement, hint handoff, read repair, anti-entropy repair, and the interaction between replication and consistency levels is essential for designing production-ready Cassandra clusters and succeeding in senior backend, database engineering, and solution architect interviews.