Cassandra Replication Interview Questions

Master Apache Cassandra Replication with interview questions covering Replication Factor, Replication Strategies, Multi-Data Center Replication, Consistency, Hint Handoff, Read Repair, Anti-Entropy Repair, Replica Placement, and production best practices.

Introduction

Replication is one of Cassandra's biggest strengths.

Unlike traditional databases that rely on a single primary server, Cassandra automatically stores multiple copies of data across different nodes, racks, and data centers.

Replication provides

  • High Availability
  • Fault Tolerance
  • Disaster Recovery
  • Zero Single Point of Failure
  • Multi-Region Support

Almost every enterprise Cassandra deployment uses replication to ensure applications continue running even if multiple servers fail.

This guide covers the most frequently asked Cassandra Replication interview questions.


Cassandra Replication Architecture

flowchart LR

Client --> Coordinator

Coordinator --> Replica1

Coordinator --> Replica2

Coordinator --> Replica3

1. What is Replication in Cassandra?

Answer

Replication is the process of storing multiple copies of the same data across different nodes.

Purpose

  • High Availability
  • Fault Tolerance
  • Disaster Recovery
  • Improved Read Performance

2. Why is Replication important?

Without replication

One Node Failure

↓

Data Loss

With replication

Node Failure

↓

Other Replicas Continue Serving Data

3. What is Replication Factor (RF)?

Replication Factor specifies the number of copies Cassandra stores.

Example

RF = 3

Means

Replica1

Replica2

Replica3

Replication Factor Example

flowchart LR

Partition --> NodeA

Partition --> NodeB

Partition --> NodeC

4. How is Replication Factor configured?

Configured while creating a keyspace.

Example

CREATE KEYSPACE banking

WITH replication = {

'class':'NetworkTopologyStrategy',

'US-East':3

};

5. What happens if RF = 1?

Only one copy exists.

Problems

  • No High Availability
  • Data Loss Possible
  • Poor Fault Tolerance

Not recommended for production.


6. What is the recommended Replication Factor?

Production Recommendation

RF = 3

Benefits

  • High Availability
  • Good Performance
  • Strong Consistency Support

7. What is Replica Placement?

Replica Placement determines where copies of data are stored.

Goals

  • Even Distribution
  • Rack Awareness
  • Data Center Awareness

Replica Placement

flowchart LR

Partition --> Rack1

Partition --> Rack2

Partition --> Rack3

8. What is SimpleStrategy?

SimpleStrategy is used for

  • Single Data Center
  • Development
  • Testing

Example

CREATE KEYSPACE demo

WITH replication = {

'class':'SimpleStrategy',

'replication_factor':3

};

9. What is NetworkTopologyStrategy?

Designed for production.

Supports

  • Multiple Data Centers
  • Rack Awareness
  • Independent Replication

Example

CREATE KEYSPACE banking

WITH replication = {

'class':'NetworkTopologyStrategy',

'US-East':3,

'Europe':3

};

10. Difference between SimpleStrategy and NetworkTopologyStrategy?

SimpleStrategy NetworkTopologyStrategy
Single DC Multi DC
No Rack Awareness Rack Aware
Development Production
Limited Availability High Availability

11. What is Multi-Data Center Replication?

Data is replicated across multiple geographical regions.

Example

US-East

↓

Europe

↓

Asia

Benefits

  • Disaster Recovery
  • Low Latency
  • Business Continuity

Multi-DC Architecture

flowchart LR

US-East

<--> Europe

Europe

<--> Asia

12. What is Rack Awareness?

Cassandra stores replicas on different racks.

Benefits

  • Hardware Failure Protection
  • Better Availability

13. What is Coordinator Node?

Coordinator Node

  • Receives Request
  • Sends to Replicas
  • Waits for Required Consistency
  • Returns Response

Any node can become coordinator.


Write Flow

flowchart LR

Client --> Coordinator --> Replica1

Coordinator --> Replica2

Coordinator --> Replica3

14. How does Cassandra replicate writes?

Steps

  1. Client sends write.
  2. Coordinator identifies replicas.
  3. Writes sent to replicas.
  4. Required consistency acknowledgments received.
  5. Success returned.

15. How are reads handled?

Coordinator

Queries replicas

Consistency achieved

Latest data returned


Read Flow

flowchart LR

Client --> Coordinator --> Replica1

Coordinator --> Replica2

Coordinator --> MergeResult --> Client

16. What is Hint Handoff?

If one replica is unavailable

Coordinator stores

Hint

Delivers later.


Hint Handoff Architecture

flowchart LR

ReplicaDown --> HintStored --> ReplicaRecovered --> ReplayHint

17. What is Read Repair?

During reads

Coordinator detects inconsistent replicas

Automatically repairs them.


Read Repair Flow

flowchart LR

ReadRequest --> ReplicaComparison --> RepairReplica --> UpdatedReplica

18. What is Anti-Entropy Repair?

Background synchronization between replicas.

Uses

Merkle Trees

Command

nodetool repair

19. Why is Repair necessary?

Replicas may become inconsistent because of

  • Node Failure
  • Network Partition
  • Hint Expiration
  • Temporary Outages

Repair synchronizes all replicas.


20. What are Merkle Trees?

Binary hash trees used to compare replica data efficiently.

Only changed data is synchronized.


21. What happens if a replica goes down?

Coordinator

Stores Hint

Other replicas continue serving requests

Repair later synchronizes data


22. Can Cassandra survive node failures?

Yes.

If enough replicas remain available according to the configured consistency level.


23. Relationship between Replication and Consistency?

Replication

Stores copies

Consistency

Determines how many copies participate.


Replication + Consistency

flowchart LR

ReplicationFactor --> ReplicaCopies --> ConsistencyLevel --> SuccessfulReadWrite

24. Example

RF = 3

Write QUORUM

Read QUORUM

Formula

Read + Write > RF
2 + 2 > 3

Strong consistency.


25. What happens if RF = 3 and ONE is used?

Write succeeds after

1 Replica

Eventually remaining replicas synchronize.


26. What happens if RF = 3 and ALL is used?

All replicas must respond.

Advantages

  • Highest Consistency

Disadvantages

  • Higher Latency
  • Lower Availability

27. What is Replica Synchronization?

Methods

  • Read Repair
  • Hint Handoff
  • Anti-Entropy Repair

28. How does Cassandra support Disaster Recovery?

Multi-Data Center Replication

Automatic Failover

Replica Availability


Disaster Recovery

flowchart LR

PrimaryDC --> SecondaryDC --> ApplicationContinues

29. Real Banking Example

Requirement

Zero Downtime

Architecture

US-East

RF=3

Europe

RF=3

Consistency

LOCAL_QUORUM

Benefits

  • Low Latency
  • Strong Local Consistency
  • Region Failure Protection

30. Real E-Commerce Example

Order Database

Replication Factor = 3

Read = ONE

Write = QUORUM

Benefits

  • Fast Reads
  • Reliable Writes
  • High Availability

Enterprise Best Practices

  • Use NetworkTopologyStrategy for production.
  • Configure RF = 3 for most workloads.
  • Distribute replicas across multiple racks.
  • Schedule nodetool repair regularly.
  • Monitor replication health continuously.
  • Use LOCAL_QUORUM for multi-data-center deployments.
  • Avoid SimpleStrategy in production.
  • Monitor hint backlog.
  • Test node failures periodically.
  • Keep all nodes time synchronized.

Quick Revision

Topic Key Point
Replication Multiple Data Copies
Replication Factor Number of Copies
RF Recommended 3
SimpleStrategy Single DC
NetworkTopologyStrategy Multi DC
Coordinator Handles Client Requests
Hint Handoff Stores Missed Writes
Read Repair Fix During Reads
Anti-Entropy Repair Background Synchronization
Merkle Tree Compare Replicas
Multi-DC Disaster Recovery

Interview Tips

Interviewers frequently ask

  • What is Replication Factor?
  • Why is RF = 3 recommended?
  • Explain SimpleStrategy vs NetworkTopologyStrategy.
  • What happens when a node fails?
  • Explain Hint Handoff.
  • Explain Read Repair.
  • Explain Anti-Entropy Repair.
  • What are Merkle Trees?
  • Explain Multi-Data Center Replication.
  • How does replication work with consistency levels?

Always explain how replication, consistency, and coordinator nodes work together to provide high availability, fault tolerance, and scalability.


Summary

Replication is a fundamental capability of Apache Cassandra that enables high availability, fault tolerance, and disaster recovery. By storing multiple copies of data across nodes, racks, and data centers, Cassandra continues serving requests even during hardware failures or network partitions.

Understanding replication factors, replication strategies, coordinator nodes, replica placement, hint handoff, read repair, anti-entropy repair, and the interaction between replication and consistency levels is essential for designing production-ready Cassandra clusters and succeeding in senior backend, database engineering, and solution architect interviews.