Kafka Replication Interview Questions and Answers
Master Kafka Replication with real-world interview questions covering leaders, followers, ISR, replication factor, leader election, failover, min.insync.replicas, acknowledgements, and production best practices.
Kafka Replication Interview Questions and Answers
Replication is one of Kafka's most important features because it provides:
- High Availability
- Fault Tolerance
- Data Durability
- Disaster Recovery
- Zero Data Loss (Proper Configuration)
Without replication, a single broker failure could permanently lose messages.
This article covers everything interviewers expect regarding Kafka Replication.
Kafka Replication Architecture
flowchart LR
Producer --> LeaderPartition["Leader Partition"]
LeaderPartition["Leader Partition"] --> Follower1["Follower 1"]
LeaderPartition["Leader Partition"] --> Follower2["Follower 2"]
Follower1["Follower 1"] --> Consumer
LeaderPartition["Leader Partition"] --> Consumer
Q1. What is Kafka Replication?
Answer
Kafka Replication is the process of maintaining multiple copies of the same partition across different brokers.
Instead of storing data on one broker only, Kafka stores replicas on multiple brokers.
Example
Partition-0
Broker 1 (Leader)
Broker 2 (Follower)
Broker 3 (Follower)
If Broker 1 fails, another replica becomes the leader.
Replication
flowchart LR
Leader --> Replica1["Replica 1"]
Leader --> Replica2["Replica 2"]
Q2. Why is Replication required?
Answer
Replication provides:
- High Availability
- Fault Tolerance
- Zero Single Point of Failure
- Continuous Processing
- Disaster Recovery
Without replication
Broker Failure
↓
Data Lost
With replication
Broker Failure
↓
Follower becomes Leader
↓
No Data Loss
Q3. What is Replication Factor?
Answer
Replication Factor determines how many copies of a partition exist.
Example
Replication Factor = 3
Partition copies
Leader
Follower
Follower
Recommendation
| Environment | Replication Factor |
|---|---|
| Development | 1 |
| Testing | 2 |
| Production | 3 |
Replication Factor
flowchart LR
Partition --> Leader
Partition --> Follower1["Follower 1"]
Partition --> Follower2["Follower 2"]
Q4. What is a Leader Replica?
Answer
Every partition has exactly one Leader.
The Leader handles:
- Producer Writes
- Consumer Reads
- Replication Coordination
All client traffic goes through the Leader.
Leader Flow
flowchart LR
Producer --> Leader
Leader --> Consumers
Leader --> Followers
Q5. What is a Follower Replica?
Answer
Follower replicas continuously copy data from the Leader.
Followers do not normally serve producer write requests.
Responsibilities:
- Replicate Data
- Stay Synchronized
- Become Leader if necessary
Followers
flowchart LR
Leader --> FollowerA["Follower A"]
Leader --> FollowerB["Follower B"]
Q6. What is ISR (In-Sync Replica)?
Answer
ISR stands for In-Sync Replica.
It is the list of replicas that are fully synchronized with the Leader.
Example
Leader
Follower 1
Follower 2
All three belong to ISR.
If a follower falls behind significantly, Kafka removes it from ISR until it catches up.
ISR
flowchart LR
Leader --> ISR
ISR --> Follower1["Follower 1"]
ISR --> Follower2["Follower 2"]
Q7. What happens when the Leader Broker crashes?
Answer
Kafka automatically performs Leader Election.
Steps
- Leader crashes.
- Controller detects failure.
- One ISR replica becomes Leader.
- Producers reconnect.
- Consumers continue reading.
Failover
sequenceDiagram
participant Producer
participant Leader
participant Follower
participant Controller
Leader-->>Controller: Broker Failure
Controller->>Follower: Promote Leader
Producer->>Follower: Continue Writes
Q8. What is Leader Election?
Answer
Leader Election selects a new Leader when the existing Leader becomes unavailable.
Kafka always prefers replicas that belong to the ISR.
Benefits
- Automatic Recovery
- High Availability
- Minimal Downtime
Leader Election
flowchart TD
LeaderFailed["Leader Failed"] --> Controller
Controller --> IsrReplica["ISR Replica"]
IsrReplica["ISR Replica"] --> NewLeader["New Leader"]
Q9. What is min.insync.replicas?
Answer
min.insync.replicas specifies the minimum number of synchronized replicas required before Kafka acknowledges a write.
Example
Replication Factor = 3
min.insync.replicas = 2
Producer with
acks=all
requires at least two ISR replicas before acknowledging success.
Benefits
- Prevents Data Loss
- Improves Durability
Write Flow
flowchart LR
Producer --> Leader
Leader --> Follower1["Follower 1"]
Leader --> Follower2["Follower 2"]
Follower1["Follower 1"] --> ACK
Follower2["Follower 2"] --> ACK
Q10. How do acknowledgements affect replication?
Answer
Kafka supports three acknowledgement modes.
| ACK Mode | Reliability |
|---|---|
| acks=0 | Lowest |
| acks=1 | Medium |
| acks=all | Highest |
For production systems:
acks=all
+
Replication Factor = 3
+
min.insync.replicas = 2
This combination provides excellent durability.
Q11. How does replication improve fault tolerance?
Answer
Replication ensures that multiple brokers contain identical partition data.
If one broker becomes unavailable:
- Another replica continues serving requests.
- No manual intervention is required.
- Consumers continue processing.
Fault Tolerance
flowchart TD
BrokerFailure["Broker Failure"] --> ReplicaPromotion["Replica Promotion"]
ReplicaPromotion["Replica Promotion"] --> ContinueProcessing["Continue Processing"]
Q12. Can consumers read from followers?
Answer
By default:
No
Consumers read only from the Leader.
Reasons:
- Strong consistency
- Ordered reads
- Simpler architecture
Some advanced configurations support follower fetching for specific scenarios, but standard Kafka consumer behavior reads from leaders.
Q13. What happens if an ISR replica falls behind?
Answer
If a follower cannot keep up with the leader:
- Kafka removes it from ISR.
- It continues replicating.
- Once synchronized, it rejoins ISR.
This prevents stale replicas from participating in leader elections.
ISR Recovery
flowchart LR
FollowerSlow["Follower Slow"] --> RemovedFromIsr["Removed From ISR"]
RemovedFromIsr["Removed From ISR"] --> CatchUp["Catch Up"]
CatchUp["Catch Up"] --> JoinIsr["Join ISR"]
Q14. What are common replication interview questions?
Answer
Interviewers frequently ask:
- Leader vs Follower?
- What is ISR?
- Replication Factor?
- Leader Election?
- Broker Failure?
acks=all?min.insync.replicas?- Why replication?
- Why consumers read leaders?
- How does Kafka prevent data loss?
Q15. What are production best practices?
Answer
Recommended configuration:
- Replication Factor = 3
acks=allmin.insync.replicas=2- Rack Awareness
- Multiple Brokers
- Monitor ISR Shrink Events
- Monitor Under-Replicated Partitions
- Monitor Leader Elections
- Test Failover Regularly
- Avoid Unclean Leader Election
Enterprise Architecture
flowchart TD
SpringBoot["Spring Boot"] --> KafkaCluster["Kafka Cluster"]
KafkaCluster["Kafka Cluster"] --> Broker1["Broker 1"]
KafkaCluster["Kafka Cluster"] --> Broker2["Broker 2"]
KafkaCluster["Kafka Cluster"] --> Broker3["Broker 3"]
Broker1["Broker 1"] --> Leader
Broker2["Broker 2"] --> Follower
Broker3["Broker 3"] --> Follower
Leader --> Consumers
Replication Lifecycle
sequenceDiagram
participant Producer
participant Leader
participant Follower1
participant Follower2
Producer->>Leader: Publish Event
Leader->>Follower1: Replicate
Leader->>Follower2: Replicate
Follower1-->>Leader: ACK
Follower2-->>Leader: ACK
Leader-->>Producer: Success
Kafka Replication Overview
mindmap
root((Kafka Replication))
Leader
Followers
ISR
Replication Factor
Leader Election
ACK
High Availability
Leader vs Follower
| Leader | Follower |
|---|---|
| Handles Writes | Replicates Data |
| Serves Reads | Copies Leader |
| Coordinates Replication | Ready for Failover |
| One Per Partition | Zero or More |
Real Banking Example
A bank processes 5 million transactions per day.
Architecture:
Transfer Service
↓
Kafka Cluster
↓
Leader Partition
↓
Follower Broker A
↓
Follower Broker B
↓
Consumer Groups
↓
Core Banking
↓
Fraud Detection
↓
Notifications
If the broker hosting the leader partition fails during business hours:
- Kafka Controller elects a new leader from the ISR.
- Producers automatically reconnect.
- Consumers continue reading from the new leader.
- Banking transactions continue with minimal interruption.
Senior Interview Tips
Interviewers commonly ask:
- Explain Kafka Replication.
- What is Replication Factor?
- What is ISR?
- Leader vs Follower?
- What happens when a Leader crashes?
- Explain Leader Election.
- Why use
acks=all? - What is
min.insync.replicas? - What are Under-Replicated Partitions?
- How do you configure Kafka for zero data loss?
Remember these key points:
- One Leader handles all reads and writes for a partition.
- Followers continuously replicate the leader's data.
- ISR contains replicas fully synchronized with the leader.
- Replication Factor 3 +
acks=all+min.insync.replicas=2is a common production configuration.
Quick Revision
- Kafka Replication creates multiple copies of partition data across brokers.
- Every partition has one leader and zero or more follower replicas.
- Producers write to the leader, and followers replicate data asynchronously.
- ISR contains replicas that are fully synchronized with the leader.
- Leader Election automatically promotes an ISR replica after a broker failure.
- Replication improves availability, durability, and fault tolerance.
- Configure
acks=allwith an appropriatemin.insync.replicasvalue for reliable writes. - Monitor ISR health, under-replicated partitions, and leader elections in production.
- Avoid unclean leader election in production environments.
- Kafka Replication is fundamental to building resilient, enterprise-grade event streaming platforms.