Kafka Replication Interview Questions and Answers

Master Kafka Replication with real-world interview questions covering leaders, followers, ISR, replication factor, leader election, failover, min.insync.replicas, acknowledgements, and production best practices.

Kafka Replication Interview Questions and Answers

Replication is one of Kafka's most important features because it provides:

  • High Availability
  • Fault Tolerance
  • Data Durability
  • Disaster Recovery
  • Zero Data Loss (Proper Configuration)

Without replication, a single broker failure could permanently lose messages.

This article covers everything interviewers expect regarding Kafka Replication.


Kafka Replication Architecture

flowchart LR

Producer --> LeaderPartition["Leader Partition"]

LeaderPartition["Leader Partition"] --> Follower1["Follower 1"]

LeaderPartition["Leader Partition"] --> Follower2["Follower 2"]

Follower1["Follower 1"] --> Consumer

LeaderPartition["Leader Partition"] --> Consumer

Q1. What is Kafka Replication?

Answer

Kafka Replication is the process of maintaining multiple copies of the same partition across different brokers.

Instead of storing data on one broker only, Kafka stores replicas on multiple brokers.

Example

Partition-0

Broker 1 (Leader)

Broker 2 (Follower)

Broker 3 (Follower)

If Broker 1 fails, another replica becomes the leader.


Replication

flowchart LR

Leader --> Replica1["Replica 1"]

Leader --> Replica2["Replica 2"]

Q2. Why is Replication required?

Answer

Replication provides:

  • High Availability
  • Fault Tolerance
  • Zero Single Point of Failure
  • Continuous Processing
  • Disaster Recovery

Without replication

Broker Failure

↓

Data Lost

With replication

Broker Failure

↓

Follower becomes Leader

↓

No Data Loss

Q3. What is Replication Factor?

Answer

Replication Factor determines how many copies of a partition exist.

Example

Replication Factor = 3

Partition copies

Leader

Follower

Follower

Recommendation

Environment Replication Factor
Development 1
Testing 2
Production 3

Replication Factor

flowchart LR

Partition --> Leader

Partition --> Follower1["Follower 1"]

Partition --> Follower2["Follower 2"]

Q4. What is a Leader Replica?

Answer

Every partition has exactly one Leader.

The Leader handles:

  • Producer Writes
  • Consumer Reads
  • Replication Coordination

All client traffic goes through the Leader.


Leader Flow

flowchart LR

Producer --> Leader

Leader --> Consumers

Leader --> Followers

Q5. What is a Follower Replica?

Answer

Follower replicas continuously copy data from the Leader.

Followers do not normally serve producer write requests.

Responsibilities:

  • Replicate Data
  • Stay Synchronized
  • Become Leader if necessary

Followers

flowchart LR

Leader --> FollowerA["Follower A"]

Leader --> FollowerB["Follower B"]

Q6. What is ISR (In-Sync Replica)?

Answer

ISR stands for In-Sync Replica.

It is the list of replicas that are fully synchronized with the Leader.

Example

Leader

Follower 1

Follower 2

All three belong to ISR.

If a follower falls behind significantly, Kafka removes it from ISR until it catches up.


ISR

flowchart LR

Leader --> ISR

ISR --> Follower1["Follower 1"]

ISR --> Follower2["Follower 2"]

Q7. What happens when the Leader Broker crashes?

Answer

Kafka automatically performs Leader Election.

Steps

  1. Leader crashes.
  2. Controller detects failure.
  3. One ISR replica becomes Leader.
  4. Producers reconnect.
  5. Consumers continue reading.

Failover

sequenceDiagram
participant Producer
participant Leader
participant Follower
participant Controller
Leader-->>Controller: Broker Failure
Controller->>Follower: Promote Leader
Producer->>Follower: Continue Writes

Q8. What is Leader Election?

Answer

Leader Election selects a new Leader when the existing Leader becomes unavailable.

Kafka always prefers replicas that belong to the ISR.

Benefits

  • Automatic Recovery
  • High Availability
  • Minimal Downtime

Leader Election

flowchart TD

LeaderFailed["Leader Failed"] --> Controller
Controller --> IsrReplica["ISR Replica"]

IsrReplica["ISR Replica"] --> NewLeader["New Leader"]

Q9. What is min.insync.replicas?

Answer

min.insync.replicas specifies the minimum number of synchronized replicas required before Kafka acknowledges a write.

Example

Replication Factor = 3

min.insync.replicas = 2

Producer with

acks=all

requires at least two ISR replicas before acknowledging success.

Benefits

  • Prevents Data Loss
  • Improves Durability

Write Flow

flowchart LR

Producer --> Leader

Leader --> Follower1["Follower 1"]

Leader --> Follower2["Follower 2"]

Follower1["Follower 1"] --> ACK

Follower2["Follower 2"] --> ACK

Q10. How do acknowledgements affect replication?

Answer

Kafka supports three acknowledgement modes.

ACK Mode Reliability
acks=0 Lowest
acks=1 Medium
acks=all Highest

For production systems:

acks=all

+

Replication Factor = 3

+

min.insync.replicas = 2

This combination provides excellent durability.


Q11. How does replication improve fault tolerance?

Answer

Replication ensures that multiple brokers contain identical partition data.

If one broker becomes unavailable:

  • Another replica continues serving requests.
  • No manual intervention is required.
  • Consumers continue processing.

Fault Tolerance

flowchart TD

BrokerFailure["Broker Failure"] --> ReplicaPromotion["Replica Promotion"]
ReplicaPromotion["Replica Promotion"] --> ContinueProcessing["Continue Processing"]

Q12. Can consumers read from followers?

Answer

By default:

No

Consumers read only from the Leader.

Reasons:

  • Strong consistency
  • Ordered reads
  • Simpler architecture

Some advanced configurations support follower fetching for specific scenarios, but standard Kafka consumer behavior reads from leaders.


Q13. What happens if an ISR replica falls behind?

Answer

If a follower cannot keep up with the leader:

  • Kafka removes it from ISR.
  • It continues replicating.
  • Once synchronized, it rejoins ISR.

This prevents stale replicas from participating in leader elections.


ISR Recovery

flowchart LR

FollowerSlow["Follower Slow"] --> RemovedFromIsr["Removed From ISR"]
RemovedFromIsr["Removed From ISR"] --> CatchUp["Catch Up"]

CatchUp["Catch Up"] --> JoinIsr["Join ISR"]

Q14. What are common replication interview questions?

Answer

Interviewers frequently ask:

  • Leader vs Follower?
  • What is ISR?
  • Replication Factor?
  • Leader Election?
  • Broker Failure?
  • acks=all?
  • min.insync.replicas?
  • Why replication?
  • Why consumers read leaders?
  • How does Kafka prevent data loss?

Q15. What are production best practices?

Answer

Recommended configuration:

  • Replication Factor = 3
  • acks=all
  • min.insync.replicas=2
  • Rack Awareness
  • Multiple Brokers
  • Monitor ISR Shrink Events
  • Monitor Under-Replicated Partitions
  • Monitor Leader Elections
  • Test Failover Regularly
  • Avoid Unclean Leader Election

Enterprise Architecture

flowchart TD

SpringBoot["Spring Boot"] --> KafkaCluster["Kafka Cluster"]

KafkaCluster["Kafka Cluster"] --> Broker1["Broker 1"]

KafkaCluster["Kafka Cluster"] --> Broker2["Broker 2"]

KafkaCluster["Kafka Cluster"] --> Broker3["Broker 3"]

Broker1["Broker 1"] --> Leader

Broker2["Broker 2"] --> Follower

Broker3["Broker 3"] --> Follower

Leader --> Consumers

Replication Lifecycle

sequenceDiagram
participant Producer
participant Leader
participant Follower1
participant Follower2
Producer->>Leader: Publish Event
Leader->>Follower1: Replicate
Leader->>Follower2: Replicate
Follower1-->>Leader: ACK
Follower2-->>Leader: ACK
Leader-->>Producer: Success

Kafka Replication Overview

mindmap
  root((Kafka Replication))
    Leader
    Followers
    ISR
    Replication Factor
    Leader Election
    ACK
    High Availability

Leader vs Follower

Leader Follower
Handles Writes Replicates Data
Serves Reads Copies Leader
Coordinates Replication Ready for Failover
One Per Partition Zero or More

Real Banking Example

A bank processes 5 million transactions per day.

Architecture:

Transfer Service

↓

Kafka Cluster

↓

Leader Partition

↓

Follower Broker A

↓

Follower Broker B

↓

Consumer Groups

↓

Core Banking

↓

Fraud Detection

↓

Notifications

If the broker hosting the leader partition fails during business hours:

  • Kafka Controller elects a new leader from the ISR.
  • Producers automatically reconnect.
  • Consumers continue reading from the new leader.
  • Banking transactions continue with minimal interruption.

Senior Interview Tips

Interviewers commonly ask:

  • Explain Kafka Replication.
  • What is Replication Factor?
  • What is ISR?
  • Leader vs Follower?
  • What happens when a Leader crashes?
  • Explain Leader Election.
  • Why use acks=all?
  • What is min.insync.replicas?
  • What are Under-Replicated Partitions?
  • How do you configure Kafka for zero data loss?

Remember these key points:

  • One Leader handles all reads and writes for a partition.
  • Followers continuously replicate the leader's data.
  • ISR contains replicas fully synchronized with the leader.
  • Replication Factor 3 + acks=all + min.insync.replicas=2 is a common production configuration.

Quick Revision

  • Kafka Replication creates multiple copies of partition data across brokers.
  • Every partition has one leader and zero or more follower replicas.
  • Producers write to the leader, and followers replicate data asynchronously.
  • ISR contains replicas that are fully synchronized with the leader.
  • Leader Election automatically promotes an ISR replica after a broker failure.
  • Replication improves availability, durability, and fault tolerance.
  • Configure acks=all with an appropriate min.insync.replicas value for reliable writes.
  • Monitor ISR health, under-replicated partitions, and leader elections in production.
  • Avoid unclean leader election in production environments.
  • Kafka Replication is fundamental to building resilient, enterprise-grade event streaming platforms.