Kafka Monitoring Interview Questions and Answers

Learn Kafka Monitoring with real-world interview questions covering consumer lag, broker metrics, JMX, Prometheus, Grafana, alerts, dashboards, and production best practices.

Kafka Monitoring Interview Questions and Answers

Monitoring is one of the most critical aspects of operating Kafka in production.

A Kafka cluster may appear healthy while:

  • Consumers are falling behind.
  • Brokers are running out of disk space.
  • Replication is unhealthy.
  • Producers are experiencing high latency.
  • Messages are not being processed on time.

A robust monitoring strategy helps detect these issues before they impact business applications.


Kafka Monitoring Architecture

flowchart LR

KafkaCluster["Kafka Cluster"] --> JmxMetrics["JMX Metrics"]

JmxMetrics["JMX Metrics"] --> Prometheus

Prometheus --> Grafana

Grafana --> AlertManager

AlertManager --> OperationsTeam["Operations Team"]

Q1. Why is Kafka Monitoring important?

Answer

Kafka Monitoring helps ensure:

  • High Availability
  • Reliable Processing
  • Low Latency
  • Healthy Replication
  • Fast Incident Detection
  • Capacity Planning

Without monitoring:

  • Consumer lag may grow unnoticed.
  • Broker failures may impact production.
  • Disk exhaustion can stop writes.
  • Replication issues may increase the risk of data loss.

Monitoring Flow

flowchart LR

KafkaCluster["Kafka Cluster"] --> Metrics
Metrics --> Dashboard

Dashboard --> Alerts

Q2. What are the most important Kafka metrics?

Answer

Critical metrics include:

  • Consumer Lag
  • Broker CPU
  • Memory Usage
  • Disk Usage
  • Network Throughput
  • Request Latency
  • Under Replicated Partitions
  • Active Controller Count
  • Offline Partitions

Metrics Overview

mindmap
  root((Kafka Metrics))
    Consumer Lag
    CPU
    Memory
    Disk
    Network
    Latency
    Replication

Q3. What is Consumer Lag?

Answer

Consumer Lag is the difference between:

  • Latest Offset in a Partition

and

  • Consumer's Committed Offset

Formula

Lag = Latest Offset - Committed Offset

Example

Latest Offset = 1200

Committed Offset = 1100

Lag = 100

Consumer Lag

flowchart LR

LatestOffset["Latest Offset"] --> Lag

CommittedOffset["Committed Offset"] --> Lag

Interview Tip

Growing lag usually indicates that consumers are unable to keep up with producers.


Q4. What causes Consumer Lag?

Answer

Common causes:

  • Slow Business Logic
  • Database Bottlenecks
  • Network Latency
  • Too Few Consumers
  • Insufficient Partitions
  • Frequent Rebalancing
  • Large Messages
  • External API Delays

Lag Causes

flowchart TD

ConsumerLag["Consumer Lag"] --> SlowConsumer["Slow Consumer"]

ConsumerLag["Consumer Lag"] --> DatabaseDelay["Database Delay"]

ConsumerLag["Consumer Lag"] --> Network

ConsumerLag["Consumer Lag"] --> Rebalancing

Q5. What are Under Replicated Partitions (URP)?

Answer

A partition becomes Under Replicated when one or more follower replicas fall behind the leader.

Example

Leader

Follower 1 ✓

Follower 2 ✗

The partition is now under replicated.

URP

flowchart LR

Leader --> Follower1["Follower 1"]

Leader --> Follower2["Follower 2"]

Follower2["Follower 2"] --> OutOfSync["Out Of Sync"]

Why it matters

High URP counts increase the risk of data loss if another broker fails.


Q6. What is Active Controller Count?

Answer

Kafka clusters should have exactly one active controller.

The controller manages:

  • Broker Membership
  • Leader Elections
  • Partition Assignments

Expected Value

1

If:

0

or

>1

the cluster requires immediate investigation.


Controller

flowchart TD

Controller --> LeaderElection["Leader Election"]

Controller --> BrokerManagement["Broker Management"]

Controller --> PartitionAssignment["Partition Assignment"]

Q7. What Broker metrics should be monitored?

Answer

Monitor:

  • CPU Usage
  • Heap Memory
  • Garbage Collection
  • Disk Usage
  • Network Traffic
  • Open File Descriptors
  • Request Queue Size
  • Produce Request Latency
  • Fetch Request Latency

Broker Monitoring

flowchart LR

Broker --> CPU

Broker --> Memory

Broker --> Disk

Broker --> Network

Q8. How do you monitor Kafka?

Answer

Common monitoring stack:

  • JMX Metrics
  • Prometheus
  • Grafana
  • AlertManager

Enterprise tools:

  • Datadog
  • Dynatrace
  • Splunk
  • New Relic

Monitoring Pipeline

flowchart LR

Kafka --> JMX
JMX --> Prometheus

Prometheus --> Grafana
Grafana --> Alerts

Q9. What alerts should be configured?

Answer

Recommended alerts:

  • High Consumer Lag
  • Broker Down
  • Under Replicated Partitions
  • Offline Partitions
  • High Disk Usage
  • High CPU
  • High Memory
  • Frequent Rebalances
  • High Produce Latency
  • High Fetch Latency

Alert Flow

flowchart TD

KafkaMetrics["Kafka Metrics"] --> Threshold
Threshold --> Alert

Alert --> OperationsTeam["Operations Team"]

Q10. How do you troubleshoot high Consumer Lag?

Answer

Investigation steps:

  1. Check consumer health.
  2. Check application logs.
  3. Check database response times.
  4. Verify partition count.
  5. Check rebalancing frequency.
  6. Check broker performance.
  7. Review downstream dependencies.

Troubleshooting

flowchart TD

HighLag["High Lag"] --> Investigate

Investigate --> FixRootCause["Fix Root Cause"]

FixRootCause["Fix Root Cause"] --> LagReduced["Lag Reduced"]

Q11. What dashboards are useful in production?

Answer

Recommended dashboards:

  • Cluster Health
  • Consumer Lag
  • Broker Metrics
  • Producer Metrics
  • Replication
  • JVM
  • Network
  • Storage

Dashboard

flowchart LR

Kafka --> GrafanaDashboard["Grafana Dashboard"]

GrafanaDashboard["Grafana Dashboard"] --> Operations

GrafanaDashboard["Grafana Dashboard"] --> Developers

Q12. How do you monitor Kafka in Spring Boot?

Answer

Spring Boot applications can expose metrics using:

  • Micrometer
  • Prometheus
  • Actuator
  • JMX

Monitor:

  • Producer Success Rate
  • Consumer Processing Time
  • Retry Count
  • Dead Letter Topic Count
  • Message Processing Errors

Spring Boot

flowchart TD

SpringBoot["Spring Boot"] --> Micrometer
Micrometer --> Prometheus

Prometheus --> Grafana

Q13. What are common monitoring mistakes?

Answer

Common mistakes include:

  • Monitoring only brokers
  • Ignoring consumer lag
  • No alerting
  • Ignoring disk usage
  • Ignoring replication health
  • Monitoring only CPU
  • Not tracking throughput
  • No capacity planning

Common Problems

flowchart TD

PoorMonitoring["Poor Monitoring"] --> HiddenIssues["Hidden Issues"]

HiddenIssues["Hidden Issues"] --> ProductionIncident["Production Incident"]

Q14. What are production monitoring best practices?

Answer

Best practices:

  • Monitor consumer lag continuously.
  • Monitor replication health.
  • Track broker availability.
  • Monitor JVM metrics.
  • Alert before disk reaches critical levels.
  • Monitor throughput trends.
  • Track rebalance frequency.
  • Build business-level dashboards.
  • Test alerts regularly.
  • Review capacity monthly.

Enterprise Monitoring

flowchart TD

KafkaCluster["Kafka Cluster"] --> Metrics

Metrics --> Prometheus

Prometheus --> Grafana

Grafana --> AlertManager

AlertManager --> Slack

AlertManager --> Email

AlertManager --> PagerDuty

Kafka Monitoring Lifecycle

sequenceDiagram
participant Kafka
participant JMX
participant Prometheus
participant Grafana
participant AlertManager
Kafka->>JMX: Publish Metrics
JMX->>Prometheus: Collect
Prometheus->>Grafana: Visualize
Prometheus->>AlertManager: Threshold Exceeded
AlertManager->>Operations Team: Notify

Kafka Monitoring Overview

mindmap
  root((Kafka Monitoring))
    Consumer Lag
    Broker Health
    Disk
    CPU
    Replication
    JMX
    Prometheus
    Grafana

Important Kafka Metrics

Metric Why Monitor It
Consumer Lag Processing Health
Under Replicated Partitions Replication Health
Offline Partitions Availability
Broker CPU Resource Utilization
Heap Memory JVM Health
Disk Usage Storage Capacity
Produce Latency Producer Performance
Fetch Latency Consumer Performance
Network Throughput Traffic Patterns
Active Controller Count Cluster Stability

Real Banking Example

A digital banking platform processes 15 million transactions daily.

Monitoring architecture:

Spring Boot Services

↓

Kafka Cluster

↓

JMX Metrics

↓

Prometheus

↓

Grafana

↓

AlertManager

↓

DevOps Team

Configured alerts:

  • Consumer Lag > 10,000
  • Disk Usage > 80%
  • Under Replicated Partitions > 0
  • Broker Down
  • Produce Latency > 200 ms

This allows the operations team to respond before customer-facing services are affected.


Senior Interview Tips

Interviewers commonly ask:

  • What is Consumer Lag?
  • How do you calculate Consumer Lag?
  • What causes Consumer Lag?
  • What are Under Replicated Partitions?
  • What is Active Controller Count?
  • Which Kafka metrics are most important?
  • How do you monitor Kafka?
  • Which monitoring tools have you used?
  • What alerts do you configure?
  • How do you troubleshoot high lag?
  • What dashboards do you build?
  • How does Spring Boot expose Kafka metrics?

Remember:

  • Consumer Lag is the most important operational metric.
  • Under Replicated Partitions indicate replication problems.
  • Exactly one Active Controller should exist in a healthy cluster.
  • Good monitoring combines infrastructure metrics with business metrics.

Quick Revision

  • Kafka Monitoring is essential for maintaining healthy production clusters.
  • Monitor consumer lag, replication health, broker resources, and request latency.
  • Consumer Lag measures how far consumers are behind producers.
  • Under Replicated Partitions indicate followers are not keeping up with leaders.
  • Active Controller Count should always be exactly one.
  • Use JMX, Prometheus, Grafana, and AlertManager for observability.
  • Configure alerts for lag, disk usage, broker failures, replication issues, and latency.
  • Spring Boot integrates with Micrometer and Actuator for application-level metrics.
  • Build dashboards for infrastructure and business KPIs.
  • Continuous monitoring enables proactive issue detection, capacity planning, and reliable Kafka operations.