Kafka Monitoring Interview Questions and Answers
Learn Kafka Monitoring with real-world interview questions covering consumer lag, broker metrics, JMX, Prometheus, Grafana, alerts, dashboards, and production best practices.
Kafka Monitoring Interview Questions and Answers
Monitoring is one of the most critical aspects of operating Kafka in production.
A Kafka cluster may appear healthy while:
- Consumers are falling behind.
- Brokers are running out of disk space.
- Replication is unhealthy.
- Producers are experiencing high latency.
- Messages are not being processed on time.
A robust monitoring strategy helps detect these issues before they impact business applications.
Kafka Monitoring Architecture
flowchart LR
KafkaCluster["Kafka Cluster"] --> JmxMetrics["JMX Metrics"]
JmxMetrics["JMX Metrics"] --> Prometheus
Prometheus --> Grafana
Grafana --> AlertManager
AlertManager --> OperationsTeam["Operations Team"]
Q1. Why is Kafka Monitoring important?
Answer
Kafka Monitoring helps ensure:
- High Availability
- Reliable Processing
- Low Latency
- Healthy Replication
- Fast Incident Detection
- Capacity Planning
Without monitoring:
- Consumer lag may grow unnoticed.
- Broker failures may impact production.
- Disk exhaustion can stop writes.
- Replication issues may increase the risk of data loss.
Monitoring Flow
flowchart LR
KafkaCluster["Kafka Cluster"] --> Metrics
Metrics --> Dashboard
Dashboard --> Alerts
Q2. What are the most important Kafka metrics?
Answer
Critical metrics include:
- Consumer Lag
- Broker CPU
- Memory Usage
- Disk Usage
- Network Throughput
- Request Latency
- Under Replicated Partitions
- Active Controller Count
- Offline Partitions
Metrics Overview
mindmap
root((Kafka Metrics))
Consumer Lag
CPU
Memory
Disk
Network
Latency
Replication
Q3. What is Consumer Lag?
Answer
Consumer Lag is the difference between:
- Latest Offset in a Partition
and
- Consumer's Committed Offset
Formula
Lag = Latest Offset - Committed Offset
Example
Latest Offset = 1200
Committed Offset = 1100
Lag = 100
Consumer Lag
flowchart LR
LatestOffset["Latest Offset"] --> Lag
CommittedOffset["Committed Offset"] --> Lag
Interview Tip
Growing lag usually indicates that consumers are unable to keep up with producers.
Q4. What causes Consumer Lag?
Answer
Common causes:
- Slow Business Logic
- Database Bottlenecks
- Network Latency
- Too Few Consumers
- Insufficient Partitions
- Frequent Rebalancing
- Large Messages
- External API Delays
Lag Causes
flowchart TD
ConsumerLag["Consumer Lag"] --> SlowConsumer["Slow Consumer"]
ConsumerLag["Consumer Lag"] --> DatabaseDelay["Database Delay"]
ConsumerLag["Consumer Lag"] --> Network
ConsumerLag["Consumer Lag"] --> Rebalancing
Q5. What are Under Replicated Partitions (URP)?
Answer
A partition becomes Under Replicated when one or more follower replicas fall behind the leader.
Example
Leader
Follower 1 ✓
Follower 2 ✗
The partition is now under replicated.
URP
flowchart LR
Leader --> Follower1["Follower 1"]
Leader --> Follower2["Follower 2"]
Follower2["Follower 2"] --> OutOfSync["Out Of Sync"]
Why it matters
High URP counts increase the risk of data loss if another broker fails.
Q6. What is Active Controller Count?
Answer
Kafka clusters should have exactly one active controller.
The controller manages:
- Broker Membership
- Leader Elections
- Partition Assignments
Expected Value
1
If:
0
or
>1
the cluster requires immediate investigation.
Controller
flowchart TD
Controller --> LeaderElection["Leader Election"]
Controller --> BrokerManagement["Broker Management"]
Controller --> PartitionAssignment["Partition Assignment"]
Q7. What Broker metrics should be monitored?
Answer
Monitor:
- CPU Usage
- Heap Memory
- Garbage Collection
- Disk Usage
- Network Traffic
- Open File Descriptors
- Request Queue Size
- Produce Request Latency
- Fetch Request Latency
Broker Monitoring
flowchart LR
Broker --> CPU
Broker --> Memory
Broker --> Disk
Broker --> Network
Q8. How do you monitor Kafka?
Answer
Common monitoring stack:
- JMX Metrics
- Prometheus
- Grafana
- AlertManager
Enterprise tools:
- Datadog
- Dynatrace
- Splunk
- New Relic
Monitoring Pipeline
flowchart LR
Kafka --> JMX
JMX --> Prometheus
Prometheus --> Grafana
Grafana --> Alerts
Q9. What alerts should be configured?
Answer
Recommended alerts:
- High Consumer Lag
- Broker Down
- Under Replicated Partitions
- Offline Partitions
- High Disk Usage
- High CPU
- High Memory
- Frequent Rebalances
- High Produce Latency
- High Fetch Latency
Alert Flow
flowchart TD
KafkaMetrics["Kafka Metrics"] --> Threshold
Threshold --> Alert
Alert --> OperationsTeam["Operations Team"]
Q10. How do you troubleshoot high Consumer Lag?
Answer
Investigation steps:
- Check consumer health.
- Check application logs.
- Check database response times.
- Verify partition count.
- Check rebalancing frequency.
- Check broker performance.
- Review downstream dependencies.
Troubleshooting
flowchart TD
HighLag["High Lag"] --> Investigate
Investigate --> FixRootCause["Fix Root Cause"]
FixRootCause["Fix Root Cause"] --> LagReduced["Lag Reduced"]
Q11. What dashboards are useful in production?
Answer
Recommended dashboards:
- Cluster Health
- Consumer Lag
- Broker Metrics
- Producer Metrics
- Replication
- JVM
- Network
- Storage
Dashboard
flowchart LR
Kafka --> GrafanaDashboard["Grafana Dashboard"]
GrafanaDashboard["Grafana Dashboard"] --> Operations
GrafanaDashboard["Grafana Dashboard"] --> Developers
Q12. How do you monitor Kafka in Spring Boot?
Answer
Spring Boot applications can expose metrics using:
- Micrometer
- Prometheus
- Actuator
- JMX
Monitor:
- Producer Success Rate
- Consumer Processing Time
- Retry Count
- Dead Letter Topic Count
- Message Processing Errors
Spring Boot
flowchart TD
SpringBoot["Spring Boot"] --> Micrometer
Micrometer --> Prometheus
Prometheus --> Grafana
Q13. What are common monitoring mistakes?
Answer
Common mistakes include:
- Monitoring only brokers
- Ignoring consumer lag
- No alerting
- Ignoring disk usage
- Ignoring replication health
- Monitoring only CPU
- Not tracking throughput
- No capacity planning
Common Problems
flowchart TD
PoorMonitoring["Poor Monitoring"] --> HiddenIssues["Hidden Issues"]
HiddenIssues["Hidden Issues"] --> ProductionIncident["Production Incident"]
Q14. What are production monitoring best practices?
Answer
Best practices:
- Monitor consumer lag continuously.
- Monitor replication health.
- Track broker availability.
- Monitor JVM metrics.
- Alert before disk reaches critical levels.
- Monitor throughput trends.
- Track rebalance frequency.
- Build business-level dashboards.
- Test alerts regularly.
- Review capacity monthly.
Enterprise Monitoring
flowchart TD
KafkaCluster["Kafka Cluster"] --> Metrics
Metrics --> Prometheus
Prometheus --> Grafana
Grafana --> AlertManager
AlertManager --> Slack
AlertManager --> Email
AlertManager --> PagerDuty
Kafka Monitoring Lifecycle
sequenceDiagram
participant Kafka
participant JMX
participant Prometheus
participant Grafana
participant AlertManager
Kafka->>JMX: Publish Metrics
JMX->>Prometheus: Collect
Prometheus->>Grafana: Visualize
Prometheus->>AlertManager: Threshold Exceeded
AlertManager->>Operations Team: Notify
Kafka Monitoring Overview
mindmap
root((Kafka Monitoring))
Consumer Lag
Broker Health
Disk
CPU
Replication
JMX
Prometheus
Grafana
Important Kafka Metrics
| Metric | Why Monitor It |
|---|---|
| Consumer Lag | Processing Health |
| Under Replicated Partitions | Replication Health |
| Offline Partitions | Availability |
| Broker CPU | Resource Utilization |
| Heap Memory | JVM Health |
| Disk Usage | Storage Capacity |
| Produce Latency | Producer Performance |
| Fetch Latency | Consumer Performance |
| Network Throughput | Traffic Patterns |
| Active Controller Count | Cluster Stability |
Real Banking Example
A digital banking platform processes 15 million transactions daily.
Monitoring architecture:
Spring Boot Services
↓
Kafka Cluster
↓
JMX Metrics
↓
Prometheus
↓
Grafana
↓
AlertManager
↓
DevOps Team
Configured alerts:
- Consumer Lag > 10,000
- Disk Usage > 80%
- Under Replicated Partitions > 0
- Broker Down
- Produce Latency > 200 ms
This allows the operations team to respond before customer-facing services are affected.
Senior Interview Tips
Interviewers commonly ask:
- What is Consumer Lag?
- How do you calculate Consumer Lag?
- What causes Consumer Lag?
- What are Under Replicated Partitions?
- What is Active Controller Count?
- Which Kafka metrics are most important?
- How do you monitor Kafka?
- Which monitoring tools have you used?
- What alerts do you configure?
- How do you troubleshoot high lag?
- What dashboards do you build?
- How does Spring Boot expose Kafka metrics?
Remember:
- Consumer Lag is the most important operational metric.
- Under Replicated Partitions indicate replication problems.
- Exactly one Active Controller should exist in a healthy cluster.
- Good monitoring combines infrastructure metrics with business metrics.
Quick Revision
- Kafka Monitoring is essential for maintaining healthy production clusters.
- Monitor consumer lag, replication health, broker resources, and request latency.
- Consumer Lag measures how far consumers are behind producers.
- Under Replicated Partitions indicate followers are not keeping up with leaders.
- Active Controller Count should always be exactly one.
- Use JMX, Prometheus, Grafana, and AlertManager for observability.
- Configure alerts for lag, disk usage, broker failures, replication issues, and latency.
- Spring Boot integrates with Micrometer and Actuator for application-level metrics.
- Build dashboards for infrastructure and business KPIs.
- Continuous monitoring enables proactive issue detection, capacity planning, and reliable Kafka operations.