Real World Production Scenarios

Master real-world production troubleshooting scenarios covering Kubernetes, Docker, AWS, CI/CD, Linux, Databases, Monitoring, Logging, DevSecOps, SRE, and enterprise production incident handling.

Introduction

One of the biggest differences between a Junior DevOps Engineer and a Senior DevOps Engineer is the ability to solve production problems under pressure.

Interviewers rarely ask only theoretical questions. Instead, they present real production incidents and expect you to explain:

  • What happened?
  • How would you investigate?
  • Which tools would you use?
  • What is your troubleshooting methodology?
  • How would you permanently fix it?

Large organizations such as Amazon, Google, Netflix, Microsoft, IBM, Adobe, Capital One, and JPMorgan Chase evaluate engineers primarily through production scenarios.

This guide covers the most common real-world production problems encountered in enterprise environments.


Learning Objectives

After completing this guide, you'll understand

  • Production Troubleshooting Methodology
  • Incident Handling
  • Root Cause Analysis
  • Kubernetes Issues
  • Docker Issues
  • Linux Issues
  • AWS Issues
  • Database Issues
  • CI/CD Issues
  • Monitoring Issues
  • Logging Issues
  • Networking Issues
  • Security Issues
  • Disaster Recovery
  • Enterprise Best Practices

Enterprise Production Environment

flowchart LR

Users --> CDN

CDN --> LoadBalancer

LoadBalancer --> Kubernetes

Kubernetes --> Microservices

Microservices --> Kafka

Microservices --> Redis

Microservices --> Database

Microservices --> Prometheus

Prometheus --> Grafana

Microservices --> Loki

Microservices --> Jaeger

Alertmanager --> PagerDuty

Production Troubleshooting Methodology

Always follow a structured approach.

flowchart LR

Alert --> Identify --> Analyze --> Mitigate --> Fix --> Verify --> RCA

Never jump directly to conclusions.


Scenario 1

Application is Down

Symptoms

  • HTTP 503
  • Health Check Failed
  • Customers Cannot Login

Investigation

  • Application Logs
  • Pod Status
  • Load Balancer
  • Database Connectivity
  • CPU
  • Memory

Possible Causes

  • Crash
  • Database Down
  • Configuration Issue
  • OOMKilled

Scenario 2

Kubernetes Pod CrashLoopBackOff

Commands

kubectl get pods

kubectl describe pod

kubectl logs pod-name

Possible Causes

  • Wrong Configuration
  • Missing Secret
  • Application Crash
  • Database Connection Failure

Scenario 3

Pod OOMKilled

Symptoms

OOMKilled

Investigate

kubectl describe pod

Solutions

  • Increase Memory Limit
  • Optimize Application
  • Analyze Heap Dump
  • Tune JVM

Scenario 4

Node Not Ready

Check

kubectl get nodes

Investigate

  • Disk Pressure
  • Memory Pressure
  • Network
  • Kubelet
  • Container Runtime

Scenario 5

High CPU Usage

Check

top

htop

Investigate

  • Infinite Loop
  • High Traffic
  • Bad SQL
  • JVM Threads

Monitor

  • Grafana
  • Prometheus

Scenario 6

Memory Leak

Symptoms

  • Memory continuously increases
  • Frequent GC
  • Slow Response

Tools

  • Heap Dump
  • VisualVM
  • Eclipse MAT
  • JProfiler

Scenario 7

Disk Full

Commands

df -h

du -sh *

find / -size +1G

Solutions

  • Log Rotation
  • Cleanup
  • Archive Logs

Scenario 8

Jenkins Build Failed

Check

  • Console Logs
  • Credentials
  • Plugins
  • Build Agent
  • Disk Space

Common Issues

  • Maven Failure
  • Docker Build Failure
  • Network Timeout

Scenario 9

Docker Container Restart Loop

Commands

docker ps

docker logs container

docker inspect container

Possible Causes

  • Wrong Environment Variables
  • Missing Dependency
  • Health Check Failure

Scenario 10

Kubernetes Deployment Failed

Check

kubectl rollout status deployment

Investigate

  • Events
  • Logs
  • Image Pull
  • Secrets
  • ConfigMaps

Scenario 11

ImagePullBackOff

Possible Causes

  • Wrong Image
  • Authentication Failure
  • Registry Down

Commands

kubectl describe pod

Scenario 12

API Suddenly Slow

Check

  • CPU
  • Memory
  • Database
  • External APIs
  • Network
  • JVM Threads

Tools

  • Grafana
  • Jaeger
  • Prometheus

Scenario 13

Database Connection Pool Exhausted

Symptoms

Cannot obtain connection

Check

  • HikariCP
  • Active Connections
  • Slow Queries
  • Locks

Solutions

  • Tune Pool
  • Optimize Queries
  • Increase Connections

Scenario 14

Slow SQL Query

Investigate

EXPLAIN ANALYZE

Check

  • Missing Index
  • Full Table Scan
  • Locks
  • Large Result Sets

Scenario 15

Deadlock

Symptoms

Deadlock detected

Solutions

  • Retry Logic
  • Smaller Transactions
  • Consistent Lock Order

Scenario 16

Kafka Consumer Lag

Check

kafka-consumer-groups

Possible Causes

  • Slow Consumers
  • High Traffic
  • Processing Delay

Scenario 17

RabbitMQ Queue Growth

Investigate

  • Consumer Down
  • Slow Processing
  • Network Delay

Monitor

  • Queue Length
  • Message Rate

Scenario 18

SSL Certificate Expired

Symptoms

HTTPS Failed

Check

openssl s_client

Solution

  • Renew Certificate
  • Restart Services

Scenario 19

DNS Resolution Failure

Commands

nslookup

dig

ping

Check

  • DNS Server
  • Network
  • Firewall

Scenario 20

AWS EC2 Instance Unreachable

Investigate

  • Security Groups
  • Route Tables
  • Instance Status
  • Disk Usage
  • CPU

Scenario 21

RDS Failover

Check

  • Replica Health
  • DNS Update
  • Application Retry Logic

Scenario 22

EKS Worker Node Failure

Check

kubectl get nodes

Investigate

  • Auto Scaling Group
  • EC2 Health
  • IAM
  • Networking

Scenario 23

S3 Access Denied

Possible Causes

  • IAM
  • Bucket Policy
  • KMS
  • ACL

Scenario 24

IAM Permission Denied

Check

  • IAM Role
  • Trust Policy
  • Resource Policy
  • SCP

Scenario 25

High Error Rate

Investigate

  • Logs
  • Metrics
  • Traces
  • Database
  • Deployment History

Scenario 26

Deployment Increased Latency

Check

  • New Release
  • Database Queries
  • JVM
  • Cache
  • External APIs

Rollback if required.


Scenario 27

Prometheus Alert Fired

Workflow

Alert

↓

Grafana Dashboard

↓

Logs

↓

Tracing

↓

Root Cause

Scenario 28

Grafana Dashboard Shows CPU Spike

Investigate

  • Recent Deployments
  • Traffic
  • Cron Jobs
  • JVM

Scenario 29

Secrets Exposed in Git

Immediate Actions

  • Rotate Secrets
  • Revoke Tokens
  • Remove History
  • Audit Access

Scenario 30

Security Vulnerability Found

Workflow

Identify CVE

↓

Prioritize

↓

Patch

↓

Deploy

↓

Verify

Scenario 31

Complete Region Failure

Disaster Recovery

flowchart LR

PrimaryRegion --> SecondaryRegion --> TrafficShift --> Recovered

Scenario 32

Load Balancer Health Check Failed

Check

  • Health Endpoint
  • Network
  • Firewall
  • Security Groups
  • Application

Scenario 33

Redis Cache Down

Impact

  • Slow APIs
  • Database Load Increased

Solutions

  • Failover
  • Restart
  • Rebuild Cache

Scenario 34

Batch Job Failed

Investigate

  • Input File
  • Database
  • Scheduler
  • Application Logs

Recovery

  • Restart Failed Batch
  • Resume Processing

Scenario 35

Production Deployment Rollback

Reasons

  • High Error Rate
  • Latency
  • Memory Leak
  • Failed Health Checks

Strategies

  • Blue-Green
  • Canary
  • Rolling Rollback

Production Incident Checklist

Before making changes

✅ Confirm customer impact

✅ Review dashboards

✅ Check logs

✅ Check traces

✅ Review recent deployments

✅ Validate infrastructure

✅ Inform stakeholders


Enterprise RCA Template

Include

  • Timeline
  • Symptoms
  • Root Cause
  • Customer Impact
  • Resolution
  • Preventive Actions
  • Lessons Learned

Production Best Practices

  • Never SSH directly without approval
  • Always collect evidence before restarting services
  • Use dashboards before logs
  • Automate repetitive recovery tasks
  • Keep runbooks updated
  • Test disaster recovery regularly
  • Practice blameless postmortems
  • Monitor business KPIs along with infrastructure
  • Validate backups frequently
  • Perform capacity planning proactively

Common Enterprise Tools

Category Tools
Monitoring Prometheus, Datadog, CloudWatch
Dashboards Grafana
Logging ELK, Loki, Splunk
Tracing Jaeger, Zipkin
Kubernetes kubectl, Lens
Docker Docker CLI
Cloud AWS Console, Azure Portal
Incident Management PagerDuty, ServiceNow
CI/CD Jenkins, GitHub Actions
Load Testing JMeter, k6

Interview Tips

During scenario-based interviews

  • Stay calm and explain your thought process.
  • Start with customer impact.
  • Verify monitoring dashboards before making assumptions.
  • Use logs and traces to narrow the problem.
  • Mention rollback options if a deployment is involved.
  • Discuss both immediate mitigation and permanent fixes.
  • Always conclude with Root Cause Analysis (RCA) and preventive actions.

Interviewers often value a structured troubleshooting methodology more than immediately identifying the correct root cause.


Summary

Real-world production troubleshooting combines technical knowledge with operational discipline. Successful engineers investigate issues methodically using metrics, logs, traces, infrastructure health, deployment history, and application behavior before taking corrective action.

Mastering production scenarios involving Kubernetes, Docker, Linux, AWS, databases, CI/CD pipelines, networking, monitoring, logging, security, and disaster recovery prepares you for senior DevOps Engineer, Site Reliability Engineer (SRE), Platform Engineer, Cloud Engineer, Technical Lead, and Solution Architect interviews.

In the next chapter, you'll learn System Design and Troubleshooting, where you'll design enterprise-grade DevOps platforms and solve complex architecture and scalability challenges.