Real World Production Scenarios
Master real-world production troubleshooting scenarios covering Kubernetes, Docker, AWS, CI/CD, Linux, Databases, Monitoring, Logging, DevSecOps, SRE, and enterprise production incident handling.
Introduction
One of the biggest differences between a Junior DevOps Engineer and a Senior DevOps Engineer is the ability to solve production problems under pressure.
Interviewers rarely ask only theoretical questions. Instead, they present real production incidents and expect you to explain:
- What happened?
- How would you investigate?
- Which tools would you use?
- What is your troubleshooting methodology?
- How would you permanently fix it?
Large organizations such as Amazon, Google, Netflix, Microsoft, IBM, Adobe, Capital One, and JPMorgan Chase evaluate engineers primarily through production scenarios.
This guide covers the most common real-world production problems encountered in enterprise environments.
Learning Objectives
After completing this guide, you'll understand
- Production Troubleshooting Methodology
- Incident Handling
- Root Cause Analysis
- Kubernetes Issues
- Docker Issues
- Linux Issues
- AWS Issues
- Database Issues
- CI/CD Issues
- Monitoring Issues
- Logging Issues
- Networking Issues
- Security Issues
- Disaster Recovery
- Enterprise Best Practices
Enterprise Production Environment
flowchart LR
Users --> CDN
CDN --> LoadBalancer
LoadBalancer --> Kubernetes
Kubernetes --> Microservices
Microservices --> Kafka
Microservices --> Redis
Microservices --> Database
Microservices --> Prometheus
Prometheus --> Grafana
Microservices --> Loki
Microservices --> Jaeger
Alertmanager --> PagerDuty
Production Troubleshooting Methodology
Always follow a structured approach.
flowchart LR
Alert --> Identify --> Analyze --> Mitigate --> Fix --> Verify --> RCA
Never jump directly to conclusions.
Scenario 1
Application is Down
Symptoms
- HTTP 503
- Health Check Failed
- Customers Cannot Login
Investigation
- Application Logs
- Pod Status
- Load Balancer
- Database Connectivity
- CPU
- Memory
Possible Causes
- Crash
- Database Down
- Configuration Issue
- OOMKilled
Scenario 2
Kubernetes Pod CrashLoopBackOff
Commands
kubectl get pods
kubectl describe pod
kubectl logs pod-name
Possible Causes
- Wrong Configuration
- Missing Secret
- Application Crash
- Database Connection Failure
Scenario 3
Pod OOMKilled
Symptoms
OOMKilled
Investigate
kubectl describe pod
Solutions
- Increase Memory Limit
- Optimize Application
- Analyze Heap Dump
- Tune JVM
Scenario 4
Node Not Ready
Check
kubectl get nodes
Investigate
- Disk Pressure
- Memory Pressure
- Network
- Kubelet
- Container Runtime
Scenario 5
High CPU Usage
Check
top
htop
Investigate
- Infinite Loop
- High Traffic
- Bad SQL
- JVM Threads
Monitor
- Grafana
- Prometheus
Scenario 6
Memory Leak
Symptoms
- Memory continuously increases
- Frequent GC
- Slow Response
Tools
- Heap Dump
- VisualVM
- Eclipse MAT
- JProfiler
Scenario 7
Disk Full
Commands
df -h
du -sh *
find / -size +1G
Solutions
- Log Rotation
- Cleanup
- Archive Logs
Scenario 8
Jenkins Build Failed
Check
- Console Logs
- Credentials
- Plugins
- Build Agent
- Disk Space
Common Issues
- Maven Failure
- Docker Build Failure
- Network Timeout
Scenario 9
Docker Container Restart Loop
Commands
docker ps
docker logs container
docker inspect container
Possible Causes
- Wrong Environment Variables
- Missing Dependency
- Health Check Failure
Scenario 10
Kubernetes Deployment Failed
Check
kubectl rollout status deployment
Investigate
- Events
- Logs
- Image Pull
- Secrets
- ConfigMaps
Scenario 11
ImagePullBackOff
Possible Causes
- Wrong Image
- Authentication Failure
- Registry Down
Commands
kubectl describe pod
Scenario 12
API Suddenly Slow
Check
- CPU
- Memory
- Database
- External APIs
- Network
- JVM Threads
Tools
- Grafana
- Jaeger
- Prometheus
Scenario 13
Database Connection Pool Exhausted
Symptoms
Cannot obtain connection
Check
- HikariCP
- Active Connections
- Slow Queries
- Locks
Solutions
- Tune Pool
- Optimize Queries
- Increase Connections
Scenario 14
Slow SQL Query
Investigate
EXPLAIN ANALYZE
Check
- Missing Index
- Full Table Scan
- Locks
- Large Result Sets
Scenario 15
Deadlock
Symptoms
Deadlock detected
Solutions
- Retry Logic
- Smaller Transactions
- Consistent Lock Order
Scenario 16
Kafka Consumer Lag
Check
kafka-consumer-groups
Possible Causes
- Slow Consumers
- High Traffic
- Processing Delay
Scenario 17
RabbitMQ Queue Growth
Investigate
- Consumer Down
- Slow Processing
- Network Delay
Monitor
- Queue Length
- Message Rate
Scenario 18
SSL Certificate Expired
Symptoms
HTTPS Failed
Check
openssl s_client
Solution
- Renew Certificate
- Restart Services
Scenario 19
DNS Resolution Failure
Commands
nslookup
dig
ping
Check
- DNS Server
- Network
- Firewall
Scenario 20
AWS EC2 Instance Unreachable
Investigate
- Security Groups
- Route Tables
- Instance Status
- Disk Usage
- CPU
Scenario 21
RDS Failover
Check
- Replica Health
- DNS Update
- Application Retry Logic
Scenario 22
EKS Worker Node Failure
Check
kubectl get nodes
Investigate
- Auto Scaling Group
- EC2 Health
- IAM
- Networking
Scenario 23
S3 Access Denied
Possible Causes
- IAM
- Bucket Policy
- KMS
- ACL
Scenario 24
IAM Permission Denied
Check
- IAM Role
- Trust Policy
- Resource Policy
- SCP
Scenario 25
High Error Rate
Investigate
- Logs
- Metrics
- Traces
- Database
- Deployment History
Scenario 26
Deployment Increased Latency
Check
- New Release
- Database Queries
- JVM
- Cache
- External APIs
Rollback if required.
Scenario 27
Prometheus Alert Fired
Workflow
Alert
↓
Grafana Dashboard
↓
Logs
↓
Tracing
↓
Root Cause
Scenario 28
Grafana Dashboard Shows CPU Spike
Investigate
- Recent Deployments
- Traffic
- Cron Jobs
- JVM
Scenario 29
Secrets Exposed in Git
Immediate Actions
- Rotate Secrets
- Revoke Tokens
- Remove History
- Audit Access
Scenario 30
Security Vulnerability Found
Workflow
Identify CVE
↓
Prioritize
↓
Patch
↓
Deploy
↓
Verify
Scenario 31
Complete Region Failure
Disaster Recovery
flowchart LR
PrimaryRegion --> SecondaryRegion --> TrafficShift --> Recovered
Scenario 32
Load Balancer Health Check Failed
Check
- Health Endpoint
- Network
- Firewall
- Security Groups
- Application
Scenario 33
Redis Cache Down
Impact
- Slow APIs
- Database Load Increased
Solutions
- Failover
- Restart
- Rebuild Cache
Scenario 34
Batch Job Failed
Investigate
- Input File
- Database
- Scheduler
- Application Logs
Recovery
- Restart Failed Batch
- Resume Processing
Scenario 35
Production Deployment Rollback
Reasons
- High Error Rate
- Latency
- Memory Leak
- Failed Health Checks
Strategies
- Blue-Green
- Canary
- Rolling Rollback
Production Incident Checklist
Before making changes
✅ Confirm customer impact
✅ Review dashboards
✅ Check logs
✅ Check traces
✅ Review recent deployments
✅ Validate infrastructure
✅ Inform stakeholders
Enterprise RCA Template
Include
- Timeline
- Symptoms
- Root Cause
- Customer Impact
- Resolution
- Preventive Actions
- Lessons Learned
Production Best Practices
- Never SSH directly without approval
- Always collect evidence before restarting services
- Use dashboards before logs
- Automate repetitive recovery tasks
- Keep runbooks updated
- Test disaster recovery regularly
- Practice blameless postmortems
- Monitor business KPIs along with infrastructure
- Validate backups frequently
- Perform capacity planning proactively
Common Enterprise Tools
| Category | Tools |
|---|---|
| Monitoring | Prometheus, Datadog, CloudWatch |
| Dashboards | Grafana |
| Logging | ELK, Loki, Splunk |
| Tracing | Jaeger, Zipkin |
| Kubernetes | kubectl, Lens |
| Docker | Docker CLI |
| Cloud | AWS Console, Azure Portal |
| Incident Management | PagerDuty, ServiceNow |
| CI/CD | Jenkins, GitHub Actions |
| Load Testing | JMeter, k6 |
Interview Tips
During scenario-based interviews
- Stay calm and explain your thought process.
- Start with customer impact.
- Verify monitoring dashboards before making assumptions.
- Use logs and traces to narrow the problem.
- Mention rollback options if a deployment is involved.
- Discuss both immediate mitigation and permanent fixes.
- Always conclude with Root Cause Analysis (RCA) and preventive actions.
Interviewers often value a structured troubleshooting methodology more than immediately identifying the correct root cause.
Summary
Real-world production troubleshooting combines technical knowledge with operational discipline. Successful engineers investigate issues methodically using metrics, logs, traces, infrastructure health, deployment history, and application behavior before taking corrective action.
Mastering production scenarios involving Kubernetes, Docker, Linux, AWS, databases, CI/CD pipelines, networking, monitoring, logging, security, and disaster recovery prepares you for senior DevOps Engineer, Site Reliability Engineer (SRE), Platform Engineer, Cloud Engineer, Technical Lead, and Solution Architect interviews.
In the next chapter, you'll learn System Design and Troubleshooting, where you'll design enterprise-grade DevOps platforms and solve complex architecture and scalability challenges.