Top 100 DevOps Interview Questions

A complete collection of the Top 100 DevOps interview questions covering Linux, Networking, Git, Docker, Kubernetes, Jenkins, CI/CD, AWS, Terraform, Ansible, Monitoring, Logging, DevSecOps, Production Support, and System Design.

Introduction

This guide contains the 100 most frequently asked DevOps interview questions collected from enterprise interviews at companies including Amazon, Microsoft, Google, IBM, Red Hat, Adobe, Oracle, JPMorgan Chase, Capital One, Walmart, Cisco, and many Fortune 500 organizations.

Questions are organized by topic and include concise, production-oriented answers suitable for interviews ranging from 3 years to 15+ years of experience.


DevOps Interview Roadmap

flowchart LR

Linux --> Networking --> Git --> Docker --> Kubernetes --> CI_CD --> AWS --> Terraform --> Ansible --> Monitoring --> Logging --> DevSecOps --> Production --> SystemDesign

Linux (1–10)

1. What is Linux and why is it widely used in DevOps?

Linux is an open-source operating system that powers most cloud servers, containers, and Kubernetes clusters because of its stability, security, and performance.


2. Explain the Linux boot process.

BIOS/UEFI → Bootloader (GRUB) → Kernel → Init/Systemd → Services → User Login


  • Hard Link shares inode.
  • Soft Link points to file path.

4. How do you find high CPU processes?

top
htop
ps -ef

5. How do you check disk usage?

df -h
du -sh *

6. How do you find large files?

find / -size +1G

7. Difference between process and thread?

  • Process has separate memory.
  • Threads share process memory.

8. What is swap memory?

Disk space used when RAM becomes full.


9. How do you kill a process?

kill PID
kill -9 PID

10. What are systemd services?

Systemd manages Linux services and system startup.


Networking (11–20)

11. Difference between TCP and UDP?

TCP is reliable; UDP is faster but unreliable.


12. What is DNS?

Translates domain names into IP addresses.


13. What is Load Balancer?

Distributes traffic across multiple servers.


14. What is Reverse Proxy?

Receives client requests and forwards them to backend servers.


15. Difference between HTTP and HTTPS?

HTTPS encrypts communication using TLS.


16. What is SSL/TLS?

Protocols that secure network communication.


17. What is CIDR?

Notation for IP address ranges.


18. What is NAT?

Converts private IPs to public IPs.


19. Difference between Layer 4 and Layer 7 Load Balancer?

  • L4 → TCP/UDP
  • L7 → HTTP/HTTPS

20. How do you troubleshoot network issues?

  • ping
  • traceroute
  • nslookup
  • dig
  • telnet
  • curl

Git (21–28)

21. Difference between Git Fetch and Git Pull?

Fetch downloads changes.

Pull downloads and merges.


22. What is Git Rebase?

Rewrites commit history.


23. What is Git Merge?

Combines branches while preserving history.


24. What is Cherry Pick?

Copies specific commits.


25. How do you resolve merge conflicts?

Edit files → git add → commit.


26. What is Git Tag?

Marks release versions.


27. What is Git Stash?

Temporarily saves changes.


28. What is Git Reset?

Moves HEAD to previous commit.


Docker (29–38)

29. What is Docker?

Container platform.


30. Difference between Image and Container?

Image is template.

Container is running instance.


31. What is Docker Volume?

Persistent storage.


32. Difference between CMD and ENTRYPOINT?

CMD provides defaults.

ENTRYPOINT defines executable.


33. What is Docker Compose?

Runs multi-container applications.


34. How do you reduce Docker image size?

  • Multi-stage builds
  • Alpine images
  • Remove unnecessary packages

35. How do you debug containers?

docker logs
docker exec
docker inspect

36. What is Docker Registry?

Stores container images.


37. Difference between COPY and ADD?

COPY copies files.

ADD also supports URLs and archives.


38. Best practices for Docker?

  • Minimal images
  • Non-root user
  • Health checks
  • Scan images

Kubernetes (39–55)

39. What is Kubernetes?

Container orchestration platform.


40. What is a Pod?

Smallest deployable unit.


41. Deployment vs StatefulSet?

Deployment → Stateless

StatefulSet → Stateful


42. What is ReplicaSet?

Maintains desired pod count.


43. What is Service?

Provides stable networking.


44. ClusterIP vs NodePort vs LoadBalancer?

Internal / External Node / Cloud LB.


45. What is Ingress?

HTTP routing into cluster.


46. Liveness vs Readiness Probe?

Alive vs Ready for traffic.


47. ConfigMap vs Secret?

Configuration vs Sensitive Data.


48. What is HPA?

Horizontal Pod Autoscaler.


49. What causes CrashLoopBackOff?

Application repeatedly crashes.


50. What causes ImagePullBackOff?

Image cannot be downloaded.


51. What is OOMKilled?

Container exceeded memory limit.


52. How do you troubleshoot pods?

kubectl logs
kubectl describe
kubectl get events

53. What is RBAC?

Role-Based Access Control.


54. What is Helm?

Kubernetes package manager.


55. What is Operator?

Automates application lifecycle.


Jenkins & CI/CD (56–65)

56. What is CI/CD?

Continuous Integration and Continuous Delivery.


57. What is Jenkins Pipeline?

Pipeline as Code using Jenkinsfile.


58. Declarative vs Scripted Pipeline?

Declarative is simpler.

Scripted is flexible.


59. Pipeline stages?

Checkout → Build → Test → Scan → Deploy


60. Blue-Green Deployment?

Two identical environments.


61. Canary Deployment?

Gradual rollout.


62. Rolling Deployment?

Incremental update.


63. What is GitOps?

Git is source of truth.


64. What is Argo CD?

GitOps deployment tool.


65. Why automate deployments?

Consistency, speed, fewer errors.


AWS & Cloud (66–78)

66. Difference between EC2 and Lambda?

VM vs Serverless.


67. What is VPC?

Virtual Private Cloud.


68. What is Auto Scaling?

Automatically adjusts capacity.


69. What is ALB?

Application Load Balancer.


70. What is IAM?

Identity and Access Management.


71. S3 vs EBS?

Object Storage vs Block Storage.


72. RDS vs Aurora?

Managed DB vs AWS optimized database.


73. What is CloudWatch?

AWS monitoring service.


74. What is EKS?

Managed Kubernetes.


75. ECS vs EKS?

Containers vs Kubernetes.


76. What is Route53?

DNS service.


77. What is CloudFront?

Content Delivery Network.


78. What is Secrets Manager?

Secure secret storage.


Terraform & Ansible (79–86)

79. What is Infrastructure as Code?

Managing infrastructure using code.


80. Terraform vs Ansible?

Provisioning vs Configuration.


81. What is Terraform State?

Tracks infrastructure.


82. What are Terraform Modules?

Reusable infrastructure components.


83. What is Ansible Inventory?

List of managed hosts.


84. Playbook vs Role?

Automation script vs reusable automation.


85. Why is idempotency important?

Safe repeated execution.


86. What is Ansible Vault?

Encrypts sensitive data.


Monitoring & Logging (87–92)

87. Three pillars of Observability?

Metrics, Logs, Traces.


88. Prometheus vs Grafana?

Collection vs Visualization.


89. What is Alertmanager?

Routes alerts.


90. ELK vs Loki?

Full-text indexing vs Label indexing.


91. Why use Correlation IDs?

Track requests across services.


92. What is Distributed Tracing?

Tracks requests across microservices.


DevSecOps (93–96)

93. What is Shift Left Security?

Security early in SDLC.


94. SAST vs DAST?

Static vs Runtime security testing.


95. What is SBOM?

Software Bill of Materials.


96. What is Zero Trust?

Never Trust, Always Verify.


Production & SRE (97–100)

97. What is High Availability?

Redundant architecture minimizing downtime.


98. Difference between RTO and RPO?

RTO = Downtime

RPO = Data Loss


99. What is RCA?

Root Cause Analysis after incidents.


100. A production API suddenly becomes slow. How would you troubleshoot?

Approach

  1. Check customer impact.
  2. Review Grafana dashboards.
  3. Verify CPU, Memory, Network.
  4. Analyze Prometheus metrics.
  5. Search logs using Correlation ID.
  6. Review Jaeger traces.
  7. Check database queries.
  8. Verify recent deployments.
  9. Rollback if necessary.
  10. Perform RCA and implement preventive measures.

Enterprise DevOps Interview Flow

flowchart LR

Requirements --> Architecture --> Implementation --> Deployment --> Monitoring --> Incident --> RCA --> Automation --> Optimization

Last-Minute Interview Checklist

✅ Linux Commands

✅ Networking Basics

✅ Git Workflow

✅ Docker Commands

✅ Kubernetes Troubleshooting

✅ CI/CD Pipelines

✅ AWS Core Services

✅ Terraform & Ansible

✅ Monitoring & Logging

✅ DevSecOps

✅ Production Support

✅ System Design

✅ Troubleshooting Methodology

✅ Root Cause Analysis


Final Interview Tips

For senior DevOps interviews:

  • Explain your thinking before giving the solution.
  • Discuss trade-offs, not just tools.
  • Mention security, scalability, reliability, and cost optimization.
  • Use real production examples whenever possible.
  • Demonstrate a structured troubleshooting process.
  • Highlight automation and observability.
  • Always include rollback and disaster recovery strategies.

Summary

These 100 interview questions cover the essential knowledge expected from modern DevOps professionals, ranging from Linux and Kubernetes to AWS, CI/CD, Infrastructure as Code, Monitoring, DevSecOps, Production Operations, and System Design.

While memorizing answers is helpful, interview success comes from understanding why technologies are used, how they interact in production, and how to troubleshoot real-world issues systematically. Strong practical experience, combined with clear communication and structured problem-solving, is what distinguishes senior DevOps engineers from the rest.