Monitoring Fundamentals
Learn Monitoring and Observability fundamentals including Metrics, Logs, Traces, Health Checks, SLI, SLO, SLA, Alerting, Dashboards, Prometheus, Grafana, Datadog, Splunk, CloudWatch, and production monitoring best practices.
Introduction
Modern distributed systems consist of microservices, containers, Kubernetes clusters, databases, cloud services, APIs, and messaging systems. When something fails, engineers need to quickly identify what failed, why it failed, where it failed, and how to recover.
Monitoring provides visibility into the health and performance of systems, while Observability enables engineers to understand internal system behavior using telemetry data.
Today, almost every enterprise uses monitoring platforms such as Prometheus, Grafana, Datadog, Splunk, ELK Stack, Dynatrace, New Relic, Amazon CloudWatch, Azure Monitor, and OpenTelemetry.
This guide introduces the fundamental monitoring concepts required for DevOps Engineers, SREs, Cloud Engineers, Platform Engineers, Java Developers, and Solution Architects.
Learning Objectives
After completing this guide, you'll understand
- What is Monitoring?
- What is Observability?
- Why Monitoring Matters
- Types of Monitoring
- Metrics
- Logs
- Traces
- Health Checks
- Dashboards
- Alerting
- SLI
- SLO
- SLA
- Monitoring Architecture
- Common Monitoring Tools
- Production Best Practices
What is Monitoring?
Monitoring is the continuous process of collecting and analyzing system information to determine whether applications and infrastructure are operating correctly.
Monitoring helps answer:
- Is the application running?
- Is the server healthy?
- Is CPU usage high?
- Are APIs responding?
- Are users experiencing failures?
Why Monitoring?
Without Monitoring
Application Failure
↓
Users Report Issue
↓
Investigation
↓
Root Cause Found
Problems
- Late Detection
- Long Downtime
- Customer Impact
- Revenue Loss
With Monitoring
flowchart LR
Application --> Monitoring --> Alert --> Engineer --> FixIssue
Benefits
- Early Detection
- Faster Recovery
- Better Reliability
- Higher Availability
What is Observability?
Observability is the ability to understand the internal state of a system by analyzing telemetry data.
Observability consists of three pillars:
- Metrics
- Logs
- Traces
Together they help answer:
- What happened?
- Why did it happen?
- Where did it happen?
Monitoring vs Observability
| Monitoring | Observability |
|---|---|
| Detect Problems | Understand Problems |
| Predefined Metrics | Deep Investigation |
| Dashboards | Root Cause Analysis |
| Alerts | Distributed Tracing |
Monitoring tells you something is wrong.
Observability tells you why it's wrong.
Three Pillars of Observability
flowchart TD
Observability --> Metrics
Observability --> Logs
Observability --> Traces
Metrics
Metrics are numerical measurements collected over time.
Examples
- CPU Usage
- Memory Usage
- Disk Usage
- Request Count
- Response Time
- Error Rate
- Network Traffic
Metrics are lightweight and ideal for dashboards and alerting.
Logs
Logs record events occurring inside applications.
Examples
Application Started
User Login
Payment Successful
Database Timeout
NullPointerException
Logs provide detailed troubleshooting information.
Traces
Traces follow a request across multiple services.
Example
flowchart LR
Client --> API --> PaymentService --> Database --> Response
Useful for
- Microservices
- Distributed Systems
- Performance Analysis
Types of Monitoring
Infrastructure Monitoring
Monitor
- CPU
- Memory
- Disk
- Network
- Servers
Application Monitoring
Monitor
- Response Time
- Error Rate
- Throughput
- Availability
Database Monitoring
Monitor
- Queries
- Connections
- Locks
- Replication
- Slow Queries
Network Monitoring
Monitor
- Packet Loss
- Bandwidth
- Latency
- Firewalls
Container Monitoring
Monitor
- Pods
- Containers
- Nodes
- Images
Cloud Monitoring
Monitor
- EC2
- S3
- Lambda
- RDS
- Kubernetes
Health Checks
Applications expose health endpoints.
Example
/actuator/health
Common checks
- Database Connectivity
- Message Queue
- Disk Space
- External APIs
- Cache Availability
Liveness Probe
Checks
Is the application alive?
Failure
↓
Restart Container
Readiness Probe
Checks
Can the application receive traffic?
Failure
↓
Remove from Load Balancer
Startup Probe
Checks
Has the application finished starting?
Useful for
- Spring Boot
- Large Applications
- Slow Startup Services
Dashboards
Dashboards visualize system health.
Typical Widgets
- CPU Usage
- Memory Usage
- Requests Per Second
- Error Rate
- Latency
- Active Users
Popular Dashboard Tools
- Grafana
- Datadog
- Kibana
- CloudWatch Dashboards
Alerting
Alerts notify engineers when thresholds are exceeded.
Example
CPU > 90%
↓
Alert
↓
Engineer
↓
Fix Issue
Alert Channels
- Slack
- Microsoft Teams
- PagerDuty
- SMS
Monitoring Architecture
flowchart LR
Application --> Metrics
Application --> Logs
Application --> Traces
Metrics --> Prometheus
Logs --> Splunk
Traces --> Jaeger
Prometheus --> Grafana
SLI
Service Level Indicator
Measures service performance.
Examples
- Availability
- Latency
- Error Rate
SLO
Service Level Objective
Target performance level.
Example
99.9% uptime
SLA
Service Level Agreement
Formal agreement between provider and customer.
Example
99.95% Availability
Financial penalties may apply if violated.
Common Monitoring Tools
| Category | Tools |
|---|---|
| Metrics | Prometheus, CloudWatch |
| Dashboards | Grafana, Datadog |
| Logs | Splunk, ELK, Loki |
| Tracing | Jaeger, Zipkin, OpenTelemetry |
| APM | Dynatrace, New Relic, AppDynamics |
| Cloud | CloudWatch, Azure Monitor |
Monitoring Workflow
flowchart LR
Application --> CollectMetrics --> StoreMetrics --> Dashboard --> Alert --> Engineer
Enterprise Monitoring Stack
flowchart LR
Application --> OpenTelemetry
OpenTelemetry --> Prometheus
OpenTelemetry --> Loki
OpenTelemetry --> Jaeger
Prometheus --> Grafana
Grafana --> Alertmanager
Benefits of Monitoring
- Detect Problems Early
- Reduce Downtime
- Improve Performance
- Better User Experience
- Capacity Planning
- Faster Incident Response
- Business Visibility
Production Best Practices
Metrics
- Monitor Business Metrics
- Monitor Infrastructure
- Monitor Applications
- Track Latency
Logs
- Structured Logging
- Centralized Logging
- Retention Policy
- Log Correlation
Traces
- Distributed Tracing
- Request Correlation
- End-to-End Visibility
Alerts
- Meaningful Alerts
- Avoid Alert Fatigue
- Severity Levels
- Escalation Policies
Dashboards
- Executive Dashboard
- Application Dashboard
- Infrastructure Dashboard
- Database Dashboard
Real-World Example
A Spring Boot application runs on Kubernetes.
- Spring Boot exposes Micrometer metrics.
- Prometheus scrapes metrics every 15 seconds.
- Grafana visualizes dashboards.
- Loki collects application logs.
- OpenTelemetry generates distributed traces.
- Alertmanager sends Slack notifications when response time exceeds 2 seconds.
- Engineers use dashboards, logs, and traces to identify a slow database query causing high latency.
- The issue is resolved before users notice a major outage.
Interview Tips
Remember these keywords
- Monitoring
- Observability
- Metrics
- Logs
- Traces
- Dashboard
- Alerting
- Health Check
- Prometheus
- Grafana
- Splunk
- Datadog
- OpenTelemetry
- SLI
- SLO
- SLA
Summary
Monitoring provides continuous visibility into application and infrastructure health, while Observability enables engineers to investigate and understand complex system behavior. Modern monitoring platforms combine metrics, logs, traces, dashboards, and intelligent alerting to help teams maintain highly available, reliable, and scalable systems.
Mastering monitoring fundamentals—including the three pillars of observability, health checks, dashboards, alerting, and service reliability metrics—forms the foundation for enterprise DevOps, SRE, Platform Engineering, Cloud Engineering, and Solution Architect roles.
In the next chapter, you'll explore Monitoring Advanced, covering Prometheus architecture, Grafana dashboards, Alertmanager, OpenTelemetry, distributed tracing, Kubernetes monitoring, log aggregation, APM, cloud-native observability, and enterprise monitoring architectures.