Monitoring Fundamentals

Learn Monitoring and Observability fundamentals including Metrics, Logs, Traces, Health Checks, SLI, SLO, SLA, Alerting, Dashboards, Prometheus, Grafana, Datadog, Splunk, CloudWatch, and production monitoring best practices.

Introduction

Modern distributed systems consist of microservices, containers, Kubernetes clusters, databases, cloud services, APIs, and messaging systems. When something fails, engineers need to quickly identify what failed, why it failed, where it failed, and how to recover.

Monitoring provides visibility into the health and performance of systems, while Observability enables engineers to understand internal system behavior using telemetry data.

Today, almost every enterprise uses monitoring platforms such as Prometheus, Grafana, Datadog, Splunk, ELK Stack, Dynatrace, New Relic, Amazon CloudWatch, Azure Monitor, and OpenTelemetry.

This guide introduces the fundamental monitoring concepts required for DevOps Engineers, SREs, Cloud Engineers, Platform Engineers, Java Developers, and Solution Architects.


Learning Objectives

After completing this guide, you'll understand

  • What is Monitoring?
  • What is Observability?
  • Why Monitoring Matters
  • Types of Monitoring
  • Metrics
  • Logs
  • Traces
  • Health Checks
  • Dashboards
  • Alerting
  • SLI
  • SLO
  • SLA
  • Monitoring Architecture
  • Common Monitoring Tools
  • Production Best Practices

What is Monitoring?

Monitoring is the continuous process of collecting and analyzing system information to determine whether applications and infrastructure are operating correctly.

Monitoring helps answer:

  • Is the application running?
  • Is the server healthy?
  • Is CPU usage high?
  • Are APIs responding?
  • Are users experiencing failures?

Why Monitoring?

Without Monitoring

Application Failure

↓

Users Report Issue

↓

Investigation

↓

Root Cause Found

Problems

  • Late Detection
  • Long Downtime
  • Customer Impact
  • Revenue Loss

With Monitoring

flowchart LR

Application --> Monitoring --> Alert --> Engineer --> FixIssue

Benefits

  • Early Detection
  • Faster Recovery
  • Better Reliability
  • Higher Availability

What is Observability?

Observability is the ability to understand the internal state of a system by analyzing telemetry data.

Observability consists of three pillars:

  • Metrics
  • Logs
  • Traces

Together they help answer:

  • What happened?
  • Why did it happen?
  • Where did it happen?

Monitoring vs Observability

Monitoring Observability
Detect Problems Understand Problems
Predefined Metrics Deep Investigation
Dashboards Root Cause Analysis
Alerts Distributed Tracing

Monitoring tells you something is wrong.

Observability tells you why it's wrong.


Three Pillars of Observability

flowchart TD

Observability --> Metrics

Observability --> Logs

Observability --> Traces

Metrics

Metrics are numerical measurements collected over time.

Examples

  • CPU Usage
  • Memory Usage
  • Disk Usage
  • Request Count
  • Response Time
  • Error Rate
  • Network Traffic

Metrics are lightweight and ideal for dashboards and alerting.


Logs

Logs record events occurring inside applications.

Examples

Application Started

User Login

Payment Successful

Database Timeout

NullPointerException

Logs provide detailed troubleshooting information.


Traces

Traces follow a request across multiple services.

Example

flowchart LR

Client --> API --> PaymentService --> Database --> Response

Useful for

  • Microservices
  • Distributed Systems
  • Performance Analysis

Types of Monitoring

Infrastructure Monitoring

Monitor

  • CPU
  • Memory
  • Disk
  • Network
  • Servers

Application Monitoring

Monitor

  • Response Time
  • Error Rate
  • Throughput
  • Availability

Database Monitoring

Monitor

  • Queries
  • Connections
  • Locks
  • Replication
  • Slow Queries

Network Monitoring

Monitor

  • Packet Loss
  • Bandwidth
  • Latency
  • Firewalls

Container Monitoring

Monitor

  • Pods
  • Containers
  • Nodes
  • Images

Cloud Monitoring

Monitor

  • EC2
  • S3
  • Lambda
  • RDS
  • Kubernetes

Health Checks

Applications expose health endpoints.

Example

/actuator/health

Common checks

  • Database Connectivity
  • Message Queue
  • Disk Space
  • External APIs
  • Cache Availability

Liveness Probe

Checks

Is the application alive?

Failure

Restart Container


Readiness Probe

Checks

Can the application receive traffic?

Failure

Remove from Load Balancer


Startup Probe

Checks

Has the application finished starting?

Useful for

  • Spring Boot
  • Large Applications
  • Slow Startup Services

Dashboards

Dashboards visualize system health.

Typical Widgets

  • CPU Usage
  • Memory Usage
  • Requests Per Second
  • Error Rate
  • Latency
  • Active Users

Popular Dashboard Tools

  • Grafana
  • Datadog
  • Kibana
  • CloudWatch Dashboards

Alerting

Alerts notify engineers when thresholds are exceeded.

Example

CPU > 90%

Alert

Engineer

Fix Issue

Alert Channels

  • Email
  • Slack
  • Microsoft Teams
  • PagerDuty
  • SMS

Monitoring Architecture

flowchart LR

Application --> Metrics

Application --> Logs

Application --> Traces

Metrics --> Prometheus

Logs --> Splunk

Traces --> Jaeger

Prometheus --> Grafana

SLI

Service Level Indicator

Measures service performance.

Examples

  • Availability
  • Latency
  • Error Rate

SLO

Service Level Objective

Target performance level.

Example

99.9% uptime


SLA

Service Level Agreement

Formal agreement between provider and customer.

Example

99.95% Availability

Financial penalties may apply if violated.


Common Monitoring Tools

Category Tools
Metrics Prometheus, CloudWatch
Dashboards Grafana, Datadog
Logs Splunk, ELK, Loki
Tracing Jaeger, Zipkin, OpenTelemetry
APM Dynatrace, New Relic, AppDynamics
Cloud CloudWatch, Azure Monitor

Monitoring Workflow

flowchart LR

Application --> CollectMetrics --> StoreMetrics --> Dashboard --> Alert --> Engineer

Enterprise Monitoring Stack

flowchart LR

Application --> OpenTelemetry

OpenTelemetry --> Prometheus

OpenTelemetry --> Loki

OpenTelemetry --> Jaeger

Prometheus --> Grafana

Grafana --> Alertmanager

Benefits of Monitoring

  • Detect Problems Early
  • Reduce Downtime
  • Improve Performance
  • Better User Experience
  • Capacity Planning
  • Faster Incident Response
  • Business Visibility

Production Best Practices

Metrics

  • Monitor Business Metrics
  • Monitor Infrastructure
  • Monitor Applications
  • Track Latency

Logs

  • Structured Logging
  • Centralized Logging
  • Retention Policy
  • Log Correlation

Traces

  • Distributed Tracing
  • Request Correlation
  • End-to-End Visibility

Alerts

  • Meaningful Alerts
  • Avoid Alert Fatigue
  • Severity Levels
  • Escalation Policies

Dashboards

  • Executive Dashboard
  • Application Dashboard
  • Infrastructure Dashboard
  • Database Dashboard

Real-World Example

A Spring Boot application runs on Kubernetes.

  1. Spring Boot exposes Micrometer metrics.
  2. Prometheus scrapes metrics every 15 seconds.
  3. Grafana visualizes dashboards.
  4. Loki collects application logs.
  5. OpenTelemetry generates distributed traces.
  6. Alertmanager sends Slack notifications when response time exceeds 2 seconds.
  7. Engineers use dashboards, logs, and traces to identify a slow database query causing high latency.
  8. The issue is resolved before users notice a major outage.

Interview Tips

Remember these keywords

  • Monitoring
  • Observability
  • Metrics
  • Logs
  • Traces
  • Dashboard
  • Alerting
  • Health Check
  • Prometheus
  • Grafana
  • Splunk
  • Datadog
  • OpenTelemetry
  • SLI
  • SLO
  • SLA

Summary

Monitoring provides continuous visibility into application and infrastructure health, while Observability enables engineers to investigate and understand complex system behavior. Modern monitoring platforms combine metrics, logs, traces, dashboards, and intelligent alerting to help teams maintain highly available, reliable, and scalable systems.

Mastering monitoring fundamentals—including the three pillars of observability, health checks, dashboards, alerting, and service reliability metrics—forms the foundation for enterprise DevOps, SRE, Platform Engineering, Cloud Engineering, and Solution Architect roles.

In the next chapter, you'll explore Monitoring Advanced, covering Prometheus architecture, Grafana dashboards, Alertmanager, OpenTelemetry, distributed tracing, Kubernetes monitoring, log aggregation, APM, cloud-native observability, and enterprise monitoring architectures.