monitoring-ops
Observability patterns - metrics, logging, tracing, alerting, and infrastructure monitoring. Use for: monitoring, observability, prometheus, grafana, metrics, alerting, structured logging, distributed tracing, opentelemetry, SLO, SLI, dashboard, health check, loki, jaeger, datado
Install
npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/monitoring-ops
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
git clone https://github.com/0xDarkMatter/claude-mods.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Monitoring Operations
Comprehensive observability patterns covering the three pillars (metrics, logging, tracing), alerting strategies, dashboard design, and infrastructure monitoring for production systems.
Three Pillars Quick Reference
Use this table to decide which observability signal fits your need:
| Pillar | Best For | Tools | Data Type |
|---|---|---|---|
| Metrics | Aggregated numeric measurements, trends, alerting on thresholds | Prometheus, Datadog, CloudWatch, StatsD | Time-series (numeric) |
| Logs | Discrete events, error details, audit trails, debugging context | Loki, ELK, CloudWatch Logs, Fluentd | Unstructured/structured text |
| Traces | Request flow across services, latency breakdown, dependency mapping | Jaeger, Tempo, Zipkin, Datadog APM | Span trees (structured) |
When to use which:
- "How many requests per second?" → Metrics (counter + rate)
- "Why did this specific request fail?" → Logs (error message + stack trace)
- "Where is the latency in this request?" → Traces (span waterfall)
- "Is the system healthy right now?" → Metrics (gauges + alerts)
- "What happened at 3:42 AM?" → Logs (timestamped event search)
- "Which downstream service caused the timeout?" → Traces (span analysis)
Correlation is key: Connect all three by embedding trace_id in log entries, recording exemplars in metrics, and linking trace spans to log queries.
Metrics Type Decision Tree
Use this tree to select the correct metric type:
What are you measuring?
│
├─ A count of events that only goes up?
│ └─ COUNTER
│ Examples: http_requests_total, errors_total, bytes_sent_total
│ Use rate() or increase() to get per-second or per-interval values
│ Never use a counter's raw value — it resets on restart
│
├─ A current value that goes up AND down?
│ └─ GAUGE
│ Examples: temperature_celsius, active_connections, queue_depth
│ Use for snapshots of current state
│ Can use avg_over_time(), max_over_time() for trends
│
├─ A distribution of values (latency, size)?
│ │
│ ├─ Need aggregatable quantiles across instances?
│ │ └─ HISTOGRAM
│ │ Examples: http_request_duration_seconds, response_size_bytes
│ │ Define buckets: [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10]
│ │ Use histogram_quantile() for percentiles (p50, p95, p99)
│ │ Aggregatable across instances (histograms can be summed)
│ │
│ └─ Need pre-calculated quantiles on a single instance?
│ └─ SUMMARY
│ Examples: go_gc_duration_seconds
│ Pre-calculates quantiles client-side
│ NOT aggregatable across instances
│ Prefer histogram unless you have a specific reason
│
└─ None of the above?
└─ INFO metric (labels only, value=1)
Examples: build_info{version="1.2.3", commit="abc123"}
Use for metadata exposed as metrics
Rule of thumb: Start with counters and histograms. Add gauges for current state. Avoid summaries unless you have a compelling reason.
Alerting Decision Tree
What type of alert do you need?
│
├─ Known threshold with a fixed boundary?
│ └─ THRESHOLD-BASED
│ Example: CPU > 90% for 5 minutes
│ Pros: Simple, predictable, easy to understand
│ Cons: Requires manual tuning, doesn't adapt to patterns
│ Best for: Resource limits, error rate spikes, queue depth
│
├─ Normal behavior varies by time/season?
│ └─ ANOMALY-BASED
│ Example: Traffic 3 standard deviations below normal for this hour
│ Pros: Adapts to patterns, catches novel failures
│ Cons: Noisy during transitions, requires training data
│ Best for: Traffic patterns, business metrics, gradual degradation
│
└─ Defined reliability targets?
└─ SLO-BASED (PREFERRED)
Example: Error budget burn rate > 14.4x for 1 hour
Pros: Aligned with user impact, reduces noise, principled
Cons: Requires SLI/SLO definition, more complex setup
Best for: User-facing services, platform reliability
Severity Levels
| Severity | Response | Examples | Routing |
|---|---|---|---|
| Critical (P1) | Page on-call immediately | Service down, data loss risk, security breach | PagerDuty high-urgency, phone call |
| Warning (P2) | Investigate within hours | Elevated error rate, disk 80% full, SLO burn rate elevated | PagerDuty low-urgency, Slack alert channel |
| Info (P3) | Review next business day | Deployment completed, certificate expiring in 30 days | Slack info channel, ticket auto-created |
When to Page vs When to Ticket
Page (wake someone up) when:
- Users are currently impacted
- Data loss is occurring or imminent
- Security incident is active
- Error budget will be exhausted within hours
Create ticket (don't page) when:
- Issue is not user-facing yet
- Automated remediation is possible
- Degradation is slow and has runway
- Issue is during business hours and can be triaged normally
Structured Logging Quick Reference
Standard JSON Log Format
{
"timestamp": "2026-03-09T14:32:01.123Z",
"level": "ERROR",
"message": "Failed to process payment",
"service": "payment-api",
"version": "1.4.2",
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
"span_id": "00f067aa0ba902b7",
"request_id": "req-abc123",
"user_id": "usr-789",
"error": {
"type": "PaymentGatewayTimeout",
"message": "Gateway response timeout after 30s",
"stack": "..."
},
"duration_ms": 30042,
"http": {
"method": "POST",
"path": "/api/v1/payments",
"status_code": 504
}
}
Log Level Decision Guide
| Level | When to Use | Examples |
|---|---|---|
| DEBUG | Development only, verbose internal state | Variable values, SQL queries, cache hits/misses |
| INFO | Normal operations worth recording | Request completed, job started/finished, config loaded |
| WARN | Degraded but still functioning | Retry succeeded, fallback used, approaching limit |
| ERROR | Operation failed, needs attention | Payment failed, API call error, constraint violation |
| FATAL | Process cannot continue, must exit | Database unreachable at startup, invalid config, OOM |
Rules:
- Never log at ERROR for expected conditions (user input validation → WARN)
- Every ERROR should be actionable — if no one will act on it, use WARN
- DEBUG should be off in production by default
- INFO should not be noisy — 1-5 log lines per request, not 50
Correlation IDs
- Generate a
request_id(UUID v4 or ULID) at the edge/gateway - Propagate through all internal services via headers (
X-Request-ID) - Include
trace_idandspan_idfrom distributed tracing - Log all three IDs in every log entry for cross-referencing
Distributed Tracing Quick Reference
Core Concepts
- Trace: End-to-end journey of a request across all services
- Span: A single unit of work (HTTP call, DB query, function execution)
- Context propagation: Passing trace/span IDs between services via headers
W3C TraceContext Header
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
│ │ │ │
│ │ │ └─ flags (01=sampled)
│ │ └─ parent span ID (16 hex)
│ └─ trace ID (32 hex)
└─ version (00)
Sampling Strategies
| Strategy | How It Works | Use When |
|---|---|---|
| Head-based (ratio) | Decide at trace start, propagate decision | Low traffic, need predictable volume |
| Always-on | Sample everything | Development, low-traffic services |
| Parent-based | Follow parent's sampling decision | Default for most services |
| Tail-based | Decide after trace completes (at Collector) | Need error/slow traces, high traffic |
Recommendation: Use parent-based + tail-based at the Collector. This captures all error traces and slow traces while controlling volume.
Trace ID in Logs
Always include trace_id in structured log entries. This enables jumping from a log line to the full trace view:
Log entry → trace_id → Jaeger/Tempo → full request waterfall
Tool Selection Matrix
| Feature | Prometheus + Grafana | Datadog | Grafana Cloud | CloudWatch |
|---|---|---|---|---|
| Cost | Free (infra costs) | \(\) (per host/metric) | \((usage-based) |\) (AWS-native) | |
| Setup complexity | High (self-managed) | Low (SaaS agent) | Medium (managed) | Low (AWS-native) |
| Metrics | Prometheus (excellent) | Built-in (excellent) | Mimir (excellent) | Built-in (good) |
| Logs | Loki (good) | Built-in (excellent) | Loki (good) | CloudWatch Logs (good) |
| Traces | Jaeger/Tempo (good) | APM (excellent) | Tempo (good) | X-Ray (adequate) |
| Alerting | Alertmanager (good) | Built-in (excellent) | Grafana Alerting (good) | CloudWatch Alarms (adequate) |
| Dashboards | Grafana (excellent) | Built-in (excellent) | Grafana (excellent) | Dashboards (adequate) |
| Retention | Configurable (unlimited) | 15 months default | Configurable | Up to 15 months |
| Multi-cloud | Yes | Yes | Yes | AWS only |
| Best for | Cost-conscious, control | Full-featured, enterprise | Open-source + managed | AWS-native shops |
Recommendation path:
- Starting out / budget-conscious: Prometheus + Grafana + Loki + Tempo (all free, self-hosted)
- Small team, want managed: Grafana Cloud free tier (10k metrics, 50GB logs, 50GB traces)
- Enterprise, need everything: Datadog (expensive but comprehensive)
- AWS-only shop: CloudWatch + X-Ray (simplest if already on AWS)
Dashboard Design
USE Method (Infrastructure)
For every resource (CPU, memory, disk, network):
| Signal | Question | Metric Example |
|---|---|---|
| Utilization | How busy is it? | node_cpu_seconds_total (% busy) |
| Saturation | How overloaded is it? | node_load1 (run queue length) |
| Errors | Are there error events? | node_network_receive_errs_total |
RED Method (Services)
For every service endpoint:
| Signal | Question | Metric Example |
|---|---|---|
| Rate | How many requests per second? | rate(http_requests_total[5m]) |
| Errors | How many are failing? | rate(http_requests_total{status=~"5.."}[5m]) |
| Duration | How long do they take? | histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m])) |
Four Golden Signals (Google SRE)
| Signal | What to Measure | Alert Threshold Guidance |
|---|---|---|
| Latency | Time to serve a request (distinguish success vs error latency) | p99 > 2x baseline |
| Traffic | Demand on the system (requests/sec, sessions, transactions) | Anomaly detection |
| Errors | Rate of failed requests (explicit 5xx, implicit policy violations) | > 0.1% of traffic |
| Saturation | How "full" the service is (CPU, memory, queue depth) | > 80% capacity |
Dashboard Layout Best Practices
- Top row: Key health indicators (error rate, latency p99, availability %)
- Second row: Traffic and throughput (requests/sec, active users)
- Third row: Resource utilization (CPU, memory, disk, network)
- Bottom rows: Detailed breakdowns (by endpoint, by status code, by region)
- Use variables: Service, environment, time range as dropdown selectors
- Include annotations: Deployments, incidents, config changes as vertical markers
Common Gotchas
| Gotcha | Why It Happens | Fix |
|---|---|---|
| Cardinality explosion | Using unbounded label values (user ID, request path, query string) | Use bounded labels only; aggregate high-cardinality data in logs, not metrics |
| Alert fatigue | Too many alerts, too sensitive thresholds, alerts on non-actionable symptoms | Require runbook for every alert; tune thresholds; use SLO-based alerting |
| Missing correlation IDs | Logs, metrics, and traces not linked together | Include trace_id in all log entries; use exemplars in metrics |
| Sampling bias | Head-based sampling drops error/slow traces at high sample rates | Use tail-based sampling at the Collector to always capture errors and slow traces |
| Log volume costs | DEBUG or verbose INFO in production, logging full request/response bodies | Set production to INFO minimum; truncate large payloads; use sampling for verbose paths |
| Metric naming inconsistency | Different teams use different naming conventions | Adopt OpenMetrics naming: namespace_subsystem_unit_suffix (e.g., http_server_request_duration_seconds) |
| Dashboard sprawl | Everyone creates dashboards, nobody maintains them | Standardize with USE/RED templates; review quarterly; delete unused dashboards |
| SLO too aggressive | Setting 99.99% availability without the budget or architecture for it | Start with 99.5% or 99.9%; tighten only when consistently meeting targets with margin |
| Missing baseline | Alerting on absolute thresholds without understanding normal behavior | Collect 2-4 weeks of baseline data before setting alert thresholds |
| Over-instrumentation | Instrumenting every function, creating too many spans/metrics | Instrument at service boundaries; use auto-instrumentation for HTTP/DB/gRPC; add manual spans selectively |
| Ignoring metric staleness | Assuming a metric that stops reporting means zero | Use absent() or up == 0 to detect missing scrapers; distinguish "zero" from "not reporting" |
| Alerting on cause not symptom | Alerting on CPU usage instead of user-facing error rate | Alert on symptoms (error rate, latency); use cause metrics (CPU, memory) for investigation |
| No retention policy | Storing all metrics/logs at full resolution forever | Define retention tiers: 15s resolution for 2 weeks, 1m for 3 months, 5m for 1 year |
| Dashboard without context | Graphs with no units, no description, no threshold lines | Add units to Y-axis, threshold lines for SLOs, panel descriptions explaining what "good" looks like |
Reference Files
| File | Contents | Lines |
|---|---|---|
| metrics-alerting.md | Prometheus, Grafana, OpenTelemetry metrics, SLI/SLO/SLA, alert routing, runbooks, uptime monitoring | ~650 |
| logging.md | Structured logging, log levels, correlation IDs, aggregation (Loki, ELK), retention, PII masking, language-specific | ~550 |
| tracing.md | OpenTelemetry, spans, context propagation, sampling, Jaeger, async tracing, DB/HTTP/gRPC instrumentation | ~600 |
| infrastructure.md | Health checks, K8s probes, Docker HEALTHCHECK, infra metrics, APM, cost optimization, incident response | ~550 |
See Also
- docker-ops — Container monitoring with cAdvisor, Docker stats, and health checks
- ci-cd-ops — Pipeline observability, deployment tracking, build metrics
- nginx-ops — Nginx access/error log parsing, request metrics, upstream monitoring
- python-observability-ops — Python-specific instrumentation with structlog, opentelemetry-python
- OpenTelemetry documentation
- Prometheus best practices
- Google SRE Book — Monitoring chapter
- Grafana dashboards library
Files (claude-mods)
-
assets
-
.gitkeep 0 B · in bundle
-
-
references
-
infrastructure.md 28.3 KB
# Infrastructure Monitoring Reference Comprehensive reference for health checks, infrastructure metrics, APM, cost optimization, capacity planning, and incident response. --- ## Health Checks ### Types of Health Checks | Type | Question It Answers | Failure Action | |------|---------------------|----------------| | **Liveness** | Is the process alive and not deadlocked? | Restart the process | | **Readiness** | Can this instance serve traffic? | Remove from load balancer | | **Startup** | Has the process finished initializing? | Wait (don't restart yet) | ### Implementation Patterns #### Basic Health Check Endpoint ```go // Go type HealthStatus struct { Status string `json:"status"` Timestamp string `json:"timestamp"` Version string `json:"version"` Checks map[string]Check `json:"checks"` } type Check struct { Status string `json:"status"` Message string `json:"message,omitempty"` Latency string `json:"latency,omitempty"` } func healthHandler(w http.ResponseWriter, r *http.Request) { health := HealthStatus{ Status: "ok", Timestamp: time.Now().UTC().Format(time.RFC3339), Version: version, Checks: make(map[string]Check), } // Check database start := time.Now() if err := db.PingContext(r.Context()); err != nil { health.Status = "degraded" health.Checks["database"] = Check{ Status: "fail", Message: err.Error(), } } else { health.Checks["database"] = Check{ Status: "ok", Latency: time.Since(start).String(), } } // Check Redis start = time.Now() if err := redis.Ping(r.Context()).Err(); err != nil { health.Status = "degraded" health.Checks["redis"] = Check{ Status: "fail", Message: err.Error(), } } else { health.Checks["redis"] = Check{ Status: "ok", Latency: time.Since(start).String(), } } statusCode := http.StatusOK if health.Status != "ok" { statusCode = http.StatusServiceUnavailable } w.Header().Set("Content-Type", "application/json") w.WriteHeader(statusCode) json.NewEncoder(w).Encode(health) } ``` ```python # Python (FastAPI) from fastapi import FastAPI, Response from datetime import datetime, timezone import asyncio app = FastAPI() @app.get("/health") async def health_check(): checks = {} status = "ok" # Database check try: start = datetime.now(timezone.utc) await db.execute("SELECT 1") checks["database"] = { "status": "ok", "latency_ms": (datetime.now(timezone.utc) - start).total_seconds() * 1000, } except Exception as e: status = "degraded" checks["database"] = {"status": "fail", "message": str(e)} # Redis check try: start = datetime.now(timezone.utc) await redis.ping() checks["redis"] = { "status": "ok", "latency_ms": (datetime.now(timezone.utc) - start).total_seconds() * 1000, } except Exception as e: status = "degraded" checks["redis"] = {"status": "fail", "message": str(e)} response_code = 200 if status == "ok" else 503 return Response( content=json.dumps({ "status": status, "timestamp": datetime.now(timezone.utc).isoformat(), "checks": checks, }), status_code=response_code, media_type="application/json", ) @app.get("/ready") async def readiness_check(): """Readiness: can we serve traffic?""" try: await db.execute("SELECT 1") return {"status": "ready"} except Exception: return Response( content='{"status": "not_ready"}', status_code=503, media_type="application/json", ) @app.get("/live") async def liveness_check(): """Liveness: is the process alive?""" return {"status": "alive"} ``` #### Health Check Response Format ```json { "status": "ok", "timestamp": "2026-03-09T14:32:01Z", "version": "1.4.2", "checks": { "database": { "status": "ok", "latency_ms": 2.3 }, "redis": { "status": "ok", "latency_ms": 0.8 }, "external_api": { "status": "degraded", "message": "Elevated latency", "latency_ms": 850 } } } ``` ### Liveness vs Readiness Decision Guide ``` Is the process able to make progress? ├─ No (deadlocked, OOM, infinite loop) │ └─ Liveness check should FAIL → container gets restarted │ └─ Yes, but... ├─ Database is temporarily unreachable │ └─ Readiness FAIL, Liveness PASS → stop sending traffic, don't restart │ ├─ Still loading initial data/cache │ └─ Startup FAIL → don't check liveness yet, wait │ └─ Everything is fine └─ All checks PASS → serve traffic normally ``` **Common mistake:** Making liveness depend on external dependencies (database, Redis). If the database is down, restarting the application won't help — it will cause a restart storm. --- ## Kubernetes Probes ### Configuration ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: api-server spec: template: spec: containers: - name: api image: api-server:1.4.2 ports: - containerPort: 8080 # Startup probe: runs first, disables liveness/readiness until passing startupProbe: httpGet: path: /health port: 8080 initialDelaySeconds: 5 periodSeconds: 5 failureThreshold: 30 # 30 * 5s = 150s max startup time successThreshold: 1 # Liveness probe: is the process alive? livenessProbe: httpGet: path: /live port: 8080 initialDelaySeconds: 0 # Starts after startup probe passes periodSeconds: 10 timeoutSeconds: 3 failureThreshold: 3 # 3 consecutive failures → restart successThreshold: 1 # Readiness probe: can it serve traffic? readinessProbe: httpGet: path: /ready port: 8080 initialDelaySeconds: 0 periodSeconds: 5 timeoutSeconds: 3 failureThreshold: 3 # 3 failures → remove from Service successThreshold: 1 resources: requests: cpu: 100m memory: 128Mi limits: cpu: 500m memory: 512Mi ``` ### Probe Types #### HTTP GET ```yaml livenessProbe: httpGet: path: /health port: 8080 httpHeaders: - name: Authorization value: Bearer internal-token ``` #### TCP Socket ```yaml # For services that don't have HTTP (databases, message brokers) livenessProbe: tcpSocket: port: 5432 periodSeconds: 10 ``` #### Exec Command ```yaml # Run a command inside the container livenessProbe: exec: command: - /bin/sh - -c - pg_isready -U postgres periodSeconds: 10 ``` #### gRPC Health Check ```yaml # gRPC health checking protocol livenessProbe: grpc: port: 50051 service: "" # Empty string checks overall server health periodSeconds: 10 ``` ### Probe Configuration Guidelines | Parameter | Liveness | Readiness | Startup | |-----------|----------|-----------|---------| | `initialDelaySeconds` | 0 (use startup probe) | 0 | 5-10 | | `periodSeconds` | 10-15 | 5-10 | 5 | | `timeoutSeconds` | 3-5 | 3-5 | 3-5 | | `failureThreshold` | 3 | 3 | 30 (generous) | | `successThreshold` | 1 | 1-2 | 1 | --- ## Docker HEALTHCHECK ```dockerfile # Dockerfile FROM node:20-slim HEALTHCHECK --interval=30s --timeout=5s --retries=3 --start-period=60s \ CMD curl -f http://localhost:8080/health || exit 1 # Or with wget (no curl in alpine) HEALTHCHECK --interval=30s --timeout=5s --retries=3 --start-period=60s \ CMD wget --no-verbose --tries=1 --spider http://localhost:8080/health || exit 1 ``` ### docker-compose Health Check ```yaml services: api: image: api-server:1.4.2 healthcheck: test: ["CMD", "curl", "-f", "http://localhost:8080/health"] interval: 30s timeout: 5s retries: 3 start_period: 60s worker: image: worker:1.2.0 depends_on: api: condition: service_healthy postgres: condition: service_healthy postgres: image: postgres:16 healthcheck: test: ["CMD-SHELL", "pg_isready -U postgres"] interval: 10s timeout: 5s retries: 5 ``` ### Health Check Parameters | Parameter | Description | Default | Recommendation | |-----------|-------------|---------|----------------| | `interval` | Time between checks | 30s | 15-30s for critical services | | `timeout` | Max time for check | 30s | 3-5s (fail fast) | | `retries` | Failures before unhealthy | 3 | 3 (avoid flapping) | | `start_period` | Grace period for startup | 0s | Set to max startup time | --- ## Uptime Monitoring ### Uptime Kuma Setup ```yaml # docker-compose.yml services: uptime-kuma: image: louislam/uptime-kuma:1 restart: unless-stopped ports: - "3001:3001" volumes: - uptime-kuma-data:/app/data labels: - "traefik.enable=true" - "traefik.http.routers.uptime.rule=Host(`status.example.com`)" volumes: uptime-kuma-data: ``` **Monitor types supported:** - HTTP(s) — status code, keyword, response time - TCP — port open check - DNS — resolution check - Docker container — running status - gRPC — health check protocol - MQTT — broker connectivity - Ping (ICMP) — network reachability - Push — heartbeat endpoint (service pushes to Uptime Kuma) ### Synthetic Monitoring Scripted checks that simulate real user behavior from multiple regions: ```javascript // k6 script for synthetic monitoring import { check, sleep } from 'k6'; import http from 'k6/http'; export const options = { scenarios: { synthetic: { executor: 'constant-vus', vus: 1, duration: '24h', gracefulStop: '0s', }, }, thresholds: { http_req_duration: ['p(95)<500'], // 95% under 500ms http_req_failed: ['rate<0.01'], // < 1% failure rate checks: ['rate>0.99'], // 99% checks pass }, }; export default function () { // Check homepage let res = http.get('https://www.example.com'); check(res, { 'homepage status 200': (r) => r.status === 200, 'homepage loads fast': (r) => r.timings.duration < 500, 'homepage has title': (r) => r.body.includes('<title>'), }); // Check API health res = http.get('https://api.example.com/health'); check(res, { 'api health 200': (r) => r.status === 200, 'api reports ok': (r) => JSON.parse(r.body).status === 'ok', }); // Check login flow res = http.post('https://api.example.com/auth/login', JSON.stringify({ email: 'synthetic-user@example.com', password: process.env.SYNTHETIC_PASSWORD, }), { headers: { 'Content-Type': 'application/json' } }); check(res, { 'login succeeds': (r) => r.status === 200, 'login returns token': (r) => JSON.parse(r.body).token !== undefined, }); sleep(60); // Check every 60 seconds } ``` ### Multi-Region Monitoring | Provider | Regions | Free Tier | Notes | |----------|---------|-----------|-------| | **Uptime Kuma** | Self-hosted (1 region) | Free | Deploy in multiple regions yourself | | **Betteruptime** | 10+ regions | 5 monitors | Status page included | | **Grafana Synthetic** | 20+ regions | Part of Grafana Cloud | k6-based scripts | | **Datadog Synthetic** | 100+ locations | 100 API tests/month | Full browser testing | | **AWS CloudWatch Synthetics** | All AWS regions | Pay per run | Canary scripts | --- ## Infrastructure Metrics ### CPU Metrics | Metric | Source | What It Shows | |--------|--------|---------------| | `node_cpu_seconds_total{mode="user"}` | node_exporter | Time in user space | | `node_cpu_seconds_total{mode="system"}` | node_exporter | Time in kernel space | | `node_cpu_seconds_total{mode="iowait"}` | node_exporter | Time waiting for I/O | | `node_cpu_seconds_total{mode="idle"}` | node_exporter | Idle time | | `node_load1` / `node_load5` / `node_load15` | node_exporter | Load average (1/5/15 min) | **Common queries:** ```promql # CPU usage percentage (all modes except idle) 1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) # CPU usage by mode sum by (mode) (rate(node_cpu_seconds_total{instance="web01:9100"}[5m])) # IO wait percentage (high = disk bottleneck) avg by (instance) (rate(node_cpu_seconds_total{mode="iowait"}[5m])) # Load average vs CPU count node_load1 / count without (cpu) (node_cpu_seconds_total{mode="idle"}) ``` ### Memory Metrics | Metric | What It Shows | |--------|---------------| | `node_memory_MemTotal_bytes` | Total physical memory | | `node_memory_MemAvailable_bytes` | Memory available for applications | | `node_memory_Cached_bytes` | Page cache (reclaimable) | | `node_memory_Buffers_bytes` | Buffer cache | | `node_memory_SwapTotal_bytes` | Total swap | | `node_memory_SwapFree_bytes` | Free swap | ```promql # Memory usage percentage 1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) # Memory breakdown node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes - node_memory_Cached_bytes - node_memory_Buffers_bytes # Swap usage (any swap usage may indicate memory pressure) 1 - (node_memory_SwapFree_bytes / node_memory_SwapTotal_bytes) ``` ### Disk Metrics ```promql # Disk usage percentage 1 - (node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"} / node_filesystem_size_bytes) # Disk I/O utilization (percentage of time doing I/O) rate(node_disk_io_time_seconds_total[5m]) # Read/write throughput rate(node_disk_read_bytes_total[5m]) rate(node_disk_written_bytes_total[5m]) # IOPS rate(node_disk_reads_completed_total[5m]) rate(node_disk_writes_completed_total[5m]) # Average I/O latency rate(node_disk_read_time_seconds_total[5m]) / rate(node_disk_reads_completed_total[5m]) ``` ### Network Metrics ```promql # Bandwidth (bytes/sec) rate(node_network_receive_bytes_total{device!="lo"}[5m]) rate(node_network_transmit_bytes_total{device!="lo"}[5m]) # Packet errors rate(node_network_receive_errs_total[5m]) rate(node_network_transmit_errs_total[5m]) # TCP connections node_netstat_Tcp_CurrEstab # Current established connections rate(node_netstat_Tcp_ActiveOpens[5m]) # New outbound connections/sec rate(node_netstat_Tcp_PassiveOpens[5m]) # New inbound connections/sec ``` --- ## Container Metrics ### cAdvisor Metrics | Metric | Description | |--------|-------------| | `container_cpu_usage_seconds_total` | Total CPU time consumed | | `container_cpu_cfs_throttled_periods_total` | CPU throttling events | | `container_memory_working_set_bytes` | Current memory (excludes cache) | | `container_memory_usage_bytes` | Total memory (includes cache) | | `container_network_receive_bytes_total` | Network inbound bytes | | `container_network_transmit_bytes_total` | Network outbound bytes | | `container_fs_usage_bytes` | Container filesystem usage | | `container_spec_memory_limit_bytes` | Memory limit | | `container_spec_cpu_quota` | CPU quota | ```promql # Container CPU usage percentage (of limit) sum by (container, pod) ( rate(container_cpu_usage_seconds_total{container!="POD",container!=""}[5m]) ) / sum by (container, pod) ( container_spec_cpu_quota / container_spec_cpu_period ) # Container memory usage percentage (of limit) container_memory_working_set_bytes{container!="POD",container!=""} / container_spec_memory_limit_bytes{container!="POD",container!=""} > 0 # CPU throttling percentage sum by (container, pod) ( rate(container_cpu_cfs_throttled_periods_total[5m]) ) / sum by (container, pod) ( rate(container_cpu_cfs_periods_total[5m]) ) # OOMKill detection increase(kube_pod_container_status_restarts_total[1h]) > 0 and kube_pod_container_status_last_terminated_reason{reason="OOMKilled"} ``` ### Kubernetes Metrics (kube-state-metrics) ```promql # Pod status kube_pod_status_phase{phase="Running"} kube_pod_status_phase{phase="Pending"} kube_pod_status_phase{phase="Failed"} # Deployment replicas kube_deployment_status_replicas_available kube_deployment_spec_replicas # HPA status kube_horizontalpodautoscaler_status_current_replicas kube_horizontalpodautoscaler_spec_max_replicas ``` --- ## Node Exporter ### Setup ```yaml # docker-compose.yml services: node-exporter: image: prom/node-exporter:v1.7.0 restart: unless-stopped ports: - "9100:9100" volumes: - /proc:/host/proc:ro - /sys:/host/sys:ro - /:/rootfs:ro command: - '--path.procfs=/host/proc' - '--path.sysfs=/host/sys' - '--path.rootfs=/rootfs' - '--collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host|etc)($$|/)' ``` ### Kubernetes DaemonSet ```yaml apiVersion: apps/v1 kind: DaemonSet metadata: name: node-exporter namespace: monitoring spec: selector: matchLabels: app: node-exporter template: metadata: labels: app: node-exporter annotations: prometheus.io/scrape: "true" prometheus.io/port: "9100" spec: hostPID: true hostNetwork: true containers: - name: node-exporter image: prom/node-exporter:v1.7.0 ports: - containerPort: 9100 hostPort: 9100 volumeMounts: - name: proc mountPath: /host/proc readOnly: true - name: sys mountPath: /host/sys readOnly: true volumes: - name: proc hostPath: path: /proc - name: sys hostPath: path: /sys tolerations: - effect: NoSchedule operator: Exists ``` --- ## APM Tools ### Comparison | Feature | Datadog APM | New Relic | Elastic APM | Sentry | |---------|-------------|-----------|-------------|--------| | **Type** | Full APM | Full APM | Full APM | Error tracking + perf | | **Pricing** | Per host ($31+/mo) | Per user + data | Free (self-host) or Cloud | Per event volume | | **Traces** | Yes | Yes | Yes | Transaction traces | | **Error tracking** | Yes | Yes | Yes | Excellent | | **Profiling** | Yes (continuous) | Yes | No | No | | **Log correlation** | Yes | Yes | Yes | Breadcrumbs | | **Dashboards** | Built-in | Built-in | Kibana | Limited | | **Setup** | Agent-based | Agent-based | Agent or OTel | SDK-based | | **Best for** | Enterprise, full stack | Full observability | Self-hosted, ELK users | Error-focused teams | ### Sentry Error Tracking ```python # Python import sentry_sdk from sentry_sdk.integrations.fastapi import FastApiIntegration sentry_sdk.init( dsn="https://key@sentry.io/project", traces_sample_rate=0.1, # 10% of transactions profiles_sample_rate=0.1, environment="production", release="1.4.2", integrations=[FastApiIntegration()], ) ``` ```javascript // Node.js const Sentry = require('@sentry/node'); Sentry.init({ dsn: 'https://key@sentry.io/project', tracesSampleRate: 0.1, environment: 'production', release: '1.4.2', }); ``` ```go // Go import "github.com/getsentry/sentry-go" sentry.Init(sentry.ClientOptions{ Dsn: "https://key@sentry.io/project", TracesSampleRate: 0.1, Environment: "production", Release: "1.4.2", }) defer sentry.Flush(2 * time.Second) ``` --- ## Cost Optimization ### Metric Cardinality Review High cardinality is the most common cost driver in metrics systems: ```promql # Find metrics with the most time series topk(20, count by (__name__) ({__name__=~".+"})) # Find labels with high cardinality count(group by (path) (http_requests_total)) # How many unique paths? count(group by (user_id) (api_calls_total)) # Unbounded! ``` **Reduction strategies:** 1. Remove unused metrics (if nobody dashboards/alerts on it, drop it) 2. Replace high-cardinality labels with bounded categories 3. Use recording rules to pre-aggregate, drop raw metrics 4. Use metric relabeling in Prometheus to drop at scrape time ```yaml # Drop unused metrics at scrape time metric_relabel_configs: - source_labels: [__name__] regex: "go_.*" # Drop Go runtime metrics if unused action: drop ``` ### Log Volume Reduction | Strategy | Savings | Implementation | |----------|---------|----------------| | Set production to INFO | 50-80% | Logger config | | Sample health check logs | 90% for /health | Middleware filter | | Truncate large payloads | 20-40% | Body size limit (4KB) | | Drop duplicate errors | 30-50% | Rate-limit per error type | | Compress in transit | 60-80% bandwidth | Enable gzip on log shipper | ### Trace Sampling | Sampling Rate | Monthly Cost (est.) | Suitability | |---------------|---------------------|-------------| | 100% | $$$$ | Development, < 100 req/s | | 10% | $$$ | Staging, medium traffic | | 1% | $$ | Production, high traffic | | Tail-based (errors + slow) | $$ | Production (recommended) | | 0.1% | $ | Very high traffic (> 100k req/s) | ### Retention Tiers | Tier | Metrics | Logs | Traces | |------|---------|------|--------| | Hot (0-14 days) | 15s resolution | Full fidelity | All sampled traces | | Warm (14-90 days) | 1m resolution | Full fidelity | Error + slow traces only | | Cold (90 days - 1 year) | 5m resolution | Compressed | None (rely on metrics) | | Archive (1-7 years) | 1h resolution | Compliance logs only | None | --- ## Capacity Planning ### Load Testing Correlation Run load tests while monitoring infrastructure metrics to establish scaling thresholds: ``` Load Test Results: ┌─────────┬──────────┬────────┬─────────┬──────────────┐ │ RPS │ p99 (ms) │ CPU % │ Mem % │ Error Rate │ ├─────────┼──────────┼────────┼─────────┼──────────────┤ │ 100 │ 45 │ 15 │ 30 │ 0% │ │ 500 │ 85 │ 35 │ 45 │ 0% │ │ 1000 │ 150 │ 55 │ 55 │ 0% │ │ 2000 │ 320 │ 75 │ 65 │ 0.1% │ │ 3000 │ 850 │ 90 │ 72 │ 1.5% │ ← degradation │ 4000 │ 2500 │ 98 │ 78 │ 12% │ ← failure └─────────┴──────────┴────────┴─────────┴──────────────┘ Scaling trigger: 75% CPU → add instance Target capacity: 2x expected peak traffic ``` ### Scaling Triggers ```yaml # Kubernetes HPA apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: api-server spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: api-server minReplicas: 3 maxReplicas: 20 metrics: - type: Resource resource: name: cpu target: type: Utilization averageUtilization: 70 # Scale up at 70% CPU - type: Resource resource: name: memory target: type: Utilization averageUtilization: 80 - type: Pods pods: metric: name: http_requests_per_second target: type: AverageValue averageValue: "1000" # Scale at 1000 RPS per pod behavior: scaleUp: stabilizationWindowSeconds: 60 policies: - type: Percent value: 50 # Max 50% increase per scale-up periodSeconds: 60 scaleDown: stabilizationWindowSeconds: 300 # Wait 5 min before scaling down policies: - type: Percent value: 25 periodSeconds: 120 ``` ### Resource Forecasting ```promql # Predict disk full in N hours predict_linear(node_filesystem_avail_bytes[7d], 30*24*3600) < 0 # "Disk will be full within 30 days" # Predict memory usage trend predict_linear( avg_over_time(container_memory_working_set_bytes[7d]), 30*24*3600 ) # Growth rate of database size rate(pg_database_size_bytes[7d]) # Convert to "GB per month" rate(pg_database_size_bytes[7d]) * 86400 * 30 / 1e9 ``` --- ## Incident Response ### Incident Lifecycle ``` Detection → Triage → Mitigate → Resolve → Postmortem │ │ │ │ │ │ │ │ │ └─ Blameless review │ │ │ └─ Root cause fix deployed │ │ └─ User impact reduced/eliminated │ └─ Severity assigned, team engaged └─ Alert fires or user reports issue ``` ### Severity Classification | Severity | Impact | Response Time | Examples | |----------|--------|---------------|---------| | **SEV1 (Critical)** | Service down, data loss, security breach | < 15 minutes | Complete outage, payment processing failure | | **SEV2 (Major)** | Significant degradation, partial outage | < 30 minutes | One region down, 50%+ error rate | | **SEV3 (Minor)** | Limited impact, workaround exists | < 4 hours | Single feature broken, elevated latency | | **SEV4 (Low)** | Minimal impact, cosmetic | Next business day | UI glitch, non-critical alert firing | ### Incident Commander Checklist ```markdown ## Initial Response (first 15 minutes) - [ ] Acknowledge the alert / report - [ ] Assess severity (SEV1-4) - [ ] Open incident channel (#inc-YYYYMMDD-description) - [ ] Page relevant team members - [ ] Post initial status update ## Triage (15-30 minutes) - [ ] Identify affected services and scope - [ ] Check recent deployments: any changes in last 2 hours? - [ ] Check dashboards for anomalies - [ ] Check external dependencies (status pages) - [ ] Determine if rollback is feasible ## Mitigation - [ ] Implement immediate fix (rollback, feature flag, scaling) - [ ] Verify user impact is reduced - [ ] Update status page - [ ] Communicate ETA for full resolution ## Resolution - [ ] Confirm root cause - [ ] Deploy fix - [ ] Verify metrics return to baseline - [ ] Clear incident status - [ ] Schedule postmortem within 48 hours ``` ### Postmortem Template ```markdown # Incident Postmortem: [TITLE] **Date:** 2026-03-09 **Duration:** 45 minutes (14:15 - 15:00 UTC) **Severity:** SEV2 **Author:** [Name] **Status:** Complete ## Summary One-paragraph description of what happened and impact. ## Impact - Users affected: ~5,000 - Revenue impact: ~$2,500 - SLO budget consumed: 3.2 hours of the monthly 43-minute budget ## Timeline (all times UTC) | Time | Event | |------|-------| | 14:12 | Deploy v1.4.3 to production | | 14:15 | Error rate alert fires (5% → 15%) | | 14:17 | On-call acknowledges, starts investigation | | 14:22 | Root cause identified: new query missing index | | 14:25 | Decision: rollback v1.4.3 | | 14:30 | Rollback complete | | 14:35 | Error rate returns to baseline | | 15:00 | All-clear declared | ## Root Cause The v1.4.3 deployment added a new API endpoint that queried the orders table without an index on `user_id + created_at`. Under load, this caused connection pool exhaustion, which cascaded to other endpoints. ## Detection Alert fired 3 minutes after deploy. Detection was effective. ## Contributing Factors 1. No load test for the new endpoint 2. Missing index not caught in code review 3. No query performance checks in CI ## Action Items | Action | Owner | Due | Status | |--------|-------|-----|--------| | Add index on orders(user_id, created_at) | @backend | 2026-03-10 | Done | | Add slow query detection to CI pipeline | @platform | 2026-03-15 | TODO | | Add load test for new endpoints to deploy checklist | @backend | 2026-03-12 | TODO | | Set up query performance alerting (> 100ms avg) | @sre | 2026-03-14 | TODO | ## Lessons Learned - What went well: Fast detection (3 min), fast rollback (8 min) - What went poorly: No pre-production load test caught the issue - Where we got lucky: Happened during business hours, not at 3 AM ``` ### Communication During Incidents | Audience | Channel | Frequency | Content | |----------|---------|-----------|---------| | Engineering | Slack #incident | Real-time | Technical details, commands run | | Management | Slack #incidents-summary | Every 15-30 min | Impact, ETA, escalation needs | | Customers | Status page | Every 15-30 min | User-facing impact, workarounds | | Support | Slack #support-escalation | On status change | Scripted responses, known workarounds | -
logging.md 26.3 KB
# Logging Reference Comprehensive reference for structured logging, log aggregation, correlation, and language-specific implementations. --- ## Structured Logging ### Why Structured Logging Unstructured logs are human-readable but machine-hostile: ``` # BAD: unstructured 2026-03-09 14:32:01 ERROR Failed to process payment for user 789: timeout after 30s # GOOD: structured JSON {"timestamp":"2026-03-09T14:32:01.123Z","level":"ERROR","message":"Failed to process payment","user_id":"789","error":"timeout after 30s","duration_ms":30042} ``` Structured logs enable: - Machine parsing and indexing - Filtering by any field (`user_id=789`, `level=ERROR`) - Aggregation and metric extraction - Correlation with traces via `trace_id` ### Standard JSON Log Format ```json { "timestamp": "2026-03-09T14:32:01.123Z", "level": "ERROR", "message": "Failed to process payment", "logger": "payment.processor", "service": "payment-api", "version": "1.4.2", "environment": "production", "host": "payment-api-7b4d9f-x2k9l", "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736", "span_id": "00f067aa0ba902b7", "request_id": "req-abc123", "user_id": "usr-789", "error": { "type": "PaymentGatewayTimeout", "message": "Gateway response timeout after 30s", "stack": "PaymentGatewayTimeout: Gateway response timeout...\n at processPayment (payment.go:142)\n at handleRequest (handler.go:87)" }, "context": { "payment_id": "pay-456", "amount_cents": 2500, "currency": "USD", "gateway": "stripe" } } ``` ### Key Conventions | Field | Type | Required | Notes | |-------|------|----------|-------| | `timestamp` | ISO 8601 string | Yes | Always UTC, millisecond precision | | `level` | string | Yes | DEBUG, INFO, WARN, ERROR, FATAL | | `message` | string | Yes | Human-readable, no variable interpolation in the key | | `service` | string | Yes | Service name (matches Prometheus job label) | | `version` | string | Yes | Application version or git SHA | | `trace_id` | string | When available | OpenTelemetry trace ID (32 hex chars) | | `span_id` | string | When available | OpenTelemetry span ID (16 hex chars) | | `request_id` | string | When available | Edge-generated request ID | | `error` | object | On errors | Include type, message, stack | | `logger` | string | Recommended | Logger name / module path | | `host` | string | Recommended | Hostname or pod name | | `environment` | string | Recommended | production, staging, development | --- ## Log Levels ### Decision Guide ``` Is the process unable to continue? ├─ Yes → FATAL │ Process must exit. Database unreachable at startup, │ invalid critical config, out of memory. │ └─ No → Did an operation fail? ├─ Yes → Is it actionable? │ ├─ Yes → ERROR │ │ Payment failed, API call returned 500, │ │ constraint violation, file not found. │ │ │ └─ No → WARN │ Expected failure, retry will handle it, │ deprecated API used, nearing limit. │ └─ No → Is it worth recording in production? ├─ Yes → INFO │ Request handled, job completed, │ config loaded, connection established. │ └─ No → DEBUG Variable values, SQL queries, cache hit/miss, internal state. ``` ### Level Details #### FATAL ```json {"level":"FATAL","message":"Cannot connect to database","error":{"type":"ConnectionRefused","message":"dial tcp 10.0.0.5:5432: connect: connection refused"},"action":"process_exit"} ``` - Process cannot start or must terminate - Always followed by `os.Exit(1)` or equivalent - Should trigger immediate alerting - Very rare in well-designed systems #### ERROR ```json {"level":"ERROR","message":"Payment processing failed","payment_id":"pay-456","user_id":"usr-789","error":{"type":"GatewayTimeout","message":"Stripe API timeout after 30s"}} ``` - Operation failed and cannot be completed - Someone should investigate (now or soon) - Every ERROR should have an associated alert or dashboard - **Not for:** User input validation failures (that's WARN or INFO) #### WARN ```json {"level":"WARN","message":"Circuit breaker opened for payment gateway","gateway":"stripe","failure_count":5,"retry_after":"30s"} ``` - System is degraded but still functioning - Worth monitoring but not necessarily immediate action - Retry succeeded, fallback activated, approaching a limit - **Not for:** Expected user errors (wrong password → INFO) #### INFO ```json {"level":"INFO","message":"Request completed","method":"GET","path":"/api/users","status":200,"duration_ms":45,"request_id":"req-abc123"} ``` - Normal operation, audit trail, business events - Should not be noisy (aim for 1-5 lines per request) - Deployments, configuration changes, job completions - **Not for:** Debugging details (use DEBUG) #### DEBUG ```json {"level":"DEBUG","message":"Cache lookup","key":"user:789","hit":true,"ttl_remaining_ms":45200} ``` - Development and troubleshooting only - Disabled in production by default - Enable per-service or per-module when debugging - SQL queries, cache operations, internal state ### Production Log Level Strategy ``` Production default: INFO Production debug: DEBUG (per-service, time-limited, via config change) Staging: DEBUG Development: DEBUG CI/Test: WARN (reduce noise in test output) ``` --- ## Correlation IDs ### Generating IDs ```go // Go: UUID v4 import "github.com/google/uuid" requestID := uuid.New().String() // "550e8400-e29b-41d4-a716-446655440000" // Go: ULID (sortable, timestamp-prefixed) import "github.com/oklog/ulid/v2" requestID := ulid.Make().String() // "01ARZ3NDEKTSV4RRFFQ69G5FAV" ``` ```python # Python: UUID v4 import uuid request_id = str(uuid.uuid4()) # Python: ULID import ulid request_id = str(ulid.new()) ``` ```javascript // Node.js: UUID v4 import { randomUUID } from 'crypto'; const requestId = randomUUID(); // Node.js: ULID import { ulid } from 'ulid'; const requestId = ulid(); ``` ### Propagation Middleware #### Go (net/http) ```go func correlationMiddleware(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { // Get or generate request ID requestID := r.Header.Get("X-Request-ID") if requestID == "" { requestID = uuid.New().String() } // Get trace context from OpenTelemetry span := trace.SpanFromContext(r.Context()) traceID := span.SpanContext().TraceID().String() spanID := span.SpanContext().SpanID().String() // Add to context ctx := context.WithValue(r.Context(), "request_id", requestID) // Add to response headers w.Header().Set("X-Request-ID", requestID) // Add to logger context logger := slog.With( "request_id", requestID, "trace_id", traceID, "span_id", spanID, ) ctx = context.WithValue(ctx, "logger", logger) next.ServeHTTP(w, r.WithContext(ctx)) }) } ``` #### Python (FastAPI) ```python import uuid from contextvars import ContextVar from fastapi import FastAPI, Request from starlette.middleware.base import BaseHTTPMiddleware request_id_var: ContextVar[str] = ContextVar("request_id", default="") class CorrelationMiddleware(BaseHTTPMiddleware): async def dispatch(self, request: Request, call_next): request_id = request.headers.get("X-Request-ID", str(uuid.uuid4())) request_id_var.set(request_id) response = await call_next(request) response.headers["X-Request-ID"] = request_id return response app = FastAPI() app.add_middleware(CorrelationMiddleware) ``` #### Node.js (Express) ```javascript import { randomUUID } from 'crypto'; import { AsyncLocalStorage } from 'async_hooks'; const asyncLocalStorage = new AsyncLocalStorage(); function correlationMiddleware(req, res, next) { const requestId = req.headers['x-request-id'] || randomUUID(); res.setHeader('X-Request-ID', requestId); asyncLocalStorage.run({ requestId }, () => { next(); }); } // Access anywhere in the request lifecycle function getRequestId() { return asyncLocalStorage.getStore()?.requestId || 'unknown'; } ``` ### HTTP Client Propagation Always forward correlation IDs when making outbound HTTP calls: ```go // Go req, _ := http.NewRequestWithContext(ctx, "GET", url, nil) req.Header.Set("X-Request-ID", getRequestID(ctx)) // OpenTelemetry propagation is automatic with instrumented HTTP client ``` ```python # Python headers = {"X-Request-ID": request_id_var.get()} response = httpx.get(url, headers=headers) ``` --- ## Request Context Logging ### Standard Request/Response Log ```go func loggingMiddleware(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { start := time.Now() wrapped := &responseWriter{ResponseWriter: w, statusCode: 200} next.ServeHTTP(wrapped, r) duration := time.Since(start) logger.InfoContext(r.Context(), "Request completed", "method", r.Method, "path", r.URL.Path, "status", wrapped.statusCode, "duration_ms", duration.Milliseconds(), "bytes_written", wrapped.bytesWritten, "remote_addr", r.RemoteAddr, "user_agent", r.UserAgent(), ) }) } ``` ### What to Log Per Request | Field | When | Notes | |-------|------|-------| | Method, path, status, duration | Always | Core request metadata | | Request ID, trace ID | Always | Correlation | | User ID | When authenticated | For audit trail | | Request body | Selectively | Only for mutations, with size limit | | Response body | Rarely | Only for debugging, never in production | | Query parameters | When relevant | Sanitize sensitive params | | IP address | For security | Respect privacy regulations | | User-Agent | For analytics | Browser/client identification | ### Body Size Limits ```go // Never log unbounded request/response bodies const maxBodyLogSize = 4096 // 4KB func truncateBody(body []byte) string { if len(body) > maxBodyLogSize { return string(body[:maxBodyLogSize]) + "... [truncated]" } return string(body) } ``` --- ## Log Aggregation ### Loki **Architecture:** Like Prometheus, but for logs. Index-free design — indexes labels only, not log content. #### Loki Configuration ```yaml # loki-config.yml auth_enabled: false server: http_listen_port: 3100 common: path_prefix: /loki storage: filesystem: chunks_directory: /loki/chunks rules_directory: /loki/rules replication_factor: 1 ring: kvstore: store: inmemory schema_config: configs: - from: 2024-01-01 store: tsdb object_store: filesystem schema: v13 index: prefix: index_ period: 24h limits_config: retention_period: 30d max_query_length: 721h max_entries_limit_per_query: 5000 ``` #### LogQL Basics ```logql # Filter by label {job="api-server"} |= "error" # JSON parsing {job="api-server"} | json | level="ERROR" # Pattern matching {job="api-server"} | json | status_code >= 500 # Rate of log lines (like Prometheus rate) rate({job="api-server"} |= "error" [5m]) # Count errors by path sum by (path) ( count_over_time({job="api-server"} | json | level="ERROR" [5m]) ) # Latency percentile from log field quantile_over_time(0.99, {job="api-server"} | json | unwrap duration_ms [5m]) # Top error messages topk(10, sum by (message) (count_over_time({job="api-server"} | json | level="ERROR" [1h])) ) ``` #### Label Design for Loki ```yaml # GOOD: Low-cardinality labels labels: job: "api-server" environment: "production" namespace: "default" # BAD: High-cardinality labels (will kill Loki performance) labels: user_id: "12345" # Millions of unique values request_id: "abc-123" # Every request is unique path: "/api/users/123" # Include path in log content, not labels ``` **Rule:** Labels in Loki are for stream selection (which container/service), not for filtering log content. Use `| json | field="value"` for content filtering. #### Promtail (Log Collector) ```yaml # promtail-config.yml server: http_listen_port: 9080 positions: filename: /tmp/positions.yaml clients: - url: http://loki:3100/loki/api/v1/push scrape_configs: # Docker container logs - job_name: docker docker_sd_configs: - host: unix:///var/run/docker.sock refresh_interval: 5s relabel_configs: - source_labels: ['__meta_docker_container_name'] target_label: container - source_labels: ['__meta_docker_container_log_stream'] target_label: stream # Kubernetes pod logs - job_name: kubernetes kubernetes_sd_configs: - role: pod pipeline_stages: - docker: {} - json: expressions: level: level trace_id: trace_id - labels: level: - timestamp: source: timestamp format: RFC3339Nano ``` ### ELK Stack (Elasticsearch, Logstash, Kibana) #### Logstash Pipeline ```ruby # logstash.conf input { beats { port => 5044 } } filter { # Parse JSON logs json { source => "message" } # Parse timestamp date { match => ["timestamp", "ISO8601"] target => "@timestamp" } # Add geoip from remote_addr if [remote_addr] { geoip { source => "remote_addr" } } # Redact sensitive fields mutate { remove_field => ["password", "token", "authorization"] } # Parse user-agent if [user_agent] { useragent { source => "user_agent" target => "ua" } } } output { elasticsearch { hosts => ["elasticsearch:9200"] index => "logs-%{[service]}-%{+YYYY.MM.dd}" } } ``` ### CloudWatch Logs ```python # Python: CloudWatch Logs with structlog import structlog import watchtower import logging # CloudWatch handler cw_handler = watchtower.CloudWatchLogHandler( log_group="production/api-server", stream_name="{hostname}-{datetime}", use_queues=True, create_log_group=True, ) # Configure structlog to output JSON structlog.configure( processors=[ structlog.processors.TimeStamper(fmt="iso"), structlog.processors.JSONRenderer() ], wrapper_class=structlog.stdlib.BoundLogger, logger_factory=structlog.stdlib.LoggerFactory(), ) logging.basicConfig(handlers=[cw_handler], level=logging.INFO) ``` #### CloudWatch Metric Filters Extract metrics from log patterns: ```json { "filterPattern": "{ $.level = \"ERROR\" }", "metricTransformations": [ { "metricName": "ErrorCount", "metricNamespace": "ApiServer", "metricValue": "1", "defaultValue": 0 } ] } ``` ```json { "filterPattern": "{ $.duration_ms > 1000 }", "metricTransformations": [ { "metricName": "SlowRequests", "metricNamespace": "ApiServer", "metricValue": "$.duration_ms" } ] } ``` --- ## Log Retention Policies ### Tiered Storage Strategy | Tier | Duration | Resolution | Storage | Cost | |------|----------|------------|---------|------| | **Hot** | 0-7 days | Full fidelity | SSD / fast storage | $$$ | | **Warm** | 7-30 days | Full fidelity | Standard storage | $$ | | **Cold** | 30-90 days | Sampled or compressed | Object storage (S3) | $ | | **Archive** | 90 days - 7 years | Compressed | Glacier / archive | ¢ | ### Compliance Retention Requirements | Regulation | Minimum Retention | Notes | |------------|-------------------|-------| | PCI DSS | 1 year (3 months immediately available) | Audit logs for card data access | | HIPAA | 6 years | Access logs for health data | | SOX | 7 years | Financial system audit trails | | GDPR | "No longer than necessary" | Right to erasure applies | | SOC 2 | 1 year typical | Security event logs | ### Cost Optimization 1. **Set appropriate log levels:** DEBUG off in production saves 50-80% volume 2. **Sample verbose paths:** Log 10% of health check requests 3. **Truncate large fields:** Limit request/response body logging to 4KB 4. **Use log-based metrics:** Extract counts/rates, then archive raw logs 5. **Compress early:** Enable gzip on log transport (Promtail, Fluentd) 6. **Delete test/staging logs aggressively:** 7-day retention for non-production --- ## Sensitive Data Handling ### PII Masking ```go // Go: mask sensitive fields before logging func maskEmail(email string) string { parts := strings.Split(email, "@") if len(parts) != 2 { return "***" } name := parts[0] if len(name) > 2 { name = name[:2] + strings.Repeat("*", len(name)-2) } return name + "@" + parts[1] } func maskCreditCard(cc string) string { if len(cc) < 4 { return "****" } return strings.Repeat("*", len(cc)-4) + cc[len(cc)-4:] } ``` ```python # Python: structlog processor for PII masking import re SENSITIVE_KEYS = {"password", "token", "secret", "authorization", "cookie", "ssn"} EMAIL_PATTERN = re.compile(r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}") CC_PATTERN = re.compile(r"\b\d{4}[\s-]?\d{4}[\s-]?\d{4}[\s-]?\d{4}\b") def mask_sensitive_data(logger, method_name, event_dict): for key, value in list(event_dict.items()): if key.lower() in SENSITIVE_KEYS: event_dict[key] = "***REDACTED***" elif isinstance(value, str): value = EMAIL_PATTERN.sub("[EMAIL]", value) value = CC_PATTERN.sub("[CREDIT_CARD]", value) event_dict[key] = value return event_dict structlog.configure( processors=[ mask_sensitive_data, structlog.processors.JSONRenderer(), ] ) ``` ### Fields to Never Log | Field | Risk | Alternative | |-------|------|-------------| | Passwords | Credential exposure | Log "password changed" event, not the value | | API keys / tokens | Service compromise | Log last 4 characters only | | Credit card numbers | PCI violation | Log last 4 digits, masked | | SSN / national ID | Identity theft | Never log, even masked | | Full request bodies with auth | Token leakage | Strip Authorization header | | Database connection strings | DB credential exposure | Log host:port only | --- ## Log-Based Metrics ### Loki Recording Rules ```yaml # loki-rules.yml groups: - name: log_metrics interval: 1m rules: - record: log:errors:rate5m expr: | sum by (service) ( rate({job=~".+"} | json | level="ERROR" [5m]) ) - record: log:requests:duration_p99_5m expr: | quantile_over_time(0.99, {job="api-server"} | json | unwrap duration_ms [5m] ) by (service) ``` ### Extracting Metrics from Logs When full metrics instrumentation isn't available, derive metrics from structured logs: ```promql # Error rate from logs (Loki) sum(rate({job="api-server"} | json | level="ERROR" [5m])) # Slow request rate from logs sum(rate({job="api-server"} | json | duration_ms > 1000 [5m])) # Unique users from logs (approximate) count( count by (user_id) ( {job="api-server"} | json | user_id != "" [1h] ) ) ``` --- ## Language-Specific Logging ### Go (slog - standard library, Go 1.21+) ```go package main import ( "context" "log/slog" "os" ) func main() { // JSON handler for production handler := slog.NewJSONHandler(os.Stdout, &slog.HandlerOptions{ Level: slog.LevelInfo, AddSource: true, // Add file:line to log entries }) logger := slog.New(handler) slog.SetDefault(logger) // Basic logging slog.Info("Server starting", "port", 8080, "version", "1.4.2") // With context (includes trace_id if using OpenTelemetry bridge) ctx := context.Background() slog.InfoContext(ctx, "Request handled", "method", "GET", "path", "/api/users", "status", 200, "duration_ms", 45, ) // Error logging with error value slog.Error("Database query failed", "error", err, "query", "SELECT * FROM users WHERE id = $1", "user_id", userID, ) // Create child logger with bound attributes userLogger := slog.With("user_id", "usr-789", "session_id", "sess-abc") userLogger.Info("User action", "action", "login") } ``` ### Python (structlog) ```python import structlog structlog.configure( processors=[ structlog.contextvars.merge_contextvars, structlog.processors.add_log_level, structlog.processors.TimeStamper(fmt="iso"), structlog.processors.StackInfoRenderer(), structlog.processors.format_exc_info, structlog.processors.JSONRenderer(), ], wrapper_class=structlog.stdlib.BoundLogger, context_class=dict, logger_factory=structlog.stdlib.LoggerFactory(), ) log = structlog.get_logger() # Basic logging log.info("server_starting", port=8080, version="1.4.2") # Bind context for the request log = log.bind(request_id="req-abc123", user_id="usr-789") log.info("request_handled", method="GET", path="/api/users", status=200, duration_ms=45) # Error with exception try: process_payment(payment_id) except Exception: log.error("payment_failed", payment_id="pay-456", exc_info=True) # Context variables (available across async calls) structlog.contextvars.bind_contextvars(request_id="req-abc123") ``` ### Node.js (pino) ```javascript import pino from 'pino'; const logger = pino({ level: process.env.LOG_LEVEL || 'info', timestamp: pino.stdTimeFunctions.isoTime, formatters: { level: (label) => ({ level: label.toUpperCase() }), }, serializers: { err: pino.stdSerializers.err, req: pino.stdSerializers.req, res: pino.stdSerializers.res, }, redact: ['req.headers.authorization', 'req.headers.cookie', 'password'], }); // Basic logging logger.info({ port: 8080, version: '1.4.2' }, 'Server starting'); // Child logger with bound context const reqLogger = logger.child({ requestId: 'req-abc123', userId: 'usr-789' }); reqLogger.info({ method: 'GET', path: '/api/users', status: 200, durationMs: 45 }, 'Request handled'); // Error logging reqLogger.error({ err, paymentId: 'pay-456' }, 'Payment failed'); // Express/Fastify integration import pinoHttp from 'pino-http'; app.use(pinoHttp({ logger })); ``` ### Rust (tracing crate) ```rust use tracing::{info, error, warn, instrument, Level}; use tracing_subscriber::{fmt, EnvFilter}; fn main() { // JSON subscriber for production tracing_subscriber::fmt() .json() .with_env_filter(EnvFilter::from_default_env()) .with_target(true) .with_thread_ids(true) .with_file(true) .with_line_number(true) .init(); info!(port = 8080, version = "1.4.2", "Server starting"); } #[instrument(skip(db), fields(user_id = %user_id))] async fn get_user(db: &Pool, user_id: &str) -> Result<User, Error> { info!("Fetching user from database"); match db.query_one("SELECT * FROM users WHERE id = $1", &[&user_id]).await { Ok(row) => { info!("User found"); Ok(User::from_row(row)) } Err(e) => { error!(error = %e, "Database query failed"); Err(e.into()) } } } ``` ### Java (Logback + Structured Logging) ```xml <!-- logback.xml --> <configuration> <appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender"> <encoder class="net.logstash.logback.encoder.LogstashEncoder"> <includeMdcKeyName>request_id</includeMdcKeyName> <includeMdcKeyName>trace_id</includeMdcKeyName> <includeMdcKeyName>user_id</includeMdcKeyName> </encoder> </appender> <root level="INFO"> <appender-ref ref="STDOUT" /> </root> </configuration> ``` ```java import org.slf4j.Logger; import org.slf4j.LoggerFactory; import org.slf4j.MDC; import net.logstash.logback.argument.StructuredArguments; import static net.logstash.logback.argument.StructuredArguments.*; Logger log = LoggerFactory.getLogger(PaymentService.class); // Set MDC for request context MDC.put("request_id", requestId); MDC.put("trace_id", traceId); MDC.put("user_id", userId); // Structured logging with key-value pairs log.info("Request handled", kv("method", "GET"), kv("path", "/api/users"), kv("status", 200), kv("duration_ms", 45)); // Error logging log.error("Payment failed", kv("payment_id", paymentId), kv("error", e.getMessage()), e); // Clean up MDC MDC.clear(); ``` --- ## Common Patterns ### Error Logging with Stack Traces Always include the full stack trace for errors, but consider truncation for very deep stacks: ```go // Go slog.Error("Operation failed", "error", err.Error(), "stack", fmt.Sprintf("%+v", err), // With pkgs/errors stack ) ``` ```python # Python - structlog handles exc_info automatically log.error("operation_failed", exc_info=True) ``` ### Audit Logging For compliance-required operations: ```json { "timestamp": "2026-03-09T14:32:01.123Z", "level": "INFO", "type": "audit", "action": "user.role.changed", "actor": {"id": "usr-admin-1", "type": "user", "ip": "10.0.0.5"}, "target": {"id": "usr-789", "type": "user"}, "changes": {"role": {"from": "viewer", "to": "editor"}}, "result": "success", "request_id": "req-abc123" } ``` ### Request/Response Logging ```go // Log request on entry, response on exit func loggingMiddleware(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { start := time.Now() // Log request (don't log body for GET, limit body size for POST) slog.InfoContext(r.Context(), "Request received", "method", r.Method, "path", r.URL.Path, "remote_addr", r.RemoteAddr, ) wrapped := wrapResponseWriter(w) next.ServeHTTP(wrapped, r) slog.InfoContext(r.Context(), "Request completed", "method", r.Method, "path", r.URL.Path, "status", wrapped.Status(), "duration_ms", time.Since(start).Milliseconds(), "bytes", wrapped.BytesWritten(), ) }) } ``` ### Health Check Log Suppression Don't fill logs with health check noise: ```go func loggingMiddleware(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { // Skip logging for health checks if r.URL.Path == "/health" || r.URL.Path == "/ready" { next.ServeHTTP(w, r) return } // ... normal logging }) } ``` -
metrics-alerting.md 27.6 KB
# Metrics and Alerting Reference Comprehensive reference for metrics collection, visualization, alerting, SLOs, and uptime monitoring. --- ## Prometheus ### Architecture Overview ``` ┌─────────────┐ ┌─────────────┐ ┌─────────────────┐ │ Application │────▶│ Prometheus │────▶│ Alertmanager │ │ /metrics │pull │ (TSDB) │push │ (routing/notif) │ └─────────────┘ └──────┬──────┘ └─────────────────┘ │query ┌──────▼──────┐ │ Grafana │ │ (dashboards) │ └─────────────┘ ``` **Key characteristics:** - Pull-based model (Prometheus scrapes targets) - Local time-series database (TSDB) - PromQL query language - Built-in alerting rules evaluated by Prometheus, routed by Alertmanager - Service discovery (Kubernetes, Consul, DNS, file-based, EC2) ### Prometheus Configuration (prometheus.yml) ```yaml global: scrape_interval: 15s # Default scrape interval evaluation_interval: 15s # Rule evaluation interval scrape_timeout: 10s # Per-scrape timeout # Alertmanager configuration alerting: alertmanagers: - static_configs: - targets: - alertmanager:9093 # Rule files rule_files: - "rules/*.yml" # Scrape targets scrape_configs: # Self-monitoring - job_name: "prometheus" static_configs: - targets: ["localhost:9090"] # Application with static targets - job_name: "api-server" metrics_path: /metrics scheme: https static_configs: - targets: ["api1:8080", "api2:8080"] labels: environment: production # Kubernetes service discovery - job_name: "kubernetes-pods" kubernetes_sd_configs: - role: pod relabel_configs: - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape] action: keep regex: true - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_path] action: replace target_label: __metrics_path__ regex: (.+) - source_labels: [__meta_kubernetes_namespace] action: replace target_label: namespace - source_labels: [__meta_kubernetes_pod_name] action: replace target_label: pod # Node exporter - job_name: "node" static_configs: - targets: ["node-exporter:9100"] ``` ### PromQL Basics #### Rate and Increase ```promql # Per-second rate over 5 minutes (use for counters) rate(http_requests_total[5m]) # Per-second rate for specific status codes rate(http_requests_total{status_code=~"5.."}[5m]) # Total increase over 1 hour (use for counters) increase(http_requests_total[1h]) # irate: instant rate using last two data points (more volatile) irate(http_requests_total[5m]) ``` **Rule:** Always use `rate()` or `increase()` with counters. Never display raw counter values. #### Aggregation Operators ```promql # Sum across all instances sum(rate(http_requests_total[5m])) # Sum by specific label sum by (method, path) (rate(http_requests_total[5m])) # Average across instances avg(node_cpu_seconds_total{mode="idle"}) # Maximum value across instances max by (instance) (node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes) # Count number of time series count(up == 1) # Top 5 by value topk(5, rate(http_requests_total[5m])) # Bottom 5 by value bottomk(5, rate(http_requests_total[5m])) ``` #### Histogram Quantiles ```promql # 99th percentile latency histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m])) ) # 95th percentile latency by service histogram_quantile(0.95, sum by (le, service) (rate(http_request_duration_seconds_bucket[5m])) ) # 50th percentile (median) histogram_quantile(0.50, sum by (le) (rate(http_request_duration_seconds_bucket[5m])) ) # Average latency from histogram sum(rate(http_request_duration_seconds_sum[5m])) / sum(rate(http_request_duration_seconds_count[5m])) ``` #### Useful Functions ```promql # Detect missing metrics (target down) absent(up{job="api-server"}) # Time since last change (staleness) time() - process_start_time_seconds # Predict value in 4 hours using linear regression predict_linear(node_filesystem_avail_bytes[6h], 4*3600) # Compare to 1 week ago rate(http_requests_total[5m]) / rate(http_requests_total[5m] offset 7d) # Clamping values clamp_min(free_disk_percentage, 0) clamp_max(cpu_usage_percentage, 100) # Label manipulation label_replace(up, "short_instance", "$1", "instance", "(.*):.*") ``` ### Recording Rules Pre-compute expensive queries for dashboards and alerts: ```yaml # rules/recording-rules.yml groups: - name: http_request_rules interval: 15s rules: # Pre-compute request rate by service and status - record: job:http_requests:rate5m expr: sum by (job, status_code) (rate(http_requests_total[5m])) # Pre-compute error rate percentage - record: job:http_request_errors:ratio5m expr: | sum by (job) (rate(http_requests_total{status_code=~"5.."}[5m])) / sum by (job) (rate(http_requests_total[5m])) # Pre-compute p99 latency - record: job:http_request_duration_seconds:p99_5m expr: | histogram_quantile(0.99, sum by (job, le) (rate(http_request_duration_seconds_bucket[5m])) ) # Pre-compute availability - record: job:availability:ratio5m expr: | 1 - ( sum by (job) (rate(http_requests_total{status_code=~"5.."}[5m])) / sum by (job) (rate(http_requests_total[5m])) ) ``` ### Alerting Rules ```yaml # rules/alerting-rules.yml groups: - name: service_alerts rules: # High error rate - alert: HighErrorRate expr: job:http_request_errors:ratio5m > 0.01 for: 5m labels: severity: warning team: backend annotations: summary: "High error rate on {{ $labels.job }}" description: "Error rate is {{ $value | humanizePercentage }} (threshold: 1%)" runbook_url: "https://runbooks.example.com/high-error-rate" dashboard_url: "https://grafana.example.com/d/service-overview?var-service={{ $labels.job }}" # Critical error rate - alert: CriticalErrorRate expr: job:http_request_errors:ratio5m > 0.05 for: 2m labels: severity: critical team: backend annotations: summary: "Critical error rate on {{ $labels.job }}" description: "Error rate is {{ $value | humanizePercentage }} (threshold: 5%)" runbook_url: "https://runbooks.example.com/critical-error-rate" # High latency - alert: HighLatencyP99 expr: job:http_request_duration_seconds:p99_5m > 2.0 for: 10m labels: severity: warning annotations: summary: "P99 latency above 2s on {{ $labels.job }}" description: "P99 latency is {{ $value | humanizeDuration }}" # Target down - alert: TargetDown expr: up == 0 for: 3m labels: severity: critical annotations: summary: "Target {{ $labels.instance }} is down" description: "Prometheus cannot scrape {{ $labels.job }}/{{ $labels.instance }}" - name: infrastructure_alerts rules: # Disk space prediction - alert: DiskWillFillIn24Hours expr: | predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}[6h], 24*3600) < 0 for: 30m labels: severity: warning annotations: summary: "Disk {{ $labels.mountpoint }} on {{ $labels.instance }} will fill within 24 hours" # High memory usage - alert: HighMemoryUsage expr: | (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) > 0.9 for: 10m labels: severity: warning annotations: summary: "Memory usage above 90% on {{ $labels.instance }}" # High CPU usage - alert: HighCPUUsage expr: | 1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) > 0.85 for: 15m labels: severity: warning annotations: summary: "CPU usage above 85% on {{ $labels.instance }}" ``` --- ## Grafana ### Dashboard JSON Structure ```json { "dashboard": { "title": "Service Overview", "uid": "service-overview", "tags": ["production", "services"], "timezone": "browser", "refresh": "30s", "time": { "from": "now-6h", "to": "now" }, "templating": { "list": [ { "name": "service", "type": "query", "datasource": "Prometheus", "query": "label_values(up, job)", "refresh": 2, "multi": true, "includeAll": true }, { "name": "interval", "type": "interval", "options": [ {"text": "1m", "value": "1m"}, {"text": "5m", "value": "5m"}, {"text": "15m", "value": "15m"} ], "current": {"text": "5m", "value": "5m"} } ] }, "panels": [] } } ``` ### Panel Types #### Time Series Panel ```json { "type": "timeseries", "title": "Request Rate", "gridPos": {"h": 8, "w": 12, "x": 0, "y": 0}, "targets": [ { "expr": "sum by (status_code) (rate(http_requests_total{job=~\"$service\"}[$interval]))", "legendFormat": "{{status_code}}" } ], "fieldConfig": { "defaults": { "unit": "reqps", "custom": { "drawStyle": "line", "fillOpacity": 10, "stacking": {"mode": "none"} } } } } ``` #### Stat Panel ```json { "type": "stat", "title": "Current Error Rate", "gridPos": {"h": 4, "w": 6, "x": 0, "y": 0}, "targets": [ { "expr": "sum(rate(http_requests_total{job=~\"$service\",status_code=~\"5..\"}[5m])) / sum(rate(http_requests_total{job=~\"$service\"}[5m]))", "instant": true } ], "fieldConfig": { "defaults": { "unit": "percentunit", "thresholds": { "steps": [ {"color": "green", "value": null}, {"color": "yellow", "value": 0.001}, {"color": "red", "value": 0.01} ] } } } } ``` #### Gauge Panel ```json { "type": "gauge", "title": "CPU Usage", "targets": [ { "expr": "1 - avg(rate(node_cpu_seconds_total{mode=\"idle\",instance=~\"$instance\"}[5m]))", "instant": true } ], "fieldConfig": { "defaults": { "unit": "percentunit", "min": 0, "max": 1, "thresholds": { "steps": [ {"color": "green", "value": null}, {"color": "yellow", "value": 0.7}, {"color": "red", "value": 0.9} ] } } } } ``` #### Table Panel ```json { "type": "table", "title": "Top Endpoints by Error Rate", "targets": [ { "expr": "topk(10, sum by (method, path) (rate(http_requests_total{status_code=~\"5..\"}[5m])))", "instant": true, "format": "table" } ], "transformations": [ {"id": "organize", "options": {"excludeByName": {"Time": true}}} ] } ``` ### Grafana Variables | Type | Use Case | Example | |------|----------|---------| | **Query** | Dynamic from datasource | `label_values(up, job)` | | **Custom** | Fixed list of values | `production,staging,development` | | **Interval** | Time range intervals | `1m,5m,15m,1h` | | **Datasource** | Multiple Prometheus instances | Type: datasource, Query: Prometheus | | **Text box** | Free-form input | Filter by custom string | ### Annotations ```json { "annotations": { "list": [ { "name": "Deployments", "datasource": "Prometheus", "enable": true, "expr": "changes(process_start_time_seconds{job=\"api-server\"}[1m]) > 0", "tagKeys": "job", "titleFormat": "Deployment: {{job}}" }, { "name": "Alerts", "datasource": "-- Grafana --", "enable": true, "type": "alert" } ] } } ``` --- ## OpenTelemetry Metrics ### Go SDK Setup ```go package main import ( "context" "log" "time" "go.opentelemetry.io/otel" "go.opentelemetry.io/otel/exporters/prometheus" "go.opentelemetry.io/otel/metric" sdkmetric "go.opentelemetry.io/otel/sdk/metric" ) func initMeterProvider() (*sdkmetric.MeterProvider, error) { exporter, err := prometheus.New() if err != nil { return nil, err } mp := sdkmetric.NewMeterProvider( sdkmetric.WithReader(exporter), ) otel.SetMeterProvider(mp) return mp, nil } func main() { mp, err := initMeterProvider() if err != nil { log.Fatal(err) } defer mp.Shutdown(context.Background()) meter := otel.Meter("myapp") // Counter requestCounter, _ := meter.Int64Counter( "http.server.request.total", metric.WithDescription("Total HTTP requests"), metric.WithUnit("{request}"), ) // Histogram latencyHistogram, _ := meter.Float64Histogram( "http.server.request.duration", metric.WithDescription("HTTP request latency"), metric.WithUnit("s"), metric.WithExplicitBucketBoundaries(0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10), ) // UpDownCounter (gauge-like) activeConnections, _ := meter.Int64UpDownCounter( "http.server.active_connections", metric.WithDescription("Active HTTP connections"), ) // Usage ctx := context.Background() requestCounter.Add(ctx, 1, metric.WithAttributes( attribute.String("method", "GET"), attribute.String("path", "/api/users"), attribute.Int("status_code", 200), )) start := time.Now() // ... handle request ... latencyHistogram.Record(ctx, time.Since(start).Seconds()) activeConnections.Add(ctx, 1) // connection opened activeConnections.Add(ctx, -1) // connection closed } ``` ### Python SDK Setup ```python from opentelemetry import metrics from opentelemetry.sdk.metrics import MeterProvider from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader from opentelemetry.exporter.prometheus import PrometheusMetricReader from prometheus_client import start_http_server # Prometheus exporter reader = PrometheusMetricReader() provider = MeterProvider(metric_readers=[reader]) metrics.set_meter_provider(provider) # Start Prometheus HTTP server on port 8000 start_http_server(8000) meter = metrics.get_meter("myapp") # Counter request_counter = meter.create_counter( name="http.server.request.total", description="Total HTTP requests", unit="{request}", ) # Histogram latency_histogram = meter.create_histogram( name="http.server.request.duration", description="HTTP request latency", unit="s", ) # UpDownCounter active_connections = meter.create_up_down_counter( name="http.server.active_connections", description="Active HTTP connections", ) # Usage request_counter.add(1, {"method": "GET", "path": "/api/users", "status_code": 200}) latency_histogram.record(0.045, {"method": "GET", "path": "/api/users"}) active_connections.add(1) ``` ### Node.js SDK Setup ```javascript const { MeterProvider } = require('@opentelemetry/sdk-metrics'); const { PrometheusExporter } = require('@opentelemetry/exporter-prometheus'); const { metrics } = require('@opentelemetry/api'); const exporter = new PrometheusExporter({ port: 9464 }); const meterProvider = new MeterProvider({ readers: [exporter], }); metrics.setGlobalMeterProvider(meterProvider); const meter = metrics.getMeter('myapp'); // Counter const requestCounter = meter.createCounter('http.server.request.total', { description: 'Total HTTP requests', unit: '{request}', }); // Histogram const latencyHistogram = meter.createHistogram('http.server.request.duration', { description: 'HTTP request latency', unit: 's', advice: { explicitBucketBoundaries: [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10], }, }); // UpDownCounter const activeConnections = meter.createUpDownCounter('http.server.active_connections', { description: 'Active HTTP connections', }); // Usage requestCounter.add(1, { method: 'GET', path: '/api/users', status_code: 200 }); latencyHistogram.record(0.045, { method: 'GET', path: '/api/users' }); activeConnections.add(1); ``` --- ## StatsD ### Protocol Format ``` <metric_name>:<value>|<type>|@<sample_rate>|#<tags> ``` | Type | Code | Example | |------|------|---------| | Counter | `c` | `page.views:1\|c` | | Gauge | `g` | `fuel.level:0.5\|g` | | Timer | `ms` | `request.duration:320\|ms` | | Set | `s` | `users.uniques:user123\|s` | | Histogram | `h` | `request.size:512\|h` (DogStatsD) | | Distribution | `d` | `request.duration:320\|d` (DogStatsD) | ### DogStatsD Extensions (Datadog) ``` # Counter with tags http.requests:1|c|#method:GET,path:/api/users,status:200 # Histogram with sample rate http.request.duration:45.2|h|@0.5|#service:api # Gauge system.cpu.usage:72.5|g|#host:web01 # Service check _sc|myservice.health|0|#env:production|m:Service is healthy ``` **When to use StatsD over Prometheus:** - Existing StatsD infrastructure - Simple counter/gauge/timer needs without complex queries - Push model required (ephemeral jobs, serverless) - Language/framework has StatsD client but no Prometheus client --- ## Custom Metrics Design ### Naming Conventions Follow OpenMetrics/Prometheus naming: ``` <namespace>_<subsystem>_<name>_<unit>_<suffix> ``` | Component | Rules | Examples | |-----------|-------|---------| | Namespace | Application or domain | `myapp`, `payment`, `auth` | | Subsystem | Component within app | `http`, `db`, `cache`, `queue` | | Name | What is measured | `request`, `connection`, `query` | | Unit | SI unit (base, not milli/micro) | `seconds`, `bytes`, `ratio` | | Suffix | Metric type | `_total` (counter), `_info` (metadata), `_bucket` (histogram) | **Good names:** ``` http_server_request_duration_seconds # histogram http_server_requests_total # counter db_connection_pool_active_connections # gauge cache_hit_ratio # gauge (0-1) queue_messages_total # counter payment_processing_duration_seconds # histogram ``` **Bad names:** ``` requestCount # No namespace, no suffix, camelCase latency_ms # Milliseconds (use seconds), no namespace errors # Vague, no namespace, no suffix HttpRequests # PascalCase ``` ### Label Best Practices **Do:** - Use labels for dimensions you will filter/aggregate by - Keep label cardinality bounded (< 100 unique values per label) - Use consistent label names across metrics (`method`, not `http_method` in some and `request_method` in others) **Don't:** - Use user IDs, email addresses, or request IDs as labels (unbounded cardinality) - Use full URL paths as labels (use route templates: `/api/users/{id}`, not `/api/users/12345`) - Use error messages as labels (unbounded text) - Create more than 5-7 labels per metric ### Avoiding Cardinality Bombs ``` # BAD: unbounded path label http_requests_total{path="/api/users/12345"} # Millions of unique series http_requests_total{path="/api/users/67890"} # GOOD: use route template http_requests_total{route="/api/users/{id}"} # One series per route # BAD: error message as label errors_total{message="connection refused to 10.0.0.5:5432"} # GOOD: error category as label errors_total{type="connection_refused", target="postgres"} ``` **Cardinality check query:** ```promql # Find high-cardinality metrics topk(10, count by (__name__) ({__name__=~".+"})) # Check specific metric cardinality count(http_requests_total) ``` --- ## SLI / SLO / SLA ### Definitions | Term | Definition | Example | |------|------------|---------| | **SLI** (Service Level Indicator) | Quantitative measure of service behavior | 99.2% of requests complete in < 500ms | | **SLO** (Service Level Objective) | Target value for an SLI | 99.5% of requests should complete in < 500ms | | **SLA** (Service Level Agreement) | Business contract with consequences | 99.9% availability or credit issued | **Relationship:** SLI measures reality → SLO sets the target → SLA defines business consequences. ### Error Budget Calculation ``` Error budget = 1 - SLO target Example: SLO = 99.9% availability Error budget = 0.1% = 43.2 minutes/month In a 30-day month: - Total minutes: 43,200 - Allowed downtime: 43.2 minutes - Allowed error requests: 0.1% of total ``` ### Burn Rate Alerting Burn rate = rate at which error budget is being consumed relative to the budget period. ``` burn_rate = error_rate / (1 - SLO_target) ``` | Burn Rate | Budget Exhaustion | Alert? | |-----------|-------------------|--------| | 1x | 30 days (full period) | No | | 2x | 15 days | No | | 6x | 5 days | Ticket (warning) | | 14.4x | 2 days | Page (critical) | | 36x | 20 hours | Page immediately | **Multi-window burn rate alert (recommended):** ```yaml # Fast burn: 14.4x burn rate over 1-hour window, confirmed by 5-minute window - alert: SLOHighBurnRate expr: | ( sum(rate(http_requests_total{status_code=~"5.."}[1h])) / sum(rate(http_requests_total[1h])) ) > (14.4 * 0.001) and ( sum(rate(http_requests_total{status_code=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) ) > (14.4 * 0.001) labels: severity: critical annotations: summary: "High error budget burn rate" # Slow burn: 6x burn rate over 6-hour window, confirmed by 30-minute window - alert: SLOSlowBurnRate expr: | ( sum(rate(http_requests_total{status_code=~"5.."}[6h])) / sum(rate(http_requests_total[6h])) ) > (6 * 0.001) and ( sum(rate(http_requests_total{status_code=~"5.."}[30m])) / sum(rate(http_requests_total[30m])) ) > (6 * 0.001) labels: severity: warning ``` ### SLO Document Template ```markdown # SLO: [Service Name] - [SLO Name] ## Overview - **Service:** payment-api - **Owner:** payments-team - **Last reviewed:** 2026-03-01 ## SLI Definition - **Type:** Availability (success rate) - **Good events:** HTTP responses with status < 500 - **Total events:** All HTTP responses - **Measurement:** `sum(rate(http_requests_total{status<500}[5m])) / sum(rate(http_requests_total[5m]))` ## SLO Target - **Target:** 99.9% - **Window:** 30 days (rolling) - **Error budget:** 0.1% = ~43 minutes of downtime ## Alerting - **Fast burn (page):** 14.4x burn rate for 1 hour - **Slow burn (ticket):** 6x burn rate for 6 hours ## Consequences of Missing SLO - Freeze non-critical deployments - Allocate sprint capacity to reliability - Review in next SLO review meeting ``` --- ## Alert Routing ### Alertmanager Configuration ```yaml # alertmanager.yml global: resolve_timeout: 5m slack_api_url: "https://hooks.slack.com/services/T00/B00/XXX" pagerduty_url: "https://events.pagerduty.com/v2/enqueue" route: receiver: "default-slack" group_by: ["alertname", "job"] group_wait: 30s # Wait before sending first notification group_interval: 5m # Wait before sending updates repeat_interval: 4h # Resend if not resolved routes: # Critical alerts → PagerDuty - match: severity: critical receiver: "pagerduty-critical" group_wait: 10s repeat_interval: 1h # Warning alerts → Slack - match: severity: warning receiver: "slack-warnings" repeat_interval: 4h # Info alerts → Slack info channel - match: severity: info receiver: "slack-info" repeat_interval: 24h # Team-specific routing - match: team: database receiver: "pagerduty-database" routes: - match: severity: critical receiver: "pagerduty-database" - match: severity: warning receiver: "slack-database" receivers: - name: "default-slack" slack_configs: - channel: "#alerts" title: '{{ .GroupLabels.alertname }}' text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}' - name: "pagerduty-critical" pagerduty_configs: - service_key: "<integration-key>" severity: critical description: '{{ .GroupLabels.alertname }}: {{ .CommonAnnotations.summary }}' details: description: '{{ .CommonAnnotations.description }}' runbook: '{{ .CommonAnnotations.runbook_url }}' - name: "slack-warnings" slack_configs: - channel: "#alerts-warning" title: ':warning: {{ .GroupLabels.alertname }}' text: '{{ .CommonAnnotations.description }}' - name: "slack-info" slack_configs: - channel: "#alerts-info" inhibit_rules: # Suppress warning if critical is already firing - source_match: severity: critical target_match: severity: warning equal: ["alertname", "job"] ``` ### Runbook Template ```markdown # Runbook: [Alert Name] ## Alert Details - **Alert:** HighErrorRate - **Severity:** Warning / Critical - **Team:** backend ## Symptom What the user/system is experiencing when this alert fires. ## Investigation Steps 1. Check the Grafana dashboard: [link] 2. Check recent deployments: `kubectl rollout history deployment/api` 3. Check error logs: `kubectl logs -l app=api --tail=100 | jq 'select(.level=="ERROR")'` 4. Check downstream dependencies: [dashboard link] ## Mitigation Immediate actions to reduce impact: 1. If caused by recent deploy: `kubectl rollout undo deployment/api` 2. If caused by downstream: Enable circuit breaker / failover 3. If caused by traffic spike: Scale horizontally ## Resolution Steps to fully resolve: 1. Identify root cause from logs/traces 2. Create fix PR 3. Deploy fix through normal pipeline 4. Verify error rate returns to baseline ## Escalation - Level 1: On-call engineer (this runbook) - Level 2: Team lead (@team-lead) - Level 3: VP Engineering (for customer-impacting incidents) ``` --- ## Uptime Monitoring ### Uptime Kuma (Self-hosted) ```yaml # docker-compose.yml services: uptime-kuma: image: louislam/uptime-kuma:1 restart: unless-stopped ports: - "3001:3001" volumes: - uptime-kuma-data:/app/data volumes: uptime-kuma-data: ``` **Features:** - HTTP(s), TCP, DNS, Docker, gRPC, MQTT monitors - Status pages (public-facing) - Notifications: Slack, Discord, Telegram, PagerDuty, email, webhooks - Certificate expiry monitoring - Multi-language support ### Synthetic Monitoring Run scripted checks from multiple regions to verify end-to-end functionality: ```javascript // Example: Grafana synthetic monitoring check import { check } from 'k6'; import http from 'k6/http'; export default function () { const res = http.get('https://api.example.com/health'); check(res, { 'status is 200': (r) => r.status === 200, 'response time < 500ms': (r) => r.timings.duration < 500, 'body contains ok': (r) => r.body.includes('"status":"ok"'), }); } ``` ### Status Pages Communicate service health to users: | Tool | Type | Features | |------|------|----------| | **Uptime Kuma** | Self-hosted | Free, built-in status page | | **Betteruptime** | SaaS | Incident management + status page | | **Cachet** | Self-hosted | PHP-based, mature | | **Instatus** | SaaS | Modern, integrations | | **Statuspage (Atlassian)** | SaaS | Enterprise, expensive | **Status page best practices:** - Show individual component status (API, database, CDN, auth) - Include historical uptime percentage (30/90 day) - Post incident updates promptly (investigating → identified → monitoring → resolved) - Subscribe option for email/SMS/RSS notifications -
tracing.md 27.1 KB
# Distributed Tracing Reference Comprehensive reference for OpenTelemetry, context propagation, sampling, and instrumentation patterns. --- ## OpenTelemetry Architecture ``` ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │ Application │ │ Application │ │ Application │ │ (SDK + API) │ │ (SDK + API) │ │ (SDK + API) │ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │ OTLP │ OTLP │ OTLP ▼ ▼ ▼ ┌─────────────────────────────────────────────────────────┐ │ OTel Collector │ │ ┌───────────┐ ┌────────────┐ ┌───────────────────┐ │ │ │ Receivers │→ │ Processors │→ │ Exporters │ │ │ │ (OTLP, │ │ (batch, │ │ (Jaeger, Tempo, │ │ │ │ Jaeger, │ │ filter, │ │ Datadog, OTLP) │ │ │ │ Zipkin) │ │ tail │ │ │ │ │ │ │ │ sampling) │ │ │ │ │ └───────────┘ └────────────┘ └───────────────────┘ │ └─────────────────────────────────────────────────────────┘ │ │ │ ▼ ▼ ▼ ┌──────────┐ ┌──────────┐ ┌──────────────┐ │ Jaeger │ │ Tempo │ │ Datadog │ │ (UI) │ │ (store) │ │ (SaaS) │ └──────────┘ └──────────┘ └──────────────┘ ``` ### Components | Component | Role | Notes | |-----------|------|-------| | **API** | Stable interfaces for instrumentation | Language-specific, vendor-neutral | | **SDK** | Implementation of the API | Configures sampling, export, processing | | **Collector** | Receives, processes, exports telemetry | Deploy as sidecar or gateway | | **Exporters** | Send data to backends | OTLP (preferred), Jaeger, Zipkin, vendor-specific | | **Auto-instrumentation** | Automatic span creation for frameworks | HTTP, gRPC, database, messaging | ### Collector Configuration ```yaml # otel-collector-config.yml receivers: otlp: protocols: grpc: endpoint: 0.0.0.0:4317 http: endpoint: 0.0.0.0:4318 processors: batch: timeout: 5s send_batch_size: 8192 send_batch_max_size: 16384 memory_limiter: check_interval: 1s limit_mib: 1024 spike_limit_mib: 256 # Tail-based sampling (decide after seeing complete trace) tail_sampling: decision_wait: 10s num_traces: 100000 policies: # Always sample errors - name: errors type: status_code status_code: status_codes: [ERROR] # Always sample slow traces (> 2s) - name: slow-traces type: latency latency: threshold_ms: 2000 # Sample 10% of everything else - name: probabilistic type: probabilistic probabilistic: sampling_percentage: 10 # Add resource attributes resource: attributes: - key: environment value: production action: upsert exporters: otlp/jaeger: endpoint: jaeger:4317 tls: insecure: true otlp/tempo: endpoint: tempo:4317 tls: insecure: true debug: verbosity: detailed service: pipelines: traces: receivers: [otlp] processors: [memory_limiter, tail_sampling, batch, resource] exporters: [otlp/jaeger] metrics: receivers: [otlp] processors: [memory_limiter, batch] exporters: [otlp/tempo] logs: receivers: [otlp] processors: [memory_limiter, batch] exporters: [debug] ``` --- ## Span Model ### Span Anatomy ``` Trace ID: 4bf92f3577b34da6a3ce929d0e0e4736 │ ├─ Span: "GET /api/orders" │ ├─ Span ID: 00f067aa0ba902b7 │ ├─ Parent: (none - root span) │ ├─ Start: 2026-03-09T14:32:01.000Z │ ├─ End: 2026-03-09T14:32:01.245Z │ ├─ Status: OK │ ├─ Attributes: │ │ http.method: GET │ │ http.url: /api/orders?user_id=789 │ │ http.status_code: 200 │ │ http.response_content_length: 4523 │ ├─ Events: │ │ └─ "cache.miss" at T+5ms {key: "orders:usr-789"} │ │ │ ├─ Span: "SELECT orders" │ │ ├─ Span ID: a1b2c3d4e5f60718 │ │ ├─ Parent: 00f067aa0ba902b7 │ │ ├─ Duration: 45ms │ │ ├─ Attributes: │ │ │ db.system: postgresql │ │ │ db.operation: SELECT │ │ │ db.statement: SELECT * FROM orders WHERE user_id = $1 │ │ │ db.rows_affected: 12 │ │ └─ Status: OK │ │ │ └─ Span: "GET payment-service/status" │ ├─ Span ID: b2c3d4e5f6071829 │ ├─ Parent: 00f067aa0ba902b7 │ ├─ Duration: 120ms │ ├─ Attributes: │ │ http.method: GET │ │ http.url: http://payment-service:8080/status │ │ http.status_code: 200 │ │ peer.service: payment-service │ └─ Status: OK ``` ### Span Attributes (Semantic Conventions) #### HTTP Spans | Attribute | Example | Notes | |-----------|---------|-------| | `http.request.method` | `GET` | HTTP method | | `url.path` | `/api/orders` | URL path | | `http.response.status_code` | `200` | Response status | | `http.request.body.size` | `1024` | Request body bytes | | `http.response.body.size` | `4523` | Response body bytes | | `server.address` | `api.example.com` | Server hostname | | `server.port` | `443` | Server port | | `network.protocol.version` | `1.1` | HTTP version | | `user_agent.original` | `Mozilla/5.0...` | User agent string | #### Database Spans | Attribute | Example | Notes | |-----------|---------|-------| | `db.system` | `postgresql` | Database type | | `db.namespace` | `myapp` | Database name | | `db.operation.name` | `SELECT` | SQL operation | | `db.query.text` | `SELECT * FROM...` | Sanitized query | | `server.address` | `db.example.com` | DB host | | `server.port` | `5432` | DB port | | `db.response.rows_affected` | `12` | Rows returned/affected | #### gRPC Spans | Attribute | Example | Notes | |-----------|---------|-------| | `rpc.system` | `grpc` | RPC system | | `rpc.service` | `myapp.UserService` | Service name | | `rpc.method` | `GetUser` | Method name | | `rpc.grpc.status_code` | `0` | gRPC status code | ### Span Status | Status | When | Notes | |--------|------|-------| | `UNSET` | Default | Operation completed, no explicit status | | `OK` | Explicitly successful | Use sparingly, UNSET is fine for success | | `ERROR` | Operation failed | Always set for 5xx responses, exceptions | --- ## SDK Setup ### Go ```go package main import ( "context" "log" "go.opentelemetry.io/otel" "go.opentelemetry.io/otel/attribute" "go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc" "go.opentelemetry.io/otel/propagation" "go.opentelemetry.io/otel/sdk/resource" sdktrace "go.opentelemetry.io/otel/sdk/trace" semconv "go.opentelemetry.io/otel/semconv/v1.24.0" "go.opentelemetry.io/otel/trace" ) func initTracer(ctx context.Context) (*sdktrace.TracerProvider, error) { // OTLP gRPC exporter (sends to Collector) exporter, err := otlptracegrpc.New(ctx, otlptracegrpc.WithEndpoint("otel-collector:4317"), otlptracegrpc.WithInsecure(), ) if err != nil { return nil, err } // Resource: describes this service res, err := resource.Merge( resource.Default(), resource.NewWithAttributes( semconv.SchemaURL, semconv.ServiceName("order-service"), semconv.ServiceVersion("1.4.2"), attribute.String("environment", "production"), ), ) if err != nil { return nil, err } // TracerProvider with batch span processor tp := sdktrace.NewTracerProvider( sdktrace.WithBatcher(exporter), sdktrace.WithResource(res), sdktrace.WithSampler(sdktrace.ParentBased( sdktrace.TraceIDRatioBased(0.1), // 10% head sampling )), ) // Set global TracerProvider and propagator otel.SetTracerProvider(tp) otel.SetTextMapPropagator(propagation.NewCompositeTextMapPropagator( propagation.TraceContext{}, // W3C TraceContext propagation.Baggage{}, // W3C Baggage )) return tp, nil } func main() { ctx := context.Background() tp, err := initTracer(ctx) if err != nil { log.Fatal(err) } defer tp.Shutdown(ctx) // Create spans tracer := otel.Tracer("order-service") ctx, span := tracer.Start(ctx, "ProcessOrder", trace.WithAttributes( attribute.String("order.id", "ord-123"), attribute.Int("order.items", 3), ), ) defer span.End() // Add events span.AddEvent("order.validated", trace.WithAttributes( attribute.Bool("has_discount", true), )) // Record errors if err := processPayment(ctx); err != nil { span.RecordError(err) span.SetStatus(codes.Error, err.Error()) } } ``` ### Go Auto-instrumentation ```go import ( "go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp" "go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc" "go.opentelemetry.io/contrib/instrumentation/github.com/jackc/pgx/v5/otelpgx" ) // HTTP server: wrap handler mux := http.NewServeMux() mux.HandleFunc("/api/orders", handleOrders) handler := otelhttp.NewHandler(mux, "server") http.ListenAndServe(":8080", handler) // HTTP client: wrap transport client := &http.Client{ Transport: otelhttp.NewTransport(http.DefaultTransport), } // gRPC server: add interceptors server := grpc.NewServer( grpc.UnaryInterceptor(otelgrpc.UnaryServerInterceptor()), grpc.StreamInterceptor(otelgrpc.StreamServerInterceptor()), ) // gRPC client: add interceptors conn, _ := grpc.Dial(addr, grpc.WithUnaryInterceptor(otelgrpc.UnaryClientInterceptor()), grpc.WithStreamInterceptor(otelgrpc.StreamClientInterceptor()), ) // pgx (PostgreSQL): add tracer config, _ := pgxpool.ParseConfig(databaseURL) config.ConnConfig.Tracer = otelpgx.NewTracer() pool, _ := pgxpool.NewWithConfig(ctx, config) ``` ### Python ```python from opentelemetry import trace from opentelemetry.sdk.trace import TracerProvider from opentelemetry.sdk.trace.export import BatchSpanProcessor from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter from opentelemetry.sdk.resources import Resource, SERVICE_NAME, SERVICE_VERSION from opentelemetry.propagators.composite import CompositePropagator from opentelemetry.propagators.textmap import DefaultTextMapPropagator from opentelemetry.trace.propagation.tracecontext import TraceContextTextMapPropagator from opentelemetry.baggage.propagation import W3CBaggagePropagator # Resource resource = Resource.create({ SERVICE_NAME: "order-service", SERVICE_VERSION: "1.4.2", "environment": "production", }) # TracerProvider provider = TracerProvider(resource=resource) provider.add_span_processor( BatchSpanProcessor( OTLPSpanExporter(endpoint="otel-collector:4317", insecure=True) ) ) trace.set_tracer_provider(provider) # Propagator from opentelemetry import propagate propagate.set_global_textmap(CompositePropagator([ TraceContextTextMapPropagator(), W3CBaggagePropagator(), ])) # Create spans tracer = trace.get_tracer("order-service") with tracer.start_as_current_span("process_order", attributes={ "order.id": "ord-123", "order.items": 3, }) as span: span.add_event("order.validated", {"has_discount": True}) try: process_payment(order) except Exception as e: span.record_exception(e) span.set_status(trace.Status(trace.StatusCode.ERROR, str(e))) raise ``` ### Python Auto-instrumentation ```bash # Install auto-instrumentation packages pip install opentelemetry-distro opentelemetry-exporter-otlp opentelemetry-bootstrap -a install # Installs all detected instrumentors # Run with auto-instrumentation opentelemetry-instrument \ --service_name order-service \ --exporter_otlp_endpoint http://otel-collector:4317 \ python app.py ``` ```python # Or configure programmatically from opentelemetry.instrumentation.flask import FlaskInstrumentor from opentelemetry.instrumentation.requests import RequestsInstrumentor from opentelemetry.instrumentation.psycopg2 import Psycopg2Instrumentor from opentelemetry.instrumentation.redis import RedisInstrumentor FlaskInstrumentor().instrument_app(app) RequestsInstrumentor().instrument() Psycopg2Instrumentor().instrument() RedisInstrumentor().instrument() ``` ### Node.js ```javascript // tracing.js - import BEFORE other modules const { NodeSDK } = require('@opentelemetry/sdk-node'); const { OTLPTraceExporter } = require('@opentelemetry/exporter-trace-otlp-grpc'); const { getNodeAutoInstrumentations } = require('@opentelemetry/auto-instrumentations-node'); const { Resource } = require('@opentelemetry/resources'); const { ATTR_SERVICE_NAME, ATTR_SERVICE_VERSION } = require('@opentelemetry/semantic-conventions'); const sdk = new NodeSDK({ resource: new Resource({ [ATTR_SERVICE_NAME]: 'order-service', [ATTR_SERVICE_VERSION]: '1.4.2', environment: 'production', }), traceExporter: new OTLPTraceExporter({ url: 'http://otel-collector:4317', }), instrumentations: [ getNodeAutoInstrumentations({ // Disable fs instrumentation (too noisy) '@opentelemetry/instrumentation-fs': { enabled: false }, }), ], }); sdk.start(); // Graceful shutdown process.on('SIGTERM', () => { sdk.shutdown().then(() => process.exit(0)); }); ``` ```javascript // Manual span creation const { trace } = require('@opentelemetry/api'); const tracer = trace.getTracer('order-service'); async function processOrder(orderId) { return tracer.startActiveSpan('process_order', { attributes: { 'order.id': orderId }, }, async (span) => { try { span.addEvent('order.validated'); await processPayment(orderId); span.setStatus({ code: SpanStatusCode.OK }); } catch (err) { span.recordException(err); span.setStatus({ code: SpanStatusCode.ERROR, message: err.message }); throw err; } finally { span.end(); } }); } ``` --- ## Context Propagation ### W3C TraceContext The standard for propagating trace context across service boundaries. **Request headers:** ``` traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01 tracestate: vendor1=value1,vendor2=value2 ``` **Format:** `version-trace_id-parent_id-trace_flags` | Field | Size | Description | |-------|------|-------------| | version | 2 hex | Always `00` | | trace_id | 32 hex | Unique trace identifier | | parent_id | 16 hex | Span ID of the caller | | trace_flags | 2 hex | `01` = sampled, `00` = not sampled | ### B3 Propagation (Zipkin) Legacy format still used by some systems: ``` # Single header (compact) b3: 4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-1 # Multi-header X-B3-TraceId: 4bf92f3577b34da6a3ce929d0e0e4736 X-B3-SpanId: 00f067aa0ba902b7 X-B3-ParentSpanId: (parent span) X-B3-Sampled: 1 ``` ### Baggage Propagate arbitrary key-value pairs across service boundaries: ``` baggage: userId=usr-789,region=us-east-1 ``` ```go // Go: set baggage bag, _ := baggage.Parse("userId=usr-789,region=us-east-1") ctx = baggage.ContextWithBaggage(ctx, bag) // Go: read baggage bag := baggage.FromContext(ctx) userId := bag.Member("userId").Value() ``` **Use sparingly:** Baggage is sent with every request. Don't put large values or sensitive data in baggage. --- ## Sampling Strategies ### Head-based Sampling Decision made at trace creation, propagated to all downstream services. | Sampler | Description | Config | |---------|-------------|--------| | `AlwaysOn` | Sample everything | Development only | | `AlwaysOff` | Sample nothing | Disable tracing | | `TraceIDRatioBased` | Sample N% of traces | `ratio: 0.1` for 10% | | `ParentBased` | Follow parent's decision | Default, wrap another sampler | ```go // Go: 10% sampling, respecting parent's decision sampler := sdktrace.ParentBased( sdktrace.TraceIDRatioBased(0.1), ) ``` ```python # Python: 10% sampling from opentelemetry.sdk.trace.sampling import TraceIdRatioBased, ParentBasedTraceIdRatio sampler = ParentBasedTraceIdRatio(0.1) ``` ### Tail-based Sampling Decision made after the trace completes. Requires the OTel Collector. **Advantages:** - Always captures error traces - Always captures slow traces - More representative sampling **Configuration (Collector):** ```yaml processors: tail_sampling: decision_wait: 10s # Wait for spans to arrive num_traces: 100000 # Max traces in memory expected_new_traces_per_sec: 1000 policies: # Always keep errors - name: errors-policy type: status_code status_code: status_codes: [ERROR] # Always keep slow traces (root span > 2s) - name: latency-policy type: latency latency: threshold_ms: 2000 # Keep all traces for specific operations - name: critical-operations type: string_attribute string_attribute: key: operation values: [payment, refund, account_deletion] # Sample 5% of remaining traces - name: probabilistic-policy type: probabilistic probabilistic: sampling_percentage: 5 # Composite: apply multiple policies with priority - name: composite-policy type: composite composite: max_total_spans_per_second: 1000 policy_order: [errors-policy, latency-policy, probabilistic-policy] rate_allocation: - policy: errors-policy percent: 50 - policy: latency-policy percent: 30 - policy: probabilistic-policy percent: 20 ``` --- ## Jaeger ### Deployment #### All-in-One (Development) ```yaml # docker-compose.yml services: jaeger: image: jaegertracing/all-in-one:1.54 ports: - "16686:16686" # UI - "4317:4317" # OTLP gRPC - "4318:4318" # OTLP HTTP - "14250:14250" # Jaeger gRPC environment: COLLECTOR_OTLP_ENABLED: true ``` #### Production (with Elasticsearch) ```yaml services: jaeger-collector: image: jaegertracing/jaeger-collector:1.54 environment: SPAN_STORAGE_TYPE: elasticsearch ES_SERVER_URLS: http://elasticsearch:9200 COLLECTOR_OTLP_ENABLED: true ports: - "4317:4317" - "14250:14250" jaeger-query: image: jaegertracing/jaeger-query:1.54 environment: SPAN_STORAGE_TYPE: elasticsearch ES_SERVER_URLS: http://elasticsearch:9200 ports: - "16686:16686" elasticsearch: image: elasticsearch:8.12.0 environment: discovery.type: single-node xpack.security.enabled: false ES_JAVA_OPTS: "-Xms512m -Xmx512m" volumes: - es-data:/usr/share/elasticsearch/data volumes: es-data: ``` ### Jaeger UI Features | Feature | Description | |---------|-------------| | **Search** | Find traces by service, operation, tags, duration, time range | | **Trace View** | Waterfall visualization of spans with timing | | **Compare** | Compare two traces side by side | | **Dependencies** | Service dependency graph (DAG) | | **Deep Dependency** | Trace-aware dependency analysis | | **Monitor** | RED metrics derived from traces | --- ## Async Trace Propagation ### Go: context.Context ```go // Context carries trace information automatically func processOrder(ctx context.Context, orderID string) error { // Start child span — automatically linked to parent via ctx ctx, span := tracer.Start(ctx, "processOrder") defer span.End() // Pass context to goroutines g, ctx := errgroup.WithContext(ctx) g.Go(func() error { return validateOrder(ctx, orderID) // ctx carries trace }) g.Go(func() error { return checkInventory(ctx, orderID) // ctx carries trace }) return g.Wait() } ``` ### Python: contextvars ```python import asyncio from opentelemetry import trace, context tracer = trace.get_tracer("order-service") async def process_order(order_id: str): with tracer.start_as_current_span("process_order") as span: # asyncio tasks automatically inherit context results = await asyncio.gather( validate_order(order_id), check_inventory(order_id), ) return results async def validate_order(order_id: str): # This span is automatically a child of process_order with tracer.start_as_current_span("validate_order"): pass ``` ### Node.js: AsyncLocalStorage ```javascript const { trace, context } = require('@opentelemetry/api'); // OpenTelemetry SDK uses AsyncLocalStorage internally // Spans are automatically propagated through async operations async function processOrder(orderId) { return tracer.startActiveSpan('process_order', async (span) => { try { // Promise.all preserves context const [validation, inventory] = await Promise.all([ validateOrder(orderId), // context propagated checkInventory(orderId), // context propagated ]); return { validation, inventory }; } finally { span.end(); } }); } ``` --- ## Database Query Tracing ### Span Attributes for Database Queries ```json { "name": "SELECT users", "attributes": { "db.system": "postgresql", "db.namespace": "myapp", "db.operation.name": "SELECT", "db.query.text": "SELECT id, name, email FROM users WHERE id = $1", "server.address": "db.example.com", "server.port": 5432, "db.response.rows_affected": 1 } } ``` ### Query Parameter Sanitization **Never log actual parameter values** — they may contain PII: ```go // GOOD: sanitized span.SetAttributes( attribute.String("db.query.text", "SELECT * FROM users WHERE id = $1 AND email = $2"), ) // BAD: contains PII span.SetAttributes( attribute.String("db.query.text", "SELECT * FROM users WHERE id = 789 AND email = 'john@example.com'"), ) ``` ### N+1 Query Detection Use traces to identify N+1 queries — they appear as many identical DB spans under one parent: ``` GET /api/orders (250ms) ├── SELECT orders WHERE user_id = $1 (5ms) ├── SELECT product WHERE id = $1 (3ms) ← N+1 ├── SELECT product WHERE id = $1 (3ms) ← N+1 ├── SELECT product WHERE id = $1 (4ms) ← N+1 ├── SELECT product WHERE id = $1 (3ms) ← N+1 └── ... (20 more identical queries) ``` **Fix:** Use `SELECT * FROM products WHERE id IN ($1, $2, ...)` or JOIN. --- ## HTTP Client/Server Tracing ### Automatic Instrumentation Most OpenTelemetry auto-instrumentation libraries create spans automatically for: - HTTP server requests (incoming) - HTTP client requests (outgoing) - gRPC server/client calls - Database queries - Redis operations - Message queue operations (Kafka, RabbitMQ) ### Custom Attributes on Auto-instrumented Spans ```go // Add business context to auto-created spans span := trace.SpanFromContext(ctx) span.SetAttributes( attribute.String("user.id", userID), attribute.String("order.id", orderID), attribute.String("tenant.id", tenantID), ) ``` ### Error Recording ```go // Proper error recording if err != nil { span.RecordError(err) // Creates an event with exception details span.SetStatus(codes.Error, err.Error()) // Marks span as error return err } // For HTTP handlers if statusCode >= 500 { span.SetStatus(codes.Error, fmt.Sprintf("HTTP %d", statusCode)) } // Note: 4xx is NOT an error from the server's perspective ``` --- ## gRPC Tracing ### Server Interceptors ```go import "go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc" server := grpc.NewServer( grpc.StatsHandler(otelgrpc.NewServerHandler()), ) ``` ### Client Interceptors ```go conn, err := grpc.Dial(addr, grpc.WithStatsHandler(otelgrpc.NewClientHandler()), ) ``` ### Streaming Spans For streaming RPCs, spans cover the entire stream lifecycle: ``` ServerStream (user.ListUsers) — 2.5s ├─ stream.message.sent (1) — T+10ms ├─ stream.message.sent (2) — T+50ms ├─ stream.message.sent (3) — T+120ms └─ stream.message.sent (4) — T+200ms ``` ### Metadata Propagation Context is automatically propagated via gRPC metadata when using OTel interceptors: ```go // Automatic: OTel interceptors inject/extract from gRPC metadata // metadata equivalent to HTTP headers: // traceparent → grpc-metadata-traceparent // tracestate → grpc-metadata-tracestate ``` --- ## Trace-Based Testing ### Asserting Span Structure ```go // Go: using in-memory exporter for testing import ( sdktrace "go.opentelemetry.io/otel/sdk/trace" "go.opentelemetry.io/otel/sdk/trace/tracetest" ) func TestOrderProcessing(t *testing.T) { exporter := tracetest.NewInMemoryExporter() tp := sdktrace.NewTracerProvider( sdktrace.WithSyncer(exporter), ) otel.SetTracerProvider(tp) // Run the operation processOrder(context.Background(), "ord-123") // Assert spans spans := exporter.GetSpans() assert.Len(t, spans, 3) rootSpan := spans[0] assert.Equal(t, "processOrder", rootSpan.Name) assert.Equal(t, codes.Ok, rootSpan.Status.Code) dbSpan := spans[1] assert.Equal(t, "SELECT orders", dbSpan.Name) assert.Equal(t, "postgresql", dbSpan.Attributes["db.system"]) // Verify parent-child relationship assert.Equal(t, rootSpan.SpanContext.SpanID(), dbSpan.Parent.SpanID()) } ``` ```python # Python: using in-memory exporter from opentelemetry.sdk.trace.export.in_memory import InMemorySpanExporter exporter = InMemorySpanExporter() provider = TracerProvider() provider.add_span_processor(SimpleSpanProcessor(exporter)) trace.set_tracer_provider(provider) # Run operation process_order("ord-123") # Assert spans = exporter.get_finished_spans() assert len(spans) == 3 assert spans[0].name == "process_order" assert spans[1].name == "SELECT orders" assert spans[1].parent.span_id == spans[0].context.span_id ``` ### Verifying Context Propagation ```go func TestContextPropagation(t *testing.T) { // Create a trace context ctx, span := tracer.Start(context.Background(), "test-root") traceID := span.SpanContext().TraceID() // Call service that makes outbound HTTP call handler.ServeHTTP(recorder, req.WithContext(ctx)) // Verify all spans share the same trace ID spans := exporter.GetSpans() for _, s := range spans { assert.Equal(t, traceID, s.SpanContext.TraceID()) } span.End() } ```
-
-
scripts
-
.gitkeep 0 B · in bundle
-
-
SKILL.md 16.1 KB
--- name: monitoring-ops description: "Observability patterns - metrics, logging, tracing, alerting, and infrastructure monitoring. Use for: monitoring, observability, prometheus, grafana, metrics, alerting, structured logging, distributed tracing, opentelemetry, SLO, SLI, dashboard, health check, loki, jaeger, datadog, pagerduty." license: MIT allowed-tools: "Read Write Bash" metadata: author: claude-mods related-skills: python-observability-ops, docker-ops, ci-cd-ops, nginx-ops --- # Monitoring Operations Comprehensive observability patterns covering the three pillars (metrics, logging, tracing), alerting strategies, dashboard design, and infrastructure monitoring for production systems. --- ## Three Pillars Quick Reference Use this table to decide which observability signal fits your need: | Pillar | Best For | Tools | Data Type | |--------|----------|-------|-----------| | **Metrics** | Aggregated numeric measurements, trends, alerting on thresholds | Prometheus, Datadog, CloudWatch, StatsD | Time-series (numeric) | | **Logs** | Discrete events, error details, audit trails, debugging context | Loki, ELK, CloudWatch Logs, Fluentd | Unstructured/structured text | | **Traces** | Request flow across services, latency breakdown, dependency mapping | Jaeger, Tempo, Zipkin, Datadog APM | Span trees (structured) | **When to use which:** - **"How many requests per second?"** → Metrics (counter + rate) - **"Why did this specific request fail?"** → Logs (error message + stack trace) - **"Where is the latency in this request?"** → Traces (span waterfall) - **"Is the system healthy right now?"** → Metrics (gauges + alerts) - **"What happened at 3:42 AM?"** → Logs (timestamped event search) - **"Which downstream service caused the timeout?"** → Traces (span analysis) **Correlation is key:** Connect all three by embedding `trace_id` in log entries, recording exemplars in metrics, and linking trace spans to log queries. --- ## Metrics Type Decision Tree Use this tree to select the correct metric type: ``` What are you measuring? │ ├─ A count of events that only goes up? │ └─ COUNTER │ Examples: http_requests_total, errors_total, bytes_sent_total │ Use rate() or increase() to get per-second or per-interval values │ Never use a counter's raw value — it resets on restart │ ├─ A current value that goes up AND down? │ └─ GAUGE │ Examples: temperature_celsius, active_connections, queue_depth │ Use for snapshots of current state │ Can use avg_over_time(), max_over_time() for trends │ ├─ A distribution of values (latency, size)? │ │ │ ├─ Need aggregatable quantiles across instances? │ │ └─ HISTOGRAM │ │ Examples: http_request_duration_seconds, response_size_bytes │ │ Define buckets: [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10] │ │ Use histogram_quantile() for percentiles (p50, p95, p99) │ │ Aggregatable across instances (histograms can be summed) │ │ │ └─ Need pre-calculated quantiles on a single instance? │ └─ SUMMARY │ Examples: go_gc_duration_seconds │ Pre-calculates quantiles client-side │ NOT aggregatable across instances │ Prefer histogram unless you have a specific reason │ └─ None of the above? └─ INFO metric (labels only, value=1) Examples: build_info{version="1.2.3", commit="abc123"} Use for metadata exposed as metrics ``` **Rule of thumb:** Start with counters and histograms. Add gauges for current state. Avoid summaries unless you have a compelling reason. --- ## Alerting Decision Tree ``` What type of alert do you need? │ ├─ Known threshold with a fixed boundary? │ └─ THRESHOLD-BASED │ Example: CPU > 90% for 5 minutes │ Pros: Simple, predictable, easy to understand │ Cons: Requires manual tuning, doesn't adapt to patterns │ Best for: Resource limits, error rate spikes, queue depth │ ├─ Normal behavior varies by time/season? │ └─ ANOMALY-BASED │ Example: Traffic 3 standard deviations below normal for this hour │ Pros: Adapts to patterns, catches novel failures │ Cons: Noisy during transitions, requires training data │ Best for: Traffic patterns, business metrics, gradual degradation │ └─ Defined reliability targets? └─ SLO-BASED (PREFERRED) Example: Error budget burn rate > 14.4x for 1 hour Pros: Aligned with user impact, reduces noise, principled Cons: Requires SLI/SLO definition, more complex setup Best for: User-facing services, platform reliability ``` ### Severity Levels | Severity | Response | Examples | Routing | |----------|----------|----------|---------| | **Critical (P1)** | Page on-call immediately | Service down, data loss risk, security breach | PagerDuty high-urgency, phone call | | **Warning (P2)** | Investigate within hours | Elevated error rate, disk 80% full, SLO burn rate elevated | PagerDuty low-urgency, Slack alert channel | | **Info (P3)** | Review next business day | Deployment completed, certificate expiring in 30 days | Slack info channel, ticket auto-created | ### When to Page vs When to Ticket **Page (wake someone up) when:** - Users are currently impacted - Data loss is occurring or imminent - Security incident is active - Error budget will be exhausted within hours **Create ticket (don't page) when:** - Issue is not user-facing yet - Automated remediation is possible - Degradation is slow and has runway - Issue is during business hours and can be triaged normally --- ## Structured Logging Quick Reference ### Standard JSON Log Format ```json { "timestamp": "2026-03-09T14:32:01.123Z", "level": "ERROR", "message": "Failed to process payment", "service": "payment-api", "version": "1.4.2", "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736", "span_id": "00f067aa0ba902b7", "request_id": "req-abc123", "user_id": "usr-789", "error": { "type": "PaymentGatewayTimeout", "message": "Gateway response timeout after 30s", "stack": "..." }, "duration_ms": 30042, "http": { "method": "POST", "path": "/api/v1/payments", "status_code": 504 } } ``` ### Log Level Decision Guide | Level | When to Use | Examples | |-------|-------------|---------| | **DEBUG** | Development only, verbose internal state | Variable values, SQL queries, cache hits/misses | | **INFO** | Normal operations worth recording | Request completed, job started/finished, config loaded | | **WARN** | Degraded but still functioning | Retry succeeded, fallback used, approaching limit | | **ERROR** | Operation failed, needs attention | Payment failed, API call error, constraint violation | | **FATAL** | Process cannot continue, must exit | Database unreachable at startup, invalid config, OOM | **Rules:** - Never log at ERROR for expected conditions (user input validation → WARN) - Every ERROR should be actionable — if no one will act on it, use WARN - DEBUG should be off in production by default - INFO should not be noisy — 1-5 log lines per request, not 50 ### Correlation IDs - Generate a `request_id` (UUID v4 or ULID) at the edge/gateway - Propagate through all internal services via headers (`X-Request-ID`) - Include `trace_id` and `span_id` from distributed tracing - Log all three IDs in every log entry for cross-referencing --- ## Distributed Tracing Quick Reference ### Core Concepts - **Trace:** End-to-end journey of a request across all services - **Span:** A single unit of work (HTTP call, DB query, function execution) - **Context propagation:** Passing trace/span IDs between services via headers ### W3C TraceContext Header ``` traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01 │ │ │ │ │ │ │ └─ flags (01=sampled) │ │ └─ parent span ID (16 hex) │ └─ trace ID (32 hex) └─ version (00) ``` ### Sampling Strategies | Strategy | How It Works | Use When | |----------|--------------|----------| | **Head-based (ratio)** | Decide at trace start, propagate decision | Low traffic, need predictable volume | | **Always-on** | Sample everything | Development, low-traffic services | | **Parent-based** | Follow parent's sampling decision | Default for most services | | **Tail-based** | Decide after trace completes (at Collector) | Need error/slow traces, high traffic | **Recommendation:** Use parent-based + tail-based at the Collector. This captures all error traces and slow traces while controlling volume. ### Trace ID in Logs Always include `trace_id` in structured log entries. This enables jumping from a log line to the full trace view: ``` Log entry → trace_id → Jaeger/Tempo → full request waterfall ``` --- ## Tool Selection Matrix | Feature | Prometheus + Grafana | Datadog | Grafana Cloud | CloudWatch | |---------|---------------------|---------|---------------|------------| | **Cost** | Free (infra costs) | $$$$ (per host/metric) | $$ (usage-based) | $$ (AWS-native) | | **Setup complexity** | High (self-managed) | Low (SaaS agent) | Medium (managed) | Low (AWS-native) | | **Metrics** | Prometheus (excellent) | Built-in (excellent) | Mimir (excellent) | Built-in (good) | | **Logs** | Loki (good) | Built-in (excellent) | Loki (good) | CloudWatch Logs (good) | | **Traces** | Jaeger/Tempo (good) | APM (excellent) | Tempo (good) | X-Ray (adequate) | | **Alerting** | Alertmanager (good) | Built-in (excellent) | Grafana Alerting (good) | CloudWatch Alarms (adequate) | | **Dashboards** | Grafana (excellent) | Built-in (excellent) | Grafana (excellent) | Dashboards (adequate) | | **Retention** | Configurable (unlimited) | 15 months default | Configurable | Up to 15 months | | **Multi-cloud** | Yes | Yes | Yes | AWS only | | **Best for** | Cost-conscious, control | Full-featured, enterprise | Open-source + managed | AWS-native shops | **Recommendation path:** - **Starting out / budget-conscious:** Prometheus + Grafana + Loki + Tempo (all free, self-hosted) - **Small team, want managed:** Grafana Cloud free tier (10k metrics, 50GB logs, 50GB traces) - **Enterprise, need everything:** Datadog (expensive but comprehensive) - **AWS-only shop:** CloudWatch + X-Ray (simplest if already on AWS) --- ## Dashboard Design ### USE Method (Infrastructure) For every resource (CPU, memory, disk, network): | Signal | Question | Metric Example | |--------|----------|----------------| | **Utilization** | How busy is it? | `node_cpu_seconds_total` (% busy) | | **Saturation** | How overloaded is it? | `node_load1` (run queue length) | | **Errors** | Are there error events? | `node_network_receive_errs_total` | ### RED Method (Services) For every service endpoint: | Signal | Question | Metric Example | |--------|----------|----------------| | **Rate** | How many requests per second? | `rate(http_requests_total[5m])` | | **Errors** | How many are failing? | `rate(http_requests_total{status=~"5.."}[5m])` | | **Duration** | How long do they take? | `histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))` | ### Four Golden Signals (Google SRE) | Signal | What to Measure | Alert Threshold Guidance | |--------|-----------------|--------------------------| | **Latency** | Time to serve a request (distinguish success vs error latency) | p99 > 2x baseline | | **Traffic** | Demand on the system (requests/sec, sessions, transactions) | Anomaly detection | | **Errors** | Rate of failed requests (explicit 5xx, implicit policy violations) | > 0.1% of traffic | | **Saturation** | How "full" the service is (CPU, memory, queue depth) | > 80% capacity | ### Dashboard Layout Best Practices 1. **Top row:** Key health indicators (error rate, latency p99, availability %) 2. **Second row:** Traffic and throughput (requests/sec, active users) 3. **Third row:** Resource utilization (CPU, memory, disk, network) 4. **Bottom rows:** Detailed breakdowns (by endpoint, by status code, by region) 5. **Use variables:** Service, environment, time range as dropdown selectors 6. **Include annotations:** Deployments, incidents, config changes as vertical markers --- ## Common Gotchas | Gotcha | Why It Happens | Fix | |--------|----------------|-----| | **Cardinality explosion** | Using unbounded label values (user ID, request path, query string) | Use bounded labels only; aggregate high-cardinality data in logs, not metrics | | **Alert fatigue** | Too many alerts, too sensitive thresholds, alerts on non-actionable symptoms | Require runbook for every alert; tune thresholds; use SLO-based alerting | | **Missing correlation IDs** | Logs, metrics, and traces not linked together | Include trace_id in all log entries; use exemplars in metrics | | **Sampling bias** | Head-based sampling drops error/slow traces at high sample rates | Use tail-based sampling at the Collector to always capture errors and slow traces | | **Log volume costs** | DEBUG or verbose INFO in production, logging full request/response bodies | Set production to INFO minimum; truncate large payloads; use sampling for verbose paths | | **Metric naming inconsistency** | Different teams use different naming conventions | Adopt OpenMetrics naming: `namespace_subsystem_unit_suffix` (e.g., `http_server_request_duration_seconds`) | | **Dashboard sprawl** | Everyone creates dashboards, nobody maintains them | Standardize with USE/RED templates; review quarterly; delete unused dashboards | | **SLO too aggressive** | Setting 99.99% availability without the budget or architecture for it | Start with 99.5% or 99.9%; tighten only when consistently meeting targets with margin | | **Missing baseline** | Alerting on absolute thresholds without understanding normal behavior | Collect 2-4 weeks of baseline data before setting alert thresholds | | **Over-instrumentation** | Instrumenting every function, creating too many spans/metrics | Instrument at service boundaries; use auto-instrumentation for HTTP/DB/gRPC; add manual spans selectively | | **Ignoring metric staleness** | Assuming a metric that stops reporting means zero | Use `absent()` or `up == 0` to detect missing scrapers; distinguish "zero" from "not reporting" | | **Alerting on cause not symptom** | Alerting on CPU usage instead of user-facing error rate | Alert on symptoms (error rate, latency); use cause metrics (CPU, memory) for investigation | | **No retention policy** | Storing all metrics/logs at full resolution forever | Define retention tiers: 15s resolution for 2 weeks, 1m for 3 months, 5m for 1 year | | **Dashboard without context** | Graphs with no units, no description, no threshold lines | Add units to Y-axis, threshold lines for SLOs, panel descriptions explaining what "good" looks like | --- ## Reference Files | File | Contents | Lines | |------|----------|-------| | [metrics-alerting.md](references/metrics-alerting.md) | Prometheus, Grafana, OpenTelemetry metrics, SLI/SLO/SLA, alert routing, runbooks, uptime monitoring | ~650 | | [logging.md](references/logging.md) | Structured logging, log levels, correlation IDs, aggregation (Loki, ELK), retention, PII masking, language-specific | ~550 | | [tracing.md](references/tracing.md) | OpenTelemetry, spans, context propagation, sampling, Jaeger, async tracing, DB/HTTP/gRPC instrumentation | ~600 | | [infrastructure.md](references/infrastructure.md) | Health checks, K8s probes, Docker HEALTHCHECK, infra metrics, APM, cost optimization, incident response | ~550 | --- ## See Also - **docker-ops** — Container monitoring with cAdvisor, Docker stats, and health checks - **ci-cd-ops** — Pipeline observability, deployment tracking, build metrics - **nginx-ops** — Nginx access/error log parsing, request metrics, upstream monitoring - **python-observability-ops** — Python-specific instrumentation with structlog, opentelemetry-python - [OpenTelemetry documentation](https://opentelemetry.io/docs/) - [Prometheus best practices](https://prometheus.io/docs/practices/) - [Google SRE Book — Monitoring chapter](https://sre.google/sre-book/monitoring-distributed-systems/) - [Grafana dashboards library](https://grafana.com/grafana/dashboards/)
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.