Claude Skill

monitoring-ops

Observability patterns - metrics, logging, tracing, alerting, and infrastructure monitoring. Use for: monitoring, observability, prometheus, grafana, metrics, alerting, structured logging, distributed tracing, opentelemetry, SLO, SLI, dashboard, health check, loki, jaeger, datado

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download 0xdarkmatter-claude-mods-skills_monitoring-ops-3dfaf0b.zip · 44 KB
Part of 0xdarkmatter/claude-mods — 94 skills

Install

skills CLI npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/monitoring-ops
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
Git git clone https://github.com/0xDarkMatter/claude-mods.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Monitoring Operations

Comprehensive observability patterns covering the three pillars (metrics, logging, tracing), alerting strategies, dashboard design, and infrastructure monitoring for production systems.


Three Pillars Quick Reference

Use this table to decide which observability signal fits your need:

Pillar Best For Tools Data Type
Metrics Aggregated numeric measurements, trends, alerting on thresholds Prometheus, Datadog, CloudWatch, StatsD Time-series (numeric)
Logs Discrete events, error details, audit trails, debugging context Loki, ELK, CloudWatch Logs, Fluentd Unstructured/structured text
Traces Request flow across services, latency breakdown, dependency mapping Jaeger, Tempo, Zipkin, Datadog APM Span trees (structured)

When to use which:

  • "How many requests per second?" → Metrics (counter + rate)
  • "Why did this specific request fail?" → Logs (error message + stack trace)
  • "Where is the latency in this request?" → Traces (span waterfall)
  • "Is the system healthy right now?" → Metrics (gauges + alerts)
  • "What happened at 3:42 AM?" → Logs (timestamped event search)
  • "Which downstream service caused the timeout?" → Traces (span analysis)

Correlation is key: Connect all three by embedding trace_id in log entries, recording exemplars in metrics, and linking trace spans to log queries.


Metrics Type Decision Tree

Use this tree to select the correct metric type:

What are you measuring?
│
├─ A count of events that only goes up?
│  └─ COUNTER
│     Examples: http_requests_total, errors_total, bytes_sent_total
│     Use rate() or increase() to get per-second or per-interval values
│     Never use a counter's raw value — it resets on restart
│
├─ A current value that goes up AND down?
│  └─ GAUGE
│     Examples: temperature_celsius, active_connections, queue_depth
│     Use for snapshots of current state
│     Can use avg_over_time(), max_over_time() for trends
│
├─ A distribution of values (latency, size)?
│  │
│  ├─ Need aggregatable quantiles across instances?
│  │  └─ HISTOGRAM
│  │     Examples: http_request_duration_seconds, response_size_bytes
│  │     Define buckets: [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10]
│  │     Use histogram_quantile() for percentiles (p50, p95, p99)
│  │     Aggregatable across instances (histograms can be summed)
│  │
│  └─ Need pre-calculated quantiles on a single instance?
│     └─ SUMMARY
│        Examples: go_gc_duration_seconds
│        Pre-calculates quantiles client-side
│        NOT aggregatable across instances
│        Prefer histogram unless you have a specific reason
│
└─ None of the above?
   └─ INFO metric (labels only, value=1)
      Examples: build_info{version="1.2.3", commit="abc123"}
      Use for metadata exposed as metrics

Rule of thumb: Start with counters and histograms. Add gauges for current state. Avoid summaries unless you have a compelling reason.


Alerting Decision Tree

What type of alert do you need?
│
├─ Known threshold with a fixed boundary?
│  └─ THRESHOLD-BASED
│     Example: CPU > 90% for 5 minutes
│     Pros: Simple, predictable, easy to understand
│     Cons: Requires manual tuning, doesn't adapt to patterns
│     Best for: Resource limits, error rate spikes, queue depth
│
├─ Normal behavior varies by time/season?
│  └─ ANOMALY-BASED
│     Example: Traffic 3 standard deviations below normal for this hour
│     Pros: Adapts to patterns, catches novel failures
│     Cons: Noisy during transitions, requires training data
│     Best for: Traffic patterns, business metrics, gradual degradation
│
└─ Defined reliability targets?
   └─ SLO-BASED (PREFERRED)
      Example: Error budget burn rate > 14.4x for 1 hour
      Pros: Aligned with user impact, reduces noise, principled
      Cons: Requires SLI/SLO definition, more complex setup
      Best for: User-facing services, platform reliability

Severity Levels

Severity Response Examples Routing
Critical (P1) Page on-call immediately Service down, data loss risk, security breach PagerDuty high-urgency, phone call
Warning (P2) Investigate within hours Elevated error rate, disk 80% full, SLO burn rate elevated PagerDuty low-urgency, Slack alert channel
Info (P3) Review next business day Deployment completed, certificate expiring in 30 days Slack info channel, ticket auto-created

When to Page vs When to Ticket

Page (wake someone up) when:

  • Users are currently impacted
  • Data loss is occurring or imminent
  • Security incident is active
  • Error budget will be exhausted within hours

Create ticket (don't page) when:

  • Issue is not user-facing yet
  • Automated remediation is possible
  • Degradation is slow and has runway
  • Issue is during business hours and can be triaged normally

Structured Logging Quick Reference

Standard JSON Log Format

{
  "timestamp": "2026-03-09T14:32:01.123Z",
  "level": "ERROR",
  "message": "Failed to process payment",
  "service": "payment-api",
  "version": "1.4.2",
  "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
  "span_id": "00f067aa0ba902b7",
  "request_id": "req-abc123",
  "user_id": "usr-789",
  "error": {
    "type": "PaymentGatewayTimeout",
    "message": "Gateway response timeout after 30s",
    "stack": "..."
  },
  "duration_ms": 30042,
  "http": {
    "method": "POST",
    "path": "/api/v1/payments",
    "status_code": 504
  }
}

Log Level Decision Guide

Level When to Use Examples
DEBUG Development only, verbose internal state Variable values, SQL queries, cache hits/misses
INFO Normal operations worth recording Request completed, job started/finished, config loaded
WARN Degraded but still functioning Retry succeeded, fallback used, approaching limit
ERROR Operation failed, needs attention Payment failed, API call error, constraint violation
FATAL Process cannot continue, must exit Database unreachable at startup, invalid config, OOM

Rules:

  • Never log at ERROR for expected conditions (user input validation → WARN)
  • Every ERROR should be actionable — if no one will act on it, use WARN
  • DEBUG should be off in production by default
  • INFO should not be noisy — 1-5 log lines per request, not 50

Correlation IDs

  • Generate a request_id (UUID v4 or ULID) at the edge/gateway
  • Propagate through all internal services via headers (X-Request-ID)
  • Include trace_id and span_id from distributed tracing
  • Log all three IDs in every log entry for cross-referencing

Distributed Tracing Quick Reference

Core Concepts

  • Trace: End-to-end journey of a request across all services
  • Span: A single unit of work (HTTP call, DB query, function execution)
  • Context propagation: Passing trace/span IDs between services via headers

W3C TraceContext Header

traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
              │  │                                  │                  │
              │  │                                  │                  └─ flags (01=sampled)
              │  │                                  └─ parent span ID (16 hex)
              │  └─ trace ID (32 hex)
              └─ version (00)

Sampling Strategies

Strategy How It Works Use When
Head-based (ratio) Decide at trace start, propagate decision Low traffic, need predictable volume
Always-on Sample everything Development, low-traffic services
Parent-based Follow parent's sampling decision Default for most services
Tail-based Decide after trace completes (at Collector) Need error/slow traces, high traffic

Recommendation: Use parent-based + tail-based at the Collector. This captures all error traces and slow traces while controlling volume.

Trace ID in Logs

Always include trace_id in structured log entries. This enables jumping from a log line to the full trace view:

Log entry → trace_id → Jaeger/Tempo → full request waterfall

Tool Selection Matrix

Feature Prometheus + Grafana Datadog Grafana Cloud CloudWatch
Cost Free (infra costs) \(\) (per host/metric) \((usage-based) |\) (AWS-native)
Setup complexity High (self-managed) Low (SaaS agent) Medium (managed) Low (AWS-native)
Metrics Prometheus (excellent) Built-in (excellent) Mimir (excellent) Built-in (good)
Logs Loki (good) Built-in (excellent) Loki (good) CloudWatch Logs (good)
Traces Jaeger/Tempo (good) APM (excellent) Tempo (good) X-Ray (adequate)
Alerting Alertmanager (good) Built-in (excellent) Grafana Alerting (good) CloudWatch Alarms (adequate)
Dashboards Grafana (excellent) Built-in (excellent) Grafana (excellent) Dashboards (adequate)
Retention Configurable (unlimited) 15 months default Configurable Up to 15 months
Multi-cloud Yes Yes Yes AWS only
Best for Cost-conscious, control Full-featured, enterprise Open-source + managed AWS-native shops

Recommendation path:

  • Starting out / budget-conscious: Prometheus + Grafana + Loki + Tempo (all free, self-hosted)
  • Small team, want managed: Grafana Cloud free tier (10k metrics, 50GB logs, 50GB traces)
  • Enterprise, need everything: Datadog (expensive but comprehensive)
  • AWS-only shop: CloudWatch + X-Ray (simplest if already on AWS)

Dashboard Design

USE Method (Infrastructure)

For every resource (CPU, memory, disk, network):

Signal Question Metric Example
Utilization How busy is it? node_cpu_seconds_total (% busy)
Saturation How overloaded is it? node_load1 (run queue length)
Errors Are there error events? node_network_receive_errs_total

RED Method (Services)

For every service endpoint:

Signal Question Metric Example
Rate How many requests per second? rate(http_requests_total[5m])
Errors How many are failing? rate(http_requests_total{status=~"5.."}[5m])
Duration How long do they take? histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))

Four Golden Signals (Google SRE)

Signal What to Measure Alert Threshold Guidance
Latency Time to serve a request (distinguish success vs error latency) p99 > 2x baseline
Traffic Demand on the system (requests/sec, sessions, transactions) Anomaly detection
Errors Rate of failed requests (explicit 5xx, implicit policy violations) > 0.1% of traffic
Saturation How "full" the service is (CPU, memory, queue depth) > 80% capacity

Dashboard Layout Best Practices

  1. Top row: Key health indicators (error rate, latency p99, availability %)
  2. Second row: Traffic and throughput (requests/sec, active users)
  3. Third row: Resource utilization (CPU, memory, disk, network)
  4. Bottom rows: Detailed breakdowns (by endpoint, by status code, by region)
  5. Use variables: Service, environment, time range as dropdown selectors
  6. Include annotations: Deployments, incidents, config changes as vertical markers

Common Gotchas

Gotcha Why It Happens Fix
Cardinality explosion Using unbounded label values (user ID, request path, query string) Use bounded labels only; aggregate high-cardinality data in logs, not metrics
Alert fatigue Too many alerts, too sensitive thresholds, alerts on non-actionable symptoms Require runbook for every alert; tune thresholds; use SLO-based alerting
Missing correlation IDs Logs, metrics, and traces not linked together Include trace_id in all log entries; use exemplars in metrics
Sampling bias Head-based sampling drops error/slow traces at high sample rates Use tail-based sampling at the Collector to always capture errors and slow traces
Log volume costs DEBUG or verbose INFO in production, logging full request/response bodies Set production to INFO minimum; truncate large payloads; use sampling for verbose paths
Metric naming inconsistency Different teams use different naming conventions Adopt OpenMetrics naming: namespace_subsystem_unit_suffix (e.g., http_server_request_duration_seconds)
Dashboard sprawl Everyone creates dashboards, nobody maintains them Standardize with USE/RED templates; review quarterly; delete unused dashboards
SLO too aggressive Setting 99.99% availability without the budget or architecture for it Start with 99.5% or 99.9%; tighten only when consistently meeting targets with margin
Missing baseline Alerting on absolute thresholds without understanding normal behavior Collect 2-4 weeks of baseline data before setting alert thresholds
Over-instrumentation Instrumenting every function, creating too many spans/metrics Instrument at service boundaries; use auto-instrumentation for HTTP/DB/gRPC; add manual spans selectively
Ignoring metric staleness Assuming a metric that stops reporting means zero Use absent() or up == 0 to detect missing scrapers; distinguish "zero" from "not reporting"
Alerting on cause not symptom Alerting on CPU usage instead of user-facing error rate Alert on symptoms (error rate, latency); use cause metrics (CPU, memory) for investigation
No retention policy Storing all metrics/logs at full resolution forever Define retention tiers: 15s resolution for 2 weeks, 1m for 3 months, 5m for 1 year
Dashboard without context Graphs with no units, no description, no threshold lines Add units to Y-axis, threshold lines for SLOs, panel descriptions explaining what "good" looks like

Reference Files

File Contents Lines
metrics-alerting.md Prometheus, Grafana, OpenTelemetry metrics, SLI/SLO/SLA, alert routing, runbooks, uptime monitoring ~650
logging.md Structured logging, log levels, correlation IDs, aggregation (Loki, ELK), retention, PII masking, language-specific ~550
tracing.md OpenTelemetry, spans, context propagation, sampling, Jaeger, async tracing, DB/HTTP/gRPC instrumentation ~600
infrastructure.md Health checks, K8s probes, Docker HEALTHCHECK, infra metrics, APM, cost optimization, incident response ~550

See Also

Files (claude-mods)
  • assets
    • .gitkeep 0 B · in bundle
  • references
    • infrastructure.md 28.3 KB
      # Infrastructure Monitoring Reference
      
      Comprehensive reference for health checks, infrastructure metrics, APM, cost optimization, capacity planning, and incident response.
      
      ---
      
      ## Health Checks
      
      ### Types of Health Checks
      
      | Type | Question It Answers | Failure Action |
      |------|---------------------|----------------|
      | **Liveness** | Is the process alive and not deadlocked? | Restart the process |
      | **Readiness** | Can this instance serve traffic? | Remove from load balancer |
      | **Startup** | Has the process finished initializing? | Wait (don't restart yet) |
      
      ### Implementation Patterns
      
      #### Basic Health Check Endpoint
      
      ```go
      // Go
      type HealthStatus struct {
          Status    string            `json:"status"`
          Timestamp string            `json:"timestamp"`
          Version   string            `json:"version"`
          Checks    map[string]Check  `json:"checks"`
      }
      
      type Check struct {
          Status  string `json:"status"`
          Message string `json:"message,omitempty"`
          Latency string `json:"latency,omitempty"`
      }
      
      func healthHandler(w http.ResponseWriter, r *http.Request) {
          health := HealthStatus{
              Status:    "ok",
              Timestamp: time.Now().UTC().Format(time.RFC3339),
              Version:   version,
              Checks:    make(map[string]Check),
          }
      
          // Check database
          start := time.Now()
          if err := db.PingContext(r.Context()); err != nil {
              health.Status = "degraded"
              health.Checks["database"] = Check{
                  Status:  "fail",
                  Message: err.Error(),
              }
          } else {
              health.Checks["database"] = Check{
                  Status:  "ok",
                  Latency: time.Since(start).String(),
              }
          }
      
          // Check Redis
          start = time.Now()
          if err := redis.Ping(r.Context()).Err(); err != nil {
              health.Status = "degraded"
              health.Checks["redis"] = Check{
                  Status:  "fail",
                  Message: err.Error(),
              }
          } else {
              health.Checks["redis"] = Check{
                  Status:  "ok",
                  Latency: time.Since(start).String(),
              }
          }
      
          statusCode := http.StatusOK
          if health.Status != "ok" {
              statusCode = http.StatusServiceUnavailable
          }
      
          w.Header().Set("Content-Type", "application/json")
          w.WriteHeader(statusCode)
          json.NewEncoder(w).Encode(health)
      }
      ```
      
      ```python
      # Python (FastAPI)
      from fastapi import FastAPI, Response
      from datetime import datetime, timezone
      import asyncio
      
      app = FastAPI()
      
      @app.get("/health")
      async def health_check():
          checks = {}
          status = "ok"
      
          # Database check
          try:
              start = datetime.now(timezone.utc)
              await db.execute("SELECT 1")
              checks["database"] = {
                  "status": "ok",
                  "latency_ms": (datetime.now(timezone.utc) - start).total_seconds() * 1000,
              }
          except Exception as e:
              status = "degraded"
              checks["database"] = {"status": "fail", "message": str(e)}
      
          # Redis check
          try:
              start = datetime.now(timezone.utc)
              await redis.ping()
              checks["redis"] = {
                  "status": "ok",
                  "latency_ms": (datetime.now(timezone.utc) - start).total_seconds() * 1000,
              }
          except Exception as e:
              status = "degraded"
              checks["redis"] = {"status": "fail", "message": str(e)}
      
          response_code = 200 if status == "ok" else 503
          return Response(
              content=json.dumps({
                  "status": status,
                  "timestamp": datetime.now(timezone.utc).isoformat(),
                  "checks": checks,
              }),
              status_code=response_code,
              media_type="application/json",
          )
      
      @app.get("/ready")
      async def readiness_check():
          """Readiness: can we serve traffic?"""
          try:
              await db.execute("SELECT 1")
              return {"status": "ready"}
          except Exception:
              return Response(
                  content='{"status": "not_ready"}',
                  status_code=503,
                  media_type="application/json",
              )
      
      @app.get("/live")
      async def liveness_check():
          """Liveness: is the process alive?"""
          return {"status": "alive"}
      ```
      
      #### Health Check Response Format
      
      ```json
      {
        "status": "ok",
        "timestamp": "2026-03-09T14:32:01Z",
        "version": "1.4.2",
        "checks": {
          "database": {
            "status": "ok",
            "latency_ms": 2.3
          },
          "redis": {
            "status": "ok",
            "latency_ms": 0.8
          },
          "external_api": {
            "status": "degraded",
            "message": "Elevated latency",
            "latency_ms": 850
          }
        }
      }
      ```
      
      ### Liveness vs Readiness Decision Guide
      
      ```
      Is the process able to make progress?
      ├─ No (deadlocked, OOM, infinite loop)
      │  └─ Liveness check should FAIL → container gets restarted
      │
      └─ Yes, but...
         ├─ Database is temporarily unreachable
         │  └─ Readiness FAIL, Liveness PASS → stop sending traffic, don't restart
         │
         ├─ Still loading initial data/cache
         │  └─ Startup FAIL → don't check liveness yet, wait
         │
         └─ Everything is fine
            └─ All checks PASS → serve traffic normally
      ```
      
      **Common mistake:** Making liveness depend on external dependencies (database, Redis). If the database is down, restarting the application won't help — it will cause a restart storm.
      
      ---
      
      ## Kubernetes Probes
      
      ### Configuration
      
      ```yaml
      apiVersion: apps/v1
      kind: Deployment
      metadata:
        name: api-server
      spec:
        template:
          spec:
            containers:
              - name: api
                image: api-server:1.4.2
                ports:
                  - containerPort: 8080
      
                # Startup probe: runs first, disables liveness/readiness until passing
                startupProbe:
                  httpGet:
                    path: /health
                    port: 8080
                  initialDelaySeconds: 5
                  periodSeconds: 5
                  failureThreshold: 30     # 30 * 5s = 150s max startup time
                  successThreshold: 1
      
                # Liveness probe: is the process alive?
                livenessProbe:
                  httpGet:
                    path: /live
                    port: 8080
                  initialDelaySeconds: 0    # Starts after startup probe passes
                  periodSeconds: 10
                  timeoutSeconds: 3
                  failureThreshold: 3       # 3 consecutive failures → restart
                  successThreshold: 1
      
                # Readiness probe: can it serve traffic?
                readinessProbe:
                  httpGet:
                    path: /ready
                    port: 8080
                  initialDelaySeconds: 0
                  periodSeconds: 5
                  timeoutSeconds: 3
                  failureThreshold: 3       # 3 failures → remove from Service
                  successThreshold: 1
      
                resources:
                  requests:
                    cpu: 100m
                    memory: 128Mi
                  limits:
                    cpu: 500m
                    memory: 512Mi
      ```
      
      ### Probe Types
      
      #### HTTP GET
      
      ```yaml
      livenessProbe:
        httpGet:
          path: /health
          port: 8080
          httpHeaders:
            - name: Authorization
              value: Bearer internal-token
      ```
      
      #### TCP Socket
      
      ```yaml
      # For services that don't have HTTP (databases, message brokers)
      livenessProbe:
        tcpSocket:
          port: 5432
        periodSeconds: 10
      ```
      
      #### Exec Command
      
      ```yaml
      # Run a command inside the container
      livenessProbe:
        exec:
          command:
            - /bin/sh
            - -c
            - pg_isready -U postgres
        periodSeconds: 10
      ```
      
      #### gRPC Health Check
      
      ```yaml
      # gRPC health checking protocol
      livenessProbe:
        grpc:
          port: 50051
          service: ""   # Empty string checks overall server health
        periodSeconds: 10
      ```
      
      ### Probe Configuration Guidelines
      
      | Parameter | Liveness | Readiness | Startup |
      |-----------|----------|-----------|---------|
      | `initialDelaySeconds` | 0 (use startup probe) | 0 | 5-10 |
      | `periodSeconds` | 10-15 | 5-10 | 5 |
      | `timeoutSeconds` | 3-5 | 3-5 | 3-5 |
      | `failureThreshold` | 3 | 3 | 30 (generous) |
      | `successThreshold` | 1 | 1-2 | 1 |
      
      ---
      
      ## Docker HEALTHCHECK
      
      ```dockerfile
      # Dockerfile
      FROM node:20-slim
      
      HEALTHCHECK --interval=30s --timeout=5s --retries=3 --start-period=60s \
        CMD curl -f http://localhost:8080/health || exit 1
      
      # Or with wget (no curl in alpine)
      HEALTHCHECK --interval=30s --timeout=5s --retries=3 --start-period=60s \
        CMD wget --no-verbose --tries=1 --spider http://localhost:8080/health || exit 1
      ```
      
      ### docker-compose Health Check
      
      ```yaml
      services:
        api:
          image: api-server:1.4.2
          healthcheck:
            test: ["CMD", "curl", "-f", "http://localhost:8080/health"]
            interval: 30s
            timeout: 5s
            retries: 3
            start_period: 60s
      
        worker:
          image: worker:1.2.0
          depends_on:
            api:
              condition: service_healthy
            postgres:
              condition: service_healthy
      
        postgres:
          image: postgres:16
          healthcheck:
            test: ["CMD-SHELL", "pg_isready -U postgres"]
            interval: 10s
            timeout: 5s
            retries: 5
      ```
      
      ### Health Check Parameters
      
      | Parameter | Description | Default | Recommendation |
      |-----------|-------------|---------|----------------|
      | `interval` | Time between checks | 30s | 15-30s for critical services |
      | `timeout` | Max time for check | 30s | 3-5s (fail fast) |
      | `retries` | Failures before unhealthy | 3 | 3 (avoid flapping) |
      | `start_period` | Grace period for startup | 0s | Set to max startup time |
      
      ---
      
      ## Uptime Monitoring
      
      ### Uptime Kuma Setup
      
      ```yaml
      # docker-compose.yml
      services:
        uptime-kuma:
          image: louislam/uptime-kuma:1
          restart: unless-stopped
          ports:
            - "3001:3001"
          volumes:
            - uptime-kuma-data:/app/data
          labels:
            - "traefik.enable=true"
            - "traefik.http.routers.uptime.rule=Host(`status.example.com`)"
      
      volumes:
        uptime-kuma-data:
      ```
      
      **Monitor types supported:**
      - HTTP(s) — status code, keyword, response time
      - TCP — port open check
      - DNS — resolution check
      - Docker container — running status
      - gRPC — health check protocol
      - MQTT — broker connectivity
      - Ping (ICMP) — network reachability
      - Push — heartbeat endpoint (service pushes to Uptime Kuma)
      
      ### Synthetic Monitoring
      
      Scripted checks that simulate real user behavior from multiple regions:
      
      ```javascript
      // k6 script for synthetic monitoring
      import { check, sleep } from 'k6';
      import http from 'k6/http';
      
      export const options = {
        scenarios: {
          synthetic: {
            executor: 'constant-vus',
            vus: 1,
            duration: '24h',
            gracefulStop: '0s',
          },
        },
        thresholds: {
          http_req_duration: ['p(95)<500'],    // 95% under 500ms
          http_req_failed: ['rate<0.01'],       // < 1% failure rate
          checks: ['rate>0.99'],                // 99% checks pass
        },
      };
      
      export default function () {
        // Check homepage
        let res = http.get('https://www.example.com');
        check(res, {
          'homepage status 200': (r) => r.status === 200,
          'homepage loads fast': (r) => r.timings.duration < 500,
          'homepage has title': (r) => r.body.includes('<title>'),
        });
      
        // Check API health
        res = http.get('https://api.example.com/health');
        check(res, {
          'api health 200': (r) => r.status === 200,
          'api reports ok': (r) => JSON.parse(r.body).status === 'ok',
        });
      
        // Check login flow
        res = http.post('https://api.example.com/auth/login', JSON.stringify({
          email: 'synthetic-user@example.com',
          password: process.env.SYNTHETIC_PASSWORD,
        }), { headers: { 'Content-Type': 'application/json' } });
        check(res, {
          'login succeeds': (r) => r.status === 200,
          'login returns token': (r) => JSON.parse(r.body).token !== undefined,
        });
      
        sleep(60); // Check every 60 seconds
      }
      ```
      
      ### Multi-Region Monitoring
      
      | Provider | Regions | Free Tier | Notes |
      |----------|---------|-----------|-------|
      | **Uptime Kuma** | Self-hosted (1 region) | Free | Deploy in multiple regions yourself |
      | **Betteruptime** | 10+ regions | 5 monitors | Status page included |
      | **Grafana Synthetic** | 20+ regions | Part of Grafana Cloud | k6-based scripts |
      | **Datadog Synthetic** | 100+ locations | 100 API tests/month | Full browser testing |
      | **AWS CloudWatch Synthetics** | All AWS regions | Pay per run | Canary scripts |
      
      ---
      
      ## Infrastructure Metrics
      
      ### CPU Metrics
      
      | Metric | Source | What It Shows |
      |--------|--------|---------------|
      | `node_cpu_seconds_total{mode="user"}` | node_exporter | Time in user space |
      | `node_cpu_seconds_total{mode="system"}` | node_exporter | Time in kernel space |
      | `node_cpu_seconds_total{mode="iowait"}` | node_exporter | Time waiting for I/O |
      | `node_cpu_seconds_total{mode="idle"}` | node_exporter | Idle time |
      | `node_load1` / `node_load5` / `node_load15` | node_exporter | Load average (1/5/15 min) |
      
      **Common queries:**
      
      ```promql
      # CPU usage percentage (all modes except idle)
      1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))
      
      # CPU usage by mode
      sum by (mode) (rate(node_cpu_seconds_total{instance="web01:9100"}[5m]))
      
      # IO wait percentage (high = disk bottleneck)
      avg by (instance) (rate(node_cpu_seconds_total{mode="iowait"}[5m]))
      
      # Load average vs CPU count
      node_load1 / count without (cpu) (node_cpu_seconds_total{mode="idle"})
      ```
      
      ### Memory Metrics
      
      | Metric | What It Shows |
      |--------|---------------|
      | `node_memory_MemTotal_bytes` | Total physical memory |
      | `node_memory_MemAvailable_bytes` | Memory available for applications |
      | `node_memory_Cached_bytes` | Page cache (reclaimable) |
      | `node_memory_Buffers_bytes` | Buffer cache |
      | `node_memory_SwapTotal_bytes` | Total swap |
      | `node_memory_SwapFree_bytes` | Free swap |
      
      ```promql
      # Memory usage percentage
      1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)
      
      # Memory breakdown
      node_memory_MemTotal_bytes
        - node_memory_MemAvailable_bytes
        - node_memory_Cached_bytes
        - node_memory_Buffers_bytes
      
      # Swap usage (any swap usage may indicate memory pressure)
      1 - (node_memory_SwapFree_bytes / node_memory_SwapTotal_bytes)
      ```
      
      ### Disk Metrics
      
      ```promql
      # Disk usage percentage
      1 - (node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"} / node_filesystem_size_bytes)
      
      # Disk I/O utilization (percentage of time doing I/O)
      rate(node_disk_io_time_seconds_total[5m])
      
      # Read/write throughput
      rate(node_disk_read_bytes_total[5m])
      rate(node_disk_written_bytes_total[5m])
      
      # IOPS
      rate(node_disk_reads_completed_total[5m])
      rate(node_disk_writes_completed_total[5m])
      
      # Average I/O latency
      rate(node_disk_read_time_seconds_total[5m]) / rate(node_disk_reads_completed_total[5m])
      ```
      
      ### Network Metrics
      
      ```promql
      # Bandwidth (bytes/sec)
      rate(node_network_receive_bytes_total{device!="lo"}[5m])
      rate(node_network_transmit_bytes_total{device!="lo"}[5m])
      
      # Packet errors
      rate(node_network_receive_errs_total[5m])
      rate(node_network_transmit_errs_total[5m])
      
      # TCP connections
      node_netstat_Tcp_CurrEstab         # Current established connections
      rate(node_netstat_Tcp_ActiveOpens[5m])  # New outbound connections/sec
      rate(node_netstat_Tcp_PassiveOpens[5m]) # New inbound connections/sec
      ```
      
      ---
      
      ## Container Metrics
      
      ### cAdvisor Metrics
      
      | Metric | Description |
      |--------|-------------|
      | `container_cpu_usage_seconds_total` | Total CPU time consumed |
      | `container_cpu_cfs_throttled_periods_total` | CPU throttling events |
      | `container_memory_working_set_bytes` | Current memory (excludes cache) |
      | `container_memory_usage_bytes` | Total memory (includes cache) |
      | `container_network_receive_bytes_total` | Network inbound bytes |
      | `container_network_transmit_bytes_total` | Network outbound bytes |
      | `container_fs_usage_bytes` | Container filesystem usage |
      | `container_spec_memory_limit_bytes` | Memory limit |
      | `container_spec_cpu_quota` | CPU quota |
      
      ```promql
      # Container CPU usage percentage (of limit)
      sum by (container, pod) (
        rate(container_cpu_usage_seconds_total{container!="POD",container!=""}[5m])
      ) / sum by (container, pod) (
        container_spec_cpu_quota / container_spec_cpu_period
      )
      
      # Container memory usage percentage (of limit)
      container_memory_working_set_bytes{container!="POD",container!=""}
      /
      container_spec_memory_limit_bytes{container!="POD",container!=""} > 0
      
      # CPU throttling percentage
      sum by (container, pod) (
        rate(container_cpu_cfs_throttled_periods_total[5m])
      ) / sum by (container, pod) (
        rate(container_cpu_cfs_periods_total[5m])
      )
      
      # OOMKill detection
      increase(kube_pod_container_status_restarts_total[1h]) > 0
      and
      kube_pod_container_status_last_terminated_reason{reason="OOMKilled"}
      ```
      
      ### Kubernetes Metrics (kube-state-metrics)
      
      ```promql
      # Pod status
      kube_pod_status_phase{phase="Running"}
      kube_pod_status_phase{phase="Pending"}
      kube_pod_status_phase{phase="Failed"}
      
      # Deployment replicas
      kube_deployment_status_replicas_available
      kube_deployment_spec_replicas
      
      # HPA status
      kube_horizontalpodautoscaler_status_current_replicas
      kube_horizontalpodautoscaler_spec_max_replicas
      ```
      
      ---
      
      ## Node Exporter
      
      ### Setup
      
      ```yaml
      # docker-compose.yml
      services:
        node-exporter:
          image: prom/node-exporter:v1.7.0
          restart: unless-stopped
          ports:
            - "9100:9100"
          volumes:
            - /proc:/host/proc:ro
            - /sys:/host/sys:ro
            - /:/rootfs:ro
          command:
            - '--path.procfs=/host/proc'
            - '--path.sysfs=/host/sys'
            - '--path.rootfs=/rootfs'
            - '--collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host|etc)($$|/)'
      ```
      
      ### Kubernetes DaemonSet
      
      ```yaml
      apiVersion: apps/v1
      kind: DaemonSet
      metadata:
        name: node-exporter
        namespace: monitoring
      spec:
        selector:
          matchLabels:
            app: node-exporter
        template:
          metadata:
            labels:
              app: node-exporter
            annotations:
              prometheus.io/scrape: "true"
              prometheus.io/port: "9100"
          spec:
            hostPID: true
            hostNetwork: true
            containers:
              - name: node-exporter
                image: prom/node-exporter:v1.7.0
                ports:
                  - containerPort: 9100
                    hostPort: 9100
                volumeMounts:
                  - name: proc
                    mountPath: /host/proc
                    readOnly: true
                  - name: sys
                    mountPath: /host/sys
                    readOnly: true
            volumes:
              - name: proc
                hostPath:
                  path: /proc
              - name: sys
                hostPath:
                  path: /sys
            tolerations:
              - effect: NoSchedule
                operator: Exists
      ```
      
      ---
      
      ## APM Tools
      
      ### Comparison
      
      | Feature | Datadog APM | New Relic | Elastic APM | Sentry |
      |---------|-------------|-----------|-------------|--------|
      | **Type** | Full APM | Full APM | Full APM | Error tracking + perf |
      | **Pricing** | Per host ($31+/mo) | Per user + data | Free (self-host) or Cloud | Per event volume |
      | **Traces** | Yes | Yes | Yes | Transaction traces |
      | **Error tracking** | Yes | Yes | Yes | Excellent |
      | **Profiling** | Yes (continuous) | Yes | No | No |
      | **Log correlation** | Yes | Yes | Yes | Breadcrumbs |
      | **Dashboards** | Built-in | Built-in | Kibana | Limited |
      | **Setup** | Agent-based | Agent-based | Agent or OTel | SDK-based |
      | **Best for** | Enterprise, full stack | Full observability | Self-hosted, ELK users | Error-focused teams |
      
      ### Sentry Error Tracking
      
      ```python
      # Python
      import sentry_sdk
      from sentry_sdk.integrations.fastapi import FastApiIntegration
      
      sentry_sdk.init(
          dsn="https://key@sentry.io/project",
          traces_sample_rate=0.1,  # 10% of transactions
          profiles_sample_rate=0.1,
          environment="production",
          release="1.4.2",
          integrations=[FastApiIntegration()],
      )
      ```
      
      ```javascript
      // Node.js
      const Sentry = require('@sentry/node');
      
      Sentry.init({
        dsn: 'https://key@sentry.io/project',
        tracesSampleRate: 0.1,
        environment: 'production',
        release: '1.4.2',
      });
      ```
      
      ```go
      // Go
      import "github.com/getsentry/sentry-go"
      
      sentry.Init(sentry.ClientOptions{
          Dsn:              "https://key@sentry.io/project",
          TracesSampleRate: 0.1,
          Environment:      "production",
          Release:          "1.4.2",
      })
      defer sentry.Flush(2 * time.Second)
      ```
      
      ---
      
      ## Cost Optimization
      
      ### Metric Cardinality Review
      
      High cardinality is the most common cost driver in metrics systems:
      
      ```promql
      # Find metrics with the most time series
      topk(20, count by (__name__) ({__name__=~".+"}))
      
      # Find labels with high cardinality
      count(group by (path) (http_requests_total))   # How many unique paths?
      count(group by (user_id) (api_calls_total))    # Unbounded!
      ```
      
      **Reduction strategies:**
      1. Remove unused metrics (if nobody dashboards/alerts on it, drop it)
      2. Replace high-cardinality labels with bounded categories
      3. Use recording rules to pre-aggregate, drop raw metrics
      4. Use metric relabeling in Prometheus to drop at scrape time
      
      ```yaml
      # Drop unused metrics at scrape time
      metric_relabel_configs:
        - source_labels: [__name__]
          regex: "go_.*"           # Drop Go runtime metrics if unused
          action: drop
      ```
      
      ### Log Volume Reduction
      
      | Strategy | Savings | Implementation |
      |----------|---------|----------------|
      | Set production to INFO | 50-80% | Logger config |
      | Sample health check logs | 90% for /health | Middleware filter |
      | Truncate large payloads | 20-40% | Body size limit (4KB) |
      | Drop duplicate errors | 30-50% | Rate-limit per error type |
      | Compress in transit | 60-80% bandwidth | Enable gzip on log shipper |
      
      ### Trace Sampling
      
      | Sampling Rate | Monthly Cost (est.) | Suitability |
      |---------------|---------------------|-------------|
      | 100% | $$$$ | Development, < 100 req/s |
      | 10% | $$$ | Staging, medium traffic |
      | 1% | $$ | Production, high traffic |
      | Tail-based (errors + slow) | $$ | Production (recommended) |
      | 0.1% | $ | Very high traffic (> 100k req/s) |
      
      ### Retention Tiers
      
      | Tier | Metrics | Logs | Traces |
      |------|---------|------|--------|
      | Hot (0-14 days) | 15s resolution | Full fidelity | All sampled traces |
      | Warm (14-90 days) | 1m resolution | Full fidelity | Error + slow traces only |
      | Cold (90 days - 1 year) | 5m resolution | Compressed | None (rely on metrics) |
      | Archive (1-7 years) | 1h resolution | Compliance logs only | None |
      
      ---
      
      ## Capacity Planning
      
      ### Load Testing Correlation
      
      Run load tests while monitoring infrastructure metrics to establish scaling thresholds:
      
      ```
      Load Test Results:
      ┌─────────┬──────────┬────────┬─────────┬──────────────┐
      │ RPS     │ p99 (ms) │ CPU %  │ Mem %   │ Error Rate   │
      ├─────────┼──────────┼────────┼─────────┼──────────────┤
      │ 100     │ 45       │ 15     │ 30      │ 0%           │
      │ 500     │ 85       │ 35     │ 45      │ 0%           │
      │ 1000    │ 150      │ 55     │ 55      │ 0%           │
      │ 2000    │ 320      │ 75     │ 65      │ 0.1%         │
      │ 3000    │ 850      │ 90     │ 72      │ 1.5%         │  ← degradation
      │ 4000    │ 2500     │ 98     │ 78      │ 12%          │  ← failure
      └─────────┴──────────┴────────┴─────────┴──────────────┘
      
      Scaling trigger: 75% CPU → add instance
      Target capacity: 2x expected peak traffic
      ```
      
      ### Scaling Triggers
      
      ```yaml
      # Kubernetes HPA
      apiVersion: autoscaling/v2
      kind: HorizontalPodAutoscaler
      metadata:
        name: api-server
      spec:
        scaleTargetRef:
          apiVersion: apps/v1
          kind: Deployment
          name: api-server
        minReplicas: 3
        maxReplicas: 20
        metrics:
          - type: Resource
            resource:
              name: cpu
              target:
                type: Utilization
                averageUtilization: 70    # Scale up at 70% CPU
          - type: Resource
            resource:
              name: memory
              target:
                type: Utilization
                averageUtilization: 80
          - type: Pods
            pods:
              metric:
                name: http_requests_per_second
              target:
                type: AverageValue
                averageValue: "1000"       # Scale at 1000 RPS per pod
        behavior:
          scaleUp:
            stabilizationWindowSeconds: 60
            policies:
              - type: Percent
                value: 50                  # Max 50% increase per scale-up
                periodSeconds: 60
          scaleDown:
            stabilizationWindowSeconds: 300  # Wait 5 min before scaling down
            policies:
              - type: Percent
                value: 25
                periodSeconds: 120
      ```
      
      ### Resource Forecasting
      
      ```promql
      # Predict disk full in N hours
      predict_linear(node_filesystem_avail_bytes[7d], 30*24*3600) < 0
      # "Disk will be full within 30 days"
      
      # Predict memory usage trend
      predict_linear(
        avg_over_time(container_memory_working_set_bytes[7d]),
        30*24*3600
      )
      
      # Growth rate of database size
      rate(pg_database_size_bytes[7d])
      # Convert to "GB per month"
      rate(pg_database_size_bytes[7d]) * 86400 * 30 / 1e9
      ```
      
      ---
      
      ## Incident Response
      
      ### Incident Lifecycle
      
      ```
      Detection → Triage → Mitigate → Resolve → Postmortem
          │          │         │          │          │
          │          │         │          │          └─ Blameless review
          │          │         │          └─ Root cause fix deployed
          │          │         └─ User impact reduced/eliminated
          │          └─ Severity assigned, team engaged
          └─ Alert fires or user reports issue
      ```
      
      ### Severity Classification
      
      | Severity | Impact | Response Time | Examples |
      |----------|--------|---------------|---------|
      | **SEV1 (Critical)** | Service down, data loss, security breach | < 15 minutes | Complete outage, payment processing failure |
      | **SEV2 (Major)** | Significant degradation, partial outage | < 30 minutes | One region down, 50%+ error rate |
      | **SEV3 (Minor)** | Limited impact, workaround exists | < 4 hours | Single feature broken, elevated latency |
      | **SEV4 (Low)** | Minimal impact, cosmetic | Next business day | UI glitch, non-critical alert firing |
      
      ### Incident Commander Checklist
      
      ```markdown
      ## Initial Response (first 15 minutes)
      - [ ] Acknowledge the alert / report
      - [ ] Assess severity (SEV1-4)
      - [ ] Open incident channel (#inc-YYYYMMDD-description)
      - [ ] Page relevant team members
      - [ ] Post initial status update
      
      ## Triage (15-30 minutes)
      - [ ] Identify affected services and scope
      - [ ] Check recent deployments: any changes in last 2 hours?
      - [ ] Check dashboards for anomalies
      - [ ] Check external dependencies (status pages)
      - [ ] Determine if rollback is feasible
      
      ## Mitigation
      - [ ] Implement immediate fix (rollback, feature flag, scaling)
      - [ ] Verify user impact is reduced
      - [ ] Update status page
      - [ ] Communicate ETA for full resolution
      
      ## Resolution
      - [ ] Confirm root cause
      - [ ] Deploy fix
      - [ ] Verify metrics return to baseline
      - [ ] Clear incident status
      - [ ] Schedule postmortem within 48 hours
      ```
      
      ### Postmortem Template
      
      ```markdown
      # Incident Postmortem: [TITLE]
      
      **Date:** 2026-03-09
      **Duration:** 45 minutes (14:15 - 15:00 UTC)
      **Severity:** SEV2
      **Author:** [Name]
      **Status:** Complete
      
      ## Summary
      One-paragraph description of what happened and impact.
      
      ## Impact
      - Users affected: ~5,000
      - Revenue impact: ~$2,500
      - SLO budget consumed: 3.2 hours of the monthly 43-minute budget
      
      ## Timeline (all times UTC)
      | Time | Event |
      |------|-------|
      | 14:12 | Deploy v1.4.3 to production |
      | 14:15 | Error rate alert fires (5% → 15%) |
      | 14:17 | On-call acknowledges, starts investigation |
      | 14:22 | Root cause identified: new query missing index |
      | 14:25 | Decision: rollback v1.4.3 |
      | 14:30 | Rollback complete |
      | 14:35 | Error rate returns to baseline |
      | 15:00 | All-clear declared |
      
      ## Root Cause
      The v1.4.3 deployment added a new API endpoint that queried the orders
      table without an index on `user_id + created_at`. Under load, this caused
      connection pool exhaustion, which cascaded to other endpoints.
      
      ## Detection
      Alert fired 3 minutes after deploy. Detection was effective.
      
      ## Contributing Factors
      1. No load test for the new endpoint
      2. Missing index not caught in code review
      3. No query performance checks in CI
      
      ## Action Items
      | Action | Owner | Due | Status |
      |--------|-------|-----|--------|
      | Add index on orders(user_id, created_at) | @backend | 2026-03-10 | Done |
      | Add slow query detection to CI pipeline | @platform | 2026-03-15 | TODO |
      | Add load test for new endpoints to deploy checklist | @backend | 2026-03-12 | TODO |
      | Set up query performance alerting (> 100ms avg) | @sre | 2026-03-14 | TODO |
      
      ## Lessons Learned
      - What went well: Fast detection (3 min), fast rollback (8 min)
      - What went poorly: No pre-production load test caught the issue
      - Where we got lucky: Happened during business hours, not at 3 AM
      ```
      
      ### Communication During Incidents
      
      | Audience | Channel | Frequency | Content |
      |----------|---------|-----------|---------|
      | Engineering | Slack #incident | Real-time | Technical details, commands run |
      | Management | Slack #incidents-summary | Every 15-30 min | Impact, ETA, escalation needs |
      | Customers | Status page | Every 15-30 min | User-facing impact, workarounds |
      | Support | Slack #support-escalation | On status change | Scripted responses, known workarounds |
      
    • logging.md 26.3 KB
      # Logging Reference
      
      Comprehensive reference for structured logging, log aggregation, correlation, and language-specific implementations.
      
      ---
      
      ## Structured Logging
      
      ### Why Structured Logging
      
      Unstructured logs are human-readable but machine-hostile:
      
      ```
      # BAD: unstructured
      2026-03-09 14:32:01 ERROR Failed to process payment for user 789: timeout after 30s
      
      # GOOD: structured JSON
      {"timestamp":"2026-03-09T14:32:01.123Z","level":"ERROR","message":"Failed to process payment","user_id":"789","error":"timeout after 30s","duration_ms":30042}
      ```
      
      Structured logs enable:
      - Machine parsing and indexing
      - Filtering by any field (`user_id=789`, `level=ERROR`)
      - Aggregation and metric extraction
      - Correlation with traces via `trace_id`
      
      ### Standard JSON Log Format
      
      ```json
      {
        "timestamp": "2026-03-09T14:32:01.123Z",
        "level": "ERROR",
        "message": "Failed to process payment",
        "logger": "payment.processor",
        "service": "payment-api",
        "version": "1.4.2",
        "environment": "production",
        "host": "payment-api-7b4d9f-x2k9l",
        "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
        "span_id": "00f067aa0ba902b7",
        "request_id": "req-abc123",
        "user_id": "usr-789",
        "error": {
          "type": "PaymentGatewayTimeout",
          "message": "Gateway response timeout after 30s",
          "stack": "PaymentGatewayTimeout: Gateway response timeout...\n  at processPayment (payment.go:142)\n  at handleRequest (handler.go:87)"
        },
        "context": {
          "payment_id": "pay-456",
          "amount_cents": 2500,
          "currency": "USD",
          "gateway": "stripe"
        }
      }
      ```
      
      ### Key Conventions
      
      | Field | Type | Required | Notes |
      |-------|------|----------|-------|
      | `timestamp` | ISO 8601 string | Yes | Always UTC, millisecond precision |
      | `level` | string | Yes | DEBUG, INFO, WARN, ERROR, FATAL |
      | `message` | string | Yes | Human-readable, no variable interpolation in the key |
      | `service` | string | Yes | Service name (matches Prometheus job label) |
      | `version` | string | Yes | Application version or git SHA |
      | `trace_id` | string | When available | OpenTelemetry trace ID (32 hex chars) |
      | `span_id` | string | When available | OpenTelemetry span ID (16 hex chars) |
      | `request_id` | string | When available | Edge-generated request ID |
      | `error` | object | On errors | Include type, message, stack |
      | `logger` | string | Recommended | Logger name / module path |
      | `host` | string | Recommended | Hostname or pod name |
      | `environment` | string | Recommended | production, staging, development |
      
      ---
      
      ## Log Levels
      
      ### Decision Guide
      
      ```
      Is the process unable to continue?
      ├─ Yes → FATAL
      │        Process must exit. Database unreachable at startup,
      │        invalid critical config, out of memory.
      │
      └─ No → Did an operation fail?
               ├─ Yes → Is it actionable?
               │        ├─ Yes → ERROR
               │        │        Payment failed, API call returned 500,
               │        │        constraint violation, file not found.
               │        │
               │        └─ No  → WARN
               │                 Expected failure, retry will handle it,
               │                 deprecated API used, nearing limit.
               │
               └─ No  → Is it worth recording in production?
                         ├─ Yes → INFO
                         │        Request handled, job completed,
                         │        config loaded, connection established.
                         │
                         └─ No  → DEBUG
                                  Variable values, SQL queries,
                                  cache hit/miss, internal state.
      ```
      
      ### Level Details
      
      #### FATAL
      
      ```json
      {"level":"FATAL","message":"Cannot connect to database","error":{"type":"ConnectionRefused","message":"dial tcp 10.0.0.5:5432: connect: connection refused"},"action":"process_exit"}
      ```
      
      - Process cannot start or must terminate
      - Always followed by `os.Exit(1)` or equivalent
      - Should trigger immediate alerting
      - Very rare in well-designed systems
      
      #### ERROR
      
      ```json
      {"level":"ERROR","message":"Payment processing failed","payment_id":"pay-456","user_id":"usr-789","error":{"type":"GatewayTimeout","message":"Stripe API timeout after 30s"}}
      ```
      
      - Operation failed and cannot be completed
      - Someone should investigate (now or soon)
      - Every ERROR should have an associated alert or dashboard
      - **Not for:** User input validation failures (that's WARN or INFO)
      
      #### WARN
      
      ```json
      {"level":"WARN","message":"Circuit breaker opened for payment gateway","gateway":"stripe","failure_count":5,"retry_after":"30s"}
      ```
      
      - System is degraded but still functioning
      - Worth monitoring but not necessarily immediate action
      - Retry succeeded, fallback activated, approaching a limit
      - **Not for:** Expected user errors (wrong password → INFO)
      
      #### INFO
      
      ```json
      {"level":"INFO","message":"Request completed","method":"GET","path":"/api/users","status":200,"duration_ms":45,"request_id":"req-abc123"}
      ```
      
      - Normal operation, audit trail, business events
      - Should not be noisy (aim for 1-5 lines per request)
      - Deployments, configuration changes, job completions
      - **Not for:** Debugging details (use DEBUG)
      
      #### DEBUG
      
      ```json
      {"level":"DEBUG","message":"Cache lookup","key":"user:789","hit":true,"ttl_remaining_ms":45200}
      ```
      
      - Development and troubleshooting only
      - Disabled in production by default
      - Enable per-service or per-module when debugging
      - SQL queries, cache operations, internal state
      
      ### Production Log Level Strategy
      
      ```
      Production default:  INFO
      Production debug:    DEBUG (per-service, time-limited, via config change)
      Staging:             DEBUG
      Development:         DEBUG
      CI/Test:             WARN (reduce noise in test output)
      ```
      
      ---
      
      ## Correlation IDs
      
      ### Generating IDs
      
      ```go
      // Go: UUID v4
      import "github.com/google/uuid"
      requestID := uuid.New().String()  // "550e8400-e29b-41d4-a716-446655440000"
      
      // Go: ULID (sortable, timestamp-prefixed)
      import "github.com/oklog/ulid/v2"
      requestID := ulid.Make().String()  // "01ARZ3NDEKTSV4RRFFQ69G5FAV"
      ```
      
      ```python
      # Python: UUID v4
      import uuid
      request_id = str(uuid.uuid4())
      
      # Python: ULID
      import ulid
      request_id = str(ulid.new())
      ```
      
      ```javascript
      // Node.js: UUID v4
      import { randomUUID } from 'crypto';
      const requestId = randomUUID();
      
      // Node.js: ULID
      import { ulid } from 'ulid';
      const requestId = ulid();
      ```
      
      ### Propagation Middleware
      
      #### Go (net/http)
      
      ```go
      func correlationMiddleware(next http.Handler) http.Handler {
          return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
              // Get or generate request ID
              requestID := r.Header.Get("X-Request-ID")
              if requestID == "" {
                  requestID = uuid.New().String()
              }
      
              // Get trace context from OpenTelemetry
              span := trace.SpanFromContext(r.Context())
              traceID := span.SpanContext().TraceID().String()
              spanID := span.SpanContext().SpanID().String()
      
              // Add to context
              ctx := context.WithValue(r.Context(), "request_id", requestID)
      
              // Add to response headers
              w.Header().Set("X-Request-ID", requestID)
      
              // Add to logger context
              logger := slog.With(
                  "request_id", requestID,
                  "trace_id", traceID,
                  "span_id", spanID,
              )
              ctx = context.WithValue(ctx, "logger", logger)
      
              next.ServeHTTP(w, r.WithContext(ctx))
          })
      }
      ```
      
      #### Python (FastAPI)
      
      ```python
      import uuid
      from contextvars import ContextVar
      from fastapi import FastAPI, Request
      from starlette.middleware.base import BaseHTTPMiddleware
      
      request_id_var: ContextVar[str] = ContextVar("request_id", default="")
      
      class CorrelationMiddleware(BaseHTTPMiddleware):
          async def dispatch(self, request: Request, call_next):
              request_id = request.headers.get("X-Request-ID", str(uuid.uuid4()))
              request_id_var.set(request_id)
      
              response = await call_next(request)
              response.headers["X-Request-ID"] = request_id
              return response
      
      app = FastAPI()
      app.add_middleware(CorrelationMiddleware)
      ```
      
      #### Node.js (Express)
      
      ```javascript
      import { randomUUID } from 'crypto';
      import { AsyncLocalStorage } from 'async_hooks';
      
      const asyncLocalStorage = new AsyncLocalStorage();
      
      function correlationMiddleware(req, res, next) {
        const requestId = req.headers['x-request-id'] || randomUUID();
        res.setHeader('X-Request-ID', requestId);
      
        asyncLocalStorage.run({ requestId }, () => {
          next();
        });
      }
      
      // Access anywhere in the request lifecycle
      function getRequestId() {
        return asyncLocalStorage.getStore()?.requestId || 'unknown';
      }
      ```
      
      ### HTTP Client Propagation
      
      Always forward correlation IDs when making outbound HTTP calls:
      
      ```go
      // Go
      req, _ := http.NewRequestWithContext(ctx, "GET", url, nil)
      req.Header.Set("X-Request-ID", getRequestID(ctx))
      // OpenTelemetry propagation is automatic with instrumented HTTP client
      ```
      
      ```python
      # Python
      headers = {"X-Request-ID": request_id_var.get()}
      response = httpx.get(url, headers=headers)
      ```
      
      ---
      
      ## Request Context Logging
      
      ### Standard Request/Response Log
      
      ```go
      func loggingMiddleware(next http.Handler) http.Handler {
          return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
              start := time.Now()
              wrapped := &responseWriter{ResponseWriter: w, statusCode: 200}
      
              next.ServeHTTP(wrapped, r)
      
              duration := time.Since(start)
              logger.InfoContext(r.Context(), "Request completed",
                  "method", r.Method,
                  "path", r.URL.Path,
                  "status", wrapped.statusCode,
                  "duration_ms", duration.Milliseconds(),
                  "bytes_written", wrapped.bytesWritten,
                  "remote_addr", r.RemoteAddr,
                  "user_agent", r.UserAgent(),
              )
          })
      }
      ```
      
      ### What to Log Per Request
      
      | Field | When | Notes |
      |-------|------|-------|
      | Method, path, status, duration | Always | Core request metadata |
      | Request ID, trace ID | Always | Correlation |
      | User ID | When authenticated | For audit trail |
      | Request body | Selectively | Only for mutations, with size limit |
      | Response body | Rarely | Only for debugging, never in production |
      | Query parameters | When relevant | Sanitize sensitive params |
      | IP address | For security | Respect privacy regulations |
      | User-Agent | For analytics | Browser/client identification |
      
      ### Body Size Limits
      
      ```go
      // Never log unbounded request/response bodies
      const maxBodyLogSize = 4096 // 4KB
      
      func truncateBody(body []byte) string {
          if len(body) > maxBodyLogSize {
              return string(body[:maxBodyLogSize]) + "... [truncated]"
          }
          return string(body)
      }
      ```
      
      ---
      
      ## Log Aggregation
      
      ### Loki
      
      **Architecture:** Like Prometheus, but for logs. Index-free design — indexes labels only, not log content.
      
      #### Loki Configuration
      
      ```yaml
      # loki-config.yml
      auth_enabled: false
      
      server:
        http_listen_port: 3100
      
      common:
        path_prefix: /loki
        storage:
          filesystem:
            chunks_directory: /loki/chunks
            rules_directory: /loki/rules
        replication_factor: 1
        ring:
          kvstore:
            store: inmemory
      
      schema_config:
        configs:
          - from: 2024-01-01
            store: tsdb
            object_store: filesystem
            schema: v13
            index:
              prefix: index_
              period: 24h
      
      limits_config:
        retention_period: 30d
        max_query_length: 721h
        max_entries_limit_per_query: 5000
      ```
      
      #### LogQL Basics
      
      ```logql
      # Filter by label
      {job="api-server"} |= "error"
      
      # JSON parsing
      {job="api-server"} | json | level="ERROR"
      
      # Pattern matching
      {job="api-server"} | json | status_code >= 500
      
      # Rate of log lines (like Prometheus rate)
      rate({job="api-server"} |= "error" [5m])
      
      # Count errors by path
      sum by (path) (
        count_over_time({job="api-server"} | json | level="ERROR" [5m])
      )
      
      # Latency percentile from log field
      quantile_over_time(0.99, {job="api-server"} | json | unwrap duration_ms [5m])
      
      # Top error messages
      topk(10,
        sum by (message) (count_over_time({job="api-server"} | json | level="ERROR" [1h]))
      )
      ```
      
      #### Label Design for Loki
      
      ```yaml
      # GOOD: Low-cardinality labels
      labels:
        job: "api-server"
        environment: "production"
        namespace: "default"
      
      # BAD: High-cardinality labels (will kill Loki performance)
      labels:
        user_id: "12345"       # Millions of unique values
        request_id: "abc-123"  # Every request is unique
        path: "/api/users/123" # Include path in log content, not labels
      ```
      
      **Rule:** Labels in Loki are for stream selection (which container/service), not for filtering log content. Use `| json | field="value"` for content filtering.
      
      #### Promtail (Log Collector)
      
      ```yaml
      # promtail-config.yml
      server:
        http_listen_port: 9080
      
      positions:
        filename: /tmp/positions.yaml
      
      clients:
        - url: http://loki:3100/loki/api/v1/push
      
      scrape_configs:
        # Docker container logs
        - job_name: docker
          docker_sd_configs:
            - host: unix:///var/run/docker.sock
              refresh_interval: 5s
          relabel_configs:
            - source_labels: ['__meta_docker_container_name']
              target_label: container
            - source_labels: ['__meta_docker_container_log_stream']
              target_label: stream
      
        # Kubernetes pod logs
        - job_name: kubernetes
          kubernetes_sd_configs:
            - role: pod
          pipeline_stages:
            - docker: {}
            - json:
                expressions:
                  level: level
                  trace_id: trace_id
            - labels:
                level:
            - timestamp:
                source: timestamp
                format: RFC3339Nano
      ```
      
      ### ELK Stack (Elasticsearch, Logstash, Kibana)
      
      #### Logstash Pipeline
      
      ```ruby
      # logstash.conf
      input {
        beats {
          port => 5044
        }
      }
      
      filter {
        # Parse JSON logs
        json {
          source => "message"
        }
      
        # Parse timestamp
        date {
          match => ["timestamp", "ISO8601"]
          target => "@timestamp"
        }
      
        # Add geoip from remote_addr
        if [remote_addr] {
          geoip {
            source => "remote_addr"
          }
        }
      
        # Redact sensitive fields
        mutate {
          remove_field => ["password", "token", "authorization"]
        }
      
        # Parse user-agent
        if [user_agent] {
          useragent {
            source => "user_agent"
            target => "ua"
          }
        }
      }
      
      output {
        elasticsearch {
          hosts => ["elasticsearch:9200"]
          index => "logs-%{[service]}-%{+YYYY.MM.dd}"
        }
      }
      ```
      
      ### CloudWatch Logs
      
      ```python
      # Python: CloudWatch Logs with structlog
      import structlog
      import watchtower
      import logging
      
      # CloudWatch handler
      cw_handler = watchtower.CloudWatchLogHandler(
          log_group="production/api-server",
          stream_name="{hostname}-{datetime}",
          use_queues=True,
          create_log_group=True,
      )
      
      # Configure structlog to output JSON
      structlog.configure(
          processors=[
              structlog.processors.TimeStamper(fmt="iso"),
              structlog.processors.JSONRenderer()
          ],
          wrapper_class=structlog.stdlib.BoundLogger,
          logger_factory=structlog.stdlib.LoggerFactory(),
      )
      
      logging.basicConfig(handlers=[cw_handler], level=logging.INFO)
      ```
      
      #### CloudWatch Metric Filters
      
      Extract metrics from log patterns:
      
      ```json
      {
        "filterPattern": "{ $.level = \"ERROR\" }",
        "metricTransformations": [
          {
            "metricName": "ErrorCount",
            "metricNamespace": "ApiServer",
            "metricValue": "1",
            "defaultValue": 0
          }
        ]
      }
      ```
      
      ```json
      {
        "filterPattern": "{ $.duration_ms > 1000 }",
        "metricTransformations": [
          {
            "metricName": "SlowRequests",
            "metricNamespace": "ApiServer",
            "metricValue": "$.duration_ms"
          }
        ]
      }
      ```
      
      ---
      
      ## Log Retention Policies
      
      ### Tiered Storage Strategy
      
      | Tier | Duration | Resolution | Storage | Cost |
      |------|----------|------------|---------|------|
      | **Hot** | 0-7 days | Full fidelity | SSD / fast storage | $$$ |
      | **Warm** | 7-30 days | Full fidelity | Standard storage | $$ |
      | **Cold** | 30-90 days | Sampled or compressed | Object storage (S3) | $ |
      | **Archive** | 90 days - 7 years | Compressed | Glacier / archive | ¢ |
      
      ### Compliance Retention Requirements
      
      | Regulation | Minimum Retention | Notes |
      |------------|-------------------|-------|
      | PCI DSS | 1 year (3 months immediately available) | Audit logs for card data access |
      | HIPAA | 6 years | Access logs for health data |
      | SOX | 7 years | Financial system audit trails |
      | GDPR | "No longer than necessary" | Right to erasure applies |
      | SOC 2 | 1 year typical | Security event logs |
      
      ### Cost Optimization
      
      1. **Set appropriate log levels:** DEBUG off in production saves 50-80% volume
      2. **Sample verbose paths:** Log 10% of health check requests
      3. **Truncate large fields:** Limit request/response body logging to 4KB
      4. **Use log-based metrics:** Extract counts/rates, then archive raw logs
      5. **Compress early:** Enable gzip on log transport (Promtail, Fluentd)
      6. **Delete test/staging logs aggressively:** 7-day retention for non-production
      
      ---
      
      ## Sensitive Data Handling
      
      ### PII Masking
      
      ```go
      // Go: mask sensitive fields before logging
      func maskEmail(email string) string {
          parts := strings.Split(email, "@")
          if len(parts) != 2 {
              return "***"
          }
          name := parts[0]
          if len(name) > 2 {
              name = name[:2] + strings.Repeat("*", len(name)-2)
          }
          return name + "@" + parts[1]
      }
      
      func maskCreditCard(cc string) string {
          if len(cc) < 4 {
              return "****"
          }
          return strings.Repeat("*", len(cc)-4) + cc[len(cc)-4:]
      }
      ```
      
      ```python
      # Python: structlog processor for PII masking
      import re
      
      SENSITIVE_KEYS = {"password", "token", "secret", "authorization", "cookie", "ssn"}
      EMAIL_PATTERN = re.compile(r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}")
      CC_PATTERN = re.compile(r"\b\d{4}[\s-]?\d{4}[\s-]?\d{4}[\s-]?\d{4}\b")
      
      def mask_sensitive_data(logger, method_name, event_dict):
          for key, value in list(event_dict.items()):
              if key.lower() in SENSITIVE_KEYS:
                  event_dict[key] = "***REDACTED***"
              elif isinstance(value, str):
                  value = EMAIL_PATTERN.sub("[EMAIL]", value)
                  value = CC_PATTERN.sub("[CREDIT_CARD]", value)
                  event_dict[key] = value
          return event_dict
      
      structlog.configure(
          processors=[
              mask_sensitive_data,
              structlog.processors.JSONRenderer(),
          ]
      )
      ```
      
      ### Fields to Never Log
      
      | Field | Risk | Alternative |
      |-------|------|-------------|
      | Passwords | Credential exposure | Log "password changed" event, not the value |
      | API keys / tokens | Service compromise | Log last 4 characters only |
      | Credit card numbers | PCI violation | Log last 4 digits, masked |
      | SSN / national ID | Identity theft | Never log, even masked |
      | Full request bodies with auth | Token leakage | Strip Authorization header |
      | Database connection strings | DB credential exposure | Log host:port only |
      
      ---
      
      ## Log-Based Metrics
      
      ### Loki Recording Rules
      
      ```yaml
      # loki-rules.yml
      groups:
        - name: log_metrics
          interval: 1m
          rules:
            - record: log:errors:rate5m
              expr: |
                sum by (service) (
                  rate({job=~".+"} | json | level="ERROR" [5m])
                )
      
            - record: log:requests:duration_p99_5m
              expr: |
                quantile_over_time(0.99,
                  {job="api-server"} | json | unwrap duration_ms [5m]
                ) by (service)
      ```
      
      ### Extracting Metrics from Logs
      
      When full metrics instrumentation isn't available, derive metrics from structured logs:
      
      ```promql
      # Error rate from logs (Loki)
      sum(rate({job="api-server"} | json | level="ERROR" [5m]))
      
      # Slow request rate from logs
      sum(rate({job="api-server"} | json | duration_ms > 1000 [5m]))
      
      # Unique users from logs (approximate)
      count(
        count by (user_id) (
          {job="api-server"} | json | user_id != "" [1h]
        )
      )
      ```
      
      ---
      
      ## Language-Specific Logging
      
      ### Go (slog - standard library, Go 1.21+)
      
      ```go
      package main
      
      import (
          "context"
          "log/slog"
          "os"
      )
      
      func main() {
          // JSON handler for production
          handler := slog.NewJSONHandler(os.Stdout, &slog.HandlerOptions{
              Level: slog.LevelInfo,
              AddSource: true,  // Add file:line to log entries
          })
          logger := slog.New(handler)
          slog.SetDefault(logger)
      
          // Basic logging
          slog.Info("Server starting", "port", 8080, "version", "1.4.2")
      
          // With context (includes trace_id if using OpenTelemetry bridge)
          ctx := context.Background()
          slog.InfoContext(ctx, "Request handled",
              "method", "GET",
              "path", "/api/users",
              "status", 200,
              "duration_ms", 45,
          )
      
          // Error logging with error value
          slog.Error("Database query failed",
              "error", err,
              "query", "SELECT * FROM users WHERE id = $1",
              "user_id", userID,
          )
      
          // Create child logger with bound attributes
          userLogger := slog.With("user_id", "usr-789", "session_id", "sess-abc")
          userLogger.Info("User action", "action", "login")
      }
      ```
      
      ### Python (structlog)
      
      ```python
      import structlog
      
      structlog.configure(
          processors=[
              structlog.contextvars.merge_contextvars,
              structlog.processors.add_log_level,
              structlog.processors.TimeStamper(fmt="iso"),
              structlog.processors.StackInfoRenderer(),
              structlog.processors.format_exc_info,
              structlog.processors.JSONRenderer(),
          ],
          wrapper_class=structlog.stdlib.BoundLogger,
          context_class=dict,
          logger_factory=structlog.stdlib.LoggerFactory(),
      )
      
      log = structlog.get_logger()
      
      # Basic logging
      log.info("server_starting", port=8080, version="1.4.2")
      
      # Bind context for the request
      log = log.bind(request_id="req-abc123", user_id="usr-789")
      log.info("request_handled", method="GET", path="/api/users", status=200, duration_ms=45)
      
      # Error with exception
      try:
          process_payment(payment_id)
      except Exception:
          log.error("payment_failed", payment_id="pay-456", exc_info=True)
      
      # Context variables (available across async calls)
      structlog.contextvars.bind_contextvars(request_id="req-abc123")
      ```
      
      ### Node.js (pino)
      
      ```javascript
      import pino from 'pino';
      
      const logger = pino({
        level: process.env.LOG_LEVEL || 'info',
        timestamp: pino.stdTimeFunctions.isoTime,
        formatters: {
          level: (label) => ({ level: label.toUpperCase() }),
        },
        serializers: {
          err: pino.stdSerializers.err,
          req: pino.stdSerializers.req,
          res: pino.stdSerializers.res,
        },
        redact: ['req.headers.authorization', 'req.headers.cookie', 'password'],
      });
      
      // Basic logging
      logger.info({ port: 8080, version: '1.4.2' }, 'Server starting');
      
      // Child logger with bound context
      const reqLogger = logger.child({ requestId: 'req-abc123', userId: 'usr-789' });
      reqLogger.info({ method: 'GET', path: '/api/users', status: 200, durationMs: 45 }, 'Request handled');
      
      // Error logging
      reqLogger.error({ err, paymentId: 'pay-456' }, 'Payment failed');
      
      // Express/Fastify integration
      import pinoHttp from 'pino-http';
      app.use(pinoHttp({ logger }));
      ```
      
      ### Rust (tracing crate)
      
      ```rust
      use tracing::{info, error, warn, instrument, Level};
      use tracing_subscriber::{fmt, EnvFilter};
      
      fn main() {
          // JSON subscriber for production
          tracing_subscriber::fmt()
              .json()
              .with_env_filter(EnvFilter::from_default_env())
              .with_target(true)
              .with_thread_ids(true)
              .with_file(true)
              .with_line_number(true)
              .init();
      
          info!(port = 8080, version = "1.4.2", "Server starting");
      }
      
      #[instrument(skip(db), fields(user_id = %user_id))]
      async fn get_user(db: &Pool, user_id: &str) -> Result<User, Error> {
          info!("Fetching user from database");
      
          match db.query_one("SELECT * FROM users WHERE id = $1", &[&user_id]).await {
              Ok(row) => {
                  info!("User found");
                  Ok(User::from_row(row))
              }
              Err(e) => {
                  error!(error = %e, "Database query failed");
                  Err(e.into())
              }
          }
      }
      ```
      
      ### Java (Logback + Structured Logging)
      
      ```xml
      <!-- logback.xml -->
      <configuration>
        <appender name="STDOUT" class="ch.qos.logback.core.ConsoleAppender">
          <encoder class="net.logstash.logback.encoder.LogstashEncoder">
            <includeMdcKeyName>request_id</includeMdcKeyName>
            <includeMdcKeyName>trace_id</includeMdcKeyName>
            <includeMdcKeyName>user_id</includeMdcKeyName>
          </encoder>
        </appender>
      
        <root level="INFO">
          <appender-ref ref="STDOUT" />
        </root>
      </configuration>
      ```
      
      ```java
      import org.slf4j.Logger;
      import org.slf4j.LoggerFactory;
      import org.slf4j.MDC;
      import net.logstash.logback.argument.StructuredArguments;
      import static net.logstash.logback.argument.StructuredArguments.*;
      
      Logger log = LoggerFactory.getLogger(PaymentService.class);
      
      // Set MDC for request context
      MDC.put("request_id", requestId);
      MDC.put("trace_id", traceId);
      MDC.put("user_id", userId);
      
      // Structured logging with key-value pairs
      log.info("Request handled", kv("method", "GET"), kv("path", "/api/users"),
               kv("status", 200), kv("duration_ms", 45));
      
      // Error logging
      log.error("Payment failed", kv("payment_id", paymentId), kv("error", e.getMessage()), e);
      
      // Clean up MDC
      MDC.clear();
      ```
      
      ---
      
      ## Common Patterns
      
      ### Error Logging with Stack Traces
      
      Always include the full stack trace for errors, but consider truncation for very deep stacks:
      
      ```go
      // Go
      slog.Error("Operation failed",
          "error", err.Error(),
          "stack", fmt.Sprintf("%+v", err),  // With pkgs/errors stack
      )
      ```
      
      ```python
      # Python - structlog handles exc_info automatically
      log.error("operation_failed", exc_info=True)
      ```
      
      ### Audit Logging
      
      For compliance-required operations:
      
      ```json
      {
        "timestamp": "2026-03-09T14:32:01.123Z",
        "level": "INFO",
        "type": "audit",
        "action": "user.role.changed",
        "actor": {"id": "usr-admin-1", "type": "user", "ip": "10.0.0.5"},
        "target": {"id": "usr-789", "type": "user"},
        "changes": {"role": {"from": "viewer", "to": "editor"}},
        "result": "success",
        "request_id": "req-abc123"
      }
      ```
      
      ### Request/Response Logging
      
      ```go
      // Log request on entry, response on exit
      func loggingMiddleware(next http.Handler) http.Handler {
          return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
              start := time.Now()
      
              // Log request (don't log body for GET, limit body size for POST)
              slog.InfoContext(r.Context(), "Request received",
                  "method", r.Method,
                  "path", r.URL.Path,
                  "remote_addr", r.RemoteAddr,
              )
      
              wrapped := wrapResponseWriter(w)
              next.ServeHTTP(wrapped, r)
      
              slog.InfoContext(r.Context(), "Request completed",
                  "method", r.Method,
                  "path", r.URL.Path,
                  "status", wrapped.Status(),
                  "duration_ms", time.Since(start).Milliseconds(),
                  "bytes", wrapped.BytesWritten(),
              )
          })
      }
      ```
      
      ### Health Check Log Suppression
      
      Don't fill logs with health check noise:
      
      ```go
      func loggingMiddleware(next http.Handler) http.Handler {
          return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
              // Skip logging for health checks
              if r.URL.Path == "/health" || r.URL.Path == "/ready" {
                  next.ServeHTTP(w, r)
                  return
              }
              // ... normal logging
          })
      }
      ```
      
    • metrics-alerting.md 27.6 KB
      # Metrics and Alerting Reference
      
      Comprehensive reference for metrics collection, visualization, alerting, SLOs, and uptime monitoring.
      
      ---
      
      ## Prometheus
      
      ### Architecture Overview
      
      ```
      ┌─────────────┐     ┌─────────────┐     ┌─────────────────┐
      │ Application │────▶│  Prometheus  │────▶│  Alertmanager   │
      │  /metrics   │pull │  (TSDB)      │push │  (routing/notif) │
      └─────────────┘     └──────┬──────┘     └─────────────────┘
                                 │query
                          ┌──────▼──────┐
                          │   Grafana    │
                          │ (dashboards) │
                          └─────────────┘
      ```
      
      **Key characteristics:**
      - Pull-based model (Prometheus scrapes targets)
      - Local time-series database (TSDB)
      - PromQL query language
      - Built-in alerting rules evaluated by Prometheus, routed by Alertmanager
      - Service discovery (Kubernetes, Consul, DNS, file-based, EC2)
      
      ### Prometheus Configuration (prometheus.yml)
      
      ```yaml
      global:
        scrape_interval: 15s          # Default scrape interval
        evaluation_interval: 15s      # Rule evaluation interval
        scrape_timeout: 10s           # Per-scrape timeout
      
      # Alertmanager configuration
      alerting:
        alertmanagers:
          - static_configs:
              - targets:
                  - alertmanager:9093
      
      # Rule files
      rule_files:
        - "rules/*.yml"
      
      # Scrape targets
      scrape_configs:
        # Self-monitoring
        - job_name: "prometheus"
          static_configs:
            - targets: ["localhost:9090"]
      
        # Application with static targets
        - job_name: "api-server"
          metrics_path: /metrics
          scheme: https
          static_configs:
            - targets: ["api1:8080", "api2:8080"]
              labels:
                environment: production
      
        # Kubernetes service discovery
        - job_name: "kubernetes-pods"
          kubernetes_sd_configs:
            - role: pod
          relabel_configs:
            - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
              action: keep
              regex: true
            - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_path]
              action: replace
              target_label: __metrics_path__
              regex: (.+)
            - source_labels: [__meta_kubernetes_namespace]
              action: replace
              target_label: namespace
            - source_labels: [__meta_kubernetes_pod_name]
              action: replace
              target_label: pod
      
        # Node exporter
        - job_name: "node"
          static_configs:
            - targets: ["node-exporter:9100"]
      ```
      
      ### PromQL Basics
      
      #### Rate and Increase
      
      ```promql
      # Per-second rate over 5 minutes (use for counters)
      rate(http_requests_total[5m])
      
      # Per-second rate for specific status codes
      rate(http_requests_total{status_code=~"5.."}[5m])
      
      # Total increase over 1 hour (use for counters)
      increase(http_requests_total[1h])
      
      # irate: instant rate using last two data points (more volatile)
      irate(http_requests_total[5m])
      ```
      
      **Rule:** Always use `rate()` or `increase()` with counters. Never display raw counter values.
      
      #### Aggregation Operators
      
      ```promql
      # Sum across all instances
      sum(rate(http_requests_total[5m]))
      
      # Sum by specific label
      sum by (method, path) (rate(http_requests_total[5m]))
      
      # Average across instances
      avg(node_cpu_seconds_total{mode="idle"})
      
      # Maximum value across instances
      max by (instance) (node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes)
      
      # Count number of time series
      count(up == 1)
      
      # Top 5 by value
      topk(5, rate(http_requests_total[5m]))
      
      # Bottom 5 by value
      bottomk(5, rate(http_requests_total[5m]))
      ```
      
      #### Histogram Quantiles
      
      ```promql
      # 99th percentile latency
      histogram_quantile(0.99,
        sum by (le) (rate(http_request_duration_seconds_bucket[5m]))
      )
      
      # 95th percentile latency by service
      histogram_quantile(0.95,
        sum by (le, service) (rate(http_request_duration_seconds_bucket[5m]))
      )
      
      # 50th percentile (median)
      histogram_quantile(0.50,
        sum by (le) (rate(http_request_duration_seconds_bucket[5m]))
      )
      
      # Average latency from histogram
      sum(rate(http_request_duration_seconds_sum[5m]))
      /
      sum(rate(http_request_duration_seconds_count[5m]))
      ```
      
      #### Useful Functions
      
      ```promql
      # Detect missing metrics (target down)
      absent(up{job="api-server"})
      
      # Time since last change (staleness)
      time() - process_start_time_seconds
      
      # Predict value in 4 hours using linear regression
      predict_linear(node_filesystem_avail_bytes[6h], 4*3600)
      
      # Compare to 1 week ago
      rate(http_requests_total[5m]) / rate(http_requests_total[5m] offset 7d)
      
      # Clamping values
      clamp_min(free_disk_percentage, 0)
      clamp_max(cpu_usage_percentage, 100)
      
      # Label manipulation
      label_replace(up, "short_instance", "$1", "instance", "(.*):.*")
      ```
      
      ### Recording Rules
      
      Pre-compute expensive queries for dashboards and alerts:
      
      ```yaml
      # rules/recording-rules.yml
      groups:
        - name: http_request_rules
          interval: 15s
          rules:
            # Pre-compute request rate by service and status
            - record: job:http_requests:rate5m
              expr: sum by (job, status_code) (rate(http_requests_total[5m]))
      
            # Pre-compute error rate percentage
            - record: job:http_request_errors:ratio5m
              expr: |
                sum by (job) (rate(http_requests_total{status_code=~"5.."}[5m]))
                /
                sum by (job) (rate(http_requests_total[5m]))
      
            # Pre-compute p99 latency
            - record: job:http_request_duration_seconds:p99_5m
              expr: |
                histogram_quantile(0.99,
                  sum by (job, le) (rate(http_request_duration_seconds_bucket[5m]))
                )
      
            # Pre-compute availability
            - record: job:availability:ratio5m
              expr: |
                1 - (
                  sum by (job) (rate(http_requests_total{status_code=~"5.."}[5m]))
                  /
                  sum by (job) (rate(http_requests_total[5m]))
                )
      ```
      
      ### Alerting Rules
      
      ```yaml
      # rules/alerting-rules.yml
      groups:
        - name: service_alerts
          rules:
            # High error rate
            - alert: HighErrorRate
              expr: job:http_request_errors:ratio5m > 0.01
              for: 5m
              labels:
                severity: warning
                team: backend
              annotations:
                summary: "High error rate on {{ $labels.job }}"
                description: "Error rate is {{ $value | humanizePercentage }} (threshold: 1%)"
                runbook_url: "https://runbooks.example.com/high-error-rate"
                dashboard_url: "https://grafana.example.com/d/service-overview?var-service={{ $labels.job }}"
      
            # Critical error rate
            - alert: CriticalErrorRate
              expr: job:http_request_errors:ratio5m > 0.05
              for: 2m
              labels:
                severity: critical
                team: backend
              annotations:
                summary: "Critical error rate on {{ $labels.job }}"
                description: "Error rate is {{ $value | humanizePercentage }} (threshold: 5%)"
                runbook_url: "https://runbooks.example.com/critical-error-rate"
      
            # High latency
            - alert: HighLatencyP99
              expr: job:http_request_duration_seconds:p99_5m > 2.0
              for: 10m
              labels:
                severity: warning
              annotations:
                summary: "P99 latency above 2s on {{ $labels.job }}"
                description: "P99 latency is {{ $value | humanizeDuration }}"
      
            # Target down
            - alert: TargetDown
              expr: up == 0
              for: 3m
              labels:
                severity: critical
              annotations:
                summary: "Target {{ $labels.instance }} is down"
                description: "Prometheus cannot scrape {{ $labels.job }}/{{ $labels.instance }}"
      
        - name: infrastructure_alerts
          rules:
            # Disk space prediction
            - alert: DiskWillFillIn24Hours
              expr: |
                predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}[6h], 24*3600) < 0
              for: 30m
              labels:
                severity: warning
              annotations:
                summary: "Disk {{ $labels.mountpoint }} on {{ $labels.instance }} will fill within 24 hours"
      
            # High memory usage
            - alert: HighMemoryUsage
              expr: |
                (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) > 0.9
              for: 10m
              labels:
                severity: warning
              annotations:
                summary: "Memory usage above 90% on {{ $labels.instance }}"
      
            # High CPU usage
            - alert: HighCPUUsage
              expr: |
                1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) > 0.85
              for: 15m
              labels:
                severity: warning
              annotations:
                summary: "CPU usage above 85% on {{ $labels.instance }}"
      ```
      
      ---
      
      ## Grafana
      
      ### Dashboard JSON Structure
      
      ```json
      {
        "dashboard": {
          "title": "Service Overview",
          "uid": "service-overview",
          "tags": ["production", "services"],
          "timezone": "browser",
          "refresh": "30s",
          "time": {
            "from": "now-6h",
            "to": "now"
          },
          "templating": {
            "list": [
              {
                "name": "service",
                "type": "query",
                "datasource": "Prometheus",
                "query": "label_values(up, job)",
                "refresh": 2,
                "multi": true,
                "includeAll": true
              },
              {
                "name": "interval",
                "type": "interval",
                "options": [
                  {"text": "1m", "value": "1m"},
                  {"text": "5m", "value": "5m"},
                  {"text": "15m", "value": "15m"}
                ],
                "current": {"text": "5m", "value": "5m"}
              }
            ]
          },
          "panels": []
        }
      }
      ```
      
      ### Panel Types
      
      #### Time Series Panel
      
      ```json
      {
        "type": "timeseries",
        "title": "Request Rate",
        "gridPos": {"h": 8, "w": 12, "x": 0, "y": 0},
        "targets": [
          {
            "expr": "sum by (status_code) (rate(http_requests_total{job=~\"$service\"}[$interval]))",
            "legendFormat": "{{status_code}}"
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "reqps",
            "custom": {
              "drawStyle": "line",
              "fillOpacity": 10,
              "stacking": {"mode": "none"}
            }
          }
        }
      }
      ```
      
      #### Stat Panel
      
      ```json
      {
        "type": "stat",
        "title": "Current Error Rate",
        "gridPos": {"h": 4, "w": 6, "x": 0, "y": 0},
        "targets": [
          {
            "expr": "sum(rate(http_requests_total{job=~\"$service\",status_code=~\"5..\"}[5m])) / sum(rate(http_requests_total{job=~\"$service\"}[5m]))",
            "instant": true
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "percentunit",
            "thresholds": {
              "steps": [
                {"color": "green", "value": null},
                {"color": "yellow", "value": 0.001},
                {"color": "red", "value": 0.01}
              ]
            }
          }
        }
      }
      ```
      
      #### Gauge Panel
      
      ```json
      {
        "type": "gauge",
        "title": "CPU Usage",
        "targets": [
          {
            "expr": "1 - avg(rate(node_cpu_seconds_total{mode=\"idle\",instance=~\"$instance\"}[5m]))",
            "instant": true
          }
        ],
        "fieldConfig": {
          "defaults": {
            "unit": "percentunit",
            "min": 0,
            "max": 1,
            "thresholds": {
              "steps": [
                {"color": "green", "value": null},
                {"color": "yellow", "value": 0.7},
                {"color": "red", "value": 0.9}
              ]
            }
          }
        }
      }
      ```
      
      #### Table Panel
      
      ```json
      {
        "type": "table",
        "title": "Top Endpoints by Error Rate",
        "targets": [
          {
            "expr": "topk(10, sum by (method, path) (rate(http_requests_total{status_code=~\"5..\"}[5m])))",
            "instant": true,
            "format": "table"
          }
        ],
        "transformations": [
          {"id": "organize", "options": {"excludeByName": {"Time": true}}}
        ]
      }
      ```
      
      ### Grafana Variables
      
      | Type | Use Case | Example |
      |------|----------|---------|
      | **Query** | Dynamic from datasource | `label_values(up, job)` |
      | **Custom** | Fixed list of values | `production,staging,development` |
      | **Interval** | Time range intervals | `1m,5m,15m,1h` |
      | **Datasource** | Multiple Prometheus instances | Type: datasource, Query: Prometheus |
      | **Text box** | Free-form input | Filter by custom string |
      
      ### Annotations
      
      ```json
      {
        "annotations": {
          "list": [
            {
              "name": "Deployments",
              "datasource": "Prometheus",
              "enable": true,
              "expr": "changes(process_start_time_seconds{job=\"api-server\"}[1m]) > 0",
              "tagKeys": "job",
              "titleFormat": "Deployment: {{job}}"
            },
            {
              "name": "Alerts",
              "datasource": "-- Grafana --",
              "enable": true,
              "type": "alert"
            }
          ]
        }
      }
      ```
      
      ---
      
      ## OpenTelemetry Metrics
      
      ### Go SDK Setup
      
      ```go
      package main
      
      import (
          "context"
          "log"
          "time"
      
          "go.opentelemetry.io/otel"
          "go.opentelemetry.io/otel/exporters/prometheus"
          "go.opentelemetry.io/otel/metric"
          sdkmetric "go.opentelemetry.io/otel/sdk/metric"
      )
      
      func initMeterProvider() (*sdkmetric.MeterProvider, error) {
          exporter, err := prometheus.New()
          if err != nil {
              return nil, err
          }
      
          mp := sdkmetric.NewMeterProvider(
              sdkmetric.WithReader(exporter),
          )
          otel.SetMeterProvider(mp)
          return mp, nil
      }
      
      func main() {
          mp, err := initMeterProvider()
          if err != nil {
              log.Fatal(err)
          }
          defer mp.Shutdown(context.Background())
      
          meter := otel.Meter("myapp")
      
          // Counter
          requestCounter, _ := meter.Int64Counter(
              "http.server.request.total",
              metric.WithDescription("Total HTTP requests"),
              metric.WithUnit("{request}"),
          )
      
          // Histogram
          latencyHistogram, _ := meter.Float64Histogram(
              "http.server.request.duration",
              metric.WithDescription("HTTP request latency"),
              metric.WithUnit("s"),
              metric.WithExplicitBucketBoundaries(0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10),
          )
      
          // UpDownCounter (gauge-like)
          activeConnections, _ := meter.Int64UpDownCounter(
              "http.server.active_connections",
              metric.WithDescription("Active HTTP connections"),
          )
      
          // Usage
          ctx := context.Background()
          requestCounter.Add(ctx, 1, metric.WithAttributes(
              attribute.String("method", "GET"),
              attribute.String("path", "/api/users"),
              attribute.Int("status_code", 200),
          ))
      
          start := time.Now()
          // ... handle request ...
          latencyHistogram.Record(ctx, time.Since(start).Seconds())
      
          activeConnections.Add(ctx, 1)   // connection opened
          activeConnections.Add(ctx, -1)  // connection closed
      }
      ```
      
      ### Python SDK Setup
      
      ```python
      from opentelemetry import metrics
      from opentelemetry.sdk.metrics import MeterProvider
      from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader
      from opentelemetry.exporter.prometheus import PrometheusMetricReader
      from prometheus_client import start_http_server
      
      # Prometheus exporter
      reader = PrometheusMetricReader()
      provider = MeterProvider(metric_readers=[reader])
      metrics.set_meter_provider(provider)
      
      # Start Prometheus HTTP server on port 8000
      start_http_server(8000)
      
      meter = metrics.get_meter("myapp")
      
      # Counter
      request_counter = meter.create_counter(
          name="http.server.request.total",
          description="Total HTTP requests",
          unit="{request}",
      )
      
      # Histogram
      latency_histogram = meter.create_histogram(
          name="http.server.request.duration",
          description="HTTP request latency",
          unit="s",
      )
      
      # UpDownCounter
      active_connections = meter.create_up_down_counter(
          name="http.server.active_connections",
          description="Active HTTP connections",
      )
      
      # Usage
      request_counter.add(1, {"method": "GET", "path": "/api/users", "status_code": 200})
      latency_histogram.record(0.045, {"method": "GET", "path": "/api/users"})
      active_connections.add(1)
      ```
      
      ### Node.js SDK Setup
      
      ```javascript
      const { MeterProvider } = require('@opentelemetry/sdk-metrics');
      const { PrometheusExporter } = require('@opentelemetry/exporter-prometheus');
      const { metrics } = require('@opentelemetry/api');
      
      const exporter = new PrometheusExporter({ port: 9464 });
      const meterProvider = new MeterProvider({
        readers: [exporter],
      });
      metrics.setGlobalMeterProvider(meterProvider);
      
      const meter = metrics.getMeter('myapp');
      
      // Counter
      const requestCounter = meter.createCounter('http.server.request.total', {
        description: 'Total HTTP requests',
        unit: '{request}',
      });
      
      // Histogram
      const latencyHistogram = meter.createHistogram('http.server.request.duration', {
        description: 'HTTP request latency',
        unit: 's',
        advice: {
          explicitBucketBoundaries: [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10],
        },
      });
      
      // UpDownCounter
      const activeConnections = meter.createUpDownCounter('http.server.active_connections', {
        description: 'Active HTTP connections',
      });
      
      // Usage
      requestCounter.add(1, { method: 'GET', path: '/api/users', status_code: 200 });
      latencyHistogram.record(0.045, { method: 'GET', path: '/api/users' });
      activeConnections.add(1);
      ```
      
      ---
      
      ## StatsD
      
      ### Protocol Format
      
      ```
      <metric_name>:<value>|<type>|@<sample_rate>|#<tags>
      ```
      
      | Type | Code | Example |
      |------|------|---------|
      | Counter | `c` | `page.views:1\|c` |
      | Gauge | `g` | `fuel.level:0.5\|g` |
      | Timer | `ms` | `request.duration:320\|ms` |
      | Set | `s` | `users.uniques:user123\|s` |
      | Histogram | `h` | `request.size:512\|h` (DogStatsD) |
      | Distribution | `d` | `request.duration:320\|d` (DogStatsD) |
      
      ### DogStatsD Extensions (Datadog)
      
      ```
      # Counter with tags
      http.requests:1|c|#method:GET,path:/api/users,status:200
      
      # Histogram with sample rate
      http.request.duration:45.2|h|@0.5|#service:api
      
      # Gauge
      system.cpu.usage:72.5|g|#host:web01
      
      # Service check
      _sc|myservice.health|0|#env:production|m:Service is healthy
      ```
      
      **When to use StatsD over Prometheus:**
      - Existing StatsD infrastructure
      - Simple counter/gauge/timer needs without complex queries
      - Push model required (ephemeral jobs, serverless)
      - Language/framework has StatsD client but no Prometheus client
      
      ---
      
      ## Custom Metrics Design
      
      ### Naming Conventions
      
      Follow OpenMetrics/Prometheus naming:
      
      ```
      <namespace>_<subsystem>_<name>_<unit>_<suffix>
      ```
      
      | Component | Rules | Examples |
      |-----------|-------|---------|
      | Namespace | Application or domain | `myapp`, `payment`, `auth` |
      | Subsystem | Component within app | `http`, `db`, `cache`, `queue` |
      | Name | What is measured | `request`, `connection`, `query` |
      | Unit | SI unit (base, not milli/micro) | `seconds`, `bytes`, `ratio` |
      | Suffix | Metric type | `_total` (counter), `_info` (metadata), `_bucket` (histogram) |
      
      **Good names:**
      ```
      http_server_request_duration_seconds          # histogram
      http_server_requests_total                    # counter
      db_connection_pool_active_connections         # gauge
      cache_hit_ratio                               # gauge (0-1)
      queue_messages_total                          # counter
      payment_processing_duration_seconds           # histogram
      ```
      
      **Bad names:**
      ```
      requestCount          # No namespace, no suffix, camelCase
      latency_ms            # Milliseconds (use seconds), no namespace
      errors                # Vague, no namespace, no suffix
      HttpRequests          # PascalCase
      ```
      
      ### Label Best Practices
      
      **Do:**
      - Use labels for dimensions you will filter/aggregate by
      - Keep label cardinality bounded (< 100 unique values per label)
      - Use consistent label names across metrics (`method`, not `http_method` in some and `request_method` in others)
      
      **Don't:**
      - Use user IDs, email addresses, or request IDs as labels (unbounded cardinality)
      - Use full URL paths as labels (use route templates: `/api/users/{id}`, not `/api/users/12345`)
      - Use error messages as labels (unbounded text)
      - Create more than 5-7 labels per metric
      
      ### Avoiding Cardinality Bombs
      
      ```
      # BAD: unbounded path label
      http_requests_total{path="/api/users/12345"}    # Millions of unique series
      http_requests_total{path="/api/users/67890"}
      
      # GOOD: use route template
      http_requests_total{route="/api/users/{id}"}    # One series per route
      
      # BAD: error message as label
      errors_total{message="connection refused to 10.0.0.5:5432"}
      
      # GOOD: error category as label
      errors_total{type="connection_refused", target="postgres"}
      ```
      
      **Cardinality check query:**
      ```promql
      # Find high-cardinality metrics
      topk(10, count by (__name__) ({__name__=~".+"}))
      
      # Check specific metric cardinality
      count(http_requests_total)
      ```
      
      ---
      
      ## SLI / SLO / SLA
      
      ### Definitions
      
      | Term | Definition | Example |
      |------|------------|---------|
      | **SLI** (Service Level Indicator) | Quantitative measure of service behavior | 99.2% of requests complete in < 500ms |
      | **SLO** (Service Level Objective) | Target value for an SLI | 99.5% of requests should complete in < 500ms |
      | **SLA** (Service Level Agreement) | Business contract with consequences | 99.9% availability or credit issued |
      
      **Relationship:** SLI measures reality → SLO sets the target → SLA defines business consequences.
      
      ### Error Budget Calculation
      
      ```
      Error budget = 1 - SLO target
      
      Example:
        SLO = 99.9% availability
        Error budget = 0.1% = 43.2 minutes/month
      
        In a 30-day month:
        - Total minutes: 43,200
        - Allowed downtime: 43.2 minutes
        - Allowed error requests: 0.1% of total
      ```
      
      ### Burn Rate Alerting
      
      Burn rate = rate at which error budget is being consumed relative to the budget period.
      
      ```
      burn_rate = error_rate / (1 - SLO_target)
      ```
      
      | Burn Rate | Budget Exhaustion | Alert? |
      |-----------|-------------------|--------|
      | 1x | 30 days (full period) | No |
      | 2x | 15 days | No |
      | 6x | 5 days | Ticket (warning) |
      | 14.4x | 2 days | Page (critical) |
      | 36x | 20 hours | Page immediately |
      
      **Multi-window burn rate alert (recommended):**
      
      ```yaml
      # Fast burn: 14.4x burn rate over 1-hour window, confirmed by 5-minute window
      - alert: SLOHighBurnRate
        expr: |
          (
            sum(rate(http_requests_total{status_code=~"5.."}[1h]))
            /
            sum(rate(http_requests_total[1h]))
          ) > (14.4 * 0.001)
          and
          (
            sum(rate(http_requests_total{status_code=~"5.."}[5m]))
            /
            sum(rate(http_requests_total[5m]))
          ) > (14.4 * 0.001)
        labels:
          severity: critical
        annotations:
          summary: "High error budget burn rate"
      
      # Slow burn: 6x burn rate over 6-hour window, confirmed by 30-minute window
      - alert: SLOSlowBurnRate
        expr: |
          (
            sum(rate(http_requests_total{status_code=~"5.."}[6h]))
            /
            sum(rate(http_requests_total[6h]))
          ) > (6 * 0.001)
          and
          (
            sum(rate(http_requests_total{status_code=~"5.."}[30m]))
            /
            sum(rate(http_requests_total[30m]))
          ) > (6 * 0.001)
        labels:
          severity: warning
      ```
      
      ### SLO Document Template
      
      ```markdown
      # SLO: [Service Name] - [SLO Name]
      
      ## Overview
      - **Service:** payment-api
      - **Owner:** payments-team
      - **Last reviewed:** 2026-03-01
      
      ## SLI Definition
      - **Type:** Availability (success rate)
      - **Good events:** HTTP responses with status < 500
      - **Total events:** All HTTP responses
      - **Measurement:** `sum(rate(http_requests_total{status<500}[5m])) / sum(rate(http_requests_total[5m]))`
      
      ## SLO Target
      - **Target:** 99.9%
      - **Window:** 30 days (rolling)
      - **Error budget:** 0.1% = ~43 minutes of downtime
      
      ## Alerting
      - **Fast burn (page):** 14.4x burn rate for 1 hour
      - **Slow burn (ticket):** 6x burn rate for 6 hours
      
      ## Consequences of Missing SLO
      - Freeze non-critical deployments
      - Allocate sprint capacity to reliability
      - Review in next SLO review meeting
      ```
      
      ---
      
      ## Alert Routing
      
      ### Alertmanager Configuration
      
      ```yaml
      # alertmanager.yml
      global:
        resolve_timeout: 5m
        slack_api_url: "https://hooks.slack.com/services/T00/B00/XXX"
        pagerduty_url: "https://events.pagerduty.com/v2/enqueue"
      
      route:
        receiver: "default-slack"
        group_by: ["alertname", "job"]
        group_wait: 30s        # Wait before sending first notification
        group_interval: 5m     # Wait before sending updates
        repeat_interval: 4h    # Resend if not resolved
      
        routes:
          # Critical alerts → PagerDuty
          - match:
              severity: critical
            receiver: "pagerduty-critical"
            group_wait: 10s
            repeat_interval: 1h
      
          # Warning alerts → Slack
          - match:
              severity: warning
            receiver: "slack-warnings"
            repeat_interval: 4h
      
          # Info alerts → Slack info channel
          - match:
              severity: info
            receiver: "slack-info"
            repeat_interval: 24h
      
          # Team-specific routing
          - match:
              team: database
            receiver: "pagerduty-database"
            routes:
              - match:
                  severity: critical
                receiver: "pagerduty-database"
              - match:
                  severity: warning
                receiver: "slack-database"
      
      receivers:
        - name: "default-slack"
          slack_configs:
            - channel: "#alerts"
              title: '{{ .GroupLabels.alertname }}'
              text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'
      
        - name: "pagerduty-critical"
          pagerduty_configs:
            - service_key: "<integration-key>"
              severity: critical
              description: '{{ .GroupLabels.alertname }}: {{ .CommonAnnotations.summary }}'
              details:
                description: '{{ .CommonAnnotations.description }}'
                runbook: '{{ .CommonAnnotations.runbook_url }}'
      
        - name: "slack-warnings"
          slack_configs:
            - channel: "#alerts-warning"
              title: ':warning: {{ .GroupLabels.alertname }}'
              text: '{{ .CommonAnnotations.description }}'
      
        - name: "slack-info"
          slack_configs:
            - channel: "#alerts-info"
      
      inhibit_rules:
        # Suppress warning if critical is already firing
        - source_match:
            severity: critical
          target_match:
            severity: warning
          equal: ["alertname", "job"]
      ```
      
      ### Runbook Template
      
      ```markdown
      # Runbook: [Alert Name]
      
      ## Alert Details
      - **Alert:** HighErrorRate
      - **Severity:** Warning / Critical
      - **Team:** backend
      
      ## Symptom
      What the user/system is experiencing when this alert fires.
      
      ## Investigation Steps
      1. Check the Grafana dashboard: [link]
      2. Check recent deployments: `kubectl rollout history deployment/api`
      3. Check error logs: `kubectl logs -l app=api --tail=100 | jq 'select(.level=="ERROR")'`
      4. Check downstream dependencies: [dashboard link]
      
      ## Mitigation
      Immediate actions to reduce impact:
      1. If caused by recent deploy: `kubectl rollout undo deployment/api`
      2. If caused by downstream: Enable circuit breaker / failover
      3. If caused by traffic spike: Scale horizontally
      
      ## Resolution
      Steps to fully resolve:
      1. Identify root cause from logs/traces
      2. Create fix PR
      3. Deploy fix through normal pipeline
      4. Verify error rate returns to baseline
      
      ## Escalation
      - Level 1: On-call engineer (this runbook)
      - Level 2: Team lead (@team-lead)
      - Level 3: VP Engineering (for customer-impacting incidents)
      ```
      
      ---
      
      ## Uptime Monitoring
      
      ### Uptime Kuma (Self-hosted)
      
      ```yaml
      # docker-compose.yml
      services:
        uptime-kuma:
          image: louislam/uptime-kuma:1
          restart: unless-stopped
          ports:
            - "3001:3001"
          volumes:
            - uptime-kuma-data:/app/data
      
      volumes:
        uptime-kuma-data:
      ```
      
      **Features:**
      - HTTP(s), TCP, DNS, Docker, gRPC, MQTT monitors
      - Status pages (public-facing)
      - Notifications: Slack, Discord, Telegram, PagerDuty, email, webhooks
      - Certificate expiry monitoring
      - Multi-language support
      
      ### Synthetic Monitoring
      
      Run scripted checks from multiple regions to verify end-to-end functionality:
      
      ```javascript
      // Example: Grafana synthetic monitoring check
      import { check } from 'k6';
      import http from 'k6/http';
      
      export default function () {
        const res = http.get('https://api.example.com/health');
        check(res, {
          'status is 200': (r) => r.status === 200,
          'response time < 500ms': (r) => r.timings.duration < 500,
          'body contains ok': (r) => r.body.includes('"status":"ok"'),
        });
      }
      ```
      
      ### Status Pages
      
      Communicate service health to users:
      
      | Tool | Type | Features |
      |------|------|----------|
      | **Uptime Kuma** | Self-hosted | Free, built-in status page |
      | **Betteruptime** | SaaS | Incident management + status page |
      | **Cachet** | Self-hosted | PHP-based, mature |
      | **Instatus** | SaaS | Modern, integrations |
      | **Statuspage (Atlassian)** | SaaS | Enterprise, expensive |
      
      **Status page best practices:**
      - Show individual component status (API, database, CDN, auth)
      - Include historical uptime percentage (30/90 day)
      - Post incident updates promptly (investigating → identified → monitoring → resolved)
      - Subscribe option for email/SMS/RSS notifications
      
    • tracing.md 27.1 KB
      # Distributed Tracing Reference
      
      Comprehensive reference for OpenTelemetry, context propagation, sampling, and instrumentation patterns.
      
      ---
      
      ## OpenTelemetry Architecture
      
      ```
      ┌──────────────┐     ┌──────────────┐     ┌──────────────┐
      │  Application │     │  Application │     │  Application │
      │  (SDK + API) │     │  (SDK + API) │     │  (SDK + API) │
      └──────┬───────┘     └──────┬───────┘     └──────┬───────┘
             │ OTLP              │ OTLP              │ OTLP
             ▼                   ▼                   ▼
      ┌─────────────────────────────────────────────────────────┐
      │                   OTel Collector                         │
      │  ┌───────────┐  ┌────────────┐  ┌───────────────────┐  │
      │  │ Receivers │→ │ Processors │→ │    Exporters      │  │
      │  │ (OTLP,    │  │ (batch,    │  │ (Jaeger, Tempo,   │  │
      │  │  Jaeger,  │  │  filter,   │  │  Datadog, OTLP)   │  │
      │  │  Zipkin)  │  │  tail      │  │                   │  │
      │  │           │  │  sampling) │  │                   │  │
      │  └───────────┘  └────────────┘  └───────────────────┘  │
      └─────────────────────────────────────────────────────────┘
             │                    │                    │
             ▼                    ▼                    ▼
      ┌──────────┐        ┌──────────┐        ┌──────────────┐
      │  Jaeger  │        │  Tempo   │        │   Datadog    │
      │  (UI)    │        │  (store) │        │   (SaaS)     │
      └──────────┘        └──────────┘        └──────────────┘
      ```
      
      ### Components
      
      | Component | Role | Notes |
      |-----------|------|-------|
      | **API** | Stable interfaces for instrumentation | Language-specific, vendor-neutral |
      | **SDK** | Implementation of the API | Configures sampling, export, processing |
      | **Collector** | Receives, processes, exports telemetry | Deploy as sidecar or gateway |
      | **Exporters** | Send data to backends | OTLP (preferred), Jaeger, Zipkin, vendor-specific |
      | **Auto-instrumentation** | Automatic span creation for frameworks | HTTP, gRPC, database, messaging |
      
      ### Collector Configuration
      
      ```yaml
      # otel-collector-config.yml
      receivers:
        otlp:
          protocols:
            grpc:
              endpoint: 0.0.0.0:4317
            http:
              endpoint: 0.0.0.0:4318
      
      processors:
        batch:
          timeout: 5s
          send_batch_size: 8192
          send_batch_max_size: 16384
      
        memory_limiter:
          check_interval: 1s
          limit_mib: 1024
          spike_limit_mib: 256
      
        # Tail-based sampling (decide after seeing complete trace)
        tail_sampling:
          decision_wait: 10s
          num_traces: 100000
          policies:
            # Always sample errors
            - name: errors
              type: status_code
              status_code:
                status_codes: [ERROR]
            # Always sample slow traces (> 2s)
            - name: slow-traces
              type: latency
              latency:
                threshold_ms: 2000
            # Sample 10% of everything else
            - name: probabilistic
              type: probabilistic
              probabilistic:
                sampling_percentage: 10
      
        # Add resource attributes
        resource:
          attributes:
            - key: environment
              value: production
              action: upsert
      
      exporters:
        otlp/jaeger:
          endpoint: jaeger:4317
          tls:
            insecure: true
      
        otlp/tempo:
          endpoint: tempo:4317
          tls:
            insecure: true
      
        debug:
          verbosity: detailed
      
      service:
        pipelines:
          traces:
            receivers: [otlp]
            processors: [memory_limiter, tail_sampling, batch, resource]
            exporters: [otlp/jaeger]
          metrics:
            receivers: [otlp]
            processors: [memory_limiter, batch]
            exporters: [otlp/tempo]
          logs:
            receivers: [otlp]
            processors: [memory_limiter, batch]
            exporters: [debug]
      ```
      
      ---
      
      ## Span Model
      
      ### Span Anatomy
      
      ```
      Trace ID: 4bf92f3577b34da6a3ce929d0e0e4736
      │
      ├─ Span: "GET /api/orders"
      │  ├─ Span ID: 00f067aa0ba902b7
      │  ├─ Parent: (none - root span)
      │  ├─ Start: 2026-03-09T14:32:01.000Z
      │  ├─ End:   2026-03-09T14:32:01.245Z
      │  ├─ Status: OK
      │  ├─ Attributes:
      │  │   http.method: GET
      │  │   http.url: /api/orders?user_id=789
      │  │   http.status_code: 200
      │  │   http.response_content_length: 4523
      │  ├─ Events:
      │  │   └─ "cache.miss" at T+5ms {key: "orders:usr-789"}
      │  │
      │  ├─ Span: "SELECT orders"
      │  │  ├─ Span ID: a1b2c3d4e5f60718
      │  │  ├─ Parent: 00f067aa0ba902b7
      │  │  ├─ Duration: 45ms
      │  │  ├─ Attributes:
      │  │  │   db.system: postgresql
      │  │  │   db.operation: SELECT
      │  │  │   db.statement: SELECT * FROM orders WHERE user_id = $1
      │  │  │   db.rows_affected: 12
      │  │  └─ Status: OK
      │  │
      │  └─ Span: "GET payment-service/status"
      │     ├─ Span ID: b2c3d4e5f6071829
      │     ├─ Parent: 00f067aa0ba902b7
      │     ├─ Duration: 120ms
      │     ├─ Attributes:
      │     │   http.method: GET
      │     │   http.url: http://payment-service:8080/status
      │     │   http.status_code: 200
      │     │   peer.service: payment-service
      │     └─ Status: OK
      ```
      
      ### Span Attributes (Semantic Conventions)
      
      #### HTTP Spans
      
      | Attribute | Example | Notes |
      |-----------|---------|-------|
      | `http.request.method` | `GET` | HTTP method |
      | `url.path` | `/api/orders` | URL path |
      | `http.response.status_code` | `200` | Response status |
      | `http.request.body.size` | `1024` | Request body bytes |
      | `http.response.body.size` | `4523` | Response body bytes |
      | `server.address` | `api.example.com` | Server hostname |
      | `server.port` | `443` | Server port |
      | `network.protocol.version` | `1.1` | HTTP version |
      | `user_agent.original` | `Mozilla/5.0...` | User agent string |
      
      #### Database Spans
      
      | Attribute | Example | Notes |
      |-----------|---------|-------|
      | `db.system` | `postgresql` | Database type |
      | `db.namespace` | `myapp` | Database name |
      | `db.operation.name` | `SELECT` | SQL operation |
      | `db.query.text` | `SELECT * FROM...` | Sanitized query |
      | `server.address` | `db.example.com` | DB host |
      | `server.port` | `5432` | DB port |
      | `db.response.rows_affected` | `12` | Rows returned/affected |
      
      #### gRPC Spans
      
      | Attribute | Example | Notes |
      |-----------|---------|-------|
      | `rpc.system` | `grpc` | RPC system |
      | `rpc.service` | `myapp.UserService` | Service name |
      | `rpc.method` | `GetUser` | Method name |
      | `rpc.grpc.status_code` | `0` | gRPC status code |
      
      ### Span Status
      
      | Status | When | Notes |
      |--------|------|-------|
      | `UNSET` | Default | Operation completed, no explicit status |
      | `OK` | Explicitly successful | Use sparingly, UNSET is fine for success |
      | `ERROR` | Operation failed | Always set for 5xx responses, exceptions |
      
      ---
      
      ## SDK Setup
      
      ### Go
      
      ```go
      package main
      
      import (
          "context"
          "log"
      
          "go.opentelemetry.io/otel"
          "go.opentelemetry.io/otel/attribute"
          "go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc"
          "go.opentelemetry.io/otel/propagation"
          "go.opentelemetry.io/otel/sdk/resource"
          sdktrace "go.opentelemetry.io/otel/sdk/trace"
          semconv "go.opentelemetry.io/otel/semconv/v1.24.0"
          "go.opentelemetry.io/otel/trace"
      )
      
      func initTracer(ctx context.Context) (*sdktrace.TracerProvider, error) {
          // OTLP gRPC exporter (sends to Collector)
          exporter, err := otlptracegrpc.New(ctx,
              otlptracegrpc.WithEndpoint("otel-collector:4317"),
              otlptracegrpc.WithInsecure(),
          )
          if err != nil {
              return nil, err
          }
      
          // Resource: describes this service
          res, err := resource.Merge(
              resource.Default(),
              resource.NewWithAttributes(
                  semconv.SchemaURL,
                  semconv.ServiceName("order-service"),
                  semconv.ServiceVersion("1.4.2"),
                  attribute.String("environment", "production"),
              ),
          )
          if err != nil {
              return nil, err
          }
      
          // TracerProvider with batch span processor
          tp := sdktrace.NewTracerProvider(
              sdktrace.WithBatcher(exporter),
              sdktrace.WithResource(res),
              sdktrace.WithSampler(sdktrace.ParentBased(
                  sdktrace.TraceIDRatioBased(0.1), // 10% head sampling
              )),
          )
      
          // Set global TracerProvider and propagator
          otel.SetTracerProvider(tp)
          otel.SetTextMapPropagator(propagation.NewCompositeTextMapPropagator(
              propagation.TraceContext{},  // W3C TraceContext
              propagation.Baggage{},       // W3C Baggage
          ))
      
          return tp, nil
      }
      
      func main() {
          ctx := context.Background()
          tp, err := initTracer(ctx)
          if err != nil {
              log.Fatal(err)
          }
          defer tp.Shutdown(ctx)
      
          // Create spans
          tracer := otel.Tracer("order-service")
      
          ctx, span := tracer.Start(ctx, "ProcessOrder",
              trace.WithAttributes(
                  attribute.String("order.id", "ord-123"),
                  attribute.Int("order.items", 3),
              ),
          )
          defer span.End()
      
          // Add events
          span.AddEvent("order.validated", trace.WithAttributes(
              attribute.Bool("has_discount", true),
          ))
      
          // Record errors
          if err := processPayment(ctx); err != nil {
              span.RecordError(err)
              span.SetStatus(codes.Error, err.Error())
          }
      }
      ```
      
      ### Go Auto-instrumentation
      
      ```go
      import (
          "go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp"
          "go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc"
          "go.opentelemetry.io/contrib/instrumentation/github.com/jackc/pgx/v5/otelpgx"
      )
      
      // HTTP server: wrap handler
      mux := http.NewServeMux()
      mux.HandleFunc("/api/orders", handleOrders)
      handler := otelhttp.NewHandler(mux, "server")
      http.ListenAndServe(":8080", handler)
      
      // HTTP client: wrap transport
      client := &http.Client{
          Transport: otelhttp.NewTransport(http.DefaultTransport),
      }
      
      // gRPC server: add interceptors
      server := grpc.NewServer(
          grpc.UnaryInterceptor(otelgrpc.UnaryServerInterceptor()),
          grpc.StreamInterceptor(otelgrpc.StreamServerInterceptor()),
      )
      
      // gRPC client: add interceptors
      conn, _ := grpc.Dial(addr,
          grpc.WithUnaryInterceptor(otelgrpc.UnaryClientInterceptor()),
          grpc.WithStreamInterceptor(otelgrpc.StreamClientInterceptor()),
      )
      
      // pgx (PostgreSQL): add tracer
      config, _ := pgxpool.ParseConfig(databaseURL)
      config.ConnConfig.Tracer = otelpgx.NewTracer()
      pool, _ := pgxpool.NewWithConfig(ctx, config)
      ```
      
      ### Python
      
      ```python
      from opentelemetry import trace
      from opentelemetry.sdk.trace import TracerProvider
      from opentelemetry.sdk.trace.export import BatchSpanProcessor
      from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
      from opentelemetry.sdk.resources import Resource, SERVICE_NAME, SERVICE_VERSION
      from opentelemetry.propagators.composite import CompositePropagator
      from opentelemetry.propagators.textmap import DefaultTextMapPropagator
      from opentelemetry.trace.propagation.tracecontext import TraceContextTextMapPropagator
      from opentelemetry.baggage.propagation import W3CBaggagePropagator
      
      # Resource
      resource = Resource.create({
          SERVICE_NAME: "order-service",
          SERVICE_VERSION: "1.4.2",
          "environment": "production",
      })
      
      # TracerProvider
      provider = TracerProvider(resource=resource)
      provider.add_span_processor(
          BatchSpanProcessor(
              OTLPSpanExporter(endpoint="otel-collector:4317", insecure=True)
          )
      )
      trace.set_tracer_provider(provider)
      
      # Propagator
      from opentelemetry import propagate
      propagate.set_global_textmap(CompositePropagator([
          TraceContextTextMapPropagator(),
          W3CBaggagePropagator(),
      ]))
      
      # Create spans
      tracer = trace.get_tracer("order-service")
      
      with tracer.start_as_current_span("process_order", attributes={
          "order.id": "ord-123",
          "order.items": 3,
      }) as span:
          span.add_event("order.validated", {"has_discount": True})
      
          try:
              process_payment(order)
          except Exception as e:
              span.record_exception(e)
              span.set_status(trace.Status(trace.StatusCode.ERROR, str(e)))
              raise
      ```
      
      ### Python Auto-instrumentation
      
      ```bash
      # Install auto-instrumentation packages
      pip install opentelemetry-distro opentelemetry-exporter-otlp
      opentelemetry-bootstrap -a install  # Installs all detected instrumentors
      
      # Run with auto-instrumentation
      opentelemetry-instrument \
          --service_name order-service \
          --exporter_otlp_endpoint http://otel-collector:4317 \
          python app.py
      ```
      
      ```python
      # Or configure programmatically
      from opentelemetry.instrumentation.flask import FlaskInstrumentor
      from opentelemetry.instrumentation.requests import RequestsInstrumentor
      from opentelemetry.instrumentation.psycopg2 import Psycopg2Instrumentor
      from opentelemetry.instrumentation.redis import RedisInstrumentor
      
      FlaskInstrumentor().instrument_app(app)
      RequestsInstrumentor().instrument()
      Psycopg2Instrumentor().instrument()
      RedisInstrumentor().instrument()
      ```
      
      ### Node.js
      
      ```javascript
      // tracing.js - import BEFORE other modules
      const { NodeSDK } = require('@opentelemetry/sdk-node');
      const { OTLPTraceExporter } = require('@opentelemetry/exporter-trace-otlp-grpc');
      const { getNodeAutoInstrumentations } = require('@opentelemetry/auto-instrumentations-node');
      const { Resource } = require('@opentelemetry/resources');
      const { ATTR_SERVICE_NAME, ATTR_SERVICE_VERSION } = require('@opentelemetry/semantic-conventions');
      
      const sdk = new NodeSDK({
        resource: new Resource({
          [ATTR_SERVICE_NAME]: 'order-service',
          [ATTR_SERVICE_VERSION]: '1.4.2',
          environment: 'production',
        }),
        traceExporter: new OTLPTraceExporter({
          url: 'http://otel-collector:4317',
        }),
        instrumentations: [
          getNodeAutoInstrumentations({
            // Disable fs instrumentation (too noisy)
            '@opentelemetry/instrumentation-fs': { enabled: false },
          }),
        ],
      });
      
      sdk.start();
      
      // Graceful shutdown
      process.on('SIGTERM', () => {
        sdk.shutdown().then(() => process.exit(0));
      });
      ```
      
      ```javascript
      // Manual span creation
      const { trace } = require('@opentelemetry/api');
      
      const tracer = trace.getTracer('order-service');
      
      async function processOrder(orderId) {
        return tracer.startActiveSpan('process_order', {
          attributes: { 'order.id': orderId },
        }, async (span) => {
          try {
            span.addEvent('order.validated');
            await processPayment(orderId);
            span.setStatus({ code: SpanStatusCode.OK });
          } catch (err) {
            span.recordException(err);
            span.setStatus({ code: SpanStatusCode.ERROR, message: err.message });
            throw err;
          } finally {
            span.end();
          }
        });
      }
      ```
      
      ---
      
      ## Context Propagation
      
      ### W3C TraceContext
      
      The standard for propagating trace context across service boundaries.
      
      **Request headers:**
      ```
      traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
      tracestate: vendor1=value1,vendor2=value2
      ```
      
      **Format:** `version-trace_id-parent_id-trace_flags`
      
      | Field | Size | Description |
      |-------|------|-------------|
      | version | 2 hex | Always `00` |
      | trace_id | 32 hex | Unique trace identifier |
      | parent_id | 16 hex | Span ID of the caller |
      | trace_flags | 2 hex | `01` = sampled, `00` = not sampled |
      
      ### B3 Propagation (Zipkin)
      
      Legacy format still used by some systems:
      
      ```
      # Single header (compact)
      b3: 4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-1
      
      # Multi-header
      X-B3-TraceId: 4bf92f3577b34da6a3ce929d0e0e4736
      X-B3-SpanId: 00f067aa0ba902b7
      X-B3-ParentSpanId: (parent span)
      X-B3-Sampled: 1
      ```
      
      ### Baggage
      
      Propagate arbitrary key-value pairs across service boundaries:
      
      ```
      baggage: userId=usr-789,region=us-east-1
      ```
      
      ```go
      // Go: set baggage
      bag, _ := baggage.Parse("userId=usr-789,region=us-east-1")
      ctx = baggage.ContextWithBaggage(ctx, bag)
      
      // Go: read baggage
      bag := baggage.FromContext(ctx)
      userId := bag.Member("userId").Value()
      ```
      
      **Use sparingly:** Baggage is sent with every request. Don't put large values or sensitive data in baggage.
      
      ---
      
      ## Sampling Strategies
      
      ### Head-based Sampling
      
      Decision made at trace creation, propagated to all downstream services.
      
      | Sampler | Description | Config |
      |---------|-------------|--------|
      | `AlwaysOn` | Sample everything | Development only |
      | `AlwaysOff` | Sample nothing | Disable tracing |
      | `TraceIDRatioBased` | Sample N% of traces | `ratio: 0.1` for 10% |
      | `ParentBased` | Follow parent's decision | Default, wrap another sampler |
      
      ```go
      // Go: 10% sampling, respecting parent's decision
      sampler := sdktrace.ParentBased(
          sdktrace.TraceIDRatioBased(0.1),
      )
      ```
      
      ```python
      # Python: 10% sampling
      from opentelemetry.sdk.trace.sampling import TraceIdRatioBased, ParentBasedTraceIdRatio
      sampler = ParentBasedTraceIdRatio(0.1)
      ```
      
      ### Tail-based Sampling
      
      Decision made after the trace completes. Requires the OTel Collector.
      
      **Advantages:**
      - Always captures error traces
      - Always captures slow traces
      - More representative sampling
      
      **Configuration (Collector):**
      
      ```yaml
      processors:
        tail_sampling:
          decision_wait: 10s         # Wait for spans to arrive
          num_traces: 100000         # Max traces in memory
          expected_new_traces_per_sec: 1000
          policies:
            # Always keep errors
            - name: errors-policy
              type: status_code
              status_code:
                status_codes: [ERROR]
      
            # Always keep slow traces (root span > 2s)
            - name: latency-policy
              type: latency
              latency:
                threshold_ms: 2000
      
            # Keep all traces for specific operations
            - name: critical-operations
              type: string_attribute
              string_attribute:
                key: operation
                values: [payment, refund, account_deletion]
      
            # Sample 5% of remaining traces
            - name: probabilistic-policy
              type: probabilistic
              probabilistic:
                sampling_percentage: 5
      
            # Composite: apply multiple policies with priority
            - name: composite-policy
              type: composite
              composite:
                max_total_spans_per_second: 1000
                policy_order: [errors-policy, latency-policy, probabilistic-policy]
                rate_allocation:
                  - policy: errors-policy
                    percent: 50
                  - policy: latency-policy
                    percent: 30
                  - policy: probabilistic-policy
                    percent: 20
      ```
      
      ---
      
      ## Jaeger
      
      ### Deployment
      
      #### All-in-One (Development)
      
      ```yaml
      # docker-compose.yml
      services:
        jaeger:
          image: jaegertracing/all-in-one:1.54
          ports:
            - "16686:16686"   # UI
            - "4317:4317"     # OTLP gRPC
            - "4318:4318"     # OTLP HTTP
            - "14250:14250"   # Jaeger gRPC
          environment:
            COLLECTOR_OTLP_ENABLED: true
      ```
      
      #### Production (with Elasticsearch)
      
      ```yaml
      services:
        jaeger-collector:
          image: jaegertracing/jaeger-collector:1.54
          environment:
            SPAN_STORAGE_TYPE: elasticsearch
            ES_SERVER_URLS: http://elasticsearch:9200
            COLLECTOR_OTLP_ENABLED: true
          ports:
            - "4317:4317"
            - "14250:14250"
      
        jaeger-query:
          image: jaegertracing/jaeger-query:1.54
          environment:
            SPAN_STORAGE_TYPE: elasticsearch
            ES_SERVER_URLS: http://elasticsearch:9200
          ports:
            - "16686:16686"
      
        elasticsearch:
          image: elasticsearch:8.12.0
          environment:
            discovery.type: single-node
            xpack.security.enabled: false
            ES_JAVA_OPTS: "-Xms512m -Xmx512m"
          volumes:
            - es-data:/usr/share/elasticsearch/data
      
      volumes:
        es-data:
      ```
      
      ### Jaeger UI Features
      
      | Feature | Description |
      |---------|-------------|
      | **Search** | Find traces by service, operation, tags, duration, time range |
      | **Trace View** | Waterfall visualization of spans with timing |
      | **Compare** | Compare two traces side by side |
      | **Dependencies** | Service dependency graph (DAG) |
      | **Deep Dependency** | Trace-aware dependency analysis |
      | **Monitor** | RED metrics derived from traces |
      
      ---
      
      ## Async Trace Propagation
      
      ### Go: context.Context
      
      ```go
      // Context carries trace information automatically
      func processOrder(ctx context.Context, orderID string) error {
          // Start child span — automatically linked to parent via ctx
          ctx, span := tracer.Start(ctx, "processOrder")
          defer span.End()
      
          // Pass context to goroutines
          g, ctx := errgroup.WithContext(ctx)
          g.Go(func() error {
              return validateOrder(ctx, orderID)  // ctx carries trace
          })
          g.Go(func() error {
              return checkInventory(ctx, orderID) // ctx carries trace
          })
          return g.Wait()
      }
      ```
      
      ### Python: contextvars
      
      ```python
      import asyncio
      from opentelemetry import trace, context
      
      tracer = trace.get_tracer("order-service")
      
      async def process_order(order_id: str):
          with tracer.start_as_current_span("process_order") as span:
              # asyncio tasks automatically inherit context
              results = await asyncio.gather(
                  validate_order(order_id),
                  check_inventory(order_id),
              )
              return results
      
      async def validate_order(order_id: str):
          # This span is automatically a child of process_order
          with tracer.start_as_current_span("validate_order"):
              pass
      ```
      
      ### Node.js: AsyncLocalStorage
      
      ```javascript
      const { trace, context } = require('@opentelemetry/api');
      
      // OpenTelemetry SDK uses AsyncLocalStorage internally
      // Spans are automatically propagated through async operations
      
      async function processOrder(orderId) {
        return tracer.startActiveSpan('process_order', async (span) => {
          try {
            // Promise.all preserves context
            const [validation, inventory] = await Promise.all([
              validateOrder(orderId),  // context propagated
              checkInventory(orderId), // context propagated
            ]);
            return { validation, inventory };
          } finally {
            span.end();
          }
        });
      }
      ```
      
      ---
      
      ## Database Query Tracing
      
      ### Span Attributes for Database Queries
      
      ```json
      {
        "name": "SELECT users",
        "attributes": {
          "db.system": "postgresql",
          "db.namespace": "myapp",
          "db.operation.name": "SELECT",
          "db.query.text": "SELECT id, name, email FROM users WHERE id = $1",
          "server.address": "db.example.com",
          "server.port": 5432,
          "db.response.rows_affected": 1
        }
      }
      ```
      
      ### Query Parameter Sanitization
      
      **Never log actual parameter values** — they may contain PII:
      
      ```go
      // GOOD: sanitized
      span.SetAttributes(
          attribute.String("db.query.text", "SELECT * FROM users WHERE id = $1 AND email = $2"),
      )
      
      // BAD: contains PII
      span.SetAttributes(
          attribute.String("db.query.text", "SELECT * FROM users WHERE id = 789 AND email = 'john@example.com'"),
      )
      ```
      
      ### N+1 Query Detection
      
      Use traces to identify N+1 queries — they appear as many identical DB spans under one parent:
      
      ```
      GET /api/orders (250ms)
      ├── SELECT orders WHERE user_id = $1 (5ms)
      ├── SELECT product WHERE id = $1 (3ms)    ← N+1
      ├── SELECT product WHERE id = $1 (3ms)    ← N+1
      ├── SELECT product WHERE id = $1 (4ms)    ← N+1
      ├── SELECT product WHERE id = $1 (3ms)    ← N+1
      └── ... (20 more identical queries)
      ```
      
      **Fix:** Use `SELECT * FROM products WHERE id IN ($1, $2, ...)` or JOIN.
      
      ---
      
      ## HTTP Client/Server Tracing
      
      ### Automatic Instrumentation
      
      Most OpenTelemetry auto-instrumentation libraries create spans automatically for:
      - HTTP server requests (incoming)
      - HTTP client requests (outgoing)
      - gRPC server/client calls
      - Database queries
      - Redis operations
      - Message queue operations (Kafka, RabbitMQ)
      
      ### Custom Attributes on Auto-instrumented Spans
      
      ```go
      // Add business context to auto-created spans
      span := trace.SpanFromContext(ctx)
      span.SetAttributes(
          attribute.String("user.id", userID),
          attribute.String("order.id", orderID),
          attribute.String("tenant.id", tenantID),
      )
      ```
      
      ### Error Recording
      
      ```go
      // Proper error recording
      if err != nil {
          span.RecordError(err)  // Creates an event with exception details
          span.SetStatus(codes.Error, err.Error())  // Marks span as error
          return err
      }
      
      // For HTTP handlers
      if statusCode >= 500 {
          span.SetStatus(codes.Error, fmt.Sprintf("HTTP %d", statusCode))
      }
      // Note: 4xx is NOT an error from the server's perspective
      ```
      
      ---
      
      ## gRPC Tracing
      
      ### Server Interceptors
      
      ```go
      import "go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc"
      
      server := grpc.NewServer(
          grpc.StatsHandler(otelgrpc.NewServerHandler()),
      )
      ```
      
      ### Client Interceptors
      
      ```go
      conn, err := grpc.Dial(addr,
          grpc.WithStatsHandler(otelgrpc.NewClientHandler()),
      )
      ```
      
      ### Streaming Spans
      
      For streaming RPCs, spans cover the entire stream lifecycle:
      
      ```
      ServerStream (user.ListUsers) — 2.5s
      ├─ stream.message.sent (1) — T+10ms
      ├─ stream.message.sent (2) — T+50ms
      ├─ stream.message.sent (3) — T+120ms
      └─ stream.message.sent (4) — T+200ms
      ```
      
      ### Metadata Propagation
      
      Context is automatically propagated via gRPC metadata when using OTel interceptors:
      
      ```go
      // Automatic: OTel interceptors inject/extract from gRPC metadata
      // metadata equivalent to HTTP headers:
      //   traceparent → grpc-metadata-traceparent
      //   tracestate  → grpc-metadata-tracestate
      ```
      
      ---
      
      ## Trace-Based Testing
      
      ### Asserting Span Structure
      
      ```go
      // Go: using in-memory exporter for testing
      import (
          sdktrace "go.opentelemetry.io/otel/sdk/trace"
          "go.opentelemetry.io/otel/sdk/trace/tracetest"
      )
      
      func TestOrderProcessing(t *testing.T) {
          exporter := tracetest.NewInMemoryExporter()
          tp := sdktrace.NewTracerProvider(
              sdktrace.WithSyncer(exporter),
          )
          otel.SetTracerProvider(tp)
      
          // Run the operation
          processOrder(context.Background(), "ord-123")
      
          // Assert spans
          spans := exporter.GetSpans()
          assert.Len(t, spans, 3)
      
          rootSpan := spans[0]
          assert.Equal(t, "processOrder", rootSpan.Name)
          assert.Equal(t, codes.Ok, rootSpan.Status.Code)
      
          dbSpan := spans[1]
          assert.Equal(t, "SELECT orders", dbSpan.Name)
          assert.Equal(t, "postgresql", dbSpan.Attributes["db.system"])
      
          // Verify parent-child relationship
          assert.Equal(t, rootSpan.SpanContext.SpanID(), dbSpan.Parent.SpanID())
      }
      ```
      
      ```python
      # Python: using in-memory exporter
      from opentelemetry.sdk.trace.export.in_memory import InMemorySpanExporter
      
      exporter = InMemorySpanExporter()
      provider = TracerProvider()
      provider.add_span_processor(SimpleSpanProcessor(exporter))
      trace.set_tracer_provider(provider)
      
      # Run operation
      process_order("ord-123")
      
      # Assert
      spans = exporter.get_finished_spans()
      assert len(spans) == 3
      assert spans[0].name == "process_order"
      assert spans[1].name == "SELECT orders"
      assert spans[1].parent.span_id == spans[0].context.span_id
      ```
      
      ### Verifying Context Propagation
      
      ```go
      func TestContextPropagation(t *testing.T) {
          // Create a trace context
          ctx, span := tracer.Start(context.Background(), "test-root")
          traceID := span.SpanContext().TraceID()
      
          // Call service that makes outbound HTTP call
          handler.ServeHTTP(recorder, req.WithContext(ctx))
      
          // Verify all spans share the same trace ID
          spans := exporter.GetSpans()
          for _, s := range spans {
              assert.Equal(t, traceID, s.SpanContext.TraceID())
          }
      
          span.End()
      }
      ```
      
  • scripts
    • .gitkeep 0 B · in bundle
  • SKILL.md 16.1 KB
    ---
    name: monitoring-ops
    description: "Observability patterns - metrics, logging, tracing, alerting, and infrastructure monitoring. Use for: monitoring, observability, prometheus, grafana, metrics, alerting, structured logging, distributed tracing, opentelemetry, SLO, SLI, dashboard, health check, loki, jaeger, datadog, pagerduty."
    license: MIT
    allowed-tools: "Read Write Bash"
    metadata:
      author: claude-mods
      related-skills: python-observability-ops, docker-ops, ci-cd-ops, nginx-ops
    ---
    
    # Monitoring Operations
    
    Comprehensive observability patterns covering the three pillars (metrics, logging, tracing), alerting strategies, dashboard design, and infrastructure monitoring for production systems.
    
    ---
    
    ## Three Pillars Quick Reference
    
    Use this table to decide which observability signal fits your need:
    
    | Pillar | Best For | Tools | Data Type |
    |--------|----------|-------|-----------|
    | **Metrics** | Aggregated numeric measurements, trends, alerting on thresholds | Prometheus, Datadog, CloudWatch, StatsD | Time-series (numeric) |
    | **Logs** | Discrete events, error details, audit trails, debugging context | Loki, ELK, CloudWatch Logs, Fluentd | Unstructured/structured text |
    | **Traces** | Request flow across services, latency breakdown, dependency mapping | Jaeger, Tempo, Zipkin, Datadog APM | Span trees (structured) |
    
    **When to use which:**
    
    - **"How many requests per second?"** → Metrics (counter + rate)
    - **"Why did this specific request fail?"** → Logs (error message + stack trace)
    - **"Where is the latency in this request?"** → Traces (span waterfall)
    - **"Is the system healthy right now?"** → Metrics (gauges + alerts)
    - **"What happened at 3:42 AM?"** → Logs (timestamped event search)
    - **"Which downstream service caused the timeout?"** → Traces (span analysis)
    
    **Correlation is key:** Connect all three by embedding `trace_id` in log entries, recording exemplars in metrics, and linking trace spans to log queries.
    
    ---
    
    ## Metrics Type Decision Tree
    
    Use this tree to select the correct metric type:
    
    ```
    What are you measuring?
    │
    ├─ A count of events that only goes up?
    │  └─ COUNTER
    │     Examples: http_requests_total, errors_total, bytes_sent_total
    │     Use rate() or increase() to get per-second or per-interval values
    │     Never use a counter's raw value — it resets on restart
    │
    ├─ A current value that goes up AND down?
    │  └─ GAUGE
    │     Examples: temperature_celsius, active_connections, queue_depth
    │     Use for snapshots of current state
    │     Can use avg_over_time(), max_over_time() for trends
    │
    ├─ A distribution of values (latency, size)?
    │  │
    │  ├─ Need aggregatable quantiles across instances?
    │  │  └─ HISTOGRAM
    │  │     Examples: http_request_duration_seconds, response_size_bytes
    │  │     Define buckets: [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10]
    │  │     Use histogram_quantile() for percentiles (p50, p95, p99)
    │  │     Aggregatable across instances (histograms can be summed)
    │  │
    │  └─ Need pre-calculated quantiles on a single instance?
    │     └─ SUMMARY
    │        Examples: go_gc_duration_seconds
    │        Pre-calculates quantiles client-side
    │        NOT aggregatable across instances
    │        Prefer histogram unless you have a specific reason
    │
    └─ None of the above?
       └─ INFO metric (labels only, value=1)
          Examples: build_info{version="1.2.3", commit="abc123"}
          Use for metadata exposed as metrics
    ```
    
    **Rule of thumb:** Start with counters and histograms. Add gauges for current state. Avoid summaries unless you have a compelling reason.
    
    ---
    
    ## Alerting Decision Tree
    
    ```
    What type of alert do you need?
    │
    ├─ Known threshold with a fixed boundary?
    │  └─ THRESHOLD-BASED
    │     Example: CPU > 90% for 5 minutes
    │     Pros: Simple, predictable, easy to understand
    │     Cons: Requires manual tuning, doesn't adapt to patterns
    │     Best for: Resource limits, error rate spikes, queue depth
    │
    ├─ Normal behavior varies by time/season?
    │  └─ ANOMALY-BASED
    │     Example: Traffic 3 standard deviations below normal for this hour
    │     Pros: Adapts to patterns, catches novel failures
    │     Cons: Noisy during transitions, requires training data
    │     Best for: Traffic patterns, business metrics, gradual degradation
    │
    └─ Defined reliability targets?
       └─ SLO-BASED (PREFERRED)
          Example: Error budget burn rate > 14.4x for 1 hour
          Pros: Aligned with user impact, reduces noise, principled
          Cons: Requires SLI/SLO definition, more complex setup
          Best for: User-facing services, platform reliability
    ```
    
    ### Severity Levels
    
    | Severity | Response | Examples | Routing |
    |----------|----------|----------|---------|
    | **Critical (P1)** | Page on-call immediately | Service down, data loss risk, security breach | PagerDuty high-urgency, phone call |
    | **Warning (P2)** | Investigate within hours | Elevated error rate, disk 80% full, SLO burn rate elevated | PagerDuty low-urgency, Slack alert channel |
    | **Info (P3)** | Review next business day | Deployment completed, certificate expiring in 30 days | Slack info channel, ticket auto-created |
    
    ### When to Page vs When to Ticket
    
    **Page (wake someone up) when:**
    - Users are currently impacted
    - Data loss is occurring or imminent
    - Security incident is active
    - Error budget will be exhausted within hours
    
    **Create ticket (don't page) when:**
    - Issue is not user-facing yet
    - Automated remediation is possible
    - Degradation is slow and has runway
    - Issue is during business hours and can be triaged normally
    
    ---
    
    ## Structured Logging Quick Reference
    
    ### Standard JSON Log Format
    
    ```json
    {
      "timestamp": "2026-03-09T14:32:01.123Z",
      "level": "ERROR",
      "message": "Failed to process payment",
      "service": "payment-api",
      "version": "1.4.2",
      "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
      "span_id": "00f067aa0ba902b7",
      "request_id": "req-abc123",
      "user_id": "usr-789",
      "error": {
        "type": "PaymentGatewayTimeout",
        "message": "Gateway response timeout after 30s",
        "stack": "..."
      },
      "duration_ms": 30042,
      "http": {
        "method": "POST",
        "path": "/api/v1/payments",
        "status_code": 504
      }
    }
    ```
    
    ### Log Level Decision Guide
    
    | Level | When to Use | Examples |
    |-------|-------------|---------|
    | **DEBUG** | Development only, verbose internal state | Variable values, SQL queries, cache hits/misses |
    | **INFO** | Normal operations worth recording | Request completed, job started/finished, config loaded |
    | **WARN** | Degraded but still functioning | Retry succeeded, fallback used, approaching limit |
    | **ERROR** | Operation failed, needs attention | Payment failed, API call error, constraint violation |
    | **FATAL** | Process cannot continue, must exit | Database unreachable at startup, invalid config, OOM |
    
    **Rules:**
    - Never log at ERROR for expected conditions (user input validation → WARN)
    - Every ERROR should be actionable — if no one will act on it, use WARN
    - DEBUG should be off in production by default
    - INFO should not be noisy — 1-5 log lines per request, not 50
    
    ### Correlation IDs
    
    - Generate a `request_id` (UUID v4 or ULID) at the edge/gateway
    - Propagate through all internal services via headers (`X-Request-ID`)
    - Include `trace_id` and `span_id` from distributed tracing
    - Log all three IDs in every log entry for cross-referencing
    
    ---
    
    ## Distributed Tracing Quick Reference
    
    ### Core Concepts
    
    - **Trace:** End-to-end journey of a request across all services
    - **Span:** A single unit of work (HTTP call, DB query, function execution)
    - **Context propagation:** Passing trace/span IDs between services via headers
    
    ### W3C TraceContext Header
    
    ```
    traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
                  │  │                                  │                  │
                  │  │                                  │                  └─ flags (01=sampled)
                  │  │                                  └─ parent span ID (16 hex)
                  │  └─ trace ID (32 hex)
                  └─ version (00)
    ```
    
    ### Sampling Strategies
    
    | Strategy | How It Works | Use When |
    |----------|--------------|----------|
    | **Head-based (ratio)** | Decide at trace start, propagate decision | Low traffic, need predictable volume |
    | **Always-on** | Sample everything | Development, low-traffic services |
    | **Parent-based** | Follow parent's sampling decision | Default for most services |
    | **Tail-based** | Decide after trace completes (at Collector) | Need error/slow traces, high traffic |
    
    **Recommendation:** Use parent-based + tail-based at the Collector. This captures all error traces and slow traces while controlling volume.
    
    ### Trace ID in Logs
    
    Always include `trace_id` in structured log entries. This enables jumping from a log line to the full trace view:
    
    ```
    Log entry → trace_id → Jaeger/Tempo → full request waterfall
    ```
    
    ---
    
    ## Tool Selection Matrix
    
    | Feature | Prometheus + Grafana | Datadog | Grafana Cloud | CloudWatch |
    |---------|---------------------|---------|---------------|------------|
    | **Cost** | Free (infra costs) | $$$$ (per host/metric) | $$ (usage-based) | $$ (AWS-native) |
    | **Setup complexity** | High (self-managed) | Low (SaaS agent) | Medium (managed) | Low (AWS-native) |
    | **Metrics** | Prometheus (excellent) | Built-in (excellent) | Mimir (excellent) | Built-in (good) |
    | **Logs** | Loki (good) | Built-in (excellent) | Loki (good) | CloudWatch Logs (good) |
    | **Traces** | Jaeger/Tempo (good) | APM (excellent) | Tempo (good) | X-Ray (adequate) |
    | **Alerting** | Alertmanager (good) | Built-in (excellent) | Grafana Alerting (good) | CloudWatch Alarms (adequate) |
    | **Dashboards** | Grafana (excellent) | Built-in (excellent) | Grafana (excellent) | Dashboards (adequate) |
    | **Retention** | Configurable (unlimited) | 15 months default | Configurable | Up to 15 months |
    | **Multi-cloud** | Yes | Yes | Yes | AWS only |
    | **Best for** | Cost-conscious, control | Full-featured, enterprise | Open-source + managed | AWS-native shops |
    
    **Recommendation path:**
    - **Starting out / budget-conscious:** Prometheus + Grafana + Loki + Tempo (all free, self-hosted)
    - **Small team, want managed:** Grafana Cloud free tier (10k metrics, 50GB logs, 50GB traces)
    - **Enterprise, need everything:** Datadog (expensive but comprehensive)
    - **AWS-only shop:** CloudWatch + X-Ray (simplest if already on AWS)
    
    ---
    
    ## Dashboard Design
    
    ### USE Method (Infrastructure)
    
    For every resource (CPU, memory, disk, network):
    
    | Signal | Question | Metric Example |
    |--------|----------|----------------|
    | **Utilization** | How busy is it? | `node_cpu_seconds_total` (% busy) |
    | **Saturation** | How overloaded is it? | `node_load1` (run queue length) |
    | **Errors** | Are there error events? | `node_network_receive_errs_total` |
    
    ### RED Method (Services)
    
    For every service endpoint:
    
    | Signal | Question | Metric Example |
    |--------|----------|----------------|
    | **Rate** | How many requests per second? | `rate(http_requests_total[5m])` |
    | **Errors** | How many are failing? | `rate(http_requests_total{status=~"5.."}[5m])` |
    | **Duration** | How long do they take? | `histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))` |
    
    ### Four Golden Signals (Google SRE)
    
    | Signal | What to Measure | Alert Threshold Guidance |
    |--------|-----------------|--------------------------|
    | **Latency** | Time to serve a request (distinguish success vs error latency) | p99 > 2x baseline |
    | **Traffic** | Demand on the system (requests/sec, sessions, transactions) | Anomaly detection |
    | **Errors** | Rate of failed requests (explicit 5xx, implicit policy violations) | > 0.1% of traffic |
    | **Saturation** | How "full" the service is (CPU, memory, queue depth) | > 80% capacity |
    
    ### Dashboard Layout Best Practices
    
    1. **Top row:** Key health indicators (error rate, latency p99, availability %)
    2. **Second row:** Traffic and throughput (requests/sec, active users)
    3. **Third row:** Resource utilization (CPU, memory, disk, network)
    4. **Bottom rows:** Detailed breakdowns (by endpoint, by status code, by region)
    5. **Use variables:** Service, environment, time range as dropdown selectors
    6. **Include annotations:** Deployments, incidents, config changes as vertical markers
    
    ---
    
    ## Common Gotchas
    
    | Gotcha | Why It Happens | Fix |
    |--------|----------------|-----|
    | **Cardinality explosion** | Using unbounded label values (user ID, request path, query string) | Use bounded labels only; aggregate high-cardinality data in logs, not metrics |
    | **Alert fatigue** | Too many alerts, too sensitive thresholds, alerts on non-actionable symptoms | Require runbook for every alert; tune thresholds; use SLO-based alerting |
    | **Missing correlation IDs** | Logs, metrics, and traces not linked together | Include trace_id in all log entries; use exemplars in metrics |
    | **Sampling bias** | Head-based sampling drops error/slow traces at high sample rates | Use tail-based sampling at the Collector to always capture errors and slow traces |
    | **Log volume costs** | DEBUG or verbose INFO in production, logging full request/response bodies | Set production to INFO minimum; truncate large payloads; use sampling for verbose paths |
    | **Metric naming inconsistency** | Different teams use different naming conventions | Adopt OpenMetrics naming: `namespace_subsystem_unit_suffix` (e.g., `http_server_request_duration_seconds`) |
    | **Dashboard sprawl** | Everyone creates dashboards, nobody maintains them | Standardize with USE/RED templates; review quarterly; delete unused dashboards |
    | **SLO too aggressive** | Setting 99.99% availability without the budget or architecture for it | Start with 99.5% or 99.9%; tighten only when consistently meeting targets with margin |
    | **Missing baseline** | Alerting on absolute thresholds without understanding normal behavior | Collect 2-4 weeks of baseline data before setting alert thresholds |
    | **Over-instrumentation** | Instrumenting every function, creating too many spans/metrics | Instrument at service boundaries; use auto-instrumentation for HTTP/DB/gRPC; add manual spans selectively |
    | **Ignoring metric staleness** | Assuming a metric that stops reporting means zero | Use `absent()` or `up == 0` to detect missing scrapers; distinguish "zero" from "not reporting" |
    | **Alerting on cause not symptom** | Alerting on CPU usage instead of user-facing error rate | Alert on symptoms (error rate, latency); use cause metrics (CPU, memory) for investigation |
    | **No retention policy** | Storing all metrics/logs at full resolution forever | Define retention tiers: 15s resolution for 2 weeks, 1m for 3 months, 5m for 1 year |
    | **Dashboard without context** | Graphs with no units, no description, no threshold lines | Add units to Y-axis, threshold lines for SLOs, panel descriptions explaining what "good" looks like |
    
    ---
    
    ## Reference Files
    
    | File | Contents | Lines |
    |------|----------|-------|
    | [metrics-alerting.md](references/metrics-alerting.md) | Prometheus, Grafana, OpenTelemetry metrics, SLI/SLO/SLA, alert routing, runbooks, uptime monitoring | ~650 |
    | [logging.md](references/logging.md) | Structured logging, log levels, correlation IDs, aggregation (Loki, ELK), retention, PII masking, language-specific | ~550 |
    | [tracing.md](references/tracing.md) | OpenTelemetry, spans, context propagation, sampling, Jaeger, async tracing, DB/HTTP/gRPC instrumentation | ~600 |
    | [infrastructure.md](references/infrastructure.md) | Health checks, K8s probes, Docker HEALTHCHECK, infra metrics, APM, cost optimization, incident response | ~550 |
    
    ---
    
    ## See Also
    
    - **docker-ops** — Container monitoring with cAdvisor, Docker stats, and health checks
    - **ci-cd-ops** — Pipeline observability, deployment tracking, build metrics
    - **nginx-ops** — Nginx access/error log parsing, request metrics, upstream monitoring
    - **python-observability-ops** — Python-specific instrumentation with structlog, opentelemetry-python
    - [OpenTelemetry documentation](https://opentelemetry.io/docs/)
    - [Prometheus best practices](https://prometheus.io/docs/practices/)
    - [Google SRE Book — Monitoring chapter](https://sre.google/sre-book/monitoring-distributed-systems/)
    - [Grafana dashboards library](https://grafana.com/grafana/dashboards/)
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related