Claude Skill

python-observability-ops

Observability patterns for Python applications. Triggers on: logging, metrics, tracing, opentelemetry, prometheus, observability, monitoring, structlog, correlation id.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download 0xdarkmatter-claude-mods-skills_python-observability-ops-3dfaf0b.zip · 11 KB
Part of 0xdarkmatter/claude-mods — 94 skills

Install

skills CLI npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/python-observability-ops
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
Git git clone https://github.com/0xDarkMatter/claude-mods.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Python Observability Patterns

Logging, metrics, and tracing for production applications.

Structured Logging with structlog

import structlog

# Configure structlog
structlog.configure(
    processors=[
        structlog.contextvars.merge_contextvars,
        structlog.processors.add_log_level,
        structlog.processors.TimeStamper(fmt="iso"),
        structlog.processors.JSONRenderer(),
    ],
    wrapper_class=structlog.make_filtering_bound_logger(logging.INFO),
    context_class=dict,
    logger_factory=structlog.PrintLoggerFactory(),
)

logger = structlog.get_logger()

# Usage
logger.info("user_created", user_id=123, email="test@example.com")
# Output: {"event": "user_created", "user_id": 123, "email": "test@example.com", "level": "info", "timestamp": "2024-01-15T10:00:00Z"}

Request Context Propagation

import structlog
from contextvars import ContextVar
from uuid import uuid4

request_id_var: ContextVar[str] = ContextVar("request_id", default="")

def bind_request_context(request_id: str | None = None):
    """Bind request ID to logging context."""
    rid = request_id or str(uuid4())
    request_id_var.set(rid)
    structlog.contextvars.bind_contextvars(request_id=rid)
    return rid

# FastAPI middleware
@app.middleware("http")
async def request_context_middleware(request, call_next):
    request_id = request.headers.get("X-Request-ID") or str(uuid4())
    bind_request_context(request_id)
    response = await call_next(request)
    response.headers["X-Request-ID"] = request_id
    structlog.contextvars.clear_contextvars()
    return response

Prometheus Metrics

from prometheus_client import Counter, Histogram, Gauge, generate_latest
from fastapi import FastAPI, Response

# Define metrics
REQUEST_COUNT = Counter(
    "http_requests_total",
    "Total HTTP requests",
    ["method", "endpoint", "status"]
)

REQUEST_LATENCY = Histogram(
    "http_request_duration_seconds",
    "HTTP request latency",
    ["method", "endpoint"],
    buckets=[0.01, 0.05, 0.1, 0.5, 1.0, 5.0]
)

ACTIVE_CONNECTIONS = Gauge(
    "active_connections",
    "Number of active connections"
)

# Middleware to record metrics
@app.middleware("http")
async def metrics_middleware(request, call_next):
    ACTIVE_CONNECTIONS.inc()
    start = time.perf_counter()

    response = await call_next(request)

    duration = time.perf_counter() - start
    REQUEST_COUNT.labels(
        method=request.method,
        endpoint=request.url.path,
        status=response.status_code
    ).inc()
    REQUEST_LATENCY.labels(
        method=request.method,
        endpoint=request.url.path
    ).observe(duration)
    ACTIVE_CONNECTIONS.dec()

    return response

# Metrics endpoint
@app.get("/metrics")
async def metrics():
    return Response(
        content=generate_latest(),
        media_type="text/plain"
    )

OpenTelemetry Tracing

from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter

# Setup
provider = TracerProvider()
processor = BatchSpanProcessor(OTLPSpanExporter(endpoint="localhost:4317"))
provider.add_span_processor(processor)
trace.set_tracer_provider(provider)

tracer = trace.get_tracer(__name__)

# Manual instrumentation
async def process_order(order_id: int):
    with tracer.start_as_current_span("process_order") as span:
        span.set_attribute("order_id", order_id)

        with tracer.start_as_current_span("validate_order"):
            await validate(order_id)

        with tracer.start_as_current_span("charge_payment"):
            await charge(order_id)

Quick Reference

Library Purpose
structlog Structured logging
prometheus-client Metrics collection
opentelemetry Distributed tracing
Metric Type Use Case
Counter Total requests, errors
Histogram Latencies, sizes
Gauge Current connections, queue size

Additional Resources

  • ./references/structured-logging.md - structlog configuration, formatters
  • ./references/metrics.md - Prometheus patterns, custom metrics
  • ./references/tracing.md - OpenTelemetry, distributed tracing

Assets

  • ./assets/logging-config.py - Production logging configuration

See Also

Prerequisites:

  • python-async-ops - Async context propagation

Related Skills:

  • python-fastapi-ops - API middleware for metrics/tracing
  • python-cli-ops - CLI logging patterns

Integration Skills:

  • python-database-ops - Database query tracing
Files (claude-mods)
  • assets
    • logging-config.py 3.2 KB
      """
      Production logging configuration for Python applications.
      
      Usage:
          from logging_config import configure_logging
          configure_logging()
      """
      
      import logging
      import sys
      from typing import Literal
      
      import structlog
      
      
      def configure_logging(
          log_level: str = "INFO",
          format: Literal["json", "console"] = "json",
          service_name: str = "app",
      ):
          """
          Configure structured logging for production.
      
          Args:
              log_level: Logging level (DEBUG, INFO, WARNING, ERROR)
              format: Output format - 'json' for production, 'console' for development
              service_name: Service name to include in logs
          """
      
          # Timestamper
          timestamper = structlog.processors.TimeStamper(fmt="iso")
      
          # Shared processors for structlog and stdlib
          shared_processors = [
              structlog.contextvars.merge_contextvars,
              structlog.stdlib.add_log_level,
              structlog.stdlib.add_logger_name,
              structlog.stdlib.PositionalArgumentsFormatter(),
              timestamper,
              structlog.processors.StackInfoRenderer(),
              structlog.processors.UnicodeDecoder(),
          ]
      
          # Add service name
          def add_service_name(_, __, event_dict):
              event_dict["service"] = service_name
              return event_dict
      
          shared_processors.insert(0, add_service_name)
      
          # Choose renderer based on format
          if format == "json":
              renderer = structlog.processors.JSONRenderer()
          else:
              renderer = structlog.dev.ConsoleRenderer(
                  colors=True,
                  exception_formatter=structlog.dev.plain_traceback,
              )
      
          # Configure structlog
          structlog.configure(
              processors=shared_processors + [
                  structlog.stdlib.ProcessorFormatter.wrap_for_formatter,
              ],
              logger_factory=structlog.stdlib.LoggerFactory(),
              wrapper_class=structlog.stdlib.BoundLogger,
              cache_logger_on_first_use=True,
          )
      
          # Configure stdlib logging
          formatter = structlog.stdlib.ProcessorFormatter(
              foreign_pre_chain=shared_processors,
              processors=[
                  structlog.stdlib.ProcessorFormatter.remove_processors_meta,
                  renderer,
              ],
          )
      
          handler = logging.StreamHandler(sys.stdout)
          handler.setFormatter(formatter)
      
          # Configure root logger
          root_logger = logging.getLogger()
          root_logger.handlers = []
          root_logger.addHandler(handler)
          root_logger.setLevel(log_level)
      
          # Quiet noisy libraries
          logging.getLogger("uvicorn.access").setLevel(logging.WARNING)
          logging.getLogger("httpx").setLevel(logging.WARNING)
          logging.getLogger("httpcore").setLevel(logging.WARNING)
          logging.getLogger("sqlalchemy.engine").setLevel(logging.WARNING)
      
      
      def get_logger(name: str = None):
          """Get a structlog logger."""
          return structlog.get_logger(name)
      
      
      # Example usage
      if __name__ == "__main__":
          # Development
          configure_logging(log_level="DEBUG", format="console", service_name="demo")
      
          logger = get_logger("example")
      
          logger.info("application_started", version="1.0.0")
          logger.debug("debug_message", data={"key": "value"})
          logger.warning("rate_limit_approaching", current=95, limit=100)
      
          try:
              raise ValueError("Something went wrong")
          except Exception:
              logger.exception("operation_failed")
      
  • references
    • metrics.md 7.5 KB
      # Prometheus Metrics Patterns
      
      Application metrics for monitoring and alerting.
      
      ## Metric Types
      
      ```python
      from prometheus_client import Counter, Histogram, Gauge, Summary, Info
      
      # Counter - only goes up (resets on restart)
      REQUEST_COUNT = Counter(
          "http_requests_total",
          "Total number of HTTP requests",
          ["method", "endpoint", "status"]
      )
      
      # Histogram - distribution of values (latency, sizes)
      REQUEST_LATENCY = Histogram(
          "http_request_duration_seconds",
          "HTTP request latency in seconds",
          ["method", "endpoint"],
          buckets=[0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5, 5.0, 10.0]
      )
      
      # Gauge - can go up and down (current state)
      ACTIVE_CONNECTIONS = Gauge(
          "active_connections",
          "Number of active connections"
      )
      
      IN_PROGRESS_REQUESTS = Gauge(
          "in_progress_requests",
          "Number of requests currently being processed",
          ["endpoint"]
      )
      
      # Summary - like histogram but calculates quantiles client-side
      RESPONSE_SIZE = Summary(
          "response_size_bytes",
          "Response size in bytes",
          ["endpoint"]
      )
      
      # Info - static labels (version, build info)
      APP_INFO = Info(
          "app",
          "Application information"
      )
      APP_INFO.info({"version": "1.0.0", "environment": "production"})
      ```
      
      ## FastAPI Integration
      
      ```python
      from fastapi import FastAPI, Request, Response
      from prometheus_client import generate_latest, CONTENT_TYPE_LATEST
      import time
      
      app = FastAPI()
      
      @app.middleware("http")
      async def metrics_middleware(request: Request, call_next):
          """Record request metrics."""
          # Track in-progress requests
          endpoint = request.url.path
          IN_PROGRESS_REQUESTS.labels(endpoint=endpoint).inc()
      
          start = time.perf_counter()
          response = await call_next(request)
          duration = time.perf_counter() - start
      
          # Record metrics
          REQUEST_COUNT.labels(
              method=request.method,
              endpoint=endpoint,
              status=response.status_code
          ).inc()
      
          REQUEST_LATENCY.labels(
              method=request.method,
              endpoint=endpoint
          ).observe(duration)
      
          IN_PROGRESS_REQUESTS.labels(endpoint=endpoint).dec()
      
          return response
      
      
      @app.get("/metrics")
      async def metrics():
          """Prometheus metrics endpoint."""
          return Response(
              content=generate_latest(),
              media_type=CONTENT_TYPE_LATEST
          )
      ```
      
      ## Business Metrics
      
      ```python
      from prometheus_client import Counter, Histogram
      
      # User actions
      USER_SIGNUPS = Counter(
          "user_signups_total",
          "Total user signups",
          ["source", "plan"]
      )
      
      USER_LOGINS = Counter(
          "user_logins_total",
          "Total user logins",
          ["method"]  # oauth, password, token
      )
      
      # Orders
      ORDERS_CREATED = Counter(
          "orders_created_total",
          "Total orders created",
          ["payment_method"]
      )
      
      ORDER_VALUE = Histogram(
          "order_value_dollars",
          "Order value distribution",
          buckets=[10, 25, 50, 100, 250, 500, 1000, 2500, 5000]
      )
      
      # Errors by type
      ERRORS = Counter(
          "errors_total",
          "Total errors by type",
          ["type", "endpoint"]
      )
      
      
      # Usage
      async def create_order(order: OrderCreate):
          try:
              result = await process_order(order)
              ORDERS_CREATED.labels(payment_method=order.payment_method).inc()
              ORDER_VALUE.observe(float(order.total))
              return result
          except PaymentError as e:
              ERRORS.labels(type="payment", endpoint="/orders").inc()
              raise
      ```
      
      ## Database Metrics
      
      ```python
      from prometheus_client import Histogram, Counter, Gauge
      from contextlib import asynccontextmanager
      
      DB_QUERY_DURATION = Histogram(
          "db_query_duration_seconds",
          "Database query duration",
          ["operation", "table"]
      )
      
      DB_CONNECTIONS_ACTIVE = Gauge(
          "db_connections_active",
          "Active database connections"
      )
      
      DB_CONNECTIONS_POOL = Gauge(
          "db_connections_pool",
          "Database connection pool size"
      )
      
      DB_ERRORS = Counter(
          "db_errors_total",
          "Database errors",
          ["operation", "error_type"]
      )
      
      
      @asynccontextmanager
      async def timed_query(operation: str, table: str):
          """Context manager to time database queries."""
          start = time.perf_counter()
          try:
              yield
          except Exception as e:
              DB_ERRORS.labels(
                  operation=operation,
                  error_type=type(e).__name__
              ).inc()
              raise
          finally:
              duration = time.perf_counter() - start
              DB_QUERY_DURATION.labels(
                  operation=operation,
                  table=table
              ).observe(duration)
      
      
      # Usage
      async def get_user(user_id: int):
          async with timed_query("select", "users"):
              return await db.execute(select(User).where(User.id == user_id))
      ```
      
      ## Cache Metrics
      
      ```python
      CACHE_HITS = Counter(
          "cache_hits_total",
          "Cache hits",
          ["cache_name"]
      )
      
      CACHE_MISSES = Counter(
          "cache_misses_total",
          "Cache misses",
          ["cache_name"]
      )
      
      CACHE_LATENCY = Histogram(
          "cache_operation_duration_seconds",
          "Cache operation latency",
          ["cache_name", "operation"]
      )
      
      
      async def cached_get(key: str, fetch_func):
          """Get from cache with metrics."""
          start = time.perf_counter()
          value = await cache.get(key)
      
          if value is not None:
              CACHE_HITS.labels(cache_name="redis").inc()
              CACHE_LATENCY.labels(cache_name="redis", operation="get").observe(
                  time.perf_counter() - start
              )
              return value
      
          CACHE_MISSES.labels(cache_name="redis").inc()
      
          # Fetch and cache
          value = await fetch_func()
          await cache.set(key, value, ttl=300)
      
          return value
      ```
      
      ## Custom Collectors
      
      ```python
      from prometheus_client import Gauge
      from prometheus_client.core import GaugeMetricFamily, REGISTRY
      
      class QueueMetricsCollector:
          """Collect queue metrics on demand."""
      
          def collect(self):
              # This runs when /metrics is scraped
              queue_sizes = get_queue_sizes()  # Your function
      
              gauge = GaugeMetricFamily(
                  "queue_size",
                  "Current queue size",
                  labels=["queue_name"]
              )
      
              for name, size in queue_sizes.items():
                  gauge.add_metric([name], size)
      
              yield gauge
      
      
      # Register collector
      REGISTRY.register(QueueMetricsCollector())
      ```
      
      ## Decorators for Metrics
      
      ```python
      from functools import wraps
      import time
      
      def count_calls(counter: Counter, labels: dict | None = None):
          """Decorator to count function calls."""
          def decorator(func):
              @wraps(func)
              async def wrapper(*args, **kwargs):
                  counter.labels(**(labels or {})).inc()
                  return await func(*args, **kwargs)
              return wrapper
          return decorator
      
      
      def time_calls(histogram: Histogram, labels: dict | None = None):
          """Decorator to time function calls."""
          def decorator(func):
              @wraps(func)
              async def wrapper(*args, **kwargs):
                  start = time.perf_counter()
                  try:
                      return await func(*args, **kwargs)
                  finally:
                      duration = time.perf_counter() - start
                      histogram.labels(**(labels or {})).observe(duration)
              return wrapper
          return decorator
      
      
      # Usage
      @count_calls(USER_SIGNUPS, {"source": "api", "plan": "free"})
      @time_calls(REQUEST_LATENCY, {"method": "POST", "endpoint": "/users"})
      async def create_user(user: UserCreate):
          return await db.create_user(user)
      ```
      
      ## Quick Reference
      
      | Metric Type | Use Case | Example |
      |-------------|----------|---------|
      | Counter | Totals | Requests, errors, signups |
      | Histogram | Distributions | Latency, request size |
      | Gauge | Current state | Active connections, queue size |
      | Summary | Quantiles | Response times (p50, p99) |
      
      | Label Cardinality | Rule |
      |-------------------|------|
      | Good | method, endpoint, status |
      | Bad | user_id, request_id |
      | Limit | < 10 unique values per label |
      
    • structured-logging.md 7.8 KB
      # Structured Logging with structlog
      
      Production logging patterns for Python applications.
      
      ## Basic Setup
      
      ```python
      import logging
      import structlog
      import sys
      
      def configure_logging(json_output: bool = True, log_level: str = "INFO"):
          """Configure structlog for production."""
      
          # Shared processors for both stdlib and structlog
          shared_processors = [
              structlog.contextvars.merge_contextvars,
              structlog.stdlib.add_log_level,
              structlog.stdlib.add_logger_name,
              structlog.stdlib.PositionalArgumentsFormatter(),
              structlog.processors.TimeStamper(fmt="iso"),
              structlog.processors.StackInfoRenderer(),
              structlog.processors.UnicodeDecoder(),
          ]
      
          if json_output:
              # Production: JSON output
              renderer = structlog.processors.JSONRenderer()
          else:
              # Development: colored console output
              renderer = structlog.dev.ConsoleRenderer(colors=True)
      
          structlog.configure(
              processors=shared_processors + [
                  structlog.stdlib.ProcessorFormatter.wrap_for_formatter,
              ],
              logger_factory=structlog.stdlib.LoggerFactory(),
              cache_logger_on_first_use=True,
          )
      
          # Configure standard library logging
          handler = logging.StreamHandler(sys.stdout)
          handler.setFormatter(structlog.stdlib.ProcessorFormatter(
              foreign_pre_chain=shared_processors,
              processors=[
                  structlog.stdlib.ProcessorFormatter.remove_processors_meta,
                  renderer,
              ],
          ))
      
          root_logger = logging.getLogger()
          root_logger.addHandler(handler)
          root_logger.setLevel(log_level)
      
      
      # Usage
      configure_logging(json_output=True, log_level="INFO")
      logger = structlog.get_logger()
      ```
      
      ## Context Variables
      
      ```python
      import structlog
      from contextvars import ContextVar
      from uuid import uuid4
      
      # Request context
      request_id_var: ContextVar[str] = ContextVar("request_id", default="")
      user_id_var: ContextVar[int | None] = ContextVar("user_id", default=None)
      
      def bind_request_context(request_id: str | None = None, user_id: int | None = None):
          """Bind context that will be included in all log messages."""
          rid = request_id or str(uuid4())
          request_id_var.set(rid)
      
          context = {"request_id": rid}
          if user_id:
              user_id_var.set(user_id)
              context["user_id"] = user_id
      
          structlog.contextvars.bind_contextvars(**context)
          return rid
      
      def clear_request_context():
          """Clear context at end of request."""
          structlog.contextvars.clear_contextvars()
      
      
      # FastAPI middleware
      from fastapi import Request
      
      @app.middleware("http")
      async def logging_middleware(request: Request, call_next):
          # Extract or generate request ID
          request_id = request.headers.get("X-Request-ID", str(uuid4()))
          bind_request_context(request_id=request_id)
      
          # Log request
          logger.info(
              "request_started",
              method=request.method,
              path=request.url.path,
              client=request.client.host if request.client else None,
          )
      
          try:
              response = await call_next(request)
              logger.info(
                  "request_completed",
                  status_code=response.status_code,
              )
              response.headers["X-Request-ID"] = request_id
              return response
          except Exception as e:
              logger.exception("request_failed", error=str(e))
              raise
          finally:
              clear_request_context()
      ```
      
      ## Exception Logging
      
      ```python
      import structlog
      
      logger = structlog.get_logger()
      
      # Log exception with context
      try:
          result = risky_operation()
      except ValueError as e:
          logger.error(
              "operation_failed",
              error=str(e),
              error_type=type(e).__name__,
          )
          raise
      
      # Log with full traceback
      try:
          result = another_operation()
      except Exception:
          logger.exception("unexpected_error")  # Includes full traceback
          raise
      
      
      # Custom exception with context
      class OrderError(Exception):
          def __init__(self, message: str, order_id: int, **context):
              super().__init__(message)
              self.order_id = order_id
              self.context = context
      
      try:
          process_order(order_id=123)
      except OrderError as e:
          logger.error(
              "order_processing_failed",
              order_id=e.order_id,
              **e.context,
          )
      ```
      
      ## Filtering Sensitive Data
      
      ```python
      import structlog
      import re
      
      def filter_sensitive_data(_, __, event_dict):
          """Remove sensitive data from logs."""
          sensitive_keys = {"password", "token", "secret", "api_key", "authorization"}
      
          def redact(data):
              if isinstance(data, dict):
                  return {
                      k: "[REDACTED]" if k.lower() in sensitive_keys else redact(v)
                      for k, v in data.items()
                  }
              elif isinstance(data, list):
                  return [redact(item) for item in data]
              elif isinstance(data, str):
                  # Redact emails
                  return re.sub(
                      r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}',
                      '[EMAIL]',
                      data
                  )
              return data
      
          return redact(event_dict)
      
      
      structlog.configure(
          processors=[
              filter_sensitive_data,
              structlog.processors.JSONRenderer(),
          ],
      )
      ```
      
      ## Log Levels and Events
      
      ```python
      logger = structlog.get_logger()
      
      # Use semantic event names
      logger.debug("cache_lookup", key="user:123", hit=True)
      logger.info("user_created", user_id=123, email="user@example.com")
      logger.warning("rate_limit_approaching", current=95, limit=100)
      logger.error("payment_failed", order_id=456, reason="insufficient_funds")
      logger.critical("database_connection_lost", host="db.example.com")
      
      # Business events
      logger.info("order_placed", order_id=789, total=99.99, items=3)
      logger.info("order_shipped", order_id=789, carrier="ups", tracking="1Z...")
      logger.info("user_login", user_id=123, method="oauth", provider="google")
      ```
      
      ## Integration with Third-Party Loggers
      
      ```python
      import structlog
      import logging
      
      # Capture logs from libraries
      logging.getLogger("sqlalchemy.engine").setLevel(logging.WARNING)
      logging.getLogger("httpx").setLevel(logging.WARNING)
      
      # Create a structlog-wrapped stdlib logger for compatibility
      def get_stdlib_logger(name: str):
          """Get a structlog logger that works with libraries expecting stdlib."""
          return structlog.wrap_logger(
              logging.getLogger(name),
              processors=[
                  structlog.stdlib.filter_by_level,
                  structlog.stdlib.add_logger_name,
                  structlog.stdlib.add_log_level,
                  structlog.processors.TimeStamper(fmt="iso"),
                  structlog.processors.JSONRenderer(),
              ]
          )
      ```
      
      ## Performance Logging
      
      ```python
      import structlog
      import time
      from contextlib import contextmanager
      
      logger = structlog.get_logger()
      
      @contextmanager
      def log_duration(event: str, **context):
          """Context manager to log operation duration."""
          start = time.perf_counter()
          try:
              yield
              duration = time.perf_counter() - start
              logger.info(
                  event,
                  duration_ms=round(duration * 1000, 2),
                  status="success",
                  **context,
              )
          except Exception as e:
              duration = time.perf_counter() - start
              logger.error(
                  event,
                  duration_ms=round(duration * 1000, 2),
                  status="error",
                  error=str(e),
                  **context,
              )
              raise
      
      
      # Usage
      with log_duration("database_query", table="users"):
          users = await db.fetch_users()
      ```
      
      ## Quick Reference
      
      | Function | Purpose |
      |----------|---------|
      | `structlog.get_logger()` | Get logger instance |
      | `bind_contextvars()` | Add context to all logs |
      | `clear_contextvars()` | Clear request context |
      | `logger.exception()` | Log with traceback |
      
      | Processor | Purpose |
      |-----------|---------|
      | `TimeStamper(fmt="iso")` | Add timestamp |
      | `add_log_level` | Add level field |
      | `JSONRenderer()` | Output as JSON |
      | `ConsoleRenderer()` | Pretty console output |
      
    • tracing.md 8 KB
      # Distributed Tracing with OpenTelemetry
      
      Trace requests across services for debugging and performance analysis.
      
      ## Setup
      
      ```python
      from opentelemetry import trace
      from opentelemetry.sdk.trace import TracerProvider
      from opentelemetry.sdk.trace.export import BatchSpanProcessor
      from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
      from opentelemetry.sdk.resources import Resource
      
      # Create resource with service info
      resource = Resource.create({
          "service.name": "my-service",
          "service.version": "1.0.0",
          "deployment.environment": "production",
      })
      
      # Create and configure tracer provider
      provider = TracerProvider(resource=resource)
      
      # Export to OTLP collector (Jaeger, Tempo, etc.)
      otlp_exporter = OTLPSpanExporter(
          endpoint="http://localhost:4317",
          insecure=True,
      )
      provider.add_span_processor(BatchSpanProcessor(otlp_exporter))
      
      # Set as global tracer provider
      trace.set_tracer_provider(provider)
      
      # Get tracer for your module
      tracer = trace.get_tracer(__name__)
      ```
      
      ## FastAPI Auto-Instrumentation
      
      ```python
      from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
      from opentelemetry.instrumentation.httpx import HTTPXClientInstrumentor
      from opentelemetry.instrumentation.sqlalchemy import SQLAlchemyInstrumentor
      from opentelemetry.instrumentation.redis import RedisInstrumentor
      
      # Instrument FastAPI
      FastAPIInstrumentor.instrument_app(app)
      
      # Instrument HTTP client
      HTTPXClientInstrumentor().instrument()
      
      # Instrument database
      SQLAlchemyInstrumentor().instrument(engine=engine)
      
      # Instrument Redis
      RedisInstrumentor().instrument()
      ```
      
      ## Manual Instrumentation
      
      ```python
      from opentelemetry import trace
      from opentelemetry.trace import Status, StatusCode
      
      tracer = trace.get_tracer(__name__)
      
      async def process_order(order_id: int):
          """Process order with detailed tracing."""
          with tracer.start_as_current_span("process_order") as span:
              # Add attributes
              span.set_attribute("order.id", order_id)
              span.set_attribute("order.type", "standard")
      
              # Nested spans
              with tracer.start_as_current_span("validate_order"):
                  order = await validate(order_id)
                  span.set_attribute("order.items", len(order.items))
      
              with tracer.start_as_current_span("check_inventory"):
                  await check_inventory(order.items)
      
              with tracer.start_as_current_span("process_payment") as payment_span:
                  try:
                      result = await charge_payment(order)
                      payment_span.set_attribute("payment.amount", float(order.total))
                  except PaymentError as e:
                      payment_span.set_status(Status(StatusCode.ERROR, str(e)))
                      payment_span.record_exception(e)
                      raise
      
              with tracer.start_as_current_span("send_confirmation"):
                  await send_email(order.customer_email)
      
              span.set_status(Status(StatusCode.OK))
              return order
      ```
      
      ## Context Propagation
      
      ```python
      from opentelemetry import trace
      from opentelemetry.propagate import inject, extract
      from opentelemetry.trace.propagation.tracecontext import TraceContextTextMapPropagator
      
      propagator = TraceContextTextMapPropagator()
      
      # Inject context into outgoing HTTP headers
      async def call_external_service(data: dict):
          headers = {}
          inject(headers)  # Adds traceparent header
      
          async with httpx.AsyncClient() as client:
              response = await client.post(
                  "https://api.example.com/process",
                  json=data,
                  headers=headers,
              )
          return response.json()
      
      
      # Extract context from incoming request (usually handled by instrumentation)
      @app.middleware("http")
      async def trace_middleware(request: Request, call_next):
          # Extract trace context from headers
          ctx = extract(dict(request.headers))
      
          with tracer.start_as_current_span(
              f"{request.method} {request.url.path}",
              context=ctx,
          ):
              return await call_next(request)
      ```
      
      ## Adding Events and Exceptions
      
      ```python
      from opentelemetry import trace
      
      tracer = trace.get_tracer(__name__)
      
      async def process_with_events():
          with tracer.start_as_current_span("process") as span:
              # Add event (point-in-time occurrence)
              span.add_event("processing_started", {
                  "items": 10,
              })
      
              try:
                  result = await heavy_processing()
                  span.add_event("processing_completed", {
                      "result_count": len(result),
                  })
              except Exception as e:
                  # Record exception in span
                  span.record_exception(e)
                  span.set_status(Status(StatusCode.ERROR, str(e)))
                  raise
      
              return result
      ```
      
      ## Span Decorator
      
      ```python
      from functools import wraps
      from opentelemetry import trace
      
      tracer = trace.get_tracer(__name__)
      
      def traced(span_name: str | None = None, attributes: dict | None = None):
          """Decorator to trace function execution."""
          def decorator(func):
              @wraps(func)
              async def async_wrapper(*args, **kwargs):
                  name = span_name or f"{func.__module__}.{func.__name__}"
                  with tracer.start_as_current_span(name) as span:
                      if attributes:
                          for key, value in attributes.items():
                              span.set_attribute(key, value)
                      try:
                          result = await func(*args, **kwargs)
                          span.set_status(Status(StatusCode.OK))
                          return result
                      except Exception as e:
                          span.record_exception(e)
                          span.set_status(Status(StatusCode.ERROR, str(e)))
                          raise
      
              @wraps(func)
              def sync_wrapper(*args, **kwargs):
                  name = span_name or f"{func.__module__}.{func.__name__}"
                  with tracer.start_as_current_span(name) as span:
                      if attributes:
                          for key, value in attributes.items():
                              span.set_attribute(key, value)
                      try:
                          result = func(*args, **kwargs)
                          span.set_status(Status(StatusCode.OK))
                          return result
                      except Exception as e:
                          span.record_exception(e)
                          span.set_status(Status(StatusCode.ERROR, str(e)))
                          raise
      
              if asyncio.iscoroutinefunction(func):
                  return async_wrapper
              return sync_wrapper
          return decorator
      
      
      # Usage
      @traced("user.create", {"component": "users"})
      async def create_user(user: UserCreate):
          return await db.create(user)
      ```
      
      ## Linking Traces to Logs
      
      ```python
      import structlog
      from opentelemetry import trace
      
      def add_trace_context(_, __, event_dict):
          """Add trace context to log entries."""
          span = trace.get_current_span()
          if span.is_recording():
              ctx = span.get_span_context()
              event_dict["trace_id"] = format(ctx.trace_id, "032x")
              event_dict["span_id"] = format(ctx.span_id, "016x")
          return event_dict
      
      
      structlog.configure(
          processors=[
              add_trace_context,
              structlog.processors.JSONRenderer(),
          ],
      )
      ```
      
      ## Sampling
      
      ```python
      from opentelemetry.sdk.trace.sampling import (
          TraceIdRatioBased,
          ParentBased,
          ALWAYS_ON,
      )
      
      # Sample 10% of traces
      sampler = TraceIdRatioBased(0.1)
      
      # Respect parent's sampling decision, default to 10%
      sampler = ParentBased(root=TraceIdRatioBased(0.1))
      
      # Always sample (development)
      sampler = ALWAYS_ON
      
      provider = TracerProvider(
          resource=resource,
          sampler=sampler,
      )
      ```
      
      ## Quick Reference
      
      | Concept | Description |
      |---------|-------------|
      | Trace | Complete request journey |
      | Span | Single operation within trace |
      | Context | Propagated trace information |
      | Attribute | Key-value metadata on span |
      | Event | Point-in-time occurrence |
      
      | Instrumentation | Package |
      |-----------------|---------|
      | FastAPI | `opentelemetry-instrumentation-fastapi` |
      | httpx | `opentelemetry-instrumentation-httpx` |
      | SQLAlchemy | `opentelemetry-instrumentation-sqlalchemy` |
      | Redis | `opentelemetry-instrumentation-redis` |
      | Celery | `opentelemetry-instrumentation-celery` |
      
  • scripts
    • .gitkeep 0 B · in bundle
  • SKILL.md 5 KB
    ---
    name: python-observability-ops
    description: "Observability patterns for Python applications. Triggers on: logging, metrics, tracing, opentelemetry, prometheus, observability, monitoring, structlog, correlation id."
    license: MIT
    compatibility: "Python 3.10+. Requires structlog, opentelemetry-api, prometheus-client."
    allowed-tools: "Read Write"
    metadata:
      author: claude-mods
      depends-on: python-async-ops
      related-skills: python-fastapi-ops, python-cli-ops
    ---
    
    # Python Observability Patterns
    
    Logging, metrics, and tracing for production applications.
    
    ## Structured Logging with structlog
    
    ```python
    import structlog
    
    # Configure structlog
    structlog.configure(
        processors=[
            structlog.contextvars.merge_contextvars,
            structlog.processors.add_log_level,
            structlog.processors.TimeStamper(fmt="iso"),
            structlog.processors.JSONRenderer(),
        ],
        wrapper_class=structlog.make_filtering_bound_logger(logging.INFO),
        context_class=dict,
        logger_factory=structlog.PrintLoggerFactory(),
    )
    
    logger = structlog.get_logger()
    
    # Usage
    logger.info("user_created", user_id=123, email="test@example.com")
    # Output: {"event": "user_created", "user_id": 123, "email": "test@example.com", "level": "info", "timestamp": "2024-01-15T10:00:00Z"}
    ```
    
    ## Request Context Propagation
    
    ```python
    import structlog
    from contextvars import ContextVar
    from uuid import uuid4
    
    request_id_var: ContextVar[str] = ContextVar("request_id", default="")
    
    def bind_request_context(request_id: str | None = None):
        """Bind request ID to logging context."""
        rid = request_id or str(uuid4())
        request_id_var.set(rid)
        structlog.contextvars.bind_contextvars(request_id=rid)
        return rid
    
    # FastAPI middleware
    @app.middleware("http")
    async def request_context_middleware(request, call_next):
        request_id = request.headers.get("X-Request-ID") or str(uuid4())
        bind_request_context(request_id)
        response = await call_next(request)
        response.headers["X-Request-ID"] = request_id
        structlog.contextvars.clear_contextvars()
        return response
    ```
    
    ## Prometheus Metrics
    
    ```python
    from prometheus_client import Counter, Histogram, Gauge, generate_latest
    from fastapi import FastAPI, Response
    
    # Define metrics
    REQUEST_COUNT = Counter(
        "http_requests_total",
        "Total HTTP requests",
        ["method", "endpoint", "status"]
    )
    
    REQUEST_LATENCY = Histogram(
        "http_request_duration_seconds",
        "HTTP request latency",
        ["method", "endpoint"],
        buckets=[0.01, 0.05, 0.1, 0.5, 1.0, 5.0]
    )
    
    ACTIVE_CONNECTIONS = Gauge(
        "active_connections",
        "Number of active connections"
    )
    
    # Middleware to record metrics
    @app.middleware("http")
    async def metrics_middleware(request, call_next):
        ACTIVE_CONNECTIONS.inc()
        start = time.perf_counter()
    
        response = await call_next(request)
    
        duration = time.perf_counter() - start
        REQUEST_COUNT.labels(
            method=request.method,
            endpoint=request.url.path,
            status=response.status_code
        ).inc()
        REQUEST_LATENCY.labels(
            method=request.method,
            endpoint=request.url.path
        ).observe(duration)
        ACTIVE_CONNECTIONS.dec()
    
        return response
    
    # Metrics endpoint
    @app.get("/metrics")
    async def metrics():
        return Response(
            content=generate_latest(),
            media_type="text/plain"
        )
    ```
    
    ## OpenTelemetry Tracing
    
    ```python
    from opentelemetry import trace
    from opentelemetry.sdk.trace import TracerProvider
    from opentelemetry.sdk.trace.export import BatchSpanProcessor
    from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
    
    # Setup
    provider = TracerProvider()
    processor = BatchSpanProcessor(OTLPSpanExporter(endpoint="localhost:4317"))
    provider.add_span_processor(processor)
    trace.set_tracer_provider(provider)
    
    tracer = trace.get_tracer(__name__)
    
    # Manual instrumentation
    async def process_order(order_id: int):
        with tracer.start_as_current_span("process_order") as span:
            span.set_attribute("order_id", order_id)
    
            with tracer.start_as_current_span("validate_order"):
                await validate(order_id)
    
            with tracer.start_as_current_span("charge_payment"):
                await charge(order_id)
    ```
    
    ## Quick Reference
    
    | Library | Purpose |
    |---------|---------|
    | structlog | Structured logging |
    | prometheus-client | Metrics collection |
    | opentelemetry | Distributed tracing |
    
    | Metric Type | Use Case |
    |-------------|----------|
    | Counter | Total requests, errors |
    | Histogram | Latencies, sizes |
    | Gauge | Current connections, queue size |
    
    ## Additional Resources
    
    - `./references/structured-logging.md` - structlog configuration, formatters
    - `./references/metrics.md` - Prometheus patterns, custom metrics
    - `./references/tracing.md` - OpenTelemetry, distributed tracing
    
    ## Assets
    
    - `./assets/logging-config.py` - Production logging configuration
    
    ---
    
    ## See Also
    
    **Prerequisites:**
    - `python-async-ops` - Async context propagation
    
    **Related Skills:**
    - `python-fastapi-ops` - API middleware for metrics/tracing
    - `python-cli-ops` - CLI logging patterns
    
    **Integration Skills:**
    - `python-database-ops` - Database query tracing
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related