Claude
Skill
python-observability-ops
Observability patterns for Python applications. Triggers on: logging, metrics, tracing, opentelemetry, prometheus, observability, monitoring, structlog, correlation id.
Virus-scanned
Reviewed automatically before listing.
Download
0xdarkmatter-claude-mods-skills_python-observability-ops-3dfaf0b.zip · 11 KB
Install
skills CLI
npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/python-observability-ops
Claude Code
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
Git
git clone https://github.com/0xDarkMatter/claude-mods.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Python Observability Patterns
Logging, metrics, and tracing for production applications.
Structured Logging with structlog
import structlog
# Configure structlog
structlog.configure(
processors=[
structlog.contextvars.merge_contextvars,
structlog.processors.add_log_level,
structlog.processors.TimeStamper(fmt="iso"),
structlog.processors.JSONRenderer(),
],
wrapper_class=structlog.make_filtering_bound_logger(logging.INFO),
context_class=dict,
logger_factory=structlog.PrintLoggerFactory(),
)
logger = structlog.get_logger()
# Usage
logger.info("user_created", user_id=123, email="test@example.com")
# Output: {"event": "user_created", "user_id": 123, "email": "test@example.com", "level": "info", "timestamp": "2024-01-15T10:00:00Z"}
Request Context Propagation
import structlog
from contextvars import ContextVar
from uuid import uuid4
request_id_var: ContextVar[str] = ContextVar("request_id", default="")
def bind_request_context(request_id: str | None = None):
"""Bind request ID to logging context."""
rid = request_id or str(uuid4())
request_id_var.set(rid)
structlog.contextvars.bind_contextvars(request_id=rid)
return rid
# FastAPI middleware
@app.middleware("http")
async def request_context_middleware(request, call_next):
request_id = request.headers.get("X-Request-ID") or str(uuid4())
bind_request_context(request_id)
response = await call_next(request)
response.headers["X-Request-ID"] = request_id
structlog.contextvars.clear_contextvars()
return response
Prometheus Metrics
from prometheus_client import Counter, Histogram, Gauge, generate_latest
from fastapi import FastAPI, Response
# Define metrics
REQUEST_COUNT = Counter(
"http_requests_total",
"Total HTTP requests",
["method", "endpoint", "status"]
)
REQUEST_LATENCY = Histogram(
"http_request_duration_seconds",
"HTTP request latency",
["method", "endpoint"],
buckets=[0.01, 0.05, 0.1, 0.5, 1.0, 5.0]
)
ACTIVE_CONNECTIONS = Gauge(
"active_connections",
"Number of active connections"
)
# Middleware to record metrics
@app.middleware("http")
async def metrics_middleware(request, call_next):
ACTIVE_CONNECTIONS.inc()
start = time.perf_counter()
response = await call_next(request)
duration = time.perf_counter() - start
REQUEST_COUNT.labels(
method=request.method,
endpoint=request.url.path,
status=response.status_code
).inc()
REQUEST_LATENCY.labels(
method=request.method,
endpoint=request.url.path
).observe(duration)
ACTIVE_CONNECTIONS.dec()
return response
# Metrics endpoint
@app.get("/metrics")
async def metrics():
return Response(
content=generate_latest(),
media_type="text/plain"
)
OpenTelemetry Tracing
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
# Setup
provider = TracerProvider()
processor = BatchSpanProcessor(OTLPSpanExporter(endpoint="localhost:4317"))
provider.add_span_processor(processor)
trace.set_tracer_provider(provider)
tracer = trace.get_tracer(__name__)
# Manual instrumentation
async def process_order(order_id: int):
with tracer.start_as_current_span("process_order") as span:
span.set_attribute("order_id", order_id)
with tracer.start_as_current_span("validate_order"):
await validate(order_id)
with tracer.start_as_current_span("charge_payment"):
await charge(order_id)
Quick Reference
| Library | Purpose |
|---|---|
| structlog | Structured logging |
| prometheus-client | Metrics collection |
| opentelemetry | Distributed tracing |
| Metric Type | Use Case |
|---|---|
| Counter | Total requests, errors |
| Histogram | Latencies, sizes |
| Gauge | Current connections, queue size |
Additional Resources
./references/structured-logging.md- structlog configuration, formatters./references/metrics.md- Prometheus patterns, custom metrics./references/tracing.md- OpenTelemetry, distributed tracing
Assets
./assets/logging-config.py- Production logging configuration
See Also
Prerequisites:
python-async-ops- Async context propagation
Related Skills:
python-fastapi-ops- API middleware for metrics/tracingpython-cli-ops- CLI logging patterns
Integration Skills:
python-database-ops- Database query tracing
Files (claude-mods)
-
assets
-
logging-config.py 3.2 KB
""" Production logging configuration for Python applications. Usage: from logging_config import configure_logging configure_logging() """ import logging import sys from typing import Literal import structlog def configure_logging( log_level: str = "INFO", format: Literal["json", "console"] = "json", service_name: str = "app", ): """ Configure structured logging for production. Args: log_level: Logging level (DEBUG, INFO, WARNING, ERROR) format: Output format - 'json' for production, 'console' for development service_name: Service name to include in logs """ # Timestamper timestamper = structlog.processors.TimeStamper(fmt="iso") # Shared processors for structlog and stdlib shared_processors = [ structlog.contextvars.merge_contextvars, structlog.stdlib.add_log_level, structlog.stdlib.add_logger_name, structlog.stdlib.PositionalArgumentsFormatter(), timestamper, structlog.processors.StackInfoRenderer(), structlog.processors.UnicodeDecoder(), ] # Add service name def add_service_name(_, __, event_dict): event_dict["service"] = service_name return event_dict shared_processors.insert(0, add_service_name) # Choose renderer based on format if format == "json": renderer = structlog.processors.JSONRenderer() else: renderer = structlog.dev.ConsoleRenderer( colors=True, exception_formatter=structlog.dev.plain_traceback, ) # Configure structlog structlog.configure( processors=shared_processors + [ structlog.stdlib.ProcessorFormatter.wrap_for_formatter, ], logger_factory=structlog.stdlib.LoggerFactory(), wrapper_class=structlog.stdlib.BoundLogger, cache_logger_on_first_use=True, ) # Configure stdlib logging formatter = structlog.stdlib.ProcessorFormatter( foreign_pre_chain=shared_processors, processors=[ structlog.stdlib.ProcessorFormatter.remove_processors_meta, renderer, ], ) handler = logging.StreamHandler(sys.stdout) handler.setFormatter(formatter) # Configure root logger root_logger = logging.getLogger() root_logger.handlers = [] root_logger.addHandler(handler) root_logger.setLevel(log_level) # Quiet noisy libraries logging.getLogger("uvicorn.access").setLevel(logging.WARNING) logging.getLogger("httpx").setLevel(logging.WARNING) logging.getLogger("httpcore").setLevel(logging.WARNING) logging.getLogger("sqlalchemy.engine").setLevel(logging.WARNING) def get_logger(name: str = None): """Get a structlog logger.""" return structlog.get_logger(name) # Example usage if __name__ == "__main__": # Development configure_logging(log_level="DEBUG", format="console", service_name="demo") logger = get_logger("example") logger.info("application_started", version="1.0.0") logger.debug("debug_message", data={"key": "value"}) logger.warning("rate_limit_approaching", current=95, limit=100) try: raise ValueError("Something went wrong") except Exception: logger.exception("operation_failed")
-
-
references
-
metrics.md 7.5 KB
# Prometheus Metrics Patterns Application metrics for monitoring and alerting. ## Metric Types ```python from prometheus_client import Counter, Histogram, Gauge, Summary, Info # Counter - only goes up (resets on restart) REQUEST_COUNT = Counter( "http_requests_total", "Total number of HTTP requests", ["method", "endpoint", "status"] ) # Histogram - distribution of values (latency, sizes) REQUEST_LATENCY = Histogram( "http_request_duration_seconds", "HTTP request latency in seconds", ["method", "endpoint"], buckets=[0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5, 5.0, 10.0] ) # Gauge - can go up and down (current state) ACTIVE_CONNECTIONS = Gauge( "active_connections", "Number of active connections" ) IN_PROGRESS_REQUESTS = Gauge( "in_progress_requests", "Number of requests currently being processed", ["endpoint"] ) # Summary - like histogram but calculates quantiles client-side RESPONSE_SIZE = Summary( "response_size_bytes", "Response size in bytes", ["endpoint"] ) # Info - static labels (version, build info) APP_INFO = Info( "app", "Application information" ) APP_INFO.info({"version": "1.0.0", "environment": "production"}) ``` ## FastAPI Integration ```python from fastapi import FastAPI, Request, Response from prometheus_client import generate_latest, CONTENT_TYPE_LATEST import time app = FastAPI() @app.middleware("http") async def metrics_middleware(request: Request, call_next): """Record request metrics.""" # Track in-progress requests endpoint = request.url.path IN_PROGRESS_REQUESTS.labels(endpoint=endpoint).inc() start = time.perf_counter() response = await call_next(request) duration = time.perf_counter() - start # Record metrics REQUEST_COUNT.labels( method=request.method, endpoint=endpoint, status=response.status_code ).inc() REQUEST_LATENCY.labels( method=request.method, endpoint=endpoint ).observe(duration) IN_PROGRESS_REQUESTS.labels(endpoint=endpoint).dec() return response @app.get("/metrics") async def metrics(): """Prometheus metrics endpoint.""" return Response( content=generate_latest(), media_type=CONTENT_TYPE_LATEST ) ``` ## Business Metrics ```python from prometheus_client import Counter, Histogram # User actions USER_SIGNUPS = Counter( "user_signups_total", "Total user signups", ["source", "plan"] ) USER_LOGINS = Counter( "user_logins_total", "Total user logins", ["method"] # oauth, password, token ) # Orders ORDERS_CREATED = Counter( "orders_created_total", "Total orders created", ["payment_method"] ) ORDER_VALUE = Histogram( "order_value_dollars", "Order value distribution", buckets=[10, 25, 50, 100, 250, 500, 1000, 2500, 5000] ) # Errors by type ERRORS = Counter( "errors_total", "Total errors by type", ["type", "endpoint"] ) # Usage async def create_order(order: OrderCreate): try: result = await process_order(order) ORDERS_CREATED.labels(payment_method=order.payment_method).inc() ORDER_VALUE.observe(float(order.total)) return result except PaymentError as e: ERRORS.labels(type="payment", endpoint="/orders").inc() raise ``` ## Database Metrics ```python from prometheus_client import Histogram, Counter, Gauge from contextlib import asynccontextmanager DB_QUERY_DURATION = Histogram( "db_query_duration_seconds", "Database query duration", ["operation", "table"] ) DB_CONNECTIONS_ACTIVE = Gauge( "db_connections_active", "Active database connections" ) DB_CONNECTIONS_POOL = Gauge( "db_connections_pool", "Database connection pool size" ) DB_ERRORS = Counter( "db_errors_total", "Database errors", ["operation", "error_type"] ) @asynccontextmanager async def timed_query(operation: str, table: str): """Context manager to time database queries.""" start = time.perf_counter() try: yield except Exception as e: DB_ERRORS.labels( operation=operation, error_type=type(e).__name__ ).inc() raise finally: duration = time.perf_counter() - start DB_QUERY_DURATION.labels( operation=operation, table=table ).observe(duration) # Usage async def get_user(user_id: int): async with timed_query("select", "users"): return await db.execute(select(User).where(User.id == user_id)) ``` ## Cache Metrics ```python CACHE_HITS = Counter( "cache_hits_total", "Cache hits", ["cache_name"] ) CACHE_MISSES = Counter( "cache_misses_total", "Cache misses", ["cache_name"] ) CACHE_LATENCY = Histogram( "cache_operation_duration_seconds", "Cache operation latency", ["cache_name", "operation"] ) async def cached_get(key: str, fetch_func): """Get from cache with metrics.""" start = time.perf_counter() value = await cache.get(key) if value is not None: CACHE_HITS.labels(cache_name="redis").inc() CACHE_LATENCY.labels(cache_name="redis", operation="get").observe( time.perf_counter() - start ) return value CACHE_MISSES.labels(cache_name="redis").inc() # Fetch and cache value = await fetch_func() await cache.set(key, value, ttl=300) return value ``` ## Custom Collectors ```python from prometheus_client import Gauge from prometheus_client.core import GaugeMetricFamily, REGISTRY class QueueMetricsCollector: """Collect queue metrics on demand.""" def collect(self): # This runs when /metrics is scraped queue_sizes = get_queue_sizes() # Your function gauge = GaugeMetricFamily( "queue_size", "Current queue size", labels=["queue_name"] ) for name, size in queue_sizes.items(): gauge.add_metric([name], size) yield gauge # Register collector REGISTRY.register(QueueMetricsCollector()) ``` ## Decorators for Metrics ```python from functools import wraps import time def count_calls(counter: Counter, labels: dict | None = None): """Decorator to count function calls.""" def decorator(func): @wraps(func) async def wrapper(*args, **kwargs): counter.labels(**(labels or {})).inc() return await func(*args, **kwargs) return wrapper return decorator def time_calls(histogram: Histogram, labels: dict | None = None): """Decorator to time function calls.""" def decorator(func): @wraps(func) async def wrapper(*args, **kwargs): start = time.perf_counter() try: return await func(*args, **kwargs) finally: duration = time.perf_counter() - start histogram.labels(**(labels or {})).observe(duration) return wrapper return decorator # Usage @count_calls(USER_SIGNUPS, {"source": "api", "plan": "free"}) @time_calls(REQUEST_LATENCY, {"method": "POST", "endpoint": "/users"}) async def create_user(user: UserCreate): return await db.create_user(user) ``` ## Quick Reference | Metric Type | Use Case | Example | |-------------|----------|---------| | Counter | Totals | Requests, errors, signups | | Histogram | Distributions | Latency, request size | | Gauge | Current state | Active connections, queue size | | Summary | Quantiles | Response times (p50, p99) | | Label Cardinality | Rule | |-------------------|------| | Good | method, endpoint, status | | Bad | user_id, request_id | | Limit | < 10 unique values per label | -
structured-logging.md 7.8 KB
# Structured Logging with structlog Production logging patterns for Python applications. ## Basic Setup ```python import logging import structlog import sys def configure_logging(json_output: bool = True, log_level: str = "INFO"): """Configure structlog for production.""" # Shared processors for both stdlib and structlog shared_processors = [ structlog.contextvars.merge_contextvars, structlog.stdlib.add_log_level, structlog.stdlib.add_logger_name, structlog.stdlib.PositionalArgumentsFormatter(), structlog.processors.TimeStamper(fmt="iso"), structlog.processors.StackInfoRenderer(), structlog.processors.UnicodeDecoder(), ] if json_output: # Production: JSON output renderer = structlog.processors.JSONRenderer() else: # Development: colored console output renderer = structlog.dev.ConsoleRenderer(colors=True) structlog.configure( processors=shared_processors + [ structlog.stdlib.ProcessorFormatter.wrap_for_formatter, ], logger_factory=structlog.stdlib.LoggerFactory(), cache_logger_on_first_use=True, ) # Configure standard library logging handler = logging.StreamHandler(sys.stdout) handler.setFormatter(structlog.stdlib.ProcessorFormatter( foreign_pre_chain=shared_processors, processors=[ structlog.stdlib.ProcessorFormatter.remove_processors_meta, renderer, ], )) root_logger = logging.getLogger() root_logger.addHandler(handler) root_logger.setLevel(log_level) # Usage configure_logging(json_output=True, log_level="INFO") logger = structlog.get_logger() ``` ## Context Variables ```python import structlog from contextvars import ContextVar from uuid import uuid4 # Request context request_id_var: ContextVar[str] = ContextVar("request_id", default="") user_id_var: ContextVar[int | None] = ContextVar("user_id", default=None) def bind_request_context(request_id: str | None = None, user_id: int | None = None): """Bind context that will be included in all log messages.""" rid = request_id or str(uuid4()) request_id_var.set(rid) context = {"request_id": rid} if user_id: user_id_var.set(user_id) context["user_id"] = user_id structlog.contextvars.bind_contextvars(**context) return rid def clear_request_context(): """Clear context at end of request.""" structlog.contextvars.clear_contextvars() # FastAPI middleware from fastapi import Request @app.middleware("http") async def logging_middleware(request: Request, call_next): # Extract or generate request ID request_id = request.headers.get("X-Request-ID", str(uuid4())) bind_request_context(request_id=request_id) # Log request logger.info( "request_started", method=request.method, path=request.url.path, client=request.client.host if request.client else None, ) try: response = await call_next(request) logger.info( "request_completed", status_code=response.status_code, ) response.headers["X-Request-ID"] = request_id return response except Exception as e: logger.exception("request_failed", error=str(e)) raise finally: clear_request_context() ``` ## Exception Logging ```python import structlog logger = structlog.get_logger() # Log exception with context try: result = risky_operation() except ValueError as e: logger.error( "operation_failed", error=str(e), error_type=type(e).__name__, ) raise # Log with full traceback try: result = another_operation() except Exception: logger.exception("unexpected_error") # Includes full traceback raise # Custom exception with context class OrderError(Exception): def __init__(self, message: str, order_id: int, **context): super().__init__(message) self.order_id = order_id self.context = context try: process_order(order_id=123) except OrderError as e: logger.error( "order_processing_failed", order_id=e.order_id, **e.context, ) ``` ## Filtering Sensitive Data ```python import structlog import re def filter_sensitive_data(_, __, event_dict): """Remove sensitive data from logs.""" sensitive_keys = {"password", "token", "secret", "api_key", "authorization"} def redact(data): if isinstance(data, dict): return { k: "[REDACTED]" if k.lower() in sensitive_keys else redact(v) for k, v in data.items() } elif isinstance(data, list): return [redact(item) for item in data] elif isinstance(data, str): # Redact emails return re.sub( r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}', '[EMAIL]', data ) return data return redact(event_dict) structlog.configure( processors=[ filter_sensitive_data, structlog.processors.JSONRenderer(), ], ) ``` ## Log Levels and Events ```python logger = structlog.get_logger() # Use semantic event names logger.debug("cache_lookup", key="user:123", hit=True) logger.info("user_created", user_id=123, email="user@example.com") logger.warning("rate_limit_approaching", current=95, limit=100) logger.error("payment_failed", order_id=456, reason="insufficient_funds") logger.critical("database_connection_lost", host="db.example.com") # Business events logger.info("order_placed", order_id=789, total=99.99, items=3) logger.info("order_shipped", order_id=789, carrier="ups", tracking="1Z...") logger.info("user_login", user_id=123, method="oauth", provider="google") ``` ## Integration with Third-Party Loggers ```python import structlog import logging # Capture logs from libraries logging.getLogger("sqlalchemy.engine").setLevel(logging.WARNING) logging.getLogger("httpx").setLevel(logging.WARNING) # Create a structlog-wrapped stdlib logger for compatibility def get_stdlib_logger(name: str): """Get a structlog logger that works with libraries expecting stdlib.""" return structlog.wrap_logger( logging.getLogger(name), processors=[ structlog.stdlib.filter_by_level, structlog.stdlib.add_logger_name, structlog.stdlib.add_log_level, structlog.processors.TimeStamper(fmt="iso"), structlog.processors.JSONRenderer(), ] ) ``` ## Performance Logging ```python import structlog import time from contextlib import contextmanager logger = structlog.get_logger() @contextmanager def log_duration(event: str, **context): """Context manager to log operation duration.""" start = time.perf_counter() try: yield duration = time.perf_counter() - start logger.info( event, duration_ms=round(duration * 1000, 2), status="success", **context, ) except Exception as e: duration = time.perf_counter() - start logger.error( event, duration_ms=round(duration * 1000, 2), status="error", error=str(e), **context, ) raise # Usage with log_duration("database_query", table="users"): users = await db.fetch_users() ``` ## Quick Reference | Function | Purpose | |----------|---------| | `structlog.get_logger()` | Get logger instance | | `bind_contextvars()` | Add context to all logs | | `clear_contextvars()` | Clear request context | | `logger.exception()` | Log with traceback | | Processor | Purpose | |-----------|---------| | `TimeStamper(fmt="iso")` | Add timestamp | | `add_log_level` | Add level field | | `JSONRenderer()` | Output as JSON | | `ConsoleRenderer()` | Pretty console output | -
tracing.md 8 KB
# Distributed Tracing with OpenTelemetry Trace requests across services for debugging and performance analysis. ## Setup ```python from opentelemetry import trace from opentelemetry.sdk.trace import TracerProvider from opentelemetry.sdk.trace.export import BatchSpanProcessor from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter from opentelemetry.sdk.resources import Resource # Create resource with service info resource = Resource.create({ "service.name": "my-service", "service.version": "1.0.0", "deployment.environment": "production", }) # Create and configure tracer provider provider = TracerProvider(resource=resource) # Export to OTLP collector (Jaeger, Tempo, etc.) otlp_exporter = OTLPSpanExporter( endpoint="http://localhost:4317", insecure=True, ) provider.add_span_processor(BatchSpanProcessor(otlp_exporter)) # Set as global tracer provider trace.set_tracer_provider(provider) # Get tracer for your module tracer = trace.get_tracer(__name__) ``` ## FastAPI Auto-Instrumentation ```python from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor from opentelemetry.instrumentation.httpx import HTTPXClientInstrumentor from opentelemetry.instrumentation.sqlalchemy import SQLAlchemyInstrumentor from opentelemetry.instrumentation.redis import RedisInstrumentor # Instrument FastAPI FastAPIInstrumentor.instrument_app(app) # Instrument HTTP client HTTPXClientInstrumentor().instrument() # Instrument database SQLAlchemyInstrumentor().instrument(engine=engine) # Instrument Redis RedisInstrumentor().instrument() ``` ## Manual Instrumentation ```python from opentelemetry import trace from opentelemetry.trace import Status, StatusCode tracer = trace.get_tracer(__name__) async def process_order(order_id: int): """Process order with detailed tracing.""" with tracer.start_as_current_span("process_order") as span: # Add attributes span.set_attribute("order.id", order_id) span.set_attribute("order.type", "standard") # Nested spans with tracer.start_as_current_span("validate_order"): order = await validate(order_id) span.set_attribute("order.items", len(order.items)) with tracer.start_as_current_span("check_inventory"): await check_inventory(order.items) with tracer.start_as_current_span("process_payment") as payment_span: try: result = await charge_payment(order) payment_span.set_attribute("payment.amount", float(order.total)) except PaymentError as e: payment_span.set_status(Status(StatusCode.ERROR, str(e))) payment_span.record_exception(e) raise with tracer.start_as_current_span("send_confirmation"): await send_email(order.customer_email) span.set_status(Status(StatusCode.OK)) return order ``` ## Context Propagation ```python from opentelemetry import trace from opentelemetry.propagate import inject, extract from opentelemetry.trace.propagation.tracecontext import TraceContextTextMapPropagator propagator = TraceContextTextMapPropagator() # Inject context into outgoing HTTP headers async def call_external_service(data: dict): headers = {} inject(headers) # Adds traceparent header async with httpx.AsyncClient() as client: response = await client.post( "https://api.example.com/process", json=data, headers=headers, ) return response.json() # Extract context from incoming request (usually handled by instrumentation) @app.middleware("http") async def trace_middleware(request: Request, call_next): # Extract trace context from headers ctx = extract(dict(request.headers)) with tracer.start_as_current_span( f"{request.method} {request.url.path}", context=ctx, ): return await call_next(request) ``` ## Adding Events and Exceptions ```python from opentelemetry import trace tracer = trace.get_tracer(__name__) async def process_with_events(): with tracer.start_as_current_span("process") as span: # Add event (point-in-time occurrence) span.add_event("processing_started", { "items": 10, }) try: result = await heavy_processing() span.add_event("processing_completed", { "result_count": len(result), }) except Exception as e: # Record exception in span span.record_exception(e) span.set_status(Status(StatusCode.ERROR, str(e))) raise return result ``` ## Span Decorator ```python from functools import wraps from opentelemetry import trace tracer = trace.get_tracer(__name__) def traced(span_name: str | None = None, attributes: dict | None = None): """Decorator to trace function execution.""" def decorator(func): @wraps(func) async def async_wrapper(*args, **kwargs): name = span_name or f"{func.__module__}.{func.__name__}" with tracer.start_as_current_span(name) as span: if attributes: for key, value in attributes.items(): span.set_attribute(key, value) try: result = await func(*args, **kwargs) span.set_status(Status(StatusCode.OK)) return result except Exception as e: span.record_exception(e) span.set_status(Status(StatusCode.ERROR, str(e))) raise @wraps(func) def sync_wrapper(*args, **kwargs): name = span_name or f"{func.__module__}.{func.__name__}" with tracer.start_as_current_span(name) as span: if attributes: for key, value in attributes.items(): span.set_attribute(key, value) try: result = func(*args, **kwargs) span.set_status(Status(StatusCode.OK)) return result except Exception as e: span.record_exception(e) span.set_status(Status(StatusCode.ERROR, str(e))) raise if asyncio.iscoroutinefunction(func): return async_wrapper return sync_wrapper return decorator # Usage @traced("user.create", {"component": "users"}) async def create_user(user: UserCreate): return await db.create(user) ``` ## Linking Traces to Logs ```python import structlog from opentelemetry import trace def add_trace_context(_, __, event_dict): """Add trace context to log entries.""" span = trace.get_current_span() if span.is_recording(): ctx = span.get_span_context() event_dict["trace_id"] = format(ctx.trace_id, "032x") event_dict["span_id"] = format(ctx.span_id, "016x") return event_dict structlog.configure( processors=[ add_trace_context, structlog.processors.JSONRenderer(), ], ) ``` ## Sampling ```python from opentelemetry.sdk.trace.sampling import ( TraceIdRatioBased, ParentBased, ALWAYS_ON, ) # Sample 10% of traces sampler = TraceIdRatioBased(0.1) # Respect parent's sampling decision, default to 10% sampler = ParentBased(root=TraceIdRatioBased(0.1)) # Always sample (development) sampler = ALWAYS_ON provider = TracerProvider( resource=resource, sampler=sampler, ) ``` ## Quick Reference | Concept | Description | |---------|-------------| | Trace | Complete request journey | | Span | Single operation within trace | | Context | Propagated trace information | | Attribute | Key-value metadata on span | | Event | Point-in-time occurrence | | Instrumentation | Package | |-----------------|---------| | FastAPI | `opentelemetry-instrumentation-fastapi` | | httpx | `opentelemetry-instrumentation-httpx` | | SQLAlchemy | `opentelemetry-instrumentation-sqlalchemy` | | Redis | `opentelemetry-instrumentation-redis` | | Celery | `opentelemetry-instrumentation-celery` |
-
-
scripts
-
.gitkeep 0 B · in bundle
-
-
SKILL.md 5 KB
--- name: python-observability-ops description: "Observability patterns for Python applications. Triggers on: logging, metrics, tracing, opentelemetry, prometheus, observability, monitoring, structlog, correlation id." license: MIT compatibility: "Python 3.10+. Requires structlog, opentelemetry-api, prometheus-client." allowed-tools: "Read Write" metadata: author: claude-mods depends-on: python-async-ops related-skills: python-fastapi-ops, python-cli-ops --- # Python Observability Patterns Logging, metrics, and tracing for production applications. ## Structured Logging with structlog ```python import structlog # Configure structlog structlog.configure( processors=[ structlog.contextvars.merge_contextvars, structlog.processors.add_log_level, structlog.processors.TimeStamper(fmt="iso"), structlog.processors.JSONRenderer(), ], wrapper_class=structlog.make_filtering_bound_logger(logging.INFO), context_class=dict, logger_factory=structlog.PrintLoggerFactory(), ) logger = structlog.get_logger() # Usage logger.info("user_created", user_id=123, email="test@example.com") # Output: {"event": "user_created", "user_id": 123, "email": "test@example.com", "level": "info", "timestamp": "2024-01-15T10:00:00Z"} ``` ## Request Context Propagation ```python import structlog from contextvars import ContextVar from uuid import uuid4 request_id_var: ContextVar[str] = ContextVar("request_id", default="") def bind_request_context(request_id: str | None = None): """Bind request ID to logging context.""" rid = request_id or str(uuid4()) request_id_var.set(rid) structlog.contextvars.bind_contextvars(request_id=rid) return rid # FastAPI middleware @app.middleware("http") async def request_context_middleware(request, call_next): request_id = request.headers.get("X-Request-ID") or str(uuid4()) bind_request_context(request_id) response = await call_next(request) response.headers["X-Request-ID"] = request_id structlog.contextvars.clear_contextvars() return response ``` ## Prometheus Metrics ```python from prometheus_client import Counter, Histogram, Gauge, generate_latest from fastapi import FastAPI, Response # Define metrics REQUEST_COUNT = Counter( "http_requests_total", "Total HTTP requests", ["method", "endpoint", "status"] ) REQUEST_LATENCY = Histogram( "http_request_duration_seconds", "HTTP request latency", ["method", "endpoint"], buckets=[0.01, 0.05, 0.1, 0.5, 1.0, 5.0] ) ACTIVE_CONNECTIONS = Gauge( "active_connections", "Number of active connections" ) # Middleware to record metrics @app.middleware("http") async def metrics_middleware(request, call_next): ACTIVE_CONNECTIONS.inc() start = time.perf_counter() response = await call_next(request) duration = time.perf_counter() - start REQUEST_COUNT.labels( method=request.method, endpoint=request.url.path, status=response.status_code ).inc() REQUEST_LATENCY.labels( method=request.method, endpoint=request.url.path ).observe(duration) ACTIVE_CONNECTIONS.dec() return response # Metrics endpoint @app.get("/metrics") async def metrics(): return Response( content=generate_latest(), media_type="text/plain" ) ``` ## OpenTelemetry Tracing ```python from opentelemetry import trace from opentelemetry.sdk.trace import TracerProvider from opentelemetry.sdk.trace.export import BatchSpanProcessor from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter # Setup provider = TracerProvider() processor = BatchSpanProcessor(OTLPSpanExporter(endpoint="localhost:4317")) provider.add_span_processor(processor) trace.set_tracer_provider(provider) tracer = trace.get_tracer(__name__) # Manual instrumentation async def process_order(order_id: int): with tracer.start_as_current_span("process_order") as span: span.set_attribute("order_id", order_id) with tracer.start_as_current_span("validate_order"): await validate(order_id) with tracer.start_as_current_span("charge_payment"): await charge(order_id) ``` ## Quick Reference | Library | Purpose | |---------|---------| | structlog | Structured logging | | prometheus-client | Metrics collection | | opentelemetry | Distributed tracing | | Metric Type | Use Case | |-------------|----------| | Counter | Total requests, errors | | Histogram | Latencies, sizes | | Gauge | Current connections, queue size | ## Additional Resources - `./references/structured-logging.md` - structlog configuration, formatters - `./references/metrics.md` - Prometheus patterns, custom metrics - `./references/tracing.md` - OpenTelemetry, distributed tracing ## Assets - `./assets/logging-config.py` - Production logging configuration --- ## See Also **Prerequisites:** - `python-async-ops` - Async context propagation **Related Skills:** - `python-fastapi-ops` - API middleware for metrics/tracing - `python-cli-ops` - CLI logging patterns **Integration Skills:** - `python-database-ops` - Database query tracing
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.