api-design-ops
API design patterns for REST, gRPC, and GraphQL. Use for: api design, REST, gRPC, GraphQL, protobuf, schema design, api versioning, pagination, rate limiting, error format, OpenAPI, API authentication, JWT, OAuth2, API gateway, webhook, idempotency.
Install
npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/api-design-ops
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
git clone https://github.com/0xDarkMatter/claude-mods.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
API Design Ops
Comprehensive API design patterns covering REST (advanced), gRPC, and GraphQL. This skill provides decision frameworks, design patterns, and implementation guidance for building production APIs.
API Style Decision Tree
What kind of API do you need?
|
+-- Internal microservice-to-microservice?
| +-- High throughput, low latency needed? --> gRPC
| +-- Streaming (real-time data, logs)? --> gRPC (bidirectional streaming)
| +-- Simple request/response, team comfort? --> REST
|
+-- Public-facing API?
| +-- Third-party developers consuming it? --> REST (widest compatibility)
| +-- Mobile app with varied data needs? --> GraphQL
| +-- Browser-only, simple CRUD? --> REST
|
+-- Frontend for your own app?
| +-- Multiple clients with different data shapes? --> GraphQL
| +-- Single client, straightforward data? --> REST
| +-- Real-time updates needed? --> GraphQL subscriptions or SSE
|
+-- IoT / embedded / constrained devices?
| +-- Binary efficiency matters? --> gRPC
| +-- HTTP-only environments? --> REST
Quick Comparison
| Concern | REST | gRPC | GraphQL |
|---|---|---|---|
| Transport | HTTP/1.1+ | HTTP/2 | HTTP (any) |
| Serialization | JSON (text) | Protobuf (binary) | JSON (text) |
| Schema | OpenAPI (optional) | .proto (required) | SDL (required) |
| Browser support | Native | Via gRPC-Web/Connect | Native |
| Caching | HTTP caching built-in | Custom | Custom (normalized) |
| Learning curve | Low | Medium | Medium-High |
| Code generation | Optional | Required | Optional but recommended |
| Streaming | SSE, WebSocket | Native (4 patterns) | Subscriptions |
| Over-fetching | Common problem | No (typed) | Solved by design |
| File uploads | Multipart native | Chunked streaming | Multipart spec (awkward) |
REST Resource Design Quick Reference
Resource Naming
GET /users # Collection
GET /users/{id} # Singleton
GET /users/{id}/orders # Sub-collection
POST /users # Create
PUT /users/{id} # Full replace
PATCH /users/{id} # Partial update
DELETE /users/{id} # Remove
# Naming rules:
# - Plural nouns for collections: /users NOT /user
# - Kebab-case for multi-word: /line-items NOT /lineItems
# - No verbs in URLs: POST /orders NOT POST /create-order
# - Max 3 levels deep: /users/{id}/orders (not /users/{id}/orders/{oid}/items/{iid}/details)
HTTP Methods and Status Codes
| Method | Success | Empty | Invalid | Not Found | Conflict |
|---|---|---|---|---|---|
| GET | 200 | 200 (empty array) | 400 | 404 | - |
| POST | 201 + Location | - | 400/422 | - | 409 |
| PUT | 200 | - | 400/422 | 404 | 409 |
| PATCH | 200 | - | 400/422 | 404 | 409 |
| DELETE | 204 | 204 (already gone) | 400 | 404 | 409 |
HATEOAS (When Worth It)
Use when: public APIs where discoverability matters, long-lived APIs, APIs that evolve frequently. Skip when: internal microservices, mobile backends, tight coupling is acceptable.
{
"id": "order-123",
"status": "shipped",
"_links": {
"self": { "href": "/orders/order-123" },
"track": { "href": "/orders/order-123/tracking" },
"cancel": { "href": "/orders/order-123", "method": "DELETE" }
}
}
Pagination Decision Tree
What's your data like?
|
+-- Stable data, UI needs "jump to page 5"?
| --> Offset pagination: ?page=5&per_page=20
| Tradeoff: Slow on large offsets (OFFSET 10000), inconsistent with inserts
|
+-- Large dataset, forward-only traversal?
| --> Cursor pagination: ?after=eyJpZCI6MTIzfQ&limit=20
| Tradeoff: No random page access, but consistent and fast
|
+-- Real-time feed, ordered by timestamp or ID?
| --> Keyset pagination: ?created_after=2024-01-01T00:00:00Z&limit=20
| Tradeoff: Requires a unique, sequential column; no page jumping
Response Envelope
{
"data": [...],
"pagination": {
"total": 1432,
"limit": 20,
"has_more": true,
"next_cursor": "eyJpZCI6MTQzMn0="
}
}
Error Response Format (RFC 7807)
All APIs should use Problem Details (RFC 7807 / RFC 9457):
{
"type": "https://api.example.com/errors/insufficient-funds",
"title": "Insufficient Funds",
"status": 422,
"detail": "Account xxxx-1234 has a balance of $10.00, but the transfer requires $25.00.",
"instance": "/transfers/txn-abc-123",
"balance": 1000,
"required": 2500
}
Field Reference
| Field | Required | Description |
|---|---|---|
type |
Yes | URI identifying the error type (stable, documentable) |
title |
Yes | Human-readable summary (same for all instances of this type) |
status |
Yes | HTTP status code |
detail |
Yes | Human-readable explanation specific to this occurrence |
instance |
No | URI identifying the specific occurrence |
| (extensions) | No | Additional machine-readable fields |
Validation Errors
{
"type": "https://api.example.com/errors/validation",
"title": "Validation Failed",
"status": 422,
"detail": "The request body contains 2 validation errors.",
"errors": [
{ "field": "email", "message": "Must be a valid email address", "code": "invalid_format" },
{ "field": "age", "message": "Must be at least 18", "code": "out_of_range", "min": 18 }
]
}
Versioning Strategies
| Strategy | Example | Pros | Cons |
|---|---|---|---|
| URL path | /v2/users |
Obvious, cacheable, easy routing | URL pollution, hard to sunset |
| Accept header | Accept: application/vnd.api.v2+json |
Clean URLs, content negotiation | Hidden, harder to test |
| Query param | /users?version=2 |
Easy to add | Pollutes query string, caching issues |
| Date-based | API-Version: 2024-01-15 |
Granular evolution (Stripe style) | Complex implementation |
Recommendation
- Public APIs: URL path versioning (
/v1/) - simplicity wins - Internal APIs: Header or no versioning (deploy in lockstep)
- Evolving APIs: Date-based (Stripe model) if you have the engineering investment
Breaking Change Rules
A breaking change is anything that can cause existing clients to fail:
- Removing a field from a response
- Renaming a field
- Changing a field's type
- Adding a required field to a request
- Changing URL structure
- Changing error formats
- Removing an endpoint
Non-breaking (safe):
- Adding optional fields to requests
- Adding fields to responses
- Adding new endpoints
- Adding new enum values (if client handles unknown values)
Rate Limiting Design
Algorithms
| Algorithm | Behavior | Use When |
|---|---|---|
| Token bucket | Allows bursts, refills at steady rate | General API rate limiting |
| Sliding window | Smooth distribution, no burst | Strict fairness needed |
| Fixed window | Simple, potential burst at boundary | Low-stakes limiting |
| Leaky bucket | Constant output rate | Queue processing |
Response Headers
X-RateLimit-Limit: 1000 # Max requests per window
X-RateLimit-Remaining: 743 # Requests left in current window
X-RateLimit-Reset: 1672531200 # Unix timestamp when window resets
Retry-After: 30 # Seconds to wait (on 429)
429 Response Body
{
"type": "https://api.example.com/errors/rate-limit-exceeded",
"title": "Rate Limit Exceeded",
"status": 429,
"detail": "You have exceeded 1000 requests per hour. Try again in 30 seconds.",
"retry_after": 30
}
Idempotency
Which Methods Need Idempotency Keys?
| Method | Idempotent by spec? | Needs key? |
|---|---|---|
| GET | Yes | No |
| PUT | Yes | No (full replacement is naturally idempotent) |
| DELETE | Yes | No |
| PATCH | No | Recommended for critical operations |
| POST | No | Yes (always for payments, orders, transfers) |
Implementation
POST /payments
Idempotency-Key: 550e8400-e29b-41d4-a716-446655440000
Content-Type: application/json
{ "amount": 2500, "currency": "usd", "customer": "cust_123" }
Server-side:
- Receive request with
Idempotency-Keyheader - Check if key exists in store (Redis, DB)
- If exists: return stored response (same status code + body)
- If not: process request, store response keyed by idempotency key
- Keys expire after 24-48 hours
Authentication Overview
| Method | Use When | Security Level |
|---|---|---|
| API Key | Server-to-server, internal, simple | Low-Medium |
| JWT (Bearer) | Stateless auth, microservices | Medium-High |
| OAuth2 + PKCE | Third-party access, user delegation | High |
| mTLS | Service mesh, zero-trust infra | Very High |
Decision Guide
Who is authenticating?
|
+-- Your own frontend? --> JWT (short-lived access + refresh token)
+-- Third-party developer? --> OAuth2 (client credentials for server, PKCE for SPA)
+-- Another internal service? --> mTLS or JWT with service accounts
+-- Quick prototype? --> API key (but plan migration)
Gotchas Table
| Gotcha | Problem | Prevention |
|---|---|---|
| Breaking changes in "non-breaking" release | Client crashes | Additive-only policy, contract tests |
| N+1 in REST APIs | 100 users = 101 queries | Compound documents, ?include=, or GraphQL |
| Over-fetching | Mobile gets 50 fields, needs 3 | Sparse fieldsets ?fields=id,name or GraphQL |
| Under-fetching | 3 requests to build one view | Composite endpoints or BFF pattern |
| CORS misconfiguration | Frontend can't reach API | Explicit allowed origins, never * with credentials |
| Missing Content-Type | 415 or silent parsing failure | Validate Content-Type on every mutation endpoint |
| Large payloads without pagination | OOM, timeouts | Always paginate collections, set max page size |
| Inconsistent date formats | Parsing hell | ISO 8601 everywhere: 2024-01-15T10:30:00Z |
| No request IDs | Impossible to debug | Generate X-Request-ID, propagate through services |
| Enum evolution | New value breaks old client | Document that enums may grow, clients must handle unknown |
| Missing idempotency | Duplicate charges, orders | Idempotency keys on all POST endpoints with side effects |
| Unbounded query complexity | GraphQL DoS | Depth limiting, cost analysis, persisted queries |
Reference Files
| File | Contents |
|---|---|
references/rest-advanced.md |
Resource modeling, PATCH strategies, caching, webhooks, bulk ops |
references/grpc.md |
Protobuf, service definitions, Go/Rust, streaming, error handling |
references/graphql.md |
Schema design, resolvers, DataLoader, federation, performance |
references/api-security.md |
JWT, OAuth2, CORS, rate limiting, OWASP API Top 10 |
Files (claude-mods)
-
assets
-
.gitkeep 0 B · in bundle
-
-
references
-
api-security.md 18.8 KB
# API Security Patterns ## Table of Contents - [API Key Management](#api-key-management) - [JWT (JSON Web Tokens)](#jwt-json-web-tokens) - [OAuth2 Flows](#oauth2-flows) - [CORS](#cors) - [Rate Limiting Implementation](#rate-limiting-implementation) - [Input Validation](#input-validation) - [API Versioning and Deprecation](#api-versioning-and-deprecation) - [Transport Security](#transport-security) - [OWASP API Security Top 10](#owasp-api-security-top-10) --- ## API Key Management ### Generation ```go import "crypto/rand" func generateAPIKey() (string, error) { // 32 bytes = 256 bits of entropy b := make([]byte, 32) if _, err := rand.Read(b); err != nil { return "", err } // Prefix for easy identification and revocation return "sk_live_" + base64.URLEncoding.EncodeToString(b), nil } ``` ### Storage ``` NEVER store API keys in plaintext. Store: hash(api_key) in database Lookup: hash(incoming_key), compare to stored hashes Display: show only last 4 chars to user ("sk_live_...a1b2") ``` ```go import "crypto/sha256" func hashAPIKey(key string) string { h := sha256.Sum256([]byte(key)) return hex.EncodeToString(h[:]) } ``` ### Scoping ```json { "key_id": "key_abc123", "name": "Production Read-Only", "permissions": ["read:users", "read:orders"], "rate_limit": 1000, "allowed_ips": ["203.0.113.0/24"], "expires_at": "2025-01-15T00:00:00Z", "created_at": "2024-01-15T00:00:00Z" } ``` ### Rotation Strategy 1. Generate new key 2. Both old and new keys work (grace period: 24-72 hours) 3. Client updates to new key 4. Old key is revoked 5. Log all key usage for audit ## JWT (JSON Web Tokens) ### Structure ``` header.payload.signature # Header { "alg": "RS256", # Algorithm (RS256, ES256 - avoid HS256 for APIs) "typ": "JWT", "kid": "key-2024-01" # Key ID for rotation } # Payload (Claims) { "iss": "https://auth.example.com", # Issuer "sub": "user-123", # Subject (user ID) "aud": "https://api.example.com", # Audience "exp": 1705312200, # Expires (15 min from now) "iat": 1705311300, # Issued at "jti": "unique-token-id", # JWT ID (for revocation) "scope": "read:users write:orders", # Permissions "org_id": "org-456" # Custom claim } ``` ### Signing and Verification (Go) ```go import "github.com/golang-jwt/jwt/v5" // Sign (auth service) func createAccessToken(userID string, scopes []string) (string, error) { claims := jwt.MapClaims{ "sub": userID, "scope": strings.Join(scopes, " "), "exp": time.Now().Add(15 * time.Minute).Unix(), "iat": time.Now().Unix(), "iss": "https://auth.example.com", } token := jwt.NewWithClaims(jwt.SigningMethodRS256, claims) token.Header["kid"] = currentKeyID return token.SignedString(privateKey) } // Verify (API service) func verifyToken(tokenString string) (*jwt.Token, error) { return jwt.Parse(tokenString, func(token *jwt.Token) (interface{}, error) { // Validate algorithm if _, ok := token.Method.(*jwt.SigningMethodRSA); !ok { return nil, fmt.Errorf("unexpected signing method: %v", token.Header["alg"]) } // Look up public key by kid kid, _ := token.Header["kid"].(string) pubKey, err := getPublicKey(kid) if err != nil { return nil, fmt.Errorf("unknown key ID: %s", kid) } return pubKey, nil }, jwt.WithValidMethods([]string{"RS256"}), jwt.WithIssuer("https://auth.example.com"), jwt.WithAudience("https://api.example.com"), ) } ``` ### Refresh Token Flow ``` 1. Login: POST /auth/login Response: { access_token (15 min), refresh_token (7 days) } 2. API calls: Authorization: Bearer <access_token> 3. Token expired (401): POST /auth/refresh Body: { refresh_token } Response: { access_token (new, 15 min), refresh_token (rotated) } 4. Refresh token expired/revoked: redirect to login ``` ### Token Storage | Environment | Access Token | Refresh Token | |-------------|-------------|---------------| | Browser SPA | Memory (JS variable) | HttpOnly Secure cookie | | Mobile app | Secure storage (Keychain/Keystore) | Secure storage | | Server-to-server | Environment variable | Environment variable | **Never store tokens in:** - localStorage (XSS vulnerable) - sessionStorage (XSS vulnerable) - Non-HttpOnly cookies (XSS vulnerable) - URL parameters (logged, cached, leaked via Referer) ## OAuth2 Flows ### Authorization Code + PKCE (SPAs, Mobile) ``` 1. Client generates: code_verifier (random 43-128 chars) code_challenge = BASE64URL(SHA256(code_verifier)) 2. Redirect to authorization server: GET /authorize? response_type=code& client_id=app-123& redirect_uri=https://app.example.com/callback& scope=read:profile write:orders& state=random-csrf-token& code_challenge=E9Melhoa2OwvFrEMTJguCHaoeK1t8URWbuGJSstw-cM& code_challenge_method=S256 3. User authenticates, consents 4. Redirect back with code: GET /callback?code=auth-code-xyz&state=random-csrf-token 5. Exchange code for tokens: POST /token { "grant_type": "authorization_code", "code": "auth-code-xyz", "redirect_uri": "https://app.example.com/callback", "client_id": "app-123", "code_verifier": "the-original-random-string" } 6. Response: { "access_token": "eyJ...", "token_type": "Bearer", "expires_in": 900, "refresh_token": "rt_...", "scope": "read:profile write:orders" } ``` ### Client Credentials (Server-to-Server) ``` POST /token Content-Type: application/x-www-form-urlencoded grant_type=client_credentials& client_id=service-abc& client_secret=secret-xyz& scope=read:users Response: { "access_token": "eyJ...", "token_type": "Bearer", "expires_in": 3600 } ``` ### Device Flow (CLI Tools, Smart TVs) ``` 1. Device requests code: POST /device/code { "client_id": "cli-app", "scope": "read:profile" } Response: { "device_code": "device-code-abc", "user_code": "ABCD-1234", "verification_uri": "https://auth.example.com/device", "expires_in": 600, "interval": 5 } 2. Display to user: "Go to https://auth.example.com/device and enter ABCD-1234" 3. Device polls (every 5 seconds): POST /token { "grant_type": "urn:ietf:params:oauth:grant-type:device_code", "device_code": "device-code-abc", "client_id": "cli-app" } While pending: { "error": "authorization_pending" } When approved: { "access_token": "eyJ...", ... } ``` ### Flow Selection Guide | Scenario | Flow | |----------|------| | SPA (browser) | Authorization Code + PKCE | | Mobile app | Authorization Code + PKCE | | Server-to-server | Client Credentials | | CLI tool | Device Flow | | Legacy (avoid) | Implicit (deprecated), ROPC (deprecated) | ## CORS ### Configuration ```go func corsMiddleware(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { origin := r.Header.Get("Origin") // Whitelist specific origins (NEVER use * with credentials) allowedOrigins := map[string]bool{ "https://app.example.com": true, "https://staging.example.com": true, } if allowedOrigins[origin] { w.Header().Set("Access-Control-Allow-Origin", origin) w.Header().Set("Access-Control-Allow-Credentials", "true") w.Header().Set("Access-Control-Allow-Methods", "GET, POST, PUT, PATCH, DELETE, OPTIONS") w.Header().Set("Access-Control-Allow-Headers", "Authorization, Content-Type, X-Request-ID, Idempotency-Key") w.Header().Set("Access-Control-Expose-Headers", "X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset") w.Header().Set("Access-Control-Max-Age", "86400") // Cache preflight for 24h } // Handle preflight if r.Method == "OPTIONS" { w.WriteHeader(http.StatusNoContent) return } next.ServeHTTP(w, r) }) } ``` ### Common CORS Mistakes | Mistake | Risk | Fix | |---------|------|-----| | `Access-Control-Allow-Origin: *` with credentials | Credential theft | Whitelist specific origins | | Reflecting `Origin` header without validation | Any origin allowed | Check against whitelist | | Missing `Vary: Origin` | Cache poisoning | Add `Vary: Origin` header | | Not handling preflight (OPTIONS) | Mutations blocked | Return 204 for OPTIONS | | Allowing all headers | Header injection | Whitelist specific headers | ## Rate Limiting Implementation ### Token Bucket (Go + Redis) ```go import "github.com/redis/go-redis/v9" type RateLimiter struct { redis *redis.Client limit int // Max tokens window time.Duration // Refill window } func (rl *RateLimiter) Allow(ctx context.Context, key string) (bool, RateLimitInfo, error) { now := time.Now().Unix() windowKey := fmt.Sprintf("ratelimit:%s:%d", key, now/int64(rl.window.Seconds())) pipe := rl.redis.Pipeline() incr := pipe.Incr(ctx, windowKey) pipe.Expire(ctx, windowKey, rl.window) _, err := pipe.Exec(ctx) if err != nil { return false, RateLimitInfo{}, err } count := incr.Val() remaining := rl.limit - int(count) if remaining < 0 { remaining = 0 } info := RateLimitInfo{ Limit: rl.limit, Remaining: remaining, Reset: time.Unix(((now/int64(rl.window.Seconds()))+1)*int64(rl.window.Seconds()), 0), } return count <= int64(rl.limit), info, nil } type RateLimitInfo struct { Limit int Remaining int Reset time.Time } ``` ### Middleware ```go func rateLimitMiddleware(limiter *RateLimiter) func(http.Handler) http.Handler { return func(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { // Key by API key, user ID, or IP key := extractRateLimitKey(r) allowed, info, err := limiter.Allow(r.Context(), key) if err != nil { http.Error(w, "Internal Server Error", 500) return } // Always set rate limit headers w.Header().Set("X-RateLimit-Limit", strconv.Itoa(info.Limit)) w.Header().Set("X-RateLimit-Remaining", strconv.Itoa(info.Remaining)) w.Header().Set("X-RateLimit-Reset", strconv.FormatInt(info.Reset.Unix(), 10)) if !allowed { retryAfter := int(time.Until(info.Reset).Seconds()) w.Header().Set("Retry-After", strconv.Itoa(retryAfter)) w.WriteHeader(http.StatusTooManyRequests) json.NewEncoder(w).Encode(map[string]interface{}{ "type": "https://api.example.com/errors/rate-limit", "title": "Rate Limit Exceeded", "status": 429, "detail": fmt.Sprintf("Rate limit of %d requests per hour exceeded", info.Limit), "retry_after": retryAfter, }) return } next.ServeHTTP(w, r) }) } } ``` ### Tiered Rate Limits | Tier | Requests/Hour | Burst | Use Case | |------|---------------|-------|----------| | Free | 100 | 10/min | Trial users | | Basic | 1,000 | 100/min | Paid individuals | | Pro | 10,000 | 500/min | Teams | | Enterprise | 100,000 | 2,000/min | Custom SLA | ## Input Validation ### Validate at the Boundary ```go // Use a validation library, not manual checks import "github.com/go-playground/validator/v10" type CreateUserRequest struct { Name string `json:"name" validate:"required,min=2,max=100"` Email string `json:"email" validate:"required,email"` Age int `json:"age" validate:"omitempty,min=13,max=150"` Website string `json:"website" validate:"omitempty,url"` Role string `json:"role" validate:"required,oneof=admin member viewer"` Password string `json:"password" validate:"required,min=8,max=128"` } var validate = validator.New() func handleCreateUser(w http.ResponseWriter, r *http.Request) { var req CreateUserRequest if err := json.NewDecoder(r.Body).Decode(&req); err != nil { respondError(w, 400, "Invalid JSON body") return } if err := validate.Struct(req); err != nil { validationErrors := err.(validator.ValidationErrors) respondValidationErrors(w, validationErrors) return } // Input is now validated - proceed } ``` ### Validation Checklist | Check | Why | |-------|-----| | Max request body size | Prevent memory exhaustion | | String length limits | Prevent storage abuse | | Enum validation | Reject unknown values | | URL validation | Prevent SSRF (whitelist schemes) | | Email format | Reject obviously invalid | | Numeric bounds | Prevent overflow, nonsensical values | | Array max length | Prevent excessive processing | | Nested object depth | Prevent deep recursion | | Content-Type validation | Ensure expected format | | UTF-8 validation | Prevent encoding attacks | ### Schema Validation (OpenAPI) ```go import "github.com/getkin/kin-openapi/openapi3filter" // Validate requests against OpenAPI spec automatically router, _ := gorillamux.NewRouter(spec) func validationMiddleware(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { route, pathParams, _ := router.FindRoute(r) input := &openapi3filter.RequestValidationInput{ Request: r, PathParams: pathParams, Route: route, } if err := openapi3filter.ValidateRequest(r.Context(), input); err != nil { respondError(w, 400, err.Error()) return } next.ServeHTTP(w, r) }) } ``` ## API Versioning and Deprecation ### Deprecation Timeline ``` 1. Announce deprecation (minimum 6 months before removal) - Add Deprecation header to responses - Update API documentation - Email API key owners 2. Warning period (3-6 months) Deprecation: true Sunset: Sat, 15 Jun 2025 00:00:00 GMT Link: <https://docs.example.com/migration-guide>; rel="deprecation" 3. Migration support - Provide migration guide - Offer parallel running of old and new versions - Log deprecated endpoint usage for targeted outreach 4. Removal - Return 410 Gone with migration info - Keep 410 response for 6+ months ``` ### Sunset Header (RFC 8594) ``` HTTP/1.1 200 OK Sunset: Sat, 15 Jun 2025 00:00:00 GMT Deprecation: true Link: <https://api.example.com/v3/users>; rel="successor-version" ``` ## Transport Security ### TLS Configuration ```go tlsConfig := &tls.Config{ MinVersion: tls.VersionTLS12, CipherSuites: []uint16{ tls.TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384, tls.TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256, tls.TLS_ECDHE_ECDSA_WITH_AES_256_GCM_SHA384, tls.TLS_ECDHE_ECDSA_WITH_AES_128_GCM_SHA256, }, PreferServerCipherSuites: true, } server := &http.Server{ TLSConfig: tlsConfig, // ... } ``` ### Security Headers ```go func securityHeaders(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { w.Header().Set("Strict-Transport-Security", "max-age=31536000; includeSubDomains") w.Header().Set("X-Content-Type-Options", "nosniff") w.Header().Set("X-Frame-Options", "DENY") w.Header().Set("Content-Security-Policy", "default-src 'none'; frame-ancestors 'none'") w.Header().Set("Cache-Control", "no-store") // For API responses with sensitive data w.Header().Set("X-Request-ID", generateRequestID()) next.ServeHTTP(w, r) }) } ``` ## OWASP API Security Top 10 ### 2023 Edition | # | Risk | Description | Prevention | |---|------|-------------|------------| | 1 | **Broken Object-Level Auth (BOLA)** | User accesses other users' objects via ID manipulation | Check ownership in every endpoint: `WHERE id = ? AND user_id = ?` | | 2 | **Broken Authentication** | Weak auth, credential stuffing, missing rate limits on login | Rate limit login, use strong password hashing (argon2id), MFA | | 3 | **Broken Object Property-Level Auth** | Mass assignment, excessive data exposure | Explicit allowlists for input fields, separate input/output DTOs | | 4 | **Unrestricted Resource Consumption** | No rate limits, unbounded queries, large payloads | Rate limiting, pagination limits, request size limits, timeouts | | 5 | **Broken Function-Level Auth** | Admin endpoints accessible to regular users | Role-based access control, deny by default, test auth on every endpoint | | 6 | **Unrestricted Access to Sensitive Business Flows** | Automated abuse (ticket scalping, spam) | Rate limiting, CAPTCHA, device fingerprinting, business logic limits | | 7 | **Server-Side Request Forgery (SSRF)** | API fetches attacker-controlled URLs | Validate/whitelist URLs, block internal networks, use allowlists | | 8 | **Security Misconfiguration** | Default configs, verbose errors, missing CORS | Harden defaults, strip stack traces in production, audit configs | | 9 | **Improper Inventory Management** | Shadow APIs, deprecated endpoints still active | API gateway, version inventory, automated discovery, sunset old versions | | 10 | **Unsafe Consumption of APIs** | Trusting third-party API responses without validation | Validate all external API responses, set timeouts, use TLS | ### BOLA Prevention (Most Common API Vulnerability) ```go // BAD: Only checks if resource exists func getOrder(w http.ResponseWriter, r *http.Request) { orderID := chi.URLParam(r, "id") order, _ := db.GetOrder(orderID) // Anyone can access any order! json.NewEncoder(w).Encode(order) } // GOOD: Checks ownership func getOrder(w http.ResponseWriter, r *http.Request) { orderID := chi.URLParam(r, "id") userID := r.Context().Value(userIDKey).(string) order, err := db.GetOrderForUser(orderID, userID) // SQL: SELECT * FROM orders WHERE id = $1 AND user_id = $2 if err != nil { respondError(w, 404, "Order not found") // 404, not 403 (don't leak existence) return } json.NewEncoder(w).Encode(order) } ``` ### Mass Assignment Prevention ```go // BAD: Binding all fields from request func updateUser(w http.ResponseWriter, r *http.Request) { var user User json.NewDecoder(r.Body).Decode(&user) // Attacker sets role=admin! db.Save(&user) } // GOOD: Explicit allowlist of updatable fields type UpdateUserInput struct { Name *string `json:"name"` Email *string `json:"email"` // role is NOT here - cannot be set via API } func updateUser(w http.ResponseWriter, r *http.Request) { var input UpdateUserInput json.NewDecoder(r.Body).Decode(&input) user, _ := db.GetUser(userID) if input.Name != nil { user.Name = *input.Name } if input.Email != nil { user.Email = *input.Email } db.Save(&user) } ``` -
graphql.md 19.6 KB
# GraphQL Patterns ## Table of Contents - [Schema Design](#schema-design) - [Resolver Patterns](#resolver-patterns) - [Authentication and Authorization](#authentication-and-authorization) - [Error Handling](#error-handling) - [Pagination](#pagination) - [Fragments, Interfaces, Unions](#fragments-interfaces-unions) - [Schema Stitching and Federation](#schema-stitching-and-federation) - [Code-First vs Schema-First](#code-first-vs-schema-first) - [Performance](#performance) - [TypeScript and GraphQL](#typescript-and-graphql) - [When GraphQL Is Overkill](#when-graphql-is-overkill) --- ## Schema Design ### Types, Queries, and Mutations ```graphql # Scalar types: String, Int, Float, Boolean, ID # Custom scalars for domain types scalar DateTime scalar Email scalar URL type User { id: ID! name: String! email: Email! avatar: URL role: UserRole! posts(first: Int = 10, after: String): PostConnection! createdAt: DateTime! } enum UserRole { ADMIN MEMBER VIEWER } # Queries - read operations type Query { user(id: ID!): User users( first: Int = 20 after: String filter: UserFilter orderBy: UserOrderBy = CREATED_AT_DESC ): UserConnection! me: User! # Current authenticated user } # Mutations - write operations type Mutation { createUser(input: CreateUserInput!): CreateUserPayload! updateUser(input: UpdateUserInput!): UpdateUserPayload! deleteUser(id: ID!): DeleteUserPayload! } # Input types (separate from output types) input CreateUserInput { name: String! email: Email! role: UserRole = MEMBER } input UpdateUserInput { id: ID! name: String email: Email role: UserRole } input UserFilter { role: UserRole search: String createdAfter: DateTime } enum UserOrderBy { CREATED_AT_ASC CREATED_AT_DESC NAME_ASC NAME_DESC } ``` ### Mutation Payloads Always return a payload type (not the entity directly): ```graphql type CreateUserPayload { user: User! clientMutationId: String # Relay convention } type UpdateUserPayload { user: User! } type DeleteUserPayload { deletedId: ID! success: Boolean! } # For operations that can partially fail type BulkDeleteUsersPayload { deletedIds: [ID!]! errors: [BulkError!]! } type BulkError { id: ID! message: String! code: ErrorCode! } ``` ### Subscriptions ```graphql type Subscription { # Simple subscription orderStatusChanged(orderId: ID!): Order! # Filtered subscription newMessage(channelId: ID!): Message! # With initial state userPresence(teamId: ID!): PresenceEvent! } enum PresenceEventType { ONLINE OFFLINE AWAY } type PresenceEvent { user: User! type: PresenceEventType! timestamp: DateTime! } ``` ## Resolver Patterns ### Basic Resolver Structure (TypeScript) ```typescript const resolvers: Resolvers = { Query: { user: async (_, { id }, context) => { return context.dataSources.users.findById(id); }, users: async (_, { first, after, filter }, context) => { return context.dataSources.users.findMany({ first, after, filter }); }, me: async (_, __, context) => { if (!context.currentUser) { throw new AuthenticationError("Not authenticated"); } return context.currentUser; }, }, Mutation: { createUser: async (_, { input }, context) => { const user = await context.dataSources.users.create(input); return { user }; }, }, // Field-level resolver (runs when field is requested) User: { posts: async (parent, { first, after }, context) => { return context.dataSources.posts.findByUserId(parent.id, { first, after }); }, // Simple field mapping (usually not needed) email: (parent) => parent.email, }, }; ``` ### The N+1 Problem and DataLoader Without DataLoader: ``` Query { users(first: 10) { posts { title } } } # 1 query for users + 10 queries for posts = 11 queries ``` With DataLoader: ```typescript import DataLoader from "dataloader"; // Create per-request DataLoader instances function createLoaders() { return { postsByUserId: new DataLoader<string, Post[]>(async (userIds) => { // Single batched query: SELECT * FROM posts WHERE user_id IN (...) const posts = await db.posts.findMany({ where: { userId: { in: [...userIds] } }, }); // Map results back to input order const postsByUser = new Map<string, Post[]>(); for (const post of posts) { const existing = postsByUser.get(post.userId) || []; existing.push(post); postsByUser.set(post.userId, existing); } return userIds.map((id) => postsByUser.get(id) || []); }), userById: new DataLoader<string, User | null>(async (ids) => { const users = await db.users.findMany({ where: { id: { in: [...ids] } }, }); const userMap = new Map(users.map((u) => [u.id, u])); return ids.map((id) => userMap.get(id) || null); }), }; } // In resolver const resolvers = { User: { posts: (parent, args, context) => { return context.loaders.postsByUserId.load(parent.id); }, }, Post: { author: (parent, args, context) => { return context.loaders.userById.load(parent.authorId); }, }, }; ``` ## Authentication and Authorization ### Context Setup ```typescript // Server setup - extract user from token const server = new ApolloServer({ typeDefs, resolvers, context: async ({ req }) => { const token = req.headers.authorization?.replace("Bearer ", ""); let currentUser = null; if (token) { try { const decoded = await verifyJWT(token); currentUser = await db.users.findById(decoded.sub); } catch { // Invalid token - currentUser remains null } } return { currentUser, loaders: createLoaders(), dataSources: createDataSources(), }; }, }); ``` ### Authorization Patterns **Directive-based (schema-level):** ```graphql directive @auth(requires: UserRole = MEMBER) on FIELD_DEFINITION | OBJECT type Query { users: [User!]! @auth(requires: ADMIN) me: User! @auth } type User { email: Email! @auth(requires: ADMIN) # Only admins see emails name: String! # Public field } ``` ```typescript // Directive implementation class AuthDirective extends SchemaDirectiveVisitor { visitFieldDefinition(field: GraphQLField<any, any>) { const requiredRole = this.args.requires; const originalResolve = field.resolve || defaultFieldResolver; field.resolve = async (parent, args, context, info) => { if (!context.currentUser) { throw new AuthenticationError("Authentication required"); } if (requiredRole && context.currentUser.role !== requiredRole) { throw new ForbiddenError("Insufficient permissions"); } return originalResolve(parent, args, context, info); }; } } ``` **Resolver-level authorization:** ```typescript const resolvers = { Mutation: { deleteUser: async (_, { id }, context) => { // Only admins or the user themselves if (context.currentUser.role !== "ADMIN" && context.currentUser.id !== id) { throw new ForbiddenError("Cannot delete other users"); } await context.dataSources.users.delete(id); return { deletedId: id, success: true }; }, }, }; ``` ## Error Handling ### GraphQL Error Format ```json { "data": { "createUser": null }, "errors": [ { "message": "Email already exists", "locations": [{ "line": 2, "column": 3 }], "path": ["createUser"], "extensions": { "code": "CONFLICT", "field": "email", "timestamp": "2024-01-15T10:30:00Z" } } ] } ``` ### Error Classification ```typescript // Custom error classes class ValidationError extends GraphQLError { constructor(message: string, field: string) { super(message, { extensions: { code: "VALIDATION_ERROR", field, }, }); } } class BusinessRuleError extends GraphQLError { constructor(message: string, rule: string) { super(message, { extensions: { code: "BUSINESS_RULE_VIOLATION", rule, }, }); } } // Usage in resolvers const resolvers = { Mutation: { createUser: async (_, { input }, context) => { if (!isValidEmail(input.email)) { throw new ValidationError("Invalid email format", "email"); } const existing = await context.dataSources.users.findByEmail(input.email); if (existing) { throw new BusinessRuleError("Email already registered", "unique_email"); } const user = await context.dataSources.users.create(input); return { user }; }, }, }; ``` ### Partial Success Pattern ```graphql type Mutation { bulkCreateUsers(inputs: [CreateUserInput!]!): BulkCreateResult! } type BulkCreateResult { users: [User!]! errors: [CreateError!]! totalRequested: Int! totalCreated: Int! } type CreateError { index: Int! # Which input failed message: String! code: String! } ``` ## Pagination ### Relay Connection Spec ```graphql type Query { users( first: Int # Forward pagination after: String # Cursor last: Int # Backward pagination before: String # Cursor ): UserConnection! } type UserConnection { edges: [UserEdge!]! pageInfo: PageInfo! totalCount: Int # Optional - expensive on large datasets } type UserEdge { node: User! cursor: String! # Opaque cursor for this edge } type PageInfo { hasNextPage: Boolean! hasPreviousPage: Boolean! startCursor: String endCursor: String } ``` ### Implementation ```typescript async function connectionFromQuery<T>( query: SelectQueryBuilder<T>, args: { first?: number; after?: string; last?: number; before?: string } ): Promise<Connection<T>> { const limit = args.first || args.last || 20; const maxLimit = 100; const effectiveLimit = Math.min(limit, maxLimit); let afterId: string | null = null; if (args.after) { afterId = Buffer.from(args.after, "base64").toString("utf8"); } if (afterId) { query = query.where("id > :afterId", { afterId }); } // Fetch one extra to determine hasNextPage const items = await query .orderBy("id", "ASC") .take(effectiveLimit + 1) .getMany(); const hasNextPage = items.length > effectiveLimit; const nodes = hasNextPage ? items.slice(0, effectiveLimit) : items; const edges = nodes.map((node) => ({ node, cursor: Buffer.from(node.id).toString("base64"), })); return { edges, pageInfo: { hasNextPage, hasPreviousPage: !!args.after, startCursor: edges[0]?.cursor || null, endCursor: edges[edges.length - 1]?.cursor || null, }, }; } ``` ### Simple Pagination (Alternative) If Relay connections are overkill: ```graphql type Query { users(limit: Int = 20, offset: Int = 0): UserList! } type UserList { items: [User!]! total: Int! hasMore: Boolean! } ``` ## Fragments, Interfaces, Unions ### Fragments (Client-Side Reuse) ```graphql # Define reusable field sets fragment UserBasic on User { id name avatar } fragment UserDetailed on User { ...UserBasic email role createdAt posts(first: 5) { edges { node { id title } } } } # Use in queries query { me { ...UserDetailed } users(first: 10) { edges { node { ...UserBasic } } } } ``` ### Interfaces (Shared Fields) ```graphql interface Node { id: ID! } interface Timestamped { createdAt: DateTime! updatedAt: DateTime! } type User implements Node & Timestamped { id: ID! name: String! createdAt: DateTime! updatedAt: DateTime! } type Post implements Node & Timestamped { id: ID! title: String! createdAt: DateTime! updatedAt: DateTime! } # Query any Node by ID type Query { node(id: ID!): Node } ``` ### Unions (Polymorphic Results) ```graphql union SearchResult = User | Post | Comment type Query { search(query: String!): [SearchResult!]! } # Client query with type-specific fields query { search(query: "graphql") { ... on User { id name } ... on Post { id title author { name } } ... on Comment { id body post { title } } } } ``` ```typescript // Resolver must include __typename const resolvers = { SearchResult: { __resolveType(obj: any) { if (obj.email) return "User"; if (obj.title) return "Post"; if (obj.body) return "Comment"; return null; }, }, }; ``` ## Schema Stitching and Federation ### Apollo Federation Split schema across microservices: ```graphql # Users service type User @key(fields: "id") { id: ID! name: String! email: String! } type Query { user(id: ID!): User me: User } ``` ```graphql # Orders service - extends User from another service type User @key(fields: "id") { id: ID! orders: [Order!]! # Added by this service } type Order @key(fields: "id") { id: ID! total: Int! status: OrderStatus! user: User! } type Query { order(id: ID!): Order } ``` ```typescript // Orders service resolver const resolvers = { User: { // Reference resolver - how to fetch User stub __resolveReference(ref: { id: string }, context: Context) { // Only need to resolve fields this service owns return { id: ref.id }; }, orders(user: { id: string }, _, context: Context) { return context.dataSources.orders.findByUserId(user.id); }, }, }; ``` ### When to Federate | Use Federation | Don't Federate | |----------------|----------------| | Multiple teams own different domains | Single team, single service | | Independent deployment needed | Monolith or simple microservices | | Schema > 500 types | Schema < 100 types | | Different scaling requirements | Uniform load | ## Code-First vs Schema-First ### Schema-First (SDL) Write `.graphql` files, generate types: ```graphql # schema.graphql type Query { user(id: ID!): User } ``` ```typescript // Generated types (via graphql-codegen) export type QueryUserArgs = { id: string }; export type QueryResolvers = { user?: Resolver<Maybe<User>, {}, Context, QueryUserArgs>; }; ``` **Pros**: Schema is the contract, readable, tooling-friendly **Cons**: Types and schema can drift, boilerplate ### Code-First Write TypeScript/Go, generate schema: ```typescript // Using Pothos (TypeScript) const builder = new SchemaBuilder<{ Context: Context; Scalars: { DateTime: { Input: Date; Output: Date } }; }>({}); const UserType = builder.objectRef<User>("User").implement({ fields: (t) => ({ id: t.exposeID("id"), name: t.exposeString("name"), email: t.exposeString("email"), posts: t.field({ type: [PostType], resolve: (user, _, context) => context.loaders.postsByUserId.load(user.id), }), }), }); builder.queryField("user", (t) => t.field({ type: UserType, nullable: true, args: { id: t.arg.id({ required: true }) }, resolve: (_, { id }, context) => context.dataSources.users.findById(id), }) ); ``` **Pros**: Single source of truth, type-safe, refactor-friendly **Cons**: Schema less visible, framework lock-in ### Recommendation - **Schema-first**: Public APIs, multi-language teams, API-design-driven - **Code-first**: TypeScript backends, rapid iteration, small teams ## Performance ### Query Complexity Analysis ```typescript import { createComplexityLimitRule } from "graphql-validation-complexity"; const server = new ApolloServer({ validationRules: [ createComplexityLimitRule(1000, { scalarCost: 1, objectCost: 2, listFactor: 10, // Multiplier for list fields formatErrorMessage: (cost: number) => `Query too complex: cost ${cost} exceeds maximum 1000`, }), ], }); ``` ### Depth Limiting ```typescript import depthLimit from "graphql-depth-limit"; const server = new ApolloServer({ validationRules: [ depthLimit(7, { ignore: ["__schema"] }), // Max 7 levels deep ], }); ``` ### Persisted Queries Lock down which queries can execute (production hardening): ```typescript // Build step: extract queries from client code // queries.json { "abc123": "query GetUser($id: ID!) { user(id: $id) { id name email } }", "def456": "query ListUsers($first: Int) { users(first: $first) { edges { node { id name } } } }" } // Server: only allow registered queries const server = new ApolloServer({ persistedQueries: { cache: new InMemoryLRUCache(), }, // In production, reject non-persisted queries allowBatchedHttpRequests: false, }); ``` ### Automatic Persisted Queries (APQ) ``` # Client sends hash first (saves bandwidth) POST /graphql { "extensions": { "persistedQuery": { "version": 1, "sha256Hash": "abc123hash..." } }, "variables": { "id": "user-123" } } # Server: "I don't have that hash" { "errors": [{ "message": "PersistedQueryNotFound" }] } # Client retries with full query (cached for future) POST /graphql { "query": "query GetUser($id: ID!) { ... }", "extensions": { "persistedQuery": { "version": 1, "sha256Hash": "abc123hash..." } } } ``` ### Response Caching ```typescript // Field-level cache hints const resolvers = { Query: { user: (_, { id }, __, info) => { info.cacheControl.setCacheHint({ maxAge: 60, scope: "PRIVATE" }); return fetchUser(id); }, products: (_, __, ___, info) => { info.cacheControl.setCacheHint({ maxAge: 300, scope: "PUBLIC" }); return fetchProducts(); }, }, }; ``` ## TypeScript and GraphQL ### Code Generation (graphql-codegen) ```yaml # codegen.yml schema: "./schema/**/*.graphql" documents: "./src/**/*.{ts,tsx}" generates: ./src/generated/types.ts: plugins: - typescript - typescript-resolvers config: contextType: "../context#Context" mappers: User: "../models#UserModel" ./src/generated/operations.ts: plugins: - typescript - typescript-operations - typescript-react-apollo # For React hooks ``` ```bash npx graphql-codegen --watch ``` ### Typed Client (urql / Apollo) ```typescript // Auto-generated hook from codegen import { useGetUserQuery } from "./generated/operations"; function UserProfile({ id }: { id: string }) { const [{ data, fetching, error }] = useGetUserQuery({ variables: { id }, }); if (fetching) return <Loading />; if (error) return <Error error={error} />; // data.user is fully typed return <h1>{data.user.name}</h1>; } ``` ## When GraphQL Is Overkill ### Skip GraphQL When - Simple CRUD with 1-2 clients (REST is simpler) - File upload heavy (REST multipart is native) - Real-time only (WebSocket/SSE is more direct) - Team has no GraphQL experience and timeline is tight - Caching is critical (HTTP caching with REST is free) - Public API for third-party devs (REST has wider tooling) ### Use GraphQL When - Multiple clients need different data shapes (mobile, web, TV) - Deep, nested data with varied access patterns - Rapid frontend iteration (no backend changes for new views) - You have a federated microservice architecture - Over-fetching or under-fetching is a real measured problem - You can invest in proper tooling (codegen, DataLoader, complexity limits) ### GraphQL Anti-Patterns | Anti-Pattern | Problem | Fix | |--------------|---------|-----| | No DataLoader | N+1 queries tank performance | Always batch with DataLoader | | No depth/complexity limits | DoS via nested queries | Set limits before production | | Huge input types | Mutations become dump trucks | Split into focused mutations | | Business logic in resolvers | Untestable, duplicated | Thin resolvers, service layer | | No error codes | Clients parse error strings | Use `extensions.code` | | Schema-per-team with no coordination | Inconsistent naming, types | Schema governance / federation | | Exposing DB schema as GraphQL schema | Coupling, security risk | Design for the client, not the DB | -
grpc.md 18.2 KB
# gRPC Patterns ## Table of Contents - [Protocol Buffers (proto3)](#protocol-buffers-proto3) - [Service Definitions](#service-definitions) - [gRPC in Go](#grpc-in-go) - [gRPC in Rust](#grpc-in-rust) - [Interceptors and Middleware](#interceptors-and-middleware) - [Error Handling](#error-handling) - [Deadlines and Cancellation](#deadlines-and-cancellation) - [Health Checking](#health-checking) - [Reflection and CLI Tools](#reflection-and-cli-tools) - [gRPC-Web and Connect](#grpc-web-and-connect) - [When gRPC Beats REST](#when-grpc-beats-rest) --- ## Protocol Buffers (proto3) ### Basic Syntax ```protobuf syntax = "proto3"; package myapi.v1; option go_package = "github.com/myorg/myapi/gen/go/myapi/v1"; // Messages message User { string id = 1; string name = 2; string email = 3; UserRole role = 4; google.protobuf.Timestamp created_at = 5; optional string bio = 6; // Explicit optional (presence tracking) repeated string tags = 7; // List map<string, string> metadata = 8; // Key-value map } // Enums (always start with 0 = UNSPECIFIED) enum UserRole { USER_ROLE_UNSPECIFIED = 0; USER_ROLE_ADMIN = 1; USER_ROLE_MEMBER = 2; USER_ROLE_VIEWER = 3; } // Oneof (mutually exclusive fields) message Notification { string id = 1; oneof channel { EmailNotification email = 2; SmsNotification sms = 3; PushNotification push = 4; } } message EmailNotification { string subject = 1; string body = 2; } message SmsNotification { string phone = 1; string text = 2; } message PushNotification { string title = 1; string body = 2; } ``` ### Well-Known Types ```protobuf import "google/protobuf/timestamp.proto"; // Timestamp import "google/protobuf/duration.proto"; // Duration import "google/protobuf/empty.proto"; // Empty (no fields) import "google/protobuf/wrappers.proto"; // Nullable primitives import "google/protobuf/struct.proto"; // Dynamic JSON-like import "google/protobuf/field_mask.proto"; // Partial updates import "google/protobuf/any.proto"; // Type-erased message message UpdateUserRequest { string id = 1; User user = 2; google.protobuf.FieldMask update_mask = 3; // Which fields to update } ``` ### Proto Design Rules | Rule | Example | |------|---------| | Field numbers are forever | Never reuse a deleted field number | | Enums start at 0 = UNSPECIFIED | `USER_ROLE_UNSPECIFIED = 0` | | Use `optional` for presence | Distinguish "not set" from default value | | Prefix enum values with type name | `USER_ROLE_ADMIN` not `ADMIN` | | Package = `org.service.v1` | Enables API versioning | | Avoid `float`/`double` for money | Use `int64` cents or `string` | | Use FieldMask for partial updates | Explicit about which fields changed | | Reserved deleted fields | `reserved 5, 6; reserved "old_field";` | ## Service Definitions ### Four Communication Patterns ```protobuf service UserService { // Unary - simple request/response rpc GetUser(GetUserRequest) returns (GetUserResponse); // Server streaming - server sends multiple responses rpc ListUsers(ListUsersRequest) returns (stream User); // Client streaming - client sends multiple requests rpc UploadUserPhotos(stream UploadPhotoRequest) returns (UploadSummary); // Bidirectional streaming - both sides stream rpc Chat(stream ChatMessage) returns (stream ChatMessage); } message GetUserRequest { string id = 1; } message GetUserResponse { User user = 1; } message ListUsersRequest { int32 page_size = 1; string page_token = 2; string filter = 3; } ``` ### Request/Response Patterns ```protobuf // Pagination (AIP-158 style) message ListUsersRequest { int32 page_size = 1; // Max items per page string page_token = 2; // Opaque token from previous response } message ListUsersResponse { repeated User users = 1; string next_page_token = 2; // Empty = no more pages int32 total_size = 3; // Optional total count } // Batch operations message BatchGetUsersRequest { repeated string ids = 1; // Max 100 } message BatchGetUsersResponse { repeated User users = 1; } ``` ## gRPC in Go ### Server Implementation ```go package main import ( "context" "log" "net" "google.golang.org/grpc" "google.golang.org/grpc/codes" "google.golang.org/grpc/status" pb "github.com/myorg/myapi/gen/go/myapi/v1" ) type userServer struct { pb.UnimplementedUserServiceServer // Forward compatibility store UserStore } func (s *userServer) GetUser(ctx context.Context, req *pb.GetUserRequest) (*pb.GetUserResponse, error) { if req.GetId() == "" { return nil, status.Error(codes.InvalidArgument, "id is required") } user, err := s.store.Get(ctx, req.GetId()) if err != nil { if errors.Is(err, ErrNotFound) { return nil, status.Errorf(codes.NotFound, "user %s not found", req.GetId()) } return nil, status.Errorf(codes.Internal, "failed to get user: %v", err) } return &pb.GetUserResponse{User: user}, nil } // Server streaming func (s *userServer) ListUsers(req *pb.ListUsersRequest, stream pb.UserService_ListUsersServer) error { users, err := s.store.List(stream.Context(), req) if err != nil { return status.Errorf(codes.Internal, "failed to list users: %v", err) } for _, user := range users { if err := stream.Send(user); err != nil { return err } } return nil } func main() { lis, err := net.Listen("tcp", ":50051") if err != nil { log.Fatalf("failed to listen: %v", err) } server := grpc.NewServer( grpc.UnaryInterceptor(loggingInterceptor), grpc.ChainUnaryInterceptor(authInterceptor, loggingInterceptor), ) pb.RegisterUserServiceServer(server, &userServer{store: NewUserStore()}) log.Println("gRPC server listening on :50051") if err := server.Serve(lis); err != nil { log.Fatalf("failed to serve: %v", err) } } ``` ### Client Usage (Go) ```go func main() { conn, err := grpc.Dial("localhost:50051", grpc.WithTransportCredentials(insecure.NewCredentials()), grpc.WithUnaryInterceptor(retryInterceptor), ) if err != nil { log.Fatalf("failed to connect: %v", err) } defer conn.Close() client := pb.NewUserServiceClient(conn) // Unary call with deadline ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() resp, err := client.GetUser(ctx, &pb.GetUserRequest{Id: "user-123"}) if err != nil { st, ok := status.FromError(err) if ok { log.Printf("gRPC error: code=%s, message=%s", st.Code(), st.Message()) } return } log.Printf("User: %s", resp.GetUser().GetName()) // Server streaming stream, err := client.ListUsers(ctx, &pb.ListUsersRequest{PageSize: 100}) if err != nil { log.Fatal(err) } for { user, err := stream.Recv() if err == io.EOF { break } if err != nil { log.Fatal(err) } log.Printf("User: %s", user.GetName()) } } ``` ## gRPC in Rust ### Server with Tonic ```toml # Cargo.toml [dependencies] tonic = "0.12" prost = "0.13" tokio = { version = "1", features = ["full"] } [build-dependencies] tonic-build = "0.12" ``` ```rust // build.rs fn main() -> Result<(), Box<dyn std::error::Error>> { tonic_build::compile_protos("proto/myapi/v1/user.proto")?; Ok(()) } ``` ```rust use tonic::{Request, Response, Status}; pub mod myapi { pub mod v1 { tonic::include_proto!("myapi.v1"); } } use myapi::v1::user_service_server::{UserService, UserServiceServer}; use myapi::v1::{GetUserRequest, GetUserResponse, User}; #[derive(Default)] pub struct MyUserService; #[tonic::async_trait] impl UserService for MyUserService { async fn get_user( &self, request: Request<GetUserRequest>, ) -> Result<Response<GetUserResponse>, Status> { let req = request.into_inner(); if req.id.is_empty() { return Err(Status::invalid_argument("id is required")); } // Fetch user from store... let user = User { id: req.id, name: "Alice".into(), email: "alice@example.com".into(), ..Default::default() }; Ok(Response::new(GetUserResponse { user: Some(user) })) } } #[tokio::main] async fn main() -> Result<(), Box<dyn std::error::Error>> { let addr = "[::1]:50051".parse()?; let service = MyUserService::default(); tonic::transport::Server::builder() .add_service(UserServiceServer::new(service)) .serve(addr) .await?; Ok(()) } ``` ### Client with Tonic ```rust use myapi::v1::user_service_client::UserServiceClient; use myapi::v1::GetUserRequest; #[tokio::main] async fn main() -> Result<(), Box<dyn std::error::Error>> { let mut client = UserServiceClient::connect("http://[::1]:50051").await?; let request = tonic::Request::new(GetUserRequest { id: "user-123".into(), }); let response = client.get_user(request).await?; println!("User: {:?}", response.into_inner().user); Ok(()) } ``` ## Interceptors and Middleware ### Go Unary Interceptor ```go func loggingInterceptor( ctx context.Context, req interface{}, info *grpc.UnaryServerInfo, handler grpc.UnaryHandler, ) (interface{}, error) { start := time.Now() // Extract metadata md, _ := metadata.FromIncomingContext(ctx) requestID := md.Get("x-request-id") resp, err := handler(ctx, req) st, _ := status.FromError(err) log.Printf("method=%s duration=%s status=%s request_id=%v", info.FullMethod, time.Since(start), st.Code(), requestID) return resp, err } func authInterceptor( ctx context.Context, req interface{}, info *grpc.UnaryServerInfo, handler grpc.UnaryHandler, ) (interface{}, error) { md, ok := metadata.FromIncomingContext(ctx) if !ok { return nil, status.Error(codes.Unauthenticated, "no metadata") } tokens := md.Get("authorization") if len(tokens) == 0 { return nil, status.Error(codes.Unauthenticated, "no token") } claims, err := validateToken(tokens[0]) if err != nil { return nil, status.Error(codes.Unauthenticated, "invalid token") } // Add claims to context ctx = context.WithValue(ctx, claimsKey, claims) return handler(ctx, req) } ``` ### Chaining Interceptors ```go server := grpc.NewServer( grpc.ChainUnaryInterceptor( recoveryInterceptor, // Panic recovery (outermost) loggingInterceptor, // Request logging metricsInterceptor, // Prometheus metrics authInterceptor, // Authentication validationInterceptor, // Request validation ), grpc.ChainStreamInterceptor( streamLoggingInterceptor, streamAuthInterceptor, ), ) ``` ## Error Handling ### gRPC Status Codes | Code | Name | Use When | |------|------|----------| | 0 | OK | Success | | 1 | CANCELLED | Client cancelled | | 2 | UNKNOWN | Unknown error (avoid - be specific) | | 3 | INVALID_ARGUMENT | Bad request (validation) | | 4 | DEADLINE_EXCEEDED | Timeout | | 5 | NOT_FOUND | Resource doesn't exist | | 6 | ALREADY_EXISTS | Conflict (duplicate) | | 7 | PERMISSION_DENIED | Authorized but not allowed | | 8 | RESOURCE_EXHAUSTED | Rate limit, quota | | 9 | FAILED_PRECONDITION | State not ready (e.g., non-empty directory) | | 10 | ABORTED | Concurrency conflict (retry) | | 11 | OUT_OF_RANGE | Seek past end | | 12 | UNIMPLEMENTED | Method not implemented | | 13 | INTERNAL | Internal server error | | 14 | UNAVAILABLE | Service down (retry with backoff) | | 16 | UNAUTHENTICATED | No valid credentials | ### Rich Error Details (Go) ```go import ( "google.golang.org/genproto/googleapis/rpc/errdetails" "google.golang.org/grpc/status" ) func (s *server) CreateUser(ctx context.Context, req *pb.CreateUserRequest) (*pb.CreateUserResponse, error) { // Validation with rich error details var violations []*errdetails.BadRequest_FieldViolation if req.GetEmail() == "" { violations = append(violations, &errdetails.BadRequest_FieldViolation{ Field: "email", Description: "Email is required", }) } if len(req.GetName()) < 2 { violations = append(violations, &errdetails.BadRequest_FieldViolation{ Field: "name", Description: "Name must be at least 2 characters", }) } if len(violations) > 0 { st := status.New(codes.InvalidArgument, "validation failed") br := &errdetails.BadRequest{FieldViolations: violations} st, _ = st.WithDetails(br) return nil, st.Err() } // ... proceed } ``` ### Mapping gRPC to HTTP Status Codes | gRPC Code | HTTP Status | |-----------|-------------| | OK | 200 | | INVALID_ARGUMENT | 400 | | UNAUTHENTICATED | 401 | | PERMISSION_DENIED | 403 | | NOT_FOUND | 404 | | ALREADY_EXISTS | 409 | | RESOURCE_EXHAUSTED | 429 | | CANCELLED | 499 | | INTERNAL | 500 | | UNIMPLEMENTED | 501 | | UNAVAILABLE | 503 | | DEADLINE_EXCEEDED | 504 | ## Deadlines and Cancellation ### Setting Deadlines (Go Client) ```go // Always set deadlines - never leave RPCs unbounded ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) defer cancel() resp, err := client.GetUser(ctx, &pb.GetUserRequest{Id: "user-123"}) if err != nil { st, _ := status.FromError(err) if st.Code() == codes.DeadlineExceeded { // Handle timeout - maybe retry with longer deadline } } ``` ### Propagating Deadlines Deadlines automatically propagate through the call chain. If service A calls service B with a 5s deadline, and A takes 2s, B gets the remaining 3s. ```go // Server-side: check remaining time deadline, ok := ctx.Deadline() if ok { remaining := time.Until(deadline) if remaining < 100*time.Millisecond { return nil, status.Error(codes.DeadlineExceeded, "insufficient time remaining") } } ``` ## Health Checking ### Standard Health Protocol ```protobuf // Built-in: grpc.health.v1.Health service Health { rpc Check(HealthCheckRequest) returns (HealthCheckResponse); rpc Watch(HealthCheckRequest) returns (stream HealthCheckResponse); } message HealthCheckRequest { string service = 1; // Empty = overall health } message HealthCheckResponse { enum ServingStatus { UNKNOWN = 0; SERVING = 1; NOT_SERVING = 2; SERVICE_UNKNOWN = 3; } ServingStatus status = 1; } ``` ### Go Implementation ```go import "google.golang.org/grpc/health" import healthpb "google.golang.org/grpc/health/grpc_health_v1" server := grpc.NewServer() healthServer := health.NewServer() healthpb.RegisterHealthServer(server, healthServer) // Set status healthServer.SetServingStatus("myapi.v1.UserService", healthpb.HealthCheckResponse_SERVING) // Kubernetes uses grpc_health_probe // livenessProbe: // exec: // command: ["/bin/grpc_health_probe", "-addr=:50051"] ``` ## Reflection and CLI Tools ### Enable Reflection ```go import "google.golang.org/grpc/reflection" server := grpc.NewServer() reflection.Register(server) // Enable for dev/staging ``` ### grpcurl (like curl for gRPC) ```bash # List services grpcurl -plaintext localhost:50051 list # Describe a service grpcurl -plaintext localhost:50051 describe myapi.v1.UserService # Call a method grpcurl -plaintext -d '{"id": "user-123"}' \ localhost:50051 myapi.v1.UserService/GetUser # Server streaming grpcurl -plaintext -d '{"page_size": 10}' \ localhost:50051 myapi.v1.UserService/ListUsers # With metadata (headers) grpcurl -plaintext \ -H 'authorization: Bearer token123' \ -d '{"id": "user-123"}' \ localhost:50051 myapi.v1.UserService/GetUser ``` ### buf (Modern Protobuf Tooling) ```bash # Lint proto files buf lint # Detect breaking changes buf breaking --against '.git#branch=main' # Generate code buf generate # buf.yaml version: v2 lint: use: - STANDARD breaking: use: - WIRE_JSON ``` ## gRPC-Web and Connect ### The Browser Problem Browsers cannot use gRPC natively (no HTTP/2 trailers, no bidirectional streaming). Solutions: | Solution | Approach | Streaming | Ecosystem | |----------|----------|-----------|-----------| | gRPC-Web | Proxy (Envoy) translates | Server-streaming only | Google official | | Connect | Native HTTP/1.1 + HTTP/2 | All patterns via HTTP/2 | Buf (connectrpc.com) | | gRPC-Gateway | Generate REST from proto | None (REST) | grpc-ecosystem | ### Connect (Recommended for New Projects) ```protobuf // Same .proto files - no changes needed service UserService { rpc GetUser(GetUserRequest) returns (GetUserResponse); } ``` ```typescript // TypeScript client (works in browser natively) import { createClient } from "@connectrpc/connect"; import { createConnectTransport } from "@connectrpc/connect-web"; import { UserService } from "./gen/myapi/v1/user_connect"; const transport = createConnectTransport({ baseUrl: "https://api.example.com", }); const client = createClient(UserService, transport); const response = await client.getUser({ id: "user-123" }); console.log(response.user?.name); ``` Connect supports three protocols simultaneously: - **Connect protocol**: Simple HTTP POST with JSON or Protobuf - **gRPC protocol**: Standard gRPC (HTTP/2) - **gRPC-Web protocol**: Browser-compatible gRPC ## When gRPC Beats REST ### Use gRPC When - Internal service-to-service communication - Performance matters (10x smaller payloads, 7x faster serialization) - You need streaming (logs, real-time feeds, file uploads) - You want a strict contract between services - Polyglot environment (generate clients for any language) - Bidirectional communication ### Use REST When - Public API consumed by third-party developers - Browser clients are primary (unless using Connect) - You need HTTP caching (CDN, browser cache) - Team is more familiar with REST - Simple CRUD with few relationships - Webhooks are a primary integration pattern ### Hybrid Approach Many production systems use both: - gRPC for internal microservice communication - REST/GraphQL for external-facing APIs - gRPC-Gateway or Connect to expose gRPC services as REST ``` [Browser] --REST/GraphQL--> [API Gateway] --gRPC--> [User Service] --gRPC--> [Order Service] --gRPC--> [Payment Service] ``` -
rest-advanced.md 13.6 KB
# REST Advanced Patterns ## Table of Contents - [Resource Modeling](#resource-modeling) - [HTTP Methods Beyond CRUD](#http-methods-beyond-crud) - [Content Negotiation](#content-negotiation) - [Pagination Implementations](#pagination-implementations) - [Filtering, Sorting, Field Selection](#filtering-sorting-field-selection) - [Bulk Operations](#bulk-operations) - [Long-Running Operations](#long-running-operations) - [HATEOAS and Hypermedia](#hateoas-and-hypermedia) - [API Documentation](#api-documentation) - [Webhook Design](#webhook-design) - [Caching](#caching) --- ## Resource Modeling ### Collections vs Singletons ``` /users # Collection - supports GET (list), POST (create) /users/{id} # Singleton - supports GET, PUT, PATCH, DELETE /users/{id}/profile # Singleton sub-resource (1:1 relationship) /users/{id}/orders # Sub-collection (1:many relationship) ``` ### Modeling Relationships **Approach 1: Sub-resources (strong ownership)** ``` GET /users/{id}/orders # Orders belong to user POST /users/{id}/orders # Create order for user ``` **Approach 2: Top-level with filters (independent entities)** ``` GET /orders?user_id={id} # Orders exist independently GET /orders/{order_id} # Direct access without user context ``` **Approach 3: Relationship endpoints (many-to-many)** ``` GET /users/{id}/roles # List user's roles PUT /users/{id}/roles/{rid} # Assign role (no body needed) DELETE /users/{id}/roles/{rid} # Remove role ``` ### When to Use Sub-Resources | Use sub-resource | Use top-level | |------------------|---------------| | Child can't exist without parent | Entity is independently meaningful | | Always accessed in parent context | Frequently queried across parents | | Moderate cardinality (< 1000) | High cardinality | | Lifecycle tied to parent | Independent lifecycle | ### Resource Naming Patterns ``` # Actions that don't map to CRUD - use sub-resources POST /orders/{id}/cancel # State transition POST /users/{id}/verify-email # Trigger action POST /reports/{id}/export # Async operation # Avoid: verbs as top-level resources POST /cancelOrder # Bad POST /send-notification # Bad # Search as a resource (when GET query string is too complex) POST /users/search { "filters": { "age_range": [18, 30], "location": { "within": "10km", "of": [lat, lng] } } } ``` ## HTTP Methods Beyond CRUD ### PATCH Strategies **JSON Merge Patch (RFC 7396)** - Simple, intuitive: ``` PATCH /users/123 Content-Type: application/merge-patch+json { "name": "New Name", "address": null } ``` - Set `name` to "New Name" - Remove `address` (null = delete) - Leave all other fields unchanged - Limitation: cannot set a field TO null vs removing it **JSON Patch (RFC 6902)** - Precise operations: ``` PATCH /users/123 Content-Type: application/json-patch+json [ { "op": "replace", "path": "/name", "value": "New Name" }, { "op": "remove", "path": "/address" }, { "op": "add", "path": "/tags/-", "value": "premium" }, { "op": "test", "path": "/version", "value": 5 } ] ``` - Supports: add, remove, replace, move, copy, test - `test` enables optimistic concurrency (apply only if value matches) - More complex but unambiguous **Recommendation**: Use JSON Merge Patch for most APIs (simpler). Use JSON Patch when you need array manipulation or atomic test-and-set. ### HEAD and OPTIONS ``` # HEAD - metadata without body (same headers as GET) HEAD /files/report.pdf # Returns: Content-Length, Content-Type, Last-Modified, ETag # Use: check existence, get size before download # OPTIONS - discover allowed methods (CORS preflight uses this) OPTIONS /users # Returns: Allow: GET, POST, HEAD, OPTIONS ``` ## Content Negotiation ### Accept Header ``` # Client requests specific format GET /users/123 Accept: application/json # JSON (default) Accept: application/xml # XML Accept: text/csv # CSV export Accept: application/pdf # PDF report # Versioning via media type Accept: application/vnd.myapi.v2+json # Version in media type ``` ### Content-Type on Requests ``` # Server must validate Content-Type on mutations POST /users Content-Type: application/json # Standard Content-Type: multipart/form-data # File uploads Content-Type: application/x-www-form-urlencoded # Form data ``` ### Implementation (Go) ```go func handleGetUser(w http.ResponseWriter, r *http.Request) { accept := r.Header.Get("Accept") user := fetchUser(r) switch { case strings.Contains(accept, "application/xml"): w.Header().Set("Content-Type", "application/xml") xml.NewEncoder(w).Encode(user) case strings.Contains(accept, "text/csv"): w.Header().Set("Content-Type", "text/csv") writeCSV(w, user) default: w.Header().Set("Content-Type", "application/json") json.NewEncoder(w).Encode(user) } } ``` ## Pagination Implementations ### Cursor-Based (Recommended for Most Cases) Encode the cursor as base64 for opacity: ```go // Encode cursor type Cursor struct { ID int64 `json:"id"` CreatedAt time.Time `json:"created_at"` } func encodeCursor(c Cursor) string { b, _ := json.Marshal(c) return base64.URLEncoding.EncodeToString(b) } func decodeCursor(s string) (Cursor, error) { b, err := base64.URLEncoding.DecodeString(s) if err != nil { return Cursor{}, err } var c Cursor return c, json.Unmarshal(b, &c) } ``` **Request/Response:** ``` GET /users?limit=20&after=eyJpZCI6MTIzLCJjcmVhdGVkX2F0IjoiMjAyNC0wMS0xNVQxMDozMDowMFoifQ== { "data": [...], "pagination": { "has_more": true, "next_cursor": "eyJpZCI6MTQzLCJjcmVhdGVkX2F0IjoiMjAyNC0wMS0xNlQwODoxNTowMFoifQ==", "prev_cursor": "eyJpZCI6MTI0LCJjcmVhdGVkX2F0IjoiMjAyNC0wMS0xNVQxMTowMDowMFoifQ==" } } ``` **SQL (keyset pagination under the hood):** ```sql SELECT * FROM users WHERE (created_at, id) > ('2024-01-15T10:30:00Z', 123) ORDER BY created_at ASC, id ASC LIMIT 21; -- fetch limit+1 to determine has_more ``` ### Link Headers (RFC 8288) ``` Link: <https://api.example.com/users?after=abc123&limit=20>; rel="next", <https://api.example.com/users?before=xyz789&limit=20>; rel="prev", <https://api.example.com/users?limit=20>; rel="first" ``` ### Total Count Considerations - `total` count requires a separate `COUNT(*)` query - expensive on large tables - Make it opt-in: `GET /users?limit=20&include_total=true` - Consider approximate counts: `SELECT reltuples FROM pg_class WHERE relname = 'users'` ## Filtering, Sorting, Field Selection ### Filtering ``` # Simple equality GET /users?status=active&role=admin # Operators (LHS brackets style - used by Stripe, Supabase) GET /users?created_at[gte]=2024-01-01&created_at[lt]=2024-02-01 GET /products?price[lte]=100&category[in]=electronics,books # Operators (filter syntax) GET /users?filter=status eq "active" and age gt 18 ``` ### Sorting ``` # Simple (comma-separated, prefix - for descending) GET /users?sort=-created_at,name # Multiple fields GET /products?sort=category,-price # category ASC, then price DESC ``` ### Sparse Fieldsets ``` # Return only specific fields (reduces payload) GET /users?fields=id,name,email GET /users/123?fields=id,name,email,profile.avatar # Related resource fields GET /orders?fields=id,total&fields[customer]=id,name ``` ## Bulk Operations ### Batch Create ``` POST /users/batch Content-Type: application/json { "items": [ { "name": "Alice", "email": "alice@example.com" }, { "name": "Bob", "email": "bob@example.com" } ] } # Response: 207 Multi-Status { "results": [ { "status": 201, "data": { "id": "u1", "name": "Alice" } }, { "status": 409, "error": { "type": "conflict", "detail": "Email already exists" } } ], "summary": { "succeeded": 1, "failed": 1 } } ``` ### Batch Actions ``` POST /users/batch-action { "action": "deactivate", "ids": ["u1", "u2", "u3"], "reason": "Account cleanup" } ``` ### Guidelines - Set a maximum batch size (100-1000 items) - Return 207 Multi-Status for partial success - Include per-item status in response - Consider async processing for large batches (return 202 + job URL) ## Long-Running Operations ### Polling Pattern ``` # Start operation POST /reports/generate { "type": "annual", "year": 2024 } # Response: 202 Accepted { "operation_id": "op-abc-123", "status": "pending", "status_url": "/operations/op-abc-123", "estimated_completion": "2024-01-15T10:35:00Z" } # Poll for status GET /operations/op-abc-123 { "operation_id": "op-abc-123", "status": "completed", # pending | running | completed | failed "progress": 100, "result_url": "/reports/rpt-xyz-789", "completed_at": "2024-01-15T10:34:12Z" } ``` ### Webhook Callback ``` POST /reports/generate { "type": "annual", "year": 2024, "callback_url": "https://myapp.com/webhooks/report-ready" } # Server POSTs to callback_url when done: { "event": "report.completed", "operation_id": "op-abc-123", "result_url": "/reports/rpt-xyz-789" } ``` ### Server-Sent Events ``` GET /operations/op-abc-123/stream Accept: text/event-stream event: progress data: {"percent": 25, "stage": "fetching data"} event: progress data: {"percent": 75, "stage": "generating charts"} event: complete data: {"result_url": "/reports/rpt-xyz-789"} ``` ## HATEOAS and Hypermedia ### When It's Worth the Complexity | Worth it | Not worth it | |----------|--------------| | Public API with many consumers | Internal microservice | | API that evolves frequently | Stable, versioned API | | Workflow-driven (state machines) | Simple CRUD | | Discoverability is a feature | Clients are tightly coupled | ### HAL (Hypertext Application Language) ```json { "id": "order-123", "status": "pending_payment", "total": 5999, "_links": { "self": { "href": "/orders/order-123" }, "pay": { "href": "/orders/order-123/pay", "method": "POST" }, "cancel": { "href": "/orders/order-123", "method": "DELETE" } }, "_embedded": { "items": [ { "product_id": "prod-456", "quantity": 2, "_links": { "product": { "href": "/products/prod-456" } } } ] } } ``` ## API Documentation ### OpenAPI 3.1 Structure ```yaml openapi: 3.1.0 info: title: My API version: 2.0.0 description: | ## Authentication All endpoints require Bearer token authentication. contact: email: api-support@example.com servers: - url: https://api.example.com/v2 description: Production - url: https://sandbox.example.com/v2 description: Sandbox paths: /users: get: summary: List users operationId: listUsers tags: [Users] parameters: - name: limit in: query schema: type: integer default: 20 maximum: 100 responses: '200': description: Success content: application/json: schema: $ref: '#/components/schemas/UserList' ``` ### Documentation Tools | Tool | Strength | |------|----------| | Redoc | Beautiful single-page docs from OpenAPI | | Swagger UI | Interactive "try it" playground | | Stoplight | Design-first with mock servers | | Mintlify | Modern docs with guides + API reference | ## Webhook Design ### Webhook Payload ```json { "id": "evt_abc123", "type": "order.completed", "created_at": "2024-01-15T10:30:00Z", "api_version": "2024-01-15", "data": { "id": "order-456", "status": "completed", "total": 5999 } } ``` ### Signature Verification ``` # Header X-Webhook-Signature: sha256=5257a869e7ecebeda32affa62cdca3fa51cad7e77a0e56ff536d0ce8e108d8bd # Compute: HMAC-SHA256(webhook_secret, raw_body) ``` ```go func verifyWebhookSignature(secret, signature string, body []byte) bool { mac := hmac.New(sha256.New, []byte(secret)) mac.Write(body) expected := "sha256=" + hex.EncodeToString(mac.Sum(nil)) return hmac.Equal([]byte(expected), []byte(signature)) } ``` ### Webhook Best Practices | Practice | Detail | |----------|--------| | Retry with backoff | 1s, 5s, 30s, 5m, 30m, 2h, 24h | | Idempotency | Include event ID, consumers must deduplicate | | Timeout | 30 second max wait for 2xx response | | Disable after failures | Disable after N consecutive failures, notify owner | | Event log | Provide UI/API to replay failed webhooks | | Thin payloads | Send IDs + event type, let consumer fetch full data | ## Caching ### ETag-Based (Strong Validation) ``` # First request GET /users/123 ETag: "a1b2c3d4" # Subsequent request GET /users/123 If-None-Match: "a1b2c3d4" # Response if unchanged: 304 Not Modified (no body) # Response if changed: 200 with new ETag ``` ### Last-Modified (Weak Validation) ``` GET /users/123 Last-Modified: Thu, 15 Jan 2024 10:30:00 GMT # Subsequent request GET /users/123 If-Modified-Since: Thu, 15 Jan 2024 10:30:00 GMT ``` ### Cache-Control Directives ``` # Public, cacheable for 1 hour Cache-Control: public, max-age=3600 # Private (user-specific), cacheable for 5 minutes Cache-Control: private, max-age=300 # No caching (real-time data) Cache-Control: no-store # Revalidate before using cache Cache-Control: no-cache # Stale-while-revalidate (serve stale, refresh in background) Cache-Control: public, max-age=60, stale-while-revalidate=300 ``` ### Caching Strategy by Resource Type | Resource Type | Strategy | Cache-Control | |---------------|----------|---------------| | Static assets | Immutable with hash | `public, max-age=31536000, immutable` | | User profile | Short-lived, private | `private, max-age=60` | | Product catalog | Medium, public | `public, max-age=300, stale-while-revalidate=600` | | Search results | No cache or very short | `no-store` or `max-age=10` | | Real-time data | No cache | `no-store` |
-
-
scripts
-
.gitkeep 0 B · in bundle
-
-
SKILL.md 11.3 KB
--- name: api-design-ops description: "API design patterns for REST, gRPC, and GraphQL. Use for: api design, REST, gRPC, GraphQL, protobuf, schema design, api versioning, pagination, rate limiting, error format, OpenAPI, API authentication, JWT, OAuth2, API gateway, webhook, idempotency." when_to_use: "Use when designing an API's contract and shape across REST, gRPC, or GraphQL — e.g. 'how should I version this API', 'design pagination and error envelopes (RFC 7807)', 'GraphQL vs gRPC tradeoffs', 'add idempotency keys to POST'. Reach for rest-ops to implement a specific REST endpoint, auth-ops for login/token flows." license: MIT allowed-tools: "Read Write Bash" metadata: author: claude-mods related-skills: rest-ops, security-ops, go-ops, rust-ops, typescript-ops --- # API Design Ops Comprehensive API design patterns covering REST (advanced), gRPC, and GraphQL. This skill provides decision frameworks, design patterns, and implementation guidance for building production APIs. ## API Style Decision Tree ``` What kind of API do you need? | +-- Internal microservice-to-microservice? | +-- High throughput, low latency needed? --> gRPC | +-- Streaming (real-time data, logs)? --> gRPC (bidirectional streaming) | +-- Simple request/response, team comfort? --> REST | +-- Public-facing API? | +-- Third-party developers consuming it? --> REST (widest compatibility) | +-- Mobile app with varied data needs? --> GraphQL | +-- Browser-only, simple CRUD? --> REST | +-- Frontend for your own app? | +-- Multiple clients with different data shapes? --> GraphQL | +-- Single client, straightforward data? --> REST | +-- Real-time updates needed? --> GraphQL subscriptions or SSE | +-- IoT / embedded / constrained devices? | +-- Binary efficiency matters? --> gRPC | +-- HTTP-only environments? --> REST ``` ### Quick Comparison | Concern | REST | gRPC | GraphQL | |---------|------|------|---------| | Transport | HTTP/1.1+ | HTTP/2 | HTTP (any) | | Serialization | JSON (text) | Protobuf (binary) | JSON (text) | | Schema | OpenAPI (optional) | .proto (required) | SDL (required) | | Browser support | Native | Via gRPC-Web/Connect | Native | | Caching | HTTP caching built-in | Custom | Custom (normalized) | | Learning curve | Low | Medium | Medium-High | | Code generation | Optional | Required | Optional but recommended | | Streaming | SSE, WebSocket | Native (4 patterns) | Subscriptions | | Over-fetching | Common problem | No (typed) | Solved by design | | File uploads | Multipart native | Chunked streaming | Multipart spec (awkward) | ## REST Resource Design Quick Reference ### Resource Naming ``` GET /users # Collection GET /users/{id} # Singleton GET /users/{id}/orders # Sub-collection POST /users # Create PUT /users/{id} # Full replace PATCH /users/{id} # Partial update DELETE /users/{id} # Remove # Naming rules: # - Plural nouns for collections: /users NOT /user # - Kebab-case for multi-word: /line-items NOT /lineItems # - No verbs in URLs: POST /orders NOT POST /create-order # - Max 3 levels deep: /users/{id}/orders (not /users/{id}/orders/{oid}/items/{iid}/details) ``` ### HTTP Methods and Status Codes | Method | Success | Empty | Invalid | Not Found | Conflict | |--------|---------|-------|---------|-----------|----------| | GET | 200 | 200 (empty array) | 400 | 404 | - | | POST | 201 + Location | - | 400/422 | - | 409 | | PUT | 200 | - | 400/422 | 404 | 409 | | PATCH | 200 | - | 400/422 | 404 | 409 | | DELETE | 204 | 204 (already gone) | 400 | 404 | 409 | ### HATEOAS (When Worth It) Use when: public APIs where discoverability matters, long-lived APIs, APIs that evolve frequently. Skip when: internal microservices, mobile backends, tight coupling is acceptable. ```json { "id": "order-123", "status": "shipped", "_links": { "self": { "href": "/orders/order-123" }, "track": { "href": "/orders/order-123/tracking" }, "cancel": { "href": "/orders/order-123", "method": "DELETE" } } } ``` ## Pagination Decision Tree ``` What's your data like? | +-- Stable data, UI needs "jump to page 5"? | --> Offset pagination: ?page=5&per_page=20 | Tradeoff: Slow on large offsets (OFFSET 10000), inconsistent with inserts | +-- Large dataset, forward-only traversal? | --> Cursor pagination: ?after=eyJpZCI6MTIzfQ&limit=20 | Tradeoff: No random page access, but consistent and fast | +-- Real-time feed, ordered by timestamp or ID? | --> Keyset pagination: ?created_after=2024-01-01T00:00:00Z&limit=20 | Tradeoff: Requires a unique, sequential column; no page jumping ``` ### Response Envelope ```json { "data": [...], "pagination": { "total": 1432, "limit": 20, "has_more": true, "next_cursor": "eyJpZCI6MTQzMn0=" } } ``` ## Error Response Format (RFC 7807) All APIs should use Problem Details (RFC 7807 / RFC 9457): ```json { "type": "https://api.example.com/errors/insufficient-funds", "title": "Insufficient Funds", "status": 422, "detail": "Account xxxx-1234 has a balance of $10.00, but the transfer requires $25.00.", "instance": "/transfers/txn-abc-123", "balance": 1000, "required": 2500 } ``` ### Field Reference | Field | Required | Description | |-------|----------|-------------| | `type` | Yes | URI identifying the error type (stable, documentable) | | `title` | Yes | Human-readable summary (same for all instances of this type) | | `status` | Yes | HTTP status code | | `detail` | Yes | Human-readable explanation specific to this occurrence | | `instance` | No | URI identifying the specific occurrence | | (extensions) | No | Additional machine-readable fields | ### Validation Errors ```json { "type": "https://api.example.com/errors/validation", "title": "Validation Failed", "status": 422, "detail": "The request body contains 2 validation errors.", "errors": [ { "field": "email", "message": "Must be a valid email address", "code": "invalid_format" }, { "field": "age", "message": "Must be at least 18", "code": "out_of_range", "min": 18 } ] } ``` ## Versioning Strategies | Strategy | Example | Pros | Cons | |----------|---------|------|------| | URL path | `/v2/users` | Obvious, cacheable, easy routing | URL pollution, hard to sunset | | Accept header | `Accept: application/vnd.api.v2+json` | Clean URLs, content negotiation | Hidden, harder to test | | Query param | `/users?version=2` | Easy to add | Pollutes query string, caching issues | | Date-based | `API-Version: 2024-01-15` | Granular evolution (Stripe style) | Complex implementation | ### Recommendation - **Public APIs**: URL path versioning (`/v1/`) - simplicity wins - **Internal APIs**: Header or no versioning (deploy in lockstep) - **Evolving APIs**: Date-based (Stripe model) if you have the engineering investment ### Breaking Change Rules A breaking change is anything that can cause existing clients to fail: - Removing a field from a response - Renaming a field - Changing a field's type - Adding a required field to a request - Changing URL structure - Changing error formats - Removing an endpoint Non-breaking (safe): - Adding optional fields to requests - Adding fields to responses - Adding new endpoints - Adding new enum values (if client handles unknown values) ## Rate Limiting Design ### Algorithms | Algorithm | Behavior | Use When | |-----------|----------|----------| | Token bucket | Allows bursts, refills at steady rate | General API rate limiting | | Sliding window | Smooth distribution, no burst | Strict fairness needed | | Fixed window | Simple, potential burst at boundary | Low-stakes limiting | | Leaky bucket | Constant output rate | Queue processing | ### Response Headers ``` X-RateLimit-Limit: 1000 # Max requests per window X-RateLimit-Remaining: 743 # Requests left in current window X-RateLimit-Reset: 1672531200 # Unix timestamp when window resets Retry-After: 30 # Seconds to wait (on 429) ``` ### 429 Response Body ```json { "type": "https://api.example.com/errors/rate-limit-exceeded", "title": "Rate Limit Exceeded", "status": 429, "detail": "You have exceeded 1000 requests per hour. Try again in 30 seconds.", "retry_after": 30 } ``` ## Idempotency ### Which Methods Need Idempotency Keys? | Method | Idempotent by spec? | Needs key? | |--------|---------------------|------------| | GET | Yes | No | | PUT | Yes | No (full replacement is naturally idempotent) | | DELETE | Yes | No | | PATCH | No | Recommended for critical operations | | POST | No | **Yes** (always for payments, orders, transfers) | ### Implementation ``` POST /payments Idempotency-Key: 550e8400-e29b-41d4-a716-446655440000 Content-Type: application/json { "amount": 2500, "currency": "usd", "customer": "cust_123" } ``` Server-side: 1. Receive request with `Idempotency-Key` header 2. Check if key exists in store (Redis, DB) 3. If exists: return stored response (same status code + body) 4. If not: process request, store response keyed by idempotency key 5. Keys expire after 24-48 hours ## Authentication Overview | Method | Use When | Security Level | |--------|----------|----------------| | API Key | Server-to-server, internal, simple | Low-Medium | | JWT (Bearer) | Stateless auth, microservices | Medium-High | | OAuth2 + PKCE | Third-party access, user delegation | High | | mTLS | Service mesh, zero-trust infra | Very High | ### Decision Guide ``` Who is authenticating? | +-- Your own frontend? --> JWT (short-lived access + refresh token) +-- Third-party developer? --> OAuth2 (client credentials for server, PKCE for SPA) +-- Another internal service? --> mTLS or JWT with service accounts +-- Quick prototype? --> API key (but plan migration) ``` ## Gotchas Table | Gotcha | Problem | Prevention | |--------|---------|------------| | Breaking changes in "non-breaking" release | Client crashes | Additive-only policy, contract tests | | N+1 in REST APIs | 100 users = 101 queries | Compound documents, `?include=`, or GraphQL | | Over-fetching | Mobile gets 50 fields, needs 3 | Sparse fieldsets `?fields=id,name` or GraphQL | | Under-fetching | 3 requests to build one view | Composite endpoints or BFF pattern | | CORS misconfiguration | Frontend can't reach API | Explicit allowed origins, never `*` with credentials | | Missing Content-Type | 415 or silent parsing failure | Validate Content-Type on every mutation endpoint | | Large payloads without pagination | OOM, timeouts | Always paginate collections, set max page size | | Inconsistent date formats | Parsing hell | ISO 8601 everywhere: `2024-01-15T10:30:00Z` | | No request IDs | Impossible to debug | Generate `X-Request-ID`, propagate through services | | Enum evolution | New value breaks old client | Document that enums may grow, clients must handle unknown | | Missing idempotency | Duplicate charges, orders | Idempotency keys on all POST endpoints with side effects | | Unbounded query complexity | GraphQL DoS | Depth limiting, cost analysis, persisted queries | ## Reference Files | File | Contents | |------|----------| | `references/rest-advanced.md` | Resource modeling, PATCH strategies, caching, webhooks, bulk ops | | `references/grpc.md` | Protobuf, service definitions, Go/Rust, streaming, error handling | | `references/graphql.md` | Schema design, resolvers, DataLoader, federation, performance | | `references/api-security.md` | JWT, OAuth2, CORS, rate limiting, OWASP API Top 10 |
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.