api-caching-strategies
Application-level caching strategies, HTTP caching, cache invalidation, and stampede prevention
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/api-caching-strategies/skills/api-caching-strategies
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Caching Strategies
Quick Guide: Choose the right caching strategy for your use case: cache-aside for read-heavy data, write-through for consistency, write-behind for write-heavy workloads. Always set TTL to prevent stale data and memory exhaustion. Use HTTP caching headers (Cache-Control, ETag, Last-Modified) for API responses. Prevent cache stampedes with locking or request coalescing. Measure cache hit rates before and after -- caching without metrics is guessing.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST set TTL on ALL cached data -- cache without TTL leads to memory exhaustion and infinitely stale data)
(You MUST use namespaced cache keys with a consistent prefix -- generic keys cause collisions across data types)
(You MUST invalidate or update cache entries on writes -- serving stale data after mutation breaks user trust)
(You MUST implement stampede prevention (locking or coalescing) for high-traffic cache keys -- concurrent misses can overwhelm your data source)
</critical_requirements>
Auto-detection: caching, cache-aside, write-through, write-behind, cache invalidation, TTL, Cache-Control, ETag, Last-Modified, stale-while-revalidate, s-maxage, cache stampede, thundering herd, in-memory cache, LRU cache, distributed cache, cache key, cache miss, cache hit, CDN caching, HTTP caching, conditional request, 304 Not Modified
When to use:
- Read-heavy endpoints fetching the same data repeatedly (cache-aside)
- Data that must stay consistent between cache and database after writes (write-through)
- API responses that benefit from HTTP caching headers (Cache-Control, ETag)
- High-traffic cache keys that risk stampede on expiration
- Reducing database load by caching expensive query results
When NOT to use:
- Data that must always be real-time fresh (caching adds staleness by definition)
- Simple CRUD with low traffic (caching complexity outweighs benefit)
- Development/debugging (caching obscures issues -- disable in dev)
- Premature optimization without measuring actual bottlenecks first
Key patterns covered:
- Cache-aside (lazy loading) with TTL
- Write-through for read-after-write consistency
- Write-behind (write-back) for write-heavy workloads
- HTTP caching: Cache-Control, ETag, Last-Modified, conditional requests
- CDN caching with s-maxage and stale-while-revalidate
- In-memory LRU caching for single-process hot data
- Cache key strategies and namespacing
- Stampede prevention (locking, request coalescing, early recomputation)
- Tag-based and pattern-based invalidation
Detailed Resources:
- examples/core.md - Cache-aside, write-through, HTTP caching, TTL patterns, in-memory caching
- examples/advanced.md - Cache invalidation, stampede prevention, distributed caching, write-behind
- reference.md - Strategy comparison, stampede comparison, cache key conventions
<decision_framework>
Decision Framework
Which Caching Strategy?
Is data read much more than written?
+-- YES --> Can you tolerate staleness?
| +-- YES --> Cache-aside with TTL
| +-- NO --> Write-through (consistent reads)
+-- NO --> Is write latency critical?
+-- YES --> Write-behind (async persist)
+-- NO --> Write-through or skip caching
Which Cache Layer?
Is the response public (same for all users)?
+-- YES --> HTTP caching (Cache-Control: public, s-maxage)
| CDN handles it, your server may never see the request
+-- NO --> Is it shared across server instances?
+-- YES --> Distributed cache (external store)
+-- NO --> Is it hot data in a single process?
+-- YES --> In-memory LRU cache
+-- NO --> Cache-aside with distributed store
TTL Selection Guide
| Data Type | Recommended TTL | Rationale |
|---|---|---|
| Static config / feature flags | 3600s+ (1h+) | Rarely changes |
| Product catalog / listings | 300-3600s (5m-1h) | Infrequent updates |
| User profile | 60-300s (1-5m) | Changes occasionally |
| Search results | 30-60s (30s-1m) | Balance freshness vs performance |
| Real-time data (prices, stock) | 5-30s | Must be near-fresh |
| Session data | 86400s (24h) | Long-lived by design |
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Cache without TTL -- memory grows unbounded and data is stale forever
- No cache invalidation on writes -- users see stale data after mutations, eroding trust
- Generic cache keys without namespacing -- collisions return wrong data to wrong callers
- No stampede prevention on high-traffic keys -- cache expiration triggers data source flood
- Caching errors or null results without separate short TTL -- a transient failure gets cached and served for the full TTL
Medium Priority Issues:
- Same TTL for all data types -- different data has different freshness requirements
- Caching user-specific data in shared caches (CDN) without
privatedirective -- leaks personal data to other users - No cache hit/miss monitoring -- impossible to know if caching is actually helping
no-cacheconfused withno-store--no-cachestill stores but revalidates;no-storeprevents any storage- Over-caching low-traffic endpoints -- adds complexity without meaningful performance gain
Gotchas & Edge Cases:
Cache-Control: no-cachedoes NOT prevent caching -- it forces revalidation before every use. Useno-storeto prevent storage entirely.- ETag takes precedence over Last-Modified during revalidation (per RFC 9110) -- always send both for best compatibility
s-maxageoverridesmax-agefor shared caches (CDNs) only -- browsers ignore itstale-while-revalidatelets CDNs serve expired content while refreshing in the background -- great for availability, but means users briefly see stale data- Distributed cache SET with TTL resets the expiration timer on every write -- calling SET again extends the TTL, which may keep stale data alive longer than expected
- In-memory LRU caches are per-process -- multiple server instances have independent caches that can serve different data for the same key
JSON.stringifyfor cache serialization dropsundefinedvalues and convertsDateobjects to strings -- use a structured serializer if needed- Cache key generation must be deterministic -- object property order in
JSON.stringifyvaries across engines; sort keys or use a canonical serializer
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md
(You MUST set TTL on ALL cached data -- cache without TTL leads to memory exhaustion and infinitely stale data)
(You MUST use namespaced cache keys with a consistent prefix -- generic keys cause collisions across data types)
(You MUST invalidate or update cache entries on writes -- serving stale data after mutation breaks user trust)
(You MUST implement stampede prevention (locking or coalescing) for high-traffic cache keys -- concurrent misses can overwhelm your data source)
Failure to follow these rules will cause stale data bugs, cache key collisions, memory exhaustion, and data source overload during traffic spikes.
</critical_reminders>
Files (skills)
-
examples
-
advanced.md 12.9 KB
# Caching Strategies - Advanced Examples > Cache invalidation, stampede prevention, distributed caching, write-behind. See [core.md](core.md) for cache-aside, HTTP caching, and TTL patterns. --- ## Tag-Based Cache Invalidation Group related cache keys by tag to invalidate entire categories of cached data at once. ```typescript // Good Example - Tag-based cache invalidation async function setWithTags( key: string, value: string, ttl: number, tags: string[], ): Promise<void> { await cacheStore.set(key, value, { ttl }); // Add key to each tag's set for (const tag of tags) { await cacheStore.sAdd(`tag:${tag}`, key); await cacheStore.expire(`tag:${tag}`, ttl); } } async function invalidateByTag(tag: string): Promise<number> { const keys = await cacheStore.sMembers(`tag:${tag}`); if (keys.length === 0) return 0; await cacheStore.del(keys); await cacheStore.del(`tag:${tag}`); return keys.length; } // Usage const PRODUCT_TTL = 3600; await setWithTags("myapp:product:123", JSON.stringify(product), PRODUCT_TTL, [ "category:electronics", "brand:apple", ]); // When electronics category changes -- invalidate all electronics products const invalidated = await invalidateByTag("category:electronics"); export { setWithTags, invalidateByTag }; ``` **Why good:** Enables invalidating related items without knowing exact keys, tag sets auto-expire with TTL, useful for category or relationship-based invalidation --- ## Pattern-Based Key Deletion For cache stores that support key scanning (pattern matching), delete all keys matching a prefix. ```typescript // Good Example - Pattern-based invalidation const SCAN_BATCH_SIZE = 100; async function invalidateByPattern(pattern: string): Promise<number> { let cursor = "0"; let totalDeleted = 0; do { // SCAN is non-blocking (unlike KEYS which blocks for large datasets) const [nextCursor, keys] = await cacheStore.scan(cursor, { match: pattern, count: SCAN_BATCH_SIZE, }); cursor = nextCursor; if (keys.length > 0) { await cacheStore.del(keys); totalDeleted += keys.length; } } while (cursor !== "0"); return totalDeleted; } // Invalidate all cached data for a specific user await invalidateByPattern("myapp:user:456:*"); export { invalidateByPattern }; ``` **Why good:** SCAN is non-blocking (unlike KEYS), processes in batches to avoid memory spikes, returns count for monitoring **When to use:** Invalidating all keys related to an entity (user sessions, user preferences, user cart). **When not to use:** Frequent invalidation of large key sets -- tag-based invalidation is more efficient. --- ## Stampede Prevention: Distributed Lock Only one request regenerates the cache; others wait or receive stale data. Uses atomic set-if-not-exists for lock acquisition. ```typescript // Good Example - Stampede prevention with distributed lock const LOCK_TTL_SECONDS = 10; const RETRY_DELAY_MS = 50; const MAX_RETRIES = 20; async function getWithLock<T>( key: string, fetchFn: () => Promise<T>, ttl: number, ): Promise<T> { // Try cache first const cached = await cacheStore.get(key); if (cached) return JSON.parse(cached) as T; // Attempt to acquire lock (atomic set-if-not-exists) const lockKey = `lock:${key}`; const acquired = await cacheStore.set(lockKey, "1", { ttl: LOCK_TTL_SECONDS, nx: true, // Only set if key does not exist }); if (acquired) { // Lock holder: fetch data, populate cache, release lock try { const data = await fetchFn(); await cacheStore.set(key, JSON.stringify(data), { ttl }); return data; } finally { await cacheStore.del(lockKey); } } // Non-holder: wait and retry for (let i = 0; i < MAX_RETRIES; i++) { await new Promise((resolve) => setTimeout(resolve, RETRY_DELAY_MS)); const retryResult = await cacheStore.get(key); if (retryResult) return JSON.parse(retryResult) as T; } // Fallback: lock holder may have failed -- fetch directly return fetchFn(); } export { getWithLock }; ``` **Why good:** Lock has TTL so it auto-releases if holder crashes, atomic `nx` prevents race conditions, fallback to direct fetch prevents permanent stalls, bounded retries prevent infinite waits --- ## Stampede Prevention: Request Coalescing (Singleflight) In-process deduplication -- concurrent requests for the same key share a single fetch instead of each making their own. ```typescript // Good Example - In-process request coalescing (singleflight pattern) const inFlightRequests = new Map<string, Promise<unknown>>(); async function coalesce<T>(key: string, fetchFn: () => Promise<T>): Promise<T> { const existing = inFlightRequests.get(key); if (existing) return existing as Promise<T>; const promise = fetchFn().finally(() => { inFlightRequests.delete(key); }); inFlightRequests.set(key, promise); return promise; } // Usage with cache-aside const PRODUCT_TTL = 3600; async function getProduct(id: string): Promise<Product> { const cacheKey = `app:product:${id}`; const cached = await cacheStore.get(cacheKey); if (cached) return JSON.parse(cached) as Product; // All concurrent requests for the same product share one fetch const product = await coalesce(cacheKey, async () => { const result = await db.query.products.findFirst({ where: eq(products.id, id), }); if (!result) throw new Error(`Product ${id} not found`); await cacheStore.set(cacheKey, JSON.stringify(result), { ttl: PRODUCT_TTL, }); return result; }); return product; } export { coalesce }; ``` **Why good:** Zero external dependencies (in-process Map), concurrent requests share one fetch (N requests = 1 database query), promise cleaned up automatically via `finally`, simpler than distributed locking **When to use:** Single-process applications or when stampedes happen within a single instance. **When not to use:** Multi-instance deployments where the stampede spans across processes (use distributed locking instead). --- ## Stampede Prevention: Probabilistic Early Recomputation Proactively refresh cache entries before they expire. Each access has an increasing probability of triggering background refresh as the TTL approaches expiration. ```typescript // Good Example - Probabilistic early recomputation const EARLY_RECOMPUTE_FACTOR = 0.1; // Start refreshing at 10% remaining TTL interface CachedEntry<T> { data: T; expiresAt: number; // Unix timestamp in ms ttlMs: number; // Original TTL for probability calculation } async function getWithEarlyRefresh<T>( key: string, fetchFn: () => Promise<T>, ttlMs: number, ): Promise<T> { const raw = await cacheStore.get(key); if (raw) { const entry = JSON.parse(raw) as CachedEntry<T>; const remainingMs = entry.expiresAt - Date.now(); const remainingFraction = remainingMs / entry.ttlMs; // Probability of refresh increases as TTL approaches expiration if ( remainingFraction < EARLY_RECOMPUTE_FACTOR && Math.random() > remainingFraction ) { // Background refresh -- don't await, serve stale data immediately refreshInBackground(key, fetchFn, ttlMs); } return entry.data; } // Cache miss -- fetch synchronously return fetchAndStore(key, fetchFn, ttlMs); } async function fetchAndStore<T>( key: string, fetchFn: () => Promise<T>, ttlMs: number, ): Promise<T> { const data = await fetchFn(); const entry: CachedEntry<T> = { data, expiresAt: Date.now() + ttlMs, ttlMs, }; const ttlSeconds = Math.ceil(ttlMs / 1000); await cacheStore.set(key, JSON.stringify(entry), { ttl: ttlSeconds }); return data; } function refreshInBackground<T>( key: string, fetchFn: () => Promise<T>, ttlMs: number, ): void { fetchAndStore(key, fetchFn, ttlMs).catch((error) => { logger.warn("Background cache refresh failed", { key, error: getErrorMessage(error), }); }); } export { getWithEarlyRefresh }; ``` **Why good:** Responses are never delayed by cache regeneration (always serves cached data), probability-based approach distributes refresh across time (avoids all entries refreshing simultaneously), background failures don't affect current request **When to use:** High-traffic keys where even brief cache misses cause noticeable load spikes. **When not to use:** Low-traffic keys where the extra complexity is not justified. --- ## Write-Behind (Write-Back) Pattern Write to cache immediately, persist to database asynchronously. Improves write latency at the cost of temporary inconsistency and data loss risk. ```typescript // Good Example - Write-behind with buffered persistence const FLUSH_INTERVAL_MS = 5_000; const MAX_BUFFER_SIZE = 100; interface PendingWrite { key: string; value: unknown; timestamp: number; } class WriteBuffer { private buffer: PendingWrite[] = []; private timer: NodeJS.Timeout | null = null; constructor( private readonly persistFn: (writes: PendingWrite[]) => Promise<void>, ) {} add(key: string, value: unknown): void { this.buffer.push({ key, value, timestamp: Date.now() }); if (this.buffer.length >= MAX_BUFFER_SIZE) { void this.flush(); } else if (!this.timer) { this.timer = setTimeout(() => void this.flush(), FLUSH_INTERVAL_MS); } } async flush(): Promise<void> { if (this.timer) { clearTimeout(this.timer); this.timer = null; } if (this.buffer.length === 0) return; const writes = [...this.buffer]; this.buffer = []; try { await this.persistFn(writes); } catch (error) { // Re-queue failed writes for retry logger.error("Write-behind flush failed, re-queuing", { count: writes.length, error: getErrorMessage(error), }); this.buffer.unshift(...writes); } } /** Call on process shutdown to persist remaining writes */ async shutdown(): Promise<void> { await this.flush(); } } // Usage const writeBuffer = new WriteBuffer(async (writes) => { // Batch insert/update in a single transaction await db.transaction(async (tx) => { for (const write of writes) { await tx.insert(analytics).values(write.value as AnalyticsEvent); } }); }); // Write is instant -- only updates cache, buffers the database write async function trackEvent(event: AnalyticsEvent): Promise<void> { const cacheKey = `app:event:${event.id}`; await cacheStore.set(cacheKey, JSON.stringify(event), { ttl: 3600 }); writeBuffer.add(cacheKey, event); } // Register shutdown handler process.on("SIGTERM", async () => { await writeBuffer.shutdown(); process.exit(0); }); export { WriteBuffer, trackEvent }; ``` **Why good:** Write latency is minimal (only cache write), batched persistence reduces database round-trips, failed writes are re-queued, shutdown handler flushes remaining buffer **When to use:** Analytics, logging, activity tracking -- write-heavy workloads where occasional data loss is acceptable. **When not to use:** Financial transactions, user data mutations, or anything requiring write durability guarantees. --- ## Cache Hit/Miss Monitoring Track cache performance to verify caching is actually helping. ```typescript // Good Example - Cache metrics wrapper interface CacheMetrics { hits: number; misses: number; errors: number; } const metrics = new Map<string, CacheMetrics>(); function getMetrics(prefix: string): CacheMetrics { if (!metrics.has(prefix)) { metrics.set(prefix, { hits: 0, misses: 0, errors: 0 }); } return metrics.get(prefix)!; } async function getCachedWithMetrics<T>( key: string, prefix: string, fetchFn: () => Promise<T>, ttl: number, ): Promise<T> { const m = getMetrics(prefix); try { const cached = await cacheStore.get(key); if (cached) { m.hits++; return JSON.parse(cached) as T; } } catch { m.errors++; } m.misses++; const data = await fetchFn(); await cacheStore.set(key, JSON.stringify(data), { ttl }); return data; } function getCacheStats(): Record<string, CacheMetrics & { hitRate: string }> { const stats: Record<string, CacheMetrics & { hitRate: string }> = {}; for (const [prefix, m] of metrics) { const total = m.hits + m.misses; const hitRate = total > 0 ? `${((m.hits / total) * 100).toFixed(1)}%` : "N/A"; stats[prefix] = { ...m, hitRate }; } return stats; } export { getCachedWithMetrics, getCacheStats }; ``` **Why good:** Per-prefix metrics show which caches are effective, hit rate calculation provides actionable data, error tracking surfaces cache connection issues **Target cache hit rates:** | Cache Type | Good Hit Rate | Action if Below | | --------------- | ------------- | -------------------------------------------- | | User profile | > 80% | Increase TTL or check invalidation frequency | | Product catalog | > 90% | Cache may be too small (increase max items) | | Search results | > 60% | Expected lower rate due to query diversity | --- ## See Also - [core.md](core.md) - Cache-aside, HTTP caching, TTL patterns, in-memory LRU caching -
core.md 12.5 KB
# Caching Strategies - Core Examples > Cache-aside, write-through, HTTP caching headers, TTL patterns, in-memory LRU caching. See [advanced.md](advanced.md) for invalidation, stampede prevention, and distributed caching patterns. --- ## Cache-Aside with Generic Wrapper A reusable cache-aside wrapper that handles cache miss, fetch, store, and TTL in one function. ```typescript // Good Example - Generic cache-aside wrapper const DEFAULT_TTL_SECONDS = 300; async function cacheable<T>( cacheKey: string, fetchFn: () => Promise<T>, ttlSeconds: number = DEFAULT_TTL_SECONDS, ): Promise<T> { const cached = await cacheStore.get(cacheKey); if (cached) return JSON.parse(cached) as T; const data = await fetchFn(); await cacheStore.set(cacheKey, JSON.stringify(data), { ttl: ttlSeconds }); return data; } // Usage const PRODUCT_TTL = 3600; const PRODUCT_PREFIX = "app:product"; async function getProduct(id: string): Promise<Product | null> { return cacheable<Product | null>( `${PRODUCT_PREFIX}:${id}`, () => db.query.products.findFirst({ where: eq(products.id, id) }), PRODUCT_TTL, ); } export { cacheable, getProduct }; ``` **Why good:** Separates caching concern from business logic, configurable TTL per call, generic works with any data type, namespaced keys ```typescript // Bad Example - Cache logic mixed with business logic, no TTL async function getProduct(id: string) { try { const cached = await cacheStore.get(id); if (cached) return JSON.parse(cached); } catch { // Silently swallow -- cache failure hides bugs } const product = await db.query.products.findFirst({ where: eq(products.id, id), }); await cacheStore.set(id, JSON.stringify(product)); // No TTL return product; } ``` **Why bad:** No TTL means data never expires, generic key collides with other entity types, silently swallowing cache errors hides connection problems, caching logic mixed with data access --- ## Cache-Aside with Error Handling When the cache store is unavailable, the application should still work -- fall through to the data source. ```typescript // Good Example - Cache failure falls through to data source const USER_TTL = 300; const USER_PREFIX = "app:user"; async function getUserById(userId: string): Promise<User | null> { const cacheKey = `${USER_PREFIX}:${userId}`; // Cache read failure is non-fatal -- fall through to database try { const cached = await cacheStore.get(cacheKey); if (cached) return JSON.parse(cached) as User; } catch (error) { logger.warn("Cache read failed, falling through to database", { key: cacheKey, error: getErrorMessage(error), }); } const user = await db.query.users.findFirst({ where: eq(users.id, userId) }); if (!user) return null; // Cache write failure is non-fatal -- data was still fetched successfully try { await cacheStore.set(cacheKey, JSON.stringify(user), { ttl: USER_TTL }); } catch (error) { logger.warn("Cache write failed", { key: cacheKey, error: getErrorMessage(error), }); } return user; } ``` **Why good:** Application degrades gracefully when cache is down, cache errors are logged (not silently swallowed), database is the source of truth --- ## Write-Through with Delete on Remove ```typescript // Good Example - Write-through: update cache on mutation, delete on remove const CACHE_TTL = 300; const USER_PREFIX = "app:user"; async function updateUser( userId: string, updates: Partial<User>, ): Promise<User> { // Update database first (source of truth) const [updatedUser] = await db .update(users) .set({ ...updates, updatedAt: new Date() }) .where(eq(users.id, userId)) .returning(); // Write-through: update cache immediately const cacheKey = `${USER_PREFIX}:${userId}`; await cacheStore.set(cacheKey, JSON.stringify(updatedUser), { ttl: CACHE_TTL, }); return updatedUser; } async function deleteUser(userId: string): Promise<void> { await db.delete(users).where(eq(users.id, userId)); // Invalidate cache on delete const cacheKey = `${USER_PREFIX}:${userId}`; await cacheStore.del(cacheKey); } export { updateUser, deleteUser }; ``` **Why good:** Cache always reflects latest database state after writes, explicit invalidation on delete prevents stale reads, TTL still set as safety net ```typescript // Bad Example - Update database but forget to update cache async function updateUser(userId: string, updates: Partial<User>) { await db.update(users).set(updates).where(eq(users.id, userId)); // Cache still has old data until TTL expires! } ``` **Why bad:** Stale cache serves outdated data until TTL expires, users see old values after saving changes --- ## HTTP Caching: Cache-Control Headers Setting appropriate Cache-Control headers on API responses reduces requests to your server entirely. ```typescript // Good Example - Cache-Control for different API endpoint types // Public list endpoint -- CDN can cache, browser caches briefly const LIST_MAX_AGE = 60; const LIST_S_MAXAGE = 300; const LIST_SWR = 60; function setPublicListHeaders(res: Response): void { res.setHeader( "Cache-Control", `public, max-age=${LIST_MAX_AGE}, s-maxage=${LIST_S_MAXAGE}, stale-while-revalidate=${LIST_SWR}`, ); } // User-specific endpoint -- private cache only, always revalidate function setPrivateHeaders(res: Response): void { res.setHeader("Cache-Control", "private, no-cache"); } // Sensitive data -- never cache function setNoCacheHeaders(res: Response): void { res.setHeader("Cache-Control", "no-store, private"); } ``` **Why good:** Different endpoints get appropriate caching strategies, `s-maxage` lets CDN cache longer than browser, `stale-while-revalidate` improves availability during refresh --- ## HTTP Caching: ETag and Conditional Requests ETags enable conditional requests -- the server can respond with 304 Not Modified when data hasn't changed, saving bandwidth and serialization cost. ```typescript // Good Example - ETag generation and conditional request handling import { createHash } from "node:crypto"; function generateETag(body: unknown): string { const content = JSON.stringify(body); const hash = createHash("md5").update(content).digest("hex"); return `"${hash}"`; } // Middleware or route handler for conditional responses async function handleConditionalGet( req: Request, res: Response, body: unknown, ): Promise<boolean> { const etag = generateETag(body); res.setHeader("ETag", etag); const clientETag = req.headers["if-none-match"]; if (clientETag === etag) { res.status(304).end(); return true; // Response already sent } return false; // Caller should send the full response } // Usage in a route handler async function getProducts(req: Request, res: Response) { const products = await db.query.products.findMany(); const notModified = await handleConditionalGet(req, res, products); if (notModified) return; res.setHeader("Cache-Control", "public, max-age=60, must-revalidate"); res.json(products); } export { generateETag, handleConditionalGet }; ``` **Why good:** 304 responses save bandwidth (no body), ETag based on content hash detects actual changes, separates conditional logic from route handler **When to use:** Endpoints where the response body is expensive to transfer but cheap to check for changes (product listings, configuration data). --- ## CDN Caching with s-maxage Use `s-maxage` to let CDNs (shared caches) cache responses longer than the browser. Combine with `stale-while-revalidate` for high availability. ```typescript // Good Example - CDN-optimized cache headers const BROWSER_MAX_AGE = 60; // Browser: 1 minute const CDN_MAX_AGE = 300; // CDN: 5 minutes const SWR_WINDOW = 60; // Serve stale for 60s while refreshing const ERROR_WINDOW = 600; // Serve stale for 10m during origin errors function setCDNCacheHeaders(res: Response): void { res.setHeader( "Cache-Control", [ "public", `max-age=${BROWSER_MAX_AGE}`, `s-maxage=${CDN_MAX_AGE}`, `stale-while-revalidate=${SWR_WINDOW}`, `stale-if-error=${ERROR_WINDOW}`, ].join(", "), ); } ``` **Why good:** `s-maxage` overrides `max-age` for CDNs only, `stale-while-revalidate` ensures users never wait for origin refresh, `stale-if-error` provides resilience during origin outages --- ## In-Memory LRU Cache For single-process hot data. Faster than a distributed cache (no network hop) but not shared across instances. ```typescript // Good Example - LRU cache for hot data (lru-cache v11+) import { LRUCache } from "lru-cache"; const MAX_ITEMS = 500; const TTL_MS = 300_000; // 5 minutes interface CachedConfig { value: string; updatedAt: Date; } const configCache = new LRUCache<string, CachedConfig>({ max: MAX_ITEMS, ttl: TTL_MS, }); function getCachedConfig(key: string): CachedConfig | undefined { return configCache.get(key); } function setCachedConfig(key: string, config: CachedConfig): void { configCache.set(key, config); } export { getCachedConfig, setCachedConfig }; ``` **Why good:** No serialization overhead (stores objects directly), automatic LRU eviction when max reached, TTL prevents staleness, type-safe with generics ```typescript // Bad Example - Unbounded Map as cache const cache = new Map<string, unknown>(); // No max size, no TTL function getCached(key: string) { return cache.get(key); // Never evicted, never expires } function setCached(key: string, value: unknown) { cache.set(key, value); // Memory grows without bound } ``` **Why bad:** No size limit means memory grows unbounded, no TTL means data is stale forever, no eviction policy means the cache fills up and never frees memory --- ## LRU Cache with Fetch Method The `fetch` option in lru-cache provides built-in cache-aside behavior -- on a miss, it calls your fetch function automatically. ```typescript // Good Example - LRU cache with automatic fetch on miss import { LRUCache } from "lru-cache"; const MAX_ITEMS = 200; const TTL_MS = 60_000; // 1 minute const featureFlagCache = new LRUCache<string, boolean>({ max: MAX_ITEMS, ttl: TTL_MS, // Called automatically on cache miss fetchMethod: async (key) => { const flag = await db.query.featureFlags.findFirst({ where: eq(featureFlags.key, key), }); return flag?.enabled ?? false; }, }); // Usage: fetch() returns cached value or calls fetchMethod async function isFeatureEnabled(flagKey: string): Promise<boolean> { const result = await featureFlagCache.fetch(flagKey); return result ?? false; } export { isFeatureEnabled }; ``` **Why good:** Cache-aside logic is built into the cache itself, no manual get/set coordination, fetch is deduplicated (concurrent calls for the same key share one fetch) --- ## TTL Strategy by Data Type ```typescript // Named constants for TTL values -- document the rationale const TTL = { /** Static config that changes via deploy */ STATIC_CONFIG: 3600, /** Product catalog updated by admin */ PRODUCT: 300, /** User profile updated by user */ USER_PROFILE: 60, /** Search results -- balance freshness vs performance */ SEARCH: 30, /** Real-time data -- near-fresh */ REALTIME: 10, /** Session data -- long-lived by design */ SESSION: 86400, } as const; export { TTL }; ``` **Why good:** Named constants with comments make TTL decisions explicit, `as const` preserves literal types, centralized TTL values prevent inconsistency across the codebase --- ## Cache Key Generation ```typescript // Good Example - Structured cache key generation import { createHash } from "node:crypto"; const CACHE_PREFIX = "myapp"; // Entity keys: prefix:type:id const cacheKeys = { user: (id: string) => `${CACHE_PREFIX}:user:${id}`, product: (id: string) => `${CACHE_PREFIX}:product:${id}`, cart: (userId: string) => `${CACHE_PREFIX}:cart:${userId}`, // Query result keys: hash the normalized query params productList: (filters: Record<string, unknown>) => { // Sort keys for deterministic hashing const sorted = JSON.stringify(filters, Object.keys(filters).sort()); const hash = createHash("md5").update(sorted).digest("hex"); return `${CACHE_PREFIX}:products:list:${hash}`; }, // Pattern for bulk invalidation userPattern: (userId: string) => `${CACHE_PREFIX}:user:${userId}:*`, } as const; export { cacheKeys, CACHE_PREFIX }; ``` **Why good:** Consistent prefix prevents cross-application collisions, hierarchical keys enable pattern-based invalidation, sorted JSON ensures deterministic hashing regardless of property insertion order --- ## See Also - [advanced.md](advanced.md) - Cache invalidation, stampede prevention, distributed caching
-
-
reference.md 3.7 KB
# Caching Strategies Quick Reference Lookup tables and comparisons. See [SKILL.md](SKILL.md) for decision frameworks, red flags, and philosophy. --- ## Cache-Control Quick Reference | Use Case | Header | Behavior | | ----------------------- | ------------------------------------------------------------- | ----------------------------------- | | Public list data | `public, max-age=60, s-maxage=300` | Browser: 1m, CDN: 5m | | Public with SWR | `public, max-age=60, s-maxage=300, stale-while-revalidate=60` | CDN serves stale during refresh | | User-specific data | `private, no-cache` | Always revalidate, no shared cache | | Sensitive data | `no-store, private` | Never cached anywhere | | Versioned static assets | `public, max-age=31536000, immutable` | Cache forever, version in URL | | API with ETag | `max-age=0, must-revalidate` | Always revalidate, 304 if unchanged | --- ## Conditional Request Flow ``` Client sends: If-None-Match: "abc123" If-Modified-Since: Tue, 01 Jan 2025 00:00:00 GMT Server checks: ETag matches? (takes precedence per RFC 9110) +-- YES --> 304 Not Modified (no body) +-- NO --> Last-Modified changed? +-- NO --> 304 Not Modified (no body) +-- YES --> 200 OK (full response) ``` --- ## Strategy Comparison | Strategy | Read Perf | Write Perf | Consistency | Complexity | Data Loss Risk | | ------------- | -------------------- | ------------------- | ---------------- | ---------- | --------------------- | | Cache-aside | Fast (on hit) | N/A (reads only) | Eventual (TTL) | Low | None | | Write-through | Fast (on hit) | Slower (dual write) | Strong | Low | None | | Write-behind | Fast (on hit) | Fast (async) | Eventual | High | Yes (buffer loss) | | HTTP caching | Fastest (no request) | N/A | Eventual (TTL) | Low | None | | In-memory LRU | Sub-ms | Sub-ms | Per-process only | Low | Yes (process restart) | --- ## Stampede Prevention Comparison | Technique | When to Use | Trade-off | | --------------------------- | ------------------------------------ | ---------------------------------- | | Distributed lock | Multi-instance, critical keys | Lock overhead, potential stalls | | Request coalescing | Single-instance, bursty traffic | In-process only, no cross-instance | | Probabilistic early refresh | High-traffic, predictable patterns | Slightly higher cache store load | | Background refresh (cron) | Periodic data, known update schedule | Stale between refreshes | --- ## Cache Key Conventions ``` {app}:{entity}:{id} -- Entity by ID {app}:{entity}:list:{hash} -- Query result by hashed params {app}:{entity}:{id}:{subentity} -- Related entity tag:{category} -- Tag set for bulk invalidation lock:{key} -- Stampede prevention lock ``` **Rules:** - Always prefix with application name (prevents cross-app collisions) - Use colons as separators (convention for hierarchical keys) - Hash complex query params for shorter keys (MD5 is fine for cache keys) - Sort object keys before hashing (ensures deterministic output) -
SKILL.md 17.3 KB
--- name: api-caching-strategies description: Application-level caching strategies, HTTP caching, cache invalidation, and stampede prevention --- # Caching Strategies > **Quick Guide:** Choose the right caching strategy for your use case: cache-aside for read-heavy data, write-through for consistency, write-behind for write-heavy workloads. Always set TTL to prevent stale data and memory exhaustion. Use HTTP caching headers (Cache-Control, ETag, Last-Modified) for API responses. Prevent cache stampedes with locking or request coalescing. Measure cache hit rates before and after -- caching without metrics is guessing. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST set TTL on ALL cached data -- cache without TTL leads to memory exhaustion and infinitely stale data)** **(You MUST use namespaced cache keys with a consistent prefix -- generic keys cause collisions across data types)** **(You MUST invalidate or update cache entries on writes -- serving stale data after mutation breaks user trust)** **(You MUST implement stampede prevention (locking or coalescing) for high-traffic cache keys -- concurrent misses can overwhelm your data source)** </critical_requirements> --- **Auto-detection:** caching, cache-aside, write-through, write-behind, cache invalidation, TTL, Cache-Control, ETag, Last-Modified, stale-while-revalidate, s-maxage, cache stampede, thundering herd, in-memory cache, LRU cache, distributed cache, cache key, cache miss, cache hit, CDN caching, HTTP caching, conditional request, 304 Not Modified **When to use:** - Read-heavy endpoints fetching the same data repeatedly (cache-aside) - Data that must stay consistent between cache and database after writes (write-through) - API responses that benefit from HTTP caching headers (Cache-Control, ETag) - High-traffic cache keys that risk stampede on expiration - Reducing database load by caching expensive query results **When NOT to use:** - Data that must always be real-time fresh (caching adds staleness by definition) - Simple CRUD with low traffic (caching complexity outweighs benefit) - Development/debugging (caching obscures issues -- disable in dev) - Premature optimization without measuring actual bottlenecks first **Key patterns covered:** - Cache-aside (lazy loading) with TTL - Write-through for read-after-write consistency - Write-behind (write-back) for write-heavy workloads - HTTP caching: Cache-Control, ETag, Last-Modified, conditional requests - CDN caching with s-maxage and stale-while-revalidate - In-memory LRU caching for single-process hot data - Cache key strategies and namespacing - Stampede prevention (locking, request coalescing, early recomputation) - Tag-based and pattern-based invalidation --- <philosophy> ## Philosophy Caching trades freshness for speed. Every caching decision is a **consistency vs performance** trade-off -- understand where your use case falls on that spectrum before choosing a strategy. **The three questions before adding caching:** 1. **Is this actually slow?** Measure first. If the uncached response is fast enough, caching adds complexity without benefit. 2. **Can I tolerate staleness?** If data must always be real-time, caching is the wrong tool. Use read replicas or materialized views instead. 3. **What is the read-to-write ratio?** Caching shines when reads vastly outnumber writes. For write-heavy workloads, consider write-behind or skip caching entirely. **Caching is a layered system.** HTTP caching (browser and CDN) reduces requests to your server. Application-level caching (in-memory or distributed) reduces requests to your database. Apply caching at the right layer for the problem. **When to use caching:** - Response times exceed acceptable thresholds and the data source is the bottleneck - The same data is fetched repeatedly across requests - Data freshness requirements allow some staleness (even 60 seconds) - Traffic is high enough that database load is a concern **When NOT to use caching:** - Data changes frequently and must always be current - Every request returns unique data (no cache reuse) - You have not measured the actual bottleneck yet (premature optimization) </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Cache-Aside (Lazy Loading) The most common application-level caching pattern. The application checks the cache first, fetches from the data source on miss, and stores the result with a TTL. ```typescript const CACHE_TTL_SECONDS = 300; const CACHE_PREFIX = "app:user"; async function getUserById(userId: string): Promise<User | null> { const cacheKey = `${CACHE_PREFIX}:${userId}`; const cached = await cacheStore.get(cacheKey); if (cached) return JSON.parse(cached) as User; const user = await db.query.users.findFirst({ where: eq(users.id, userId) }); if (!user) return null; await cacheStore.set(cacheKey, JSON.stringify(user), { ttl: CACHE_TTL_SECONDS, }); return user; } ``` **Why good:** TTL prevents stale data, namespaced keys prevent collisions, early return on cache hit, cache is populated lazily (only data that is actually requested gets cached) ```typescript // BAD: No TTL, generic key async function getUser(id: string) { const cached = await cacheStore.get(id); // No prefix -- collides with other data types if (cached) return JSON.parse(cached); const user = await db.query.users.findFirst({ where: eq(users.id, id) }); await cacheStore.set(id, JSON.stringify(user)); // No TTL -- never expires return user; } ``` **Why bad:** No TTL means infinite staleness and eventual memory exhaustion, generic key collides with other entity types using the same ID format **When to use:** Read-heavy endpoints where data changes infrequently relative to reads. See [examples/core.md](examples/core.md) for generic cache wrapper and error handling patterns. --- ### Pattern 2: Write-Through Update both cache and data source on every write. Guarantees read-after-write consistency without waiting for TTL expiration. ```typescript const CACHE_TTL_SECONDS = 300; async function updateUser( userId: string, updates: Partial<User>, ): Promise<User> { const updatedUser = await db .update(users) .set({ ...updates, updatedAt: new Date() }) .where(eq(users.id, userId)) .returning(); const cacheKey = `${CACHE_PREFIX}:${userId}`; await cacheStore.set(cacheKey, JSON.stringify(updatedUser[0]), { ttl: CACHE_TTL_SECONDS, }); return updatedUser[0]; } ``` **Why good:** Cache always reflects latest database state, no stale reads after updates **When to use:** Data that is read frequently after writes and must be consistent. **When not to use:** Write-heavy workloads where the overhead of updating cache on every write is too expensive. See [examples/core.md](examples/core.md) for write-through with delete invalidation. --- ### Pattern 3: HTTP Caching Headers Set appropriate Cache-Control, ETag, and Last-Modified headers on API responses to leverage browser and CDN caching. This reduces requests to your server entirely. ```typescript // API endpoint setting caching headers function setCacheHeaders( res: Response, body: unknown, options: { maxAge: number; isPublic: boolean }, ) { const etag = generateETag(body); const scope = options.isPublic ? "public" : "private"; res.setHeader( "Cache-Control", `${scope}, max-age=${options.maxAge}, must-revalidate`, ); res.setHeader("ETag", etag); res.setHeader("Last-Modified", new Date().toUTCString()); } ``` **Key header combinations for APIs:** | Use Case | Cache-Control | Why | | ------------------ | ------------------------------------- | ---------------------------------- | | Public list data | `public, max-age=60, s-maxage=300` | CDN caches longer than browser | | User-specific data | `private, max-age=0, must-revalidate` | No shared cache, always revalidate | | Sensitive data | `no-store, private` | Never cache anywhere | | Immutable assets | `public, max-age=31536000, immutable` | Version in URL, cache forever | See [examples/core.md](examples/core.md) for conditional request handling (If-None-Match / 304) and CDN patterns. --- ### Pattern 4: In-Memory LRU Cache For single-process hot data that does not need to be shared across instances. Faster than a distributed cache (no network hop) but lost on process restart and not shared between workers. ```typescript import { LRUCache } from "lru-cache"; const MAX_ITEMS = 500; const TTL_MS = 300_000; // 5 minutes const configCache = new LRUCache<string, AppConfig>({ max: MAX_ITEMS, ttl: TTL_MS, }); function getConfig(key: string): AppConfig | undefined { return configCache.get(key); } function setConfig(key: string, value: AppConfig): void { configCache.set(key, value); } ``` **Why good:** No serialization overhead (stores objects directly), automatic eviction of least-recently-used items, TTL prevents staleness **When to use:** Configuration, feature flags, frequently accessed reference data in single-process applications. **When not to use:** Data that must be shared across multiple server instances (use a distributed cache instead). See [examples/core.md](examples/core.md) for LRU cache with size tracking and fetch method. --- ### Pattern 5: Cache Key Strategies Consistent, namespaced cache keys prevent collisions and enable pattern-based invalidation. ```typescript const CACHE_PREFIX = "myapp"; // Entity keys: prefix:type:id const userKey = (id: string) => `${CACHE_PREFIX}:user:${id}`; const productKey = (id: string) => `${CACHE_PREFIX}:product:${id}`; // Query result keys: prefix:type:list:hash const productListKey = (filters: ProductFilters) => { const normalized = JSON.stringify({ category: filters.category ?? "all", page: filters.page ?? 1, sort: filters.sort ?? "created", }); const hash = createHash("md5").update(normalized).digest("hex"); return `${CACHE_PREFIX}:products:list:${hash}`; }; // User-scoped keys const userCartKey = (userId: string) => `${CACHE_PREFIX}:cart:${userId}`; ``` **Why good:** Hierarchical structure prevents collisions, hashing complex queries keeps keys short, pattern prefix enables bulk invalidation See [examples/core.md](examples/core.md) for cache key generation patterns. --- ### Pattern 6: Cache Invalidation Invalidation is the hardest part of caching. Choose the simplest strategy that meets your consistency requirements. **Direct invalidation** -- delete the exact key on mutation: ```typescript async function updateProduct(id: string, data: ProductUpdate) { await db.update(products).set(data).where(eq(products.id, id)); await cacheStore.del(productKey(id)); } ``` **Tag-based invalidation** -- group related keys by tag for bulk invalidation: ```typescript // Invalidate all electronics products when category changes await invalidateByTag("category:electronics"); ``` **TTL-based expiration** -- let the cache expire naturally when eventual consistency is acceptable: ```typescript const SHORT_TTL = 60; // 1 minute for frequently changing data const MEDIUM_TTL = 300; // 5 minutes for user data const LONG_TTL = 3600; // 1 hour for reference data ``` See [examples/advanced.md](examples/advanced.md) for tag-based invalidation implementation and pattern-based key deletion. --- ### Pattern 7: Stampede Prevention When a popular cache key expires, many concurrent requests may all miss the cache simultaneously and flood the data source. This is a cache stampede (thundering herd). **Locking** -- only one request regenerates the cache; others wait or get stale data: ```typescript async function getWithLock( key: string, fetchFn: () => Promise<string>, ): Promise<string> { const cached = await cacheStore.get(key); if (cached) return cached; const lockKey = `lock:${key}`; const acquired = await cacheStore.setNX(lockKey, "1", { ttl: LOCK_TTL_SECONDS, }); if (!acquired) { await delay(RETRY_DELAY_MS); return getWithLock(key, fetchFn); // Retry after brief wait } try { const value = await fetchFn(); await cacheStore.set(key, value, { ttl: CACHE_TTL_SECONDS }); return value; } finally { await cacheStore.del(lockKey); } } ``` **Why good:** Only one request hits the data source, others wait briefly, lock auto-expires if holder crashes See [examples/advanced.md](examples/advanced.md) for request coalescing (singleflight) and probabilistic early recomputation. </patterns> --- **Detailed Resources:** - [examples/core.md](examples/core.md) - Cache-aside, write-through, HTTP caching, TTL patterns, in-memory caching - [examples/advanced.md](examples/advanced.md) - Cache invalidation, stampede prevention, distributed caching, write-behind - [reference.md](reference.md) - Strategy comparison, stampede comparison, cache key conventions --- <decision_framework> ## Decision Framework ### Which Caching Strategy? ``` Is data read much more than written? +-- YES --> Can you tolerate staleness? | +-- YES --> Cache-aside with TTL | +-- NO --> Write-through (consistent reads) +-- NO --> Is write latency critical? +-- YES --> Write-behind (async persist) +-- NO --> Write-through or skip caching ``` ### Which Cache Layer? ``` Is the response public (same for all users)? +-- YES --> HTTP caching (Cache-Control: public, s-maxage) | CDN handles it, your server may never see the request +-- NO --> Is it shared across server instances? +-- YES --> Distributed cache (external store) +-- NO --> Is it hot data in a single process? +-- YES --> In-memory LRU cache +-- NO --> Cache-aside with distributed store ``` ### TTL Selection Guide | Data Type | Recommended TTL | Rationale | | ------------------------------ | ----------------- | -------------------------------- | | Static config / feature flags | 3600s+ (1h+) | Rarely changes | | Product catalog / listings | 300-3600s (5m-1h) | Infrequent updates | | User profile | 60-300s (1-5m) | Changes occasionally | | Search results | 30-60s (30s-1m) | Balance freshness vs performance | | Real-time data (prices, stock) | 5-30s | Must be near-fresh | | Session data | 86400s (24h) | Long-lived by design | </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Cache without TTL -- memory grows unbounded and data is stale forever - No cache invalidation on writes -- users see stale data after mutations, eroding trust - Generic cache keys without namespacing -- collisions return wrong data to wrong callers - No stampede prevention on high-traffic keys -- cache expiration triggers data source flood - Caching errors or null results without separate short TTL -- a transient failure gets cached and served for the full TTL **Medium Priority Issues:** - Same TTL for all data types -- different data has different freshness requirements - Caching user-specific data in shared caches (CDN) without `private` directive -- leaks personal data to other users - No cache hit/miss monitoring -- impossible to know if caching is actually helping - `no-cache` confused with `no-store` -- `no-cache` still stores but revalidates; `no-store` prevents any storage - Over-caching low-traffic endpoints -- adds complexity without meaningful performance gain **Gotchas & Edge Cases:** - `Cache-Control: no-cache` does NOT prevent caching -- it forces revalidation before every use. Use `no-store` to prevent storage entirely. - ETag takes precedence over Last-Modified during revalidation (per RFC 9110) -- always send both for best compatibility - `s-maxage` overrides `max-age` for shared caches (CDNs) only -- browsers ignore it - `stale-while-revalidate` lets CDNs serve expired content while refreshing in the background -- great for availability, but means users briefly see stale data - Distributed cache SET with TTL resets the expiration timer on every write -- calling SET again extends the TTL, which may keep stale data alive longer than expected - In-memory LRU caches are per-process -- multiple server instances have independent caches that can serve different data for the same key - `JSON.stringify` for cache serialization drops `undefined` values and converts `Date` objects to strings -- use a structured serializer if needed - Cache key generation must be deterministic -- object property order in `JSON.stringify` varies across engines; sort keys or use a canonical serializer </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** **(You MUST set TTL on ALL cached data -- cache without TTL leads to memory exhaustion and infinitely stale data)** **(You MUST use namespaced cache keys with a consistent prefix -- generic keys cause collisions across data types)** **(You MUST invalidate or update cache entries on writes -- serving stale data after mutation breaks user trust)** **(You MUST implement stampede prevention (locking or coalescing) for high-traffic cache keys -- concurrent misses can overwhelm your data source)** **Failure to follow these rules will cause stale data bugs, cache key collisions, memory exhaustion, and data source overload during traffic spikes.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.