Claude Skill

api-caching-strategies

Application-level caching strategies, HTTP caching, cache invalidation, and stampede prevention

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agents-inc-skills-dist_plugins_api-caching-strategies_skills_api-caching-strategies-3a51ef5.zip · 16 KB
Part of agents-inc/skills — 130 skills

Install

skills CLI npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/api-caching-strategies/skills/api-caching-strategies
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
Git git clone https://github.com/agents-inc/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Caching Strategies

Quick Guide: Choose the right caching strategy for your use case: cache-aside for read-heavy data, write-through for consistency, write-behind for write-heavy workloads. Always set TTL to prevent stale data and memory exhaustion. Use HTTP caching headers (Cache-Control, ETag, Last-Modified) for API responses. Prevent cache stampedes with locking or request coalescing. Measure cache hit rates before and after -- caching without metrics is guessing.


<critical_requirements>

CRITICAL: Before Using This Skill

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST set TTL on ALL cached data -- cache without TTL leads to memory exhaustion and infinitely stale data)

(You MUST use namespaced cache keys with a consistent prefix -- generic keys cause collisions across data types)

(You MUST invalidate or update cache entries on writes -- serving stale data after mutation breaks user trust)

(You MUST implement stampede prevention (locking or coalescing) for high-traffic cache keys -- concurrent misses can overwhelm your data source)

</critical_requirements>


Auto-detection: caching, cache-aside, write-through, write-behind, cache invalidation, TTL, Cache-Control, ETag, Last-Modified, stale-while-revalidate, s-maxage, cache stampede, thundering herd, in-memory cache, LRU cache, distributed cache, cache key, cache miss, cache hit, CDN caching, HTTP caching, conditional request, 304 Not Modified

When to use:

  • Read-heavy endpoints fetching the same data repeatedly (cache-aside)
  • Data that must stay consistent between cache and database after writes (write-through)
  • API responses that benefit from HTTP caching headers (Cache-Control, ETag)
  • High-traffic cache keys that risk stampede on expiration
  • Reducing database load by caching expensive query results

When NOT to use:

  • Data that must always be real-time fresh (caching adds staleness by definition)
  • Simple CRUD with low traffic (caching complexity outweighs benefit)
  • Development/debugging (caching obscures issues -- disable in dev)
  • Premature optimization without measuring actual bottlenecks first

Key patterns covered:

  • Cache-aside (lazy loading) with TTL
  • Write-through for read-after-write consistency
  • Write-behind (write-back) for write-heavy workloads
  • HTTP caching: Cache-Control, ETag, Last-Modified, conditional requests
  • CDN caching with s-maxage and stale-while-revalidate
  • In-memory LRU caching for single-process hot data
  • Cache key strategies and namespacing
  • Stampede prevention (locking, request coalescing, early recomputation)
  • Tag-based and pattern-based invalidation



Detailed Resources:

  • examples/core.md - Cache-aside, write-through, HTTP caching, TTL patterns, in-memory caching
  • examples/advanced.md - Cache invalidation, stampede prevention, distributed caching, write-behind
  • reference.md - Strategy comparison, stampede comparison, cache key conventions

<decision_framework>

Decision Framework

Which Caching Strategy?

Is data read much more than written?
+-- YES --> Can you tolerate staleness?
|   +-- YES --> Cache-aside with TTL
|   +-- NO  --> Write-through (consistent reads)
+-- NO  --> Is write latency critical?
    +-- YES --> Write-behind (async persist)
    +-- NO  --> Write-through or skip caching

Which Cache Layer?

Is the response public (same for all users)?
+-- YES --> HTTP caching (Cache-Control: public, s-maxage)
|          CDN handles it, your server may never see the request
+-- NO  --> Is it shared across server instances?
    +-- YES --> Distributed cache (external store)
    +-- NO  --> Is it hot data in a single process?
        +-- YES --> In-memory LRU cache
        +-- NO  --> Cache-aside with distributed store

TTL Selection Guide

Data Type Recommended TTL Rationale
Static config / feature flags 3600s+ (1h+) Rarely changes
Product catalog / listings 300-3600s (5m-1h) Infrequent updates
User profile 60-300s (1-5m) Changes occasionally
Search results 30-60s (30s-1m) Balance freshness vs performance
Real-time data (prices, stock) 5-30s Must be near-fresh
Session data 86400s (24h) Long-lived by design

</decision_framework>


<red_flags>

RED FLAGS

High Priority Issues:

  • Cache without TTL -- memory grows unbounded and data is stale forever
  • No cache invalidation on writes -- users see stale data after mutations, eroding trust
  • Generic cache keys without namespacing -- collisions return wrong data to wrong callers
  • No stampede prevention on high-traffic keys -- cache expiration triggers data source flood
  • Caching errors or null results without separate short TTL -- a transient failure gets cached and served for the full TTL

Medium Priority Issues:

  • Same TTL for all data types -- different data has different freshness requirements
  • Caching user-specific data in shared caches (CDN) without private directive -- leaks personal data to other users
  • No cache hit/miss monitoring -- impossible to know if caching is actually helping
  • no-cache confused with no-store -- no-cache still stores but revalidates; no-store prevents any storage
  • Over-caching low-traffic endpoints -- adds complexity without meaningful performance gain

Gotchas & Edge Cases:

  • Cache-Control: no-cache does NOT prevent caching -- it forces revalidation before every use. Use no-store to prevent storage entirely.
  • ETag takes precedence over Last-Modified during revalidation (per RFC 9110) -- always send both for best compatibility
  • s-maxage overrides max-age for shared caches (CDNs) only -- browsers ignore it
  • stale-while-revalidate lets CDNs serve expired content while refreshing in the background -- great for availability, but means users briefly see stale data
  • Distributed cache SET with TTL resets the expiration timer on every write -- calling SET again extends the TTL, which may keep stale data alive longer than expected
  • In-memory LRU caches are per-process -- multiple server instances have independent caches that can serve different data for the same key
  • JSON.stringify for cache serialization drops undefined values and converts Date objects to strings -- use a structured serializer if needed
  • Cache key generation must be deterministic -- object property order in JSON.stringify varies across engines; sort keys or use a canonical serializer

</red_flags>


<critical_reminders>

CRITICAL REMINDERS

All code must follow project conventions in CLAUDE.md

(You MUST set TTL on ALL cached data -- cache without TTL leads to memory exhaustion and infinitely stale data)

(You MUST use namespaced cache keys with a consistent prefix -- generic keys cause collisions across data types)

(You MUST invalidate or update cache entries on writes -- serving stale data after mutation breaks user trust)

(You MUST implement stampede prevention (locking or coalescing) for high-traffic cache keys -- concurrent misses can overwhelm your data source)

Failure to follow these rules will cause stale data bugs, cache key collisions, memory exhaustion, and data source overload during traffic spikes.

</critical_reminders>

Files (skills)
  • examples
    • advanced.md 12.9 KB
      # Caching Strategies - Advanced Examples
      
      > Cache invalidation, stampede prevention, distributed caching, write-behind. See [core.md](core.md) for cache-aside, HTTP caching, and TTL patterns.
      
      ---
      
      ## Tag-Based Cache Invalidation
      
      Group related cache keys by tag to invalidate entire categories of cached data at once.
      
      ```typescript
      // Good Example - Tag-based cache invalidation
      async function setWithTags(
        key: string,
        value: string,
        ttl: number,
        tags: string[],
      ): Promise<void> {
        await cacheStore.set(key, value, { ttl });
      
        // Add key to each tag's set
        for (const tag of tags) {
          await cacheStore.sAdd(`tag:${tag}`, key);
          await cacheStore.expire(`tag:${tag}`, ttl);
        }
      }
      
      async function invalidateByTag(tag: string): Promise<number> {
        const keys = await cacheStore.sMembers(`tag:${tag}`);
        if (keys.length === 0) return 0;
      
        await cacheStore.del(keys);
        await cacheStore.del(`tag:${tag}`);
        return keys.length;
      }
      
      // Usage
      const PRODUCT_TTL = 3600;
      
      await setWithTags("myapp:product:123", JSON.stringify(product), PRODUCT_TTL, [
        "category:electronics",
        "brand:apple",
      ]);
      
      // When electronics category changes -- invalidate all electronics products
      const invalidated = await invalidateByTag("category:electronics");
      
      export { setWithTags, invalidateByTag };
      ```
      
      **Why good:** Enables invalidating related items without knowing exact keys, tag sets auto-expire with TTL, useful for category or relationship-based invalidation
      
      ---
      
      ## Pattern-Based Key Deletion
      
      For cache stores that support key scanning (pattern matching), delete all keys matching a prefix.
      
      ```typescript
      // Good Example - Pattern-based invalidation
      const SCAN_BATCH_SIZE = 100;
      
      async function invalidateByPattern(pattern: string): Promise<number> {
        let cursor = "0";
        let totalDeleted = 0;
      
        do {
          // SCAN is non-blocking (unlike KEYS which blocks for large datasets)
          const [nextCursor, keys] = await cacheStore.scan(cursor, {
            match: pattern,
            count: SCAN_BATCH_SIZE,
          });
          cursor = nextCursor;
      
          if (keys.length > 0) {
            await cacheStore.del(keys);
            totalDeleted += keys.length;
          }
        } while (cursor !== "0");
      
        return totalDeleted;
      }
      
      // Invalidate all cached data for a specific user
      await invalidateByPattern("myapp:user:456:*");
      
      export { invalidateByPattern };
      ```
      
      **Why good:** SCAN is non-blocking (unlike KEYS), processes in batches to avoid memory spikes, returns count for monitoring
      
      **When to use:** Invalidating all keys related to an entity (user sessions, user preferences, user cart). **When not to use:** Frequent invalidation of large key sets -- tag-based invalidation is more efficient.
      
      ---
      
      ## Stampede Prevention: Distributed Lock
      
      Only one request regenerates the cache; others wait or receive stale data. Uses atomic set-if-not-exists for lock acquisition.
      
      ```typescript
      // Good Example - Stampede prevention with distributed lock
      const LOCK_TTL_SECONDS = 10;
      const RETRY_DELAY_MS = 50;
      const MAX_RETRIES = 20;
      
      async function getWithLock<T>(
        key: string,
        fetchFn: () => Promise<T>,
        ttl: number,
      ): Promise<T> {
        // Try cache first
        const cached = await cacheStore.get(key);
        if (cached) return JSON.parse(cached) as T;
      
        // Attempt to acquire lock (atomic set-if-not-exists)
        const lockKey = `lock:${key}`;
        const acquired = await cacheStore.set(lockKey, "1", {
          ttl: LOCK_TTL_SECONDS,
          nx: true, // Only set if key does not exist
        });
      
        if (acquired) {
          // Lock holder: fetch data, populate cache, release lock
          try {
            const data = await fetchFn();
            await cacheStore.set(key, JSON.stringify(data), { ttl });
            return data;
          } finally {
            await cacheStore.del(lockKey);
          }
        }
      
        // Non-holder: wait and retry
        for (let i = 0; i < MAX_RETRIES; i++) {
          await new Promise((resolve) => setTimeout(resolve, RETRY_DELAY_MS));
          const retryResult = await cacheStore.get(key);
          if (retryResult) return JSON.parse(retryResult) as T;
        }
      
        // Fallback: lock holder may have failed -- fetch directly
        return fetchFn();
      }
      
      export { getWithLock };
      ```
      
      **Why good:** Lock has TTL so it auto-releases if holder crashes, atomic `nx` prevents race conditions, fallback to direct fetch prevents permanent stalls, bounded retries prevent infinite waits
      
      ---
      
      ## Stampede Prevention: Request Coalescing (Singleflight)
      
      In-process deduplication -- concurrent requests for the same key share a single fetch instead of each making their own.
      
      ```typescript
      // Good Example - In-process request coalescing (singleflight pattern)
      const inFlightRequests = new Map<string, Promise<unknown>>();
      
      async function coalesce<T>(key: string, fetchFn: () => Promise<T>): Promise<T> {
        const existing = inFlightRequests.get(key);
        if (existing) return existing as Promise<T>;
      
        const promise = fetchFn().finally(() => {
          inFlightRequests.delete(key);
        });
      
        inFlightRequests.set(key, promise);
        return promise;
      }
      
      // Usage with cache-aside
      const PRODUCT_TTL = 3600;
      
      async function getProduct(id: string): Promise<Product> {
        const cacheKey = `app:product:${id}`;
        const cached = await cacheStore.get(cacheKey);
        if (cached) return JSON.parse(cached) as Product;
      
        // All concurrent requests for the same product share one fetch
        const product = await coalesce(cacheKey, async () => {
          const result = await db.query.products.findFirst({
            where: eq(products.id, id),
          });
          if (!result) throw new Error(`Product ${id} not found`);
      
          await cacheStore.set(cacheKey, JSON.stringify(result), {
            ttl: PRODUCT_TTL,
          });
          return result;
        });
      
        return product;
      }
      
      export { coalesce };
      ```
      
      **Why good:** Zero external dependencies (in-process Map), concurrent requests share one fetch (N requests = 1 database query), promise cleaned up automatically via `finally`, simpler than distributed locking
      
      **When to use:** Single-process applications or when stampedes happen within a single instance. **When not to use:** Multi-instance deployments where the stampede spans across processes (use distributed locking instead).
      
      ---
      
      ## Stampede Prevention: Probabilistic Early Recomputation
      
      Proactively refresh cache entries before they expire. Each access has an increasing probability of triggering background refresh as the TTL approaches expiration.
      
      ```typescript
      // Good Example - Probabilistic early recomputation
      const EARLY_RECOMPUTE_FACTOR = 0.1; // Start refreshing at 10% remaining TTL
      
      interface CachedEntry<T> {
        data: T;
        expiresAt: number; // Unix timestamp in ms
        ttlMs: number; // Original TTL for probability calculation
      }
      
      async function getWithEarlyRefresh<T>(
        key: string,
        fetchFn: () => Promise<T>,
        ttlMs: number,
      ): Promise<T> {
        const raw = await cacheStore.get(key);
      
        if (raw) {
          const entry = JSON.parse(raw) as CachedEntry<T>;
          const remainingMs = entry.expiresAt - Date.now();
          const remainingFraction = remainingMs / entry.ttlMs;
      
          // Probability of refresh increases as TTL approaches expiration
          if (
            remainingFraction < EARLY_RECOMPUTE_FACTOR &&
            Math.random() > remainingFraction
          ) {
            // Background refresh -- don't await, serve stale data immediately
            refreshInBackground(key, fetchFn, ttlMs);
          }
      
          return entry.data;
        }
      
        // Cache miss -- fetch synchronously
        return fetchAndStore(key, fetchFn, ttlMs);
      }
      
      async function fetchAndStore<T>(
        key: string,
        fetchFn: () => Promise<T>,
        ttlMs: number,
      ): Promise<T> {
        const data = await fetchFn();
        const entry: CachedEntry<T> = {
          data,
          expiresAt: Date.now() + ttlMs,
          ttlMs,
        };
        const ttlSeconds = Math.ceil(ttlMs / 1000);
        await cacheStore.set(key, JSON.stringify(entry), { ttl: ttlSeconds });
        return data;
      }
      
      function refreshInBackground<T>(
        key: string,
        fetchFn: () => Promise<T>,
        ttlMs: number,
      ): void {
        fetchAndStore(key, fetchFn, ttlMs).catch((error) => {
          logger.warn("Background cache refresh failed", {
            key,
            error: getErrorMessage(error),
          });
        });
      }
      
      export { getWithEarlyRefresh };
      ```
      
      **Why good:** Responses are never delayed by cache regeneration (always serves cached data), probability-based approach distributes refresh across time (avoids all entries refreshing simultaneously), background failures don't affect current request
      
      **When to use:** High-traffic keys where even brief cache misses cause noticeable load spikes. **When not to use:** Low-traffic keys where the extra complexity is not justified.
      
      ---
      
      ## Write-Behind (Write-Back) Pattern
      
      Write to cache immediately, persist to database asynchronously. Improves write latency at the cost of temporary inconsistency and data loss risk.
      
      ```typescript
      // Good Example - Write-behind with buffered persistence
      const FLUSH_INTERVAL_MS = 5_000;
      const MAX_BUFFER_SIZE = 100;
      
      interface PendingWrite {
        key: string;
        value: unknown;
        timestamp: number;
      }
      
      class WriteBuffer {
        private buffer: PendingWrite[] = [];
        private timer: NodeJS.Timeout | null = null;
      
        constructor(
          private readonly persistFn: (writes: PendingWrite[]) => Promise<void>,
        ) {}
      
        add(key: string, value: unknown): void {
          this.buffer.push({ key, value, timestamp: Date.now() });
      
          if (this.buffer.length >= MAX_BUFFER_SIZE) {
            void this.flush();
          } else if (!this.timer) {
            this.timer = setTimeout(() => void this.flush(), FLUSH_INTERVAL_MS);
          }
        }
      
        async flush(): Promise<void> {
          if (this.timer) {
            clearTimeout(this.timer);
            this.timer = null;
          }
      
          if (this.buffer.length === 0) return;
      
          const writes = [...this.buffer];
          this.buffer = [];
      
          try {
            await this.persistFn(writes);
          } catch (error) {
            // Re-queue failed writes for retry
            logger.error("Write-behind flush failed, re-queuing", {
              count: writes.length,
              error: getErrorMessage(error),
            });
            this.buffer.unshift(...writes);
          }
        }
      
        /** Call on process shutdown to persist remaining writes */
        async shutdown(): Promise<void> {
          await this.flush();
        }
      }
      
      // Usage
      const writeBuffer = new WriteBuffer(async (writes) => {
        // Batch insert/update in a single transaction
        await db.transaction(async (tx) => {
          for (const write of writes) {
            await tx.insert(analytics).values(write.value as AnalyticsEvent);
          }
        });
      });
      
      // Write is instant -- only updates cache, buffers the database write
      async function trackEvent(event: AnalyticsEvent): Promise<void> {
        const cacheKey = `app:event:${event.id}`;
        await cacheStore.set(cacheKey, JSON.stringify(event), { ttl: 3600 });
        writeBuffer.add(cacheKey, event);
      }
      
      // Register shutdown handler
      process.on("SIGTERM", async () => {
        await writeBuffer.shutdown();
        process.exit(0);
      });
      
      export { WriteBuffer, trackEvent };
      ```
      
      **Why good:** Write latency is minimal (only cache write), batched persistence reduces database round-trips, failed writes are re-queued, shutdown handler flushes remaining buffer
      
      **When to use:** Analytics, logging, activity tracking -- write-heavy workloads where occasional data loss is acceptable. **When not to use:** Financial transactions, user data mutations, or anything requiring write durability guarantees.
      
      ---
      
      ## Cache Hit/Miss Monitoring
      
      Track cache performance to verify caching is actually helping.
      
      ```typescript
      // Good Example - Cache metrics wrapper
      interface CacheMetrics {
        hits: number;
        misses: number;
        errors: number;
      }
      
      const metrics = new Map<string, CacheMetrics>();
      
      function getMetrics(prefix: string): CacheMetrics {
        if (!metrics.has(prefix)) {
          metrics.set(prefix, { hits: 0, misses: 0, errors: 0 });
        }
        return metrics.get(prefix)!;
      }
      
      async function getCachedWithMetrics<T>(
        key: string,
        prefix: string,
        fetchFn: () => Promise<T>,
        ttl: number,
      ): Promise<T> {
        const m = getMetrics(prefix);
      
        try {
          const cached = await cacheStore.get(key);
          if (cached) {
            m.hits++;
            return JSON.parse(cached) as T;
          }
        } catch {
          m.errors++;
        }
      
        m.misses++;
        const data = await fetchFn();
        await cacheStore.set(key, JSON.stringify(data), { ttl });
        return data;
      }
      
      function getCacheStats(): Record<string, CacheMetrics & { hitRate: string }> {
        const stats: Record<string, CacheMetrics & { hitRate: string }> = {};
        for (const [prefix, m] of metrics) {
          const total = m.hits + m.misses;
          const hitRate =
            total > 0 ? `${((m.hits / total) * 100).toFixed(1)}%` : "N/A";
          stats[prefix] = { ...m, hitRate };
        }
        return stats;
      }
      
      export { getCachedWithMetrics, getCacheStats };
      ```
      
      **Why good:** Per-prefix metrics show which caches are effective, hit rate calculation provides actionable data, error tracking surfaces cache connection issues
      
      **Target cache hit rates:**
      
      | Cache Type      | Good Hit Rate | Action if Below                              |
      | --------------- | ------------- | -------------------------------------------- |
      | User profile    | > 80%         | Increase TTL or check invalidation frequency |
      | Product catalog | > 90%         | Cache may be too small (increase max items)  |
      | Search results  | > 60%         | Expected lower rate due to query diversity   |
      
      ---
      
      ## See Also
      
      - [core.md](core.md) - Cache-aside, HTTP caching, TTL patterns, in-memory LRU caching
      
    • core.md 12.5 KB
      # Caching Strategies - Core Examples
      
      > Cache-aside, write-through, HTTP caching headers, TTL patterns, in-memory LRU caching. See [advanced.md](advanced.md) for invalidation, stampede prevention, and distributed caching patterns.
      
      ---
      
      ## Cache-Aside with Generic Wrapper
      
      A reusable cache-aside wrapper that handles cache miss, fetch, store, and TTL in one function.
      
      ```typescript
      // Good Example - Generic cache-aside wrapper
      const DEFAULT_TTL_SECONDS = 300;
      
      async function cacheable<T>(
        cacheKey: string,
        fetchFn: () => Promise<T>,
        ttlSeconds: number = DEFAULT_TTL_SECONDS,
      ): Promise<T> {
        const cached = await cacheStore.get(cacheKey);
        if (cached) return JSON.parse(cached) as T;
      
        const data = await fetchFn();
        await cacheStore.set(cacheKey, JSON.stringify(data), { ttl: ttlSeconds });
        return data;
      }
      
      // Usage
      const PRODUCT_TTL = 3600;
      const PRODUCT_PREFIX = "app:product";
      
      async function getProduct(id: string): Promise<Product | null> {
        return cacheable<Product | null>(
          `${PRODUCT_PREFIX}:${id}`,
          () => db.query.products.findFirst({ where: eq(products.id, id) }),
          PRODUCT_TTL,
        );
      }
      
      export { cacheable, getProduct };
      ```
      
      **Why good:** Separates caching concern from business logic, configurable TTL per call, generic works with any data type, namespaced keys
      
      ```typescript
      // Bad Example - Cache logic mixed with business logic, no TTL
      async function getProduct(id: string) {
        try {
          const cached = await cacheStore.get(id);
          if (cached) return JSON.parse(cached);
        } catch {
          // Silently swallow -- cache failure hides bugs
        }
      
        const product = await db.query.products.findFirst({
          where: eq(products.id, id),
        });
        await cacheStore.set(id, JSON.stringify(product)); // No TTL
        return product;
      }
      ```
      
      **Why bad:** No TTL means data never expires, generic key collides with other entity types, silently swallowing cache errors hides connection problems, caching logic mixed with data access
      
      ---
      
      ## Cache-Aside with Error Handling
      
      When the cache store is unavailable, the application should still work -- fall through to the data source.
      
      ```typescript
      // Good Example - Cache failure falls through to data source
      const USER_TTL = 300;
      const USER_PREFIX = "app:user";
      
      async function getUserById(userId: string): Promise<User | null> {
        const cacheKey = `${USER_PREFIX}:${userId}`;
      
        // Cache read failure is non-fatal -- fall through to database
        try {
          const cached = await cacheStore.get(cacheKey);
          if (cached) return JSON.parse(cached) as User;
        } catch (error) {
          logger.warn("Cache read failed, falling through to database", {
            key: cacheKey,
            error: getErrorMessage(error),
          });
        }
      
        const user = await db.query.users.findFirst({ where: eq(users.id, userId) });
        if (!user) return null;
      
        // Cache write failure is non-fatal -- data was still fetched successfully
        try {
          await cacheStore.set(cacheKey, JSON.stringify(user), { ttl: USER_TTL });
        } catch (error) {
          logger.warn("Cache write failed", {
            key: cacheKey,
            error: getErrorMessage(error),
          });
        }
      
        return user;
      }
      ```
      
      **Why good:** Application degrades gracefully when cache is down, cache errors are logged (not silently swallowed), database is the source of truth
      
      ---
      
      ## Write-Through with Delete on Remove
      
      ```typescript
      // Good Example - Write-through: update cache on mutation, delete on remove
      const CACHE_TTL = 300;
      const USER_PREFIX = "app:user";
      
      async function updateUser(
        userId: string,
        updates: Partial<User>,
      ): Promise<User> {
        // Update database first (source of truth)
        const [updatedUser] = await db
          .update(users)
          .set({ ...updates, updatedAt: new Date() })
          .where(eq(users.id, userId))
          .returning();
      
        // Write-through: update cache immediately
        const cacheKey = `${USER_PREFIX}:${userId}`;
        await cacheStore.set(cacheKey, JSON.stringify(updatedUser), {
          ttl: CACHE_TTL,
        });
      
        return updatedUser;
      }
      
      async function deleteUser(userId: string): Promise<void> {
        await db.delete(users).where(eq(users.id, userId));
      
        // Invalidate cache on delete
        const cacheKey = `${USER_PREFIX}:${userId}`;
        await cacheStore.del(cacheKey);
      }
      
      export { updateUser, deleteUser };
      ```
      
      **Why good:** Cache always reflects latest database state after writes, explicit invalidation on delete prevents stale reads, TTL still set as safety net
      
      ```typescript
      // Bad Example - Update database but forget to update cache
      async function updateUser(userId: string, updates: Partial<User>) {
        await db.update(users).set(updates).where(eq(users.id, userId));
        // Cache still has old data until TTL expires!
      }
      ```
      
      **Why bad:** Stale cache serves outdated data until TTL expires, users see old values after saving changes
      
      ---
      
      ## HTTP Caching: Cache-Control Headers
      
      Setting appropriate Cache-Control headers on API responses reduces requests to your server entirely.
      
      ```typescript
      // Good Example - Cache-Control for different API endpoint types
      
      // Public list endpoint -- CDN can cache, browser caches briefly
      const LIST_MAX_AGE = 60;
      const LIST_S_MAXAGE = 300;
      const LIST_SWR = 60;
      
      function setPublicListHeaders(res: Response): void {
        res.setHeader(
          "Cache-Control",
          `public, max-age=${LIST_MAX_AGE}, s-maxage=${LIST_S_MAXAGE}, stale-while-revalidate=${LIST_SWR}`,
        );
      }
      
      // User-specific endpoint -- private cache only, always revalidate
      function setPrivateHeaders(res: Response): void {
        res.setHeader("Cache-Control", "private, no-cache");
      }
      
      // Sensitive data -- never cache
      function setNoCacheHeaders(res: Response): void {
        res.setHeader("Cache-Control", "no-store, private");
      }
      ```
      
      **Why good:** Different endpoints get appropriate caching strategies, `s-maxage` lets CDN cache longer than browser, `stale-while-revalidate` improves availability during refresh
      
      ---
      
      ## HTTP Caching: ETag and Conditional Requests
      
      ETags enable conditional requests -- the server can respond with 304 Not Modified when data hasn't changed, saving bandwidth and serialization cost.
      
      ```typescript
      // Good Example - ETag generation and conditional request handling
      import { createHash } from "node:crypto";
      
      function generateETag(body: unknown): string {
        const content = JSON.stringify(body);
        const hash = createHash("md5").update(content).digest("hex");
        return `"${hash}"`;
      }
      
      // Middleware or route handler for conditional responses
      async function handleConditionalGet(
        req: Request,
        res: Response,
        body: unknown,
      ): Promise<boolean> {
        const etag = generateETag(body);
        res.setHeader("ETag", etag);
      
        const clientETag = req.headers["if-none-match"];
        if (clientETag === etag) {
          res.status(304).end();
          return true; // Response already sent
        }
      
        return false; // Caller should send the full response
      }
      
      // Usage in a route handler
      async function getProducts(req: Request, res: Response) {
        const products = await db.query.products.findMany();
      
        const notModified = await handleConditionalGet(req, res, products);
        if (notModified) return;
      
        res.setHeader("Cache-Control", "public, max-age=60, must-revalidate");
        res.json(products);
      }
      
      export { generateETag, handleConditionalGet };
      ```
      
      **Why good:** 304 responses save bandwidth (no body), ETag based on content hash detects actual changes, separates conditional logic from route handler
      
      **When to use:** Endpoints where the response body is expensive to transfer but cheap to check for changes (product listings, configuration data).
      
      ---
      
      ## CDN Caching with s-maxage
      
      Use `s-maxage` to let CDNs (shared caches) cache responses longer than the browser. Combine with `stale-while-revalidate` for high availability.
      
      ```typescript
      // Good Example - CDN-optimized cache headers
      const BROWSER_MAX_AGE = 60; // Browser: 1 minute
      const CDN_MAX_AGE = 300; // CDN: 5 minutes
      const SWR_WINDOW = 60; // Serve stale for 60s while refreshing
      const ERROR_WINDOW = 600; // Serve stale for 10m during origin errors
      
      function setCDNCacheHeaders(res: Response): void {
        res.setHeader(
          "Cache-Control",
          [
            "public",
            `max-age=${BROWSER_MAX_AGE}`,
            `s-maxage=${CDN_MAX_AGE}`,
            `stale-while-revalidate=${SWR_WINDOW}`,
            `stale-if-error=${ERROR_WINDOW}`,
          ].join(", "),
        );
      }
      ```
      
      **Why good:** `s-maxage` overrides `max-age` for CDNs only, `stale-while-revalidate` ensures users never wait for origin refresh, `stale-if-error` provides resilience during origin outages
      
      ---
      
      ## In-Memory LRU Cache
      
      For single-process hot data. Faster than a distributed cache (no network hop) but not shared across instances.
      
      ```typescript
      // Good Example - LRU cache for hot data (lru-cache v11+)
      import { LRUCache } from "lru-cache";
      
      const MAX_ITEMS = 500;
      const TTL_MS = 300_000; // 5 minutes
      
      interface CachedConfig {
        value: string;
        updatedAt: Date;
      }
      
      const configCache = new LRUCache<string, CachedConfig>({
        max: MAX_ITEMS,
        ttl: TTL_MS,
      });
      
      function getCachedConfig(key: string): CachedConfig | undefined {
        return configCache.get(key);
      }
      
      function setCachedConfig(key: string, config: CachedConfig): void {
        configCache.set(key, config);
      }
      
      export { getCachedConfig, setCachedConfig };
      ```
      
      **Why good:** No serialization overhead (stores objects directly), automatic LRU eviction when max reached, TTL prevents staleness, type-safe with generics
      
      ```typescript
      // Bad Example - Unbounded Map as cache
      const cache = new Map<string, unknown>(); // No max size, no TTL
      
      function getCached(key: string) {
        return cache.get(key); // Never evicted, never expires
      }
      
      function setCached(key: string, value: unknown) {
        cache.set(key, value); // Memory grows without bound
      }
      ```
      
      **Why bad:** No size limit means memory grows unbounded, no TTL means data is stale forever, no eviction policy means the cache fills up and never frees memory
      
      ---
      
      ## LRU Cache with Fetch Method
      
      The `fetch` option in lru-cache provides built-in cache-aside behavior -- on a miss, it calls your fetch function automatically.
      
      ```typescript
      // Good Example - LRU cache with automatic fetch on miss
      import { LRUCache } from "lru-cache";
      
      const MAX_ITEMS = 200;
      const TTL_MS = 60_000; // 1 minute
      
      const featureFlagCache = new LRUCache<string, boolean>({
        max: MAX_ITEMS,
        ttl: TTL_MS,
        // Called automatically on cache miss
        fetchMethod: async (key) => {
          const flag = await db.query.featureFlags.findFirst({
            where: eq(featureFlags.key, key),
          });
          return flag?.enabled ?? false;
        },
      });
      
      // Usage: fetch() returns cached value or calls fetchMethod
      async function isFeatureEnabled(flagKey: string): Promise<boolean> {
        const result = await featureFlagCache.fetch(flagKey);
        return result ?? false;
      }
      
      export { isFeatureEnabled };
      ```
      
      **Why good:** Cache-aside logic is built into the cache itself, no manual get/set coordination, fetch is deduplicated (concurrent calls for the same key share one fetch)
      
      ---
      
      ## TTL Strategy by Data Type
      
      ```typescript
      // Named constants for TTL values -- document the rationale
      const TTL = {
        /** Static config that changes via deploy */
        STATIC_CONFIG: 3600,
      
        /** Product catalog updated by admin */
        PRODUCT: 300,
      
        /** User profile updated by user */
        USER_PROFILE: 60,
      
        /** Search results -- balance freshness vs performance */
        SEARCH: 30,
      
        /** Real-time data -- near-fresh */
        REALTIME: 10,
      
        /** Session data -- long-lived by design */
        SESSION: 86400,
      } as const;
      
      export { TTL };
      ```
      
      **Why good:** Named constants with comments make TTL decisions explicit, `as const` preserves literal types, centralized TTL values prevent inconsistency across the codebase
      
      ---
      
      ## Cache Key Generation
      
      ```typescript
      // Good Example - Structured cache key generation
      import { createHash } from "node:crypto";
      
      const CACHE_PREFIX = "myapp";
      
      // Entity keys: prefix:type:id
      const cacheKeys = {
        user: (id: string) => `${CACHE_PREFIX}:user:${id}`,
        product: (id: string) => `${CACHE_PREFIX}:product:${id}`,
        cart: (userId: string) => `${CACHE_PREFIX}:cart:${userId}`,
      
        // Query result keys: hash the normalized query params
        productList: (filters: Record<string, unknown>) => {
          // Sort keys for deterministic hashing
          const sorted = JSON.stringify(filters, Object.keys(filters).sort());
          const hash = createHash("md5").update(sorted).digest("hex");
          return `${CACHE_PREFIX}:products:list:${hash}`;
        },
      
        // Pattern for bulk invalidation
        userPattern: (userId: string) => `${CACHE_PREFIX}:user:${userId}:*`,
      } as const;
      
      export { cacheKeys, CACHE_PREFIX };
      ```
      
      **Why good:** Consistent prefix prevents cross-application collisions, hierarchical keys enable pattern-based invalidation, sorted JSON ensures deterministic hashing regardless of property insertion order
      
      ---
      
      ## See Also
      
      - [advanced.md](advanced.md) - Cache invalidation, stampede prevention, distributed caching
      
  • reference.md 3.7 KB
    # Caching Strategies Quick Reference
    
    Lookup tables and comparisons. See [SKILL.md](SKILL.md) for decision frameworks, red flags, and philosophy.
    
    ---
    
    ## Cache-Control Quick Reference
    
    | Use Case                | Header                                                        | Behavior                            |
    | ----------------------- | ------------------------------------------------------------- | ----------------------------------- |
    | Public list data        | `public, max-age=60, s-maxage=300`                            | Browser: 1m, CDN: 5m                |
    | Public with SWR         | `public, max-age=60, s-maxage=300, stale-while-revalidate=60` | CDN serves stale during refresh     |
    | User-specific data      | `private, no-cache`                                           | Always revalidate, no shared cache  |
    | Sensitive data          | `no-store, private`                                           | Never cached anywhere               |
    | Versioned static assets | `public, max-age=31536000, immutable`                         | Cache forever, version in URL       |
    | API with ETag           | `max-age=0, must-revalidate`                                  | Always revalidate, 304 if unchanged |
    
    ---
    
    ## Conditional Request Flow
    
    ```
    Client sends:    If-None-Match: "abc123"
                     If-Modified-Since: Tue, 01 Jan 2025 00:00:00 GMT
    
    Server checks:  ETag matches? (takes precedence per RFC 9110)
    +-- YES --> 304 Not Modified (no body)
    +-- NO  --> Last-Modified changed?
        +-- NO  --> 304 Not Modified (no body)
        +-- YES --> 200 OK (full response)
    ```
    
    ---
    
    ## Strategy Comparison
    
    | Strategy      | Read Perf            | Write Perf          | Consistency      | Complexity | Data Loss Risk        |
    | ------------- | -------------------- | ------------------- | ---------------- | ---------- | --------------------- |
    | Cache-aside   | Fast (on hit)        | N/A (reads only)    | Eventual (TTL)   | Low        | None                  |
    | Write-through | Fast (on hit)        | Slower (dual write) | Strong           | Low        | None                  |
    | Write-behind  | Fast (on hit)        | Fast (async)        | Eventual         | High       | Yes (buffer loss)     |
    | HTTP caching  | Fastest (no request) | N/A                 | Eventual (TTL)   | Low        | None                  |
    | In-memory LRU | Sub-ms               | Sub-ms              | Per-process only | Low        | Yes (process restart) |
    
    ---
    
    ## Stampede Prevention Comparison
    
    | Technique                   | When to Use                          | Trade-off                          |
    | --------------------------- | ------------------------------------ | ---------------------------------- |
    | Distributed lock            | Multi-instance, critical keys        | Lock overhead, potential stalls    |
    | Request coalescing          | Single-instance, bursty traffic      | In-process only, no cross-instance |
    | Probabilistic early refresh | High-traffic, predictable patterns   | Slightly higher cache store load   |
    | Background refresh (cron)   | Periodic data, known update schedule | Stale between refreshes            |
    
    ---
    
    ## Cache Key Conventions
    
    ```
    {app}:{entity}:{id}              -- Entity by ID
    {app}:{entity}:list:{hash}       -- Query result by hashed params
    {app}:{entity}:{id}:{subentity}  -- Related entity
    tag:{category}                   -- Tag set for bulk invalidation
    lock:{key}                       -- Stampede prevention lock
    ```
    
    **Rules:**
    
    - Always prefix with application name (prevents cross-app collisions)
    - Use colons as separators (convention for hierarchical keys)
    - Hash complex query params for shorter keys (MD5 is fine for cache keys)
    - Sort object keys before hashing (ensures deterministic output)
    
  • SKILL.md 17.3 KB
    ---
    name: api-caching-strategies
    description: Application-level caching strategies, HTTP caching, cache invalidation, and stampede prevention
    ---
    
    # Caching Strategies
    
    > **Quick Guide:** Choose the right caching strategy for your use case: cache-aside for read-heavy data, write-through for consistency, write-behind for write-heavy workloads. Always set TTL to prevent stale data and memory exhaustion. Use HTTP caching headers (Cache-Control, ETag, Last-Modified) for API responses. Prevent cache stampedes with locking or request coalescing. Measure cache hit rates before and after -- caching without metrics is guessing.
    
    ---
    
    <critical_requirements>
    
    ## CRITICAL: Before Using This Skill
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST set TTL on ALL cached data -- cache without TTL leads to memory exhaustion and infinitely stale data)**
    
    **(You MUST use namespaced cache keys with a consistent prefix -- generic keys cause collisions across data types)**
    
    **(You MUST invalidate or update cache entries on writes -- serving stale data after mutation breaks user trust)**
    
    **(You MUST implement stampede prevention (locking or coalescing) for high-traffic cache keys -- concurrent misses can overwhelm your data source)**
    
    </critical_requirements>
    
    ---
    
    **Auto-detection:** caching, cache-aside, write-through, write-behind, cache invalidation, TTL, Cache-Control, ETag, Last-Modified, stale-while-revalidate, s-maxage, cache stampede, thundering herd, in-memory cache, LRU cache, distributed cache, cache key, cache miss, cache hit, CDN caching, HTTP caching, conditional request, 304 Not Modified
    
    **When to use:**
    
    - Read-heavy endpoints fetching the same data repeatedly (cache-aside)
    - Data that must stay consistent between cache and database after writes (write-through)
    - API responses that benefit from HTTP caching headers (Cache-Control, ETag)
    - High-traffic cache keys that risk stampede on expiration
    - Reducing database load by caching expensive query results
    
    **When NOT to use:**
    
    - Data that must always be real-time fresh (caching adds staleness by definition)
    - Simple CRUD with low traffic (caching complexity outweighs benefit)
    - Development/debugging (caching obscures issues -- disable in dev)
    - Premature optimization without measuring actual bottlenecks first
    
    **Key patterns covered:**
    
    - Cache-aside (lazy loading) with TTL
    - Write-through for read-after-write consistency
    - Write-behind (write-back) for write-heavy workloads
    - HTTP caching: Cache-Control, ETag, Last-Modified, conditional requests
    - CDN caching with s-maxage and stale-while-revalidate
    - In-memory LRU caching for single-process hot data
    - Cache key strategies and namespacing
    - Stampede prevention (locking, request coalescing, early recomputation)
    - Tag-based and pattern-based invalidation
    
    ---
    
    <philosophy>
    
    ## Philosophy
    
    Caching trades freshness for speed. Every caching decision is a **consistency vs performance** trade-off -- understand where your use case falls on that spectrum before choosing a strategy.
    
    **The three questions before adding caching:**
    
    1. **Is this actually slow?** Measure first. If the uncached response is fast enough, caching adds complexity without benefit.
    2. **Can I tolerate staleness?** If data must always be real-time, caching is the wrong tool. Use read replicas or materialized views instead.
    3. **What is the read-to-write ratio?** Caching shines when reads vastly outnumber writes. For write-heavy workloads, consider write-behind or skip caching entirely.
    
    **Caching is a layered system.** HTTP caching (browser and CDN) reduces requests to your server. Application-level caching (in-memory or distributed) reduces requests to your database. Apply caching at the right layer for the problem.
    
    **When to use caching:**
    
    - Response times exceed acceptable thresholds and the data source is the bottleneck
    - The same data is fetched repeatedly across requests
    - Data freshness requirements allow some staleness (even 60 seconds)
    - Traffic is high enough that database load is a concern
    
    **When NOT to use caching:**
    
    - Data changes frequently and must always be current
    - Every request returns unique data (no cache reuse)
    - You have not measured the actual bottleneck yet (premature optimization)
    
    </philosophy>
    
    ---
    
    <patterns>
    
    ## Core Patterns
    
    ### Pattern 1: Cache-Aside (Lazy Loading)
    
    The most common application-level caching pattern. The application checks the cache first, fetches from the data source on miss, and stores the result with a TTL.
    
    ```typescript
    const CACHE_TTL_SECONDS = 300;
    const CACHE_PREFIX = "app:user";
    
    async function getUserById(userId: string): Promise<User | null> {
      const cacheKey = `${CACHE_PREFIX}:${userId}`;
      const cached = await cacheStore.get(cacheKey);
      if (cached) return JSON.parse(cached) as User;
    
      const user = await db.query.users.findFirst({ where: eq(users.id, userId) });
      if (!user) return null;
    
      await cacheStore.set(cacheKey, JSON.stringify(user), {
        ttl: CACHE_TTL_SECONDS,
      });
      return user;
    }
    ```
    
    **Why good:** TTL prevents stale data, namespaced keys prevent collisions, early return on cache hit, cache is populated lazily (only data that is actually requested gets cached)
    
    ```typescript
    // BAD: No TTL, generic key
    async function getUser(id: string) {
      const cached = await cacheStore.get(id); // No prefix -- collides with other data types
      if (cached) return JSON.parse(cached);
      const user = await db.query.users.findFirst({ where: eq(users.id, id) });
      await cacheStore.set(id, JSON.stringify(user)); // No TTL -- never expires
      return user;
    }
    ```
    
    **Why bad:** No TTL means infinite staleness and eventual memory exhaustion, generic key collides with other entity types using the same ID format
    
    **When to use:** Read-heavy endpoints where data changes infrequently relative to reads.
    
    See [examples/core.md](examples/core.md) for generic cache wrapper and error handling patterns.
    
    ---
    
    ### Pattern 2: Write-Through
    
    Update both cache and data source on every write. Guarantees read-after-write consistency without waiting for TTL expiration.
    
    ```typescript
    const CACHE_TTL_SECONDS = 300;
    
    async function updateUser(
      userId: string,
      updates: Partial<User>,
    ): Promise<User> {
      const updatedUser = await db
        .update(users)
        .set({ ...updates, updatedAt: new Date() })
        .where(eq(users.id, userId))
        .returning();
    
      const cacheKey = `${CACHE_PREFIX}:${userId}`;
      await cacheStore.set(cacheKey, JSON.stringify(updatedUser[0]), {
        ttl: CACHE_TTL_SECONDS,
      });
    
      return updatedUser[0];
    }
    ```
    
    **Why good:** Cache always reflects latest database state, no stale reads after updates
    
    **When to use:** Data that is read frequently after writes and must be consistent. **When not to use:** Write-heavy workloads where the overhead of updating cache on every write is too expensive.
    
    See [examples/core.md](examples/core.md) for write-through with delete invalidation.
    
    ---
    
    ### Pattern 3: HTTP Caching Headers
    
    Set appropriate Cache-Control, ETag, and Last-Modified headers on API responses to leverage browser and CDN caching. This reduces requests to your server entirely.
    
    ```typescript
    // API endpoint setting caching headers
    function setCacheHeaders(
      res: Response,
      body: unknown,
      options: { maxAge: number; isPublic: boolean },
    ) {
      const etag = generateETag(body);
      const scope = options.isPublic ? "public" : "private";
    
      res.setHeader(
        "Cache-Control",
        `${scope}, max-age=${options.maxAge}, must-revalidate`,
      );
      res.setHeader("ETag", etag);
      res.setHeader("Last-Modified", new Date().toUTCString());
    }
    ```
    
    **Key header combinations for APIs:**
    
    | Use Case           | Cache-Control                         | Why                                |
    | ------------------ | ------------------------------------- | ---------------------------------- |
    | Public list data   | `public, max-age=60, s-maxage=300`    | CDN caches longer than browser     |
    | User-specific data | `private, max-age=0, must-revalidate` | No shared cache, always revalidate |
    | Sensitive data     | `no-store, private`                   | Never cache anywhere               |
    | Immutable assets   | `public, max-age=31536000, immutable` | Version in URL, cache forever      |
    
    See [examples/core.md](examples/core.md) for conditional request handling (If-None-Match / 304) and CDN patterns.
    
    ---
    
    ### Pattern 4: In-Memory LRU Cache
    
    For single-process hot data that does not need to be shared across instances. Faster than a distributed cache (no network hop) but lost on process restart and not shared between workers.
    
    ```typescript
    import { LRUCache } from "lru-cache";
    
    const MAX_ITEMS = 500;
    const TTL_MS = 300_000; // 5 minutes
    
    const configCache = new LRUCache<string, AppConfig>({
      max: MAX_ITEMS,
      ttl: TTL_MS,
    });
    
    function getConfig(key: string): AppConfig | undefined {
      return configCache.get(key);
    }
    
    function setConfig(key: string, value: AppConfig): void {
      configCache.set(key, value);
    }
    ```
    
    **Why good:** No serialization overhead (stores objects directly), automatic eviction of least-recently-used items, TTL prevents staleness
    
    **When to use:** Configuration, feature flags, frequently accessed reference data in single-process applications. **When not to use:** Data that must be shared across multiple server instances (use a distributed cache instead).
    
    See [examples/core.md](examples/core.md) for LRU cache with size tracking and fetch method.
    
    ---
    
    ### Pattern 5: Cache Key Strategies
    
    Consistent, namespaced cache keys prevent collisions and enable pattern-based invalidation.
    
    ```typescript
    const CACHE_PREFIX = "myapp";
    
    // Entity keys: prefix:type:id
    const userKey = (id: string) => `${CACHE_PREFIX}:user:${id}`;
    const productKey = (id: string) => `${CACHE_PREFIX}:product:${id}`;
    
    // Query result keys: prefix:type:list:hash
    const productListKey = (filters: ProductFilters) => {
      const normalized = JSON.stringify({
        category: filters.category ?? "all",
        page: filters.page ?? 1,
        sort: filters.sort ?? "created",
      });
      const hash = createHash("md5").update(normalized).digest("hex");
      return `${CACHE_PREFIX}:products:list:${hash}`;
    };
    
    // User-scoped keys
    const userCartKey = (userId: string) => `${CACHE_PREFIX}:cart:${userId}`;
    ```
    
    **Why good:** Hierarchical structure prevents collisions, hashing complex queries keeps keys short, pattern prefix enables bulk invalidation
    
    See [examples/core.md](examples/core.md) for cache key generation patterns.
    
    ---
    
    ### Pattern 6: Cache Invalidation
    
    Invalidation is the hardest part of caching. Choose the simplest strategy that meets your consistency requirements.
    
    **Direct invalidation** -- delete the exact key on mutation:
    
    ```typescript
    async function updateProduct(id: string, data: ProductUpdate) {
      await db.update(products).set(data).where(eq(products.id, id));
      await cacheStore.del(productKey(id));
    }
    ```
    
    **Tag-based invalidation** -- group related keys by tag for bulk invalidation:
    
    ```typescript
    // Invalidate all electronics products when category changes
    await invalidateByTag("category:electronics");
    ```
    
    **TTL-based expiration** -- let the cache expire naturally when eventual consistency is acceptable:
    
    ```typescript
    const SHORT_TTL = 60; // 1 minute for frequently changing data
    const MEDIUM_TTL = 300; // 5 minutes for user data
    const LONG_TTL = 3600; // 1 hour for reference data
    ```
    
    See [examples/advanced.md](examples/advanced.md) for tag-based invalidation implementation and pattern-based key deletion.
    
    ---
    
    ### Pattern 7: Stampede Prevention
    
    When a popular cache key expires, many concurrent requests may all miss the cache simultaneously and flood the data source. This is a cache stampede (thundering herd).
    
    **Locking** -- only one request regenerates the cache; others wait or get stale data:
    
    ```typescript
    async function getWithLock(
      key: string,
      fetchFn: () => Promise<string>,
    ): Promise<string> {
      const cached = await cacheStore.get(key);
      if (cached) return cached;
    
      const lockKey = `lock:${key}`;
      const acquired = await cacheStore.setNX(lockKey, "1", {
        ttl: LOCK_TTL_SECONDS,
      });
    
      if (!acquired) {
        await delay(RETRY_DELAY_MS);
        return getWithLock(key, fetchFn); // Retry after brief wait
      }
    
      try {
        const value = await fetchFn();
        await cacheStore.set(key, value, { ttl: CACHE_TTL_SECONDS });
        return value;
      } finally {
        await cacheStore.del(lockKey);
      }
    }
    ```
    
    **Why good:** Only one request hits the data source, others wait briefly, lock auto-expires if holder crashes
    
    See [examples/advanced.md](examples/advanced.md) for request coalescing (singleflight) and probabilistic early recomputation.
    
    </patterns>
    
    ---
    
    **Detailed Resources:**
    
    - [examples/core.md](examples/core.md) - Cache-aside, write-through, HTTP caching, TTL patterns, in-memory caching
    - [examples/advanced.md](examples/advanced.md) - Cache invalidation, stampede prevention, distributed caching, write-behind
    - [reference.md](reference.md) - Strategy comparison, stampede comparison, cache key conventions
    
    ---
    
    <decision_framework>
    
    ## Decision Framework
    
    ### Which Caching Strategy?
    
    ```
    Is data read much more than written?
    +-- YES --> Can you tolerate staleness?
    |   +-- YES --> Cache-aside with TTL
    |   +-- NO  --> Write-through (consistent reads)
    +-- NO  --> Is write latency critical?
        +-- YES --> Write-behind (async persist)
        +-- NO  --> Write-through or skip caching
    ```
    
    ### Which Cache Layer?
    
    ```
    Is the response public (same for all users)?
    +-- YES --> HTTP caching (Cache-Control: public, s-maxage)
    |          CDN handles it, your server may never see the request
    +-- NO  --> Is it shared across server instances?
        +-- YES --> Distributed cache (external store)
        +-- NO  --> Is it hot data in a single process?
            +-- YES --> In-memory LRU cache
            +-- NO  --> Cache-aside with distributed store
    ```
    
    ### TTL Selection Guide
    
    | Data Type                      | Recommended TTL   | Rationale                        |
    | ------------------------------ | ----------------- | -------------------------------- |
    | Static config / feature flags  | 3600s+ (1h+)      | Rarely changes                   |
    | Product catalog / listings     | 300-3600s (5m-1h) | Infrequent updates               |
    | User profile                   | 60-300s (1-5m)    | Changes occasionally             |
    | Search results                 | 30-60s (30s-1m)   | Balance freshness vs performance |
    | Real-time data (prices, stock) | 5-30s             | Must be near-fresh               |
    | Session data                   | 86400s (24h)      | Long-lived by design             |
    
    </decision_framework>
    
    ---
    
    <red_flags>
    
    ## RED FLAGS
    
    **High Priority Issues:**
    
    - Cache without TTL -- memory grows unbounded and data is stale forever
    - No cache invalidation on writes -- users see stale data after mutations, eroding trust
    - Generic cache keys without namespacing -- collisions return wrong data to wrong callers
    - No stampede prevention on high-traffic keys -- cache expiration triggers data source flood
    - Caching errors or null results without separate short TTL -- a transient failure gets cached and served for the full TTL
    
    **Medium Priority Issues:**
    
    - Same TTL for all data types -- different data has different freshness requirements
    - Caching user-specific data in shared caches (CDN) without `private` directive -- leaks personal data to other users
    - No cache hit/miss monitoring -- impossible to know if caching is actually helping
    - `no-cache` confused with `no-store` -- `no-cache` still stores but revalidates; `no-store` prevents any storage
    - Over-caching low-traffic endpoints -- adds complexity without meaningful performance gain
    
    **Gotchas & Edge Cases:**
    
    - `Cache-Control: no-cache` does NOT prevent caching -- it forces revalidation before every use. Use `no-store` to prevent storage entirely.
    - ETag takes precedence over Last-Modified during revalidation (per RFC 9110) -- always send both for best compatibility
    - `s-maxage` overrides `max-age` for shared caches (CDNs) only -- browsers ignore it
    - `stale-while-revalidate` lets CDNs serve expired content while refreshing in the background -- great for availability, but means users briefly see stale data
    - Distributed cache SET with TTL resets the expiration timer on every write -- calling SET again extends the TTL, which may keep stale data alive longer than expected
    - In-memory LRU caches are per-process -- multiple server instances have independent caches that can serve different data for the same key
    - `JSON.stringify` for cache serialization drops `undefined` values and converts `Date` objects to strings -- use a structured serializer if needed
    - Cache key generation must be deterministic -- object property order in `JSON.stringify` varies across engines; sort keys or use a canonical serializer
    
    </red_flags>
    
    ---
    
    <critical_reminders>
    
    ## CRITICAL REMINDERS
    
    > **All code must follow project conventions in CLAUDE.md**
    
    **(You MUST set TTL on ALL cached data -- cache without TTL leads to memory exhaustion and infinitely stale data)**
    
    **(You MUST use namespaced cache keys with a consistent prefix -- generic keys cause collisions across data types)**
    
    **(You MUST invalidate or update cache entries on writes -- serving stale data after mutation breaks user trust)**
    
    **(You MUST implement stampede prevention (locking or coalescing) for high-traffic cache keys -- concurrent misses can overwhelm your data source)**
    
    **Failure to follow these rules will cause stale data bugs, cache key collisions, memory exhaustion, and data source overload during traffic spikes.**
    
    </critical_reminders>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related