Claude Skill

cost-tracking

Use when metering and capping AI or cloud app spend — tokens read from the response `usage` object, priced off a dated rate table, ledgered per user/tenant/feature, with alerts and a hard cap before the bill. NOT cash runway (that is `finance-ops`), NOT cost-per-unit margin (that

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download ericrisco-rsc-harness-skills_cost-tracking-953fef5.zip · 12 KB
Part of ericrisco/rsc-harness — 46 skills

Install

skills CLI npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/cost-tracking
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
Git git clone https://github.com/ericrisco/rsc-harness.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Cost tracking

You are building the money meter for an AI or cloud app: every model call gets a token count, a price, a ledger row, and a budget check — so a cap fires before the invoice, not when finance forwards it in a panic. Model-API spend roughly doubled from $3.5B to $8.4B between late 2024 and mid 2025 (firecrawl.dev best-llm-observability-tools, accessed 2026-06-02); the bill is now big enough to need a guardrail, not a spreadsheet at month-end.

The chain you build, in order — each step feeds the next:

  1. Capture — read tokens from the response, not from a pre-send guess.
  2. Price — multiply tokens by a versioned, dated rate table.
  3. Ledger — append one idempotent row per request, tagged for attribution.
  4. Budget — roll the ledger up against a soft and a hard threshold.
  5. Guard — alert, degrade, or refuse before the threshold becomes an invoice.

The one rule that organizes everything: bill against the response usage object. Anything you compute before the call is an estimate — good only for the pre-flight cap check, never for the ledger.

What this skill produces

A checkable cost setup: a pricing table (each model row carries effective_date + source), an append-only ledger schema (idempotency key + attribution keys), and a budget with both a soft and a hard threshold. scripts/verify.sh lints those artifacts (last section). Prose alone is not a deliverable — emit the config.

Capture: read usage, do not estimate

Pre-send token counts (tiktoken, Anthropic client.messages.count_tokens()) are estimates. They exist for the pre-flight cap check — "will this request likely breach the budget?" — and nothing else. The truth lands in the response: input, output, cached, and reasoning tokens, plus audio/image tokens where the modality applies. Bill the ledger off that object.

Two provider facts that bite if you assume otherwise:

  • Anthropic count_tokens() returns input tokens only, is free, has its own rate limit, and is still an estimate. Anthropic is not tiktoken-compatible — do not reuse an OpenAI tokenizer to price Claude (platform.claude.com token-counting; github.com/anthropics/anthropic-tokenizer-typescript, accessed 2026-06-02).
  • Output and reasoning tokens are the expensive half — output runs 4-5x the input rate. A meter that only counts input is wrong by most of the bill.
# Bad: pricing off a pre-send character/word guess. Wrong, and ignores output.
est_tokens = len(prompt) // 4
cost = est_tokens * rate_in            # output + reasoning never counted

# Good: capture every field the response actually reports, then price that.
resp = client.messages.create(model=model, messages=msgs, max_tokens=1024)
u = resp.usage
record = {
    "input_tokens":  u.input_tokens,
    "output_tokens": u.output_tokens,
    "cache_read_tokens":     getattr(u, "cache_read_input_tokens", 0),
    "cache_write_tokens":    getattr(u, "cache_creation_input_tokens", 0),
    # OpenAI exposes cached as usage.prompt_tokens_details.cached_tokens
}
cost = price(model, record)            # see "pricing is data" below

Wrap the SDK call once so capture cannot be skipped. A meter you have to remember to call is a meter that's already missing rows (langfuse.com token-and-cost-tracking, accessed 2026-06-02).

Pricing is data, never literals

Rates drift fast and silently mis-bill when stale. Keep prices in a versioned table — one row per model, each with effective_date and source — loaded as data. Never write a rate as a literal in business logic. Look up by model and fail loud on an unknown model; never default to $0, or a new model silently bills as free and the leak is invisible.

# pricing.yaml — perishable. Verify against source before trusting. Dated 2026-06-02.
models:
  - model: claude-haiku-4.5
    effective_date: 2026-06-02
    source: cloudzero.com/blog/claude-api-pricing
    input_per_mtok: 1.00
    output_per_mtok: 5.00
    cache_read_per_mtok: 0.10
  - model: claude-sonnet-4.6
    effective_date: 2026-06-02
    source: cloudzero.com/blog/claude-api-pricing
    input_per_mtok: 3.00
    output_per_mtok: 15.00
    cache_read_per_mtok: 0.30
  - model: claude-opus-4.7
    effective_date: 2026-06-02
    source: cloudzero.com/blog/claude-api-pricing
    input_per_mtok: 5.00
    output_per_mtok: 25.00
    cache_read_per_mtok: 0.50
  - model: gpt-5.5
    effective_date: 2026-06-02
    source: openai.com/api/pricing
    input_per_mtok: 5.00
    output_per_mtok: 40.00
    cache_read_per_mtok: 0.50   # OpenAI cached input = 90% off standard input

These numbers are a snapshot, not a constant — model names and rates move month to month. Dated 2026-06-02 from the sources above. The dated per-provider snapshots, the usage-field map per provider, and refresh instructions live in references/pricing-tables.md; read it before you trust a rate.

Ledger: append-only, idempotent, attributed

One row per request, append-only. Two things must be on every row or the ledger lies:

  • A request/idempotency key. Retries and SDK auto-retries fire the same logical call twice; without a key the row is written twice and you double-count spend.
  • At least one attribution key (user, tenant, feature, model). A per-org total can tell you the bill is high; it cannot tell you which feature or customer is the leak. Attribution is the difference between "spend is up" and "the summarize-document feature on the enterprise tenant tripled."
-- append-only; (request_id) is the idempotency key — upsert, never plain insert
CREATE TABLE llm_cost_ledger (
  request_id     TEXT PRIMARY KEY,           -- idempotency: retries collapse to one row
  ts             TIMESTAMPTZ NOT NULL,
  model          TEXT NOT NULL,
  input_tokens   INTEGER NOT NULL,
  output_tokens  INTEGER NOT NULL,
  cached_tokens  INTEGER NOT NULL DEFAULT 0,
  cost_usd       NUMERIC(12,6) NOT NULL,      -- priced from the table above
  user_id        TEXT,                        -- attribution keys
  tenant_id      TEXT,
  feature        TEXT
);

This is an operational ledger, not the accounting record — categorizing the spend into the books is bookkeeping, and the same rows feed the cost numerator in unit-economics and one input line in finance-ops. "Cost per active user" on the behavior side is analytics; charging customers for metered usage is stripe.

Budgets, alerts, caps

A budget needs a soft state (alert + degrade) and a hard state (refuse). Roll the ledger up per window (day/month) and per attribution key, then branch:

Spend vs budget State Action
< 50% normal log only
50% / 80% warn fire alert to the same pipe as cloud alerts; no behavior change
100% soft cap degrade — downshift to a cheaper model, drop optional/enrichment calls, shrink context
over hard cap hard cap refuse the request with a typed error (BudgetExceededError), not a silent failure

Two distinct checks, do not conflate them:

  • Pre-flight gate uses the estimate (pre-send token count × rate) to refuse a request that would obviously blow the hard cap — this is the only legitimate use of an estimate.
  • Post-hoc reconciliation rolls up the ledger (real usage) against the budget. When someone asks "why is the bill 3x the estimate," you compare provider-billed usage to your ledgered usage — the gap is almost always uncounted output/reasoning/cache-write tokens or missing rows from un-wrapped call sites.

The three levers that move the number

Don't assert savings — show the break-even. Pricing per fact-checked sources accessed 2026-06-02 (platform.claude.com prompt-caching; finout.io anthropic-api-pricing).

  • Prompt caching. Cache reads cost 0.1x base input. The 5-min write costs 1.25x, the 1-hour write 2x. So a 5-min entry pays for itself after roughly one cache hit: 1.25 + 0.1·h (cached) beats 1·(1+h) (uncached) once h ≥ 1. Caching a stable system prompt across a session is almost always net cheaper.
  • Batch API. 50% off on both OpenAI and Anthropic, async within 24h, and it stacks with caching — combined up to ~95% off. Use it for anything not user-facing-realtime: evals, backfills, nightly summaries.
  • Model downshift. The largest lever. Route easy requests to Haiku/cheap models and reserve Opus/GPT-5.5 for hard ones — at 5x the rate, downshifting the routable half of traffic dwarfs a few percent of caching.

"We'll add caching later" without measuring the hit rate is a guess, not a lever. Instrument cache_read_tokens in the ledger first, then you know.

Cloud spend is the slow backstop

App-level metering is your real-time guard. Cloud billing alerts are a delayed backstop — useful, but never the thing standing between you and a runaway loop.

  • AWS Cost Anomaly Detection is ML-based, runs ~3x/day with up to 24h data delay → SNS → Lambda. AWS Budgets adds threshold alerts on the same pipe.
  • GCP/Azure are threshold-only. GCP's budget → Pub/Sub → Cloud Function is the one that can actually pause or throttle a workload programmatically.

Route every cloud alert into the same alert pipe as your app-level budget alerts so there's one place to look. The 24h delay is exactly why the in-app cap exists: by the time AWS notices the anomaly, the loop already spent the money. Recipes and the alert-routing pattern are in references/cloud-caps.md; the cap plumbing in depth is aws-essentials / gcp-essentials.

Build vs buy

You want Use Trade-off
Zero code change, fastest setup Helicone (proxy, ~2-min) adds a network hop / latency
SDK-level capture + a ready cost table Langfuse (MIT, ships model+tokenizer cost table) you wire the SDK, but no proxy hop
Full control / custom attribution / typed caps DIY ledger (this skill) you own pricing-table freshness and capture coverage

(firecrawl.dev best-llm-observability-tools; guptadeepak.com top-5-llm-observability-platforms-2026, accessed 2026-06-02.) Buy the proxy/platform when you want spend visibility fast; build the ledger when caps and per-feature attribution must live inside your own logic.

Anti-patterns

Anti-pattern Why it's wrong Do instead
Pricing literals in business logic a rate change silently mis-bills everything versioned table, each row dated + sourced
Billing off the pre-send estimate estimates ignore output/reasoning/cache; off by most of the bill price the response usage object
No idempotency key on ledger rows retries double-count spend request_id PRIMARY KEY, upsert not insert
Unknown model defaults to $0 a new model bills as free; leak is invisible fail loud on a model absent from the table
Only a soft alert, no hard cap alert fires, loop keeps spending a hard cap that refuses with a typed error
Org-total budget, no attribution "spend is up" — but you can't find the leak tag every row by user/tenant/feature
Counting input tokens only output+reasoning are 4-5x the cost — the expensive half capture all token fields from usage
Trusting cloud alerts for real-time control ~24h delay; the loop already spent it app-level cap is the guard; cloud is the backstop
"Add caching later" with no measurement savings unproven; may not even hit instrument cache_read_tokens, compute break-even

Verify the artifact

scripts/verify.sh [path] lints a candidate cost config/ledger (yaml/json/ts) and fails if: a pricing entry lacks effective_date or source; a model referenced in logic is missing from the table; the ledger schema lacks an idempotency/request key or any attribution key; the budget declares no soft+hard pair; or cost looks derived from a len()/char estimate instead of a usage field. It is read-only and exits 0 on a clean config and on no config found — no false failure.

Files (rsc-harness)
  • evals
    • cases.yaml 3.1 KB
      skill: cost-tracking
      
      should_trigger:
        - prompt: "Wrap our OpenAI calls so we record token cost per request."
          why: Core capture — per-request token + cost capture around an SDK call is the skill's first step.
        - prompt: "Set a $500/month cap on the Claude API that hard-stops when we hit it."
          why: Budget plus a hard cap that refuses — the guard layer this skill owns.
        - prompt: "Our bill is 3x what we estimated — reconcile the billed usage against what we computed."
          why: Non-obvious — reconciliation of provider-billed vs ledgered usage is a direct tell for this skill, not generic debugging.
        - prompt: "Attribute our LLM spend per tenant so we know which customer is expensive."
          why: Ledger attribution — per-tenant cost rollup is exactly the attribution-key design here.
        - prompt: "Pon un tope de gasto a la API y avísame al 80%."
          why: Spanish — budget plus a soft alert threshold (80%), the alert/cap branch of the skill.
        - prompt: "Is prompt caching actually cheaper for our agent, and where's the break-even?"
          why: The levers section — caching break-even math is owned here, not asserted.
        - prompt: "A runaway agent loop burned through our token budget overnight with no guardrail."
          why: The surprise-bill scenario the hard cap exists to prevent; routes straight to budgets+caps.
      
      should_not_trigger:
        - prompt: "What's our cash runway at this burn rate?"
          route_to: finance-ops
          why: Company-level cash, not the AI/cloud line item; cost-tracking is just one input feeding it.
        - prompt: "Compute the gross margin per subscription."
          route_to: unit-economics
          why: Margin math — cost-tracking only produces the cost numerator, it does not divide by units.
        - prompt: "Record this AWS invoice as an expense in the books."
          route_to: bookkeeping
          why: The accounting record, not the operational real-time meter.
        - prompt: "Decide what to charge customers for the pro plan."
          route_to: pricing
          why: Willingness-to-pay / what to charge, not controlling what we spend.
        - prompt: "Chart our monthly active users and retention curve."
          route_to: analytics
          why: Product/behavior telemetry, a different meter; the dollar capture is here but user-behavior charting is not.
      
      capability:
        - scenario: "We have a Next.js app calling Claude + OpenAI with no cost visibility and a scary bill. Give us a metering + budget-cap design."
          must_include:
            - reads cost from the response `usage` object, not pre-send estimates (estimates only for the pre-flight cap check)
            - versioned pricing table where each model row carries effective_date + source, loaded as data not literals
            - loud failure on a model absent from the pricing table (never default to $0)
            - append-only ledger with an idempotency/request key and at least one attribution key (user/tenant/feature)
            - budget with BOTH a soft threshold (alert/degrade) and a hard threshold (refuse with a typed error)
            - the three levers (prompt caching with ~1-hit break-even, batch API 50% stacking with caching, model downshift)
            - notes cloud billing alerts (AWS CAD ~24h delay, GCP Pub/Sub) are a delayed backstop, not the real-time guard
      
    • README.md 1.1 KB
      # Evals for cost-tracking
      
      These cases are eyeballed, not automated. Read `cases.yaml` and check three things by hand. First, every `should_trigger` prompt should make this skill the obvious pick — pay attention to the non-obvious ones (the "bill is 3x the estimate / reconcile billed vs response usage" line and the Spanish budget-cap line), since those are where a thinner trigger set would miss. Second, every `should_not_trigger` prompt should route cleanly to the named sibling (finance-ops, unit-economics, bookkeeping, pricing, analytics) — if any of those feels like it could legitimately land here, the description boundary needs tightening. Third, run the `capability` scenario against a generated answer and confirm each `must_include` rubric item is actually present, not hand-waved (especially: cost read from `usage`, dated+sourced pricing table, idempotency + attribution on the ledger, and a soft *and* hard cap). Note that `scripts/verify.sh` checks the emitted artifact (the pricing/ledger/budget config), not the prose — a great write-up that emits no checkable config has not finished the job.
      
  • references
    • cloud-caps.md 2.6 KB
      # Cloud spend caps — recipes (accessed 2026-06-02)
      
      App-level metering is the real-time guard. Everything here is the **slow backstop** — useful for catching what the app meter missed, never the thing standing between you and a runaway loop. The recurring caveat: cloud billing data lags, so by the time the cloud notices, the money is already spent.
      
      ## AWS
      
      ### Cost Anomaly Detection (ML-based)
      - ML model, runs roughly **3x/day**, with **up to 24h data delay**.
      - Emits to **SNS** → fan out to a **Lambda** that posts into your alert pipe (or revokes a key / disables a feature flag).
      - Source: aws.amazon.com/aws-cost-management/aws-cost-anomaly-detection; docs.aws.amazon.com/cost-management/.../manage-ad.html. Accessed 2026-06-02.
      
      ### Budgets (threshold)
      - Set 50/80/100% threshold alerts on a monthly budget → SNS → same Lambda.
      - Pairs with CAD: Budgets catch the *expected* threshold; CAD catches the *unexpected* spike.
      
      ```text
      AWS Budgets / Cost Anomaly Detection
              │ (SNS, up to ~24h delay)
              ▼
           Lambda  ──►  shared alert pipe (Slack/PagerDuty/webhook)
                    └─► optional: disable feature flag / rotate-revoke API key
      ```
      
      ## GCP
      
      - Budgets are **threshold-only** for alerting, but the budget can publish to **Pub/Sub**.
      - Pub/Sub → **Cloud Function** is the one path that can actually **pause or throttle a workload** programmatically (e.g. scale a service to zero, disable a billing account binding).
      - Source: costimizer.ai gcp guide. Accessed 2026-06-02.
      
      ```text
      GCP Budget ──► Pub/Sub ──► Cloud Function
                                  ├─► shared alert pipe
                                  └─► throttle / pause workload (real control)
      ```
      
      ## Azure
      
      - Threshold-only budgets with action groups. Same shape as GCP minus the clean Pub/Sub→pause path; wire the action group into the shared alert pipe.
      
      ## Alert-routing pattern (all providers)
      
      Route **every** cloud alert into the **same pipe** as your in-app budget alerts, so there is one place to look and one severity ladder. Tag each alert with its source (`aws-cad`, `gcp-budget`, `app-meter`) so reconciliation can tell which guard fired.
      
      ## The delay caveat, stated plainly
      
      | Guard | Latency | Can it stop spend? |
      |---|---|---|
      | App-level cap (this skill) | real-time | yes — refuse the request now |
      | AWS CAD / Budgets | up to ~24h | no — alert only (Lambda can revoke after the fact) |
      | GCP budget → Pub/Sub → Function | minutes-to-hours | partial — can pause a workload, but on lagged data |
      
      Design accordingly: the cloud cap is insurance against what your app meter failed to count; it is not the meter.
      
    • pricing-tables.md 3.2 KB
      # Pricing tables — PERISHABLE snapshot (dated 2026-06-02)
      
      These rates move month to month. **Treat this file as data with an expiry, not a constant.** Every number below carries a source; re-fetch before trusting. The whole point of the skill is that prices live in a dated table, never as literals in logic.
      
      ## Anthropic (Claude) — per 1M tokens
      
      | Model | Input | Output | Cache read | Source |
      |---|---|---|---|---|
      | Haiku 4.5 | $1.00 | $5.00 | $0.10 | cloudzero.com/blog/claude-api-pricing |
      | Sonnet 4.6 | $3.00 | $15.00 | $0.30 | cloudzero.com/blog/claude-api-pricing |
      | Opus 4.7 | $5.00 | $25.00 | $0.50 | evolink.ai/blog/claude-api-pricing-guide-2026 |
      
      Cache writes: 5-minute = 1.25x base input; 1-hour = 2x base input. Cache reads = 0.1x base input. Break-even on a 5-min cache entry: after ~1 hit (platform.claude.com prompt-caching; finout.io anthropic-api-pricing). All accessed 2026-06-02.
      
      ## OpenAI — per 1M tokens
      
      | Model | Input | Cached input | Source |
      |---|---|---|---|
      | GPT-5.5 | $5.00 | $0.50 | openai.com/api/pricing |
      | GPT-5.4 | $2.50 | $0.25 | devtk.ai/en/blog/openai-api-pricing-guide-2026 |
      
      Cached input = **90% off** standard input across the line. Output rates vary by model — read them off the source, do not assume. Model names and exact rates drift fast; this is precisely why the rate table is data and never a literal. Accessed 2026-06-02.
      
      ## Bedrock note
      
      AWS Bedrock re-prices the same underlying models (often with its own regional/commitment rates) and bills on its own line. Do not assume Bedrock matches first-party Anthropic/OpenAI rates — pull the Bedrock model rate separately and tag the table row `source: bedrock`. Cloud-side caps for Bedrock spend belong in `cloud-caps.md`.
      
      ## Batch API discount
      
      50% off on both OpenAI and Anthropic, returned async within 24h, and **stacks with prompt caching** (combined up to ~95% off). Source: finout.io anthropic-api-pricing, accessed 2026-06-02. Use for non-realtime work (evals, backfills, nightly jobs).
      
      ## The `usage` object — field map per provider
      
      Bill against these, not pre-send estimates.
      
      | Concept | Anthropic | OpenAI |
      |---|---|---|
      | Input tokens | `usage.input_tokens` | `usage.prompt_tokens` |
      | Output tokens | `usage.output_tokens` | `usage.completion_tokens` |
      | Cached (read) | `usage.cache_read_input_tokens` | `usage.prompt_tokens_details.cached_tokens` |
      | Cache write | `usage.cache_creation_input_tokens` | n/a (no explicit write line) |
      | Reasoning | included in output | `usage.completion_tokens_details.reasoning_tokens` |
      
      Pre-flight estimate tools (estimates only): Anthropic `client.messages.count_tokens()` returns input tokens, is free, separate rate limit, **not** tiktoken-compatible; OpenAI/most use `tiktoken` locally. Sources: platform.claude.com token-counting; langfuse.com token-and-cost-tracking. Accessed 2026-06-02.
      
      ## How to refresh this file
      
      1. Pull each provider's current pricing page (sources above).
      2. Update the rate and set `effective_date` to today on every changed row.
      3. Add any new model name as a new row — never leave a model out, or the lookup will hit the "unknown model" loud-fail path (which is correct behavior).
      4. Re-run `scripts/verify.sh` against your `pricing.yaml` to confirm every row still carries `effective_date` + `source`.
      
  • scripts
    • verify.sh 5.3 KB
      #!/usr/bin/env bash
      #
      # verify.sh - lint a candidate cost-tracking config/ledger artifact.
      #
      # Usage:
      #   cd <dir containing the config>
      #   ./verify.sh [path-to-config]      # yaml | yml | json | ts | js
      #
      # Read-only: never mutates anything. Checks the artifact for:
      #   1. every pricing entry carries effective_date AND source
      #   2. no model referenced in logic is absent from the pricing table
      #   3. the ledger schema has an idempotency/request key + >=1 attribution key
      #   4. the budget declares BOTH a soft and a hard threshold
      #   5. cost is derived from a usage/response field, not a len()/char estimate
      #
      # Emits [ ok ]/[fail] per check, then PASS or FAIL.
      # Exits 0 on a clean config AND on no config found (no false failure).
      # Exits 1 with the missing pieces on any violation.
      #
      # Portable: stock macOS bash 3.2 + grep + awk. No associative arrays, no mapfile.
      
      set -euo pipefail
      
      if [ -t 1 ]; then
        RED=$'\033[31m'; GREEN=$'\033[32m'; YELLOW=$'\033[33m'; RESET=$'\033[0m'
      else
        RED=''; GREEN=''; YELLOW=''; RESET=''
      fi
      
      failed=0
      fail() { printf '%s[fail]%s %s\n' "$RED" "$RESET" "$*"; failed=1; }
      ok()   { printf '%s[ ok ]%s %s\n' "$GREEN" "$RESET" "$*"; }
      
      # --- locate the config -------------------------------------------------------
      file="${1:-}"
      if [ -z "$file" ]; then
        for cand in cost-tracking.yaml cost-tracking.yml pricing.yaml pricing.yml \
                    cost-config.yaml cost-config.yml cost-tracking.json cost-tracking.ts; do
          if [ -f "$cand" ]; then file="$cand"; break; fi
        done
      fi
      
      if [ -z "$file" ] || [ ! -f "$file" ]; then
        printf '%s[skip]%s no cost-tracking config found (cost-tracking.{yaml,json,ts} / pricing.yaml) - nothing to verify.\n' "$YELLOW" "$RESET"
        exit 0
      fi
      
      printf -- '----- linting %s\n' "$file"
      
      # case-insensitive fixed-ish grep helper
      has() { grep -Eiq "$1" "$file"; }
      
      # --- 1. pricing entries carry effective_date AND source ----------------------
      mentions_pricing=0
      if has 'pric|input_per_mtok|in_per_mtok|rate_in|cost_per'; then mentions_pricing=1; fi
      if [ "$mentions_pricing" -eq 1 ]; then
        miss=""
        has 'effective_date|effective-date|effectiveDate' || miss="$miss effective_date"
        has 'source' || miss="$miss source"
        if [ -z "$miss" ]; then
          ok "pricing entries carry effective_date + source"
        else
          fail "pricing table missing:$miss (each model row must be dated and sourced)"
        fi
      else
        printf '%s[skip]%s no pricing table in this file (check 1)\n' "$YELLOW" "$RESET"
      fi
      
      # --- 2. no model referenced in logic absent from the table -------------------
      # Heuristic: collect model-ish identifiers (claude-*/gpt-*/sonnet/opus/haiku/o-series)
      # that appear OUTSIDE a "model:" / "- model:" declaration line, and warn if the
      # file declares no model rows at all while still naming models.
      declared=$(grep -Eio '(^|[^a-z])model[[:space:]]*[:=]' "$file" | wc -l | tr -d ' ')
      named=$(grep -Eio '(claude-[a-z0-9.]+|gpt-[a-z0-9.]+|sonnet|opus|haiku|gemini-[a-z0-9.]+)' "$file" \
              | sort -u | wc -l | tr -d ' ')
      if [ "$named" -gt 0 ]; then
        if [ "$declared" -eq 0 ]; then
          fail "models are named but no 'model:' table rows declared - unknown models would price as \$0"
        else
          ok "models named ($named) are declared in 'model:' rows ($declared) - lookup can fail-loud on unknowns"
        fi
      else
        printf '%s[skip]%s no model identifiers found (check 2)\n' "$YELLOW" "$RESET"
      fi
      
      # --- 3. ledger: idempotency key + >=1 attribution key ------------------------
      mentions_ledger=0
      if has 'ledger|usage|input_tokens|output_tokens|cost_usd'; then mentions_ledger=1; fi
      if [ "$mentions_ledger" -eq 1 ]; then
        miss=""
        has 'request_id|idempotency|request-id|requestId|dedup' || miss="$miss idempotency/request key"
        has 'user_id|tenant_id|tenant|feature|user-id|tenantId|userId' || miss="$miss attribution-key(user/tenant/feature)"
        if [ -z "$miss" ]; then
          ok "ledger has an idempotency key + at least one attribution key"
        else
          fail "ledger missing:$miss"
        fi
      else
        printf '%s[skip]%s no ledger schema in this file (check 3)\n' "$YELLOW" "$RESET"
      fi
      
      # --- 4. budget: both a soft and a hard threshold -----------------------------
      mentions_budget=0
      if has 'budget|threshold|cap|alert'; then mentions_budget=1; fi
      if [ "$mentions_budget" -eq 1 ]; then
        miss=""
        has 'soft|alert|warn|degrade|downshift' || miss="$miss soft(alert/degrade)"
        has 'hard|refuse|hard_cap|hard-cap|block|reject' || miss="$miss hard(refuse)"
        if [ -z "$miss" ]; then
          ok "budget declares both a soft and a hard threshold"
        else
          fail "budget missing:$miss threshold"
        fi
      else
        printf '%s[skip]%s no budget config in this file (check 4)\n' "$YELLOW" "$RESET"
      fi
      
      # --- 5. cost derived from usage, not a len()/char estimate -------------------
      if has 'usage|input_tokens|output_tokens|cached_tokens|prompt_tokens'; then
        if grep -Eiq 'len\(|\.length|char_count|charcount|//[[:space:]]*4|/[[:space:]]*4' "$file" \
           && ! has 'estimate|pre-?flight|preflight'; then
          fail "cost may be derived from a len()/char estimate - bill against the response usage object"
        else
          ok "cost derived from a usage/response field, not a char estimate"
        fi
      else
        printf '%s[skip]%s no usage-based cost field found (check 5)\n' "$YELLOW" "$RESET"
      fi
      
      echo
      if [ "$failed" -ne 0 ]; then
        printf '%sFAIL:%s cost-tracking config is missing required pieces (see [fail] lines above).\n' "$RED" "$RESET"
        exit 1
      fi
      printf '%sPASS:%s cost-tracking config carries the required guardrails.\n' "$GREEN" "$RESET"
      
  • SKILL.md 12.4 KB
    ---
    name: cost-tracking
    description: "Use when metering and capping AI or cloud app spend — tokens read from the response `usage` object, priced off a dated rate table, ledgered per user/tenant/feature, with alerts and a hard cap before the bill. NOT cash runway (that is `finance-ops`), NOT cost-per-unit margin (that is `unit-economics`), NOT booking the spend (that is `bookkeeping`)."
    tags: ["cost-tracking", "token-accounting", "llm-spend", "budgets", "alerts", "prompt-caching", "billing-guardrails", "finops"]
    recommends: ["finance-ops", "unit-economics", "analytics", "stripe", "aws-essentials", "gcp-essentials"]
    origin: risco
    ---
    
    # Cost tracking
    
    You are building the money meter for an AI or cloud app: every model call gets a token count, a price, a ledger row, and a budget check — so a cap fires *before* the invoice, not when finance forwards it in a panic. Model-API spend roughly doubled from $3.5B to $8.4B between late 2024 and mid 2025 (firecrawl.dev best-llm-observability-tools, accessed 2026-06-02); the bill is now big enough to need a guardrail, not a spreadsheet at month-end.
    
    The chain you build, in order — each step feeds the next:
    
    1. **Capture** — read tokens from the response, not from a pre-send guess.
    2. **Price** — multiply tokens by a *versioned, dated* rate table.
    3. **Ledger** — append one idempotent row per request, tagged for attribution.
    4. **Budget** — roll the ledger up against a soft and a hard threshold.
    5. **Guard** — alert, degrade, or refuse before the threshold becomes an invoice.
    
    The one rule that organizes everything: **bill against the response `usage` object. Anything you compute before the call is an estimate — good only for the pre-flight cap check, never for the ledger.**
    
    ## What this skill produces
    
    A checkable cost setup: a **pricing table** (each model row carries `effective_date` + `source`), an **append-only ledger** schema (idempotency key + attribution keys), and a **budget** with both a soft and a hard threshold. `scripts/verify.sh` lints those artifacts (last section). Prose alone is not a deliverable — emit the config.
    
    ## Capture: read `usage`, do not estimate
    
    Pre-send token counts (`tiktoken`, Anthropic `client.messages.count_tokens()`) are estimates. They exist for the *pre-flight cap check* — "will this request likely breach the budget?" — and nothing else. The truth lands in the response: input, output, **cached**, and **reasoning** tokens, plus audio/image tokens where the modality applies. Bill the ledger off that object.
    
    Two provider facts that bite if you assume otherwise:
    - Anthropic `count_tokens()` returns *input* tokens only, is free, has its own rate limit, and is still an estimate. Anthropic is **not** tiktoken-compatible — do not reuse an OpenAI tokenizer to price Claude (platform.claude.com token-counting; github.com/anthropics/anthropic-tokenizer-typescript, accessed 2026-06-02).
    - Output and reasoning tokens are the expensive half — output runs 4-5x the input rate. A meter that only counts input is wrong by most of the bill.
    
    ```python
    # Bad: pricing off a pre-send character/word guess. Wrong, and ignores output.
    est_tokens = len(prompt) // 4
    cost = est_tokens * rate_in            # output + reasoning never counted
    
    # Good: capture every field the response actually reports, then price that.
    resp = client.messages.create(model=model, messages=msgs, max_tokens=1024)
    u = resp.usage
    record = {
        "input_tokens":  u.input_tokens,
        "output_tokens": u.output_tokens,
        "cache_read_tokens":     getattr(u, "cache_read_input_tokens", 0),
        "cache_write_tokens":    getattr(u, "cache_creation_input_tokens", 0),
        # OpenAI exposes cached as usage.prompt_tokens_details.cached_tokens
    }
    cost = price(model, record)            # see "pricing is data" below
    ```
    
    Wrap the SDK call **once** so capture cannot be skipped. A meter you have to remember to call is a meter that's already missing rows (langfuse.com token-and-cost-tracking, accessed 2026-06-02).
    
    ## Pricing is data, never literals
    
    Rates drift fast and silently mis-bill when stale. Keep prices in a **versioned table** — one row per model, each with `effective_date` and `source` — loaded as data. Never write a rate as a literal in business logic. Look up by model and **fail loud on an unknown model; never default to $0**, or a new model silently bills as free and the leak is invisible.
    
    ```yaml
    # pricing.yaml — perishable. Verify against source before trusting. Dated 2026-06-02.
    models:
      - model: claude-haiku-4.5
        effective_date: 2026-06-02
        source: cloudzero.com/blog/claude-api-pricing
        input_per_mtok: 1.00
        output_per_mtok: 5.00
        cache_read_per_mtok: 0.10
      - model: claude-sonnet-4.6
        effective_date: 2026-06-02
        source: cloudzero.com/blog/claude-api-pricing
        input_per_mtok: 3.00
        output_per_mtok: 15.00
        cache_read_per_mtok: 0.30
      - model: claude-opus-4.7
        effective_date: 2026-06-02
        source: cloudzero.com/blog/claude-api-pricing
        input_per_mtok: 5.00
        output_per_mtok: 25.00
        cache_read_per_mtok: 0.50
      - model: gpt-5.5
        effective_date: 2026-06-02
        source: openai.com/api/pricing
        input_per_mtok: 5.00
        output_per_mtok: 40.00
        cache_read_per_mtok: 0.50   # OpenAI cached input = 90% off standard input
    ```
    
    These numbers are a **snapshot, not a constant** — model names and rates move month to month. Dated 2026-06-02 from the sources above. The dated per-provider snapshots, the `usage`-field map per provider, and refresh instructions live in `references/pricing-tables.md`; read it before you trust a rate.
    
    ## Ledger: append-only, idempotent, attributed
    
    One row per request, append-only. Two things must be on every row or the ledger lies:
    
    - A **request/idempotency key.** Retries and SDK auto-retries fire the same logical call twice; without a key the row is written twice and you double-count spend.
    - At least one **attribution key** (user, tenant, feature, model). A per-org total can tell you the bill is high; it cannot tell you *which feature or customer* is the leak. Attribution is the difference between "spend is up" and "the summarize-document feature on the enterprise tenant tripled."
    
    ```sql
    -- append-only; (request_id) is the idempotency key — upsert, never plain insert
    CREATE TABLE llm_cost_ledger (
      request_id     TEXT PRIMARY KEY,           -- idempotency: retries collapse to one row
      ts             TIMESTAMPTZ NOT NULL,
      model          TEXT NOT NULL,
      input_tokens   INTEGER NOT NULL,
      output_tokens  INTEGER NOT NULL,
      cached_tokens  INTEGER NOT NULL DEFAULT 0,
      cost_usd       NUMERIC(12,6) NOT NULL,      -- priced from the table above
      user_id        TEXT,                        -- attribution keys
      tenant_id      TEXT,
      feature        TEXT
    );
    ```
    
    This is an *operational* ledger, not the accounting record — categorizing the spend into the books is `bookkeeping`, and the same rows feed the cost numerator in `unit-economics` and one input line in `finance-ops`. "Cost per active user" on the behavior side is `analytics`; charging customers for metered usage is `stripe`.
    
    ## Budgets, alerts, caps
    
    A budget needs a **soft** state (alert + degrade) and a **hard** state (refuse). Roll the ledger up per window (day/month) and per attribution key, then branch:
    
    | Spend vs budget | State | Action |
    |---|---|---|
    | < 50% | normal | log only |
    | 50% / 80% | warn | fire alert to the same pipe as cloud alerts; no behavior change |
    | 100% | **soft cap** | degrade — downshift to a cheaper model, drop optional/enrichment calls, shrink context |
    | over hard cap | **hard cap** | refuse the request with a typed error (`BudgetExceededError`), not a silent failure |
    
    Two distinct checks, do not conflate them:
    - **Pre-flight gate** uses the *estimate* (pre-send token count × rate) to refuse a request that would obviously blow the hard cap — this is the only legitimate use of an estimate.
    - **Post-hoc reconciliation** rolls up the *ledger* (real `usage`) against the budget. When someone asks "why is the bill 3x the estimate," you compare provider-billed usage to your ledgered `usage` — the gap is almost always uncounted output/reasoning/cache-write tokens or missing rows from un-wrapped call sites.
    
    ## The three levers that move the number
    
    Don't assert savings — show the break-even. Pricing per fact-checked sources accessed 2026-06-02 (platform.claude.com prompt-caching; finout.io anthropic-api-pricing).
    
    - **Prompt caching.** Cache *reads* cost 0.1x base input. The 5-min write costs 1.25x, the 1-hour write 2x. So a 5-min entry pays for itself after roughly **one** cache hit: `1.25 + 0.1·h` (cached) beats `1·(1+h)` (uncached) once `h ≥ 1`. Caching a stable system prompt across a session is almost always net cheaper.
    - **Batch API.** 50% off on both OpenAI and Anthropic, async within 24h, and it **stacks with caching** — combined up to ~95% off. Use it for anything not user-facing-realtime: evals, backfills, nightly summaries.
    - **Model downshift.** The largest lever. Route easy requests to Haiku/cheap models and reserve Opus/GPT-5.5 for hard ones — at 5x the rate, downshifting the routable half of traffic dwarfs a few percent of caching.
    
    "We'll add caching later" without measuring the hit rate is a guess, not a lever. Instrument `cache_read_tokens` in the ledger first, then you know.
    
    ## Cloud spend is the slow backstop
    
    App-level metering is your real-time guard. Cloud billing alerts are a **delayed backstop** — useful, but never the thing standing between you and a runaway loop.
    
    - **AWS** Cost Anomaly Detection is ML-based, runs ~3x/day with up to 24h data delay → SNS → Lambda. AWS Budgets adds threshold alerts on the same pipe.
    - **GCP/Azure** are threshold-only. GCP's budget → Pub/Sub → Cloud Function is the one that can actually *pause or throttle* a workload programmatically.
    
    Route every cloud alert into the **same alert pipe** as your app-level budget alerts so there's one place to look. The 24h delay is exactly why the in-app cap exists: by the time AWS notices the anomaly, the loop already spent the money. Recipes and the alert-routing pattern are in `references/cloud-caps.md`; the cap plumbing in depth is `aws-essentials` / `gcp-essentials`.
    
    ## Build vs buy
    
    | You want | Use | Trade-off |
    |---|---|---|
    | Zero code change, fastest setup | Helicone (proxy, ~2-min) | adds a network hop / latency |
    | SDK-level capture + a ready cost table | Langfuse (MIT, ships model+tokenizer cost table) | you wire the SDK, but no proxy hop |
    | Full control / custom attribution / typed caps | DIY ledger (this skill) | you own pricing-table freshness and capture coverage |
    
    (firecrawl.dev best-llm-observability-tools; guptadeepak.com top-5-llm-observability-platforms-2026, accessed 2026-06-02.) Buy the proxy/platform when you want spend *visibility* fast; build the ledger when caps and per-feature attribution must live inside your own logic.
    
    ## Anti-patterns
    
    | Anti-pattern | Why it's wrong | Do instead |
    |---|---|---|
    | Pricing literals in business logic | a rate change silently mis-bills everything | versioned table, each row dated + sourced |
    | Billing off the pre-send estimate | estimates ignore output/reasoning/cache; off by most of the bill | price the response `usage` object |
    | No idempotency key on ledger rows | retries double-count spend | `request_id` PRIMARY KEY, upsert not insert |
    | Unknown model defaults to $0 | a new model bills as free; leak is invisible | fail loud on a model absent from the table |
    | Only a soft alert, no hard cap | alert fires, loop keeps spending | a hard cap that refuses with a typed error |
    | Org-total budget, no attribution | "spend is up" — but you can't find the leak | tag every row by user/tenant/feature |
    | Counting input tokens only | output+reasoning are 4-5x the cost — the expensive half | capture all token fields from `usage` |
    | Trusting cloud alerts for real-time control | ~24h delay; the loop already spent it | app-level cap is the guard; cloud is the backstop |
    | "Add caching later" with no measurement | savings unproven; may not even hit | instrument `cache_read_tokens`, compute break-even |
    
    ## Verify the artifact
    
    `scripts/verify.sh [path]` lints a candidate cost config/ledger (yaml/json/ts) and fails if: a pricing entry lacks `effective_date` or `source`; a model referenced in logic is missing from the table; the ledger schema lacks an idempotency/request key or any attribution key; the budget declares no soft+hard pair; or cost looks derived from a `len()`/char estimate instead of a `usage` field. It is read-only and exits 0 on a clean config and on no config found — no false failure.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related