cost-tracking
Use when metering and capping AI or cloud app spend — tokens read from the response `usage` object, priced off a dated rate table, ledgered per user/tenant/feature, with alerts and a hard cap before the bill. NOT cash runway (that is `finance-ops`), NOT cost-per-unit margin (that
Install
npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/cost-tracking
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
git clone https://github.com/ericrisco/rsc-harness.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Cost tracking
You are building the money meter for an AI or cloud app: every model call gets a token count, a price, a ledger row, and a budget check — so a cap fires before the invoice, not when finance forwards it in a panic. Model-API spend roughly doubled from $3.5B to $8.4B between late 2024 and mid 2025 (firecrawl.dev best-llm-observability-tools, accessed 2026-06-02); the bill is now big enough to need a guardrail, not a spreadsheet at month-end.
The chain you build, in order — each step feeds the next:
- Capture — read tokens from the response, not from a pre-send guess.
- Price — multiply tokens by a versioned, dated rate table.
- Ledger — append one idempotent row per request, tagged for attribution.
- Budget — roll the ledger up against a soft and a hard threshold.
- Guard — alert, degrade, or refuse before the threshold becomes an invoice.
The one rule that organizes everything: bill against the response usage object. Anything you compute before the call is an estimate — good only for the pre-flight cap check, never for the ledger.
What this skill produces
A checkable cost setup: a pricing table (each model row carries effective_date + source), an append-only ledger schema (idempotency key + attribution keys), and a budget with both a soft and a hard threshold. scripts/verify.sh lints those artifacts (last section). Prose alone is not a deliverable — emit the config.
Capture: read usage, do not estimate
Pre-send token counts (tiktoken, Anthropic client.messages.count_tokens()) are estimates. They exist for the pre-flight cap check — "will this request likely breach the budget?" — and nothing else. The truth lands in the response: input, output, cached, and reasoning tokens, plus audio/image tokens where the modality applies. Bill the ledger off that object.
Two provider facts that bite if you assume otherwise:
- Anthropic
count_tokens()returns input tokens only, is free, has its own rate limit, and is still an estimate. Anthropic is not tiktoken-compatible — do not reuse an OpenAI tokenizer to price Claude (platform.claude.com token-counting; github.com/anthropics/anthropic-tokenizer-typescript, accessed 2026-06-02). - Output and reasoning tokens are the expensive half — output runs 4-5x the input rate. A meter that only counts input is wrong by most of the bill.
# Bad: pricing off a pre-send character/word guess. Wrong, and ignores output.
est_tokens = len(prompt) // 4
cost = est_tokens * rate_in # output + reasoning never counted
# Good: capture every field the response actually reports, then price that.
resp = client.messages.create(model=model, messages=msgs, max_tokens=1024)
u = resp.usage
record = {
"input_tokens": u.input_tokens,
"output_tokens": u.output_tokens,
"cache_read_tokens": getattr(u, "cache_read_input_tokens", 0),
"cache_write_tokens": getattr(u, "cache_creation_input_tokens", 0),
# OpenAI exposes cached as usage.prompt_tokens_details.cached_tokens
}
cost = price(model, record) # see "pricing is data" below
Wrap the SDK call once so capture cannot be skipped. A meter you have to remember to call is a meter that's already missing rows (langfuse.com token-and-cost-tracking, accessed 2026-06-02).
Pricing is data, never literals
Rates drift fast and silently mis-bill when stale. Keep prices in a versioned table — one row per model, each with effective_date and source — loaded as data. Never write a rate as a literal in business logic. Look up by model and fail loud on an unknown model; never default to $0, or a new model silently bills as free and the leak is invisible.
# pricing.yaml — perishable. Verify against source before trusting. Dated 2026-06-02.
models:
- model: claude-haiku-4.5
effective_date: 2026-06-02
source: cloudzero.com/blog/claude-api-pricing
input_per_mtok: 1.00
output_per_mtok: 5.00
cache_read_per_mtok: 0.10
- model: claude-sonnet-4.6
effective_date: 2026-06-02
source: cloudzero.com/blog/claude-api-pricing
input_per_mtok: 3.00
output_per_mtok: 15.00
cache_read_per_mtok: 0.30
- model: claude-opus-4.7
effective_date: 2026-06-02
source: cloudzero.com/blog/claude-api-pricing
input_per_mtok: 5.00
output_per_mtok: 25.00
cache_read_per_mtok: 0.50
- model: gpt-5.5
effective_date: 2026-06-02
source: openai.com/api/pricing
input_per_mtok: 5.00
output_per_mtok: 40.00
cache_read_per_mtok: 0.50 # OpenAI cached input = 90% off standard input
These numbers are a snapshot, not a constant — model names and rates move month to month. Dated 2026-06-02 from the sources above. The dated per-provider snapshots, the usage-field map per provider, and refresh instructions live in references/pricing-tables.md; read it before you trust a rate.
Ledger: append-only, idempotent, attributed
One row per request, append-only. Two things must be on every row or the ledger lies:
- A request/idempotency key. Retries and SDK auto-retries fire the same logical call twice; without a key the row is written twice and you double-count spend.
- At least one attribution key (user, tenant, feature, model). A per-org total can tell you the bill is high; it cannot tell you which feature or customer is the leak. Attribution is the difference between "spend is up" and "the summarize-document feature on the enterprise tenant tripled."
-- append-only; (request_id) is the idempotency key — upsert, never plain insert
CREATE TABLE llm_cost_ledger (
request_id TEXT PRIMARY KEY, -- idempotency: retries collapse to one row
ts TIMESTAMPTZ NOT NULL,
model TEXT NOT NULL,
input_tokens INTEGER NOT NULL,
output_tokens INTEGER NOT NULL,
cached_tokens INTEGER NOT NULL DEFAULT 0,
cost_usd NUMERIC(12,6) NOT NULL, -- priced from the table above
user_id TEXT, -- attribution keys
tenant_id TEXT,
feature TEXT
);
This is an operational ledger, not the accounting record — categorizing the spend into the books is bookkeeping, and the same rows feed the cost numerator in unit-economics and one input line in finance-ops. "Cost per active user" on the behavior side is analytics; charging customers for metered usage is stripe.
Budgets, alerts, caps
A budget needs a soft state (alert + degrade) and a hard state (refuse). Roll the ledger up per window (day/month) and per attribution key, then branch:
| Spend vs budget | State | Action |
|---|---|---|
| < 50% | normal | log only |
| 50% / 80% | warn | fire alert to the same pipe as cloud alerts; no behavior change |
| 100% | soft cap | degrade — downshift to a cheaper model, drop optional/enrichment calls, shrink context |
| over hard cap | hard cap | refuse the request with a typed error (BudgetExceededError), not a silent failure |
Two distinct checks, do not conflate them:
- Pre-flight gate uses the estimate (pre-send token count × rate) to refuse a request that would obviously blow the hard cap — this is the only legitimate use of an estimate.
- Post-hoc reconciliation rolls up the ledger (real
usage) against the budget. When someone asks "why is the bill 3x the estimate," you compare provider-billed usage to your ledgeredusage— the gap is almost always uncounted output/reasoning/cache-write tokens or missing rows from un-wrapped call sites.
The three levers that move the number
Don't assert savings — show the break-even. Pricing per fact-checked sources accessed 2026-06-02 (platform.claude.com prompt-caching; finout.io anthropic-api-pricing).
- Prompt caching. Cache reads cost 0.1x base input. The 5-min write costs 1.25x, the 1-hour write 2x. So a 5-min entry pays for itself after roughly one cache hit:
1.25 + 0.1·h(cached) beats1·(1+h)(uncached) onceh ≥ 1. Caching a stable system prompt across a session is almost always net cheaper. - Batch API. 50% off on both OpenAI and Anthropic, async within 24h, and it stacks with caching — combined up to ~95% off. Use it for anything not user-facing-realtime: evals, backfills, nightly summaries.
- Model downshift. The largest lever. Route easy requests to Haiku/cheap models and reserve Opus/GPT-5.5 for hard ones — at 5x the rate, downshifting the routable half of traffic dwarfs a few percent of caching.
"We'll add caching later" without measuring the hit rate is a guess, not a lever. Instrument cache_read_tokens in the ledger first, then you know.
Cloud spend is the slow backstop
App-level metering is your real-time guard. Cloud billing alerts are a delayed backstop — useful, but never the thing standing between you and a runaway loop.
- AWS Cost Anomaly Detection is ML-based, runs ~3x/day with up to 24h data delay → SNS → Lambda. AWS Budgets adds threshold alerts on the same pipe.
- GCP/Azure are threshold-only. GCP's budget → Pub/Sub → Cloud Function is the one that can actually pause or throttle a workload programmatically.
Route every cloud alert into the same alert pipe as your app-level budget alerts so there's one place to look. The 24h delay is exactly why the in-app cap exists: by the time AWS notices the anomaly, the loop already spent the money. Recipes and the alert-routing pattern are in references/cloud-caps.md; the cap plumbing in depth is aws-essentials / gcp-essentials.
Build vs buy
| You want | Use | Trade-off |
|---|---|---|
| Zero code change, fastest setup | Helicone (proxy, ~2-min) | adds a network hop / latency |
| SDK-level capture + a ready cost table | Langfuse (MIT, ships model+tokenizer cost table) | you wire the SDK, but no proxy hop |
| Full control / custom attribution / typed caps | DIY ledger (this skill) | you own pricing-table freshness and capture coverage |
(firecrawl.dev best-llm-observability-tools; guptadeepak.com top-5-llm-observability-platforms-2026, accessed 2026-06-02.) Buy the proxy/platform when you want spend visibility fast; build the ledger when caps and per-feature attribution must live inside your own logic.
Anti-patterns
| Anti-pattern | Why it's wrong | Do instead |
|---|---|---|
| Pricing literals in business logic | a rate change silently mis-bills everything | versioned table, each row dated + sourced |
| Billing off the pre-send estimate | estimates ignore output/reasoning/cache; off by most of the bill | price the response usage object |
| No idempotency key on ledger rows | retries double-count spend | request_id PRIMARY KEY, upsert not insert |
| Unknown model defaults to $0 | a new model bills as free; leak is invisible | fail loud on a model absent from the table |
| Only a soft alert, no hard cap | alert fires, loop keeps spending | a hard cap that refuses with a typed error |
| Org-total budget, no attribution | "spend is up" — but you can't find the leak | tag every row by user/tenant/feature |
| Counting input tokens only | output+reasoning are 4-5x the cost — the expensive half | capture all token fields from usage |
| Trusting cloud alerts for real-time control | ~24h delay; the loop already spent it | app-level cap is the guard; cloud is the backstop |
| "Add caching later" with no measurement | savings unproven; may not even hit | instrument cache_read_tokens, compute break-even |
Verify the artifact
scripts/verify.sh [path] lints a candidate cost config/ledger (yaml/json/ts) and fails if: a pricing entry lacks effective_date or source; a model referenced in logic is missing from the table; the ledger schema lacks an idempotency/request key or any attribution key; the budget declares no soft+hard pair; or cost looks derived from a len()/char estimate instead of a usage field. It is read-only and exits 0 on a clean config and on no config found — no false failure.
Files (rsc-harness)
-
evals
-
cases.yaml 3.1 KB
skill: cost-tracking should_trigger: - prompt: "Wrap our OpenAI calls so we record token cost per request." why: Core capture — per-request token + cost capture around an SDK call is the skill's first step. - prompt: "Set a $500/month cap on the Claude API that hard-stops when we hit it." why: Budget plus a hard cap that refuses — the guard layer this skill owns. - prompt: "Our bill is 3x what we estimated — reconcile the billed usage against what we computed." why: Non-obvious — reconciliation of provider-billed vs ledgered usage is a direct tell for this skill, not generic debugging. - prompt: "Attribute our LLM spend per tenant so we know which customer is expensive." why: Ledger attribution — per-tenant cost rollup is exactly the attribution-key design here. - prompt: "Pon un tope de gasto a la API y avísame al 80%." why: Spanish — budget plus a soft alert threshold (80%), the alert/cap branch of the skill. - prompt: "Is prompt caching actually cheaper for our agent, and where's the break-even?" why: The levers section — caching break-even math is owned here, not asserted. - prompt: "A runaway agent loop burned through our token budget overnight with no guardrail." why: The surprise-bill scenario the hard cap exists to prevent; routes straight to budgets+caps. should_not_trigger: - prompt: "What's our cash runway at this burn rate?" route_to: finance-ops why: Company-level cash, not the AI/cloud line item; cost-tracking is just one input feeding it. - prompt: "Compute the gross margin per subscription." route_to: unit-economics why: Margin math — cost-tracking only produces the cost numerator, it does not divide by units. - prompt: "Record this AWS invoice as an expense in the books." route_to: bookkeeping why: The accounting record, not the operational real-time meter. - prompt: "Decide what to charge customers for the pro plan." route_to: pricing why: Willingness-to-pay / what to charge, not controlling what we spend. - prompt: "Chart our monthly active users and retention curve." route_to: analytics why: Product/behavior telemetry, a different meter; the dollar capture is here but user-behavior charting is not. capability: - scenario: "We have a Next.js app calling Claude + OpenAI with no cost visibility and a scary bill. Give us a metering + budget-cap design." must_include: - reads cost from the response `usage` object, not pre-send estimates (estimates only for the pre-flight cap check) - versioned pricing table where each model row carries effective_date + source, loaded as data not literals - loud failure on a model absent from the pricing table (never default to $0) - append-only ledger with an idempotency/request key and at least one attribution key (user/tenant/feature) - budget with BOTH a soft threshold (alert/degrade) and a hard threshold (refuse with a typed error) - the three levers (prompt caching with ~1-hit break-even, batch API 50% stacking with caching, model downshift) - notes cloud billing alerts (AWS CAD ~24h delay, GCP Pub/Sub) are a delayed backstop, not the real-time guard -
README.md 1.1 KB
# Evals for cost-tracking These cases are eyeballed, not automated. Read `cases.yaml` and check three things by hand. First, every `should_trigger` prompt should make this skill the obvious pick — pay attention to the non-obvious ones (the "bill is 3x the estimate / reconcile billed vs response usage" line and the Spanish budget-cap line), since those are where a thinner trigger set would miss. Second, every `should_not_trigger` prompt should route cleanly to the named sibling (finance-ops, unit-economics, bookkeeping, pricing, analytics) — if any of those feels like it could legitimately land here, the description boundary needs tightening. Third, run the `capability` scenario against a generated answer and confirm each `must_include` rubric item is actually present, not hand-waved (especially: cost read from `usage`, dated+sourced pricing table, idempotency + attribution on the ledger, and a soft *and* hard cap). Note that `scripts/verify.sh` checks the emitted artifact (the pricing/ledger/budget config), not the prose — a great write-up that emits no checkable config has not finished the job.
-
-
references
-
cloud-caps.md 2.6 KB
# Cloud spend caps — recipes (accessed 2026-06-02) App-level metering is the real-time guard. Everything here is the **slow backstop** — useful for catching what the app meter missed, never the thing standing between you and a runaway loop. The recurring caveat: cloud billing data lags, so by the time the cloud notices, the money is already spent. ## AWS ### Cost Anomaly Detection (ML-based) - ML model, runs roughly **3x/day**, with **up to 24h data delay**. - Emits to **SNS** → fan out to a **Lambda** that posts into your alert pipe (or revokes a key / disables a feature flag). - Source: aws.amazon.com/aws-cost-management/aws-cost-anomaly-detection; docs.aws.amazon.com/cost-management/.../manage-ad.html. Accessed 2026-06-02. ### Budgets (threshold) - Set 50/80/100% threshold alerts on a monthly budget → SNS → same Lambda. - Pairs with CAD: Budgets catch the *expected* threshold; CAD catches the *unexpected* spike. ```text AWS Budgets / Cost Anomaly Detection │ (SNS, up to ~24h delay) ▼ Lambda ──► shared alert pipe (Slack/PagerDuty/webhook) └─► optional: disable feature flag / rotate-revoke API key ``` ## GCP - Budgets are **threshold-only** for alerting, but the budget can publish to **Pub/Sub**. - Pub/Sub → **Cloud Function** is the one path that can actually **pause or throttle a workload** programmatically (e.g. scale a service to zero, disable a billing account binding). - Source: costimizer.ai gcp guide. Accessed 2026-06-02. ```text GCP Budget ──► Pub/Sub ──► Cloud Function ├─► shared alert pipe └─► throttle / pause workload (real control) ``` ## Azure - Threshold-only budgets with action groups. Same shape as GCP minus the clean Pub/Sub→pause path; wire the action group into the shared alert pipe. ## Alert-routing pattern (all providers) Route **every** cloud alert into the **same pipe** as your in-app budget alerts, so there is one place to look and one severity ladder. Tag each alert with its source (`aws-cad`, `gcp-budget`, `app-meter`) so reconciliation can tell which guard fired. ## The delay caveat, stated plainly | Guard | Latency | Can it stop spend? | |---|---|---| | App-level cap (this skill) | real-time | yes — refuse the request now | | AWS CAD / Budgets | up to ~24h | no — alert only (Lambda can revoke after the fact) | | GCP budget → Pub/Sub → Function | minutes-to-hours | partial — can pause a workload, but on lagged data | Design accordingly: the cloud cap is insurance against what your app meter failed to count; it is not the meter. -
pricing-tables.md 3.2 KB
# Pricing tables — PERISHABLE snapshot (dated 2026-06-02) These rates move month to month. **Treat this file as data with an expiry, not a constant.** Every number below carries a source; re-fetch before trusting. The whole point of the skill is that prices live in a dated table, never as literals in logic. ## Anthropic (Claude) — per 1M tokens | Model | Input | Output | Cache read | Source | |---|---|---|---|---| | Haiku 4.5 | $1.00 | $5.00 | $0.10 | cloudzero.com/blog/claude-api-pricing | | Sonnet 4.6 | $3.00 | $15.00 | $0.30 | cloudzero.com/blog/claude-api-pricing | | Opus 4.7 | $5.00 | $25.00 | $0.50 | evolink.ai/blog/claude-api-pricing-guide-2026 | Cache writes: 5-minute = 1.25x base input; 1-hour = 2x base input. Cache reads = 0.1x base input. Break-even on a 5-min cache entry: after ~1 hit (platform.claude.com prompt-caching; finout.io anthropic-api-pricing). All accessed 2026-06-02. ## OpenAI — per 1M tokens | Model | Input | Cached input | Source | |---|---|---|---| | GPT-5.5 | $5.00 | $0.50 | openai.com/api/pricing | | GPT-5.4 | $2.50 | $0.25 | devtk.ai/en/blog/openai-api-pricing-guide-2026 | Cached input = **90% off** standard input across the line. Output rates vary by model — read them off the source, do not assume. Model names and exact rates drift fast; this is precisely why the rate table is data and never a literal. Accessed 2026-06-02. ## Bedrock note AWS Bedrock re-prices the same underlying models (often with its own regional/commitment rates) and bills on its own line. Do not assume Bedrock matches first-party Anthropic/OpenAI rates — pull the Bedrock model rate separately and tag the table row `source: bedrock`. Cloud-side caps for Bedrock spend belong in `cloud-caps.md`. ## Batch API discount 50% off on both OpenAI and Anthropic, returned async within 24h, and **stacks with prompt caching** (combined up to ~95% off). Source: finout.io anthropic-api-pricing, accessed 2026-06-02. Use for non-realtime work (evals, backfills, nightly jobs). ## The `usage` object — field map per provider Bill against these, not pre-send estimates. | Concept | Anthropic | OpenAI | |---|---|---| | Input tokens | `usage.input_tokens` | `usage.prompt_tokens` | | Output tokens | `usage.output_tokens` | `usage.completion_tokens` | | Cached (read) | `usage.cache_read_input_tokens` | `usage.prompt_tokens_details.cached_tokens` | | Cache write | `usage.cache_creation_input_tokens` | n/a (no explicit write line) | | Reasoning | included in output | `usage.completion_tokens_details.reasoning_tokens` | Pre-flight estimate tools (estimates only): Anthropic `client.messages.count_tokens()` returns input tokens, is free, separate rate limit, **not** tiktoken-compatible; OpenAI/most use `tiktoken` locally. Sources: platform.claude.com token-counting; langfuse.com token-and-cost-tracking. Accessed 2026-06-02. ## How to refresh this file 1. Pull each provider's current pricing page (sources above). 2. Update the rate and set `effective_date` to today on every changed row. 3. Add any new model name as a new row — never leave a model out, or the lookup will hit the "unknown model" loud-fail path (which is correct behavior). 4. Re-run `scripts/verify.sh` against your `pricing.yaml` to confirm every row still carries `effective_date` + `source`.
-
-
scripts
-
verify.sh 5.3 KB
#!/usr/bin/env bash # # verify.sh - lint a candidate cost-tracking config/ledger artifact. # # Usage: # cd <dir containing the config> # ./verify.sh [path-to-config] # yaml | yml | json | ts | js # # Read-only: never mutates anything. Checks the artifact for: # 1. every pricing entry carries effective_date AND source # 2. no model referenced in logic is absent from the pricing table # 3. the ledger schema has an idempotency/request key + >=1 attribution key # 4. the budget declares BOTH a soft and a hard threshold # 5. cost is derived from a usage/response field, not a len()/char estimate # # Emits [ ok ]/[fail] per check, then PASS or FAIL. # Exits 0 on a clean config AND on no config found (no false failure). # Exits 1 with the missing pieces on any violation. # # Portable: stock macOS bash 3.2 + grep + awk. No associative arrays, no mapfile. set -euo pipefail if [ -t 1 ]; then RED=$'\033[31m'; GREEN=$'\033[32m'; YELLOW=$'\033[33m'; RESET=$'\033[0m' else RED=''; GREEN=''; YELLOW=''; RESET='' fi failed=0 fail() { printf '%s[fail]%s %s\n' "$RED" "$RESET" "$*"; failed=1; } ok() { printf '%s[ ok ]%s %s\n' "$GREEN" "$RESET" "$*"; } # --- locate the config ------------------------------------------------------- file="${1:-}" if [ -z "$file" ]; then for cand in cost-tracking.yaml cost-tracking.yml pricing.yaml pricing.yml \ cost-config.yaml cost-config.yml cost-tracking.json cost-tracking.ts; do if [ -f "$cand" ]; then file="$cand"; break; fi done fi if [ -z "$file" ] || [ ! -f "$file" ]; then printf '%s[skip]%s no cost-tracking config found (cost-tracking.{yaml,json,ts} / pricing.yaml) - nothing to verify.\n' "$YELLOW" "$RESET" exit 0 fi printf -- '----- linting %s\n' "$file" # case-insensitive fixed-ish grep helper has() { grep -Eiq "$1" "$file"; } # --- 1. pricing entries carry effective_date AND source ---------------------- mentions_pricing=0 if has 'pric|input_per_mtok|in_per_mtok|rate_in|cost_per'; then mentions_pricing=1; fi if [ "$mentions_pricing" -eq 1 ]; then miss="" has 'effective_date|effective-date|effectiveDate' || miss="$miss effective_date" has 'source' || miss="$miss source" if [ -z "$miss" ]; then ok "pricing entries carry effective_date + source" else fail "pricing table missing:$miss (each model row must be dated and sourced)" fi else printf '%s[skip]%s no pricing table in this file (check 1)\n' "$YELLOW" "$RESET" fi # --- 2. no model referenced in logic absent from the table ------------------- # Heuristic: collect model-ish identifiers (claude-*/gpt-*/sonnet/opus/haiku/o-series) # that appear OUTSIDE a "model:" / "- model:" declaration line, and warn if the # file declares no model rows at all while still naming models. declared=$(grep -Eio '(^|[^a-z])model[[:space:]]*[:=]' "$file" | wc -l | tr -d ' ') named=$(grep -Eio '(claude-[a-z0-9.]+|gpt-[a-z0-9.]+|sonnet|opus|haiku|gemini-[a-z0-9.]+)' "$file" \ | sort -u | wc -l | tr -d ' ') if [ "$named" -gt 0 ]; then if [ "$declared" -eq 0 ]; then fail "models are named but no 'model:' table rows declared - unknown models would price as \$0" else ok "models named ($named) are declared in 'model:' rows ($declared) - lookup can fail-loud on unknowns" fi else printf '%s[skip]%s no model identifiers found (check 2)\n' "$YELLOW" "$RESET" fi # --- 3. ledger: idempotency key + >=1 attribution key ------------------------ mentions_ledger=0 if has 'ledger|usage|input_tokens|output_tokens|cost_usd'; then mentions_ledger=1; fi if [ "$mentions_ledger" -eq 1 ]; then miss="" has 'request_id|idempotency|request-id|requestId|dedup' || miss="$miss idempotency/request key" has 'user_id|tenant_id|tenant|feature|user-id|tenantId|userId' || miss="$miss attribution-key(user/tenant/feature)" if [ -z "$miss" ]; then ok "ledger has an idempotency key + at least one attribution key" else fail "ledger missing:$miss" fi else printf '%s[skip]%s no ledger schema in this file (check 3)\n' "$YELLOW" "$RESET" fi # --- 4. budget: both a soft and a hard threshold ----------------------------- mentions_budget=0 if has 'budget|threshold|cap|alert'; then mentions_budget=1; fi if [ "$mentions_budget" -eq 1 ]; then miss="" has 'soft|alert|warn|degrade|downshift' || miss="$miss soft(alert/degrade)" has 'hard|refuse|hard_cap|hard-cap|block|reject' || miss="$miss hard(refuse)" if [ -z "$miss" ]; then ok "budget declares both a soft and a hard threshold" else fail "budget missing:$miss threshold" fi else printf '%s[skip]%s no budget config in this file (check 4)\n' "$YELLOW" "$RESET" fi # --- 5. cost derived from usage, not a len()/char estimate ------------------- if has 'usage|input_tokens|output_tokens|cached_tokens|prompt_tokens'; then if grep -Eiq 'len\(|\.length|char_count|charcount|//[[:space:]]*4|/[[:space:]]*4' "$file" \ && ! has 'estimate|pre-?flight|preflight'; then fail "cost may be derived from a len()/char estimate - bill against the response usage object" else ok "cost derived from a usage/response field, not a char estimate" fi else printf '%s[skip]%s no usage-based cost field found (check 5)\n' "$YELLOW" "$RESET" fi echo if [ "$failed" -ne 0 ]; then printf '%sFAIL:%s cost-tracking config is missing required pieces (see [fail] lines above).\n' "$RED" "$RESET" exit 1 fi printf '%sPASS:%s cost-tracking config carries the required guardrails.\n' "$GREEN" "$RESET"
-
-
SKILL.md 12.4 KB
--- name: cost-tracking description: "Use when metering and capping AI or cloud app spend — tokens read from the response `usage` object, priced off a dated rate table, ledgered per user/tenant/feature, with alerts and a hard cap before the bill. NOT cash runway (that is `finance-ops`), NOT cost-per-unit margin (that is `unit-economics`), NOT booking the spend (that is `bookkeeping`)." tags: ["cost-tracking", "token-accounting", "llm-spend", "budgets", "alerts", "prompt-caching", "billing-guardrails", "finops"] recommends: ["finance-ops", "unit-economics", "analytics", "stripe", "aws-essentials", "gcp-essentials"] origin: risco --- # Cost tracking You are building the money meter for an AI or cloud app: every model call gets a token count, a price, a ledger row, and a budget check — so a cap fires *before* the invoice, not when finance forwards it in a panic. Model-API spend roughly doubled from $3.5B to $8.4B between late 2024 and mid 2025 (firecrawl.dev best-llm-observability-tools, accessed 2026-06-02); the bill is now big enough to need a guardrail, not a spreadsheet at month-end. The chain you build, in order — each step feeds the next: 1. **Capture** — read tokens from the response, not from a pre-send guess. 2. **Price** — multiply tokens by a *versioned, dated* rate table. 3. **Ledger** — append one idempotent row per request, tagged for attribution. 4. **Budget** — roll the ledger up against a soft and a hard threshold. 5. **Guard** — alert, degrade, or refuse before the threshold becomes an invoice. The one rule that organizes everything: **bill against the response `usage` object. Anything you compute before the call is an estimate — good only for the pre-flight cap check, never for the ledger.** ## What this skill produces A checkable cost setup: a **pricing table** (each model row carries `effective_date` + `source`), an **append-only ledger** schema (idempotency key + attribution keys), and a **budget** with both a soft and a hard threshold. `scripts/verify.sh` lints those artifacts (last section). Prose alone is not a deliverable — emit the config. ## Capture: read `usage`, do not estimate Pre-send token counts (`tiktoken`, Anthropic `client.messages.count_tokens()`) are estimates. They exist for the *pre-flight cap check* — "will this request likely breach the budget?" — and nothing else. The truth lands in the response: input, output, **cached**, and **reasoning** tokens, plus audio/image tokens where the modality applies. Bill the ledger off that object. Two provider facts that bite if you assume otherwise: - Anthropic `count_tokens()` returns *input* tokens only, is free, has its own rate limit, and is still an estimate. Anthropic is **not** tiktoken-compatible — do not reuse an OpenAI tokenizer to price Claude (platform.claude.com token-counting; github.com/anthropics/anthropic-tokenizer-typescript, accessed 2026-06-02). - Output and reasoning tokens are the expensive half — output runs 4-5x the input rate. A meter that only counts input is wrong by most of the bill. ```python # Bad: pricing off a pre-send character/word guess. Wrong, and ignores output. est_tokens = len(prompt) // 4 cost = est_tokens * rate_in # output + reasoning never counted # Good: capture every field the response actually reports, then price that. resp = client.messages.create(model=model, messages=msgs, max_tokens=1024) u = resp.usage record = { "input_tokens": u.input_tokens, "output_tokens": u.output_tokens, "cache_read_tokens": getattr(u, "cache_read_input_tokens", 0), "cache_write_tokens": getattr(u, "cache_creation_input_tokens", 0), # OpenAI exposes cached as usage.prompt_tokens_details.cached_tokens } cost = price(model, record) # see "pricing is data" below ``` Wrap the SDK call **once** so capture cannot be skipped. A meter you have to remember to call is a meter that's already missing rows (langfuse.com token-and-cost-tracking, accessed 2026-06-02). ## Pricing is data, never literals Rates drift fast and silently mis-bill when stale. Keep prices in a **versioned table** — one row per model, each with `effective_date` and `source` — loaded as data. Never write a rate as a literal in business logic. Look up by model and **fail loud on an unknown model; never default to $0**, or a new model silently bills as free and the leak is invisible. ```yaml # pricing.yaml — perishable. Verify against source before trusting. Dated 2026-06-02. models: - model: claude-haiku-4.5 effective_date: 2026-06-02 source: cloudzero.com/blog/claude-api-pricing input_per_mtok: 1.00 output_per_mtok: 5.00 cache_read_per_mtok: 0.10 - model: claude-sonnet-4.6 effective_date: 2026-06-02 source: cloudzero.com/blog/claude-api-pricing input_per_mtok: 3.00 output_per_mtok: 15.00 cache_read_per_mtok: 0.30 - model: claude-opus-4.7 effective_date: 2026-06-02 source: cloudzero.com/blog/claude-api-pricing input_per_mtok: 5.00 output_per_mtok: 25.00 cache_read_per_mtok: 0.50 - model: gpt-5.5 effective_date: 2026-06-02 source: openai.com/api/pricing input_per_mtok: 5.00 output_per_mtok: 40.00 cache_read_per_mtok: 0.50 # OpenAI cached input = 90% off standard input ``` These numbers are a **snapshot, not a constant** — model names and rates move month to month. Dated 2026-06-02 from the sources above. The dated per-provider snapshots, the `usage`-field map per provider, and refresh instructions live in `references/pricing-tables.md`; read it before you trust a rate. ## Ledger: append-only, idempotent, attributed One row per request, append-only. Two things must be on every row or the ledger lies: - A **request/idempotency key.** Retries and SDK auto-retries fire the same logical call twice; without a key the row is written twice and you double-count spend. - At least one **attribution key** (user, tenant, feature, model). A per-org total can tell you the bill is high; it cannot tell you *which feature or customer* is the leak. Attribution is the difference between "spend is up" and "the summarize-document feature on the enterprise tenant tripled." ```sql -- append-only; (request_id) is the idempotency key — upsert, never plain insert CREATE TABLE llm_cost_ledger ( request_id TEXT PRIMARY KEY, -- idempotency: retries collapse to one row ts TIMESTAMPTZ NOT NULL, model TEXT NOT NULL, input_tokens INTEGER NOT NULL, output_tokens INTEGER NOT NULL, cached_tokens INTEGER NOT NULL DEFAULT 0, cost_usd NUMERIC(12,6) NOT NULL, -- priced from the table above user_id TEXT, -- attribution keys tenant_id TEXT, feature TEXT ); ``` This is an *operational* ledger, not the accounting record — categorizing the spend into the books is `bookkeeping`, and the same rows feed the cost numerator in `unit-economics` and one input line in `finance-ops`. "Cost per active user" on the behavior side is `analytics`; charging customers for metered usage is `stripe`. ## Budgets, alerts, caps A budget needs a **soft** state (alert + degrade) and a **hard** state (refuse). Roll the ledger up per window (day/month) and per attribution key, then branch: | Spend vs budget | State | Action | |---|---|---| | < 50% | normal | log only | | 50% / 80% | warn | fire alert to the same pipe as cloud alerts; no behavior change | | 100% | **soft cap** | degrade — downshift to a cheaper model, drop optional/enrichment calls, shrink context | | over hard cap | **hard cap** | refuse the request with a typed error (`BudgetExceededError`), not a silent failure | Two distinct checks, do not conflate them: - **Pre-flight gate** uses the *estimate* (pre-send token count × rate) to refuse a request that would obviously blow the hard cap — this is the only legitimate use of an estimate. - **Post-hoc reconciliation** rolls up the *ledger* (real `usage`) against the budget. When someone asks "why is the bill 3x the estimate," you compare provider-billed usage to your ledgered `usage` — the gap is almost always uncounted output/reasoning/cache-write tokens or missing rows from un-wrapped call sites. ## The three levers that move the number Don't assert savings — show the break-even. Pricing per fact-checked sources accessed 2026-06-02 (platform.claude.com prompt-caching; finout.io anthropic-api-pricing). - **Prompt caching.** Cache *reads* cost 0.1x base input. The 5-min write costs 1.25x, the 1-hour write 2x. So a 5-min entry pays for itself after roughly **one** cache hit: `1.25 + 0.1·h` (cached) beats `1·(1+h)` (uncached) once `h ≥ 1`. Caching a stable system prompt across a session is almost always net cheaper. - **Batch API.** 50% off on both OpenAI and Anthropic, async within 24h, and it **stacks with caching** — combined up to ~95% off. Use it for anything not user-facing-realtime: evals, backfills, nightly summaries. - **Model downshift.** The largest lever. Route easy requests to Haiku/cheap models and reserve Opus/GPT-5.5 for hard ones — at 5x the rate, downshifting the routable half of traffic dwarfs a few percent of caching. "We'll add caching later" without measuring the hit rate is a guess, not a lever. Instrument `cache_read_tokens` in the ledger first, then you know. ## Cloud spend is the slow backstop App-level metering is your real-time guard. Cloud billing alerts are a **delayed backstop** — useful, but never the thing standing between you and a runaway loop. - **AWS** Cost Anomaly Detection is ML-based, runs ~3x/day with up to 24h data delay → SNS → Lambda. AWS Budgets adds threshold alerts on the same pipe. - **GCP/Azure** are threshold-only. GCP's budget → Pub/Sub → Cloud Function is the one that can actually *pause or throttle* a workload programmatically. Route every cloud alert into the **same alert pipe** as your app-level budget alerts so there's one place to look. The 24h delay is exactly why the in-app cap exists: by the time AWS notices the anomaly, the loop already spent the money. Recipes and the alert-routing pattern are in `references/cloud-caps.md`; the cap plumbing in depth is `aws-essentials` / `gcp-essentials`. ## Build vs buy | You want | Use | Trade-off | |---|---|---| | Zero code change, fastest setup | Helicone (proxy, ~2-min) | adds a network hop / latency | | SDK-level capture + a ready cost table | Langfuse (MIT, ships model+tokenizer cost table) | you wire the SDK, but no proxy hop | | Full control / custom attribution / typed caps | DIY ledger (this skill) | you own pricing-table freshness and capture coverage | (firecrawl.dev best-llm-observability-tools; guptadeepak.com top-5-llm-observability-platforms-2026, accessed 2026-06-02.) Buy the proxy/platform when you want spend *visibility* fast; build the ledger when caps and per-feature attribution must live inside your own logic. ## Anti-patterns | Anti-pattern | Why it's wrong | Do instead | |---|---|---| | Pricing literals in business logic | a rate change silently mis-bills everything | versioned table, each row dated + sourced | | Billing off the pre-send estimate | estimates ignore output/reasoning/cache; off by most of the bill | price the response `usage` object | | No idempotency key on ledger rows | retries double-count spend | `request_id` PRIMARY KEY, upsert not insert | | Unknown model defaults to $0 | a new model bills as free; leak is invisible | fail loud on a model absent from the table | | Only a soft alert, no hard cap | alert fires, loop keeps spending | a hard cap that refuses with a typed error | | Org-total budget, no attribution | "spend is up" — but you can't find the leak | tag every row by user/tenant/feature | | Counting input tokens only | output+reasoning are 4-5x the cost — the expensive half | capture all token fields from `usage` | | Trusting cloud alerts for real-time control | ~24h delay; the loop already spent it | app-level cap is the guard; cloud is the backstop | | "Add caching later" with no measurement | savings unproven; may not even hit | instrument `cache_read_tokens`, compute break-even | ## Verify the artifact `scripts/verify.sh [path]` lints a candidate cost config/ledger (yaml/json/ts) and fails if: a pricing entry lacks `effective_date` or `source`; a model referenced in logic is missing from the table; the ledger schema lacks an idempotency/request key or any attribution key; the budget declares no soft+hard pair; or cost looks derived from a `len()`/char estimate instead of a `usage` field. It is read-only and exits 0 on a clean config and on no config found — no false failure.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.