claude-api-ops
Building applications ON Claude - the Anthropic API and Claude Agent SDK. Use for: anthropic api, claude api, messages api, tool use, function calling, prompt caching, agent sdk, claude-agent-sdk, structured output, json schema output, batches api, extended thinking, adaptive thi
Install
npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/claude-api-ops
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
git clone https://github.com/0xDarkMatter/claude-mods.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Claude API Operations
Building applications and agents on Anthropic's API: the Messages API, tool use, prompt caching, structured outputs, batches, thinking/effort, and the Claude Agent SDK. For developers writing apps against the API — not for using Claude Code itself.
API surfaces move fast. Model IDs, parameters, and betas in this skill were
verified against platform.claude.com (2026-08). When in doubt — especially for
"latest model" or pricing questions — verify with WebFetch against
https://platform.claude.com/docs/en/about-claude/models/overview.md or query
the Models API (client.models.list()).
Current Models (verified 2026-08)
| Model | ID (exact, no date suffix) | Context | Max Output | Input $/MTok | Output $/MTok |
|---|---|---|---|---|---|
| Claude Fable 5 | claude-fable-5 |
1M | 128K | $10.00 | $50.00 |
| Claude Opus 5 | claude-opus-5 |
1M | 128K | $5.00 | $25.00 |
| Claude Sonnet 5 | claude-sonnet-5 |
1M | 128K | $2.00 | $10.00 |
| Claude Haiku 4.5 | claude-haiku-4-5 |
200K | 64K | $1.00 | $5.00 |
Use these alias IDs verbatim. Never append date suffixes (claude-sonnet-5-20260630
is wrong → 404). Haiku 4.5 is the one current model with a dated snapshot id
(claude-haiku-4-5-20251001) behind its alias; from the 4.6 generation on, the dateless
id is the pinned snapshot.
Legacy (still available, no longer current): claude-opus-4-8, claude-opus-4-7,
claude-opus-4-6, claude-opus-4-5, claude-sonnet-4-6, claude-sonnet-4-5. Migrating
off one: https://platform.claude.com/docs/en/models/opus-5/migration-guide.md (or run
/claude-api migrate in Claude Code). Live capability lookup:
client.models.retrieve("claude-opus-5") → .max_input_tokens, .max_tokens,
.capabilities dict.
Model Selection Decision Tree
What is the workload?
│
├─ Hardest problems, long-horizon agents, deep research, ceiling intelligence
│ └─ claude-fable-5 (premium ceiling) or claude-opus-5 (default flagship)
│
├─ Agentic coding, tool-heavy workflows, production assistants
│ └─ claude-opus-5 (quality) or claude-sonnet-5 (speed/cost balance)
│
├─ High-volume production: summarization, RAG answers, extraction
│ └─ claude-sonnet-5
│
├─ Classification, routing, simple Q&A, latency-critical
│ └─ claude-haiku-4-5
│
└─ Subagents inside a larger system
└─ One tier below the orchestrator (Opus loop → Sonnet/Haiku workers)
Tiering rule: route by task difficulty, not by uniform default. An Opus orchestrator dispatching Haiku classifiers is routinely 5-10x cheaper than Opus-everywhere with no quality loss on the simple legs.
Which Surface? (API vs Agent SDK vs Batches)
| Need | Use | Why |
|---|---|---|
| One request → one response (classify, summarize, extract, Q&A) | Messages API | Simplest; full control |
| Multi-step pipeline, your code controls the logic | Messages API + tool use | You own the loop |
| Custom agent with your own tools, your infra | Messages API + tool use (manual loop or SDK tool runner) | Max flexibility |
| Agent that reads/edits files, runs commands, searches — without building tools | Claude Agent SDK | Claude Code's tools + agent loop as a library |
| CI/CD automation, coding agents, production agent apps | Claude Agent SDK | Built-in tools, hooks, sessions, MCP |
| Large non-urgent workloads (eval runs, backfills, bulk extraction) | Batches API | 50% discount, ≤24h turnaround |
| Hosted agent, Anthropic runs loop + sandbox | Managed Agents (beta) | No infra; see official docs |
Rule of thumb: start at the simplest tier. Reach for an agent only when the task is genuinely open-ended (multi-step, hard to fully specify, errors recoverable, value justifies cost).
Messages API Quick Start
Everything goes through POST /v1/messages. Headers: x-api-key,
anthropic-version: 2023-06-01, content-type: application/json.
# pip install anthropic
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
response = client.messages.create(
model="claude-opus-5",
max_tokens=16000,
system="You are a concise technical assistant.",
messages=[{"role": "user", "content": "Explain CRDTs in one paragraph."}],
)
for block in response.content: # content is a list of typed blocks
if block.type == "text": # always check .type before .text
print(block.text)
print(response.stop_reason, response.usage.input_tokens, response.usage.output_tokens)
// npm install @anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const response = await client.messages.create({
model: "claude-opus-5",
max_tokens: 16000,
messages: [{ role: "user", content: "Explain CRDTs in one paragraph." }],
});
for (const block of response.content) {
if (block.type === "text") console.log(block.text); // narrow the union first
}
Streaming (default to it for long outputs — non-streaming above ~16K
max_tokens risks SDK HTTP timeouts):
with client.messages.stream(model="claude-opus-5", max_tokens=64000,
messages=[{"role": "user", "content": "Write a long report"}]) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
final = stream.get_final_message() # full Message after streaming
Full params, response shape, stop reasons, errors, retries, rate limits: references/messages-api.md
Thinking & Effort (quick reference)
- Adaptive thinking is ON BY DEFAULT on Fable 5 / Opus 5 / Sonnet 5 — send no
thinkingfield and you still get (and pay for) thinking. On the legacy 4.6–4.8 models it stays off until you setthinking: {"type": "adaptive"}. - Manual budgets are gone.
{"type": "enabled", "budget_tokens": N}returns a 400 on Opus 4.7 and every later model (Opus 5, Sonnet 5, Fable 5 included); deprecated on Opus 4.6 / Sonnet 4.6. Control depth witheffort, not tokens. - Turning thinking off: Sonnet 5 accepts
{"type": "disabled"}. Opus 5 accepts it only at efforthighor below — pairing it withxhigh/maxis a 400. Fable 5 rejects it outright; thinking there is unconditional, so budget for it. - Effort (GA):
output_config: {"effort": "low" | "medium" | "high" | "xhigh" | "max"}— nested inoutput_config, not top-level. Defaulthigh(identical to omitting it).xhigh: Fable 5, Opus 5, Opus 4.8/4.7, Sonnet 5.max: those plus Opus 4.6 and Sonnet 4.6. Haiku 4.5 does not supporteffortat all. - Sampling params removed on Opus 4.7 and later (so Opus 5, Sonnet 5, Fable 5):
temperature,top_p,top_kall return 400 — and the Python SDK v1.0+ doesn't define them, so passing them raisesTypeError. Steer with prompting + effort. - Forced tool_choice is fine with adaptive thinking. The auto/none-only
restriction applies to manual extended thinking (
{"type": "enabled"}) only; adaptive mode — including the models where it's on by default — accepts{"type": "any"}and{"type": "tool", ...}. - Thinking text is omitted by default on Fable 5 / Opus 5 / Sonnet 5 / Opus 4.8 /
4.7 — opt in with
thinking: {"type": "adaptive", "display": "summarized"}if you surface reasoning to users. Either way the blocks are billed, and must be echoed back unmodified (emptythinkingfield included) in a tool-use loop, or the next request 400s.
Details and gotchas: references/structured-outputs.md (thinking interplay) and references/messages-api.md.
Tool Use (quick reference)
tools = [{
"name": "get_weather",
"description": "Get current weather. Call when the user asks about weather conditions.",
"input_schema": {
"type": "object",
"properties": {"location": {"type": "string", "description": "City, e.g. Paris"}},
"required": ["location"],
},
}]
response = client.messages.create(model="claude-opus-5", max_tokens=16000,
tools=tools, messages=messages)
if response.stop_reason == "tool_use":
... # execute, send tool_result back, loop
tool_choice: {"type": "auto"} (default) | {"type": "any"} | {"type": "tool", "name": "..."} | {"type": "none"}. Add
"disable_parallel_tool_use": true to force at most one call per response.
The agentic loop, parallel tool results, pause_turn, is_error, server-side
tools, and SDK tool runners: references/tool-use.md
Cost Optimization Checklist
Work top-down; each item is independent:
- Right-size the model. Haiku for classification/routing, Sonnet for volume work, Opus/Fable for the hard 10%. Largest single lever.
- Prompt caching on stable prefixes (system prompt, tool defs, big docs):
cache_control: {"type": "ephemeral"}. Reads cost ~0.1x; up to 90% savings. Verify withusage.cache_read_input_tokens > 0— zero means a silent invalidator (timestamp in system prompt, unsorted JSON, varying tools). - Batches API for anything that can wait ≤24h: flat 50% off all tokens, stacks with caching.
- Cap output: set
max_tokensto what you need (256 for classification); stream + generous cap for long generation. - Tune effort down where quality allows:
mediumis often the sweet spot;lowfor subagents and simple tasks. - Count before sending:
client.messages.count_tokens(...)(never tiktoken — it's OpenAI's tokenizer and undercounts Claude by 15-20%). - Keep prefixes stable: order requests
tools→system→messages, volatile content last; don't swap tool sets or models mid-conversation.
Mechanics, breakpoints, TTLs, batch lifecycle, tiering math: references/caching-and-cost.md
Context Engineering
Prompt engineering asks what to write in the prompt. Context engineering asks what earns a place in the window on this call — including everything that lands there without you typing it: tool definitions, tool results, retrieved documents, prior turns, thinking blocks. It is iterative (every inference) where prompt engineering is discrete (written once). Target: the smallest set of high-signal tokens that gets the outcome.
The budget is real because attention degrades with length (context rot — n² pairwise relationships), not just because tokens cost money. A 1M window is a capacity, not a target.
The three tiers
Every candidate fact lives in exactly one place. Choosing deliberately is most of the job.
| Tier | Where | Cost | Use when |
|---|---|---|---|
| 1 — In context | tools / system / messages, every call |
Paid every turn (≈0.1× cached) | It steers most turns |
| 2 — On disk, read on demand | A file the agent can read; only the path stays in context | Paid only when read | The agent can tell from a name that it needs this |
| 3 — Retrieved | Index / search tool behind a query | Paid only on a hit, plus a relevance gamble | The corpus is too large to enumerate |
When a prompt is too big, demote before you delete — a path is ~10 tokens; the file it names may be 10,000.
This repo already runs on the tier-1/tier-2 split. A skill's description is
always resident (tier 1, so it must carry the routing signal); SKILL.md loads on a
match; references/*.md load only when cited and needed. "Description is the
trigger", "body under 500 lines", "one concept per reference", "every reference must
be cited" are context-engineering rules wearing authoring clothes.
Cache-aware prompt architecture
Requests render tools → system → messages, and the cache is a prefix match.
So static prefix first, volatile content last — put the cache_control breakpoint
at the end of the stable part and let per-request content fall after it.
Reordering a prompt destroys the cache silently: no error, just a different prefix
hash, cache_read_input_tokens: 0, and a 1.25–2× bill where you expected 0.1×. The
usage block is the only symptom, which is why asserting cache_read_input_tokens > 0
in staging is a real test.
The compaction decision
Under modern prompt caching, keeping the full history has been measured to beat summarisation on cost, latency AND recall at the same time. A 2026 production-tutor evaluation (660 turns, 11 configurations) put keep-everything at 92–100% fact recall, $0.11/turn and 17 s TTFT, against 38–58% recall, $0.24/turn and 21 s for its clear-plus-summarise preset. Summarising rewrites the cached prefix and forfeits the 0.1× discount — the cheap move is usually to append. (That study ran on a non-Claude model; what transfers is the mechanism, and Claude's flat 0.1× cache read makes it stronger, not weaker. Full caveats in references/compaction.md.)
So: compact only as a deliberate response to a named constraint.
| Constraint | Diagnose | Try first |
|---|---|---|
| Context ceiling — it will not fit | Projected tokens > window | Cap tool output → payloads to files → server-side clearing |
| Cost ceiling — the bill is unacceptable | Compare against cached cost, not uncached | Verify the cache is hitting → tier down → cap tool output |
| Latency target — TTFT too slow at depth | Confirm growth is in the prefix | Cap tool output → lower effort → stream |
Capping tool output at the tool boundary is the underrated lever: it shrinks context without rewriting the cached prefix (the same study measured −38% cost/turn with no recall loss). Clearing and summarising both break the cache; they are what people reach for first and should reach for last.
First-party clearing is context_management (beta context-management-2025-06-27):
clear_tool_uses_20250919 and clear_thinking_20251015, applied server-side. Always
set clear_at_least — it stops a trigger paying a full cache re-write to save a
handful of tokens. Pair with the memory tool so durable conclusions are written out
before raw material is cleared.
Agentic specifics
- Tool results are the growth term, not the system prompt. Design tools to return decisions, not dumps.
- Summarise vs write-to-file: needed later in full → write to a file, return the path. Only the conclusion matters → summarise at the tool boundary (free of cache cost, unlike rewriting history after the fact).
- Sub-agents are context isolation, not just parallelism: 80K tokens of exploration are billed once inside the child and discarded; the parent sees a ~1–2K-token distillation. Costs: cold cache in the child, a lossy hand-off. Skip it when the subtask needs most of the parent's context to make sense.
Full doctrine — tiers, progressive disclosure, instrumentation:
references/context-engineering.md.
Compaction economics, context_management parameters, memory tool:
references/compaction.md.
For Claude Code's own context surface see the claude-code-ops skill; for
prompts re-sent on a cadence, loop-ops; for cross-provider fan-out, fleetflow.
Claude Agent SDK (quick reference)
# pip install claude-agent-sdk (Python >= 3.10)
import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions
async def main():
async for message in query(
prompt="Find and fix the bug in auth.py",
options=ClaudeAgentOptions(allowed_tools=["Read", "Edit", "Bash"]),
):
if hasattr(message, "result"):
print(message.result)
asyncio.run(main())
// npm install @anthropic-ai/claude-agent-sdk
import { query } from "@anthropic-ai/claude-agent-sdk";
for await (const message of query({
prompt: "Find and fix the bug in auth.ts",
options: { allowedTools: ["Read", "Edit", "Bash"] },
})) {
if ("result" in message) console.log(message.result);
}
Built-in tools (Read/Write/Edit/Bash/Glob/Grep/WebSearch/WebFetch/...), hooks
(PreToolUse, PostToolUse, ...), subagents, MCP servers, sessions
(resume/fork), permission modes, and the SDK-vs-raw-API decision:
references/agent-sdk.md
Common Pitfalls
| Pitfall | Symptom | Fix |
|---|---|---|
| Date-suffixed or guessed model ID | 404 not_found_error |
Use exact alias IDs from the table above |
budget_tokens on Opus 4.7+ (incl. Opus 5 / Sonnet 5 / Fable 5) |
400 | thinking: {"type": "adaptive"} + effort |
| Assuming thinking is opt-in on Fable 5 / Opus 5 / Sonnet 5 | Unexpected thinking tokens billed | Adaptive thinking is on by default there; Fable 5 can't be disabled at all |
thinking: {"type": "disabled"} at xhigh/max on Opus 5 |
400 | Drop effort to high or below, or leave thinking on |
temperature/top_p/top_k on Opus 4.7+ |
400 (or TypeError on Python SDK v1.0+) |
Remove; steer via prompt + effort |
effort on Haiku 4.5 |
400 | Haiku 4.5 doesn't support the parameter |
Rebuilding assistant turns in a tool loop (dropping empty thinking blocks) |
400 "thinking blocks cannot be modified" | Echo the content list back exactly as received |
| Assistant-turn prefill on Opus 4.7+ models | 400 | output_config.format or system-prompt instruction |
| Cache marker on <minimum prefix | Silent no-cache (cache_creation_input_tokens: 0) |
Min 512-4096 tokens depending on model (see caching ref) |
Not handling stop_reason: "tool_use" |
Agent "stops" after first tool call | Loop: execute tools, append tool_result, re-request |
Missing tool_result for a tool_use id |
400 on follow-up | One tool_result per tool_use block, ids matching |
Non-streaming with max_tokens > ~16K |
SDK timeout / ValueError |
Stream + get_final_message() / finalMessage() |
output_format top-level param |
Deprecated | output_config: {"format": {...}} |
| tiktoken for Claude token counts | 15-20%+ undercount | messages.count_tokens endpoint |
| String-matching error messages | Fragile retries | Typed exceptions: anthropic.RateLimitError etc. |
Raw string-matching tool input |
Breaks on escaping changes | Always json.loads() / use parsed block.input |
| Compacting by reflex on a long conversation | Higher cost, worse recall than doing nothing | Name the constraint first; under caching, appending usually wins (see Context Engineering) |
clear_tool_uses without clear_at_least |
A full cache re-write to reclaim a few hundred tokens | Set clear_at_least so each cache break is worth taking |
Resources & Verification
This skill ships a staleness verifier and two copy-and-adapt starter assets. The model table and pricing above are the facts most likely to drift — run the verifier when you suspect they're stale.
scripts/check-model-table.py — guards the Current Models table (this file)
and the per-model prompt-cache minimum table
(references/caching-and-cost.md) against drift.
Two modes per the resource protocol §7:
# Structural (default, no network): every row well-formed, ids carry no date
# suffix, prices numeric, the two files agree on the model lineup. It also
# guards the cache-economics constants that are stated in more than one file
# (0.1x read, 1.25x/2x writes, 4 breakpoints, 20-block lookback, the
# context-management beta id), asserts each doctrine reference carries a
# "verified <ISO date>" stamp, and checks SKILL.md <-> references/ citation
# integrity in both directions. It then scans every file in the skill for model
# ids: an id in neither the table nor the Legacy list is flagged "unknown", and
# a LEGACY id sitting where a reader would copy it (model=..., "model": ...,
# --model ...) is flagged "retired" - append a `legacy-ok` comment to that line
# for a deliberate migration example. Exit 4 on any contradiction.
python skills/claude-api-ops/scripts/check-model-table.py --offline
python skills/claude-api-ops/scripts/check-model-table.py --offline --json | python -m json.tool
# Live (advisory, needs ANTHROPIC_API_KEY): curls the Models API and compares
# its id set against the documented ids. Exit 10 if a documented id is gone or a
# newer alias id is missing from the table; exit 7 (not a failure) if the key is
# unset or the API is unreachable. Live mode checks model-ID coverage ONLY — the
# API returns no pricing, so pricing/context drift stays an --offline + docs concern.
ANTHROPIC_API_KEY=sk-... python skills/claude-api-ops/scripts/check-model-table.py --live
scripts/context-budget.py — append-vs-compact calculator. Models both paths
in dollars over the turns you actually have left, checks the context ceiling
first, and exits 10 when cost favours compaction, 0 when appending wins:
# Short session — appending is cheaper (exit 0)
python skills/claude-api-ops/scripts/context-budget.py \
--history-tokens 25000 --turns-remaining 5 --base-rate 0.30
# Deep session — cost favours compaction (exit 10)
python skills/claude-api-ops/scripts/context-budget.py \
--history-tokens 120000 --turns-remaining 40 --base-rate 2.00 --json
Two results worth knowing before you trust it: the break-even turn count is scale-invariant (history size and price cancel out — it tracks the summary ratio, not how big or costly the conversation is), and it prices cost only. Recall loss is not in the model, so "compact" means cheaper, not better.
assets/cached-agent-loop.py — the cache-aware sibling of the minimal loop
below, and the executable form of the Context Engineering section: breakpoint at
the end of the static prefix, a rolling breakpoint on the newest turn, an
intermediate anchor every ~15 blocks so long tool-heavy turns don't jump the
20-block lookback, tool output capped at the boundary, and a per-turn
cache_read_input_tokens check that warns when the prefix silently changed.
Copy it when the agent is long-running; copy agentic-loop.py when it isn't.
The footgun it encodes: cache_control is a key on a content block, so it can
only be set on a dict. Appending response.content verbatim (SDK block
objects) or using the "content": "a string" shorthand leaves nowhere to put a
marker — every breakpoint aimed at those turns is discarded with no error and no
warning. Normalise content to dict blocks before placing breakpoints.
assets/recall-probe.py — the "measure it on your workload" harness:
plants a fact, buries it under N turns, probes for it, and reports recall, cost
per turn and TTFT for append vs compact. Makes real API calls, so start
small (--turns 6 --trials 1). Replace the synthetic filler turns with traffic
from your own logs — that is the point of running it.
assets/agentic-loop.py — a minimal, runnable tool-use loop (define a tool,
call messages.create, loop while stop_reason == "tool_use", append
tool_result, re-request until end_turn). Copy it as the starting point when
building a manual agent loop; the >>> ADAPT marks show what to change.
assets/output-schema.json — a known-good structured-outputs request body in
the canonical output_config.format shape (with additionalProperties: false
and a required array). Copy and reshape schema.properties when adding JSON
outputs; see references/structured-outputs.md
for the rules. (Supported on every current model — Fable 5, Opus 5, Sonnet 5,
Haiku 4.5 — and the legacy 4.5–4.8 line.)
Reference Files
| File | Covers |
|---|---|
| references/messages-api.md | Params, response shape, streaming events, stop reasons, error handling, retries, rate limits |
| references/tool-use.md | Tool definitions, tool_choice, parallel tools, agentic loop, tool results, server tools, tool runners |
| references/caching-and-cost.md | Prompt caching mechanics, Batches API, token counting, model tiering economics |
| references/structured-outputs.md | output_config.format, schema rules/limits, strict tools, parse() helpers, thinking interplay |
| references/agent-sdk.md | Python + TS Agent SDK, ClaudeAgentOptions, hooks, MCP, sessions, SDK vs raw API |
| references/context-engineering.md | Context budget, the three tiers, progressive disclosure, cache-aware ordering, tool-result bloat, sub-agents as isolation, instrumentation |
| references/compaction.md | When compaction is justified, break-even arithmetic, context_management edits, memory tool, how to compact well |
Live Documentation
When cached facts may be stale, WebFetch (append .md for clean markdown):
- Models/pricing:
https://platform.claude.com/docs/en/about-claude/models/overview.md - Messages API:
https://platform.claude.com/docs/en/api/messages - Tool use:
https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview.md - Prompt caching:
https://platform.claude.com/docs/en/build-with-claude/prompt-caching.md - Structured outputs:
https://platform.claude.com/docs/en/build-with-claude/structured-outputs.md - Batches:
https://platform.claude.com/docs/en/build-with-claude/batch-processing.md - Agent SDK:
https://code.claude.com/docs/en/agent-sdk/overview - Context editing:
https://platform.claude.com/docs/en/build-with-claude/context-editing - Context engineering:
https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
Files (claude-mods)
-
assets
-
.gitkeep 0 B · in bundle
-
agentic-loop.py 4.4 KB
#!/usr/bin/env python3 """Minimal, correct tool-use agentic loop on the Anthropic Messages API. The canonical pattern: define a tool, call messages.create, and keep looping while stop_reason == "tool_use" — execute each requested tool, append a tool_result, and re-request until stop_reason == "end_turn". Run: pip install anthropic (then: export ANTHROPIC_API_KEY=sk-...) python agentic-loop.py Copy this file and adapt the >>> ADAPT marks for your own tools. Reflects the current API (model claude-opus-5, typed content blocks). """ # The Anthropic SDK accepts plain dict literals for tools/messages at runtime # (as the official docs show), but its strict TypedDict stubs over-narrow them. # Silence those false positives so this starter stays readable; real apps may # prefer the SDK's typed params (anthropic.types.ToolParam, MessageParam). # pyright: reportArgumentType=false import anthropic client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from the environment MODEL = "claude-opus-5" # >>> ADAPT: pick a tier (see the skill's model table) # --- 1. Define your tool(s) ------------------------------------------------ # The input_schema is JSON Schema. Write a description that says WHEN to call. TOOLS = [ { "name": "get_weather", # >>> ADAPT "description": "Get the current weather for a city. " "Call this whenever the user asks about weather conditions.", "input_schema": { "type": "object", "properties": { "location": {"type": "string", "description": "City, e.g. 'Paris'"}, }, "required": ["location"], }, } ] # --- 2. Implement each tool ------------------------------------------------ # Map tool name -> a Python callable. Never trust the arguments blindly: they # come from the model. Validate before doing anything with side effects. def get_weather(location: str) -> str: # >>> ADAPT: real implementation return f"It is 21°C and sunny in {location}." TOOL_IMPLS = {"get_weather": get_weather} def run_tool(name: str, tool_input: dict) -> str: """Dispatch a tool call, returning a string result for the model.""" impl = TOOL_IMPLS.get(name) if impl is None: return f"ERROR: unknown tool {name!r}" try: return impl(**tool_input) except Exception as exc: # surface failures back to the model, don't crash return f"ERROR running {name}: {exc}" # --- 3. The loop ----------------------------------------------------------- def agent(user_prompt: str, max_turns: int = 10) -> str: # The conversation is a growing list of message dicts we own and replay. messages = [{"role": "user", "content": user_prompt}] for _ in range(max_turns): response = client.messages.create( model=MODEL, max_tokens=4096, tools=TOOLS, messages=messages, ) # Append the assistant turn VERBATIM — content is a list of typed # blocks (text and/or tool_use). It must go back as-is next request. messages.append({"role": "assistant", "content": response.content}) # If the model didn't ask for a tool, we're done — return its text. if response.stop_reason != "tool_use": return "".join( block.text for block in response.content if block.type == "text" ) # Otherwise: execute EVERY tool_use block and collect tool_result # blocks (the model may request several tools in parallel). tool_results = [] for block in response.content: if block.type != "tool_use": continue # skip text/thinking blocks result_text = run_tool(block.name, block.input) tool_results.append({ "type": "tool_result", "tool_use_id": block.id, # MUST echo the matching id "content": result_text, # "is_error": True, # set when the tool failed }) # Feed results back as a single user turn, then loop to re-request. messages.append({"role": "user", "content": tool_results}) return "Stopped: hit max_turns without an end_turn." if __name__ == "__main__": answer = agent("What's the weather in Tokyo right now?") print(answer) # For debugging, inspect the assembled transcript: # print(json.dumps(..., indent=2, default=str)) -
cached-agent-loop.py 14.1 KB
#!/usr/bin/env python3 """A cache-correct agentic loop — the context-engineering doctrine, executable. assets/agentic-loop.py is the MINIMAL correct loop: it shows stop_reason handling and nothing else, on purpose. This file is its cache-aware sibling. Same loop, plus the four things that decide whether a long-running agent costs 0.1x or 1.25x per turn (see SKILL.md "Context Engineering"): 1. STATIC PREFIX FIRST, VOLATILE LAST. tools -> system -> messages is the render order, and the cache is a prefix match. The breakpoint goes at the end of the stable part; anything per-request falls after it. 2. A ROLLING BREAKPOINT on the newest turn, so hits accrue as the conversation grows -- and never more than MAX_BREAKPOINTS of them. 3. AN INTERMEDIATE BREAKPOINT every ~15 blocks. A breakpoint searches backward at most 20 content blocks; a turn that appends more than that jumps the window and silently misses. 4. TOOL OUTPUT CAPPED AT THE BOUNDARY. The cheapest context lever there is: it shrinks context WITHOUT rewriting the cached prefix, so unlike compaction it costs nothing in cache terms. 5. CONTENT NORMALISED TO DICT BLOCKS. cache_control is a key on a content block, so it can only be set on a dict. `response.content` is a list of SDK block OBJECTS and "content": "a string" has no blocks at all -- append either verbatim and every marker aimed at it is silently discarded. An earlier version of this file did exactly that and placed ZERO message breakpoints in a realistic conversation, with no error and no warning. to_blocks() is what makes points 2 and 3 actually take effect. And the one assertion that matters: cache_read_input_tokens > 0. A broken cache produces no error, no warning, and no symptom other than the bill -- this loop prints the usage every turn and complains when the cache misses. Run: pip install anthropic (then: export ANTHROPIC_API_KEY=sk-...) python cached-agent-loop.py Copy this file and adapt the >>> ADAPT marks. Cache facts verified against platform.claude.com 2026-08-30; see references/caching-and-cost.md. """ # The Anthropic SDK accepts plain dict literals for tools/messages at runtime # (as the official docs show), but its strict TypedDict stubs over-narrow them. # Silence those false positives so this starter stays readable. # pyright: reportArgumentType=false import anthropic client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from the environment MODEL = "claude-opus-5" # >>> ADAPT: pick a tier (see the skill's model table) # --- Cache tuning knobs (documented values, not guesses) -------------------- MAX_BREAKPOINTS = 4 # hard API limit: 4 cache_control markers per request LOOKBACK_BLOCKS = 20 # a breakpoint searches back at most 20 content blocks BREAKPOINT_EVERY = 15 # ...so re-anchor before that window is exhausted TOOL_RESULT_CAP = 4000 # >>> ADAPT: chars kept per tool result (lever #4) # The minimum cacheable prefix is MODEL-DEPENDENT (512-4096 tokens). Below it # the marker is silently ignored -- cache_creation_input_tokens stays 0 and # there is no error. If your system prompt is small, caching it buys nothing; # check the per-model table in references/caching-and-cost.md before tuning. # === 1. THE STATIC PREFIX ==================================================== # Everything here must be byte-identical on every request. The classic silent # invalidators live in exactly this block: a datetime.now(), a request id, a # per-user name, a conditionally-appended section, an unsorted json.dumps. # >>> ADAPT: put your real instructions/docs here. Keep them CONSTANT. SYSTEM_PROMPT = """You are a careful assistant with tool access. Prefer calling a tool over guessing. Answer concisely once you have the facts. """ + ("Reference material the agent needs on most turns goes here. " * 200) # Tools render at position 0, ahead of system. Changing ANY tool definition # invalidates the entire cache -- so build this list once, statically. Never # assemble it per-user or per-request. TOOLS = [ { "name": "get_weather", # >>> ADAPT "description": "Get current weather for a city. Call when asked about weather.", "input_schema": { "type": "object", "properties": {"location": {"type": "string", "description": "City, e.g. Paris"}}, "required": ["location"], }, }, ] def run_tool(name: str, tool_input: dict) -> str: """>>> ADAPT: dispatch to your real implementations.""" if name == "get_weather": return f"18C, light rain in {tool_input.get('location', 'unknown')}." return f"No such tool: {name}" def capped(text: str, limit: int = TOOL_RESULT_CAP) -> str: """Lever #4 — cap at the TOOL BOUNDARY, before the result enters context. This is the important asymmetry: truncating here shortens what gets APPENDED, so the cached prefix behind it is untouched. Summarising the same content *after* it is already in the history is compaction, and pays a full cache write. If the caller may need the full payload, write it to a file and return the path instead of truncating (tier 2 -- see references/context-engineering.md §5.2). """ if len(text) <= limit: return text return (text[:limit] + f"\n... [truncated {len(text) - limit} chars. " f"Re-run with a narrower query, or read the full payload from disk.]") # === 2/3. BREAKPOINT PLACEMENT ============================================== def to_blocks(content) -> list: """Normalise any message content into a list of plain dict blocks. THIS IS LOAD-BEARING, not tidiness. cache_control is a key on a content block, so a breakpoint can only be attached to a dict. Two shapes in normal use are NOT dicts and will silently refuse every marker: * `response.content` — SDK block OBJECTS (TextBlock, ToolUseBlock, ...). Appending them verbatim, as the minimal loop does, is idiomatic and correct for a loop that never caches. Here it means every assistant turn is un-markable. * `"content": "a plain string"` — the shorthand form has no blocks at all, so there is nowhere to put a marker. Either one produces NO error and NO warning: the request simply caches less than you think. Normalising at append time is what keeps the breakpoint logic below sound. """ if isinstance(content, str): return [{"type": "text", "text": content}] out = [] for block in content: if isinstance(block, dict): out.append(block) elif hasattr(block, "model_dump"): # pydantic v2 (current SDK) out.append(block.model_dump(exclude_none=True)) elif hasattr(block, "dict"): # pydantic v1 out.append(block.dict(exclude_none=True)) else: # last resort out.append(dict(vars(block))) return out def _blocks(message) -> list: content = message["content"] return content if isinstance(content, list) else [] def place_message_breakpoints(messages: list) -> int: """Re-anchor rolling cache breakpoints across the message list, in place. Returns the number of markers actually placed — check it. Silently placing zero is the failure this function exists to prevent. Rules encoded here: * The newest turn always carries a breakpoint, so the next request can read everything up to it (hits accrue as the conversation grows). * An extra anchor every BREAKPOINT_EVERY blocks, because the backward search stops after LOOKBACK_BLOCKS (20). A tool-heavy turn that appends 30 blocks would otherwise jump clean over the previous entry. * At most MAX_BREAKPOINTS - 1 markers here; the system block owns the fourth. Exceeding 4 is an API error, so oldest markers are dropped. * A marker can only sit on a dict block. If the chosen position is not one (an un-normalised SDK object, say), walk BACKWARD to the nearest dict rather than dropping the anchor — dropping it silently is exactly how a loop ends up with no caching at all. Run content through to_blocks() and this path never triggers. """ assert BREAKPOINT_EVERY < LOOKBACK_BLOCKS, "anchor must fall inside the window" # Clear existing markers, then re-place: idempotent, so the loop can call # this every turn without accumulating stale markers. for msg in messages: for block in _blocks(msg): if isinstance(block, dict): block.pop("cache_control", None) # Walk the flattened block stream, marking a candidate every N blocks and # always marking the final block. Positions count EVERY block (the API sees # them all), even ones that cannot themselves carry a marker. flat = [(mi, bi) for mi, msg in enumerate(messages) for bi, _ in enumerate(_blocks(msg))] if not flat: # Every message used the plain-string shorthand: nothing can be cached. print(" WARNING: no content blocks to anchor a cache breakpoint on.\n" " Message content is in string shorthand - run it through\n" " to_blocks() so cache_control has somewhere to attach.") return 0 def settable(idx: int): """Nearest dict block at or before idx, within the lookback window.""" for k in range(idx, max(-1, idx - LOOKBACK_BLOCKS), -1): mi, bi = flat[k] if isinstance(_blocks(messages[mi])[bi], dict): return k return None chosen: list[int] = [] for i in list(range(BREAKPOINT_EVERY - 1, len(flat), BREAKPOINT_EVERY)) + [len(flat) - 1]: k = settable(i) if k is not None and k not in chosen: chosen.append(k) # Keep the most recent ones; the system breakpoint consumes one of the 4. placed = 0 for k in chosen[-(MAX_BREAKPOINTS - 1):]: mi, bi = flat[k] block = _blocks(messages[mi])[bi] if isinstance(block, dict): block["cache_control"] = {"type": "ephemeral"} placed += 1 # >>> ADAPT: {"type": "ephemeral", "ttl": "1h"} if turns are minutes # apart. 1h writes cost 2x vs 1.25x, so it needs 3+ reads to pay off. if placed == 0: print(" WARNING: placed 0 message cache breakpoints - the conversation\n" " will re-read from the system breakpoint only. Normalise\n" " message content with to_blocks().") return placed def report_usage(usage, turn: int, first_turn: bool) -> None: """The only symptom a broken cache produces is in here. Look at it.""" read = getattr(usage, "cache_read_input_tokens", 0) or 0 written = getattr(usage, "cache_creation_input_tokens", 0) or 0 print(f" [turn {turn}] uncached={usage.input_tokens} " f"cache_write={written} cache_read={read} out={usage.output_tokens}") if not first_turn and read == 0: # Not an exception on purpose: this is a cost bug, not a crash. In CI or # staging, promote it -- an assert here catches a stray timestamp in the # system prompt before it reaches production. print(" WARNING: cache_read_input_tokens == 0 on a repeat request.\n" " The prefix changed. Check for a timestamp/uuid/per-user\n" " value in SYSTEM_PROMPT, a reordered block, an unsorted\n" " json.dumps, a swapped model, or a mutated tool list.\n" " See references/caching-and-cost.md, 'Silent invalidators'.") def main() -> None: # >>> ADAPT: your opening request. Volatile content belongs HERE, after the # cached system block -- never interpolated into SYSTEM_PROMPT. # Block form, not the "content": "..." string shorthand — a string has no # block for cache_control to attach to. messages = [{"role": "user", "content": [{"type": "text", "text": "What's the weather in Paris and in Oslo?"}]}] for turn in range(1, 21): # bounded: never ship an unbounded agent loop place_message_breakpoints(messages) # returns the count; 0 means no caching response = client.messages.create( model=MODEL, max_tokens=4096, # The static prefix, with the breakpoint at its end. Tools render # ahead of system, so this one marker caches BOTH together. system=[{"type": "text", "text": SYSTEM_PROMPT, "cache_control": {"type": "ephemeral"}}], tools=TOOLS, messages=messages, ) report_usage(response.usage, turn, first_turn=(turn == 1)) if response.stop_reason != "tool_use": for block in response.content: if block.type == "text": print(block.text) return # to_blocks(), not response.content verbatim: SDK block objects cannot # carry cache_control, so appending them raw silently un-caches every # assistant turn (no error, just a bigger bill). messages.append({"role": "assistant", "content": to_blocks(response.content)}) tool_results = [] for block in response.content: if block.type != "tool_use": continue try: # block.input is already parsed -- never string-match the raw text. output = run_tool(block.name, block.input) is_error = False except Exception as exc: # tool failures are data, not crashes output, is_error = f"Tool failed: {exc}", True tool_results.append({ "type": "tool_result", "tool_use_id": block.id, # one result per tool_use, ids matching "content": capped(output), # <- lever #4, at the boundary "is_error": is_error, }) messages.append({"role": "user", "content": tool_results}) print("Hit the turn ceiling without an end_turn - check the tool loop.") if __name__ == "__main__": main() -
output-schema.json 1.4 KB
{ "_comment": "Canonical structured-outputs request body for the Anthropic Messages API. Pass the `output_config` object below in client.messages.create(...). This is the current `output_config.format` shape — NOT the deprecated top-level `output_format` param. Adapt `schema.properties` and `required` to your data; keep `additionalProperties: false` so the model can't add stray keys. Supported on Fable 5, Opus 5, Sonnet 5, Haiku 4.5, and the legacy Opus 4.8/4.7/4.6/4.5 and Sonnet 4.6/4.5 line — use plain alias ids.", "model": "claude-opus-5", "max_tokens": 1024, "messages": [ { "role": "user", "content": "Extract the contact: Jane Doe (jane@example.com) wants the Enterprise plan and asked for a demo." } ], "output_config": { "format": { "type": "json_schema", "schema": { "type": "object", "properties": { "name": { "type": "string", "description": "Full name of the contact" }, "email": { "type": "string", "format": "email" }, "plan": { "type": "string", "enum": ["Free", "Pro", "Enterprise"] }, "demo_requested": { "type": "boolean", "description": "Whether the contact asked for a demo" } }, "required": ["name", "email", "plan", "demo_requested"], "additionalProperties": false } } } } -
recall-probe.py 10 KB
#!/usr/bin/env python3 """Measure what compaction actually costs you — on YOUR workload. references/compaction.md says the keep-everything baseline is far stronger than teams assume, and that you should verify that on your own traffic before compacting. This is the harness for doing so. It is the published methodology, runnable: plant a fact, bury it under N turns, probe for it, and compare strategies on the three axes people reach for compaction to improve. plant (turn 0) -> N filler turns -> probe ("what was the code?") | recall hit/miss + $ per turn + time to first token Strategies compared out of the box: append keep the full history every turn (the baseline) compact summarise everything older than --keep-recent turns, then continue from the summary (what most agent frameworks do by default) The point is NOT to reproduce a published number. It is to find out whether YOUR workload behaves like the studied one, because the answer decides whether compaction is buying you anything. A strategy that wins on cost and loses on recall has not won. Run: pip install anthropic (then: export ANTHROPIC_API_KEY=sk-...) python recall-probe.py --turns 10 --trials 3 Costs real money — it makes (turns + 2) x trials x strategies API calls. Start with --turns 6 --trials 1 on a cheap model to sanity-check the wiring. >>> ADAPT: replace FILLER_TURNS with real turns from your own application. Synthetic filler is the weakest part of any harness like this — a planted fact survives generic chit-chat far more easily than it survives your actual tool-heavy traffic, so results on the shipped filler will flatter both strategies. """ # pyright: reportArgumentType=false import argparse import statistics import sys import time import anthropic client = anthropic.Anthropic() MODEL = "claude-haiku-4-5" # >>> ADAPT: probe on the tier you actually ship SYSTEM_PROMPT = ("You are a helpful assistant. Answer concisely. " "When asked to recall a specific detail from earlier in the " "conversation, quote it exactly.") # >>> ADAPT: the fact to plant. Make it arbitrary and unguessable, so a correct # answer proves recall rather than a plausible reconstruction from priors. SECRET_LABEL = "deployment code" SECRET_VALUE = "TANGERINE-4471-QUAY" PLANT = (f"Before we start: the {SECRET_LABEL} for this session is " f"{SECRET_VALUE}. Acknowledge it and we'll move on.") PROBE = f"What was the {SECRET_LABEL} I gave you at the start? Answer with just the code." # >>> ADAPT: replace with turns sampled from your own logs. FILLER_TURNS = [ "Explain the difference between a process and a thread.", "Now give me a short example of a race condition.", "How would you fix it with a mutex?", "What's the tradeoff versus a lock-free queue?", "Summarise when I should prefer each.", "What about async I/O — where does that fit?", "Give me a rule of thumb for choosing a concurrency model.", "What metrics would tell me I chose wrong?", "How would I load-test that?", "What's a common mistake teams make here?", ] COMPACT_INSTRUCTION = ( # Per Anthropic's guidance: maximise RECALL first, then tune precision. A # compaction prompt written for brevity silently drops the one detail the # next 40 turns needed -- which is precisely what this harness measures. "Summarise the conversation so far. Preserve every concrete detail, " "identifier, number, name and code mentioned — completeness matters more " "than brevity. Write it as notes for someone continuing the conversation." ) def send(messages, cache=True): """One request. Returns (text, usage, time_to_first_token_seconds).""" system = [{"type": "text", "text": SYSTEM_PROMPT}] if cache: # Cache the static prefix -- otherwise 'append' is measured without the # very mechanism that makes it competitive, and the comparison is rigged. system[0]["cache_control"] = {"type": "ephemeral"} start = time.perf_counter() ttft = None chunks = [] with client.messages.stream(model=MODEL, max_tokens=1024, system=system, messages=messages) as stream: for text in stream.text_stream: if ttft is None: ttft = time.perf_counter() - start chunks.append(text) final = stream.get_final_message() return "".join(chunks), final.usage, (ttft if ttft is not None else 0.0) def cost_usd(usage, in_rate: float, out_rate: float) -> float: """Bill the three input buckets at their real multipliers, plus output.""" read = getattr(usage, "cache_read_input_tokens", 0) or 0 write = getattr(usage, "cache_creation_input_tokens", 0) or 0 return (usage.input_tokens * in_rate + write * in_rate * 1.25 # 5-minute-TTL cache write + read * in_rate * 0.10 # cache read, flat across models + usage.output_tokens * out_rate) / 1_000_000 def run_trial(strategy: str, turns: int, keep_recent: int, in_rate: float, out_rate: float, compact_every: int) -> dict: messages = [{"role": "user", "content": PLANT}] total_cost, ttfts, compactions = 0.0, [], 0 turns_since_compaction = 0 text, usage, ttft = send(messages) total_cost += cost_usd(usage, in_rate, out_rate) ttfts.append(ttft) messages.append({"role": "assistant", "content": text}) for i in range(turns): messages.append({"role": "user", "content": FILLER_TURNS[i % len(FILLER_TURNS)]}) text, usage, ttft = send(messages) total_cost += cost_usd(usage, in_rate, out_rate) ttfts.append(ttft) messages.append({"role": "assistant", "content": text}) turns_since_compaction += 1 if (strategy == "compact" and turns_since_compaction >= compact_every and len(messages) > 2 * keep_recent + 1): # Summarise everything except the most recent keep_recent exchanges, # then continue from the summary. This is the default behaviour of # most agent frameworks -- and the thing under test. # # FAIRNESS, and the reason for the cadence: an earlier version of # this harness compacted on EVERY turn once the history exceeded # keep_recent, so the compact arm paid a summarisation call per turn # (21 API calls against append's 12, over 10 turns). That inflates # the arm under test and would "confirm" the keep-everything result # regardless of what the data said. Real systems trigger on a token # threshold (context_management's `trigger`); a turn cadence is the # closest honest approximation without burning a count_tokens call # every turn. head, tail = messages[:-2 * keep_recent], messages[-2 * keep_recent:] summary_text, usage, _ = send( head + [{"role": "user", "content": COMPACT_INSTRUCTION}]) total_cost += cost_usd(usage, in_rate, out_rate) messages = ([{"role": "user", "content": f"[Summary of earlier conversation]\n{summary_text}"}, {"role": "assistant", "content": "Understood, continuing."}] + tail) compactions += 1 turns_since_compaction = 0 messages.append({"role": "user", "content": PROBE}) answer, usage, ttft = send(messages) total_cost += cost_usd(usage, in_rate, out_rate) ttfts.append(ttft) # Normalised substring match. Deliberately generous: a strategy that cannot # pass THIS has lost the fact outright, not merely paraphrased it. recalled = SECRET_VALUE.lower() in answer.lower() return {"recalled": recalled, "cost": total_cost, "compactions": compactions, "ttft": statistics.mean(ttfts), "answer": answer.strip()[:120]} def main() -> int: ap = argparse.ArgumentParser( description="Plant-a-fact recall probe: append vs compact, on your workload.") ap.add_argument("--turns", type=int, default=10, help="filler turns before the probe") ap.add_argument("--trials", type=int, default=3, help="repeats per strategy") ap.add_argument("--keep-recent", type=int, default=2, help="exchanges kept verbatim by the compact strategy") ap.add_argument("--compact-every", type=int, default=5, help="turns between compactions (default: 5). Setting this " "to 1 compacts every turn, which inflates the compact " "arm and rigs the result toward keep-everything - only " "do it if you are deliberately modelling that.") ap.add_argument("--in-rate", type=float, default=1.00, help="input $/MTok") ap.add_argument("--out-rate", type=float, default=5.00, help="output $/MTok") args = ap.parse_args() print(f"model={MODEL} turns={args.turns} trials={args.trials} compact_every={args.compact_every}\n") for strategy in ("append", "compact"): results = [run_trial(strategy, args.turns, args.keep_recent, args.in_rate, args.out_rate, args.compact_every) for _ in range(args.trials)] hits = sum(r["recalled"] for r in results) per_turn = statistics.mean(r["cost"] for r in results) / (args.turns + 2) print(f"{strategy:8s} recall {hits}/{args.trials} " f"${per_turn:.5f}/turn " f"ttft {statistics.mean(r['ttft'] for r in results):.2f}s " f"compactions {statistics.mean(r['compactions'] for r in results):.1f}") for r in results: if not r["recalled"]: print(f" miss -> {r['answer']!r}") print("\nIf append wins on all three, compaction is costing you twice. If " "compact wins on cost alone, price the recall loss before adopting it.") return 0 if __name__ == "__main__": try: sys.exit(main()) except KeyboardInterrupt: sys.exit(2)
-
-
references
-
agent-sdk.md 9.7 KB
# Claude Agent SDK Reference The Agent SDK is **Claude Code as a library**: the same agent loop, built-in tools (file ops, bash, search, web), context management, and permission system, programmable from Python or TypeScript. Verified against code.claude.com/docs/en/agent-sdk (2026-06). | | Python | TypeScript | |---|---|---| | Package | `claude-agent-sdk` (`pip install claude-agent-sdk`) | `@anthropic-ai/claude-agent-sdk` (`npm install @anthropic-ai/claude-agent-sdk`) | | Runtime | Python ≥ 3.10 | Node (bundles a native Claude Code binary — no separate install) | | Entry point | `query(prompt, options=ClaudeAgentOptions(...))` → async iterator | `query({ prompt, options })` → async iterator | | Option naming | `snake_case` (`allowed_tools`) | `camelCase` (`allowedTools`) | Auth: `ANTHROPIC_API_KEY`, or third-party providers via env flags (`CLAUDE_CODE_USE_BEDROCK=1`, `CLAUDE_CODE_USE_VERTEX=1`, `CLAUDE_CODE_USE_FOUNDRY=1`, `CLAUDE_CODE_USE_ANTHROPIC_AWS=1` + `ANTHROPIC_AWS_WORKSPACE_ID`). Note: from June 15, 2026, Agent SDK / `claude -p` usage on subscription plans draws from a separate monthly Agent SDK credit — production apps should use API keys. ## When SDK vs raw API | Signal | Choice | |---|---| | Agent must read/edit files, run commands, search a codebase | **Agent SDK** — tools are built in, loop is handled | | CI/CD automation, "fix the failing test", repo-scale refactors | **Agent SDK** | | You want hooks, permission gating, session resume out of the box | **Agent SDK** | | Single call: classify/summarize/extract | **Messages API** — SDK is overkill | | You need exact control of every request (params, caching breakpoints, message shapes) | **Messages API + tool use** | | Your tools are pure in-process functions, no filesystem | **Messages API + tool runner** is lighter | | Hosted agent, no infra at all | **Managed Agents** (REST, Anthropic-run sandbox) | The difference in code: ```python # Messages API: you implement the loop response = client.messages.create(...) while response.stop_reason == "tool_use": result = your_tool_executor(...) response = client.messages.create(...) # Agent SDK: the loop, tools, and context management are inside query() async for message in query(prompt="Fix the bug in auth.py"): print(message) ``` ## Minimal agents ```python import asyncio from claude_agent_sdk import query, ClaudeAgentOptions async def main(): async for message in query( prompt="Find all TODO comments and create a summary", options=ClaudeAgentOptions(allowed_tools=["Read", "Glob", "Grep"]), ): if hasattr(message, "result"): # final ResultMessage print(message.result) asyncio.run(main()) ``` ```typescript import { query } from "@anthropic-ai/claude-agent-sdk"; for await (const message of query({ prompt: "Find all TODO comments and create a summary", options: { allowedTools: ["Read", "Glob", "Grep"] }, })) { if ("result" in message) console.log(message.result); } ``` The iterator yields typed messages: system messages (`subtype: "init"` carries `session_id`), assistant/tool activity, and a final result message (`message.result` / `ResultMessage`). ## ClaudeAgentOptions (key fields) | Python | TypeScript | Purpose | |---|---|---| | `allowed_tools` | `allowedTools` | Pre-approved tools (no permission prompt). Include `"Agent"` to auto-approve subagent spawns | | `disallowed_tools` | `disallowedTools` | Hard-blocked tools | | `permission_mode` | `permissionMode` | e.g. `"default"`, `"acceptEdits"`, `"bypassPermissions"`, `"plan"` | | `system_prompt` | `systemPrompt` | Replace or extend the system prompt | | `mcp_servers` | `mcpServers` | MCP server map (see below) | | `hooks` | `hooks` | Lifecycle callbacks (see below) | | `agents` | `agents` | Named subagent definitions | | `resume` | `resume` | Session ID to continue with full context | | `cwd` | `cwd` | Working directory for the agent | | `model` | `model` | Override model | | `max_turns` | `maxTurns` | Cap agent-loop iterations | | `setting_sources` | `settingSources` | Restrict which filesystem config loads (`.claude/`, `~/.claude/`) | | `plugins` | `plugins` | Programmatic plugin loading | ### Built-in tools `Read`, `Write`, `Edit`, `Bash`, `Glob`, `Grep`, `WebSearch`, `WebFetch`, `Monitor` (watch a background process), `AskUserQuestion` (clarifying questions with options), `Agent` (spawn subagents). A read-only agent is just `allowed_tools=["Read", "Glob", "Grep"]`. ## Hooks Callbacks at lifecycle points: `PreToolUse`, `PostToolUse`, `Stop`, `SessionStart`, `SessionEnd`, `UserPromptSubmit`, and more. Use them to audit, block, or transform agent behavior. ```python from claude_agent_sdk import query, ClaudeAgentOptions, HookMatcher from datetime import datetime async def log_file_change(input_data, tool_use_id, context): path = input_data.get("tool_input", {}).get("file_path", "unknown") with open("./audit.log", "a") as f: f.write(f"{datetime.now()}: modified {path}\n") return {} # empty dict = allow; hooks can also block/modify options = ClaudeAgentOptions( permission_mode="acceptEdits", hooks={"PostToolUse": [HookMatcher(matcher="Edit|Write", hooks=[log_file_change])]}, ) ``` ```typescript import { query, HookCallback } from "@anthropic-ai/claude-agent-sdk"; import { appendFile } from "fs/promises"; const logFileChange: HookCallback = async (input) => { const filePath = (input as any).tool_input?.file_path ?? "unknown"; await appendFile("./audit.log", `${new Date().toISOString()}: modified ${filePath}\n`); return {}; }; const options = { permissionMode: "acceptEdits" as const, hooks: { PostToolUse: [{ matcher: "Edit|Write", hooks: [logFileChange] }] }, }; ``` `matcher` is a regex over tool names. A `PreToolUse` hook returning a deny decision blocks the call — this is the programmatic equivalent of a human approval gate. ## MCP servers Wire any MCP server (stdio command or remote) into the agent: ```python options = ClaudeAgentOptions( mcp_servers={ "playwright": {"command": "npx", "args": ["@playwright/mcp@latest"]}, }, ) ``` ```typescript const options = { mcpServers: { playwright: { command: "npx", args: ["@playwright/mcp@latest"] }, }, }; ``` MCP tools surface as `mcp__<server>__<tool>` — add them to `allowed_tools` to pre-approve. This is how you give an agent databases, browsers, and third-party APIs without writing tool plumbing. ## Subagents ```python from claude_agent_sdk import query, ClaudeAgentOptions, AgentDefinition options = ClaudeAgentOptions( allowed_tools=["Read", "Glob", "Grep", "Agent"], # Agent tool approves spawns agents={ "code-reviewer": AgentDefinition( description="Expert code reviewer for quality and security reviews.", prompt="Analyze code quality and suggest improvements.", tools=["Read", "Glob", "Grep"], ), }, ) # prompt: "Use the code-reviewer agent to review this codebase" ``` Messages emitted inside a subagent carry `parent_tool_use_id` so you can attribute output to the spawning call. Use subagents to isolate context (reviewer doesn't pollute the main transcript) and to parallelize independent legs. ## Sessions: resume and fork Session state is JSONL on your filesystem. Capture the session ID from the init message, resume later with full context: ```python from claude_agent_sdk import query, ClaudeAgentOptions, SystemMessage, ResultMessage session_id = None async for message in query(prompt="Read the authentication module", options=ClaudeAgentOptions(allowed_tools=["Read", "Glob"])): if isinstance(message, SystemMessage) and message.subtype == "init": session_id = message.data["session_id"] async for message in query(prompt="Now find all places that call it", options=ClaudeAgentOptions(resume=session_id)): if isinstance(message, ResultMessage): print(message.result) ``` ```typescript let sessionId: string | undefined; for await (const message of query({ prompt: "Read the authentication module", options: { allowedTools: ["Read", "Glob"] } })) { if (message.type === "system" && message.subtype === "init") { sessionId = message.session_id; } } for await (const message of query({ prompt: "Now find all places that call it", options: { resume: sessionId } })) { if ("result" in message) console.log(message.result); } ``` Sessions can also be forked to explore alternative approaches from the same context point. ## Filesystem configuration With default options the SDK loads Claude Code's filesystem config from `.claude/` (project) and `~/.claude/` (user): skills (`.claude/skills/*/SKILL.md`), commands, `CLAUDE.md` memory, plugins. Restrict with `setting_sources` / `settingSources` when you want a hermetic agent (CI) that ignores developer-machine state. ## Production patterns - **CI agent:** `permission_mode="bypassPermissions"` (or a tight `allowed_tools` list) + `max_turns` cap + hooks for audit logging. Never bypass permissions on a machine with credentials you don't want the agent exercising. - **Approval gate:** `PreToolUse` hook on `Bash|Write|Edit` that checks the input and returns deny for out-of-policy actions. - **Observability:** log every message from the iterator; hook `PostToolUse` for tool-level metrics; final `ResultMessage` includes cost/usage data. - **Prototype → production:** a common path is Agent SDK locally (works on your filesystem), then Managed Agents for hosted production (Anthropic runs the sandbox; custom tools become event round-trips). - Workflows translate 1:1 with the Claude Code CLI (`claude -p`) — anything you can do interactively you can automate via the SDK. -
caching-and-cost.md 11.1 KB
# Caching, Batches & Cost Reference The three big cost levers, in order of typical impact: model tiering, prompt caching, Batches API. They stack — a cached Haiku batch request can cost ~2-3% of an uncached Opus interactive request for the same tokens. ## Prompt Caching ### The one invariant **Caching is a prefix match.** The cache key is the exact bytes of the rendered prompt up to each `cache_control` breakpoint. One byte changed anywhere in the prefix invalidates everything after it. Render order is `tools` → `system` → `messages` — a breakpoint on the last system block caches tools + system together. ### Syntax ```python response = client.messages.create( model="claude-opus-5", max_tokens=16000, system=[{ "type": "text", "text": LARGE_STABLE_PROMPT, # 50KB of docs, instructions... "cache_control": {"type": "ephemeral"}, # 5-minute TTL (default) # "cache_control": {"type": "ephemeral", "ttl": "1h"} # 1-hour TTL }], messages=[{"role": "user", "content": question}], ) ``` ```typescript const response = await client.messages.create({ model: "claude-opus-5", max_tokens: 16000, system: [{ type: "text", text: LARGE_STABLE_PROMPT, cache_control: { type: "ephemeral" } }], messages: [{ role: "user", content: question }], }); ``` Simplest option — top-level auto-caching (caches the last cacheable block, no per-block markers): ```python client.messages.create(model="claude-opus-5", max_tokens=16000, cache_control={"type": "ephemeral"}, system=big_doc, messages=[...]) ``` Rules: - Max **4** `cache_control` breakpoints per request. - Valid on system text blocks, tool definitions, and message content blocks (`text`, `image`, `tool_use`, `tool_result`, `document`). - **Minimum cacheable prefix is model-dependent** — below it the marker is silently ignored (no error, just `cache_creation_input_tokens: 0`): | Model | Minimum prefix tokens | |---|---:| | Opus 4.6 / 4.5, Haiku 4.5 | 4096 | | Opus 4.7 | 2048 | | Sonnet 5, Opus 4.8, Sonnet 4.6, Sonnet 4.5 | 1024 | | Opus 5, Fable 5 | 512 | Note the shape: the newest models cache *sooner* (512) while Haiku 4.5 — the cheapest model, and the one you'd most want to cache in bulk — needs the largest prefix (4096). Sizing a shared prefix for Haiku covers every other model too. ### Pricing & break-even | Operation | Cost vs base input | |---|---| | Cache write, 5-min TTL | 1.25x | | Cache write, 1-hour TTL | 2x | | Cache read | ~0.1x | Break-even: 5-min TTL pays off at the **2nd** request (1.25 + 0.1 = 1.35x vs 2x); 1-hour TTL needs **3+** requests (2 + 0.2 = 2.2x vs 3x). Steady-state savings on a large cached prefix approach 90%. ### Multi-turn / agent placement - **Multi-turn:** put the breakpoint on the last content block of the latest turn; earlier breakpoints remain valid read points, so hits accrue as the conversation grows. Top-level auto-caching does this for you. - **Shared prefix, varying question:** breakpoint at the end of the *shared* part, not the end of the prompt — otherwise every request writes a distinct entry and nothing is ever read. - **20-block lookback:** a breakpoint searches backward at most 20 content blocks for a prior entry. Agent turns adding >20 blocks (many tool_use/tool_result pairs) silently miss — add an intermediate breakpoint every ~15 blocks in long turns. - **Concurrent fan-out:** an entry becomes readable only once the first response starts streaming. Fire 1 request, await first token, then fire the other N-1 so they read the fresh cache. ### Silent invalidators (audit checklist) If `usage.cache_read_input_tokens` stays 0 across identical-prefix requests, grep the prompt-assembly path for: | Pattern | Why it kills the cache | |---|---| | `datetime.now()` / `Date.now()` in the system prompt | New prefix every request | | `uuid4()` / request IDs early in content | Same | | `json.dumps(d)` without `sort_keys=True` | Non-deterministic bytes | | Per-user IDs interpolated into the system prompt | No cross-user sharing | | Conditional system sections (`if flag: system += ...`) | Each flag combo = distinct prefix | | Tool set built per-user / unsorted | Tools render at position 0 — full invalidation | | Switching models mid-conversation | Caches are model-scoped | Fix: move volatile content **after** the last breakpoint (or into the latest user message), serialize deterministically, freeze the system prompt and tool list. ### Verifying ```python u = response.usage print(u.cache_creation_input_tokens) # written this request (paid 1.25-2x) print(u.cache_read_input_tokens) # served from cache (paid ~0.1x) print(u.input_tokens) # uncached remainder (full price) # total prompt = input + cache_creation + cache_read ``` ### What invalidates which level (verified 2026-08-30) Changes cascade **downward**: the level that changes, and every level after it, is invalidated. ✓ = that level's cache survives, ✘ = it is invalidated. | What changes | Tools | System | Messages | Note | |---|:--:|:--:|:--:|---| | Tool definitions (names, descriptions, params) | ✘ | ✘ | ✘ | Tools render at position 0 — a full rebuild | | Model switch | ✘ | ✘ | ✘ | Caches are model-scoped | | Web search toggle | ✓ | ✘ | ✘ | Modifies the system prompt | | Citations toggle | ✓ | ✘ | ✘ | Modifies the system prompt | | `speed: "fast"` ↔ standard | ✓ | ✘ | ✘ | Invalidates system + messages | | `tool_choice` | ✓ | ✓ | ✘ | Affects message blocks only | | Images added/removed anywhere | ✓ | ✓ | ✘ | Affects message blocks only | | Thinking config (mode, `budget_tokens`) | model-specific | model-specific | ✘ | Always invalidates messages; tools/system too on models that render it ahead of them | | `output_config.effort` | model-specific | model-specific | ✘ | Same shape as thinking. Setting effort explicitly to the model's default is equivalent to omitting it and does **not** invalidate | | Non-tool results w/ extended thinking | ✓ | ✓ | model-specific | Opus 4.5+ / Sonnet 4.6+ preserve thinking blocks (✓); earlier Opus/Sonnet and all Haiku strip them, dropping following messages from cache | The trap: `tool_choice`, images, thinking and effort are all commonly toggled per-request and all invalidate the **messages** cache every time. In a long agent loop, where the messages cache holds most of the tokens, "just flipping `tool_choice`" is not free — keep it constant across a conversation. Source: `https://platform.claude.com/docs/en/build-with-claude/prompt-caching.md` ## Batches API `POST /v1/messages/batches` — asynchronous Messages requests at a **flat 50% discount on all token usage** (stacks with prompt caching). | Fact | Value | |---|---| | Max batch size | 100,000 requests or 256 MB | | Turnaround | Usually <1 hour; max 24h | | Results retention | 29 days | | Feature support | All Messages features (tools, vision, caching, structured outputs, thinking) — no streaming | | Rate limits | Separate pool from interactive traffic | ```python import anthropic, time from anthropic.types.message_create_params import MessageCreateParamsNonStreaming from anthropic.types.messages.batch_create_params import Request client = anthropic.Anthropic() # 1. Create batch = client.messages.batches.create(requests=[ Request(custom_id=f"item-{i}", params=MessageCreateParamsNonStreaming( model="claude-haiku-4-5", max_tokens=64, messages=[{"role": "user", "content": f"Classify sentiment (one word): {text}"}])) for i, text in enumerate(texts) ]) # 2. Poll while True: batch = client.messages.batches.retrieve(batch.id) if batch.processing_status == "ended": break time.sleep(60) print(batch.request_counts) # succeeded / errored / canceled / expired # 3. Results (order not guaranteed — key on custom_id) for result in client.messages.batches.results(batch.id): if result.result.type == "succeeded": msg = result.result.message text = next((b.text for b in msg.content if b.type == "text"), "") elif result.result.type == "errored": # error.type == "invalid_request" → fix and resubmit; otherwise safe to retry ... ``` ```typescript const batch = await client.messages.batches.create({ requests: [{ custom_id: "request-1", params: { model: "claude-sonnet-5", max_tokens: 1024, messages: [{ role: "user", content: "Summarize..." }] }, }], }); // poll batches.retrieve(batch.id) until processing_status === "ended" for await (const result of await client.messages.batches.results(batch.id)) { if (result.result.type === "succeeded") { /* result.result.message */ } } ``` Batch gotchas: - `custom_id` is your only join key — results stream in completion order, not submission order. - Result types: `succeeded` | `errored` | `canceled` | `expired`. Resubmit `expired`; inspect `errored` (validation vs server error). - Cancel is async: `batches.cancel(id)` → status `"canceling"`; some requests may still complete. - Caching inside batches works, but hit rates are best-effort (requests run concurrently) — put a shared cached `system` block on every request and use the 1-hour TTL. - Per-request params are full `MessageCreateParams` minus `stream`. Use batches for: eval suites, backfills, bulk extraction/classification, nightly report generation, regenerating embeddings-adjacent metadata — any workload where minutes-to-hours latency is fine. ## Model Tiering Economics Worked example, 10M input + 1M output tokens/day: | Strategy | Cost/day | |---|---| | Everything Opus 5 | 10×$5 + 1×$25 = **$75** | | Everything Sonnet 5 | 10×$2 + 1×$10 = **$30** | | Route: 80% Haiku, 15% Sonnet, 5% Opus | input 8×$1 + 1.5×$2 + 0.5×$5 = $13.5; output 0.8×$5 + 0.15×$10 + 0.05×$25 = $6.75 = **~$20** | | Same + cached system prompts (70% of input cached at ~0.1x) | input → ~$5 = **~$12** | | Same + batchable share moved to Batches | **lower still (50% off that share)** | Patterns: - **Router**: a Haiku call classifies difficulty, dispatches to the right model. - **Cascade**: try Haiku; escalate to Sonnet/Opus only when confidence is low or validation fails (works well with structured outputs as the validator). - **Subagents**: keep the orchestrator on Opus, push parallel/simple legs to Haiku/Sonnet. Spawn a separate request per subagent (don't switch models mid-conversation — it kills the cache). - **Effort tuning**: on supported models, dropping `output_config.effort` from `high` to `medium` often cuts output tokens substantially at minor quality cost — cheaper than switching model tier for borderline workloads. ## Token Counting for Cost Estimation ```python count = client.messages.count_tokens( model="claude-opus-5", system=system, tools=tools, messages=messages) est_input_cost = count.input_tokens * 5.00 / 1_000_000 # Opus 5 input rate ``` - Free endpoint; counts include tools and system. - Tool use adds a hidden system prompt (~290-800 tokens depending on model and `tool_choice`). - Token counts are model-specific — re-baseline when migrating models; don't apply blanket multipliers. -
compaction.md 13.7 KB
# Compaction & Context Editing Reference > **What this file owns:** the decision to *discard or rewrite* context — when it is > justified, what each option costs, and the first-party APIs that do it > (`context_management` edits, the memory tool). > > **Adjacent files:** [context-engineering.md](context-engineering.md) owns the budget > and the three tiers; [caching-and-cost.md](caching-and-cost.md) owns cache mechanics. > > Facts verified against platform.claude.com and anthropic.com **2026-08-30**. --- ## 1. The headline: compaction is not the default **Compaction** = summarising the message history and reinitiating a context window from the summary. The instinct is that long conversations must be summarised or they get expensive and the model gets confused. **Under modern prompt caching that instinct is measurably wrong more often than it is right.** A 2026 evaluation of a production AI tutor (Bouchard, Solano & Vaid, Towards AI — 660 turns across 11 configurations) compared keeping the full history against a range of compaction strategies (their production preset of clearing + summarisation, context reset, prompt compression, selective retention): | Strategy | Fact recall at turn 11 | Cost / turn | Time to first token | |---|---:|---:|---:| | **Full history (keep everything)** | **92–100%** | **$0.11** | **17 s** | | Production preset (clear + summarise) | 38–58% | $0.24 | 21 s | | Context reset | 17% | — | — | Keep-everything won **cost, latency and recall simultaneously** — the three axes compaction is usually reached for. Summarisation was more than twice the cost per turn, slower to first token, and forgot roughly half of a fact planted ten turns earlier. **Read the caveats, they matter:** - The study ran on **Gemini 3.5 Flash via LangChain**, not Claude. What transfers is the *mechanism* — a cached prefix is billed at a fraction of base input, while summarising rewrites that prefix and forfeits the discount. Claude's cache economics make the mechanism *stronger*, not weaker (§3). The specific percentages are theirs, not a Claude benchmark. - "Keep everything" is not "ignore context entirely." The same study found that **capping tool outputs at a fixed size cut cost per turn 38% with no recall loss** — because a cap shrinks context *without rewriting the cached prefix*. Shaping what enters context is cheap; rewriting what is already there is not. **The rule:** compaction is a deliberate response to a **named constraint**, not a reflex. If you cannot name which of the three constraints in §2 you are hitting, do not compact — cap tool output and append instead. --- ## 2. The three constraints that justify compaction Name one before you compact. Each has a different remedy, and reaching for the wrong one is how teams pay summarisation costs to solve a problem summarisation does not fix. ### Constraint A — Context ceiling *The conversation will not fit.* The hard one; it is not negotiable by budget. - **Diagnose:** projected tokens at turn N exceed the model's window (1M on the current flagships, 200K on Haiku 4.5 — see the model table in `SKILL.md`). - **Cheapest remedies first:** cap tool outputs → move payloads to files (tier 2) → server-side clearing (§3) → summarising compaction (§5). - **Note:** a 1M window makes this constraint *rare* for conversational work and still common for long agentic runs, where tool results dominate. ### Constraint B — Cost ceiling *It fits, but the bill is unacceptable.* This is the constraint most often assumed and least often real, because the comparison is usually made against **uncached** costs. - **Diagnose honestly:** compare `cache_read_input_tokens × 0.1 × base_rate` against the cost of a summarisation call **plus** the cache write it forces on the next request. Do the arithmetic before assuming (§3). - **Cheapest remedies first:** verify the cache is actually hitting (a zero `cache_read_input_tokens` is a bug, not a reason to compact) → tier the model down → cap tool outputs → 1-hour TTL if the gaps between turns are the problem. ### Constraint C — Latency target *Time-to-first-token is too slow at depth.* Real, and the weakest case for summarisation. - **Diagnose:** measure TTFT against context length; confirm the growth is in the prefix rather than in thinking or tool round-trips. - **Caution:** the tutor study measured compaction as *slower* (21 s vs 17 s) — the summarisation call is itself a round-trip, and the next request re-writes the cache. Compaction buys latency only when it is amortised over many subsequent turns. - **Cheapest remedies first:** cap tool outputs → tier down → reduce `effort` → stream (TTFT is a streaming problem before it is a context problem). ### What each option costs you | Option | Recall | Cache | Latency | Reversible? | |---|---|---|---|---| | **Append (do nothing)** | Full | Preserved | Grows with length | n/a | | **Cap tool output at the boundary** | Loses untruncated detail only | **Preserved** | Improves | No, but loss is bounded and predictable | | **Write payload to file, keep the path** | Full (one round-trip away) | Preserved | One extra hop when read | Yes | | **Server-side clearing** (`context_management`) | Cleared results unrecoverable to the model unless saved to memory | **Invalidated at the clear point** | Improves after the re-write | No | | **Summarising compaction** | Lossy, and the loss is unpredictable | **Invalidated wholesale** | One extra call now, faster later | No | Read that table top-down: the first three preserve the cache, and the two that break it are the two people reach for first. --- ## 3. The break-even arithmetic Claude's cache read is **0.1× the base input rate** on every model; writes are **1.25×** (5-minute TTL) or **2×** (1-hour TTL). So the per-turn cost of carrying an N-token history is: ``` append : N × 0.1 × base_rate (steady state, cache hit) compact : summarisation call (input ≈ N, output ≈ S) + S × 1.25 × base_rate (cache write of the new prefix) + S × 0.1 × base_rate per later turn (steady state on the summary) + everything before the rewrite, forfeited ``` Compaction only wins once `0.1 × N` exceeds `0.1 × S` by enough to repay the summarisation call **and** the forfeited prefix — i.e. when the history is very large, the summary is very small, and the session continues for many turns afterwards. A compaction just before the conversation ends is pure loss. **The published threshold, adapted to Claude.** The tutor study put the crossover at a *cached input price above ~$0.55 per million tokens* — above that, summarisation starts to compete. Because a Claude cache read is 0.1× base input, that translates to: ``` cache-read rate = 0.1 × base input rate crossover ≈ $0.55 / MTok cached → base input rate ≳ $5.50 / MTok ``` Derive it from the current price table in `SKILL.md` rather than memorising a model list: a model whose **base input rate is under ~$5.50/MTok has cache reads cheap enough that appending stays ahead**, and only the premium tier approaches the crossover. Re-run this arithmetic when prices change — it is two multiplications. This is why the honest first question is never "should we compact?" but **"is the cache hitting?"** A broken cache makes appending look 10× worse than it is and makes compaction look like the fix, when the actual fix is a stray timestamp in the system prompt. --- ## 4. First-party option: server-side context editing `context_management` clears content **server-side, before the prompt reaches the model**. Your client keeps the full unmodified history — there is no client state to synchronise. Beta header: `context-management-2025-06-27`. ### Strategy: clear tool uses ```python response = client.beta.messages.create( model="claude-opus-5", max_tokens=4096, messages=messages, tools=tools, betas=["context-management-2025-06-27"], context_management={"edits": [{ "type": "clear_tool_uses_20250919", "trigger": {"type": "input_tokens", "value": 30000}, # when to fire "keep": {"type": "tool_uses", "value": 3}, # recent pairs kept "clear_at_least":{"type": "input_tokens", "value": 5000}, # min cleared per fire "exclude_tools": ["web_search"], # never clear these }]}, ) ``` | Parameter | Default | Meaning | |---|---|---| | `trigger` | 100,000 input tokens | Threshold that activates clearing (`input_tokens` or `tool_uses`) | | `keep` | 3 tool uses | Most recent tool-use/result pairs preserved | | `clear_at_least` | none | Minimum tokens cleared per activation — **the cache-economics knob** | | `exclude_tools` | none | Tools whose results are never cleared | | `clear_tool_inputs` | `false` | Also clear the tool *call parameters*, not just results | Oldest results go first; cleared content is replaced with a placeholder. ### Strategy: clear thinking blocks ```python context_management={"edits": [ {"type": "clear_thinking_20251015", "keep": {"type": "thinking_turns", "value": 2}}, {"type": "clear_tool_uses_20250919", "trigger": {"type": "input_tokens", "value": 50000}}, ]} ``` `keep` takes `{"type": "thinking_turns", "value": N}` or `"all"` (maximises cache hits). When combining strategies, **thinking-block clearing must be listed first**. Default behaviour differs by model generation: Opus 4.5+ and Sonnet 4.6+ keep all prior thinking; earlier Opus/Sonnet and all Haiku models keep only the last turn. ### The cache interaction — the part that decides whether this pays - **Clearing invalidates the cached prefix from the clear point onward.** Each activation costs a cache write on the next request. - Therefore **set `clear_at_least`**. Without it, a trigger can fire and clear a trivial amount, paying a full cache re-write to save a few hundred tokens. It exists precisely to make each cache break worth taking. - Thinking-block clearing is the mirror image: **keeping** thinking preserves the cache; **clearing** it invalidates at the clearing point. `"all"` is the cache-optimal setting. ### Inspecting and previewing Responses report what was applied: ```json {"context_management": {"applied_edits": [ {"type": "clear_tool_uses_20250919", "cleared_tool_uses": 8, "cleared_input_tokens": 50000} ]}} ``` Preview before committing to a configuration — `count_tokens` accepts the same `context_management` block and returns both the post-clearing `input_tokens` and `context_management.original_input_tokens`. ### Reported results Anthropic's internal agentic-search evaluation: **context editing alone +29% over baseline; context editing with the memory tool +39%.** In a 100-turn web-search evaluation, context editing let agents complete workflows that would otherwise fail on context exhaustion, **reducing token consumption by 84%**. Note the shape of that claim — the gains are on *long-horizon agentic search*, where Constraint A is genuinely binding. It is not evidence that clearing helps a twelve-turn conversation. --- ## 5. First-party option: the memory tool `memory_20250818` gives Claude a directory of files it can create, read, update and delete, persisting **across** conversations. ```python tools=[{"type": "memory_20250818", "name": "memory"}], context_management={"edits": [{"type": "clear_tool_uses_20250919"}]} ``` The pairing is the point: Claude is warned as context approaches a clearing threshold, and can **write the durable conclusion to memory before the raw material is cleared**. Clearing without memory throws information away; clearing with memory demotes it from tier 1 to tier 2 (→ [context-engineering.md](context-engineering.md) §2). --- ## 6. If you must compact: do it well When a named constraint genuinely demands summarisation: - **Maximise recall first, then tune precision.** Anthropic's guidance is to start with a compaction prompt that captures every relevant piece of information, then iterate to trim. A compaction prompt tuned for brevity first will silently drop the one detail the next 40 turns needed. - **Compact at a natural boundary** — a finished sub-task, a landed change — not at an arbitrary token count mid-reasoning. - **Keep the last N turns verbatim** alongside the summary. Recent turns are where reference resolution ("that file", "the second one") lives. - **Write the durable facts to a file first** (memory tool or `NOTES.md`), so the summary is a convenience rather than the sole record. - **Compact once, deep** rather than repeatedly and shallowly. Every pass is a cache write and a lossy re-encoding; summaries of summaries degrade fast. - **Measure it.** Plant a fact early, probe for it later, and compare against the keep-everything baseline on *your* workload. The tutor study's headline result is that the baseline is much stronger than teams assume — including, quite possibly, yours. --- ## 7. Sources - Anthropic — *Effective context engineering for AI agents*: `https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents` - Anthropic — *Managing context on the Claude Developer Platform* (29% / 39% / 84%): `https://claude.com/blog/context-management` - Context editing API: `https://platform.claude.com/docs/en/build-with-claude/context-editing` - Memory tool: `https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool` - Prompt caching (multipliers, prefix rules): `https://platform.claude.com/docs/en/build-with-claude/prompt-caching.md` - Bouchard, Solano & Vaid (Towards AI), *Context Engineering in 2026: Why We Stopped Compacting Our Agent's Context*, AI Engineer World's Fair, Aug 2026: `https://www.louisbouchard.ai/context-engineering-2026/` -
context-engineering.md 15.4 KB
# Context Engineering Reference > **What this file owns:** the discipline of deciding *what the model sees on every > call* — the context budget, the three placement tiers, and the agentic-loop > specifics (tool-result bloat, sub-agents as isolation, note-taking). > > **Adjacent files, so this one doesn't duplicate them:** > [compaction.md](compaction.md) owns the *decision to discard* context and the > `context_management` / memory-tool APIs. > [caching-and-cost.md](caching-and-cost.md) owns prompt-cache *mechanics* > (breakpoints, TTLs, minimum prefixes, invalidation table). > > Facts verified against platform.claude.com and anthropic.com **2026-08-30**. --- ## 1. The frame: context is a budget, not a container Prompt engineering asks "what do I write in the prompt?". **Context engineering asks "what earns a place in the window on *this* call?"** — including everything that lands there without you typing it: tool definitions, tool results, retrieved documents, prior turns, system reminders, thinking blocks. Anthropic's framing (*Effective context engineering for AI agents*): the goal is "the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome." Prompt engineering is discrete — you write it once. Context engineering is **iterative**: it happens on every single inference. ### Why the budget is real - **Attention is finite.** Transformers form n² pairwise relationships for n tokens. Accuracy degrades as token count grows — the failure mode Anthropic names **context rot**. A 1M-token window is a *capacity*, not a *target*. - **Training-distribution thinness.** Models have seen far fewer long-range, context-wide dependencies than short ones, so long-context reasoning is the weakest part of the envelope, not merely the slowest. - **Every token is billed on every turn.** In a multi-turn loop the same prefix is re-sent each call. Uncached, a 50K-token prefix over 40 turns is 2M input tokens. (Cached, it is ~10% of that — see §4, and it is exactly why the compaction instinct is often wrong. → [compaction.md](compaction.md).) **Working rule.** Before adding anything to a prompt, ask: *does this change the model's next action?* If not, it belongs on disk or behind a retrieval call, not in the window. --- ## 2. The three tiers Every piece of information the agent could use lives in exactly one of three places. Choosing the tier deliberately is most of context engineering. | Tier | Where it lives | Latency to use | Cost profile | Use when | |---|---|---|---|---| | **1 — In context** | Rendered into `tools` / `system` / `messages` every call | Zero | Paid every turn (≈0.1× when cached) | It steers *most* turns: the task spec, the invariants, the current working set | | **2 — On disk, read on demand** | A file the agent can `Read`/`Grep` when it decides to | One tool round-trip | Paid only when read | It steers *some* turns and the agent can tell when it needs it from the filename alone | | **3 — Retrieved** | Index / vector store / API behind a search tool | One+ round-trips, plus a relevance gamble | Paid only when hit | The corpus is too large to enumerate and relevance is query-dependent | ### Choosing between tiers - **Tier 1 is the expensive tier.** It costs on every call whether or not the turn needed it. Reserve it for what is load-bearing on the majority of turns. - **Tier 2 is the default for "might need it".** Anthropic's just-in-time framing: let the agent maintain lightweight identifiers (file paths, queries, links) and hydrate them at runtime. A path is ~10 tokens; the file it names may be 10,000. The trade-off is honest — "runtime exploration is slower than retrieving pre-computed data." - **Tier 3 pays for scale with a relevance risk.** If retrieval misses, the model does not know what it did not see. Prefer tier 2 whenever the candidate set is small enough to name. - **Hybrids win in practice.** Pre-load the handful of things that steer every turn (tier 1), leave the long tail addressable (tier 2/3). ### The demotion test When a prompt is too big, demote rather than delete. For each block ask: 1. Does it change the next action on **most** turns? → stays tier 1. 2. Can the agent tell from a **name** that it needs this? → tier 2 (write it to a file, leave the path in context). 3. Neither, but it is occasionally essential? → tier 3 (index it, ship a search tool). Deleting outright is the fourth option and the only irreversible one. --- ## 3. Progressive disclosure — this repo already runs on it **Claude Code skills are a tier-1/tier-2 split you can read on disk.** That is not an analogy; it is the same mechanism: | Layer | Tier | Loaded | |---|---|---| | Skill `name` + `description` frontmatter | 1 | Always — every skill's description sits in the session prompt | | `SKILL.md` body | 2 | When the router matches the description | | `references/*.md` (this file) | 2 | Only when `SKILL.md` cites it and the task needs it | | `scripts/*` output | 2 | Only when executed | This is why the repo's own rules read the way they do, and the rules are context engineering rules wearing authoring clothes: - **"The `description` is the trigger"** — the description is the only always-resident text, so it must carry the routing signal and nothing else. - **"Keep the body under 500 lines"** — the body is what gets pulled in wholesale on a match; an oversized body spends the budget of every task that touches the skill. - **"One concept per reference file"** — a file is the unit of loading. Two concepts in one file means loading both to get either. - **"Every reference must be cited from `SKILL.md`"** — an uncited file is unreachable. In tier terms: a tier-2 artefact with no pointer in tier 1 does not exist. The same shape generalises to any agent you build: a small always-on spec, a set of named-and-addressable documents, and a retrieval path for the long tail. See [SKILL-CREATION-PROTOCOL.md](../../../docs/SKILL-CREATION-PROTOCOL.md) Step 3 and [SKILL-RESOURCE-PROTOCOL.md](../../../docs/SKILL-RESOURCE-PROTOCOL.md) §1. --- ## 4. Cache-aware prompt architecture The cache is what makes a large tier-1 affordable — and the cache is a **prefix match**, so *ordering is architecture*. ### The layout rule Requests render in a fixed order — `tools` → `system` → `messages` — and each level builds on the previous. Therefore: ``` tools ─┐ system ├─ static, byte-identical across calls ← cache_control breakpoint here (stable docs) ┘ messages ← volatile: the turn, the retrieved chunk, the timestamp ``` **Static prefix first, volatile content last.** Anything that varies per request must sit *after* the last breakpoint, or it changes the prefix bytes and every cached token behind it is re-billed at write price. ### Why reordering silently destroys the cache There is no error. Moving a block, adding a conditional system section, or interpolating a user id early in the prompt produces a *different prefix hash*, which is simply a miss: `cache_read_input_tokens: 0`, `cache_creation_input_tokens: <all of it>`, a 1.25–2× bill instead of 0.1×, and no diagnostic anywhere. The only signal is the usage block — which is why "assert `cache_read_input_tokens > 0` in staging" is a real test, not a nicety. Two ordering traps specific to agent loops: - **A breakpoint searches backward at most 20 content blocks.** A turn that appends more than 20 blocks (many `tool_use`/`tool_result` pairs) jumps the window and silently misses. Add an intermediate breakpoint roughly every 15 blocks in long turns. - **A cache entry only becomes readable once the first response begins streaming.** Fanning out N parallel requests against a cold shared prefix writes N entries and reads none. Fire one, await first token, then fire the rest. The invalidation table, per-model minimum prefixes, TTL pricing and the full silent-invalidator checklist live in [caching-and-cost.md](caching-and-cost.md) — this section is the *shape* rule only. ### The consequence for compaction Rewriting history to make it shorter changes the prefix. A summarisation pass therefore pays a full cache write on the next call *and* discards every cached token before it. That is the mechanism behind the counter-intuitive finding in [compaction.md](compaction.md): under caching, the cheap thing is usually to **append**, not to **rewrite**. --- ## 5. Multi-turn and agentic specifics ### 5.1 Tool results are where the budget actually goes In an agent loop, the growth term is almost never the system prompt — it is tool output. A file read, a search result, an HTTP response: each lands verbatim and stays for the rest of the session. **Design tools to return decisions, not dumps.** Anthropic's guidance: tools should be "self-contained, robust to error, and extremely clear with respect to their intended use", with "minimal overlap in functionality". A bloated tool set costs tokens at position 0 *and* creates ambiguous decision points. Concrete levers, cheapest first: | Lever | What it does | Cost | |---|---|---| | **Cap tool output at a fixed size** | Truncate/paginate at the tool boundary, before the result enters context | Free — and critically, it **shrinks context without rewriting the prefix**, so the cache survives | | **Return a handle, not the payload** | Tool writes to a file, returns the path + a 200-token précis | One extra round-trip if the agent needs the full text | | **Filter at the source** | `grep`-shaped tools instead of `cat`-shaped ones | Design-time only | | **Clear stale results** | `context_management` server-side clearing | Invalidates the cache at the clear point → [compaction.md](compaction.md) | The capping lever is the one to reach for first: a Towards AI evaluation (AI Engineer World's Fair, August 2026) measured **38% lower cost per turn** from fixed-size tool-output caps alone, with no loss of recall — precisely because it does not touch the cached prefix. ### 5.2 Summarise a tool result, or write it to a file? | Situation | Do this | |---|---| | Result is large and needed **later, in full** (a fetched spec, a big file) | **Write to a file**, return the path. Lossless, addressable, and the path costs ~10 tokens per turn | | Result is large and only its **conclusion** matters (a 300-row query → "4 rows failed") | **Summarise at the tool boundary** — before it enters context, not after | | Result is large, and you cannot tell which parts matter yet | **Both**: file for fidelity, précis in context, path in the précis | | Result is small, or is the thing the user asked for | **Leave it alone.** Compression has a floor; do not spend a round-trip to save 200 tokens | The distinction that matters: **summarising at the tool boundary is free of cache cost** (the shorter result is what gets appended, and nothing before it moves). Summarising *after the fact* — rewriting history that is already in the prefix — is compaction, and pays the full cache-write penalty. ### 5.3 Structured note-taking (persistent memory) Have the agent maintain an external `NOTES.md` / progress file and pull it back in when needed — "persistent memory with minimal overhead". It is a tier-2 artefact that the agent itself authors, and it survives context resets, which is what makes it the natural companion to any clearing strategy: write the durable conclusion out *before* the raw material is cleared. The first-party version of this is the memory tool → [compaction.md](compaction.md) §4. ### 5.4 Sub-agents as context isolation A sub-agent is not primarily a parallelism device — **it is a second context window whose contents never touch yours.** ``` orchestrator context: task spec + plan + N × (≈1-2K token summary) ↓ spawn ↑ return sub-agent context: the 80K tokens of exploration nobody else needs ``` The exploration — dozens of file reads, failed greps, dead ends — is billed once, inside the sub-agent, and then discarded. The orchestrator sees only the distilled result; Anthropic's guidance puts that return payload at roughly **1,000–2,000 tokens**. Use it when: - A subtask generates far more intermediate context than conclusion (search, triage, audit, "find where X is implemented"). - You want a genuinely independent opinion — a fresh window cannot be primed by the orchestrator's earlier wrong turn. (Adversarial verification depends on this.) - The subtask's tool set is large and irrelevant to the main loop — it renders at position 0 in the sub-agent's prompt, not yours. Do **not** use it when the subtask needs most of the orchestrator's context to make sense: you will pay to reconstruct that context in the child, and lose fidelity in the hand-off. The hand-off is a lossy channel by design; if the summary has to carry everything, isolation was the wrong tool. Costs to price in: the sub-agent's prefix is a **cold cache** (a fresh window shares nothing with the parent), the hand-off is lossy, and errors are harder to attribute. Tier the model down for the isolated leg — an Opus orchestrator with Haiku/Sonnet sub-agents is the standard shape (see [caching-and-cost.md](caching-and-cost.md), "Model Tiering Economics"). --- ## 6. Instrumentation — what to measure You cannot engineer a budget you cannot see. Log per request: | Signal | Source | What it tells you | |---|---|---| | `usage.input_tokens` | response | Uncached remainder — the part after your last breakpoint | | `usage.cache_read_input_tokens` | response | Cache is working. **Zero across identical-prefix calls = a silent invalidator** | | `usage.cache_creation_input_tokens` | response | What you paid write price for this turn | | `usage.output_tokens` | response | The other half of the bill | | Pre-flight estimate | `client.messages.count_tokens(...)` | Free; counts tools + system. Never `tiktoken` (OpenAI's tokenizer, 15–20% undercount on Claude) | | Growth per turn | your own diff | Which tool is the growth term — almost always one of them dominates | The single most valuable alarm: **`cache_read_input_tokens == 0` on a request whose prefix should be unchanged.** It is the only symptom a broken cache produces. --- ## 7. Cross-references | Concern | Where | |---|---| | Compaction decision, `context_management`, memory tool | [compaction.md](compaction.md) | | Cache breakpoints, TTLs, minimums, invalidation table, batches | [caching-and-cost.md](caching-and-cost.md) | | Tool definitions, agentic loop, `tool_result` mechanics | [tool-use.md](tool-use.md) | | Sub-agents in the Agent SDK (`agents` option) | [agent-sdk.md](agent-sdk.md) | | Claude Code's own context surface (CLAUDE.md, skills, hooks) | `claude-code-ops` skill | | Scheduled/autonomous loops that re-send a prompt on a cadence | `loop-ops` skill | | Cross-provider fleets and adversarial verify | `fleetflow` skill | ## 8. Sources - Anthropic — *Effective context engineering for AI agents*: `https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents` - Anthropic — *Managing context on the Claude Developer Platform*: `https://claude.com/blog/context-management` - Prompt caching (ordering, 20-block lookback, concurrency): `https://platform.claude.com/docs/en/build-with-claude/prompt-caching.md` - Context editing: `https://platform.claude.com/docs/en/build-with-claude/context-editing` - Bouchard, Solano & Vaid (Towards AI), *Context Engineering in 2026*, AI Engineer World's Fair, Aug 2026: `https://www.louisbouchard.ai/context-engineering-2026/` -
messages-api.md 11.7 KB
# Messages API Reference `POST https://api.anthropic.com/v1/messages` — the single endpoint everything runs through. Tools, structured outputs, thinking, and caching are all features of this endpoint, not separate APIs. ## Required Headers | Header | Value | |---|---| | `x-api-key` | Your API key (`sk-ant-...`) | | `anthropic-version` | `2023-06-01` | | `content-type` | `application/json` | | `anthropic-beta` | Comma-separated beta IDs, only for beta features | OAuth bearer tokens go on `Authorization: Bearer <token>` instead of `x-api-key` (plus `anthropic-beta: oauth-2025-04-20`). Setting both `ANTHROPIC_API_KEY` and `ANTHROPIC_AUTH_TOKEN` makes the SDK send both headers and the API rejects the request. ## Request Parameters | Param | Type | Required | Notes | |---|---|---|---| | `model` | string | yes | Exact alias ID, e.g. `claude-opus-5` — no date suffixes | | `max_tokens` | int | yes | Hard output cap. Default sensibly: ~16000 non-streaming, ~64000 streaming, ~256 classification | | `messages` | array | yes | Alternating `user`/`assistant` turns; first must be `user`. Consecutive same-role messages are merged | | `system` | string \| block[] | no | System prompt. Block-list form required for `cache_control` | | `tools` | array | no | Custom + server tool definitions (see tool-use.md) | | `tool_choice` | object | no | `auto` (default) / `any` / `tool` / `none` | | `thinking` | object | no | Adaptive; **on by default** on Fable 5 / Opus 5 / Sonnet 5. `{"type": "enabled", "budget_tokens": N}` 400s on Opus 4.7+ | | `output_config` | object | no | `{"effort": "...", "format": {...}, "task_budget": {...}}` | | `stop_sequences` | string[] | no | Custom stop strings | | `stream` | bool | no | SSE streaming | | `metadata` | object | no | `{"user_id": "..."}` — opaque end-user id for abuse detection | | `temperature` / `top_p` / `top_k` | number | no | **Removed on Opus 4.7 and later (400)** — Opus 5, Sonnet 5, Fable 5 included. On earlier 4.x: at most one of temperature/top_p | | `cache_control` | object | no | Top-level auto-caching: caches the last cacheable block | | `container` | string | no | Reuse a code-execution container id | | `mcp_servers` | array | no | Remote MCP connector (beta `mcp-client-2025-11-20`) | ### Message content blocks `content` is either a plain string or an array of blocks: ```json {"role": "user", "content": [ {"type": "text", "text": "What's in this image?"}, {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": "<b64>"}}, {"type": "image", "source": {"type": "url", "url": "https://example.com/img.png"}}, {"type": "document", "source": {"type": "base64", "media_type": "application/pdf", "data": "<b64>"}}, {"type": "tool_result", "tool_use_id": "toolu_...", "content": "..."} ]} ``` ## Response Shape ```json { "id": "msg_01...", "type": "message", "role": "assistant", "model": "claude-opus-5", "content": [ {"type": "thinking", "thinking": "...", "signature": "..."}, {"type": "text", "text": "Hello!"}, {"type": "tool_use", "id": "toolu_01...", "name": "get_weather", "input": {"location": "Paris"}} ], "stop_reason": "end_turn", "stop_sequence": null, "usage": { "input_tokens": 1024, "output_tokens": 256, "cache_creation_input_tokens": 0, "cache_read_input_tokens": 0 } } ``` `content` is a **list of typed blocks** — never index `content[0].text` blindly (a `thinking` block may come first). Filter by `.type`. ## Stop Reasons | `stop_reason` | Meaning | What to do | |---|---|---| | `end_turn` | Finished naturally | Done | | `max_tokens` | Hit the `max_tokens` cap | Raise the cap or stream; output may be truncated mid-thought | | `stop_sequence` | Hit a custom stop string | `stop_sequence` field has which one | | `tool_use` | Claude wants tool(s) executed | Execute each `tool_use` block, send `tool_result`(s), re-request | | `pause_turn` | Server-side tool loop hit its iteration limit | Append the assistant turn and re-send unchanged — server resumes; do NOT add a "continue" user message | | `refusal` | Safety refusal | Check `stop_details` (`category`: "cyber"/"bio"/"reasoning_extraction" (Fable 5)/null, `explanation`); don't retry same prompt | | `model_context_window_exceeded` | Context window exhausted (distinct from max_tokens) | Compact, truncate, or split the conversation | ```python if response.stop_reason == "refusal" and response.stop_details: print(response.stop_details.category, response.stop_details.explanation) ``` ## Multi-Turn Conversations The API is stateless — send the full history every request: ```python messages = [] def chat(user_msg: str) -> str: messages.append({"role": "user", "content": user_msg}) r = client.messages.create(model="claude-opus-5", max_tokens=16000, messages=messages) # Append the FULL content list (preserves tool_use/thinking/compaction blocks) messages.append({"role": "assistant", "content": r.content}) return next(b.text for b in r.content if b.type == "text") ``` For conversations that may exceed context: server-side **compaction** (beta header `compact-2026-01-12`, `context_management: {"edits": [{"type": "compact_20260112"}]}` on `client.beta.messages.create`). Critical: append `response.content` back verbatim — compaction blocks must be preserved or state is silently lost. ## Streaming ```python with client.messages.stream(model="claude-opus-5", max_tokens=64000, messages=[...]) as stream: for text in stream.text_stream: print(text, end="", flush=True) final = stream.get_final_message() print(final.usage.output_tokens) ``` ```typescript const stream = client.messages.stream({ model: "claude-opus-5", max_tokens: 64000, messages }); stream.on("text", (delta) => process.stdout.write(delta)); const final = await stream.finalMessage(); // never wrap .on() in new Promise() ``` ### SSE event sequence ``` message_start → message metadata (id, model, usage so far) content_block_start → index + block type (text / thinking / tool_use) content_block_delta → text_delta | thinking_delta | input_json_delta content_block_stop → block finished message_delta → stop_reason + final usage message_stop → stream done ``` Tool inputs stream as `input_json_delta` (partial JSON strings) — accumulate and parse at `content_block_stop`, or just use `get_final_message()` / `finalMessage()` which assembles parsed blocks for you. **Why stream:** non-streaming requests with large `max_tokens` exceed HTTP timeouts (the Python SDK raises `ValueError` for non-streaming requests it estimates will run >~10 min). Default to streaming for anything long. ## Error Handling | HTTP | `error.type` | Retryable | Typical cause | |---|---|---|---| | 400 | `invalid_request_error` | no | Bad params: removed sampling params, `budget_tokens` on 4.7+, prefill on 4.7+, modified thinking blocks, role ordering | | 401 | `authentication_error` | no | Missing/invalid key; both key + token set | | 403 | `permission_error` | no | Key lacks model/feature access | | 404 | `not_found_error` | no | Bad model ID (date-suffix mistake) or endpoint | | 413 | `request_too_large` | no | Body over size limit — shrink images/history | | 429 | `rate_limit_error` | yes | RPM/ITPM/OTPM exceeded — honor `retry-after` | | 500 | `api_error` | yes | Transient server issue | | 529 | `overloaded_error` | yes | Capacity — backoff; consider another model | Error envelope: ```json {"type": "error", "error": {"type": "rate_limit_error", "message": "..."}, "request_id": "req_011CSH..."} ``` Log `request_id` (also `response._request_id` on SDK success objects) when reporting issues to Anthropic. ### Typed exceptions — never string-match messages ```python import anthropic try: r = client.messages.create(...) except anthropic.BadRequestError as e: # 400 raise # don't retry client errors except anthropic.RateLimitError as e: # 429 wait = int(e.response.headers.get("retry-after", "60")) except anthropic.APIStatusError as e: # catch-all with .status_code / .type if e.status_code >= 500: ... # retryable except anthropic.APIConnectionError: ... # network — retryable ``` ```typescript try { await client.messages.create({...}); } catch (err) { if (err instanceof Anthropic.RateLimitError) { /* backoff */ } else if (err instanceof Anthropic.APIError) { console.error(err.status, err.message); } } ``` All subclasses expose `.type` (e.g. `"overloaded_error"`) for finer classification than the status code (e.g. `billing_error` vs `permission_error`, both 403). ### Retries The official SDKs **auto-retry** connection errors, 408/409/429 and >=500 with exponential backoff — default `max_retries=2`. Configure per client (`anthropic.Anthropic(max_retries=5)`) or per call (`client.with_options(max_retries=5, timeout=20.0).messages.create(...)`). Only hand-roll retry logic when you need behavior beyond that (e.g. queue + jitter across many workers): ```python import random, time def call_with_retry(client, max_retries=5, base=1.0, cap=60.0, **kwargs): last = None for attempt in range(max_retries): try: return client.messages.create(**kwargs) except anthropic.RateLimitError as e: last = e except anthropic.APIStatusError as e: if e.status_code < 500: raise # 4xx (except 429) is not retryable last = e time.sleep(min(base * 2 ** attempt + random.random(), cap)) raise last ``` Default request timeout is 10 minutes (`timeout=` on the client or `with_options`). On timeout: `anthropic.APITimeoutError`, retried per `max_retries`. ## Rate Limits Limits are per-organization, per-model-class, measured three ways: - **RPM** — requests per minute - **ITPM** — input tokens per minute (cache reads often discounted/exempt — check headers) - **OTPM** — output tokens per minute Tiers scale with cumulative spend (Tier 1-4, then custom/scale). Check live limits in Console or response headers: | Header | Meaning | |---|---| | `retry-after` | Seconds to wait (on 429) | | `anthropic-ratelimit-requests-limit` / `-remaining` / `-reset` | RPM state | | `anthropic-ratelimit-input-tokens-*` / `-output-tokens-*` | ITPM / OTPM state | Practical guidance: - Treat 429 as backpressure: honor `retry-after`, add jitter, cap concurrency. - Long-running agent fleets: budget OTPM, not just RPM — output is usually the binding constraint. - Batches API has separate, much higher throughput and doesn't draw from interactive rate limits — move bulk traffic there. - 529 `overloaded_error` is capacity, not your quota — backoff and/or fail over to a different model tier. ## Token Counting `POST /v1/messages/count_tokens` — free, model-specific, counts a request without running it: ```python n = client.messages.count_tokens( model="claude-opus-5", system=system, tools=tools, messages=[{"role": "user", "content": text}], ).input_tokens ``` Never estimate with `tiktoken` (OpenAI tokenizer; 15-20% undercount on prose, worse on code). Token counts differ **between Claude models** too — count against the model you'll run. ## Vision & Documents - Images: `{"type": "image", "source": {...}}` blocks — base64, URL, or Files API `{"type": "file", "file_id": ...}`. Opus 4.7+ supports high-res input (up to 2576px long edge, pixel-accurate coordinates; up to ~3x image tokens). - PDFs: `{"type": "document", "source": {...}}` — base64, URL, plain text, or file_id. Optional `citations: {"enabled": true}`. - Files API (beta `files-api-2025-04-14`): upload once (`client.beta.files.upload(...)`), reference by `file_id` across requests. 500 MB/file, 100 GB/org. -
structured-outputs.md 7.4 KB
# Structured Outputs Reference Two related features, same constrained-sampling mechanism: | Feature | Parameter | Constrains | |---|---|---| | **JSON outputs** | `output_config: {"format": {...}}` | Claude's response text (guaranteed valid JSON matching your schema) | | **Strict tool use** | `strict: true` on a tool definition | The `input` of tool calls | They can be combined in one request. Supported on Fable 5, Opus 5, Sonnet 5, and the legacy Opus 4.8/4.7/4.6/4.5 and Sonnet 4.6/4.5 line, plus Haiku 4.5 (the compatibility list names its dated snapshot `claude-haiku-4-5-20251001`; the docs' own examples use plain aliases throughout, so keep using `claude-haiku-4-5`). Where a model lacks support, enforce output shape via system-prompt instructions or strict tool use instead. **Naming:** the canonical parameter is `output_config.format`. The older top-level `output_format` parameter (and the `structured-outputs-2025-11-13` beta header) is **deprecated** — still accepted during a transition window, and still used as a convenience kwarg by some SDK `parse()` methods, but write new code against `output_config`. ## JSON Outputs — raw schema ```python import json, anthropic client = anthropic.Anthropic() response = client.messages.create( model="claude-opus-5", max_tokens=16000, messages=[{"role": "user", "content": "Extract: John Smith (john@example.com) wants the Enterprise plan."}], output_config={ "format": { "type": "json_schema", "schema": { "type": "object", "properties": { "name": {"type": "string"}, "email": {"type": "string", "format": "email"}, "plan": {"type": "string", "enum": ["Free", "Pro", "Enterprise"]}, }, "required": ["name", "email", "plan"], "additionalProperties": False, }, } }, ) text = next(b.text for b in response.content if b.type == "text") data = json.loads(text) # guaranteed valid against the schema (unless refusal/max_tokens) ``` cURL shape: ```json { "model": "claude-opus-5", "max_tokens": 1024, "output_config": { "format": {"type": "json_schema", "schema": { ... }} }, "messages": [{"role": "user", "content": "..."}] } ``` ## SDK helpers — `parse()` (recommended) ```python from pydantic import BaseModel class ContactInfo(BaseModel): name: str email: str plan: str demo_requested: bool response = client.messages.parse( model="claude-opus-5", max_tokens=16000, messages=[{"role": "user", "content": "Extract: Jane Doe (jane@co.com), Enterprise, wants a demo."}], output_format=ContactInfo, # parse() convenience kwarg ) contact = response.parsed_output # validated ContactInfo instance ``` ```typescript import { z } from "zod"; import { zodOutputFormat } from "@anthropic-ai/sdk/helpers/zod"; const ContactInfo = z.object({ name: z.string(), email: z.string(), plan: z.string(), demo_requested: z.boolean(), }); const response = await client.messages.parse({ model: "claude-opus-5", max_tokens: 16000, output_config: { format: zodOutputFormat(ContactInfo) }, messages: [{ role: "user", content: "Extract: ..." }], }); console.log(response.parsed_output!.name); // null if parsing failed — guard it ``` The SDKs strip unsupported schema constraints (e.g. `minLength`) before sending and validate them client-side instead. ## Strict Tool Use ```python tools = [{ "name": "book_flight", "description": "Book a flight", "strict": True, "input_schema": { "type": "object", "properties": { "destination": {"type": "string"}, "date": {"type": "string", "format": "date"}, "passengers": {"type": "integer", "enum": [1, 2, 3, 4, 5, 6, 7, 8]}, }, "required": ["destination", "date", "passengers"], "additionalProperties": False, }, }] ``` - Per-tool opt-in; non-strict tools don't count toward complexity limits. - Max **20 strict tools** per request. - Guarantees the `tool_use.input` validates exactly — no missing required fields, no type drift. ## JSON Schema: supported vs not **Supported:** object/array/string/integer/number/boolean/null; `enum` (scalars only); `const`; `anyOf`/`allOf` (no `allOf` + `$ref` combo); internal `$ref`/`$defs`; `default`; `required`; `additionalProperties: false` (mandatory on every object); string `format` (`date-time`, `time`, `date`, `duration`, `email`, `hostname`, `uri`, `ipv4`, `ipv6`, `uuid`); array `minItems` 0 or 1 only; simple regex `pattern`. **Not supported:** recursive schemas; external `$ref`; numeric constraints (`minimum`/`maximum`/`multipleOf`); string length constraints (`minLength`/`maxLength`); array constraints beyond `minItems` 0/1; regex backreferences, lookahead/lookbehind, `\b`; `additionalProperties` anything but `false`. **Complexity limits:** 20 strict tools; 24 optional parameters total across all schemas; 16 union-typed (`anyOf`) parameters; grammar compilation timeout 180s ("Schema is too complex"). Reduce by flattening nesting, making params required, splitting across requests. ## Operational notes - **First-request latency:** new schemas compile a grammar on first use; cached for 24h (keyed on schema + tool set; name/description changes don't invalidate). - **Prompt cache interplay:** changing `output_config.format` invalidates the prompt cache; the feature also injects an extra system prompt (more input tokens). - **Failure modes:** `stop_reason: "refusal"` → output may not match the schema; `stop_reason: "max_tokens"` → JSON may be truncated/incomplete — raise `max_tokens` and check before parsing. - **Incompatible with:** citations (400) and assistant-message prefilling. **Works with:** batches, streaming, token counting, extended/adaptive thinking. - Don't put PHI/PII in schema property names, enum values, or patterns — schemas are cached separately from ZDR-handled message content. ## Structured outputs vs tool-use extraction Before structured outputs existed, the standard extraction trick was a forced tool call (`tool_choice: {"type": "tool", "name": "record_result"}`) with the target shape as `input_schema`. Decision now: | Want | Use | |---|---| | The *final answer* as guaranteed JSON | `output_config.format` | | Valid *parameters* for a real action/function | tool + `strict: true` | | Extraction under **manual** extended thinking (`type: "enabled"`) | `output_config.format` — forced `tool_choice` is a 400 in that mode | | Extraction mid-agentic-loop (model also has other tools) | A strict "report/record" tool keeps the loop uniform | | Legacy prefill (`{"name": "` assistant prefill) | Dead on Opus 4.7+ models (400) — migrate to `output_config.format` | ## Thinking interplay - `output_config.format` **works with adaptive/extended thinking** — the model thinks, then the final text block conforms to the schema. - Forced tool extraction is rejected only under **manual** extended thinking (`{"type": "enabled"}` + `tool_choice: any/tool` = 400). Adaptive thinking — what every current model uses — accepts forced tool choice, so this is no longer a reason to avoid it; prefer `output_config.format` because it targets the *answer* rather than a tool's arguments. - Effort and format coexist in `output_config`: `output_config={"effort": "medium", "format": {...}}`. -
tool-use.md 11 KB
# Tool Use Reference Tool use is a feature of `POST /v1/messages` — you pass `tools`, Claude responds with `tool_use` content blocks, you execute and return `tool_result` blocks. **Client tools** run in your code; **server tools** (web_search, code_execution, web_fetch) run on Anthropic's infrastructure. ## Tool Definition ```json { "name": "get_weather", "description": "Get current weather for a location. Call this when the user asks about weather conditions, temperature, or forecasts.", "input_schema": { "type": "object", "properties": { "location": {"type": "string", "description": "City and state, e.g. San Francisco, CA"}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "description": "Temperature unit"} }, "required": ["location"] } } ``` Rules that actually move the needle: - **Descriptions are the routing signal.** Be prescriptive about *when* to call, not just what it does ("Call this when the user asks about current prices or recent events"). Recent Opus models reach for tools more conservatively — trigger conditions in the description give measurable lift. - Describe every property; use `enum` for closed sets; only truly-required params in `required`. - Specific names beat generic ones: `get_current_weather` > `weather`. - Keep the tool set focused — too many similar tools degrades selection. For large tool libraries, use the server-side **tool search tool** (loads schemas on demand, preserves the prompt cache by appending rather than swapping). - **Sample calls in definitions**: you can include example invocations (`input_examples` on supported surfaces / examples embedded in the description) to demonstrate parameter formats and cut parameter errors — most useful for complex nested schemas. - `strict: true` on a tool guarantees the emitted `input` validates against the schema exactly (see structured-outputs.md; max 20 strict tools/request). ## tool_choice | Value | Behavior | |---|---| | `{"type": "auto"}` | Claude decides (default when tools are present) | | `{"type": "any"}` | Must call at least one tool | | `{"type": "tool", "name": "get_weather"}` | Must call that specific tool | | `{"type": "none"}` | Cannot call tools (definitions stay in context) | Any variant accepts `"disable_parallel_tool_use": true` to cap at one tool call per response. Gotchas: - **Manual extended thinking (`{"type": "enabled"}`) + `any`/`tool` = 400.** Only `auto`/`none` are compatible with it. **Adaptive thinking has no such limit** — including the models where it is on by default (Fable 5, Opus 5, Sonnet 5), forced tool choice is accepted. - `any`/`tool` add more tool-use system-prompt tokens than `auto`/`none` (a measured ~410 vs ~290 on Opus 4.8; re-measure with `count_tokens` rather than carrying the figure across model generations). - Changing `tool_choice` between requests does **not** invalidate the tools+system prompt cache (message cache only). ## Parallel Tool Use By default Claude may emit **multiple `tool_use` blocks in one response**. Execute them all (concurrently if safe), then return **all results in a single user message** — one `tool_result` per `tool_use`, ids matching, results may be in any order but must all be present: ```python tool_results = [] for block in response.content: if block.type == "tool_use": result = execute_tool(block.name, block.input) # block.input is parsed dict tool_results.append({ "type": "tool_result", "tool_use_id": block.id, "content": result, }) messages.append({"role": "assistant", "content": response.content}) messages.append({"role": "user", "content": tool_results}) ``` A follow-up request missing a `tool_result` for any outstanding `tool_use` id is a 400. ## The Agentic Loop (manual) Use the manual loop when you need approval gates, custom logging, or conditional execution: ```python import anthropic client = anthropic.Anthropic() messages = [{"role": "user", "content": user_input}] while True: response = client.messages.create( model="claude-opus-5", max_tokens=16000, tools=tools, messages=messages, ) if response.stop_reason == "end_turn": break if response.stop_reason == "pause_turn": # Server-side tool loop hit its iteration limit: append and re-send. # Do NOT inject a "continue" user message — the API resumes automatically. messages.append({"role": "assistant", "content": response.content}) continue if response.stop_reason == "tool_use": messages.append({"role": "assistant", "content": response.content}) results = [] for block in response.content: if block.type == "tool_use": try: out = execute_tool(block.name, block.input) results.append({"type": "tool_result", "tool_use_id": block.id, "content": out}) except Exception as e: results.append({"type": "tool_result", "tool_use_id": block.id, "content": f"Error: {e}", "is_error": True}) messages.append({"role": "user", "content": results}) continue break # max_tokens / refusal / stop_sequence — handle per stop_reason final_text = next((b.text for b in response.content if b.type == "text"), "") ``` ```typescript import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic(); const messages: Anthropic.MessageParam[] = [{ role: "user", content: userInput }]; while (true) { const response = await client.messages.create({ model: "claude-opus-5", max_tokens: 16000, tools, messages, }); if (response.stop_reason === "end_turn") break; if (response.stop_reason === "pause_turn") { messages.push({ role: "assistant", content: response.content }); continue; } const toolUses = response.content.filter( (b): b is Anthropic.ToolUseBlock => b.type === "tool_use", ); messages.push({ role: "assistant", content: response.content }); const results: Anthropic.ToolResultBlockParam[] = []; for (const t of toolUses) { results.push({ type: "tool_result", tool_use_id: t.id, content: await executeTool(t.name, t.input) }); } messages.push({ role: "user", content: results }); } ``` Loop invariants: 1. Append the **full** `response.content` as the assistant turn (preserves `tool_use` + `thinking` blocks; thinking `signature` must round-trip untouched). 2. One `tool_result` per `tool_use`, matching `tool_use_id`. 3. Tool results go in a **user** message. 4. Add a max-iterations guard (e.g. 10-20) so a confused model can't loop forever. 5. Parse `block.input` as structured data — never regex the serialized JSON (escaping varies across models). ## Tool Result Shapes ```json {"type": "tool_result", "tool_use_id": "toolu_01...", "content": "plain string"} {"type": "tool_result", "tool_use_id": "toolu_01...", "content": [{"type": "text", "text": "..."}, {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": "..."}}]} {"type": "tool_result", "tool_use_id": "toolu_01...", "content": "Error: location 'xyz' not found. Provide a valid city.", "is_error": true} ``` `is_error: true` tells Claude the execution failed — it will typically adjust approach or ask for clarification. Return **informative** error strings, not stack traces. ## SDK Tool Runners (beta) — skip the manual loop The runners handle call → execute → feed-back → repeat automatically. ```python from anthropic import beta_tool import anthropic client = anthropic.Anthropic() @beta_tool def get_weather(location: str, unit: str = "celsius") -> str: """Get current weather for a location. Args: location: City and state, e.g. San Francisco, CA. unit: "celsius" or "fahrenheit". """ return f"22°C and sunny in {location}" runner = client.beta.messages.tool_runner( model="claude-opus-5", max_tokens=16000, tools=[get_weather], messages=[{"role": "user", "content": "Weather in Paris?"}], ) for message in runner: # iterates messages until Claude stops calling tools print(message) ``` ```typescript import Anthropic from "@anthropic-ai/sdk"; import { betaZodTool } from "@anthropic-ai/sdk/helpers/beta/zod"; import { z } from "zod"; const getWeather = betaZodTool({ name: "get_weather", description: "Get current weather for a location", inputSchema: z.object({ location: z.string().describe("City and state, e.g. San Francisco, CA"), }), run: async ({ location }) => `22°C and sunny in ${location}`, }); const finalMessage = await client.beta.messages.toolRunner({ model: "claude-opus-5", max_tokens: 16000, tools: [getWeather], messages: [{ role: "user", content: "Weather in Paris?" }], }); ``` Schemas are generated from the function signature/docstring (Python) or Zod schema (TS). The runner executes tools **automatically** — for destructive side effects (email, payments, deletes), validate inside the tool function or use the manual loop with a human-approval gate. ## Server-Side Tools Declared in `tools`, executed by Anthropic — no client handling: | Tool | Type string | Notes | |---|---|---| | Web search | `web_search_20260209` | Per-search pricing; `allowed_domains`, `blocked_domains`, `max_uses` | | Web fetch | `web_fetch_20260209` | Fetch URL content with citations | | Code execution | `code_execution_20260120` | Sandboxed Python 3.11 (pandas/numpy/matplotlib preinstalled), no internet; container reusable via response `container.id` | | Tool search | `tool_search_tool_bm25_20251119` / `..._regex_20251119` | On-demand tool discovery for large libraries | | Memory | `memory_20250818` | Client-executed but Anthropic-defined schema | | Bash / text editor | `bash_20250124` / `text_editor_20250728` (name `str_replace_based_edit_tool`) | Anthropic-defined, **you** execute | Server tool runs may return `stop_reason: "pause_turn"` when the server-side loop hits its iteration limit (default 10) — append the assistant turn and re-request to resume (see loop above). Cap continuations (~5) to avoid infinite resumes. `web_search_20260209` / `web_fetch_20260209` include **dynamic filtering** (model filters results in a sandbox before they hit context) automatically — don't add a separate `code_execution` tool just for that; only declare code_execution when you need it independently. ## Bash vs Dedicated Tools (design) A bash tool gives breadth but the harness sees only an opaque command string. Promote an action to a dedicated tool when you need to: - **Gate** it (send_email behind confirmation — easy as a tool, impossible as `bash -c "curl ..."`) - **Validate** invariants (an edit tool can reject writes to files changed since last read) - **Render** it specially (question-asking as a modal) - **Parallelize** safely (read-only tools marked parallel-safe; bash must serialize) Start with bash for breadth; promote when you need to gate, render, audit, or parallelize.
-
-
scripts
-
.gitkeep 0 B · in bundle
-
check-model-table.py 32.2 KB
#!/usr/bin/env python3 """Staleness verifier for the claude-api-ops model + cache-minimum tables. Guards the two fast-moving fact tables in this skill against silent drift: - the "Current Models" table in SKILL.md (ids, pricing, context, output) - the per-model prompt-cache minimum table in references/caching-and-cost.md Two modes (protocol SKILL-RESOURCE-PROTOCOL.md §7): --offline (default): parse both tables, assert internal consistency. No network. Exit 4 (VALIDATION) on a malformed/contradictory row. --live: curl the Anthropic Models API and compare its model-id set against the documented ids. Advisory only. Live-mode scope limit: the Models API returns model IDs but NOT pricing, context, or output limits. --live therefore verifies model-ID existence/coverage ONLY: - a documented id absent from the live list -> DRIFT (retired/typo) - a live id newer than anything documented -> DRIFT (table lacks a new model) Pricing/context/output drift is out of scope for --live (the API can't confirm it); --offline guards their well-formedness, and the SKILL.md "Live Documentation" links remain the human cross-check for pricing. Usage: check-model-table.py [--offline | --live] [--json] [--skill-dir DIR] [-q] Input: reads SKILL.md and references/caching-and-cost.md (resolved relative to this script, or --skill-dir) Output: stdout = data only (JSON envelope under --json, else a plain summary) Stderr: headers, progress, notes, errors Exit: 0 ok/consistent, 2 usage, 3 not-found, 4 validation (malformed/contradictory), 5 missing-dep (curl, --live only), 7 unavailable (no key / API unreachable), 10 drift (live id-set disagrees with the table) Examples: check-model-table.py --offline check-model-table.py --offline --json | python -m json.tool ANTHROPIC_API_KEY=sk-... check-model-table.py --live check-model-table.py --live # exits 7 (advisory) when the key is unset """ from __future__ import annotations import argparse import json import os import re import subprocess import sys from pathlib import Path from typing import NoReturn # Windows consoles default to cp1252; force UTF-8 so em-dashes/§ in notes don't # raise UnicodeEncodeError or print mojibake (matches the repo's standard fix). for _stream in (sys.stdout, sys.stderr): try: _stream.reconfigure(encoding="utf-8") # type: ignore[attr-defined] except (AttributeError, ValueError): pass class Term: """Tiny ANSI helper mirroring skills/_lib/term.sh (term.sh is bash-only; per TERMINAL-DESIGN.md §9 the Python port is inline with matching keys/glyphs). Honors FORCE_COLOR / NO_COLOR / TERM_ASCII; color tracks the bound stream's TTY, and glyphs fall back to ASCII on TERM_ASCII or a non-UTF stream encoding.""" _C = {"green": "\033[32m", "yellow": "\033[33m", "orange": "\033[38;5;208m", "red": "\033[31m", "cyan": "\033[36m", "dim": "\033[2m", "off": "\033[0m"} _GLYPH = {"ok": "✓", "bad": "✗", "warn": "▲", "skip": "—", "na": "—", "unknown": "?"} _ASCII = {"ok": "+", "bad": "x", "warn": "!", "skip": "-", "na": "-", "unknown": "?"} _MARK_COLOR = {"ok": "green", "bad": "red", "warn": "orange", "skip": "dim", "na": "dim", "unknown": "yellow"} def __init__(self, stream=sys.stderr): enc = (getattr(stream, "encoding", "") or "").lower() self.ascii = (os.environ.get("TERM_ASCII") == "1" or os.environ.get("FLEET_ASCII") == "1" or "utf" not in enc) if os.environ.get("FORCE_COLOR"): self.color = True elif (os.environ.get("NO_COLOR") is not None or os.environ.get("TERM") == "dumb" or not getattr(stream, "isatty", lambda: False)()): self.color = False else: self.color = True def c(self, name, text): return f"{self._C.get(name, '')}{text}{self._C['off']}" if self.color else text def mark(self, state): return self.c(self._MARK_COLOR.get(state, ""), (self._ASCII if self.ascii else self._GLYPH).get(state, ".")) def hdr(self, text): return self.c("cyan", f"=== {text} ===") TERM = Term(sys.stderr) EXIT_OK = 0 EXIT_USAGE = 2 EXIT_NOT_FOUND = 3 EXIT_VALIDATION = 4 EXIT_MISSING_DEP = 5 EXIT_UNAVAILABLE = 7 EXIT_DRIFT = 10 SCHEMA = "claude-mods.claude-api-ops.model-table/v1" MODELS_API = "https://api.anthropic.com/v1/models?limit=1000" ANTHROPIC_VERSION = "2023-06-01" # --- Cache-economics constants (offline cross-file tripwire) ----------------- # These numbers are stated in more than one file, so a future edit can desync # them silently. This table is the single source of truth: each entry names the # canonical value, a regex that must match in EVERY listed file, and why. # Changing a real-world value means changing it here AND in every listed file -- # that is the point: a one-file edit trips this check instead of rotting. # # Deliberately offline-only. A --live parse of the docs page would be brittle # (prose/format churn -> spurious exit 10), and SKILL-RESOURCE-PROTOCOL.md §7 is # explicit that a gate which cries wolf is a gate everyone learns to ignore. # The human cross-check stays the "Live Documentation" link in SKILL.md. CACHE_CONSTANTS = [ { "key": "cache_read_multiplier", "value": "0.1x base input", "pattern": r"0\.1\s*[x×]", "files": ["SKILL.md", "references/caching-and-cost.md", "references/context-engineering.md", "references/compaction.md"], }, { "key": "cache_write_5m_multiplier", "value": "1.25x base input", "pattern": r"1\.25\s*[x×]", "files": ["references/caching-and-cost.md", "references/compaction.md"], }, { "key": "cache_write_1h_multiplier", "value": "2x base input", # Anchored to the TTL in either order so a bare "2x" elsewhere can't # satisfy it. Note "2×" (U+00D7) has no trailing \b -- don't add one. "pattern": (r"(?<![\d.])2\s*[x×][^\n]{0,40}1-hour" r"|1-hour[^\n]{0,40}(?<![\d.])2\s*[x×]"), "files": ["references/caching-and-cost.md", "references/compaction.md"], }, { "key": "max_breakpoints", "value": "4 cache_control breakpoints per request", "pattern": r"\*\*4\*\*\s*`?cache_control`?\s*breakpoints", "files": ["references/caching-and-cost.md"], }, { "key": "lookback_blocks", "value": "20-block backward lookback per breakpoint", "pattern": r"\b20\b[^.\n]{0,48}\bblock", "files": ["references/caching-and-cost.md", "references/context-engineering.md"], }, { # context-budget.py encodes the same multipliers as executable constants. # Prose and code drifting apart is the exact failure this guards: a doc # that says 0.1x beside a calculator that bills 0.15x. "key": "cache_multipliers_in_calculator", "value": "context-budget.py constants match the documented multipliers", "pattern": (r"CACHE_READ_MULTIPLIER\s*=\s*0\.1\b[\s\S]{0,200}?" r"CACHE_WRITE_5M\s*=\s*1\.25\b[\s\S]{0,200}?" r"CACHE_WRITE_1H\s*=\s*2\.0\b"), "files": ["scripts/context-budget.py"], }, { "key": "context_management_beta", "value": "context-management-2025-06-27 beta header", "pattern": r"context-management-2025-06-27", "files": ["SKILL.md", "references/compaction.md"], }, ] # Docs whose facts were verified on a date must SAY so, so staleness is visible # to a reader rather than only to this script. DATE_STAMPED_DOCS = ["references/context-engineering.md", "references/compaction.md"] DATE_STAMP_RE = re.compile(r"verified[^\n]{0,80}?(\d{4}-\d{2}-\d{2})", re.I) # A well-formed alias id: claude-<word>-<digit>... and NO date suffix. # Accepts claude-opus-5, claude-fable-5, claude-sonnet-5, claude-haiku-4-5. ID_RE = re.compile(r"^claude-[a-z]+-\d+(?:-\d+)?$") # A date suffix looks like an 8-digit run (e.g. -20251114). DATE_SUFFIX_RE = re.compile(r"-\d{8}$") # Any Claude model id appearing in the skill body, dated snapshots included. ANY_ID_RE = re.compile(r"\bclaude-[a-z]+-\d+(?:-\d+)?(?:-\d{8})?\b") # A model id being ASSIGNED - what a reader copies. Covers model="x", model: "x", # "model": "x", MODEL = "x", and --model x. ASSIGN_RE = re.compile( r'(?:"?\bmodel"?\s*[:=]\s*["\']([a-z0-9.\-]+)["\'])' r'|(?:--model[=\s]+([a-z0-9.\-]+))', re.IGNORECASE, ) # Escape hatch: a line carrying this marker may name a legacy id deliberately # (e.g. a migration before/after snippet). LEGACY_OK = "legacy-ok" # tests/ is fixture material - it deliberately writes malformed ids to prove the # verifier rejects them, so scanning it would be self-defeating. SCAN_SKIP_DIRS = {"tests", "__pycache__", ".git"} SCAN_SUFFIXES = {".md", ".py", ".json", ".sh", ".ts", ".js", ".yaml", ".yml", ".txt"} def note(msg: str, quiet: bool) -> None: if not quiet: print(msg, file=sys.stderr) def fail_validation(message: str, details: dict, json_mode: bool) -> NoReturn: if json_mode: print(json.dumps({"error": {"code": "VALIDATION", "message": message, "details": details}})) print(f"{TERM.mark('bad')} ERROR: {message}", file=sys.stderr) for k, v in details.items(): print(f" {k}: {v}", file=sys.stderr) sys.exit(EXIT_VALIDATION) # --------------------------------------------------------------------------- # Parsing # --------------------------------------------------------------------------- def split_row(line: str) -> list[str]: """Split a markdown table row into trimmed cells (drops outer pipes).""" cells = [c.strip() for c in line.strip().strip("|").split("|")] return cells def is_separator(cells: list[str]) -> bool: return all(re.fullmatch(r":?-{2,}:?", c or "") for c in cells) and bool(cells) def parse_model_table(text: str) -> tuple[list[dict], list[str]]: """Parse the SKILL.md 'Current Models' table. Columns: Model | ID | Context | Max Output | Input $/MTok | Output $/MTok Returns one dict per data row. """ lines = text.splitlines() # Locate the header row that contains the ID column and a price column. start = None for i, line in enumerate(lines): low = line.lower() if line.lstrip().startswith("|") and "id" in low and "context" in low and "output" in low: start = i break if start is None: return [], [] header = split_row(lines[start]) rows: list[dict] = [] # Expect a separator row next, then data rows until a non-table line. j = start + 1 if j < len(lines) and is_separator(split_row(lines[j])): j += 1 while j < len(lines): line = lines[j] if not line.lstrip().startswith("|"): break cells = split_row(line) if is_separator(cells): j += 1 continue if len(cells) >= 6: rows.append({ "name": cells[0], "id_cell": cells[1], "context": cells[2], "max_output": cells[3], "input_price": cells[4], "output_price": cells[5], }) j += 1 return rows, header def parse_legacy_ids(text: str) -> list[str]: """Extract the backticked ids from SKILL.md's 'Legacy (still available...)' paragraph. Parsed rather than hard-coded so retiring a model is a one-line doc edit, not a code change.""" m = re.search(r"\*\*Legacy \(still available[^*]*\)\:\*\*(.+?)(?:\n\n|\Z)", text, re.S) if not m: return [] return [i for i in re.findall(r"`([^`]+)`", m.group(1)) if ID_RE.match(i)] def scan_body_ids(skill_dir: Path, current: set[str], legacy: set[str]) -> list[dict]: """Flag model ids used in the skill body that a reader must not copy. Two distinct faults, because they have different severities in practice: unknown - an id in neither the current table nor the legacy list. A typo or a hallucinated model; nothing resolves it. retired - a LEGACY id sitting in an assignable position (model=..., MODEL = ..., "model": ..., --model ...). Naming a legacy model in prose is legitimate; shipping one as the value a reader copies is not. Append the `legacy-ok` marker to that line to allow it. """ findings: list[dict] = [] for path in sorted(skill_dir.rglob("*")): if not path.is_file() or path.suffix.lower() not in SCAN_SUFFIXES: continue if SCAN_SKIP_DIRS & set(p.name for p in path.relative_to(skill_dir).parents): continue rel = path.relative_to(skill_dir).as_posix() try: lines = path.read_text(encoding="utf-8").splitlines() except UnicodeDecodeError: continue for n, line in enumerate(lines, 1): if LEGACY_OK in line: continue # Strip a dated snapshot to its alias before membership testing: a # dated id is a real, resolvable snapshot of a known model, and the # docs discuss several deliberately. for raw in ANY_ID_RE.findall(line): base = DATE_SUFFIX_RE.sub("", raw) if base not in current and base not in legacy: findings.append({"file": rel, "line": n, "id": raw, "fault": "unknown"}) for a, b in ASSIGN_RE.findall(line): val = DATE_SUFFIX_RE.sub("", (a or b).strip()) if val in legacy: findings.append({"file": rel, "line": n, "id": val, "fault": "retired"}) return findings def parse_cache_min_table(text: str) -> list[dict]: """Parse the caching-and-cost.md 'Minimum prefix tokens' table. Columns: Model | Minimum prefix tokens. The Model cell holds friendly names (possibly several comma-separated), not ids. """ lines = text.splitlines() start = None for i, line in enumerate(lines): low = line.lower() if line.lstrip().startswith("|") and "model" in low and "minimum" in low and "prefix" in low: start = i break if start is None: return [] rows: list[dict] = [] j = start + 1 if j < len(lines) and is_separator(split_row(lines[j])): j += 1 while j < len(lines): line = lines[j] if not line.lstrip().startswith("|"): break cells = split_row(line) if is_separator(cells): j += 1 continue if len(cells) >= 2: rows.append({"names": cells[0], "min_tokens": cells[1]}) j += 1 return rows # --------------------------------------------------------------------------- # Offline validation # --------------------------------------------------------------------------- PRICE_RE = re.compile(r"^\$\d+(?:\.\d+)?$") SIZE_RE = re.compile(r"^\d+(?:\.\d+)?[KM]$") def clean_id(id_cell: str) -> str: """Strip backtick code fences from an ID cell.""" return id_cell.strip().strip("`").strip() def validate_cache_constants(skill_dir: Path, json_mode: bool, quiet: bool) -> list[dict]: """Assert every file that states a cache-economics constant states the same one. Presence-based on purpose: if an edit changes 0.1x to 0.15x in one file, that file stops matching and this trips. Returns one result row per constant. """ rows: list[dict] = [] for const in CACHE_CONSTANTS: rx = re.compile(const["pattern"]) missing: list[str] = [] for rel in const["files"]: path = skill_dir / rel if not path.is_file(): fail_validation("file referenced by a cache constant is missing", {"constant": const["key"], "file": rel}, json_mode) if not rx.search(path.read_text(encoding="utf-8")): missing.append(rel) if missing: fail_validation( f"cache constant {const['key']!r} not stated consistently across files", {"expected": const["value"], "pattern": const["pattern"], "files_missing_it": ", ".join(missing), "hint": ("either the doc drifted, or the real value changed and " "CACHE_CONSTANTS in this script needs updating too")}, json_mode) rows.append({"key": const["key"], "value": const["value"], "files": const["files"], "consistent": True}) note(f" {len(rows)} cache constants consistent across files", quiet) return rows def validate_date_stamps(skill_dir: Path, json_mode: bool, quiet: bool) -> list[dict]: """Every fact-heavy doctrine reference must carry a 'verified <ISO date>' stamp.""" rows: list[dict] = [] for rel in DATE_STAMPED_DOCS: path = skill_dir / rel if not path.is_file(): fail_validation("date-stamped doc is missing", {"file": rel}, json_mode) m = DATE_STAMP_RE.search(path.read_text(encoding="utf-8")) if not m: fail_validation( "doc encodes fast-moving facts but carries no verification date", {"file": rel, "hint": "add a line like 'Facts verified against ... YYYY-MM-DD'"}, json_mode) rows.append({"file": rel, "verified": m.group(1)}) note(f" {len(rows)} doctrine docs carry a verification date", quiet) return rows def validate_cited_references(skill_dir: Path, json_mode: bool, quiet: bool) -> list[str]: """Every references/*.md on disk must be cited from SKILL.md, and vice versa. SKILL-RESOURCE-PROTOCOL.md §1: an uncited reference is dead weight the router never finds; a cited-but-absent one is a broken link. """ skill_text = (skill_dir / "SKILL.md").read_text(encoding="utf-8") on_disk = sorted(p.name for p in (skill_dir / "references").glob("*.md")) cited = set(re.findall(r"\(references/([A-Za-z0-9._-]+\.md)\)", skill_text)) uncited = [n for n in on_disk if n not in cited] ghosts = sorted(n for n in cited if n not in on_disk) if uncited or ghosts: fail_validation( "SKILL.md and references/ disagree", {"on_disk_but_uncited": ", ".join(uncited) or "(none)", "cited_but_missing": ", ".join(ghosts) or "(none)"}, json_mode) # Same rule for scripts/ and assets/ (SKILL-RESOURCE-PROTOCOL.md §2.8: every # resource is cited from SKILL.md). Those are referenced as backticked paths # or worked invocations rather than markdown links, so match on the basename. shipped: list[str] = [] for sub in ("scripts", "assets"): for p in sorted((skill_dir / sub).glob("*")): if not p.is_file() or p.name.startswith(".") or p.suffix == ".pyc": continue if p.name not in skill_text: fail_validation( f"{sub}/ file is not cited from SKILL.md", {"file": f"{sub}/{p.name}", "hint": "an uncited resource is dead weight the router never " "finds - cite it with a worked invocation"}, json_mode) shipped.append(f"{sub}/{p.name}") note(f" {len(on_disk)} reference files + {len(shipped)} scripts/assets, " "all cited from SKILL.md", quiet) return on_disk + shipped def validate_offline(skill_dir: Path, json_mode: bool, quiet: bool) -> dict: skill_md = skill_dir / "SKILL.md" cache_md = skill_dir / "references" / "caching-and-cost.md" for p in (skill_md, cache_md): if not p.is_file(): if json_mode: print(json.dumps({"error": {"code": "NOT_FOUND", "message": f"missing file: {p}", "details": {}}})) print(f"ERROR: required file not found: {p}", file=sys.stderr) sys.exit(EXIT_NOT_FOUND) note(TERM.hdr("offline model-table consistency check"), quiet) model_rows, _ = parse_model_table(skill_md.read_text(encoding="utf-8")) if not model_rows: fail_validation("could not locate a non-empty Current Models table in SKILL.md", {"file": str(skill_md)}, json_mode) documented_ids: list[str] = [] models_out: list[dict] = [] for row in model_rows: mid = clean_id(row["id_cell"]) problems = [] if not ID_RE.match(mid): problems.append("id does not match claude-[a-z]+-<digits>") if DATE_SUFFIX_RE.search(mid): problems.append("id carries a date suffix (should be a bare alias)") if not PRICE_RE.match(row["input_price"]): problems.append(f"input price not numeric: {row['input_price']!r}") if not PRICE_RE.match(row["output_price"]): problems.append(f"output price not numeric: {row['output_price']!r}") if not SIZE_RE.match(row["context"]): problems.append(f"context not a size (e.g. 1M/200K): {row['context']!r}") if not SIZE_RE.match(row["max_output"]): problems.append(f"max output not a size: {row['max_output']!r}") if problems: fail_validation(f"malformed model row: {row['name']!r}", {"id": mid, "problems": "; ".join(problems)}, json_mode) documented_ids.append(mid) models_out.append({ "name": row["name"], "id": mid, "context": row["context"], "max_output": row["max_output"], "input_price": row["input_price"], "output_price": row["output_price"], }) # No duplicate ids. dupes = {x for x in documented_ids if documented_ids.count(x) > 1} if dupes: fail_validation("duplicate model ids in the table", {"ids": ", ".join(sorted(dupes))}, json_mode) # Cache-minimum table. cache_rows = parse_cache_min_table(cache_md.read_text(encoding="utf-8")) if not cache_rows: fail_validation("could not locate the cache-minimum table in caching-and-cost.md", {"file": str(cache_md)}, json_mode) for crow in cache_rows: if not re.fullmatch(r"\d+", crow["min_tokens"]): fail_validation("cache-minimum value is not an integer", {"row": crow["names"], "value": crow["min_tokens"]}, json_mode) # Cross-file consistency: every model NAME (e.g. "Opus 4.8", "Fable 5", # "Sonnet 4.6", "Haiku 4.5") in the model table must appear in the cache # table's name set, so the two files agree on the model lineup. cache_blob = " ".join(c["names"] for c in cache_rows).lower() missing_in_cache: list[str] = [] for m in models_out: # Derive the short family+version token, e.g. "Claude Opus 4.8" -> "opus 4.8". short = re.sub(r"^claude\s+", "", m["name"], flags=re.I).strip().lower() if short not in cache_blob: missing_in_cache.append(m["name"]) if missing_in_cache: fail_validation( "model(s) in SKILL.md absent from the cache-minimum table — files contradict", {"missing": ", ".join(missing_in_cache), "hint": "every documented model needs a prompt-cache minimum row"}, json_mode) # Body scan: the tables can be perfectly consistent while a code sample # ships a retired model. Guard what readers copy, not just what they read. legacy_ids = set(parse_legacy_ids(skill_md.read_text(encoding="utf-8"))) body = scan_body_ids(skill_dir, set(documented_ids), legacy_ids) if body: detail = {} for f in body[:12]: detail[f"{f['file']}:{f['line']}"] = f"{f['fault']}: {f['id']}" if len(body) > 12: detail["..."] = f"{len(body) - 12} more" fail_validation( f"{len(body)} model-id problem(s) in the skill body", {**detail, "hint": "'retired' = a legacy id in an assignable position; retarget " "it at a current model or append the 'legacy-ok' marker. " "'unknown' = an id in neither the model table nor the legacy list."}, json_mode) note(f" {len(models_out)} model rows, all well-formed", quiet) note(f" {len(cache_rows)} cache-minimum rows, all integer", quiet) note(" cross-file model lineup consistent", quiet) note(f" {len(legacy_ids)} legacy id(s) tracked; no retired/unknown id in the body", quiet) # Context-engineering layer: constants stated in several files, verification # date stamps, and SKILL.md <-> references/ citation integrity. constants = validate_cache_constants(skill_dir, json_mode, quiet) stamps = validate_date_stamps(skill_dir, json_mode, quiet) refs = validate_cited_references(skill_dir, json_mode, quiet) note(f"{TERM.mark('ok')} OK: tables internally consistent.", quiet) return { "mode": "offline", "models": models_out, "documented_ids": documented_ids, "cache_min_rows": cache_rows, "cache_constants": constants, "date_stamps": stamps, "reference_files": refs, "legacy_ids": sorted(legacy_ids), "consistent": True, } # --------------------------------------------------------------------------- # Live validation # --------------------------------------------------------------------------- def fetch_live_ids(quiet: bool) -> list[str] | None: """Return the live model-id list, or None if unavailable (advisory).""" key = os.environ.get("ANTHROPIC_API_KEY", "").strip() if not key: note("NOTE: ANTHROPIC_API_KEY is unset - skipping live check (advisory).", quiet) return None cmd = [ "curl", "-fsS", "--max-time", "20", "-H", f"x-api-key: {key}", "-H", f"anthropic-version: {ANTHROPIC_VERSION}", MODELS_API, ] try: proc = subprocess.run(cmd, capture_output=True, text=True, timeout=30) except (subprocess.TimeoutExpired, OSError) as exc: note(f"NOTE: Models API call failed ({exc}) — advisory, not a failure.", quiet) return None if proc.returncode != 0: note(f"NOTE: Models API unreachable (curl exit {proc.returncode}) — advisory.", quiet) if proc.stderr.strip(): note(f" {proc.stderr.strip().splitlines()[-1]}", quiet) return None try: payload = json.loads(proc.stdout) except json.JSONDecodeError: note("NOTE: Models API returned non-JSON — advisory, not a failure.", quiet) return None data = payload.get("data") if not isinstance(data, list): note("NOTE: Models API JSON missing 'data' list — advisory.", quiet) return None return [m.get("id", "") for m in data if isinstance(m, dict) and m.get("id")] def validate_live(skill_dir: Path, json_mode: bool, quiet: bool) -> dict: if not _have("curl"): if json_mode: print(json.dumps({"error": {"code": "PRECONDITION", "message": "curl required for --live", "details": {}}})) print("ERROR: curl is required for --live", file=sys.stderr) sys.exit(EXIT_MISSING_DEP) # Reuse offline parse for the documented id set (also validates well-formedness). note(TERM.hdr("live model-id coverage check"), quiet) skill_md = skill_dir / "SKILL.md" if not skill_md.is_file(): print(f"ERROR: required file not found: {skill_md}", file=sys.stderr) sys.exit(EXIT_NOT_FOUND) parsed = parse_model_table(skill_md.read_text(encoding="utf-8")) if not parsed or not parsed[0]: fail_validation("could not parse the model table for live comparison", {"file": str(skill_md)}, json_mode) documented = [clean_id(r["id_cell"]) for r in parsed[0]] live = fetch_live_ids(quiet) if live is None: # Advisory: not a failure. Exit 7. if json_mode: print(json.dumps({"data": {"mode": "live", "status": "unavailable", "documented_ids": documented, "live_ids": None}, "meta": {"schema": SCHEMA, "status": "unavailable"}})) sys.exit(EXIT_UNAVAILABLE) live_set = set(live) doc_set = set(documented) # A documented id absent from the live list = drift (retired/typo). missing = sorted(doc_set - live_set) # A live id NEWER than anything documented = drift (table lacks a new model). # Restrict "newer" to well-formed alias ids so we ignore date-suffixed and # snapshot variants the docs intentionally don't list. live_alias = {m for m in live_set if ID_RE.match(m) and not DATE_SUFFIX_RE.search(m)} new_models = sorted(live_alias - doc_set) drift = bool(missing or new_models) result = { "mode": "live", "status": "drift" if drift else "ok", "documented_ids": documented, "live_ids": sorted(live_set), "missing_from_live": missing, "new_in_live": new_models, } if drift: if missing: note(f"{TERM.mark('bad')} {TERM.c('red', 'DRIFT: documented id(s) absent from live Models API:')}", quiet) for m in missing: note(f" {TERM.c('red', '-')} {m}", quiet) if new_models: note(f"{TERM.mark('bad')} {TERM.c('red', 'DRIFT: live Models API has alias id(s) the table lacks:')}", quiet) for m in new_models: note(f" {TERM.c('green', '+')} {m}", quiet) if json_mode: print(json.dumps({"data": result, "meta": {"schema": SCHEMA, "status": "drift"}})) else: print("DRIFT: model-id table disagrees with the live Models API " f"(missing={missing}, new={new_models})") sys.exit(EXIT_DRIFT) note("OK: every documented id exists live; no newer alias id missing from the table.", quiet) return result def _have(tool: str) -> bool: from shutil import which return which(tool) is not None # --------------------------------------------------------------------------- # Main # --------------------------------------------------------------------------- def main(argv: list[str]) -> int: parser = argparse.ArgumentParser( prog="check-model-table.py", add_help=True, description="Staleness verifier for the claude-api-ops model + cache tables.", epilog=( "EXAMPLES:\n" " check-model-table.py --offline\n" " check-model-table.py --offline --json | python -m json.tool\n" " ANTHROPIC_API_KEY=sk-... check-model-table.py --live\n" " check-model-table.py --live # exits 7 (advisory) when key unset\n" ), formatter_class=argparse.RawDescriptionHelpFormatter, ) mode = parser.add_mutually_exclusive_group() mode.add_argument("--offline", action="store_true", help="parse + assert internal consistency, no network (default)") mode.add_argument("--live", action="store_true", help="compare documented ids against the live Models API (advisory)") parser.add_argument("--json", action="store_true", help="emit the JSON envelope on stdout") parser.add_argument("--skill-dir", default=None, help="skill root (default: parent of this script's dir)") parser.add_argument("-q", "--quiet", action="store_true", help="suppress stderr progress/notes") args = parser.parse_args(argv) if args.skill_dir: skill_dir = Path(args.skill_dir).resolve() else: skill_dir = Path(__file__).resolve().parent.parent if not skill_dir.is_dir(): print(f"ERROR: skill dir not found: {skill_dir}", file=sys.stderr) return EXIT_NOT_FOUND if args.live: result = validate_live(skill_dir, args.json, args.quiet) else: result = validate_offline(skill_dir, args.json, args.quiet) if args.json: print(json.dumps({"data": result, "meta": {"schema": SCHEMA, "status": "ok"}})) return EXIT_OK if __name__ == "__main__": try: sys.exit(main(sys.argv[1:])) except KeyboardInterrupt: sys.exit(EXIT_USAGE) -
context-budget.py 17.9 KB
#!/usr/bin/env python3 """Append-vs-compact calculator for a cached multi-turn Claude conversation. Answers the one question the context-engineering doctrine says to ask before compacting: over the turns you actually have left, is rewriting the history cheaper than carrying it? Under prompt caching the answer is usually no, and the arithmetic is short enough that people skip it and guess wrong. Models both paths in dollars (see references/compaction.md §3): append : history rides the cache at CACHE_READ_MULTIPLIER of base input, every remaining turn, growing by --growth-per-turn. compact : summarisation call(s) (cached read in, summary out) + a cache WRITE of the new prefix each time + the remaining turns on the summary, and every token cached before a rewrite is forfeited. Compaction RECURS when the history grows. With --growth-per-turn set, the summary climbs back toward the original size and a real system compacts again; this models that cadence rather than charging a single one-shot rewrite. Assuming one compaction understated its cost by ~2.2x on a 60-turn, 4k-per-turn session. Also checks the hard constraint first: if the projected history overflows the context window, cost is moot and compaction (or offloading) is forced. Two things to understand before trusting the verdict: * The break-even turn count is SCALE-INVARIANT. History size and price both cancel out of fixed_cost / per_turn_saving -- it is driven only by the summary ratio, the output/input price ratio, and the write multiplier. So "compact" vs "append" is almost entirely a question of how many turns you have left, not how big or expensive the conversation is. * RECALL IS NOT PRICED. This models dollars only, which is one of the three constraints in references/compaction.md. The measured recall cost of summarisation (92-100% -> 38-58% on a planted-fact probe) does not appear anywhere in these numbers. A "compact" verdict means compaction is cheaper, NOT that it is right. Probe your own recall before acting on it -- assets/recall-probe.py does exactly that. Usage: context-budget.py --history-tokens N --turns-remaining N [OPTIONS] Input: argv only; no stdin, no network, no files read or written. Output: stdout = data only (JSON envelope under --json, else a plain summary) Stderr: headers, workings, notes Exit: 0 append wins (keep everything), 2 usage, 4 validation (bad numbers), 10 compaction indicated (cost crossover or context ceiling) Examples: # Short session -> append wins (exit 0) context-budget.py --history-tokens 25000 --turns-remaining 5 --base-rate 0.30 # 40 turns left on a 120K history -> cost favours compaction (exit 10) context-budget.py --history-tokens 120000 --turns-remaining 40 --base-rate 2.00 # Premium tier, long session, history still growing each turn context-budget.py --history-tokens 300000 --turns-remaining 60 \ --base-rate 10.00 --growth-per-turn 4000 --ttl 1h # Machine-readable, for an agent deciding mid-run context-budget.py --history-tokens 900000 --turns-remaining 10 \ --base-rate 5.00 --json | python -m json.tool """ from __future__ import annotations import argparse import json import os import sys # Windows consoles default to cp1252; force UTF-8 so the section glyphs in the # human framing don't raise UnicodeEncodeError (matches check-model-table.py). for _stream in (sys.stdout, sys.stderr): try: _stream.reconfigure(encoding="utf-8") # type: ignore[attr-defined] except (AttributeError, ValueError): pass class Term: """Tiny ANSI helper mirroring skills/_lib/term.sh (term.sh is bash-only; per TERMINAL-DESIGN.md §9 the Python port is inline with matching keys/glyphs). Honors FORCE_COLOR / NO_COLOR / TERM_ASCII; color tracks the bound stream's TTY, and glyphs fall back to ASCII on TERM_ASCII or a non-UTF stream encoding.""" _C = {"green": "\033[32m", "yellow": "\033[33m", "orange": "\033[38;5;208m", "red": "\033[31m", "cyan": "\033[36m", "dim": "\033[2m", "off": "\033[0m"} _GLYPH = {"ok": "✓", "bad": "✗", "warn": "▲", "skip": "—", "na": "—", "unknown": "?"} _ASCII = {"ok": "+", "bad": "x", "warn": "!", "skip": "-", "na": "-", "unknown": "?"} _MARK_COLOR = {"ok": "green", "bad": "red", "warn": "orange", "skip": "dim", "na": "dim", "unknown": "yellow"} def __init__(self, stream=sys.stderr): enc = (getattr(stream, "encoding", "") or "").lower() self.ascii = (os.environ.get("TERM_ASCII") == "1" or os.environ.get("FLEET_ASCII") == "1" or "utf" not in enc) if os.environ.get("FORCE_COLOR"): self.color = True elif (os.environ.get("NO_COLOR") is not None or os.environ.get("TERM") == "dumb" or not getattr(stream, "isatty", lambda: False)()): self.color = False else: self.color = True def c(self, name, text): return f"{self._C.get(name, '')}{text}{self._C['off']}" if self.color else text def mark(self, state): return self.c(self._MARK_COLOR.get(state, ""), (self._ASCII if self.ascii else self._GLYPH).get(state, ".")) def hdr(self, text): return self.c("cyan", f"=== {text} ===") TERM = Term(sys.stderr) EXIT_OK = 0 EXIT_USAGE = 2 EXIT_VALIDATION = 4 EXIT_COMPACT = 10 SCHEMA = "claude-mods.claude-api-ops.context-budget/v1" # --- Cache economics (verified against platform.claude.com 2026-08-30) ------- # Guarded by check-model-table.py --offline: these literals are cross-checked # against the prose in SKILL.md and references/*.md, so a one-file edit trips # CI instead of leaving the docs and this calculator quietly disagreeing. CACHE_READ_MULTIPLIER = 0.1 # cache read, every model, flat CACHE_WRITE_5M = 1.25 # cache write, 5-minute TTL CACHE_WRITE_1H = 2.0 # cache write, 1-hour TTL # Every model in the current lineup prices output at exactly 5x its input rate # (10/50, 5/25, 2/10, 1/5), so --output-rate defaults to 5x --base-rate rather # than forcing the caller to look it up. Override it if that ever stops holding. OUTPUT_RATE_RATIO = 5.0 PER_MTOK = 1_000_000.0 def note(msg: str, quiet: bool) -> None: if not quiet: print(msg, file=sys.stderr) def fail(code: str, message: str, details: dict, json_mode: bool, exit_code: int): if json_mode: print(json.dumps({"error": {"code": code, "message": message, "details": details}})) print(f"{TERM.mark('bad')} ERROR: {message}", file=sys.stderr) for k, v in details.items(): print(f" {k}: {v}", file=sys.stderr) sys.exit(exit_code) def cached_carry_cost(start_tokens: float, turns: int, growth: float, rate: float) -> float: """Cost of re-sending a growing history at cache-read price for `turns` turns.""" total_tokens = sum(start_tokens + growth * t for t in range(turns)) return total_tokens * CACHE_READ_MULTIPLIER * rate / PER_MTOK def main(argv: list[str]) -> int: parser = argparse.ArgumentParser( prog="context-budget.py", add_help=True, description="Append-vs-compact calculator for a cached Claude conversation.", epilog=( "EXAMPLES:\n" " context-budget.py --history-tokens 25000 --turns-remaining 5 " "--base-rate 0.30 # append wins (exit 0)\n" " context-budget.py --history-tokens 120000 --turns-remaining 40 " "--base-rate 2.00 # cost favours compaction (exit 10)\n" " context-budget.py --history-tokens 300000 --turns-remaining 60 " "--base-rate 10.00 --growth-per-turn 4000 --ttl 1h\n" " context-budget.py --history-tokens 900000 --turns-remaining 10 " "--base-rate 5.00 --json\n" "\nEXIT: 0 append wins, 2 usage, 4 bad numbers, 10 compaction indicated\n" ), formatter_class=argparse.RawDescriptionHelpFormatter, ) parser.add_argument("--history-tokens", type=float, required=True, help="current conversation/prefix size in tokens") parser.add_argument("--turns-remaining", type=int, required=True, help="turns you expect AFTER this decision (a compaction " "just before the end is pure loss)") parser.add_argument("--base-rate", type=float, default=5.00, help="base INPUT price in $/MTok (default: 5.00)") parser.add_argument("--output-rate", type=float, default=None, help=f"output price in $/MTok (default: {OUTPUT_RATE_RATIO}x " "--base-rate, which holds across the current lineup)") parser.add_argument("--summary-tokens", type=float, default=None, help="expected size of the summary (default: 10%% of history)") parser.add_argument("--growth-per-turn", type=float, default=0.0, help="tokens the history grows each turn (default: 0)") parser.add_argument("--ttl", choices=["5m", "1h"], default="5m", help="cache TTL, sets the write multiplier (default: 5m)") parser.add_argument("--context-window", type=float, default=1_000_000, help="model context window in tokens (default: 1000000)") parser.add_argument("--json", action="store_true", help="emit the JSON envelope on stdout") parser.add_argument("-q", "--quiet", action="store_true", help="suppress stderr framing/workings") args = parser.parse_args(argv) jm, quiet = args.json, args.quiet # --- validate (agents fabricate plausible inputs; §6 of the resource protocol) if args.history_tokens <= 0: fail("VALIDATION", "--history-tokens must be positive", {"got": args.history_tokens}, jm, EXIT_VALIDATION) if args.turns_remaining < 0: fail("VALIDATION", "--turns-remaining cannot be negative", {"got": args.turns_remaining}, jm, EXIT_VALIDATION) if args.base_rate <= 0: fail("VALIDATION", "--base-rate must be positive", {"got": args.base_rate}, jm, EXIT_VALIDATION) if args.growth_per_turn < 0: fail("VALIDATION", "--growth-per-turn cannot be negative", {"got": args.growth_per_turn}, jm, EXIT_VALIDATION) if args.context_window <= 0: fail("VALIDATION", "--context-window must be positive", {"got": args.context_window}, jm, EXIT_VALIDATION) rate = args.base_rate out_rate = args.output_rate if args.output_rate is not None else rate * OUTPUT_RATE_RATIO if out_rate <= 0: fail("VALIDATION", "--output-rate must be positive", {"got": out_rate}, jm, EXIT_VALIDATION) summary = (args.summary_tokens if args.summary_tokens is not None else args.history_tokens * 0.10) if summary <= 0: fail("VALIDATION", "--summary-tokens must be positive", {"got": summary}, jm, EXIT_VALIDATION) if summary >= args.history_tokens: fail("VALIDATION", "--summary-tokens must be smaller than the history " "(a 'summary' that isn't smaller saves nothing)", {"summary": summary, "history": args.history_tokens}, jm, EXIT_VALIDATION) write_mult = CACHE_WRITE_1H if args.ttl == "1h" else CACHE_WRITE_5M turns = args.turns_remaining note(TERM.hdr("append vs compact"), quiet) # --- Constraint A: does it even fit? Checked first; cost is moot if not. projected = args.history_tokens + args.growth_per_turn * turns overflows = projected > args.context_window # --- Path 1: append. Carry the growing history at cache-read price. cost_append = cached_carry_cost(args.history_tokens, turns, args.growth_per_turn, rate) # --- Path 2: compact. Summarise now, then carry the summary. # a) one summarisation call: history read from cache, summary generated per_summarise = (args.history_tokens * CACHE_READ_MULTIPLIER * rate / PER_MTOK + summary * out_rate / PER_MTOK) # b) writing the new (shorter) prefix into the cache, once per compaction per_rewrite = summary * write_mult * rate / PER_MTOK # c) HOW MANY compactions. With growth, the summary climbs back to the # original size after (history - summary)/growth turns and a real # system compacts again. Charging a single one-shot rewrite flatters # compaction badly on long agentic sessions (~2.2x on 60 turns at # 4k/turn), which is the regime where people actually reach for it. if args.growth_per_turn > 0: regrow_turns = (args.history_tokens - summary) / args.growth_per_turn compactions = max(1, 1 + int(turns / regrow_turns)) if turns > 0 else 0 else: compactions = 1 if turns > 0 else 0 cost_summarise = per_summarise * compactions cost_rewrite = per_rewrite * compactions # d) carrying the summary for the remaining turns, growing as before cost_carry = cached_carry_cost(summary, turns, args.growth_per_turn, rate) cost_compact = cost_summarise + cost_rewrite + cost_carry delta = cost_append - cost_compact # >0 means compaction is cheaper # Break-even: turns needed to repay ONE compaction's fixed cost. With # growth the cycle repeats, so this is a per-cycle repayment period, not a # whole-session verdict -- the verdict below compares the full modelled # costs including every compaction. Per-turn saving is the token difference # carried at cache-read price. per_turn_saving = ((args.history_tokens - summary) * CACHE_READ_MULTIPLIER * rate / PER_MTOK) fixed_cost = per_summarise + per_rewrite # one compaction's fixed cost breakeven = (fixed_cost / per_turn_saving) if per_turn_saving > 0 else float("inf") if overflows: verdict, reason = "compact", "context ceiling: projected history overflows the window" elif delta > 0: verdict, reason = "compact", "cost ceiling: compaction is cheaper over the remaining turns" else: verdict, reason = "append", "keep everything: appending is cheaper over the remaining turns" note(f" projected history at turn {turns}: {projected:,.0f} tokens " f"(window {args.context_window:,.0f})", quiet) note(f" compactions modelled: {compactions}" + (" (history regrows at --growth-per-turn)" if args.growth_per_turn > 0 else " (no growth given -> one-shot)"), quiet) note(f" append ${cost_append:.4f} compact ${cost_compact:.4f}" f" (summarise ${cost_summarise:.4f} + rewrite ${cost_rewrite:.4f}" f" + carry ${cost_carry:.4f})", quiet) note(f" break-even at ~{breakeven:.1f} turns per compaction cycle " f"(you have {turns} turns, {compactions} compaction(s) modelled)", quiet) note(" note: break-even is scale-invariant - size and price cancel out; it " "tracks the summary ratio, not how big or costly the conversation is.", quiet) if verdict == "append": note(f"{TERM.mark('ok')} APPEND. {reason}.", quiet) note(" Cheaper levers before compaction: cap tool output at the boundary, " "move payloads to files, verify the cache is actually hitting.", quiet) else: note(f"{TERM.mark('warn')} COMPACT. {reason}.", quiet) note(" Still try the cache-preserving levers first (tool-output caps, " "payloads to files) - they do not rewrite the prefix.", quiet) note(" RECALL IS NOT PRICED HERE. Cheaper is not the same as better: " "summarisation has measured 92-100% -> 38-58% on planted-fact " "recall. Probe yours (assets/recall-probe.py) before acting.", quiet) data = { "verdict": verdict, "reason": reason, "context_ceiling_hit": overflows, "inputs": { "history_tokens": args.history_tokens, "turns_remaining": turns, "base_rate_per_mtok": rate, "output_rate_per_mtok": out_rate, "summary_tokens": summary, "growth_per_turn": args.growth_per_turn, "ttl": args.ttl, "context_window": args.context_window, }, "multipliers": { "cache_read": CACHE_READ_MULTIPLIER, "cache_write": write_mult, }, "cost_usd": { "append": round(cost_append, 6), "compact": round(cost_compact, 6), "compact_summarise_call": round(cost_summarise, 6), "compact_cache_rewrite": round(cost_rewrite, 6), "compact_carry": round(cost_carry, 6), "delta_append_minus_compact": round(delta, 6), }, "breakeven_turns_per_compaction": ( round(breakeven, 2) if breakeven != float("inf") else None), "compactions_modelled": compactions, "projected_history_tokens": projected, "caveats": [ "models cost only - recall loss from summarisation is not priced", "break-even turns is scale-invariant: driven by the summary ratio, " "the output/input price ratio and the write multiplier - not by " "history size or price level", "break-even is the repayment period for ONE compaction; the verdict " "compares full modelled costs across all modelled compactions", ], } if jm: print(json.dumps({"data": data, "meta": {"schema": SCHEMA, "status": verdict}})) else: print(f"{verdict}\tappend=${cost_append:.4f}\tcompact=${cost_compact:.4f}" f"\tbreakeven_turns_per_compaction={breakeven:.1f}" f" compactions={compactions}") return EXIT_COMPACT if verdict == "compact" else EXIT_OK if __name__ == "__main__": try: sys.exit(main(sys.argv[1:])) except KeyboardInterrupt: sys.exit(EXIT_USAGE)
-
-
tests
-
run.sh 23.6 KB
#!/usr/bin/env bash # Self-test for claude-api-ops — fully offline: no network, no Anthropic API. # # Wraps the skill's §7 staleness verifier (scripts/check-model-table.py), which # guards the two fast-moving fact tables (SKILL.md "Current Models" and # references/caching-and-cost.md cache-minimums) against silent drift. Contract # (py_compile + --help), offline happy path against the shipped skill, the # --json §7 envelope, and a NEGATIVE proving the verifier actually rejects a bad # model id (a date-suffixed alias — exactly what SKILL.md forbids). --live is # NEVER invoked: it hits the Models API, and a network blip must never fail a PR. # # Usage: bash tests/run.sh # Exit: 0 all pass, 1 one or more failures set -uo pipefail HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SKILL="$(dirname "$HERE")" V="$SKILL/scripts/check-model-table.py" # Pick a python that actually executes (the Windows Store python3 stub exists # on PATH but exits non-zero; probe by running it). PYTHON="" for c in python python3 py; do if command -v "$c" >/dev/null 2>&1 && "$c" -c "" >/dev/null 2>&1; then PYTHON="$c"; break; fi done SB="$(mktemp -d)"; trap 'rm -rf "$SB"' EXIT PASS=0; FAIL=0 ok() { PASS=$((PASS+1)); printf ' PASS %s\n' "$1"; } no() { FAIL=$((FAIL+1)); printf ' FAIL %s\n' "$1"; } expect_exit() { [[ "$2" == "$3" ]] && ok "$1 (exit $3)" || no "$1 (want $2 got $3)"; } expect_has() { case "$3" in *"$2"*) ok "$1";; *) no "$1 (missing '$2')";; esac; } echo "=== claude-api-ops self-test ===" if [[ -z "$PYTHON" ]]; then echo " SKIP no working python (verifier is python) — cannot test" [[ "$FAIL" -eq 0 ]] || exit 1 exit 0 fi # ── contract ────────────────────────────────────────────────────────────────── echo "-- contract --" "$PYTHON" -m py_compile "$V" 2>/dev/null && ok "py_compile check-model-table.py" || no "py_compile check-model-table.py" "$PYTHON" "$V" --help >/dev/null 2>&1; expect_exit "--help exits 0" 0 $? out="$("$PYTHON" "$V" --help 2>&1)" expect_has "--help has EXAMPLES" "EXAMPLES" "$out" "$PYTHON" "$V" --bogus >/dev/null 2>&1; expect_exit "unknown flag -> 2" 2 $? # ── offline structural mode (§7 seam: --offline default, --live advisory) ───── echo "-- offline structural --" "$PYTHON" "$V" --offline >/dev/null 2>&1; expect_exit "--offline clean on shipped skill" 0 $? out="$("$PYTHON" "$V" --offline --json 2>/dev/null)" expect_has "--offline --json envelope schema" '"schema": "claude-mods.claude-api-ops.model-table/v1"' "$out" expect_has "--offline --json consistent" '"consistent": true' "$out" # ── negative: a date-suffixed alias must be rejected (exit 4 VALIDATION) ────── # Run the verifier from a doctored copy; never mutate the shipped skill. echo "-- negative --" cp -r "$SKILL" "$SB/copy" # Append the date suffix SKILL.md explicitly forbids ("Never append date # suffixes"). The verifier's DATE_SUFFIX_RE must flag it as VALIDATION drift. # Target the table cell uniquely (the prose never pairs "Opus 5 |" with the # backticked id) so the edit is surgical. If the lineup changes and this string # stops matching, the copy stays clean and the exit-4 assertion below fails loudly # rather than silently passing on an unmodified file. "$PYTHON" - "$SB/copy/SKILL.md" <<'PY' import pathlib, sys p = pathlib.Path(sys.argv[1]) t = p.read_text(encoding="utf-8") src = "Opus 5 | `claude-opus-5`" assert src in t, "negative-test fixture no longer matches SKILL.md model table" t = t.replace(src, "Opus 5 | `claude-opus-5-20260724`") p.write_text(t, encoding="utf-8") PY "$PYTHON" "$SB/copy/scripts/check-model-table.py" --offline >"$SB/neg.out" 2>&1 expect_exit "--offline flags date-suffixed id -> 4" 4 $? expect_has "finding names the date suffix" "date suffix" "$(cat "$SB/neg.out")" # ── context-engineering layer (offline cross-file tripwires) ───────────────── # The verifier also guards the cache-economics constants that are now stated in # more than one file, the verification date stamps on the doctrine references, # and SKILL.md <-> references/ citation integrity. Each gets a NEGATIVE below: # a passing check proves nothing unless it can be made to fail. echo "-- context-engineering checks --" out="$("$PYTHON" "$V" --offline --json -q 2>/dev/null)" expect_has "--json reports cache_constants" '"cache_constants"' "$out" expect_has "--json reports date_stamps" '"date_stamps"' "$out" expect_has "--json reports reference_files" '"reference_files"' "$out" expect_has "--json names the context-management beta" 'context-management-2025-06-27' "$out" # Fresh sandbox copy per negative: validate_offline runs the model-table checks # BEFORE these, so a copy poisoned by an earlier negative would exit 4 on the # wrong finding and the assertion would pass for the wrong reason. n=0 fresh_copy() { n=$((n+1)); rm -rf "$SB/c$n"; cp -r "$SKILL" "$SB/c$n"; echo "$SB/c$n"; } # NEGATIVE 1: desync a cache constant in exactly one file (0.1x -> 0.15x in # compaction.md). The real drift this guards: someone updates a multiplier in # the doc they happen to be editing and leaves the other three stale. C="$(fresh_copy)" "$PYTHON" - "$C/references/compaction.md" <<'PY' import pathlib, re, sys p = pathlib.Path(sys.argv[1]); t = p.read_text(encoding="utf-8") # Cover every spelling the verifier's pattern accepts: "0.1x", "0.1×", "0.1 ×". p.write_text(re.sub(r"0\.1(\s*)([x×])", r"0.15\1\2", t), encoding="utf-8") PY "$PYTHON" "$C/scripts/check-model-table.py" --offline >"$SB/n1.out" 2>&1 expect_exit "desynced cache multiplier -> 4" 4 $? expect_has "finding names the constant" "cache_read_multiplier" "$(cat "$SB/n1.out")" # NEGATIVE 2: strip the verification date stamp from a doctrine reference. C="$(fresh_copy)" "$PYTHON" - "$C/references/context-engineering.md" <<'PY' import pathlib, re, sys p = pathlib.Path(sys.argv[1]); t = p.read_text(encoding="utf-8") p.write_text(re.sub(r"(?i)verified", "checked", t), encoding="utf-8") PY "$PYTHON" "$C/scripts/check-model-table.py" --offline >"$SB/n2.out" 2>&1 expect_exit "missing verification date stamp -> 4" 4 $? expect_has "finding names the undated file" "context-engineering.md" "$(cat "$SB/n2.out")" # NEGATIVE 3: an uncited reference file (dead weight the router never finds -- # SKILL-RESOURCE-PROTOCOL.md §1). C="$(fresh_copy)" printf '# orphan\n' > "$C/references/orphan-doc.md" "$PYTHON" "$C/scripts/check-model-table.py" --offline >"$SB/n3.out" 2>&1 expect_exit "uncited reference file -> 4" 4 $? expect_has "finding names the orphan" "orphan-doc.md" "$(cat "$SB/n3.out")" # NEGATIVE 4: a cited-but-missing reference (broken link in SKILL.md). C="$(fresh_copy)" rm -f "$C/references/compaction.md" "$PYTHON" "$C/scripts/check-model-table.py" --offline >"$SB/n4.out" 2>&1 expect_exit "cited-but-missing reference -> 4" 4 $? # NEGATIVE 5: an uncited script/asset (same protocol rule as references, but # those are cited by basename in prose rather than by markdown link). C="$(fresh_copy)" printf '# orphan\n' > "$C/assets/orphan-asset.py" "$PYTHON" "$C/scripts/check-model-table.py" --offline >"$SB/n5.out" 2>&1 expect_exit "uncited asset -> 4" 4 $? expect_has "finding names the uncited asset" "orphan-asset.py" "$(cat "$SB/n5.out")" # ── context-budget calculator ──────────────────────────────────────────────── # The doctrine's break-even arithmetic, executable. Its VERDICT is the contract: # exit 0 = append wins, exit 10 = cost favours compaction. Both directions are # asserted, because a calculator that can only say one thing is not a calculator. echo "-- context-budget --" CB="$SKILL/scripts/context-budget.py" "$PYTHON" -m py_compile "$CB" 2>/dev/null && ok "py_compile context-budget.py" || no "py_compile context-budget.py" "$PYTHON" "$CB" --help >/dev/null 2>&1; expect_exit "context-budget --help exits 0" 0 $? expect_has "context-budget --help has EXAMPLES" "EXAMPLES" "$("$PYTHON" "$CB" --help 2>&1)" "$PYTHON" "$CB" --bogus >/dev/null 2>&1; expect_exit "context-budget unknown flag -> 2" 2 $? # Short session: the doctrine's default answer. Must be exit 0 (append). "$PYTHON" "$CB" --history-tokens 25000 --turns-remaining 5 --base-rate 0.30 -q >/dev/null 2>&1 expect_exit "short session -> append (0)" 0 $? # Deep session: enough remaining turns to repay the rewrite. Must be exit 10. "$PYTHON" "$CB" --history-tokens 120000 --turns-remaining 40 --base-rate 2.00 -q >/dev/null 2>&1 expect_exit "deep session -> compaction indicated (10)" 10 $? # Context ceiling binds regardless of cost. "$PYTHON" "$CB" --history-tokens 900000 --turns-remaining 10 --growth-per-turn 50000 \ --context-window 1000000 -q >/dev/null 2>&1 expect_exit "context ceiling -> 10" 10 $? out="$("$PYTHON" "$CB" --history-tokens 25000 --turns-remaining 5 --base-rate 0.30 --json -q 2>/dev/null)" expect_has "context-budget --json envelope schema" '"schema": "claude-mods.claude-api-ops.context-budget/v1"' "$out" expect_has "context-budget --json verdict" '"verdict": "append"' "$out" # The recall caveat must ride along in the machine-readable output: an agent # acting on the verdict alone would otherwise treat "cheaper" as "better". expect_has "context-budget --json carries the recall caveat" 'recall loss' "$out" # REGRESSION: the calculator originally charged ONE compaction. With growth the # summary regrows and a real system compacts again -- assuming one-shot # understated compaction's fixed cost ~2.2x on a 60-turn, 4k/turn session, i.e. # it was biased TOWARD compaction, the opposite of the doctrine's default. out="$("$PYTHON" "$CB" --history-tokens 120000 --turns-remaining 60 --base-rate 2.00 --growth-per-turn 4000 --json -q 2>/dev/null)" expect_has "growing session models >1 compaction" '"compactions_modelled": 3' "$out" out="$("$PYTHON" "$CB" --history-tokens 120000 --turns-remaining 60 --base-rate 2.00 --json -q 2>/dev/null)" expect_has "no growth -> one-shot compaction" '"compactions_modelled": 1' "$out" # Break-even is a per-compaction repayment period, not a session verdict; the # key name has to say so or readers compare it against total turns and misread. expect_has "break-even key is scoped per compaction" '"breakeven_turns_per_compaction"' "$out" # Input validation (resource protocol §6 - agents fabricate plausible inputs). for bad in "--history-tokens -5 --turns-remaining 10" \ "--history-tokens 1000 --turns-remaining -1" \ "--history-tokens 1000 --turns-remaining 5 --base-rate 0" \ "--history-tokens 1000 --turns-remaining 5 --summary-tokens 5000"; do "$PYTHON" "$CB" $bad >/dev/null 2>&1 expect_exit "rejects bad input ($bad)" 4 $? done # ── cache-correct loop asset ───────────────────────────────────────────────── # The asset's only real logic is breakpoint placement, and getting it wrong is # silent (a missed cache costs money and raises no error). Exercise it directly # with the anthropic SDK stubbed out - no network, no SDK install needed. echo "-- cached-agent-loop --" "$PYTHON" -m py_compile "$SKILL/assets/cached-agent-loop.py" 2>/dev/null \ && ok "py_compile cached-agent-loop.py" || no "py_compile cached-agent-loop.py" "$PYTHON" -m py_compile "$SKILL/assets/recall-probe.py" 2>/dev/null \ && ok "py_compile recall-probe.py" || no "py_compile recall-probe.py" "$PYTHON" - "$SKILL/assets/cached-agent-loop.py" >"$SB/bp.out" 2>&1 <<'PY' import sys, types, importlib.util stub = types.ModuleType("anthropic"); stub.Anthropic = lambda *a, **k: None sys.modules["anthropic"] = stub spec = importlib.util.spec_from_file_location("loop", sys.argv[1]) m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m) def marks(msgs): return [(i, j) for i, msg in enumerate(msgs) for j, b in enumerate(msg["content"]) if isinstance(b, dict) and "cache_control" in b] # The newest turn must always carry a breakpoint, or hits never accrue. msgs = [{"role": "user", "content": [{"type": "text", "text": "hi"}]}] m.place_message_breakpoints(msgs) assert marks(msgs) == [(0, 0)], f"short: {marks(msgs)}" # A tool-heavy turn appending 40 blocks must not exceed the API's 4-breakpoint # limit (one is spent on the system block) and must keep consecutive anchors # inside the 20-block backward search, or the lookback silently misses. msgs = [{"role": "user", "content": [{"type": "text", "text": f"b{i}"} for i in range(40)]}] m.place_message_breakpoints(msgs) got = marks(msgs) assert len(got) <= m.MAX_BREAKPOINTS - 1, f"too many breakpoints: {got}" assert got[-1] == (0, 39), f"newest block unmarked: {got}" gaps = [got[i + 1][1] - got[i][1] for i in range(len(got) - 1)] assert all(g <= m.LOOKBACK_BLOCKS for g in gaps), f"anchor gap exceeds lookback: {gaps}" assert m.BREAKPOINT_EVERY < m.LOOKBACK_BLOCKS, "anchor spacing must fit the window" # Idempotent: called every turn, it must not accumulate stale markers. before = marks(msgs); m.place_message_breakpoints(msgs) assert marks(msgs) == before, "not idempotent" # Tool output is capped at the boundary (the cache-preserving lever). assert len(m.capped("x" * 99999)) < 99999, "capped() did not truncate" assert m.capped("short") == "short", "capped() mangled a short result" # === REGRESSION: the shipped loop once placed ZERO breakpoints =============== # In a real loop, assistant turns are `response.content` -- SDK block OBJECTS, # not dicts -- and cache_control can only be set on a dict. The first version # of this asset skipped non-dict blocks silently, so every anchor landed on an # assistant block and was dropped: no error, no warning, no caching, full price # forever. The original test missed it because it only ever fed dicts. These # cases feed the shapes a real loop actually produces. class SDKBlock: # stands in for anthropic.types.TextBlock def __init__(self, t): self.type = "text"; self.text = t def model_dump(self, exclude_none=False): return {"type": "text", "text": self.text} def mixed(n_turns): msgs = [] for k in range(n_turns): msgs.append({"role": "user", "content": [{"type": "text", "text": f"u{k}-{j}"} for j in range(4)]}) msgs.append({"role": "assistant", "content": [SDKBlock(f"a{k}-{j}") for j in range(4)]}) return msgs # Un-normalised SDK objects: must still place markers (walks back to a dict). msgs = mixed(6) placed = m.place_message_breakpoints(msgs) assert placed > 0, "REGRESSION: zero breakpoints placed on an SDK-object conversation" assert placed <= m.MAX_BREAKPOINTS - 1, f"too many breakpoints: {placed}" # Normalised via to_blocks() -- the path main() takes. The newest block must # carry a marker, and anchors must stay inside the 20-block lookback. msgs = [{"role": m0["role"], "content": m.to_blocks(m0["content"])} for m0 in mixed(6)] placed = m.place_message_breakpoints(msgs) got = marks(msgs) assert placed == len(got), f"reported {placed} but set {len(got)}" assert got[-1] == (len(msgs) - 1, len(msgs[-1]["content"]) - 1), f"newest block unmarked: {got}" flat_pos = [] for mi, msg in enumerate(msgs): for bi in range(len(msg["content"])): flat_pos.append((mi, bi)) idxs = [flat_pos.index(g) for g in got] gaps = [idxs[i + 1] - idxs[i] for i in range(len(idxs) - 1)] assert all(g <= m.LOOKBACK_BLOCKS for g in gaps), f"anchor gap exceeds lookback: {gaps}" # String-shorthand content has no block to attach a marker to: warn, don't crash. assert m.place_message_breakpoints([{"role": "user", "content": "plain string"}]) == 0 # to_blocks() normalises all three shapes a caller may hand it. assert m.to_blocks("hi") == [{"type": "text", "text": "hi"}] assert m.to_blocks([{"type": "text", "text": "hi"}]) == [{"type": "text", "text": "hi"}] assert m.to_blocks([SDKBlock("hi")]) == [{"type": "text", "text": "hi"}] print("OK") PY expect_has "breakpoint placement, capping and idempotence" "OK" "$(cat "$SB/bp.out")" # ── negative: a retired model id in a code sample (exit 4 VALIDATION) ───────── # The regression this check exists for: on 2026-08-30 the tables were retargeted # at the Claude 5 lineup and went green, then a sibling branch landed an asset # still pinned to claude-opus-4-8. Tables were consistent; the sample was wrong. echo "-- negative: retired id in a sample --" rm -rf "$SB/copy2"; cp -r "$SKILL" "$SB/copy2" "$PYTHON" - "$SB/copy2" <<'PY' import pathlib, sys p = pathlib.Path(sys.argv[1]) / "assets/cached-agent-loop.py" t = p.read_text(encoding="utf-8") assert 'MODEL = "claude-opus-5"' in t, "fixture no longer matches cached-agent-loop.py" p.write_text(t.replace('MODEL = "claude-opus-5"', 'MODEL = "claude-opus-4-8"'), encoding="utf-8") PY "$PYTHON" "$SB/copy2/scripts/check-model-table.py" --offline >"$SB/neg2.out" 2>&1 expect_exit "retired id in a sample -> 4" 4 $? expect_has "finding names the file and fault" "cached-agent-loop.py" "$(cat "$SB/neg2.out")" expect_has "finding classifies it retired" "retired" "$(cat "$SB/neg2.out")" # The escape hatch must actually release the gate, or authors will delete it. "$PYTHON" - "$SB/copy2" <<'PY' import pathlib, sys p = pathlib.Path(sys.argv[1]) / "assets/cached-agent-loop.py" p.write_text(p.read_text(encoding="utf-8").replace( 'MODEL = "claude-opus-4-8"', 'MODEL = "claude-opus-4-8" # legacy-ok'), encoding="utf-8") PY "$PYTHON" "$SB/copy2/scripts/check-model-table.py" --offline >/dev/null 2>&1 expect_exit "legacy-ok marker releases the gate" 0 $? # An id in neither the table nor the legacy list is a typo or a hallucination. rm -rf "$SB/copy3"; cp -r "$SKILL" "$SB/copy3" printf '\nUse claude-opus-7 for this.\n' >> "$SB/copy3/references/tool-use.md" "$PYTHON" "$SB/copy3/scripts/check-model-table.py" --offline >"$SB/neg3.out" 2>&1 expect_exit "unknown model id -> 4" 4 $? expect_has "finding classifies it unknown" "unknown" "$(cat "$SB/neg3.out")" # Runtime warnings must be ASCII: this asset prints to a console that is cp1252 # by default on Windows, where an em-dash renders as a replacement character. "$PYTHON" - "$SKILL/assets/cached-agent-loop.py" "$SKILL/assets/recall-probe.py" >"$SB/ascii.out" 2>&1 <<'PY' import ast, pathlib, sys bad = [] for f in sys.argv[1:]: tree = ast.parse(pathlib.Path(f).read_text(encoding="utf-8")) for node in ast.walk(tree): if isinstance(node, ast.Call) and getattr(node.func, "id", "") == "print": for a in ast.walk(node): if isinstance(a, ast.Constant) and isinstance(a.value, str) \ and any(ord(c) > 127 for c in a.value): bad.append((pathlib.Path(f).name, a.value[:40])) print("CLEAN" if not bad else f"NON-ASCII IN PRINT: {bad}") PY expect_has "asset runtime output is ASCII-safe" "CLEAN" "$(cat "$SB/ascii.out")" # ── recall-probe fairness ──────────────────────────────────────────────────── # REGRESSION: the harness originally compacted on EVERY turn once the history # exceeded keep_recent, so the compact arm paid a summarisation call per turn # (21 API calls vs append's 12 over 10 turns). That inflates the arm under test # and would "confirm" the keep-everything result whatever the data said. The # cadence knob is the fix; assert it exists and is actually consulted. echo "-- recall-probe fairness --" RP="$SKILL/assets/recall-probe.py" grep -q 'compact_every' "$RP" && ok "recall-probe has a compaction cadence" \ || no "recall-probe has a compaction cadence" grep -q 'turns_since_compaction >= compact_every' "$RP" \ && ok "cadence actually gates compaction" || no "cadence actually gates compaction" grep -q '"--compact-every"' "$RP" && ok "cadence is user-tunable" || no "cadence is user-tunable" # Default must not be 1 -- that is the rigged configuration. "$PYTHON" - "$RP" >"$SB/rp.out" 2>&1 <<'PY' import ast, pathlib, sys tree = ast.parse(pathlib.Path(sys.argv[1]).read_text(encoding="utf-8")) for node in ast.walk(tree): if isinstance(node, ast.Call) and getattr(node.func, "attr", "") == "add_argument": if node.args and getattr(node.args[0], "value", "") == "--compact-every": d = [k.value.value for k in node.keywords if k.arg == "default"] print("DEFAULT", d[0] if d else "none") PY out="$(cat "$SB/rp.out")" case "$out" in "DEFAULT 1"|"DEFAULT none"|"") no "compact cadence default is unbiased (got '$out')" ;; *) ok "compact cadence default is unbiased ($out)" ;; esac # ── SKILL.md sanity ─────────────────────────────────────────────────────────── echo "-- SKILL.md --" # CONTRACT (frontmatter shape): this suite asserts that SKILL.md's frontmatter # keeps `name: claude-api-ops` and a `when_to_use:` field, and that # len(description) + len(when_to_use) stays within the repo's 1000-char per-skill # cap enforced by tests/validate.sh. A description-trim or frontmatter cleanup # lane that removes `when_to_use` from this skill WILL break CI here -- that is # deliberate, and stated here so the edit site is not the first place you find out. grep -q '^name: claude-api-ops$' "$SKILL/SKILL.md" && ok "frontmatter name" || no "frontmatter name" grep -q '^when_to_use: ' "$SKILL/SKILL.md" && ok "frontmatter when_to_use present" || no "frontmatter when_to_use present" grep -q 'check-model-table.py' "$SKILL/SKILL.md" && ok "verifier cited from SKILL.md" || no "verifier cited from SKILL.md" # Description budget: mirrors tests/validate.sh's hard cap so this skill fails # in its own suite rather than only in the catalog-wide gate. combined="$("$PYTHON" - "$SKILL/SKILL.md" <<'PY' import pathlib, re, sys t = pathlib.Path(sys.argv[1]).read_text(encoding="utf-8") fm = t.split("---")[1] def field(k): m = re.search(r'^%s: "(.*)"$' % k, fm, re.M) return m.group(1) if m else "" print(len(field("description")) + len(field("when_to_use"))) PY )" if [[ "$combined" -gt 0 && "$combined" -le 1000 ]]; then ok "description + when_to_use within 1000-char cap ($combined)" else no "description + when_to_use out of range (got '$combined', cap 1000)" fi # Body-size limit: SKILL-CREATION-PROTOCOL.md Step 3 caps the body at 500 lines; # depth belongs in references/*.md. lines="$(wc -l < "$SKILL/SKILL.md" | tr -d ' ')" [[ "$lines" -lt 500 ]] && ok "SKILL.md body under 500 lines ($lines)" \ || no "SKILL.md body is $lines lines (limit 500)" # The context-engineering content itself is cited and reachable. echo "-- context-engineering content --" grep -q '^## Context Engineering$' "$SKILL/SKILL.md" \ && ok "Context Engineering section present" || no "Context Engineering section present" for r in context-engineering compaction; do [[ -f "$SKILL/references/$r.md" ]] && ok "references/$r.md exists" || no "references/$r.md exists" grep -q "(references/$r.md)" "$SKILL/SKILL.md" \ && ok "references/$r.md cited from SKILL.md" || no "references/$r.md cited from SKILL.md" done # The load-bearing, counter-intuitive claim must survive edits: compaction is a # response to a NAMED constraint, not a default. grep -qi 'named constraint' "$SKILL/references/compaction.md" \ && ok "compaction doctrine states the named-constraint rule" \ || no "compaction doctrine states the named-constraint rule" grep -q 'clear_at_least' "$SKILL/references/compaction.md" \ && ok "context_management params documented" || no "context_management params documented" # The three tiers are the spine of the doctrine reference. grep -q 'Tier 1' "$SKILL/references/context-engineering.md" \ && ok "three-tier model documented" || no "three-tier model documented" echo "" echo "=== $PASS passed, $FAIL failed ===" [[ "$FAIL" -eq 0 ]] || exit 1 exit 0
-
-
SKILL.md 26.4 KB
--- name: claude-api-ops description: "Building applications ON Claude - the Anthropic API and Claude Agent SDK. Use for: anthropic api, claude api, messages api, tool use, function calling, prompt caching, agent sdk, claude-agent-sdk, structured output, json schema output, batches api, extended thinking, adaptive thinking, model selection, claude pricing, build claude agent, anthropic sdk, stop_reason handling, streaming claude, token counting, cache_control, output_config, tool_choice, agentic loop, rate limits anthropic, context engineering, context window budget, compaction, context editing, context_management, clear_tool_uses, memory tool, context rot, tool result bloat, subagent context isolation." when_to_use: "Use when building applications on the Anthropic API or Claude Agent SDK — e.g. 'add tool use to my Claude app', 'set up prompt caching', 'which Claude model should I use', 'handle stop_reason / streaming', 'should I compact this agent context'." license: MIT allowed-tools: "Read Write Bash WebFetch" metadata: author: claude-mods related-skills: mcp-ops --- # Claude API Operations Building applications and agents on Anthropic's API: the Messages API, tool use, prompt caching, structured outputs, batches, thinking/effort, and the Claude Agent SDK. For developers writing apps *against* the API — not for using Claude Code itself. **API surfaces move fast.** Model IDs, parameters, and betas in this skill were verified against platform.claude.com (2026-08). When in doubt — especially for "latest model" or pricing questions — verify with WebFetch against `https://platform.claude.com/docs/en/about-claude/models/overview.md` or query the Models API (`client.models.list()`). ## Current Models (verified 2026-08) | Model | ID (exact, no date suffix) | Context | Max Output | Input $/MTok | Output $/MTok | |---|---|---|---|---|---| | Claude Fable 5 | `claude-fable-5` | 1M | 128K | $10.00 | $50.00 | | Claude Opus 5 | `claude-opus-5` | 1M | 128K | $5.00 | $25.00 | | Claude Sonnet 5 | `claude-sonnet-5` | 1M | 128K | $2.00 | $10.00 | | Claude Haiku 4.5 | `claude-haiku-4-5` | 200K | 64K | $1.00 | $5.00 | Use these alias IDs verbatim. **Never append date suffixes** (`claude-sonnet-5-20260630` is wrong → 404). Haiku 4.5 is the one current model with a *dated* snapshot id (`claude-haiku-4-5-20251001`) behind its alias; from the 4.6 generation on, the dateless id **is** the pinned snapshot. **Legacy (still available, no longer current):** `claude-opus-4-8`, `claude-opus-4-7`, `claude-opus-4-6`, `claude-opus-4-5`, `claude-sonnet-4-6`, `claude-sonnet-4-5`. Migrating off one: `https://platform.claude.com/docs/en/models/opus-5/migration-guide.md` (or run `/claude-api migrate` in Claude Code). Live capability lookup: `client.models.retrieve("claude-opus-5")` → `.max_input_tokens`, `.max_tokens`, `.capabilities` dict. ## Model Selection Decision Tree ``` What is the workload? │ ├─ Hardest problems, long-horizon agents, deep research, ceiling intelligence │ └─ claude-fable-5 (premium ceiling) or claude-opus-5 (default flagship) │ ├─ Agentic coding, tool-heavy workflows, production assistants │ └─ claude-opus-5 (quality) or claude-sonnet-5 (speed/cost balance) │ ├─ High-volume production: summarization, RAG answers, extraction │ └─ claude-sonnet-5 │ ├─ Classification, routing, simple Q&A, latency-critical │ └─ claude-haiku-4-5 │ └─ Subagents inside a larger system └─ One tier below the orchestrator (Opus loop → Sonnet/Haiku workers) ``` Tiering rule: route by task difficulty, not by uniform default. An Opus orchestrator dispatching Haiku classifiers is routinely 5-10x cheaper than Opus-everywhere with no quality loss on the simple legs. ## Which Surface? (API vs Agent SDK vs Batches) | Need | Use | Why | |---|---|---| | One request → one response (classify, summarize, extract, Q&A) | **Messages API** | Simplest; full control | | Multi-step pipeline, your code controls the logic | **Messages API + tool use** | You own the loop | | Custom agent with your own tools, your infra | **Messages API + tool use** (manual loop or SDK tool runner) | Max flexibility | | Agent that reads/edits files, runs commands, searches — without building tools | **Claude Agent SDK** | Claude Code's tools + agent loop as a library | | CI/CD automation, coding agents, production agent apps | **Claude Agent SDK** | Built-in tools, hooks, sessions, MCP | | Large non-urgent workloads (eval runs, backfills, bulk extraction) | **Batches API** | 50% discount, ≤24h turnaround | | Hosted agent, Anthropic runs loop + sandbox | **Managed Agents** (beta) | No infra; see official docs | Rule of thumb: start at the simplest tier. Reach for an agent only when the task is genuinely open-ended (multi-step, hard to fully specify, errors recoverable, value justifies cost). ## Messages API Quick Start Everything goes through `POST /v1/messages`. Headers: `x-api-key`, `anthropic-version: 2023-06-01`, `content-type: application/json`. ```python # pip install anthropic import anthropic client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY response = client.messages.create( model="claude-opus-5", max_tokens=16000, system="You are a concise technical assistant.", messages=[{"role": "user", "content": "Explain CRDTs in one paragraph."}], ) for block in response.content: # content is a list of typed blocks if block.type == "text": # always check .type before .text print(block.text) print(response.stop_reason, response.usage.input_tokens, response.usage.output_tokens) ``` ```typescript // npm install @anthropic-ai/sdk import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic(); const response = await client.messages.create({ model: "claude-opus-5", max_tokens: 16000, messages: [{ role: "user", content: "Explain CRDTs in one paragraph." }], }); for (const block of response.content) { if (block.type === "text") console.log(block.text); // narrow the union first } ``` Streaming (default to it for long outputs — non-streaming above ~16K `max_tokens` risks SDK HTTP timeouts): ```python with client.messages.stream(model="claude-opus-5", max_tokens=64000, messages=[{"role": "user", "content": "Write a long report"}]) as stream: for text in stream.text_stream: print(text, end="", flush=True) final = stream.get_final_message() # full Message after streaming ``` Full params, response shape, stop reasons, errors, retries, rate limits: [references/messages-api.md](references/messages-api.md) ## Thinking & Effort (quick reference) - **Adaptive thinking is ON BY DEFAULT on Fable 5 / Opus 5 / Sonnet 5** — send no `thinking` field and you still get (and pay for) thinking. On the legacy 4.6–4.8 models it stays off until you set `thinking: {"type": "adaptive"}`. - **Manual budgets are gone.** `{"type": "enabled", "budget_tokens": N}` returns a **400 on Opus 4.7 and every later model** (Opus 5, Sonnet 5, Fable 5 included); deprecated on Opus 4.6 / Sonnet 4.6. Control depth with `effort`, not tokens. - **Turning thinking off:** Sonnet 5 accepts `{"type": "disabled"}`. Opus 5 accepts it only at effort `high` or below — pairing it with `xhigh`/`max` is a **400**. Fable 5 **rejects it outright**; thinking there is unconditional, so budget for it. - **Effort (GA):** `output_config: {"effort": "low" | "medium" | "high" | "xhigh" | "max"}` — nested in `output_config`, not top-level. Default `high` (identical to omitting it). `xhigh`: Fable 5, Opus 5, Opus 4.8/4.7, **Sonnet 5**. `max`: those plus Opus 4.6 and Sonnet 4.6. Haiku 4.5 does not support `effort` at all. - **Sampling params removed on Opus 4.7 and later** (so Opus 5, Sonnet 5, Fable 5): `temperature`, `top_p`, `top_k` all return 400 — and the Python SDK v1.0+ doesn't define them, so passing them raises `TypeError`. Steer with prompting + effort. - **Forced tool_choice is fine with adaptive thinking.** The auto/none-only restriction applies to *manual* extended thinking (`{"type": "enabled"}`) only; adaptive mode — including the models where it's on by default — accepts `{"type": "any"}` and `{"type": "tool", ...}`. - Thinking text is **omitted by default** on Fable 5 / Opus 5 / Sonnet 5 / Opus 4.8 / 4.7 — opt in with `thinking: {"type": "adaptive", "display": "summarized"}` if you surface reasoning to users. Either way the blocks are billed, and must be echoed back **unmodified** (empty `thinking` field included) in a tool-use loop, or the next request 400s. Details and gotchas: [references/structured-outputs.md](references/structured-outputs.md) (thinking interplay) and [references/messages-api.md](references/messages-api.md). ## Tool Use (quick reference) ```python tools = [{ "name": "get_weather", "description": "Get current weather. Call when the user asks about weather conditions.", "input_schema": { "type": "object", "properties": {"location": {"type": "string", "description": "City, e.g. Paris"}}, "required": ["location"], }, }] response = client.messages.create(model="claude-opus-5", max_tokens=16000, tools=tools, messages=messages) if response.stop_reason == "tool_use": ... # execute, send tool_result back, loop ``` `tool_choice`: `{"type": "auto"}` (default) | `{"type": "any"}` | `{"type": "tool", "name": "..."}` | `{"type": "none"}`. Add `"disable_parallel_tool_use": true` to force at most one call per response. The agentic loop, parallel tool results, `pause_turn`, `is_error`, server-side tools, and SDK tool runners: [references/tool-use.md](references/tool-use.md) ## Cost Optimization Checklist Work top-down; each item is independent: - [ ] **Right-size the model.** Haiku for classification/routing, Sonnet for volume work, Opus/Fable for the hard 10%. Largest single lever. - [ ] **Prompt caching** on stable prefixes (system prompt, tool defs, big docs): `cache_control: {"type": "ephemeral"}`. Reads cost ~0.1x; up to 90% savings. Verify with `usage.cache_read_input_tokens > 0` — zero means a silent invalidator (timestamp in system prompt, unsorted JSON, varying tools). - [ ] **Batches API** for anything that can wait ≤24h: flat 50% off all tokens, stacks with caching. - [ ] **Cap output**: set `max_tokens` to what you need (256 for classification); stream + generous cap for long generation. - [ ] **Tune effort down** where quality allows: `medium` is often the sweet spot; `low` for subagents and simple tasks. - [ ] **Count before sending**: `client.messages.count_tokens(...)` (never tiktoken — it's OpenAI's tokenizer and undercounts Claude by 15-20%). - [ ] **Keep prefixes stable**: order requests `tools` → `system` → `messages`, volatile content last; don't swap tool sets or models mid-conversation. Mechanics, breakpoints, TTLs, batch lifecycle, tiering math: [references/caching-and-cost.md](references/caching-and-cost.md) ## Context Engineering Prompt engineering asks what to write in the prompt. **Context engineering asks what earns a place in the window on *this* call** — including everything that lands there without you typing it: tool definitions, tool results, retrieved documents, prior turns, thinking blocks. It is iterative (every inference) where prompt engineering is discrete (written once). Target: the smallest set of high-signal tokens that gets the outcome. The budget is real because attention degrades with length (**context rot** — n² pairwise relationships), not just because tokens cost money. A 1M window is a capacity, not a target. ### The three tiers Every candidate fact lives in exactly one place. Choosing deliberately is most of the job. | Tier | Where | Cost | Use when | |---|---|---|---| | **1 — In context** | `tools` / `system` / `messages`, every call | Paid every turn (≈0.1× cached) | It steers *most* turns | | **2 — On disk, read on demand** | A file the agent can read; only the **path** stays in context | Paid only when read | The agent can tell from a *name* that it needs this | | **3 — Retrieved** | Index / search tool behind a query | Paid only on a hit, plus a relevance gamble | The corpus is too large to enumerate | When a prompt is too big, **demote before you delete** — a path is ~10 tokens; the file it names may be 10,000. **This repo already runs on the tier-1/tier-2 split.** A skill's `description` is always resident (tier 1, so it must carry the routing signal); `SKILL.md` loads on a match; `references/*.md` load only when cited and needed. "Description is the trigger", "body under 500 lines", "one concept per reference", "every reference must be cited" are context-engineering rules wearing authoring clothes. ### Cache-aware prompt architecture Requests render `tools` → `system` → `messages`, and the cache is a **prefix match**. So **static prefix first, volatile content last** — put the `cache_control` breakpoint at the end of the stable part and let per-request content fall after it. Reordering a prompt destroys the cache **silently**: no error, just a different prefix hash, `cache_read_input_tokens: 0`, and a 1.25–2× bill where you expected 0.1×. The usage block is the only symptom, which is why asserting `cache_read_input_tokens > 0` in staging is a real test. ### The compaction decision **Under modern prompt caching, keeping the full history has been measured to beat summarisation on cost, latency AND recall at the same time.** A 2026 production-tutor evaluation (660 turns, 11 configurations) put keep-everything at 92–100% fact recall, $0.11/turn and 17 s TTFT, against 38–58% recall, $0.24/turn and 21 s for its clear-plus-summarise preset. Summarising rewrites the cached prefix and forfeits the 0.1× discount — the cheap move is usually to **append**. (That study ran on a non-Claude model; what transfers is the *mechanism*, and Claude's flat 0.1× cache read makes it stronger, not weaker. Full caveats in [references/compaction.md](references/compaction.md).) So: **compact only as a deliberate response to a named constraint.** | Constraint | Diagnose | Try first | |---|---|---| | **Context ceiling** — it will not fit | Projected tokens > window | Cap tool output → payloads to files → server-side clearing | | **Cost ceiling** — the bill is unacceptable | Compare against *cached* cost, not uncached | **Verify the cache is hitting** → tier down → cap tool output | | **Latency target** — TTFT too slow at depth | Confirm growth is in the prefix | Cap tool output → lower `effort` → stream | Capping tool output at the tool boundary is the underrated lever: it shrinks context **without rewriting the cached prefix** (the same study measured −38% cost/turn with no recall loss). Clearing and summarising both break the cache; they are what people reach for first and should reach for last. First-party clearing is `context_management` (beta `context-management-2025-06-27`): `clear_tool_uses_20250919` and `clear_thinking_20251015`, applied server-side. Always set `clear_at_least` — it stops a trigger paying a full cache re-write to save a handful of tokens. Pair with the memory tool so durable conclusions are written out before raw material is cleared. ### Agentic specifics - **Tool results are the growth term**, not the system prompt. Design tools to return decisions, not dumps. - **Summarise vs write-to-file:** needed later *in full* → write to a file, return the path. Only the *conclusion* matters → summarise **at the tool boundary** (free of cache cost, unlike rewriting history after the fact). - **Sub-agents are context isolation**, not just parallelism: 80K tokens of exploration are billed once inside the child and discarded; the parent sees a ~1–2K-token distillation. Costs: cold cache in the child, a lossy hand-off. Skip it when the subtask needs most of the parent's context to make sense. Full doctrine — tiers, progressive disclosure, instrumentation: [references/context-engineering.md](references/context-engineering.md). Compaction economics, `context_management` parameters, memory tool: [references/compaction.md](references/compaction.md). For Claude Code's own context surface see the `claude-code-ops` skill; for prompts re-sent on a cadence, `loop-ops`; for cross-provider fan-out, `fleetflow`. ## Claude Agent SDK (quick reference) ```python # pip install claude-agent-sdk (Python >= 3.10) import asyncio from claude_agent_sdk import query, ClaudeAgentOptions async def main(): async for message in query( prompt="Find and fix the bug in auth.py", options=ClaudeAgentOptions(allowed_tools=["Read", "Edit", "Bash"]), ): if hasattr(message, "result"): print(message.result) asyncio.run(main()) ``` ```typescript // npm install @anthropic-ai/claude-agent-sdk import { query } from "@anthropic-ai/claude-agent-sdk"; for await (const message of query({ prompt: "Find and fix the bug in auth.ts", options: { allowedTools: ["Read", "Edit", "Bash"] }, })) { if ("result" in message) console.log(message.result); } ``` Built-in tools (Read/Write/Edit/Bash/Glob/Grep/WebSearch/WebFetch/...), hooks (`PreToolUse`, `PostToolUse`, ...), subagents, MCP servers, sessions (resume/fork), permission modes, and the SDK-vs-raw-API decision: [references/agent-sdk.md](references/agent-sdk.md) ## Common Pitfalls | Pitfall | Symptom | Fix | |---|---|---| | Date-suffixed or guessed model ID | 404 `not_found_error` | Use exact alias IDs from the table above | | `budget_tokens` on Opus 4.7+ (incl. Opus 5 / Sonnet 5 / Fable 5) | 400 | `thinking: {"type": "adaptive"}` + `effort` | | Assuming thinking is opt-in on Fable 5 / Opus 5 / Sonnet 5 | Unexpected thinking tokens billed | Adaptive thinking is on by default there; Fable 5 can't be disabled at all | | `thinking: {"type": "disabled"}` at `xhigh`/`max` on Opus 5 | 400 | Drop effort to `high` or below, or leave thinking on | | `temperature`/`top_p`/`top_k` on Opus 4.7+ | 400 (or `TypeError` on Python SDK v1.0+) | Remove; steer via prompt + `effort` | | `effort` on Haiku 4.5 | 400 | Haiku 4.5 doesn't support the parameter | | Rebuilding assistant turns in a tool loop (dropping empty `thinking` blocks) | 400 "thinking blocks cannot be modified" | Echo the content list back exactly as received | | Assistant-turn prefill on Opus 4.7+ models | 400 | `output_config.format` or system-prompt instruction | | Cache marker on <minimum prefix | Silent no-cache (`cache_creation_input_tokens: 0`) | Min 512-4096 tokens depending on model (see caching ref) | | Not handling `stop_reason: "tool_use"` | Agent "stops" after first tool call | Loop: execute tools, append `tool_result`, re-request | | Missing `tool_result` for a `tool_use` id | 400 on follow-up | One `tool_result` per `tool_use` block, ids matching | | Non-streaming with `max_tokens` > ~16K | SDK timeout / `ValueError` | Stream + `get_final_message()` / `finalMessage()` | | `output_format` top-level param | Deprecated | `output_config: {"format": {...}}` | | tiktoken for Claude token counts | 15-20%+ undercount | `messages.count_tokens` endpoint | | String-matching error messages | Fragile retries | Typed exceptions: `anthropic.RateLimitError` etc. | | Raw string-matching tool `input` | Breaks on escaping changes | Always `json.loads()` / use parsed `block.input` | | Compacting by reflex on a long conversation | Higher cost, worse recall than doing nothing | Name the constraint first; under caching, appending usually wins (see Context Engineering) | | `clear_tool_uses` without `clear_at_least` | A full cache re-write to reclaim a few hundred tokens | Set `clear_at_least` so each cache break is worth taking | ## Resources & Verification This skill ships a staleness verifier and two copy-and-adapt starter assets. The model table and pricing above are the facts most likely to drift — run the verifier when you suspect they're stale. **`scripts/check-model-table.py`** — guards the Current Models table (this file) and the per-model prompt-cache minimum table ([references/caching-and-cost.md](references/caching-and-cost.md)) against drift. Two modes per the [resource protocol §7](../../docs/SKILL-RESOURCE-PROTOCOL.md): ```bash # Structural (default, no network): every row well-formed, ids carry no date # suffix, prices numeric, the two files agree on the model lineup. It also # guards the cache-economics constants that are stated in more than one file # (0.1x read, 1.25x/2x writes, 4 breakpoints, 20-block lookback, the # context-management beta id), asserts each doctrine reference carries a # "verified <ISO date>" stamp, and checks SKILL.md <-> references/ citation # integrity in both directions. It then scans every file in the skill for model # ids: an id in neither the table nor the Legacy list is flagged "unknown", and # a LEGACY id sitting where a reader would copy it (model=..., "model": ..., # --model ...) is flagged "retired" - append a `legacy-ok` comment to that line # for a deliberate migration example. Exit 4 on any contradiction. python skills/claude-api-ops/scripts/check-model-table.py --offline python skills/claude-api-ops/scripts/check-model-table.py --offline --json | python -m json.tool # Live (advisory, needs ANTHROPIC_API_KEY): curls the Models API and compares # its id set against the documented ids. Exit 10 if a documented id is gone or a # newer alias id is missing from the table; exit 7 (not a failure) if the key is # unset or the API is unreachable. Live mode checks model-ID coverage ONLY — the # API returns no pricing, so pricing/context drift stays an --offline + docs concern. ANTHROPIC_API_KEY=sk-... python skills/claude-api-ops/scripts/check-model-table.py --live ``` **`scripts/context-budget.py`** — append-vs-compact calculator. Models both paths in dollars over the turns you actually have left, checks the context ceiling first, and exits **10** when cost favours compaction, **0** when appending wins: ```bash # Short session — appending is cheaper (exit 0) python skills/claude-api-ops/scripts/context-budget.py \ --history-tokens 25000 --turns-remaining 5 --base-rate 0.30 # Deep session — cost favours compaction (exit 10) python skills/claude-api-ops/scripts/context-budget.py \ --history-tokens 120000 --turns-remaining 40 --base-rate 2.00 --json ``` Two results worth knowing before you trust it: the break-even turn count is **scale-invariant** (history size and price cancel out — it tracks the summary ratio, not how big or costly the conversation is), and it prices **cost only**. Recall loss is not in the model, so "compact" means cheaper, not better. **`assets/cached-agent-loop.py`** — the cache-aware sibling of the minimal loop below, and the executable form of the Context Engineering section: breakpoint at the end of the static prefix, a rolling breakpoint on the newest turn, an intermediate anchor every ~15 blocks so long tool-heavy turns don't jump the 20-block lookback, tool output capped at the boundary, and a per-turn `cache_read_input_tokens` check that warns when the prefix silently changed. Copy it when the agent is long-running; copy `agentic-loop.py` when it isn't. The footgun it encodes: `cache_control` is a key on a content block, so it can only be set on a **dict**. Appending `response.content` verbatim (SDK block objects) or using the `"content": "a string"` shorthand leaves nowhere to put a marker — every breakpoint aimed at those turns is discarded with no error and no warning. Normalise content to dict blocks before placing breakpoints. **`assets/recall-probe.py`** — the "measure it on your workload" harness: plants a fact, buries it under N turns, probes for it, and reports recall, cost per turn and TTFT for **append** vs **compact**. Makes real API calls, so start small (`--turns 6 --trials 1`). Replace the synthetic filler turns with traffic from your own logs — that is the point of running it. **`assets/agentic-loop.py`** — a minimal, runnable tool-use loop (define a tool, call `messages.create`, loop while `stop_reason == "tool_use"`, append `tool_result`, re-request until `end_turn`). Copy it as the starting point when building a manual agent loop; the `>>> ADAPT` marks show what to change. **`assets/output-schema.json`** — a known-good structured-outputs request body in the canonical `output_config.format` shape (with `additionalProperties: false` and a `required` array). Copy and reshape `schema.properties` when adding JSON outputs; see [references/structured-outputs.md](references/structured-outputs.md) for the rules. (Supported on every current model — Fable 5, Opus 5, Sonnet 5, Haiku 4.5 — and the legacy 4.5–4.8 line.) ## Reference Files | File | Covers | |---|---| | [references/messages-api.md](references/messages-api.md) | Params, response shape, streaming events, stop reasons, error handling, retries, rate limits | | [references/tool-use.md](references/tool-use.md) | Tool definitions, tool_choice, parallel tools, agentic loop, tool results, server tools, tool runners | | [references/caching-and-cost.md](references/caching-and-cost.md) | Prompt caching mechanics, Batches API, token counting, model tiering economics | | [references/structured-outputs.md](references/structured-outputs.md) | output_config.format, schema rules/limits, strict tools, parse() helpers, thinking interplay | | [references/agent-sdk.md](references/agent-sdk.md) | Python + TS Agent SDK, ClaudeAgentOptions, hooks, MCP, sessions, SDK vs raw API | | [references/context-engineering.md](references/context-engineering.md) | Context budget, the three tiers, progressive disclosure, cache-aware ordering, tool-result bloat, sub-agents as isolation, instrumentation | | [references/compaction.md](references/compaction.md) | When compaction is justified, break-even arithmetic, context_management edits, memory tool, how to compact well | ## Live Documentation When cached facts may be stale, WebFetch (append `.md` for clean markdown): - Models/pricing: `https://platform.claude.com/docs/en/about-claude/models/overview.md` - Messages API: `https://platform.claude.com/docs/en/api/messages` - Tool use: `https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview.md` - Prompt caching: `https://platform.claude.com/docs/en/build-with-claude/prompt-caching.md` - Structured outputs: `https://platform.claude.com/docs/en/build-with-claude/structured-outputs.md` - Batches: `https://platform.claude.com/docs/en/build-with-claude/batch-processing.md` - Agent SDK: `https://code.claude.com/docs/en/agent-sdk/overview` - Context editing: `https://platform.claude.com/docs/en/build-with-claude/context-editing` - Context engineering: `https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents`
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.