Claude Skill

claude-api-ops

Building applications ON Claude - the Anthropic API and Claude Agent SDK. Use for: anthropic api, claude api, messages api, tool use, function calling, prompt caching, agent sdk, claude-agent-sdk, structured output, json schema output, batches api, extended thinking, adaptive thi

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download 0xdarkmatter-claude-mods-skills_claude-api-ops-3dfaf0b.zip · 83 KB
Part of 0xdarkmatter/claude-mods — 94 skills

Install

skills CLI npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/claude-api-ops
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
Git git clone https://github.com/0xDarkMatter/claude-mods.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Claude API Operations

Building applications and agents on Anthropic's API: the Messages API, tool use, prompt caching, structured outputs, batches, thinking/effort, and the Claude Agent SDK. For developers writing apps against the API — not for using Claude Code itself.

API surfaces move fast. Model IDs, parameters, and betas in this skill were verified against platform.claude.com (2026-08). When in doubt — especially for "latest model" or pricing questions — verify with WebFetch against https://platform.claude.com/docs/en/about-claude/models/overview.md or query the Models API (client.models.list()).

Current Models (verified 2026-08)

Model ID (exact, no date suffix) Context Max Output Input $/MTok Output $/MTok
Claude Fable 5 claude-fable-5 1M 128K $10.00 $50.00
Claude Opus 5 claude-opus-5 1M 128K $5.00 $25.00
Claude Sonnet 5 claude-sonnet-5 1M 128K $2.00 $10.00
Claude Haiku 4.5 claude-haiku-4-5 200K 64K $1.00 $5.00

Use these alias IDs verbatim. Never append date suffixes (claude-sonnet-5-20260630 is wrong → 404). Haiku 4.5 is the one current model with a dated snapshot id (claude-haiku-4-5-20251001) behind its alias; from the 4.6 generation on, the dateless id is the pinned snapshot.

Legacy (still available, no longer current): claude-opus-4-8, claude-opus-4-7, claude-opus-4-6, claude-opus-4-5, claude-sonnet-4-6, claude-sonnet-4-5. Migrating off one: https://platform.claude.com/docs/en/models/opus-5/migration-guide.md (or run /claude-api migrate in Claude Code). Live capability lookup: client.models.retrieve("claude-opus-5") → .max_input_tokens, .max_tokens, .capabilities dict.

Model Selection Decision Tree

What is the workload?
│
├─ Hardest problems, long-horizon agents, deep research, ceiling intelligence
│  └─ claude-fable-5 (premium ceiling) or claude-opus-5 (default flagship)
│
├─ Agentic coding, tool-heavy workflows, production assistants
│  └─ claude-opus-5 (quality) or claude-sonnet-5 (speed/cost balance)
│
├─ High-volume production: summarization, RAG answers, extraction
│  └─ claude-sonnet-5
│
├─ Classification, routing, simple Q&A, latency-critical
│  └─ claude-haiku-4-5
│
└─ Subagents inside a larger system
   └─ One tier below the orchestrator (Opus loop → Sonnet/Haiku workers)

Tiering rule: route by task difficulty, not by uniform default. An Opus orchestrator dispatching Haiku classifiers is routinely 5-10x cheaper than Opus-everywhere with no quality loss on the simple legs.

Which Surface? (API vs Agent SDK vs Batches)

Need Use Why
One request → one response (classify, summarize, extract, Q&A) Messages API Simplest; full control
Multi-step pipeline, your code controls the logic Messages API + tool use You own the loop
Custom agent with your own tools, your infra Messages API + tool use (manual loop or SDK tool runner) Max flexibility
Agent that reads/edits files, runs commands, searches — without building tools Claude Agent SDK Claude Code's tools + agent loop as a library
CI/CD automation, coding agents, production agent apps Claude Agent SDK Built-in tools, hooks, sessions, MCP
Large non-urgent workloads (eval runs, backfills, bulk extraction) Batches API 50% discount, ≤24h turnaround
Hosted agent, Anthropic runs loop + sandbox Managed Agents (beta) No infra; see official docs

Rule of thumb: start at the simplest tier. Reach for an agent only when the task is genuinely open-ended (multi-step, hard to fully specify, errors recoverable, value justifies cost).

Messages API Quick Start

Everything goes through POST /v1/messages. Headers: x-api-key, anthropic-version: 2023-06-01, content-type: application/json.

# pip install anthropic
import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=16000,
    system="You are a concise technical assistant.",
    messages=[{"role": "user", "content": "Explain CRDTs in one paragraph."}],
)
for block in response.content:        # content is a list of typed blocks
    if block.type == "text":          # always check .type before .text
        print(block.text)
print(response.stop_reason, response.usage.input_tokens, response.usage.output_tokens)
// npm install @anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();

const response = await client.messages.create({
  model: "claude-opus-5",
  max_tokens: 16000,
  messages: [{ role: "user", content: "Explain CRDTs in one paragraph." }],
});
for (const block of response.content) {
  if (block.type === "text") console.log(block.text);  // narrow the union first
}

Streaming (default to it for long outputs — non-streaming above ~16K max_tokens risks SDK HTTP timeouts):

with client.messages.stream(model="claude-opus-5", max_tokens=64000,
                            messages=[{"role": "user", "content": "Write a long report"}]) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    final = stream.get_final_message()   # full Message after streaming

Full params, response shape, stop reasons, errors, retries, rate limits: references/messages-api.md

Thinking & Effort (quick reference)

  • Adaptive thinking is ON BY DEFAULT on Fable 5 / Opus 5 / Sonnet 5 — send no thinking field and you still get (and pay for) thinking. On the legacy 4.6–4.8 models it stays off until you set thinking: {"type": "adaptive"}.
  • Manual budgets are gone. {"type": "enabled", "budget_tokens": N} returns a 400 on Opus 4.7 and every later model (Opus 5, Sonnet 5, Fable 5 included); deprecated on Opus 4.6 / Sonnet 4.6. Control depth with effort, not tokens.
  • Turning thinking off: Sonnet 5 accepts {"type": "disabled"}. Opus 5 accepts it only at effort high or below — pairing it with xhigh/max is a 400. Fable 5 rejects it outright; thinking there is unconditional, so budget for it.
  • Effort (GA): output_config: {"effort": "low" | "medium" | "high" | "xhigh" | "max"} — nested in output_config, not top-level. Default high (identical to omitting it). xhigh: Fable 5, Opus 5, Opus 4.8/4.7, Sonnet 5. max: those plus Opus 4.6 and Sonnet 4.6. Haiku 4.5 does not support effort at all.
  • Sampling params removed on Opus 4.7 and later (so Opus 5, Sonnet 5, Fable 5): temperature, top_p, top_k all return 400 — and the Python SDK v1.0+ doesn't define them, so passing them raises TypeError. Steer with prompting + effort.
  • Forced tool_choice is fine with adaptive thinking. The auto/none-only restriction applies to manual extended thinking ({"type": "enabled"}) only; adaptive mode — including the models where it's on by default — accepts {"type": "any"} and {"type": "tool", ...}.
  • Thinking text is omitted by default on Fable 5 / Opus 5 / Sonnet 5 / Opus 4.8 / 4.7 — opt in with thinking: {"type": "adaptive", "display": "summarized"} if you surface reasoning to users. Either way the blocks are billed, and must be echoed back unmodified (empty thinking field included) in a tool-use loop, or the next request 400s.

Details and gotchas: references/structured-outputs.md (thinking interplay) and references/messages-api.md.

Tool Use (quick reference)

tools = [{
    "name": "get_weather",
    "description": "Get current weather. Call when the user asks about weather conditions.",
    "input_schema": {
        "type": "object",
        "properties": {"location": {"type": "string", "description": "City, e.g. Paris"}},
        "required": ["location"],
    },
}]
response = client.messages.create(model="claude-opus-5", max_tokens=16000,
                                  tools=tools, messages=messages)
if response.stop_reason == "tool_use":
    ...  # execute, send tool_result back, loop

tool_choice: {"type": "auto"} (default) | {"type": "any"} | {"type": "tool", "name": "..."} | {"type": "none"}. Add "disable_parallel_tool_use": true to force at most one call per response.

The agentic loop, parallel tool results, pause_turn, is_error, server-side tools, and SDK tool runners: references/tool-use.md

Cost Optimization Checklist

Work top-down; each item is independent:

  • Right-size the model. Haiku for classification/routing, Sonnet for volume work, Opus/Fable for the hard 10%. Largest single lever.
  • Prompt caching on stable prefixes (system prompt, tool defs, big docs): cache_control: {"type": "ephemeral"}. Reads cost ~0.1x; up to 90% savings. Verify with usage.cache_read_input_tokens > 0 — zero means a silent invalidator (timestamp in system prompt, unsorted JSON, varying tools).
  • Batches API for anything that can wait ≤24h: flat 50% off all tokens, stacks with caching.
  • Cap output: set max_tokens to what you need (256 for classification); stream + generous cap for long generation.
  • Tune effort down where quality allows: medium is often the sweet spot; low for subagents and simple tasks.
  • Count before sending: client.messages.count_tokens(...) (never tiktoken — it's OpenAI's tokenizer and undercounts Claude by 15-20%).
  • Keep prefixes stable: order requests tools → system → messages, volatile content last; don't swap tool sets or models mid-conversation.

Mechanics, breakpoints, TTLs, batch lifecycle, tiering math: references/caching-and-cost.md

Context Engineering

Prompt engineering asks what to write in the prompt. Context engineering asks what earns a place in the window on this call — including everything that lands there without you typing it: tool definitions, tool results, retrieved documents, prior turns, thinking blocks. It is iterative (every inference) where prompt engineering is discrete (written once). Target: the smallest set of high-signal tokens that gets the outcome.

The budget is real because attention degrades with length (context rot — n² pairwise relationships), not just because tokens cost money. A 1M window is a capacity, not a target.

The three tiers

Every candidate fact lives in exactly one place. Choosing deliberately is most of the job.

Tier Where Cost Use when
1 — In context tools / system / messages, every call Paid every turn (≈0.1× cached) It steers most turns
2 — On disk, read on demand A file the agent can read; only the path stays in context Paid only when read The agent can tell from a name that it needs this
3 — Retrieved Index / search tool behind a query Paid only on a hit, plus a relevance gamble The corpus is too large to enumerate

When a prompt is too big, demote before you delete — a path is ~10 tokens; the file it names may be 10,000.

This repo already runs on the tier-1/tier-2 split. A skill's description is always resident (tier 1, so it must carry the routing signal); SKILL.md loads on a match; references/*.md load only when cited and needed. "Description is the trigger", "body under 500 lines", "one concept per reference", "every reference must be cited" are context-engineering rules wearing authoring clothes.

Cache-aware prompt architecture

Requests render tools → system → messages, and the cache is a prefix match. So static prefix first, volatile content last — put the cache_control breakpoint at the end of the stable part and let per-request content fall after it.

Reordering a prompt destroys the cache silently: no error, just a different prefix hash, cache_read_input_tokens: 0, and a 1.25–2× bill where you expected 0.1×. The usage block is the only symptom, which is why asserting cache_read_input_tokens > 0 in staging is a real test.

The compaction decision

Under modern prompt caching, keeping the full history has been measured to beat summarisation on cost, latency AND recall at the same time. A 2026 production-tutor evaluation (660 turns, 11 configurations) put keep-everything at 92–100% fact recall, $0.11/turn and 17 s TTFT, against 38–58% recall, $0.24/turn and 21 s for its clear-plus-summarise preset. Summarising rewrites the cached prefix and forfeits the 0.1× discount — the cheap move is usually to append. (That study ran on a non-Claude model; what transfers is the mechanism, and Claude's flat 0.1× cache read makes it stronger, not weaker. Full caveats in references/compaction.md.)

So: compact only as a deliberate response to a named constraint.

Constraint Diagnose Try first
Context ceiling — it will not fit Projected tokens > window Cap tool output → payloads to files → server-side clearing
Cost ceiling — the bill is unacceptable Compare against cached cost, not uncached Verify the cache is hitting → tier down → cap tool output
Latency target — TTFT too slow at depth Confirm growth is in the prefix Cap tool output → lower effort → stream

Capping tool output at the tool boundary is the underrated lever: it shrinks context without rewriting the cached prefix (the same study measured −38% cost/turn with no recall loss). Clearing and summarising both break the cache; they are what people reach for first and should reach for last.

First-party clearing is context_management (beta context-management-2025-06-27): clear_tool_uses_20250919 and clear_thinking_20251015, applied server-side. Always set clear_at_least — it stops a trigger paying a full cache re-write to save a handful of tokens. Pair with the memory tool so durable conclusions are written out before raw material is cleared.

Agentic specifics

  • Tool results are the growth term, not the system prompt. Design tools to return decisions, not dumps.
  • Summarise vs write-to-file: needed later in full → write to a file, return the path. Only the conclusion matters → summarise at the tool boundary (free of cache cost, unlike rewriting history after the fact).
  • Sub-agents are context isolation, not just parallelism: 80K tokens of exploration are billed once inside the child and discarded; the parent sees a ~1–2K-token distillation. Costs: cold cache in the child, a lossy hand-off. Skip it when the subtask needs most of the parent's context to make sense.

Full doctrine — tiers, progressive disclosure, instrumentation: references/context-engineering.md. Compaction economics, context_management parameters, memory tool: references/compaction.md. For Claude Code's own context surface see the claude-code-ops skill; for prompts re-sent on a cadence, loop-ops; for cross-provider fan-out, fleetflow.

Claude Agent SDK (quick reference)

# pip install claude-agent-sdk   (Python >= 3.10)
import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions

async def main():
    async for message in query(
        prompt="Find and fix the bug in auth.py",
        options=ClaudeAgentOptions(allowed_tools=["Read", "Edit", "Bash"]),
    ):
        if hasattr(message, "result"):
            print(message.result)

asyncio.run(main())
// npm install @anthropic-ai/claude-agent-sdk
import { query } from "@anthropic-ai/claude-agent-sdk";

for await (const message of query({
  prompt: "Find and fix the bug in auth.ts",
  options: { allowedTools: ["Read", "Edit", "Bash"] },
})) {
  if ("result" in message) console.log(message.result);
}

Built-in tools (Read/Write/Edit/Bash/Glob/Grep/WebSearch/WebFetch/...), hooks (PreToolUse, PostToolUse, ...), subagents, MCP servers, sessions (resume/fork), permission modes, and the SDK-vs-raw-API decision: references/agent-sdk.md

Common Pitfalls

Pitfall Symptom Fix
Date-suffixed or guessed model ID 404 not_found_error Use exact alias IDs from the table above
budget_tokens on Opus 4.7+ (incl. Opus 5 / Sonnet 5 / Fable 5) 400 thinking: {"type": "adaptive"} + effort
Assuming thinking is opt-in on Fable 5 / Opus 5 / Sonnet 5 Unexpected thinking tokens billed Adaptive thinking is on by default there; Fable 5 can't be disabled at all
thinking: {"type": "disabled"} at xhigh/max on Opus 5 400 Drop effort to high or below, or leave thinking on
temperature/top_p/top_k on Opus 4.7+ 400 (or TypeError on Python SDK v1.0+) Remove; steer via prompt + effort
effort on Haiku 4.5 400 Haiku 4.5 doesn't support the parameter
Rebuilding assistant turns in a tool loop (dropping empty thinking blocks) 400 "thinking blocks cannot be modified" Echo the content list back exactly as received
Assistant-turn prefill on Opus 4.7+ models 400 output_config.format or system-prompt instruction
Cache marker on <minimum prefix Silent no-cache (cache_creation_input_tokens: 0) Min 512-4096 tokens depending on model (see caching ref)
Not handling stop_reason: "tool_use" Agent "stops" after first tool call Loop: execute tools, append tool_result, re-request
Missing tool_result for a tool_use id 400 on follow-up One tool_result per tool_use block, ids matching
Non-streaming with max_tokens > ~16K SDK timeout / ValueError Stream + get_final_message() / finalMessage()
output_format top-level param Deprecated output_config: {"format": {...}}
tiktoken for Claude token counts 15-20%+ undercount messages.count_tokens endpoint
String-matching error messages Fragile retries Typed exceptions: anthropic.RateLimitError etc.
Raw string-matching tool input Breaks on escaping changes Always json.loads() / use parsed block.input
Compacting by reflex on a long conversation Higher cost, worse recall than doing nothing Name the constraint first; under caching, appending usually wins (see Context Engineering)
clear_tool_uses without clear_at_least A full cache re-write to reclaim a few hundred tokens Set clear_at_least so each cache break is worth taking

Resources & Verification

This skill ships a staleness verifier and two copy-and-adapt starter assets. The model table and pricing above are the facts most likely to drift — run the verifier when you suspect they're stale.

scripts/check-model-table.py — guards the Current Models table (this file) and the per-model prompt-cache minimum table (references/caching-and-cost.md) against drift. Two modes per the resource protocol §7:

# Structural (default, no network): every row well-formed, ids carry no date
# suffix, prices numeric, the two files agree on the model lineup. It also
# guards the cache-economics constants that are stated in more than one file
# (0.1x read, 1.25x/2x writes, 4 breakpoints, 20-block lookback, the
# context-management beta id), asserts each doctrine reference carries a
# "verified <ISO date>" stamp, and checks SKILL.md <-> references/ citation
# integrity in both directions. It then scans every file in the skill for model
# ids: an id in neither the table nor the Legacy list is flagged "unknown", and
# a LEGACY id sitting where a reader would copy it (model=..., "model": ...,
# --model ...) is flagged "retired" - append a `legacy-ok` comment to that line
# for a deliberate migration example. Exit 4 on any contradiction.
python skills/claude-api-ops/scripts/check-model-table.py --offline
python skills/claude-api-ops/scripts/check-model-table.py --offline --json | python -m json.tool

# Live (advisory, needs ANTHROPIC_API_KEY): curls the Models API and compares
# its id set against the documented ids. Exit 10 if a documented id is gone or a
# newer alias id is missing from the table; exit 7 (not a failure) if the key is
# unset or the API is unreachable. Live mode checks model-ID coverage ONLY — the
# API returns no pricing, so pricing/context drift stays an --offline + docs concern.
ANTHROPIC_API_KEY=sk-... python skills/claude-api-ops/scripts/check-model-table.py --live

scripts/context-budget.py — append-vs-compact calculator. Models both paths in dollars over the turns you actually have left, checks the context ceiling first, and exits 10 when cost favours compaction, 0 when appending wins:

# Short session — appending is cheaper (exit 0)
python skills/claude-api-ops/scripts/context-budget.py \
    --history-tokens 25000 --turns-remaining 5 --base-rate 0.30

# Deep session — cost favours compaction (exit 10)
python skills/claude-api-ops/scripts/context-budget.py \
    --history-tokens 120000 --turns-remaining 40 --base-rate 2.00 --json

Two results worth knowing before you trust it: the break-even turn count is scale-invariant (history size and price cancel out — it tracks the summary ratio, not how big or costly the conversation is), and it prices cost only. Recall loss is not in the model, so "compact" means cheaper, not better.

assets/cached-agent-loop.py — the cache-aware sibling of the minimal loop below, and the executable form of the Context Engineering section: breakpoint at the end of the static prefix, a rolling breakpoint on the newest turn, an intermediate anchor every ~15 blocks so long tool-heavy turns don't jump the 20-block lookback, tool output capped at the boundary, and a per-turn cache_read_input_tokens check that warns when the prefix silently changed. Copy it when the agent is long-running; copy agentic-loop.py when it isn't.

The footgun it encodes: cache_control is a key on a content block, so it can only be set on a dict. Appending response.content verbatim (SDK block objects) or using the "content": "a string" shorthand leaves nowhere to put a marker — every breakpoint aimed at those turns is discarded with no error and no warning. Normalise content to dict blocks before placing breakpoints.

assets/recall-probe.py — the "measure it on your workload" harness: plants a fact, buries it under N turns, probes for it, and reports recall, cost per turn and TTFT for append vs compact. Makes real API calls, so start small (--turns 6 --trials 1). Replace the synthetic filler turns with traffic from your own logs — that is the point of running it.

assets/agentic-loop.py — a minimal, runnable tool-use loop (define a tool, call messages.create, loop while stop_reason == "tool_use", append tool_result, re-request until end_turn). Copy it as the starting point when building a manual agent loop; the >>> ADAPT marks show what to change.

assets/output-schema.json — a known-good structured-outputs request body in the canonical output_config.format shape (with additionalProperties: false and a required array). Copy and reshape schema.properties when adding JSON outputs; see references/structured-outputs.md for the rules. (Supported on every current model — Fable 5, Opus 5, Sonnet 5, Haiku 4.5 — and the legacy 4.5–4.8 line.)

Reference Files

File Covers
references/messages-api.md Params, response shape, streaming events, stop reasons, error handling, retries, rate limits
references/tool-use.md Tool definitions, tool_choice, parallel tools, agentic loop, tool results, server tools, tool runners
references/caching-and-cost.md Prompt caching mechanics, Batches API, token counting, model tiering economics
references/structured-outputs.md output_config.format, schema rules/limits, strict tools, parse() helpers, thinking interplay
references/agent-sdk.md Python + TS Agent SDK, ClaudeAgentOptions, hooks, MCP, sessions, SDK vs raw API
references/context-engineering.md Context budget, the three tiers, progressive disclosure, cache-aware ordering, tool-result bloat, sub-agents as isolation, instrumentation
references/compaction.md When compaction is justified, break-even arithmetic, context_management edits, memory tool, how to compact well

Live Documentation

When cached facts may be stale, WebFetch (append .md for clean markdown):

  • Models/pricing: https://platform.claude.com/docs/en/about-claude/models/overview.md
  • Messages API: https://platform.claude.com/docs/en/api/messages
  • Tool use: https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview.md
  • Prompt caching: https://platform.claude.com/docs/en/build-with-claude/prompt-caching.md
  • Structured outputs: https://platform.claude.com/docs/en/build-with-claude/structured-outputs.md
  • Batches: https://platform.claude.com/docs/en/build-with-claude/batch-processing.md
  • Agent SDK: https://code.claude.com/docs/en/agent-sdk/overview
  • Context editing: https://platform.claude.com/docs/en/build-with-claude/context-editing
  • Context engineering: https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
Files (claude-mods)
  • assets
    • .gitkeep 0 B · in bundle
    • agentic-loop.py 4.4 KB
      #!/usr/bin/env python3
      """Minimal, correct tool-use agentic loop on the Anthropic Messages API.
      
      The canonical pattern: define a tool, call messages.create, and keep looping
      while stop_reason == "tool_use" — execute each requested tool, append a
      tool_result, and re-request until stop_reason == "end_turn".
      
      Run:  pip install anthropic   (then: export ANTHROPIC_API_KEY=sk-...)
            python agentic-loop.py
      
      Copy this file and adapt the >>> ADAPT marks for your own tools.
      Reflects the current API (model claude-opus-5, typed content blocks).
      """
      # The Anthropic SDK accepts plain dict literals for tools/messages at runtime
      # (as the official docs show), but its strict TypedDict stubs over-narrow them.
      # Silence those false positives so this starter stays readable; real apps may
      # prefer the SDK's typed params (anthropic.types.ToolParam, MessageParam).
      # pyright: reportArgumentType=false
      import anthropic
      
      client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from the environment
      
      MODEL = "claude-opus-5"  # >>> ADAPT: pick a tier (see the skill's model table)
      
      
      # --- 1. Define your tool(s) ------------------------------------------------
      # The input_schema is JSON Schema. Write a description that says WHEN to call.
      TOOLS = [
          {
              "name": "get_weather",  # >>> ADAPT
              "description": "Get the current weather for a city. "
                             "Call this whenever the user asks about weather conditions.",
              "input_schema": {
                  "type": "object",
                  "properties": {
                      "location": {"type": "string", "description": "City, e.g. 'Paris'"},
                  },
                  "required": ["location"],
              },
          }
      ]
      
      
      # --- 2. Implement each tool ------------------------------------------------
      # Map tool name -> a Python callable. Never trust the arguments blindly: they
      # come from the model. Validate before doing anything with side effects.
      def get_weather(location: str) -> str:  # >>> ADAPT: real implementation
          return f"It is 21°C and sunny in {location}."
      
      
      TOOL_IMPLS = {"get_weather": get_weather}
      
      
      def run_tool(name: str, tool_input: dict) -> str:
          """Dispatch a tool call, returning a string result for the model."""
          impl = TOOL_IMPLS.get(name)
          if impl is None:
              return f"ERROR: unknown tool {name!r}"
          try:
              return impl(**tool_input)
          except Exception as exc:  # surface failures back to the model, don't crash
              return f"ERROR running {name}: {exc}"
      
      
      # --- 3. The loop -----------------------------------------------------------
      def agent(user_prompt: str, max_turns: int = 10) -> str:
          # The conversation is a growing list of message dicts we own and replay.
          messages = [{"role": "user", "content": user_prompt}]
      
          for _ in range(max_turns):
              response = client.messages.create(
                  model=MODEL,
                  max_tokens=4096,
                  tools=TOOLS,
                  messages=messages,
              )
      
              # Append the assistant turn VERBATIM — content is a list of typed
              # blocks (text and/or tool_use). It must go back as-is next request.
              messages.append({"role": "assistant", "content": response.content})
      
              # If the model didn't ask for a tool, we're done — return its text.
              if response.stop_reason != "tool_use":
                  return "".join(
                      block.text for block in response.content if block.type == "text"
                  )
      
              # Otherwise: execute EVERY tool_use block and collect tool_result
              # blocks (the model may request several tools in parallel).
              tool_results = []
              for block in response.content:
                  if block.type != "tool_use":
                      continue  # skip text/thinking blocks
                  result_text = run_tool(block.name, block.input)
                  tool_results.append({
                      "type": "tool_result",
                      "tool_use_id": block.id,   # MUST echo the matching id
                      "content": result_text,
                      # "is_error": True,        # set when the tool failed
                  })
      
              # Feed results back as a single user turn, then loop to re-request.
              messages.append({"role": "user", "content": tool_results})
      
          return "Stopped: hit max_turns without an end_turn."
      
      
      if __name__ == "__main__":
          answer = agent("What's the weather in Tokyo right now?")
          print(answer)
          # For debugging, inspect the assembled transcript:
          # print(json.dumps(..., indent=2, default=str))
      
    • cached-agent-loop.py 14.1 KB
      #!/usr/bin/env python3
      """A cache-correct agentic loop — the context-engineering doctrine, executable.
      
      assets/agentic-loop.py is the MINIMAL correct loop: it shows stop_reason
      handling and nothing else, on purpose. This file is its cache-aware sibling.
      Same loop, plus the four things that decide whether a long-running agent costs
      0.1x or 1.25x per turn (see SKILL.md "Context Engineering"):
      
        1. STATIC PREFIX FIRST, VOLATILE LAST. tools -> system -> messages is the
           render order, and the cache is a prefix match. The breakpoint goes at the
           end of the stable part; anything per-request falls after it.
        2. A ROLLING BREAKPOINT on the newest turn, so hits accrue as the
           conversation grows -- and never more than MAX_BREAKPOINTS of them.
        3. AN INTERMEDIATE BREAKPOINT every ~15 blocks. A breakpoint searches
           backward at most 20 content blocks; a turn that appends more than that
           jumps the window and silently misses.
        4. TOOL OUTPUT CAPPED AT THE BOUNDARY. The cheapest context lever there is:
           it shrinks context WITHOUT rewriting the cached prefix, so unlike
           compaction it costs nothing in cache terms.
        5. CONTENT NORMALISED TO DICT BLOCKS. cache_control is a key on a content
           block, so it can only be set on a dict. `response.content` is a list of
           SDK block OBJECTS and "content": "a string" has no blocks at all -- append
           either verbatim and every marker aimed at it is silently discarded. An
           earlier version of this file did exactly that and placed ZERO message
           breakpoints in a realistic conversation, with no error and no warning.
           to_blocks() is what makes points 2 and 3 actually take effect.
      
      And the one assertion that matters: cache_read_input_tokens > 0. A broken
      cache produces no error, no warning, and no symptom other than the bill --
      this loop prints the usage every turn and complains when the cache misses.
      
      Run:  pip install anthropic   (then: export ANTHROPIC_API_KEY=sk-...)
            python cached-agent-loop.py
      
      Copy this file and adapt the >>> ADAPT marks. Cache facts verified against
      platform.claude.com 2026-08-30; see references/caching-and-cost.md.
      """
      # The Anthropic SDK accepts plain dict literals for tools/messages at runtime
      # (as the official docs show), but its strict TypedDict stubs over-narrow them.
      # Silence those false positives so this starter stays readable.
      # pyright: reportArgumentType=false
      import anthropic
      
      client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from the environment
      
      MODEL = "claude-opus-5"  # >>> ADAPT: pick a tier (see the skill's model table)
      
      # --- Cache tuning knobs (documented values, not guesses) --------------------
      MAX_BREAKPOINTS = 4      # hard API limit: 4 cache_control markers per request
      LOOKBACK_BLOCKS = 20     # a breakpoint searches back at most 20 content blocks
      BREAKPOINT_EVERY = 15    # ...so re-anchor before that window is exhausted
      TOOL_RESULT_CAP = 4000   # >>> ADAPT: chars kept per tool result (lever #4)
      
      # The minimum cacheable prefix is MODEL-DEPENDENT (512-4096 tokens). Below it
      # the marker is silently ignored -- cache_creation_input_tokens stays 0 and
      # there is no error. If your system prompt is small, caching it buys nothing;
      # check the per-model table in references/caching-and-cost.md before tuning.
      
      
      # === 1. THE STATIC PREFIX ====================================================
      # Everything here must be byte-identical on every request. The classic silent
      # invalidators live in exactly this block: a datetime.now(), a request id, a
      # per-user name, a conditionally-appended section, an unsorted json.dumps.
      # >>> ADAPT: put your real instructions/docs here. Keep them CONSTANT.
      SYSTEM_PROMPT = """You are a careful assistant with tool access.
      Prefer calling a tool over guessing. Answer concisely once you have the facts.
      """ + ("Reference material the agent needs on most turns goes here. " * 200)
      
      # Tools render at position 0, ahead of system. Changing ANY tool definition
      # invalidates the entire cache -- so build this list once, statically. Never
      # assemble it per-user or per-request.
      TOOLS = [
          {
              "name": "get_weather",  # >>> ADAPT
              "description": "Get current weather for a city. Call when asked about weather.",
              "input_schema": {
                  "type": "object",
                  "properties": {"location": {"type": "string",
                                              "description": "City, e.g. Paris"}},
                  "required": ["location"],
              },
          },
      ]
      
      
      def run_tool(name: str, tool_input: dict) -> str:
          """>>> ADAPT: dispatch to your real implementations."""
          if name == "get_weather":
              return f"18C, light rain in {tool_input.get('location', 'unknown')}."
          return f"No such tool: {name}"
      
      
      def capped(text: str, limit: int = TOOL_RESULT_CAP) -> str:
          """Lever #4 — cap at the TOOL BOUNDARY, before the result enters context.
      
          This is the important asymmetry: truncating here shortens what gets
          APPENDED, so the cached prefix behind it is untouched. Summarising the same
          content *after* it is already in the history is compaction, and pays a full
          cache write. If the caller may need the full payload, write it to a file
          and return the path instead of truncating (tier 2 -- see
          references/context-engineering.md §5.2).
          """
          if len(text) <= limit:
              return text
          return (text[:limit]
                  + f"\n... [truncated {len(text) - limit} chars. "
                    f"Re-run with a narrower query, or read the full payload from disk.]")
      
      
      # === 2/3. BREAKPOINT PLACEMENT ==============================================
      def to_blocks(content) -> list:
          """Normalise any message content into a list of plain dict blocks.
      
          THIS IS LOAD-BEARING, not tidiness. cache_control is a key on a content
          block, so a breakpoint can only be attached to a dict. Two shapes in normal
          use are NOT dicts and will silently refuse every marker:
      
            * `response.content` — SDK block OBJECTS (TextBlock, ToolUseBlock, ...).
              Appending them verbatim, as the minimal loop does, is idiomatic and
              correct for a loop that never caches. Here it means every assistant
              turn is un-markable.
            * `"content": "a plain string"` — the shorthand form has no blocks at
              all, so there is nowhere to put a marker.
      
          Either one produces NO error and NO warning: the request simply caches
          less than you think. Normalising at append time is what keeps the
          breakpoint logic below sound.
          """
          if isinstance(content, str):
              return [{"type": "text", "text": content}]
          out = []
          for block in content:
              if isinstance(block, dict):
                  out.append(block)
              elif hasattr(block, "model_dump"):        # pydantic v2 (current SDK)
                  out.append(block.model_dump(exclude_none=True))
              elif hasattr(block, "dict"):              # pydantic v1
                  out.append(block.dict(exclude_none=True))
              else:                                     # last resort
                  out.append(dict(vars(block)))
          return out
      
      
      def _blocks(message) -> list:
          content = message["content"]
          return content if isinstance(content, list) else []
      
      
      def place_message_breakpoints(messages: list) -> int:
          """Re-anchor rolling cache breakpoints across the message list, in place.
      
          Returns the number of markers actually placed — check it. Silently placing
          zero is the failure this function exists to prevent.
      
          Rules encoded here:
            * The newest turn always carries a breakpoint, so the next request can
              read everything up to it (hits accrue as the conversation grows).
            * An extra anchor every BREAKPOINT_EVERY blocks, because the backward
              search stops after LOOKBACK_BLOCKS (20). A tool-heavy turn that appends
              30 blocks would otherwise jump clean over the previous entry.
            * At most MAX_BREAKPOINTS - 1 markers here; the system block owns the
              fourth. Exceeding 4 is an API error, so oldest markers are dropped.
            * A marker can only sit on a dict block. If the chosen position is not
              one (an un-normalised SDK object, say), walk BACKWARD to the nearest
              dict rather than dropping the anchor — dropping it silently is exactly
              how a loop ends up with no caching at all. Run content through
              to_blocks() and this path never triggers.
          """
          assert BREAKPOINT_EVERY < LOOKBACK_BLOCKS, "anchor must fall inside the window"
      
          # Clear existing markers, then re-place: idempotent, so the loop can call
          # this every turn without accumulating stale markers.
          for msg in messages:
              for block in _blocks(msg):
                  if isinstance(block, dict):
                      block.pop("cache_control", None)
      
          # Walk the flattened block stream, marking a candidate every N blocks and
          # always marking the final block. Positions count EVERY block (the API sees
          # them all), even ones that cannot themselves carry a marker.
          flat = [(mi, bi) for mi, msg in enumerate(messages)
                  for bi, _ in enumerate(_blocks(msg))]
          if not flat:
              # Every message used the plain-string shorthand: nothing can be cached.
              print("  WARNING: no content blocks to anchor a cache breakpoint on.\n"
                    "           Message content is in string shorthand - run it through\n"
                    "           to_blocks() so cache_control has somewhere to attach.")
              return 0
      
          def settable(idx: int):
              """Nearest dict block at or before idx, within the lookback window."""
              for k in range(idx, max(-1, idx - LOOKBACK_BLOCKS), -1):
                  mi, bi = flat[k]
                  if isinstance(_blocks(messages[mi])[bi], dict):
                      return k
              return None
      
          chosen: list[int] = []
          for i in list(range(BREAKPOINT_EVERY - 1, len(flat), BREAKPOINT_EVERY)) + [len(flat) - 1]:
              k = settable(i)
              if k is not None and k not in chosen:
                  chosen.append(k)
      
          # Keep the most recent ones; the system breakpoint consumes one of the 4.
          placed = 0
          for k in chosen[-(MAX_BREAKPOINTS - 1):]:
              mi, bi = flat[k]
              block = _blocks(messages[mi])[bi]
              if isinstance(block, dict):
                  block["cache_control"] = {"type": "ephemeral"}
                  placed += 1
                  # >>> ADAPT: {"type": "ephemeral", "ttl": "1h"} if turns are minutes
                  # apart. 1h writes cost 2x vs 1.25x, so it needs 3+ reads to pay off.
          if placed == 0:
              print("  WARNING: placed 0 message cache breakpoints - the conversation\n"
                    "           will re-read from the system breakpoint only. Normalise\n"
                    "           message content with to_blocks().")
          return placed
      
      
      def report_usage(usage, turn: int, first_turn: bool) -> None:
          """The only symptom a broken cache produces is in here. Look at it."""
          read = getattr(usage, "cache_read_input_tokens", 0) or 0
          written = getattr(usage, "cache_creation_input_tokens", 0) or 0
          print(f"  [turn {turn}] uncached={usage.input_tokens} "
                f"cache_write={written} cache_read={read} out={usage.output_tokens}")
          if not first_turn and read == 0:
              # Not an exception on purpose: this is a cost bug, not a crash. In CI or
              # staging, promote it -- an assert here catches a stray timestamp in the
              # system prompt before it reaches production.
              print("  WARNING: cache_read_input_tokens == 0 on a repeat request.\n"
                    "           The prefix changed. Check for a timestamp/uuid/per-user\n"
                    "           value in SYSTEM_PROMPT, a reordered block, an unsorted\n"
                    "           json.dumps, a swapped model, or a mutated tool list.\n"
                    "           See references/caching-and-cost.md, 'Silent invalidators'.")
      
      
      def main() -> None:
          # >>> ADAPT: your opening request. Volatile content belongs HERE, after the
          # cached system block -- never interpolated into SYSTEM_PROMPT.
          # Block form, not the "content": "..." string shorthand — a string has no
          # block for cache_control to attach to.
          messages = [{"role": "user",
                       "content": [{"type": "text",
                                    "text": "What's the weather in Paris and in Oslo?"}]}]
      
          for turn in range(1, 21):  # bounded: never ship an unbounded agent loop
              place_message_breakpoints(messages)  # returns the count; 0 means no caching
      
              response = client.messages.create(
                  model=MODEL,
                  max_tokens=4096,
                  # The static prefix, with the breakpoint at its end. Tools render
                  # ahead of system, so this one marker caches BOTH together.
                  system=[{"type": "text", "text": SYSTEM_PROMPT,
                           "cache_control": {"type": "ephemeral"}}],
                  tools=TOOLS,
                  messages=messages,
              )
              report_usage(response.usage, turn, first_turn=(turn == 1))
      
              if response.stop_reason != "tool_use":
                  for block in response.content:
                      if block.type == "text":
                          print(block.text)
                  return
      
              # to_blocks(), not response.content verbatim: SDK block objects cannot
              # carry cache_control, so appending them raw silently un-caches every
              # assistant turn (no error, just a bigger bill).
              messages.append({"role": "assistant",
                               "content": to_blocks(response.content)})
      
              tool_results = []
              for block in response.content:
                  if block.type != "tool_use":
                      continue
                  try:
                      # block.input is already parsed -- never string-match the raw text.
                      output = run_tool(block.name, block.input)
                      is_error = False
                  except Exception as exc:  # tool failures are data, not crashes
                      output, is_error = f"Tool failed: {exc}", True
                  tool_results.append({
                      "type": "tool_result",
                      "tool_use_id": block.id,      # one result per tool_use, ids matching
                      "content": capped(output),    # <- lever #4, at the boundary
                      "is_error": is_error,
                  })
              messages.append({"role": "user", "content": tool_results})
      
          print("Hit the turn ceiling without an end_turn - check the tool loop.")
      
      
      if __name__ == "__main__":
          main()
      
    • output-schema.json 1.4 KB
      {
        "_comment": "Canonical structured-outputs request body for the Anthropic Messages API. Pass the `output_config` object below in client.messages.create(...). This is the current `output_config.format` shape — NOT the deprecated top-level `output_format` param. Adapt `schema.properties` and `required` to your data; keep `additionalProperties: false` so the model can't add stray keys. Supported on Fable 5, Opus 5, Sonnet 5, Haiku 4.5, and the legacy Opus 4.8/4.7/4.6/4.5 and Sonnet 4.6/4.5 line — use plain alias ids.",
        "model": "claude-opus-5",
        "max_tokens": 1024,
        "messages": [
          {
            "role": "user",
            "content": "Extract the contact: Jane Doe (jane@example.com) wants the Enterprise plan and asked for a demo."
          }
        ],
        "output_config": {
          "format": {
            "type": "json_schema",
            "schema": {
              "type": "object",
              "properties": {
                "name": {
                  "type": "string",
                  "description": "Full name of the contact"
                },
                "email": {
                  "type": "string",
                  "format": "email"
                },
                "plan": {
                  "type": "string",
                  "enum": ["Free", "Pro", "Enterprise"]
                },
                "demo_requested": {
                  "type": "boolean",
                  "description": "Whether the contact asked for a demo"
                }
              },
              "required": ["name", "email", "plan", "demo_requested"],
              "additionalProperties": false
            }
          }
        }
      }
      
    • recall-probe.py 10 KB
      #!/usr/bin/env python3
      """Measure what compaction actually costs you — on YOUR workload.
      
      references/compaction.md says the keep-everything baseline is far stronger
      than teams assume, and that you should verify that on your own traffic before
      compacting. This is the harness for doing so. It is the published methodology,
      runnable: plant a fact, bury it under N turns, probe for it, and compare
      strategies on the three axes people reach for compaction to improve.
      
        plant (turn 0)  ->  N filler turns  ->  probe ("what was the code?")
                                                |
                              recall hit/miss + $ per turn + time to first token
      
      Strategies compared out of the box:
        append   keep the full history every turn (the baseline)
        compact  summarise everything older than --keep-recent turns, then continue
                 from the summary (what most agent frameworks do by default)
      
      The point is NOT to reproduce a published number. It is to find out whether
      YOUR workload behaves like the studied one, because the answer decides whether
      compaction is buying you anything. A strategy that wins on cost and loses on
      recall has not won.
      
      Run:  pip install anthropic   (then: export ANTHROPIC_API_KEY=sk-...)
            python recall-probe.py --turns 10 --trials 3
      
      Costs real money — it makes (turns + 2) x trials x strategies API calls.
      Start with --turns 6 --trials 1 on a cheap model to sanity-check the wiring.
      
      >>> ADAPT: replace FILLER_TURNS with real turns from your own application.
      Synthetic filler is the weakest part of any harness like this — a planted fact
      survives generic chit-chat far more easily than it survives your actual
      tool-heavy traffic, so results on the shipped filler will flatter both
      strategies.
      """
      # pyright: reportArgumentType=false
      import argparse
      import statistics
      import sys
      import time
      
      import anthropic
      
      client = anthropic.Anthropic()
      
      MODEL = "claude-haiku-4-5"  # >>> ADAPT: probe on the tier you actually ship
      
      SYSTEM_PROMPT = ("You are a helpful assistant. Answer concisely. "
                       "When asked to recall a specific detail from earlier in the "
                       "conversation, quote it exactly.")
      
      # >>> ADAPT: the fact to plant. Make it arbitrary and unguessable, so a correct
      # answer proves recall rather than a plausible reconstruction from priors.
      SECRET_LABEL = "deployment code"
      SECRET_VALUE = "TANGERINE-4471-QUAY"
      
      PLANT = (f"Before we start: the {SECRET_LABEL} for this session is "
               f"{SECRET_VALUE}. Acknowledge it and we'll move on.")
      PROBE = f"What was the {SECRET_LABEL} I gave you at the start? Answer with just the code."
      
      # >>> ADAPT: replace with turns sampled from your own logs.
      FILLER_TURNS = [
          "Explain the difference between a process and a thread.",
          "Now give me a short example of a race condition.",
          "How would you fix it with a mutex?",
          "What's the tradeoff versus a lock-free queue?",
          "Summarise when I should prefer each.",
          "What about async I/O — where does that fit?",
          "Give me a rule of thumb for choosing a concurrency model.",
          "What metrics would tell me I chose wrong?",
          "How would I load-test that?",
          "What's a common mistake teams make here?",
      ]
      
      COMPACT_INSTRUCTION = (
          # Per Anthropic's guidance: maximise RECALL first, then tune precision. A
          # compaction prompt written for brevity silently drops the one detail the
          # next 40 turns needed -- which is precisely what this harness measures.
          "Summarise the conversation so far. Preserve every concrete detail, "
          "identifier, number, name and code mentioned — completeness matters more "
          "than brevity. Write it as notes for someone continuing the conversation."
      )
      
      
      def send(messages, cache=True):
          """One request. Returns (text, usage, time_to_first_token_seconds)."""
          system = [{"type": "text", "text": SYSTEM_PROMPT}]
          if cache:
              # Cache the static prefix -- otherwise 'append' is measured without the
              # very mechanism that makes it competitive, and the comparison is rigged.
              system[0]["cache_control"] = {"type": "ephemeral"}
          start = time.perf_counter()
          ttft = None
          chunks = []
          with client.messages.stream(model=MODEL, max_tokens=1024,
                                      system=system, messages=messages) as stream:
              for text in stream.text_stream:
                  if ttft is None:
                      ttft = time.perf_counter() - start
                  chunks.append(text)
              final = stream.get_final_message()
          return "".join(chunks), final.usage, (ttft if ttft is not None else 0.0)
      
      
      def cost_usd(usage, in_rate: float, out_rate: float) -> float:
          """Bill the three input buckets at their real multipliers, plus output."""
          read = getattr(usage, "cache_read_input_tokens", 0) or 0
          write = getattr(usage, "cache_creation_input_tokens", 0) or 0
          return (usage.input_tokens * in_rate
                  + write * in_rate * 1.25      # 5-minute-TTL cache write
                  + read * in_rate * 0.10       # cache read, flat across models
                  + usage.output_tokens * out_rate) / 1_000_000
      
      
      def run_trial(strategy: str, turns: int, keep_recent: int,
                    in_rate: float, out_rate: float, compact_every: int) -> dict:
          messages = [{"role": "user", "content": PLANT}]
          total_cost, ttfts, compactions = 0.0, [], 0
          turns_since_compaction = 0
      
          text, usage, ttft = send(messages)
          total_cost += cost_usd(usage, in_rate, out_rate)
          ttfts.append(ttft)
          messages.append({"role": "assistant", "content": text})
      
          for i in range(turns):
              messages.append({"role": "user", "content": FILLER_TURNS[i % len(FILLER_TURNS)]})
              text, usage, ttft = send(messages)
              total_cost += cost_usd(usage, in_rate, out_rate)
              ttfts.append(ttft)
              messages.append({"role": "assistant", "content": text})
      
              turns_since_compaction += 1
              if (strategy == "compact"
                      and turns_since_compaction >= compact_every
                      and len(messages) > 2 * keep_recent + 1):
                  # Summarise everything except the most recent keep_recent exchanges,
                  # then continue from the summary. This is the default behaviour of
                  # most agent frameworks -- and the thing under test.
                  #
                  # FAIRNESS, and the reason for the cadence: an earlier version of
                  # this harness compacted on EVERY turn once the history exceeded
                  # keep_recent, so the compact arm paid a summarisation call per turn
                  # (21 API calls against append's 12, over 10 turns). That inflates
                  # the arm under test and would "confirm" the keep-everything result
                  # regardless of what the data said. Real systems trigger on a token
                  # threshold (context_management's `trigger`); a turn cadence is the
                  # closest honest approximation without burning a count_tokens call
                  # every turn.
                  head, tail = messages[:-2 * keep_recent], messages[-2 * keep_recent:]
                  summary_text, usage, _ = send(
                      head + [{"role": "user", "content": COMPACT_INSTRUCTION}])
                  total_cost += cost_usd(usage, in_rate, out_rate)
                  messages = ([{"role": "user",
                                "content": f"[Summary of earlier conversation]\n{summary_text}"},
                               {"role": "assistant", "content": "Understood, continuing."}]
                              + tail)
                  compactions += 1
                  turns_since_compaction = 0
      
          messages.append({"role": "user", "content": PROBE})
          answer, usage, ttft = send(messages)
          total_cost += cost_usd(usage, in_rate, out_rate)
          ttfts.append(ttft)
      
          # Normalised substring match. Deliberately generous: a strategy that cannot
          # pass THIS has lost the fact outright, not merely paraphrased it.
          recalled = SECRET_VALUE.lower() in answer.lower()
          return {"recalled": recalled, "cost": total_cost, "compactions": compactions,
                  "ttft": statistics.mean(ttfts), "answer": answer.strip()[:120]}
      
      
      def main() -> int:
          ap = argparse.ArgumentParser(
              description="Plant-a-fact recall probe: append vs compact, on your workload.")
          ap.add_argument("--turns", type=int, default=10, help="filler turns before the probe")
          ap.add_argument("--trials", type=int, default=3, help="repeats per strategy")
          ap.add_argument("--keep-recent", type=int, default=2,
                          help="exchanges kept verbatim by the compact strategy")
          ap.add_argument("--compact-every", type=int, default=5,
                          help="turns between compactions (default: 5). Setting this "
                               "to 1 compacts every turn, which inflates the compact "
                               "arm and rigs the result toward keep-everything - only "
                               "do it if you are deliberately modelling that.")
          ap.add_argument("--in-rate", type=float, default=1.00, help="input $/MTok")
          ap.add_argument("--out-rate", type=float, default=5.00, help="output $/MTok")
          args = ap.parse_args()
      
          print(f"model={MODEL} turns={args.turns} trials={args.trials} compact_every={args.compact_every}\n")
          for strategy in ("append", "compact"):
              results = [run_trial(strategy, args.turns, args.keep_recent,
                                   args.in_rate, args.out_rate, args.compact_every)
                         for _ in range(args.trials)]
              hits = sum(r["recalled"] for r in results)
              per_turn = statistics.mean(r["cost"] for r in results) / (args.turns + 2)
              print(f"{strategy:8s} recall {hits}/{args.trials}  "
                    f"${per_turn:.5f}/turn  "
                    f"ttft {statistics.mean(r['ttft'] for r in results):.2f}s  "
                    f"compactions {statistics.mean(r['compactions'] for r in results):.1f}")
              for r in results:
                  if not r["recalled"]:
                      print(f"           miss -> {r['answer']!r}")
      
          print("\nIf append wins on all three, compaction is costing you twice. If "
                "compact wins on cost alone, price the recall loss before adopting it.")
          return 0
      
      
      if __name__ == "__main__":
          try:
              sys.exit(main())
          except KeyboardInterrupt:
              sys.exit(2)
      
  • references
    • agent-sdk.md 9.7 KB
      # Claude Agent SDK Reference
      
      The Agent SDK is **Claude Code as a library**: the same agent loop, built-in
      tools (file ops, bash, search, web), context management, and permission system,
      programmable from Python or TypeScript. Verified against
      code.claude.com/docs/en/agent-sdk (2026-06).
      
      | | Python | TypeScript |
      |---|---|---|
      | Package | `claude-agent-sdk` (`pip install claude-agent-sdk`) | `@anthropic-ai/claude-agent-sdk` (`npm install @anthropic-ai/claude-agent-sdk`) |
      | Runtime | Python ≥ 3.10 | Node (bundles a native Claude Code binary — no separate install) |
      | Entry point | `query(prompt, options=ClaudeAgentOptions(...))` → async iterator | `query({ prompt, options })` → async iterator |
      | Option naming | `snake_case` (`allowed_tools`) | `camelCase` (`allowedTools`) |
      
      Auth: `ANTHROPIC_API_KEY`, or third-party providers via env flags
      (`CLAUDE_CODE_USE_BEDROCK=1`, `CLAUDE_CODE_USE_VERTEX=1`,
      `CLAUDE_CODE_USE_FOUNDRY=1`, `CLAUDE_CODE_USE_ANTHROPIC_AWS=1` +
      `ANTHROPIC_AWS_WORKSPACE_ID`). Note: from June 15, 2026, Agent SDK / `claude -p`
      usage on subscription plans draws from a separate monthly Agent SDK credit —
      production apps should use API keys.
      
      ## When SDK vs raw API
      
      | Signal | Choice |
      |---|---|
      | Agent must read/edit files, run commands, search a codebase | **Agent SDK** — tools are built in, loop is handled |
      | CI/CD automation, "fix the failing test", repo-scale refactors | **Agent SDK** |
      | You want hooks, permission gating, session resume out of the box | **Agent SDK** |
      | Single call: classify/summarize/extract | **Messages API** — SDK is overkill |
      | You need exact control of every request (params, caching breakpoints, message shapes) | **Messages API + tool use** |
      | Your tools are pure in-process functions, no filesystem | **Messages API + tool runner** is lighter |
      | Hosted agent, no infra at all | **Managed Agents** (REST, Anthropic-run sandbox) |
      
      The difference in code:
      
      ```python
      # Messages API: you implement the loop
      response = client.messages.create(...)
      while response.stop_reason == "tool_use":
          result = your_tool_executor(...)
          response = client.messages.create(...)
      
      # Agent SDK: the loop, tools, and context management are inside query()
      async for message in query(prompt="Fix the bug in auth.py"):
          print(message)
      ```
      
      ## Minimal agents
      
      ```python
      import asyncio
      from claude_agent_sdk import query, ClaudeAgentOptions
      
      async def main():
          async for message in query(
              prompt="Find all TODO comments and create a summary",
              options=ClaudeAgentOptions(allowed_tools=["Read", "Glob", "Grep"]),
          ):
              if hasattr(message, "result"):       # final ResultMessage
                  print(message.result)
      
      asyncio.run(main())
      ```
      
      ```typescript
      import { query } from "@anthropic-ai/claude-agent-sdk";
      
      for await (const message of query({
        prompt: "Find all TODO comments and create a summary",
        options: { allowedTools: ["Read", "Glob", "Grep"] },
      })) {
        if ("result" in message) console.log(message.result);
      }
      ```
      
      The iterator yields typed messages: system messages (`subtype: "init"` carries
      `session_id`), assistant/tool activity, and a final result message
      (`message.result` / `ResultMessage`).
      
      ## ClaudeAgentOptions (key fields)
      
      | Python | TypeScript | Purpose |
      |---|---|---|
      | `allowed_tools` | `allowedTools` | Pre-approved tools (no permission prompt). Include `"Agent"` to auto-approve subagent spawns |
      | `disallowed_tools` | `disallowedTools` | Hard-blocked tools |
      | `permission_mode` | `permissionMode` | e.g. `"default"`, `"acceptEdits"`, `"bypassPermissions"`, `"plan"` |
      | `system_prompt` | `systemPrompt` | Replace or extend the system prompt |
      | `mcp_servers` | `mcpServers` | MCP server map (see below) |
      | `hooks` | `hooks` | Lifecycle callbacks (see below) |
      | `agents` | `agents` | Named subagent definitions |
      | `resume` | `resume` | Session ID to continue with full context |
      | `cwd` | `cwd` | Working directory for the agent |
      | `model` | `model` | Override model |
      | `max_turns` | `maxTurns` | Cap agent-loop iterations |
      | `setting_sources` | `settingSources` | Restrict which filesystem config loads (`.claude/`, `~/.claude/`) |
      | `plugins` | `plugins` | Programmatic plugin loading |
      
      ### Built-in tools
      
      `Read`, `Write`, `Edit`, `Bash`, `Glob`, `Grep`, `WebSearch`, `WebFetch`,
      `Monitor` (watch a background process), `AskUserQuestion` (clarifying
      questions with options), `Agent` (spawn subagents). A read-only agent is just
      `allowed_tools=["Read", "Glob", "Grep"]`.
      
      ## Hooks
      
      Callbacks at lifecycle points: `PreToolUse`, `PostToolUse`, `Stop`,
      `SessionStart`, `SessionEnd`, `UserPromptSubmit`, and more. Use them to audit,
      block, or transform agent behavior.
      
      ```python
      from claude_agent_sdk import query, ClaudeAgentOptions, HookMatcher
      from datetime import datetime
      
      async def log_file_change(input_data, tool_use_id, context):
          path = input_data.get("tool_input", {}).get("file_path", "unknown")
          with open("./audit.log", "a") as f:
              f.write(f"{datetime.now()}: modified {path}\n")
          return {}        # empty dict = allow; hooks can also block/modify
      
      options = ClaudeAgentOptions(
          permission_mode="acceptEdits",
          hooks={"PostToolUse": [HookMatcher(matcher="Edit|Write", hooks=[log_file_change])]},
      )
      ```
      
      ```typescript
      import { query, HookCallback } from "@anthropic-ai/claude-agent-sdk";
      import { appendFile } from "fs/promises";
      
      const logFileChange: HookCallback = async (input) => {
        const filePath = (input as any).tool_input?.file_path ?? "unknown";
        await appendFile("./audit.log", `${new Date().toISOString()}: modified ${filePath}\n`);
        return {};
      };
      
      const options = {
        permissionMode: "acceptEdits" as const,
        hooks: { PostToolUse: [{ matcher: "Edit|Write", hooks: [logFileChange] }] },
      };
      ```
      
      `matcher` is a regex over tool names. A `PreToolUse` hook returning a deny
      decision blocks the call — this is the programmatic equivalent of a human
      approval gate.
      
      ## MCP servers
      
      Wire any MCP server (stdio command or remote) into the agent:
      
      ```python
      options = ClaudeAgentOptions(
          mcp_servers={
              "playwright": {"command": "npx", "args": ["@playwright/mcp@latest"]},
          },
      )
      ```
      
      ```typescript
      const options = {
        mcpServers: {
          playwright: { command: "npx", args: ["@playwright/mcp@latest"] },
        },
      };
      ```
      
      MCP tools surface as `mcp__<server>__<tool>` — add them to
      `allowed_tools` to pre-approve. This is how you give an agent databases,
      browsers, and third-party APIs without writing tool plumbing.
      
      ## Subagents
      
      ```python
      from claude_agent_sdk import query, ClaudeAgentOptions, AgentDefinition
      
      options = ClaudeAgentOptions(
          allowed_tools=["Read", "Glob", "Grep", "Agent"],   # Agent tool approves spawns
          agents={
              "code-reviewer": AgentDefinition(
                  description="Expert code reviewer for quality and security reviews.",
                  prompt="Analyze code quality and suggest improvements.",
                  tools=["Read", "Glob", "Grep"],
              ),
          },
      )
      # prompt: "Use the code-reviewer agent to review this codebase"
      ```
      
      Messages emitted inside a subagent carry `parent_tool_use_id` so you can
      attribute output to the spawning call. Use subagents to isolate context
      (reviewer doesn't pollute the main transcript) and to parallelize independent
      legs.
      
      ## Sessions: resume and fork
      
      Session state is JSONL on your filesystem. Capture the session ID from the
      init message, resume later with full context:
      
      ```python
      from claude_agent_sdk import query, ClaudeAgentOptions, SystemMessage, ResultMessage
      
      session_id = None
      async for message in query(prompt="Read the authentication module",
                                 options=ClaudeAgentOptions(allowed_tools=["Read", "Glob"])):
          if isinstance(message, SystemMessage) and message.subtype == "init":
              session_id = message.data["session_id"]
      
      async for message in query(prompt="Now find all places that call it",
                                 options=ClaudeAgentOptions(resume=session_id)):
          if isinstance(message, ResultMessage):
              print(message.result)
      ```
      
      ```typescript
      let sessionId: string | undefined;
      for await (const message of query({ prompt: "Read the authentication module",
                                          options: { allowedTools: ["Read", "Glob"] } })) {
        if (message.type === "system" && message.subtype === "init") {
          sessionId = message.session_id;
        }
      }
      for await (const message of query({ prompt: "Now find all places that call it",
                                          options: { resume: sessionId } })) {
        if ("result" in message) console.log(message.result);
      }
      ```
      
      Sessions can also be forked to explore alternative approaches from the same
      context point.
      
      ## Filesystem configuration
      
      With default options the SDK loads Claude Code's filesystem config from
      `.claude/` (project) and `~/.claude/` (user): skills
      (`.claude/skills/*/SKILL.md`), commands, `CLAUDE.md` memory, plugins. Restrict
      with `setting_sources` / `settingSources` when you want a hermetic agent (CI)
      that ignores developer-machine state.
      
      ## Production patterns
      
      - **CI agent:** `permission_mode="bypassPermissions"` (or a tight
        `allowed_tools` list) + `max_turns` cap + hooks for audit logging. Never
        bypass permissions on a machine with credentials you don't want the agent
        exercising.
      - **Approval gate:** `PreToolUse` hook on `Bash|Write|Edit` that checks the
        input and returns deny for out-of-policy actions.
      - **Observability:** log every message from the iterator; hook
        `PostToolUse` for tool-level metrics; final `ResultMessage` includes
        cost/usage data.
      - **Prototype → production:** a common path is Agent SDK locally (works on
        your filesystem), then Managed Agents for hosted production (Anthropic runs
        the sandbox; custom tools become event round-trips).
      - Workflows translate 1:1 with the Claude Code CLI (`claude -p`) — anything
        you can do interactively you can automate via the SDK.
      
    • caching-and-cost.md 11.1 KB
      # Caching, Batches & Cost Reference
      
      The three big cost levers, in order of typical impact: model tiering, prompt
      caching, Batches API. They stack — a cached Haiku batch request can cost ~2-3%
      of an uncached Opus interactive request for the same tokens.
      
      ## Prompt Caching
      
      ### The one invariant
      
      **Caching is a prefix match.** The cache key is the exact bytes of the
      rendered prompt up to each `cache_control` breakpoint. One byte changed
      anywhere in the prefix invalidates everything after it. Render order is
      `tools` → `system` → `messages` — a breakpoint on the last system block caches
      tools + system together.
      
      ### Syntax
      
      ```python
      response = client.messages.create(
          model="claude-opus-5",
          max_tokens=16000,
          system=[{
              "type": "text",
              "text": LARGE_STABLE_PROMPT,                      # 50KB of docs, instructions...
              "cache_control": {"type": "ephemeral"},           # 5-minute TTL (default)
              # "cache_control": {"type": "ephemeral", "ttl": "1h"}  # 1-hour TTL
          }],
          messages=[{"role": "user", "content": question}],
      )
      ```
      
      ```typescript
      const response = await client.messages.create({
        model: "claude-opus-5",
        max_tokens: 16000,
        system: [{ type: "text", text: LARGE_STABLE_PROMPT,
                   cache_control: { type: "ephemeral" } }],
        messages: [{ role: "user", content: question }],
      });
      ```
      
      Simplest option — top-level auto-caching (caches the last cacheable block, no
      per-block markers):
      
      ```python
      client.messages.create(model="claude-opus-5", max_tokens=16000,
                             cache_control={"type": "ephemeral"},
                             system=big_doc, messages=[...])
      ```
      
      Rules:
      
      - Max **4** `cache_control` breakpoints per request.
      - Valid on system text blocks, tool definitions, and message content blocks
        (`text`, `image`, `tool_use`, `tool_result`, `document`).
      - **Minimum cacheable prefix is model-dependent** — below it the marker is
        silently ignored (no error, just `cache_creation_input_tokens: 0`):
      
      | Model | Minimum prefix tokens |
      |---|---:|
      | Opus 4.6 / 4.5, Haiku 4.5 | 4096 |
      | Opus 4.7 | 2048 |
      | Sonnet 5, Opus 4.8, Sonnet 4.6, Sonnet 4.5 | 1024 |
      | Opus 5, Fable 5 | 512 |
      
      Note the shape: the newest models cache *sooner* (512) while Haiku 4.5 — the
      cheapest model, and the one you'd most want to cache in bulk — needs the largest
      prefix (4096). Sizing a shared prefix for Haiku covers every other model too.
      
      ### Pricing & break-even
      
      | Operation | Cost vs base input |
      |---|---|
      | Cache write, 5-min TTL | 1.25x |
      | Cache write, 1-hour TTL | 2x |
      | Cache read | ~0.1x |
      
      Break-even: 5-min TTL pays off at the **2nd** request (1.25 + 0.1 = 1.35x vs
      2x); 1-hour TTL needs **3+** requests (2 + 0.2 = 2.2x vs 3x). Steady-state
      savings on a large cached prefix approach 90%.
      
      ### Multi-turn / agent placement
      
      - **Multi-turn:** put the breakpoint on the last content block of the latest
        turn; earlier breakpoints remain valid read points, so hits accrue as the
        conversation grows. Top-level auto-caching does this for you.
      - **Shared prefix, varying question:** breakpoint at the end of the *shared*
        part, not the end of the prompt — otherwise every request writes a distinct
        entry and nothing is ever read.
      - **20-block lookback:** a breakpoint searches backward at most 20 content
        blocks for a prior entry. Agent turns adding >20 blocks (many
        tool_use/tool_result pairs) silently miss — add an intermediate breakpoint
        every ~15 blocks in long turns.
      - **Concurrent fan-out:** an entry becomes readable only once the first
        response starts streaming. Fire 1 request, await first token, then fire the
        other N-1 so they read the fresh cache.
      
      ### Silent invalidators (audit checklist)
      
      If `usage.cache_read_input_tokens` stays 0 across identical-prefix requests,
      grep the prompt-assembly path for:
      
      | Pattern | Why it kills the cache |
      |---|---|
      | `datetime.now()` / `Date.now()` in the system prompt | New prefix every request |
      | `uuid4()` / request IDs early in content | Same |
      | `json.dumps(d)` without `sort_keys=True` | Non-deterministic bytes |
      | Per-user IDs interpolated into the system prompt | No cross-user sharing |
      | Conditional system sections (`if flag: system += ...`) | Each flag combo = distinct prefix |
      | Tool set built per-user / unsorted | Tools render at position 0 — full invalidation |
      | Switching models mid-conversation | Caches are model-scoped |
      
      Fix: move volatile content **after** the last breakpoint (or into the latest
      user message), serialize deterministically, freeze the system prompt and tool
      list.
      
      ### Verifying
      
      ```python
      u = response.usage
      print(u.cache_creation_input_tokens)  # written this request (paid 1.25-2x)
      print(u.cache_read_input_tokens)      # served from cache (paid ~0.1x)
      print(u.input_tokens)                 # uncached remainder (full price)
      # total prompt = input + cache_creation + cache_read
      ```
      
      ### What invalidates which level (verified 2026-08-30)
      
      Changes cascade **downward**: the level that changes, and every level after it,
      is invalidated. ✓ = that level's cache survives, ✘ = it is invalidated.
      
      | What changes | Tools | System | Messages | Note |
      |---|:--:|:--:|:--:|---|
      | Tool definitions (names, descriptions, params) | ✘ | ✘ | ✘ | Tools render at position 0 — a full rebuild |
      | Model switch | ✘ | ✘ | ✘ | Caches are model-scoped |
      | Web search toggle | ✓ | ✘ | ✘ | Modifies the system prompt |
      | Citations toggle | ✓ | ✘ | ✘ | Modifies the system prompt |
      | `speed: "fast"` ↔ standard | ✓ | ✘ | ✘ | Invalidates system + messages |
      | `tool_choice` | ✓ | ✓ | ✘ | Affects message blocks only |
      | Images added/removed anywhere | ✓ | ✓ | ✘ | Affects message blocks only |
      | Thinking config (mode, `budget_tokens`) | model-specific | model-specific | ✘ | Always invalidates messages; tools/system too on models that render it ahead of them |
      | `output_config.effort` | model-specific | model-specific | ✘ | Same shape as thinking. Setting effort explicitly to the model's default is equivalent to omitting it and does **not** invalidate |
      | Non-tool results w/ extended thinking | ✓ | ✓ | model-specific | Opus 4.5+ / Sonnet 4.6+ preserve thinking blocks (✓); earlier Opus/Sonnet and all Haiku strip them, dropping following messages from cache |
      
      The trap: `tool_choice`, images, thinking and effort are all commonly toggled
      per-request and all invalidate the **messages** cache every time. In a long agent
      loop, where the messages cache holds most of the tokens, "just flipping
      `tool_choice`" is not free — keep it constant across a conversation.
      
      Source: `https://platform.claude.com/docs/en/build-with-claude/prompt-caching.md`
      
      ## Batches API
      
      `POST /v1/messages/batches` — asynchronous Messages requests at a **flat 50%
      discount on all token usage** (stacks with prompt caching).
      
      | Fact | Value |
      |---|---|
      | Max batch size | 100,000 requests or 256 MB |
      | Turnaround | Usually <1 hour; max 24h |
      | Results retention | 29 days |
      | Feature support | All Messages features (tools, vision, caching, structured outputs, thinking) — no streaming |
      | Rate limits | Separate pool from interactive traffic |
      
      ```python
      import anthropic, time
      from anthropic.types.message_create_params import MessageCreateParamsNonStreaming
      from anthropic.types.messages.batch_create_params import Request
      
      client = anthropic.Anthropic()
      
      # 1. Create
      batch = client.messages.batches.create(requests=[
          Request(custom_id=f"item-{i}",
                  params=MessageCreateParamsNonStreaming(
                      model="claude-haiku-4-5", max_tokens=64,
                      messages=[{"role": "user",
                                 "content": f"Classify sentiment (one word): {text}"}]))
          for i, text in enumerate(texts)
      ])
      
      # 2. Poll
      while True:
          batch = client.messages.batches.retrieve(batch.id)
          if batch.processing_status == "ended":
              break
          time.sleep(60)
      print(batch.request_counts)  # succeeded / errored / canceled / expired
      
      # 3. Results (order not guaranteed — key on custom_id)
      for result in client.messages.batches.results(batch.id):
          if result.result.type == "succeeded":
              msg = result.result.message
              text = next((b.text for b in msg.content if b.type == "text"), "")
          elif result.result.type == "errored":
              # error.type == "invalid_request" → fix and resubmit; otherwise safe to retry
              ...
      ```
      
      ```typescript
      const batch = await client.messages.batches.create({
        requests: [{
          custom_id: "request-1",
          params: { model: "claude-sonnet-5", max_tokens: 1024,
                    messages: [{ role: "user", content: "Summarize..." }] },
        }],
      });
      // poll batches.retrieve(batch.id) until processing_status === "ended"
      for await (const result of await client.messages.batches.results(batch.id)) {
        if (result.result.type === "succeeded") { /* result.result.message */ }
      }
      ```
      
      Batch gotchas:
      
      - `custom_id` is your only join key — results stream in completion order, not
        submission order.
      - Result types: `succeeded` | `errored` | `canceled` | `expired`. Resubmit
        `expired`; inspect `errored` (validation vs server error).
      - Cancel is async: `batches.cancel(id)` → status `"canceling"`; some requests
        may still complete.
      - Caching inside batches works, but hit rates are best-effort (requests run
        concurrently) — put a shared cached `system` block on every request and use
        the 1-hour TTL.
      - Per-request params are full `MessageCreateParams` minus `stream`.
      
      Use batches for: eval suites, backfills, bulk extraction/classification,
      nightly report generation, regenerating embeddings-adjacent metadata — any
      workload where minutes-to-hours latency is fine.
      
      ## Model Tiering Economics
      
      Worked example, 10M input + 1M output tokens/day:
      
      | Strategy | Cost/day |
      |---|---|
      | Everything Opus 5 | 10×$5 + 1×$25 = **$75** |
      | Everything Sonnet 5 | 10×$2 + 1×$10 = **$30** |
      | Route: 80% Haiku, 15% Sonnet, 5% Opus | input 8×$1 + 1.5×$2 + 0.5×$5 = $13.5; output 0.8×$5 + 0.15×$10 + 0.05×$25 = $6.75 = **~$20** |
      | Same + cached system prompts (70% of input cached at ~0.1x) | input → ~$5 = **~$12** |
      | Same + batchable share moved to Batches | **lower still (50% off that share)** |
      
      Patterns:
      
      - **Router**: a Haiku call classifies difficulty, dispatches to the right model.
      - **Cascade**: try Haiku; escalate to Sonnet/Opus only when confidence is low
        or validation fails (works well with structured outputs as the validator).
      - **Subagents**: keep the orchestrator on Opus, push parallel/simple legs to
        Haiku/Sonnet. Spawn a separate request per subagent (don't switch models
        mid-conversation — it kills the cache).
      - **Effort tuning**: on supported models, dropping `output_config.effort` from
        `high` to `medium` often cuts output tokens substantially at minor quality
        cost — cheaper than switching model tier for borderline workloads.
      
      ## Token Counting for Cost Estimation
      
      ```python
      count = client.messages.count_tokens(
          model="claude-opus-5", system=system, tools=tools, messages=messages)
      est_input_cost = count.input_tokens * 5.00 / 1_000_000   # Opus 5 input rate
      ```
      
      - Free endpoint; counts include tools and system.
      - Tool use adds a hidden system prompt (~290-800 tokens depending on model and
        `tool_choice`).
      - Token counts are model-specific — re-baseline when migrating models; don't
        apply blanket multipliers.
      
    • compaction.md 13.7 KB
      # Compaction & Context Editing Reference
      
      > **What this file owns:** the decision to *discard or rewrite* context — when it is
      > justified, what each option costs, and the first-party APIs that do it
      > (`context_management` edits, the memory tool).
      >
      > **Adjacent files:** [context-engineering.md](context-engineering.md) owns the budget
      > and the three tiers; [caching-and-cost.md](caching-and-cost.md) owns cache mechanics.
      >
      > Facts verified against platform.claude.com and anthropic.com **2026-08-30**.
      
      ---
      
      ## 1. The headline: compaction is not the default
      
      **Compaction** = summarising the message history and reinitiating a context window
      from the summary.
      
      The instinct is that long conversations must be summarised or they get expensive and
      the model gets confused. **Under modern prompt caching that instinct is measurably
      wrong more often than it is right.**
      
      A 2026 evaluation of a production AI tutor (Bouchard, Solano & Vaid, Towards AI —
      660 turns across 11 configurations) compared keeping the full history against a
      range of compaction strategies (their production preset of clearing + summarisation,
      context reset, prompt compression, selective retention):
      
      | Strategy | Fact recall at turn 11 | Cost / turn | Time to first token |
      |---|---:|---:|---:|
      | **Full history (keep everything)** | **92–100%** | **$0.11** | **17 s** |
      | Production preset (clear + summarise) | 38–58% | $0.24 | 21 s |
      | Context reset | 17% | — | — |
      
      Keep-everything won **cost, latency and recall simultaneously** — the three axes
      compaction is usually reached for. Summarisation was more than twice the cost per
      turn, slower to first token, and forgot roughly half of a fact planted ten turns
      earlier.
      
      **Read the caveats, they matter:**
      
      - The study ran on **Gemini 3.5 Flash via LangChain**, not Claude. What transfers is
        the *mechanism* — a cached prefix is billed at a fraction of base input, while
        summarising rewrites that prefix and forfeits the discount. Claude's cache economics
        make the mechanism *stronger*, not weaker (§3). The specific percentages are theirs,
        not a Claude benchmark.
      - "Keep everything" is not "ignore context entirely." The same study found that
        **capping tool outputs at a fixed size cut cost per turn 38% with no recall loss** —
        because a cap shrinks context *without rewriting the cached prefix*. Shaping what
        enters context is cheap; rewriting what is already there is not.
      
      **The rule:** compaction is a deliberate response to a **named constraint**, not a
      reflex. If you cannot name which of the three constraints in §2 you are hitting, do
      not compact — cap tool output and append instead.
      
      ---
      
      ## 2. The three constraints that justify compaction
      
      Name one before you compact. Each has a different remedy, and reaching for the wrong
      one is how teams pay summarisation costs to solve a problem summarisation does not fix.
      
      ### Constraint A — Context ceiling
      
      *The conversation will not fit.* The hard one; it is not negotiable by budget.
      
      - **Diagnose:** projected tokens at turn N exceed the model's window (1M on the
        current flagships, 200K on Haiku 4.5 — see the model table in `SKILL.md`).
      - **Cheapest remedies first:** cap tool outputs → move payloads to files (tier 2) →
        server-side clearing (§3) → summarising compaction (§5).
      - **Note:** a 1M window makes this constraint *rare* for conversational work and
        still common for long agentic runs, where tool results dominate.
      
      ### Constraint B — Cost ceiling
      
      *It fits, but the bill is unacceptable.* This is the constraint most often assumed
      and least often real, because the comparison is usually made against **uncached**
      costs.
      
      - **Diagnose honestly:** compare `cache_read_input_tokens × 0.1 × base_rate` against
        the cost of a summarisation call **plus** the cache write it forces on the next
        request. Do the arithmetic before assuming (§3).
      - **Cheapest remedies first:** verify the cache is actually hitting (a zero
        `cache_read_input_tokens` is a bug, not a reason to compact) → tier the model down →
        cap tool outputs → 1-hour TTL if the gaps between turns are the problem.
      
      ### Constraint C — Latency target
      
      *Time-to-first-token is too slow at depth.* Real, and the weakest case for
      summarisation.
      
      - **Diagnose:** measure TTFT against context length; confirm the growth is in the
        prefix rather than in thinking or tool round-trips.
      - **Caution:** the tutor study measured compaction as *slower* (21 s vs 17 s) — the
        summarisation call is itself a round-trip, and the next request re-writes the cache.
        Compaction buys latency only when it is amortised over many subsequent turns.
      - **Cheapest remedies first:** cap tool outputs → tier down → reduce `effort` →
        stream (TTFT is a streaming problem before it is a context problem).
      
      ### What each option costs you
      
      | Option | Recall | Cache | Latency | Reversible? |
      |---|---|---|---|---|
      | **Append (do nothing)** | Full | Preserved | Grows with length | n/a |
      | **Cap tool output at the boundary** | Loses untruncated detail only | **Preserved** | Improves | No, but loss is bounded and predictable |
      | **Write payload to file, keep the path** | Full (one round-trip away) | Preserved | One extra hop when read | Yes |
      | **Server-side clearing** (`context_management`) | Cleared results unrecoverable to the model unless saved to memory | **Invalidated at the clear point** | Improves after the re-write | No |
      | **Summarising compaction** | Lossy, and the loss is unpredictable | **Invalidated wholesale** | One extra call now, faster later | No |
      
      Read that table top-down: the first three preserve the cache, and the two that break
      it are the two people reach for first.
      
      ---
      
      ## 3. The break-even arithmetic
      
      Claude's cache read is **0.1× the base input rate** on every model; writes are
      **1.25×** (5-minute TTL) or **2×** (1-hour TTL).
      
      So the per-turn cost of carrying an N-token history is:
      
      ```
      append   :  N × 0.1 × base_rate                    (steady state, cache hit)
      compact  :  summarisation call (input ≈ N, output ≈ S)
               +  S × 1.25 × base_rate                   (cache write of the new prefix)
               +  S × 0.1  × base_rate  per later turn   (steady state on the summary)
               +  everything before the rewrite, forfeited
      ```
      
      Compaction only wins once `0.1 × N` exceeds `0.1 × S` by enough to repay the
      summarisation call **and** the forfeited prefix — i.e. when the history is very large,
      the summary is very small, and the session continues for many turns afterwards. A
      compaction just before the conversation ends is pure loss.
      
      **The published threshold, adapted to Claude.** The tutor study put the crossover at
      a *cached input price above ~$0.55 per million tokens* — above that, summarisation
      starts to compete. Because a Claude cache read is 0.1× base input, that translates to:
      
      ```
      cache-read rate  =  0.1 × base input rate
      crossover        ≈  $0.55 / MTok cached
                       →  base input rate  ≳  $5.50 / MTok
      ```
      
      Derive it from the current price table in `SKILL.md` rather than memorising a model
      list: a model whose **base input rate is under ~$5.50/MTok has cache reads cheap
      enough that appending stays ahead**, and only the premium tier approaches the
      crossover. Re-run this arithmetic when prices change — it is two multiplications.
      
      This is why the honest first question is never "should we compact?" but **"is the
      cache hitting?"** A broken cache makes appending look 10× worse than it is and makes
      compaction look like the fix, when the actual fix is a stray timestamp in the system
      prompt.
      
      ---
      
      ## 4. First-party option: server-side context editing
      
      `context_management` clears content **server-side, before the prompt reaches the
      model**. Your client keeps the full unmodified history — there is no client state to
      synchronise.
      
      Beta header: `context-management-2025-06-27`.
      
      ### Strategy: clear tool uses
      
      ```python
      response = client.beta.messages.create(
          model="claude-opus-5",
          max_tokens=4096,
          messages=messages,
          tools=tools,
          betas=["context-management-2025-06-27"],
          context_management={"edits": [{
              "type": "clear_tool_uses_20250919",
              "trigger":       {"type": "input_tokens", "value": 30000},  # when to fire
              "keep":          {"type": "tool_uses",    "value": 3},      # recent pairs kept
              "clear_at_least":{"type": "input_tokens", "value": 5000},   # min cleared per fire
              "exclude_tools": ["web_search"],                            # never clear these
          }]},
      )
      ```
      
      | Parameter | Default | Meaning |
      |---|---|---|
      | `trigger` | 100,000 input tokens | Threshold that activates clearing (`input_tokens` or `tool_uses`) |
      | `keep` | 3 tool uses | Most recent tool-use/result pairs preserved |
      | `clear_at_least` | none | Minimum tokens cleared per activation — **the cache-economics knob** |
      | `exclude_tools` | none | Tools whose results are never cleared |
      | `clear_tool_inputs` | `false` | Also clear the tool *call parameters*, not just results |
      
      Oldest results go first; cleared content is replaced with a placeholder.
      
      ### Strategy: clear thinking blocks
      
      ```python
      context_management={"edits": [
          {"type": "clear_thinking_20251015", "keep": {"type": "thinking_turns", "value": 2}},
          {"type": "clear_tool_uses_20250919", "trigger": {"type": "input_tokens", "value": 50000}},
      ]}
      ```
      
      `keep` takes `{"type": "thinking_turns", "value": N}` or `"all"` (maximises cache
      hits). When combining strategies, **thinking-block clearing must be listed first**.
      
      Default behaviour differs by model generation: Opus 4.5+ and Sonnet 4.6+ keep all
      prior thinking; earlier Opus/Sonnet and all Haiku models keep only the last turn.
      
      ### The cache interaction — the part that decides whether this pays
      
      - **Clearing invalidates the cached prefix from the clear point onward.** Each
        activation costs a cache write on the next request.
      - Therefore **set `clear_at_least`**. Without it, a trigger can fire and clear a
        trivial amount, paying a full cache re-write to save a few hundred tokens. It exists
        precisely to make each cache break worth taking.
      - Thinking-block clearing is the mirror image: **keeping** thinking preserves the
        cache; **clearing** it invalidates at the clearing point. `"all"` is the
        cache-optimal setting.
      
      ### Inspecting and previewing
      
      Responses report what was applied:
      
      ```json
      {"context_management": {"applied_edits": [
        {"type": "clear_tool_uses_20250919", "cleared_tool_uses": 8, "cleared_input_tokens": 50000}
      ]}}
      ```
      
      Preview before committing to a configuration — `count_tokens` accepts the same
      `context_management` block and returns both the post-clearing `input_tokens` and
      `context_management.original_input_tokens`.
      
      ### Reported results
      
      Anthropic's internal agentic-search evaluation: **context editing alone +29% over
      baseline; context editing with the memory tool +39%.** In a 100-turn web-search
      evaluation, context editing let agents complete workflows that would otherwise fail
      on context exhaustion, **reducing token consumption by 84%**.
      
      Note the shape of that claim — the gains are on *long-horizon agentic search*, where
      Constraint A is genuinely binding. It is not evidence that clearing helps a
      twelve-turn conversation.
      
      ---
      
      ## 5. First-party option: the memory tool
      
      `memory_20250818` gives Claude a directory of files it can create, read, update and
      delete, persisting **across** conversations.
      
      ```python
      tools=[{"type": "memory_20250818", "name": "memory"}],
      context_management={"edits": [{"type": "clear_tool_uses_20250919"}]}
      ```
      
      The pairing is the point: Claude is warned as context approaches a clearing
      threshold, and can **write the durable conclusion to memory before the raw material is
      cleared**. Clearing without memory throws information away; clearing with memory
      demotes it from tier 1 to tier 2
      (→ [context-engineering.md](context-engineering.md) §2).
      
      ---
      
      ## 6. If you must compact: do it well
      
      When a named constraint genuinely demands summarisation:
      
      - **Maximise recall first, then tune precision.** Anthropic's guidance is to start
        with a compaction prompt that captures every relevant piece of information, then
        iterate to trim. A compaction prompt tuned for brevity first will silently drop the
        one detail the next 40 turns needed.
      - **Compact at a natural boundary** — a finished sub-task, a landed change — not at an
        arbitrary token count mid-reasoning.
      - **Keep the last N turns verbatim** alongside the summary. Recent turns are where
        reference resolution ("that file", "the second one") lives.
      - **Write the durable facts to a file first** (memory tool or `NOTES.md`), so the
        summary is a convenience rather than the sole record.
      - **Compact once, deep** rather than repeatedly and shallowly. Every pass is a
        cache write and a lossy re-encoding; summaries of summaries degrade fast.
      - **Measure it.** Plant a fact early, probe for it later, and compare against the
        keep-everything baseline on *your* workload. The tutor study's headline result is
        that the baseline is much stronger than teams assume — including, quite possibly,
        yours.
      
      ---
      
      ## 7. Sources
      
      - Anthropic — *Effective context engineering for AI agents*:
        `https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents`
      - Anthropic — *Managing context on the Claude Developer Platform* (29% / 39% / 84%):
        `https://claude.com/blog/context-management`
      - Context editing API: `https://platform.claude.com/docs/en/build-with-claude/context-editing`
      - Memory tool: `https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool`
      - Prompt caching (multipliers, prefix rules):
        `https://platform.claude.com/docs/en/build-with-claude/prompt-caching.md`
      - Bouchard, Solano & Vaid (Towards AI), *Context Engineering in 2026: Why We Stopped
        Compacting Our Agent's Context*, AI Engineer World's Fair, Aug 2026:
        `https://www.louisbouchard.ai/context-engineering-2026/`
      
    • context-engineering.md 15.4 KB
      # Context Engineering Reference
      
      > **What this file owns:** the discipline of deciding *what the model sees on every
      > call* — the context budget, the three placement tiers, and the agentic-loop
      > specifics (tool-result bloat, sub-agents as isolation, note-taking).
      >
      > **Adjacent files, so this one doesn't duplicate them:**
      > [compaction.md](compaction.md) owns the *decision to discard* context and the
      > `context_management` / memory-tool APIs.
      > [caching-and-cost.md](caching-and-cost.md) owns prompt-cache *mechanics*
      > (breakpoints, TTLs, minimum prefixes, invalidation table).
      >
      > Facts verified against platform.claude.com and anthropic.com **2026-08-30**.
      
      ---
      
      ## 1. The frame: context is a budget, not a container
      
      Prompt engineering asks "what do I write in the prompt?". **Context engineering asks
      "what earns a place in the window on *this* call?"** — including everything that lands
      there without you typing it: tool definitions, tool results, retrieved documents,
      prior turns, system reminders, thinking blocks.
      
      Anthropic's framing (*Effective context engineering for AI agents*): the goal is
      "the smallest possible set of high-signal tokens that maximize the likelihood of some
      desired outcome." Prompt engineering is discrete — you write it once. Context
      engineering is **iterative**: it happens on every single inference.
      
      ### Why the budget is real
      
      - **Attention is finite.** Transformers form n² pairwise relationships for n tokens.
        Accuracy degrades as token count grows — the failure mode Anthropic names
        **context rot**. A 1M-token window is a *capacity*, not a *target*.
      - **Training-distribution thinness.** Models have seen far fewer long-range,
        context-wide dependencies than short ones, so long-context reasoning is the
        weakest part of the envelope, not merely the slowest.
      - **Every token is billed on every turn.** In a multi-turn loop the same prefix is
        re-sent each call. Uncached, a 50K-token prefix over 40 turns is 2M input tokens.
        (Cached, it is ~10% of that — see §4, and it is exactly why the compaction
        instinct is often wrong. → [compaction.md](compaction.md).)
      
      **Working rule.** Before adding anything to a prompt, ask: *does this change the
      model's next action?* If not, it belongs on disk or behind a retrieval call, not in
      the window.
      
      ---
      
      ## 2. The three tiers
      
      Every piece of information the agent could use lives in exactly one of three places.
      Choosing the tier deliberately is most of context engineering.
      
      | Tier | Where it lives | Latency to use | Cost profile | Use when |
      |---|---|---|---|---|
      | **1 — In context** | Rendered into `tools` / `system` / `messages` every call | Zero | Paid every turn (≈0.1× when cached) | It steers *most* turns: the task spec, the invariants, the current working set |
      | **2 — On disk, read on demand** | A file the agent can `Read`/`Grep` when it decides to | One tool round-trip | Paid only when read | It steers *some* turns and the agent can tell when it needs it from the filename alone |
      | **3 — Retrieved** | Index / vector store / API behind a search tool | One+ round-trips, plus a relevance gamble | Paid only when hit | The corpus is too large to enumerate and relevance is query-dependent |
      
      ### Choosing between tiers
      
      - **Tier 1 is the expensive tier.** It costs on every call whether or not the turn
        needed it. Reserve it for what is load-bearing on the majority of turns.
      - **Tier 2 is the default for "might need it".** Anthropic's just-in-time framing:
        let the agent maintain lightweight identifiers (file paths, queries, links) and
        hydrate them at runtime. A path is ~10 tokens; the file it names may be 10,000.
        The trade-off is honest — "runtime exploration is slower than retrieving
        pre-computed data."
      - **Tier 3 pays for scale with a relevance risk.** If retrieval misses, the model
        does not know what it did not see. Prefer tier 2 whenever the candidate set is
        small enough to name.
      - **Hybrids win in practice.** Pre-load the handful of things that steer every turn
        (tier 1), leave the long tail addressable (tier 2/3).
      
      ### The demotion test
      
      When a prompt is too big, demote rather than delete. For each block ask:
      
      1. Does it change the next action on **most** turns? → stays tier 1.
      2. Can the agent tell from a **name** that it needs this? → tier 2 (write it to a
         file, leave the path in context).
      3. Neither, but it is occasionally essential? → tier 3 (index it, ship a search tool).
      
      Deleting outright is the fourth option and the only irreversible one.
      
      ---
      
      ## 3. Progressive disclosure — this repo already runs on it
      
      **Claude Code skills are a tier-1/tier-2 split you can read on disk.** That is not an
      analogy; it is the same mechanism:
      
      | Layer | Tier | Loaded |
      |---|---|---|
      | Skill `name` + `description` frontmatter | 1 | Always — every skill's description sits in the session prompt |
      | `SKILL.md` body | 2 | When the router matches the description |
      | `references/*.md` (this file) | 2 | Only when `SKILL.md` cites it and the task needs it |
      | `scripts/*` output | 2 | Only when executed |
      
      This is why the repo's own rules read the way they do, and the rules are context
      engineering rules wearing authoring clothes:
      
      - **"The `description` is the trigger"** — the description is the only always-resident
        text, so it must carry the routing signal and nothing else.
      - **"Keep the body under 500 lines"** — the body is what gets pulled in wholesale on a
        match; an oversized body spends the budget of every task that touches the skill.
      - **"One concept per reference file"** — a file is the unit of loading. Two concepts
        in one file means loading both to get either.
      - **"Every reference must be cited from `SKILL.md`"** — an uncited file is
        unreachable. In tier terms: a tier-2 artefact with no pointer in tier 1 does not
        exist.
      
      The same shape generalises to any agent you build: a small always-on spec, a set of
      named-and-addressable documents, and a retrieval path for the long tail.
      See [SKILL-CREATION-PROTOCOL.md](../../../docs/SKILL-CREATION-PROTOCOL.md) Step 3 and
      [SKILL-RESOURCE-PROTOCOL.md](../../../docs/SKILL-RESOURCE-PROTOCOL.md) §1.
      
      ---
      
      ## 4. Cache-aware prompt architecture
      
      The cache is what makes a large tier-1 affordable — and the cache is a **prefix
      match**, so *ordering is architecture*.
      
      ### The layout rule
      
      Requests render in a fixed order — `tools` → `system` → `messages` — and each level
      builds on the previous. Therefore:
      
      ```
      tools        ─┐
      system        ├─ static, byte-identical across calls   ← cache_control breakpoint here
      (stable docs) ┘
      messages      ← volatile: the turn, the retrieved chunk, the timestamp
      ```
      
      **Static prefix first, volatile content last.** Anything that varies per request must
      sit *after* the last breakpoint, or it changes the prefix bytes and every cached
      token behind it is re-billed at write price.
      
      ### Why reordering silently destroys the cache
      
      There is no error. Moving a block, adding a conditional system section, or
      interpolating a user id early in the prompt produces a *different prefix hash*, which
      is simply a miss: `cache_read_input_tokens: 0`, `cache_creation_input_tokens: <all of
      it>`, a 1.25–2× bill instead of 0.1×, and no diagnostic anywhere. The only signal is
      the usage block — which is why "assert `cache_read_input_tokens > 0` in staging" is a
      real test, not a nicety.
      
      Two ordering traps specific to agent loops:
      
      - **A breakpoint searches backward at most 20 content blocks.** A turn that appends
        more than 20 blocks (many `tool_use`/`tool_result` pairs) jumps the window and
        silently misses. Add an intermediate breakpoint roughly every 15 blocks in long
        turns.
      - **A cache entry only becomes readable once the first response begins streaming.**
        Fanning out N parallel requests against a cold shared prefix writes N entries and
        reads none. Fire one, await first token, then fire the rest.
      
      The invalidation table, per-model minimum prefixes, TTL pricing and the full
      silent-invalidator checklist live in
      [caching-and-cost.md](caching-and-cost.md) — this section is the *shape* rule only.
      
      ### The consequence for compaction
      
      Rewriting history to make it shorter changes the prefix. A summarisation pass
      therefore pays a full cache write on the next call *and* discards every cached token
      before it. That is the mechanism behind the counter-intuitive finding in
      [compaction.md](compaction.md): under caching, the cheap thing is usually to
      **append**, not to **rewrite**.
      
      ---
      
      ## 5. Multi-turn and agentic specifics
      
      ### 5.1 Tool results are where the budget actually goes
      
      In an agent loop, the growth term is almost never the system prompt — it is tool
      output. A file read, a search result, an HTTP response: each lands verbatim and stays
      for the rest of the session.
      
      **Design tools to return decisions, not dumps.** Anthropic's guidance: tools should
      be "self-contained, robust to error, and extremely clear with respect to their
      intended use", with "minimal overlap in functionality". A bloated tool set costs
      tokens at position 0 *and* creates ambiguous decision points.
      
      Concrete levers, cheapest first:
      
      | Lever | What it does | Cost |
      |---|---|---|
      | **Cap tool output at a fixed size** | Truncate/paginate at the tool boundary, before the result enters context | Free — and critically, it **shrinks context without rewriting the prefix**, so the cache survives |
      | **Return a handle, not the payload** | Tool writes to a file, returns the path + a 200-token précis | One extra round-trip if the agent needs the full text |
      | **Filter at the source** | `grep`-shaped tools instead of `cat`-shaped ones | Design-time only |
      | **Clear stale results** | `context_management` server-side clearing | Invalidates the cache at the clear point → [compaction.md](compaction.md) |
      
      The capping lever is the one to reach for first: a Towards AI evaluation
      (AI Engineer World's Fair, August 2026) measured **38% lower cost per turn** from
      fixed-size tool-output caps alone, with no loss of recall — precisely because it does
      not touch the cached prefix.
      
      ### 5.2 Summarise a tool result, or write it to a file?
      
      | Situation | Do this |
      |---|---|
      | Result is large and needed **later, in full** (a fetched spec, a big file) | **Write to a file**, return the path. Lossless, addressable, and the path costs ~10 tokens per turn |
      | Result is large and only its **conclusion** matters (a 300-row query → "4 rows failed") | **Summarise at the tool boundary** — before it enters context, not after |
      | Result is large, and you cannot tell which parts matter yet | **Both**: file for fidelity, précis in context, path in the précis |
      | Result is small, or is the thing the user asked for | **Leave it alone.** Compression has a floor; do not spend a round-trip to save 200 tokens |
      
      The distinction that matters: **summarising at the tool boundary is free of cache
      cost** (the shorter result is what gets appended, and nothing before it moves).
      Summarising *after the fact* — rewriting history that is already in the prefix — is
      compaction, and pays the full cache-write penalty.
      
      ### 5.3 Structured note-taking (persistent memory)
      
      Have the agent maintain an external `NOTES.md` / progress file and pull it back in
      when needed — "persistent memory with minimal overhead". It is a tier-2 artefact that
      the agent itself authors, and it survives context resets, which is what makes it the
      natural companion to any clearing strategy: write the durable conclusion out
      *before* the raw material is cleared. The first-party version of this is the memory
      tool → [compaction.md](compaction.md) §4.
      
      ### 5.4 Sub-agents as context isolation
      
      A sub-agent is not primarily a parallelism device — **it is a second context window
      whose contents never touch yours.**
      
      ```
      orchestrator context:  task spec + plan + N × (≈1-2K token summary)
            ↓ spawn                                   ↑ return
      sub-agent context:     the 80K tokens of exploration nobody else needs
      ```
      
      The exploration — dozens of file reads, failed greps, dead ends — is billed once,
      inside the sub-agent, and then discarded. The orchestrator sees only the distilled
      result; Anthropic's guidance puts that return payload at roughly **1,000–2,000
      tokens**.
      
      Use it when:
      
      - A subtask generates far more intermediate context than conclusion (search, triage,
        audit, "find where X is implemented").
      - You want a genuinely independent opinion — a fresh window cannot be primed by the
        orchestrator's earlier wrong turn. (Adversarial verification depends on this.)
      - The subtask's tool set is large and irrelevant to the main loop — it renders at
        position 0 in the sub-agent's prompt, not yours.
      
      Do **not** use it when the subtask needs most of the orchestrator's context to make
      sense: you will pay to reconstruct that context in the child, and lose fidelity in
      the hand-off. The hand-off is a lossy channel by design; if the summary has to carry
      everything, isolation was the wrong tool.
      
      Costs to price in: the sub-agent's prefix is a **cold cache** (a fresh window shares
      nothing with the parent), the hand-off is lossy, and errors are harder to attribute.
      Tier the model down for the isolated leg — an Opus orchestrator with Haiku/Sonnet
      sub-agents is the standard shape (see [caching-and-cost.md](caching-and-cost.md),
      "Model Tiering Economics").
      
      ---
      
      ## 6. Instrumentation — what to measure
      
      You cannot engineer a budget you cannot see. Log per request:
      
      | Signal | Source | What it tells you |
      |---|---|---|
      | `usage.input_tokens` | response | Uncached remainder — the part after your last breakpoint |
      | `usage.cache_read_input_tokens` | response | Cache is working. **Zero across identical-prefix calls = a silent invalidator** |
      | `usage.cache_creation_input_tokens` | response | What you paid write price for this turn |
      | `usage.output_tokens` | response | The other half of the bill |
      | Pre-flight estimate | `client.messages.count_tokens(...)` | Free; counts tools + system. Never `tiktoken` (OpenAI's tokenizer, 15–20% undercount on Claude) |
      | Growth per turn | your own diff | Which tool is the growth term — almost always one of them dominates |
      
      The single most valuable alarm: **`cache_read_input_tokens == 0` on a request whose
      prefix should be unchanged.** It is the only symptom a broken cache produces.
      
      ---
      
      ## 7. Cross-references
      
      | Concern | Where |
      |---|---|
      | Compaction decision, `context_management`, memory tool | [compaction.md](compaction.md) |
      | Cache breakpoints, TTLs, minimums, invalidation table, batches | [caching-and-cost.md](caching-and-cost.md) |
      | Tool definitions, agentic loop, `tool_result` mechanics | [tool-use.md](tool-use.md) |
      | Sub-agents in the Agent SDK (`agents` option) | [agent-sdk.md](agent-sdk.md) |
      | Claude Code's own context surface (CLAUDE.md, skills, hooks) | `claude-code-ops` skill |
      | Scheduled/autonomous loops that re-send a prompt on a cadence | `loop-ops` skill |
      | Cross-provider fleets and adversarial verify | `fleetflow` skill |
      
      ## 8. Sources
      
      - Anthropic — *Effective context engineering for AI agents*:
        `https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents`
      - Anthropic — *Managing context on the Claude Developer Platform*:
        `https://claude.com/blog/context-management`
      - Prompt caching (ordering, 20-block lookback, concurrency):
        `https://platform.claude.com/docs/en/build-with-claude/prompt-caching.md`
      - Context editing: `https://platform.claude.com/docs/en/build-with-claude/context-editing`
      - Bouchard, Solano & Vaid (Towards AI), *Context Engineering in 2026*, AI Engineer
        World's Fair, Aug 2026: `https://www.louisbouchard.ai/context-engineering-2026/`
      
    • messages-api.md 11.7 KB
      # Messages API Reference
      
      `POST https://api.anthropic.com/v1/messages` — the single endpoint everything
      runs through. Tools, structured outputs, thinking, and caching are all features
      of this endpoint, not separate APIs.
      
      ## Required Headers
      
      | Header | Value |
      |---|---|
      | `x-api-key` | Your API key (`sk-ant-...`) |
      | `anthropic-version` | `2023-06-01` |
      | `content-type` | `application/json` |
      | `anthropic-beta` | Comma-separated beta IDs, only for beta features |
      
      OAuth bearer tokens go on `Authorization: Bearer <token>` instead of
      `x-api-key` (plus `anthropic-beta: oauth-2025-04-20`). Setting both
      `ANTHROPIC_API_KEY` and `ANTHROPIC_AUTH_TOKEN` makes the SDK send both headers
      and the API rejects the request.
      
      ## Request Parameters
      
      | Param | Type | Required | Notes |
      |---|---|---|---|
      | `model` | string | yes | Exact alias ID, e.g. `claude-opus-5` — no date suffixes |
      | `max_tokens` | int | yes | Hard output cap. Default sensibly: ~16000 non-streaming, ~64000 streaming, ~256 classification |
      | `messages` | array | yes | Alternating `user`/`assistant` turns; first must be `user`. Consecutive same-role messages are merged |
      | `system` | string \| block[] | no | System prompt. Block-list form required for `cache_control` |
      | `tools` | array | no | Custom + server tool definitions (see tool-use.md) |
      | `tool_choice` | object | no | `auto` (default) / `any` / `tool` / `none` |
      | `thinking` | object | no | Adaptive; **on by default** on Fable 5 / Opus 5 / Sonnet 5. `{"type": "enabled", "budget_tokens": N}` 400s on Opus 4.7+ |
      | `output_config` | object | no | `{"effort": "...", "format": {...}, "task_budget": {...}}` |
      | `stop_sequences` | string[] | no | Custom stop strings |
      | `stream` | bool | no | SSE streaming |
      | `metadata` | object | no | `{"user_id": "..."}` — opaque end-user id for abuse detection |
      | `temperature` / `top_p` / `top_k` | number | no | **Removed on Opus 4.7 and later (400)** — Opus 5, Sonnet 5, Fable 5 included. On earlier 4.x: at most one of temperature/top_p |
      | `cache_control` | object | no | Top-level auto-caching: caches the last cacheable block |
      | `container` | string | no | Reuse a code-execution container id |
      | `mcp_servers` | array | no | Remote MCP connector (beta `mcp-client-2025-11-20`) |
      
      ### Message content blocks
      
      `content` is either a plain string or an array of blocks:
      
      ```json
      {"role": "user", "content": [
        {"type": "text", "text": "What's in this image?"},
        {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": "<b64>"}},
        {"type": "image", "source": {"type": "url", "url": "https://example.com/img.png"}},
        {"type": "document", "source": {"type": "base64", "media_type": "application/pdf", "data": "<b64>"}},
        {"type": "tool_result", "tool_use_id": "toolu_...", "content": "..."}
      ]}
      ```
      
      ## Response Shape
      
      ```json
      {
        "id": "msg_01...",
        "type": "message",
        "role": "assistant",
        "model": "claude-opus-5",
        "content": [
          {"type": "thinking", "thinking": "...", "signature": "..."},
          {"type": "text", "text": "Hello!"},
          {"type": "tool_use", "id": "toolu_01...", "name": "get_weather", "input": {"location": "Paris"}}
        ],
        "stop_reason": "end_turn",
        "stop_sequence": null,
        "usage": {
          "input_tokens": 1024,
          "output_tokens": 256,
          "cache_creation_input_tokens": 0,
          "cache_read_input_tokens": 0
        }
      }
      ```
      
      `content` is a **list of typed blocks** — never index `content[0].text` blindly
      (a `thinking` block may come first). Filter by `.type`.
      
      ## Stop Reasons
      
      | `stop_reason` | Meaning | What to do |
      |---|---|---|
      | `end_turn` | Finished naturally | Done |
      | `max_tokens` | Hit the `max_tokens` cap | Raise the cap or stream; output may be truncated mid-thought |
      | `stop_sequence` | Hit a custom stop string | `stop_sequence` field has which one |
      | `tool_use` | Claude wants tool(s) executed | Execute each `tool_use` block, send `tool_result`(s), re-request |
      | `pause_turn` | Server-side tool loop hit its iteration limit | Append the assistant turn and re-send unchanged — server resumes; do NOT add a "continue" user message |
      | `refusal` | Safety refusal | Check `stop_details` (`category`: "cyber"/"bio"/"reasoning_extraction" (Fable 5)/null, `explanation`); don't retry same prompt |
      | `model_context_window_exceeded` | Context window exhausted (distinct from max_tokens) | Compact, truncate, or split the conversation |
      
      ```python
      if response.stop_reason == "refusal" and response.stop_details:
          print(response.stop_details.category, response.stop_details.explanation)
      ```
      
      ## Multi-Turn Conversations
      
      The API is stateless — send the full history every request:
      
      ```python
      messages = []
      def chat(user_msg: str) -> str:
          messages.append({"role": "user", "content": user_msg})
          r = client.messages.create(model="claude-opus-5", max_tokens=16000, messages=messages)
          # Append the FULL content list (preserves tool_use/thinking/compaction blocks)
          messages.append({"role": "assistant", "content": r.content})
          return next(b.text for b in r.content if b.type == "text")
      ```
      
      For conversations that may exceed context: server-side **compaction** (beta
      header `compact-2026-01-12`, `context_management: {"edits": [{"type":
      "compact_20260112"}]}` on `client.beta.messages.create`). Critical: append
      `response.content` back verbatim — compaction blocks must be preserved or
      state is silently lost.
      
      ## Streaming
      
      ```python
      with client.messages.stream(model="claude-opus-5", max_tokens=64000,
                                  messages=[...]) as stream:
          for text in stream.text_stream:
              print(text, end="", flush=True)
          final = stream.get_final_message()
      print(final.usage.output_tokens)
      ```
      
      ```typescript
      const stream = client.messages.stream({ model: "claude-opus-5", max_tokens: 64000, messages });
      stream.on("text", (delta) => process.stdout.write(delta));
      const final = await stream.finalMessage();   // never wrap .on() in new Promise()
      ```
      
      ### SSE event sequence
      
      ```
      message_start          → message metadata (id, model, usage so far)
      content_block_start    → index + block type (text / thinking / tool_use)
      content_block_delta    → text_delta | thinking_delta | input_json_delta
      content_block_stop     → block finished
      message_delta          → stop_reason + final usage
      message_stop           → stream done
      ```
      
      Tool inputs stream as `input_json_delta` (partial JSON strings) — accumulate
      and parse at `content_block_stop`, or just use `get_final_message()` /
      `finalMessage()` which assembles parsed blocks for you.
      
      **Why stream:** non-streaming requests with large `max_tokens` exceed HTTP
      timeouts (the Python SDK raises `ValueError` for non-streaming requests it
      estimates will run >~10 min). Default to streaming for anything long.
      
      ## Error Handling
      
      | HTTP | `error.type` | Retryable | Typical cause |
      |---|---|---|---|
      | 400 | `invalid_request_error` | no | Bad params: removed sampling params, `budget_tokens` on 4.7+, prefill on 4.7+, modified thinking blocks, role ordering |
      | 401 | `authentication_error` | no | Missing/invalid key; both key + token set |
      | 403 | `permission_error` | no | Key lacks model/feature access |
      | 404 | `not_found_error` | no | Bad model ID (date-suffix mistake) or endpoint |
      | 413 | `request_too_large` | no | Body over size limit — shrink images/history |
      | 429 | `rate_limit_error` | yes | RPM/ITPM/OTPM exceeded — honor `retry-after` |
      | 500 | `api_error` | yes | Transient server issue |
      | 529 | `overloaded_error` | yes | Capacity — backoff; consider another model |
      
      Error envelope:
      
      ```json
      {"type": "error",
       "error": {"type": "rate_limit_error", "message": "..."},
       "request_id": "req_011CSH..."}
      ```
      
      Log `request_id` (also `response._request_id` on SDK success objects) when
      reporting issues to Anthropic.
      
      ### Typed exceptions — never string-match messages
      
      ```python
      import anthropic
      try:
          r = client.messages.create(...)
      except anthropic.BadRequestError as e:      # 400
          raise                                    # don't retry client errors
      except anthropic.RateLimitError as e:       # 429
          wait = int(e.response.headers.get("retry-after", "60"))
      except anthropic.APIStatusError as e:        # catch-all with .status_code / .type
          if e.status_code >= 500: ...             # retryable
      except anthropic.APIConnectionError:
          ...                                      # network — retryable
      ```
      
      ```typescript
      try {
        await client.messages.create({...});
      } catch (err) {
        if (err instanceof Anthropic.RateLimitError) { /* backoff */ }
        else if (err instanceof Anthropic.APIError) { console.error(err.status, err.message); }
      }
      ```
      
      All subclasses expose `.type` (e.g. `"overloaded_error"`) for finer
      classification than the status code (e.g. `billing_error` vs
      `permission_error`, both 403).
      
      ### Retries
      
      The official SDKs **auto-retry** connection errors, 408/409/429 and >=500 with
      exponential backoff — default `max_retries=2`. Configure per client
      (`anthropic.Anthropic(max_retries=5)`) or per call
      (`client.with_options(max_retries=5, timeout=20.0).messages.create(...)`).
      Only hand-roll retry logic when you need behavior beyond that (e.g. queue +
      jitter across many workers):
      
      ```python
      import random, time
      
      def call_with_retry(client, max_retries=5, base=1.0, cap=60.0, **kwargs):
          last = None
          for attempt in range(max_retries):
              try:
                  return client.messages.create(**kwargs)
              except anthropic.RateLimitError as e:
                  last = e
              except anthropic.APIStatusError as e:
                  if e.status_code < 500:
                      raise          # 4xx (except 429) is not retryable
                  last = e
              time.sleep(min(base * 2 ** attempt + random.random(), cap))
          raise last
      ```
      
      Default request timeout is 10 minutes (`timeout=` on the client or
      `with_options`). On timeout: `anthropic.APITimeoutError`, retried per
      `max_retries`.
      
      ## Rate Limits
      
      Limits are per-organization, per-model-class, measured three ways:
      
      - **RPM** — requests per minute
      - **ITPM** — input tokens per minute (cache reads often discounted/exempt — check headers)
      - **OTPM** — output tokens per minute
      
      Tiers scale with cumulative spend (Tier 1-4, then custom/scale). Check live
      limits in Console or response headers:
      
      | Header | Meaning |
      |---|---|
      | `retry-after` | Seconds to wait (on 429) |
      | `anthropic-ratelimit-requests-limit` / `-remaining` / `-reset` | RPM state |
      | `anthropic-ratelimit-input-tokens-*` / `-output-tokens-*` | ITPM / OTPM state |
      
      Practical guidance:
      
      - Treat 429 as backpressure: honor `retry-after`, add jitter, cap concurrency.
      - Long-running agent fleets: budget OTPM, not just RPM — output is usually the
        binding constraint.
      - Batches API has separate, much higher throughput and doesn't draw from
        interactive rate limits — move bulk traffic there.
      - 529 `overloaded_error` is capacity, not your quota — backoff and/or fail over
        to a different model tier.
      
      ## Token Counting
      
      `POST /v1/messages/count_tokens` — free, model-specific, counts a request
      without running it:
      
      ```python
      n = client.messages.count_tokens(
          model="claude-opus-5",
          system=system, tools=tools,
          messages=[{"role": "user", "content": text}],
      ).input_tokens
      ```
      
      Never estimate with `tiktoken` (OpenAI tokenizer; 15-20% undercount on prose,
      worse on code). Token counts differ **between Claude models** too — count
      against the model you'll run.
      
      ## Vision & Documents
      
      - Images: `{"type": "image", "source": {...}}` blocks — base64, URL, or Files
        API `{"type": "file", "file_id": ...}`. Opus 4.7+ supports high-res input
        (up to 2576px long edge, pixel-accurate coordinates; up to ~3x image tokens).
      - PDFs: `{"type": "document", "source": {...}}` — base64, URL, plain text, or
        file_id. Optional `citations: {"enabled": true}`.
      - Files API (beta `files-api-2025-04-14`): upload once
        (`client.beta.files.upload(...)`), reference by `file_id` across requests.
        500 MB/file, 100 GB/org.
      
    • structured-outputs.md 7.4 KB
      # Structured Outputs Reference
      
      Two related features, same constrained-sampling mechanism:
      
      | Feature | Parameter | Constrains |
      |---|---|---|
      | **JSON outputs** | `output_config: {"format": {...}}` | Claude's response text (guaranteed valid JSON matching your schema) |
      | **Strict tool use** | `strict: true` on a tool definition | The `input` of tool calls |
      
      They can be combined in one request. Supported on Fable 5, Opus 5, Sonnet 5, and
      the legacy Opus 4.8/4.7/4.6/4.5 and Sonnet 4.6/4.5 line, plus Haiku 4.5 (the
      compatibility list names its dated snapshot `claude-haiku-4-5-20251001`; the docs'
      own examples use plain aliases throughout, so keep using `claude-haiku-4-5`).
      Where a model lacks support, enforce output shape via system-prompt instructions
      or strict tool use instead.
      
      **Naming:** the canonical parameter is `output_config.format`. The older
      top-level `output_format` parameter (and the `structured-outputs-2025-11-13`
      beta header) is **deprecated** — still accepted during a transition window,
      and still used as a convenience kwarg by some SDK `parse()` methods, but write
      new code against `output_config`.
      
      ## JSON Outputs — raw schema
      
      ```python
      import json, anthropic
      
      client = anthropic.Anthropic()
      response = client.messages.create(
          model="claude-opus-5",
          max_tokens=16000,
          messages=[{"role": "user",
                     "content": "Extract: John Smith (john@example.com) wants the Enterprise plan."}],
          output_config={
              "format": {
                  "type": "json_schema",
                  "schema": {
                      "type": "object",
                      "properties": {
                          "name":  {"type": "string"},
                          "email": {"type": "string", "format": "email"},
                          "plan":  {"type": "string", "enum": ["Free", "Pro", "Enterprise"]},
                      },
                      "required": ["name", "email", "plan"],
                      "additionalProperties": False,
                  },
              }
          },
      )
      text = next(b.text for b in response.content if b.type == "text")
      data = json.loads(text)   # guaranteed valid against the schema (unless refusal/max_tokens)
      ```
      
      cURL shape:
      
      ```json
      {
        "model": "claude-opus-5",
        "max_tokens": 1024,
        "output_config": {
          "format": {"type": "json_schema", "schema": { ... }}
        },
        "messages": [{"role": "user", "content": "..."}]
      }
      ```
      
      ## SDK helpers — `parse()` (recommended)
      
      ```python
      from pydantic import BaseModel
      
      class ContactInfo(BaseModel):
          name: str
          email: str
          plan: str
          demo_requested: bool
      
      response = client.messages.parse(
          model="claude-opus-5",
          max_tokens=16000,
          messages=[{"role": "user", "content": "Extract: Jane Doe (jane@co.com), Enterprise, wants a demo."}],
          output_format=ContactInfo,          # parse() convenience kwarg
      )
      contact = response.parsed_output        # validated ContactInfo instance
      ```
      
      ```typescript
      import { z } from "zod";
      import { zodOutputFormat } from "@anthropic-ai/sdk/helpers/zod";
      
      const ContactInfo = z.object({
        name: z.string(),
        email: z.string(),
        plan: z.string(),
        demo_requested: z.boolean(),
      });
      
      const response = await client.messages.parse({
        model: "claude-opus-5",
        max_tokens: 16000,
        output_config: { format: zodOutputFormat(ContactInfo) },
        messages: [{ role: "user", content: "Extract: ..." }],
      });
      console.log(response.parsed_output!.name);  // null if parsing failed — guard it
      ```
      
      The SDKs strip unsupported schema constraints (e.g. `minLength`) before
      sending and validate them client-side instead.
      
      ## Strict Tool Use
      
      ```python
      tools = [{
          "name": "book_flight",
          "description": "Book a flight",
          "strict": True,
          "input_schema": {
              "type": "object",
              "properties": {
                  "destination": {"type": "string"},
                  "date":        {"type": "string", "format": "date"},
                  "passengers":  {"type": "integer", "enum": [1, 2, 3, 4, 5, 6, 7, 8]},
              },
              "required": ["destination", "date", "passengers"],
              "additionalProperties": False,
          },
      }]
      ```
      
      - Per-tool opt-in; non-strict tools don't count toward complexity limits.
      - Max **20 strict tools** per request.
      - Guarantees the `tool_use.input` validates exactly — no missing required
        fields, no type drift.
      
      ## JSON Schema: supported vs not
      
      **Supported:** object/array/string/integer/number/boolean/null; `enum`
      (scalars only); `const`; `anyOf`/`allOf` (no `allOf` + `$ref` combo); internal
      `$ref`/`$defs`; `default`; `required`; `additionalProperties: false`
      (mandatory on every object); string `format` (`date-time`, `time`, `date`,
      `duration`, `email`, `hostname`, `uri`, `ipv4`, `ipv6`, `uuid`); array
      `minItems` 0 or 1 only; simple regex `pattern`.
      
      **Not supported:** recursive schemas; external `$ref`; numeric constraints
      (`minimum`/`maximum`/`multipleOf`); string length constraints
      (`minLength`/`maxLength`); array constraints beyond `minItems` 0/1;
      regex backreferences, lookahead/lookbehind, `\b`; `additionalProperties`
      anything but `false`.
      
      **Complexity limits:** 20 strict tools; 24 optional parameters total across
      all schemas; 16 union-typed (`anyOf`) parameters; grammar compilation timeout
      180s ("Schema is too complex"). Reduce by flattening nesting, making params
      required, splitting across requests.
      
      ## Operational notes
      
      - **First-request latency:** new schemas compile a grammar on first use;
        cached for 24h (keyed on schema + tool set; name/description changes don't
        invalidate).
      - **Prompt cache interplay:** changing `output_config.format` invalidates the
        prompt cache; the feature also injects an extra system prompt (more input
        tokens).
      - **Failure modes:** `stop_reason: "refusal"` → output may not match the
        schema; `stop_reason: "max_tokens"` → JSON may be truncated/incomplete —
        raise `max_tokens` and check before parsing.
      - **Incompatible with:** citations (400) and assistant-message prefilling.
        **Works with:** batches, streaming, token counting, extended/adaptive
        thinking.
      - Don't put PHI/PII in schema property names, enum values, or patterns —
        schemas are cached separately from ZDR-handled message content.
      
      ## Structured outputs vs tool-use extraction
      
      Before structured outputs existed, the standard extraction trick was a forced
      tool call (`tool_choice: {"type": "tool", "name": "record_result"}`) with the
      target shape as `input_schema`. Decision now:
      
      | Want | Use |
      |---|---|
      | The *final answer* as guaranteed JSON | `output_config.format` |
      | Valid *parameters* for a real action/function | tool + `strict: true` |
      | Extraction under **manual** extended thinking (`type: "enabled"`) | `output_config.format` — forced `tool_choice` is a 400 in that mode |
      | Extraction mid-agentic-loop (model also has other tools) | A strict "report/record" tool keeps the loop uniform |
      | Legacy prefill (`{"name": "` assistant prefill) | Dead on Opus 4.7+ models (400) — migrate to `output_config.format` |
      
      ## Thinking interplay
      
      - `output_config.format` **works with adaptive/extended thinking** — the model
        thinks, then the final text block conforms to the schema.
      - Forced tool extraction is rejected only under **manual** extended thinking
        (`{"type": "enabled"}` + `tool_choice: any/tool` = 400). Adaptive thinking —
        what every current model uses — accepts forced tool choice, so this is no
        longer a reason to avoid it; prefer `output_config.format` because it targets
        the *answer* rather than a tool's arguments.
      - Effort and format coexist in `output_config`:
        `output_config={"effort": "medium", "format": {...}}`.
      
    • tool-use.md 11 KB
      # Tool Use Reference
      
      Tool use is a feature of `POST /v1/messages` — you pass `tools`, Claude
      responds with `tool_use` content blocks, you execute and return `tool_result`
      blocks. **Client tools** run in your code; **server tools** (web_search,
      code_execution, web_fetch) run on Anthropic's infrastructure.
      
      ## Tool Definition
      
      ```json
      {
        "name": "get_weather",
        "description": "Get current weather for a location. Call this when the user asks about weather conditions, temperature, or forecasts.",
        "input_schema": {
          "type": "object",
          "properties": {
            "location": {"type": "string", "description": "City and state, e.g. San Francisco, CA"},
            "unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "description": "Temperature unit"}
          },
          "required": ["location"]
        }
      }
      ```
      
      Rules that actually move the needle:
      
      - **Descriptions are the routing signal.** Be prescriptive about *when* to
        call, not just what it does ("Call this when the user asks about current
        prices or recent events"). Recent Opus models reach for tools more
        conservatively — trigger conditions in the description give measurable lift.
      - Describe every property; use `enum` for closed sets; only truly-required
        params in `required`.
      - Specific names beat generic ones: `get_current_weather` > `weather`.
      - Keep the tool set focused — too many similar tools degrades selection. For
        large tool libraries, use the server-side **tool search tool** (loads schemas
        on demand, preserves the prompt cache by appending rather than swapping).
      - **Sample calls in definitions**: you can include example invocations
        (`input_examples` on supported surfaces / examples embedded in the
        description) to demonstrate parameter formats and cut parameter errors —
        most useful for complex nested schemas.
      - `strict: true` on a tool guarantees the emitted `input` validates against
        the schema exactly (see structured-outputs.md; max 20 strict tools/request).
      
      ## tool_choice
      
      | Value | Behavior |
      |---|---|
      | `{"type": "auto"}` | Claude decides (default when tools are present) |
      | `{"type": "any"}` | Must call at least one tool |
      | `{"type": "tool", "name": "get_weather"}` | Must call that specific tool |
      | `{"type": "none"}` | Cannot call tools (definitions stay in context) |
      
      Any variant accepts `"disable_parallel_tool_use": true` to cap at one tool
      call per response.
      
      Gotchas:
      
      - **Manual extended thinking (`{"type": "enabled"}`) + `any`/`tool` = 400.**
        Only `auto`/`none` are compatible with it. **Adaptive thinking has no such
        limit** — including the models where it is on by default (Fable 5, Opus 5,
        Sonnet 5), forced tool choice is accepted.
      - `any`/`tool` add more tool-use system-prompt tokens than `auto`/`none` (a
        measured ~410 vs ~290 on Opus 4.8; re-measure with `count_tokens` rather than
        carrying the figure across model generations).
      - Changing `tool_choice` between requests does **not** invalidate the
        tools+system prompt cache (message cache only).
      
      ## Parallel Tool Use
      
      By default Claude may emit **multiple `tool_use` blocks in one response**.
      Execute them all (concurrently if safe), then return **all results in a single
      user message** — one `tool_result` per `tool_use`, ids matching, results may
      be in any order but must all be present:
      
      ```python
      tool_results = []
      for block in response.content:
          if block.type == "tool_use":
              result = execute_tool(block.name, block.input)   # block.input is parsed dict
              tool_results.append({
                  "type": "tool_result",
                  "tool_use_id": block.id,
                  "content": result,
              })
      messages.append({"role": "assistant", "content": response.content})
      messages.append({"role": "user", "content": tool_results})
      ```
      
      A follow-up request missing a `tool_result` for any outstanding `tool_use` id
      is a 400.
      
      ## The Agentic Loop (manual)
      
      Use the manual loop when you need approval gates, custom logging, or
      conditional execution:
      
      ```python
      import anthropic
      
      client = anthropic.Anthropic()
      messages = [{"role": "user", "content": user_input}]
      
      while True:
          response = client.messages.create(
              model="claude-opus-5",
              max_tokens=16000,
              tools=tools,
              messages=messages,
          )
      
          if response.stop_reason == "end_turn":
              break
      
          if response.stop_reason == "pause_turn":
              # Server-side tool loop hit its iteration limit: append and re-send.
              # Do NOT inject a "continue" user message — the API resumes automatically.
              messages.append({"role": "assistant", "content": response.content})
              continue
      
          if response.stop_reason == "tool_use":
              messages.append({"role": "assistant", "content": response.content})
              results = []
              for block in response.content:
                  if block.type == "tool_use":
                      try:
                          out = execute_tool(block.name, block.input)
                          results.append({"type": "tool_result",
                                          "tool_use_id": block.id, "content": out})
                      except Exception as e:
                          results.append({"type": "tool_result",
                                          "tool_use_id": block.id,
                                          "content": f"Error: {e}", "is_error": True})
              messages.append({"role": "user", "content": results})
              continue
      
          break  # max_tokens / refusal / stop_sequence — handle per stop_reason
      
      final_text = next((b.text for b in response.content if b.type == "text"), "")
      ```
      
      ```typescript
      import Anthropic from "@anthropic-ai/sdk";
      
      const client = new Anthropic();
      const messages: Anthropic.MessageParam[] = [{ role: "user", content: userInput }];
      
      while (true) {
        const response = await client.messages.create({
          model: "claude-opus-5", max_tokens: 16000, tools, messages,
        });
      
        if (response.stop_reason === "end_turn") break;
      
        if (response.stop_reason === "pause_turn") {
          messages.push({ role: "assistant", content: response.content });
          continue;
        }
      
        const toolUses = response.content.filter(
          (b): b is Anthropic.ToolUseBlock => b.type === "tool_use",
        );
        messages.push({ role: "assistant", content: response.content });
      
        const results: Anthropic.ToolResultBlockParam[] = [];
        for (const t of toolUses) {
          results.push({ type: "tool_result", tool_use_id: t.id,
                         content: await executeTool(t.name, t.input) });
        }
        messages.push({ role: "user", content: results });
      }
      ```
      
      Loop invariants:
      
      1. Append the **full** `response.content` as the assistant turn (preserves
         `tool_use` + `thinking` blocks; thinking `signature` must round-trip
         untouched).
      2. One `tool_result` per `tool_use`, matching `tool_use_id`.
      3. Tool results go in a **user** message.
      4. Add a max-iterations guard (e.g. 10-20) so a confused model can't loop forever.
      5. Parse `block.input` as structured data — never regex the serialized JSON
         (escaping varies across models).
      
      ## Tool Result Shapes
      
      ```json
      {"type": "tool_result", "tool_use_id": "toolu_01...", "content": "plain string"}
      
      {"type": "tool_result", "tool_use_id": "toolu_01...",
       "content": [{"type": "text", "text": "..."},
                   {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": "..."}}]}
      
      {"type": "tool_result", "tool_use_id": "toolu_01...",
       "content": "Error: location 'xyz' not found. Provide a valid city.",
       "is_error": true}
      ```
      
      `is_error: true` tells Claude the execution failed — it will typically adjust
      approach or ask for clarification. Return **informative** error strings, not
      stack traces.
      
      ## SDK Tool Runners (beta) — skip the manual loop
      
      The runners handle call → execute → feed-back → repeat automatically.
      
      ```python
      from anthropic import beta_tool
      import anthropic
      
      client = anthropic.Anthropic()
      
      @beta_tool
      def get_weather(location: str, unit: str = "celsius") -> str:
          """Get current weather for a location.
      
          Args:
              location: City and state, e.g. San Francisco, CA.
              unit: "celsius" or "fahrenheit".
          """
          return f"22°C and sunny in {location}"
      
      runner = client.beta.messages.tool_runner(
          model="claude-opus-5", max_tokens=16000,
          tools=[get_weather],
          messages=[{"role": "user", "content": "Weather in Paris?"}],
      )
      for message in runner:        # iterates messages until Claude stops calling tools
          print(message)
      ```
      
      ```typescript
      import Anthropic from "@anthropic-ai/sdk";
      import { betaZodTool } from "@anthropic-ai/sdk/helpers/beta/zod";
      import { z } from "zod";
      
      const getWeather = betaZodTool({
        name: "get_weather",
        description: "Get current weather for a location",
        inputSchema: z.object({
          location: z.string().describe("City and state, e.g. San Francisco, CA"),
        }),
        run: async ({ location }) => `22°C and sunny in ${location}`,
      });
      
      const finalMessage = await client.beta.messages.toolRunner({
        model: "claude-opus-5", max_tokens: 16000,
        tools: [getWeather],
        messages: [{ role: "user", content: "Weather in Paris?" }],
      });
      ```
      
      Schemas are generated from the function signature/docstring (Python) or Zod
      schema (TS). The runner executes tools **automatically** — for destructive
      side effects (email, payments, deletes), validate inside the tool function or
      use the manual loop with a human-approval gate.
      
      ## Server-Side Tools
      
      Declared in `tools`, executed by Anthropic — no client handling:
      
      | Tool | Type string | Notes |
      |---|---|---|
      | Web search | `web_search_20260209` | Per-search pricing; `allowed_domains`, `blocked_domains`, `max_uses` |
      | Web fetch | `web_fetch_20260209` | Fetch URL content with citations |
      | Code execution | `code_execution_20260120` | Sandboxed Python 3.11 (pandas/numpy/matplotlib preinstalled), no internet; container reusable via response `container.id` |
      | Tool search | `tool_search_tool_bm25_20251119` / `..._regex_20251119` | On-demand tool discovery for large libraries |
      | Memory | `memory_20250818` | Client-executed but Anthropic-defined schema |
      | Bash / text editor | `bash_20250124` / `text_editor_20250728` (name `str_replace_based_edit_tool`) | Anthropic-defined, **you** execute |
      
      Server tool runs may return `stop_reason: "pause_turn"` when the server-side
      loop hits its iteration limit (default 10) — append the assistant turn and
      re-request to resume (see loop above). Cap continuations (~5) to avoid
      infinite resumes.
      
      `web_search_20260209` / `web_fetch_20260209` include **dynamic filtering**
      (model filters results in a sandbox before they hit context) automatically —
      don't add a separate `code_execution` tool just for that; only declare
      code_execution when you need it independently.
      
      ## Bash vs Dedicated Tools (design)
      
      A bash tool gives breadth but the harness sees only an opaque command string.
      Promote an action to a dedicated tool when you need to:
      
      - **Gate** it (send_email behind confirmation — easy as a tool, impossible as `bash -c "curl ..."`)
      - **Validate** invariants (an edit tool can reject writes to files changed since last read)
      - **Render** it specially (question-asking as a modal)
      - **Parallelize** safely (read-only tools marked parallel-safe; bash must serialize)
      
      Start with bash for breadth; promote when you need to gate, render, audit, or
      parallelize.
      
  • scripts
    • .gitkeep 0 B · in bundle
    • check-model-table.py 32.2 KB
      #!/usr/bin/env python3
      """Staleness verifier for the claude-api-ops model + cache-minimum tables.
      
      Guards the two fast-moving fact tables in this skill against silent drift:
        - the "Current Models" table in SKILL.md (ids, pricing, context, output)
        - the per-model prompt-cache minimum table in references/caching-and-cost.md
      
      Two modes (protocol SKILL-RESOURCE-PROTOCOL.md §7):
        --offline (default): parse both tables, assert internal consistency. No network.
                             Exit 4 (VALIDATION) on a malformed/contradictory row.
        --live:              curl the Anthropic Models API and compare its model-id set
                             against the documented ids. Advisory only.
      
      Live-mode scope limit: the Models API returns model IDs but NOT pricing, context,
      or output limits. --live therefore verifies model-ID existence/coverage ONLY:
        - a documented id absent from the live list  -> DRIFT (retired/typo)
        - a live id newer than anything documented    -> DRIFT (table lacks a new model)
      Pricing/context/output drift is out of scope for --live (the API can't confirm it);
      --offline guards their well-formedness, and the SKILL.md "Live Documentation" links
      remain the human cross-check for pricing.
      
      Usage:   check-model-table.py [--offline | --live] [--json] [--skill-dir DIR] [-q]
      Input:   reads SKILL.md and references/caching-and-cost.md (resolved relative to
               this script, or --skill-dir)
      Output:  stdout = data only (JSON envelope under --json, else a plain summary)
      Stderr:  headers, progress, notes, errors
      Exit:    0 ok/consistent, 2 usage, 3 not-found, 4 validation (malformed/contradictory),
               5 missing-dep (curl, --live only), 7 unavailable (no key / API unreachable),
               10 drift (live id-set disagrees with the table)
      
      Examples:
        check-model-table.py --offline
        check-model-table.py --offline --json | python -m json.tool
        ANTHROPIC_API_KEY=sk-... check-model-table.py --live
        check-model-table.py --live   # exits 7 (advisory) when the key is unset
      """
      from __future__ import annotations
      
      import argparse
      import json
      import os
      import re
      import subprocess
      import sys
      from pathlib import Path
      from typing import NoReturn
      
      # Windows consoles default to cp1252; force UTF-8 so em-dashes/§ in notes don't
      # raise UnicodeEncodeError or print mojibake (matches the repo's standard fix).
      for _stream in (sys.stdout, sys.stderr):
          try:
              _stream.reconfigure(encoding="utf-8")  # type: ignore[attr-defined]
          except (AttributeError, ValueError):
              pass
      
      class Term:
          """Tiny ANSI helper mirroring skills/_lib/term.sh (term.sh is bash-only; per
          TERMINAL-DESIGN.md §9 the Python port is inline with matching keys/glyphs).
          Honors FORCE_COLOR / NO_COLOR / TERM_ASCII; color tracks the bound stream's TTY,
          and glyphs fall back to ASCII on TERM_ASCII or a non-UTF stream encoding."""
      
          _C = {"green": "\033[32m", "yellow": "\033[33m", "orange": "\033[38;5;208m",
                "red": "\033[31m", "cyan": "\033[36m", "dim": "\033[2m", "off": "\033[0m"}
          _GLYPH = {"ok": "✓", "bad": "✗", "warn": "▲", "skip": "—", "na": "—", "unknown": "?"}
          _ASCII = {"ok": "+", "bad": "x", "warn": "!", "skip": "-", "na": "-", "unknown": "?"}
          _MARK_COLOR = {"ok": "green", "bad": "red", "warn": "orange", "skip": "dim",
                         "na": "dim", "unknown": "yellow"}
      
          def __init__(self, stream=sys.stderr):
              enc = (getattr(stream, "encoding", "") or "").lower()
              self.ascii = (os.environ.get("TERM_ASCII") == "1"
                            or os.environ.get("FLEET_ASCII") == "1" or "utf" not in enc)
              if os.environ.get("FORCE_COLOR"):
                  self.color = True
              elif (os.environ.get("NO_COLOR") is not None or os.environ.get("TERM") == "dumb"
                    or not getattr(stream, "isatty", lambda: False)()):
                  self.color = False
              else:
                  self.color = True
      
          def c(self, name, text):
              return f"{self._C.get(name, '')}{text}{self._C['off']}" if self.color else text
      
          def mark(self, state):
              return self.c(self._MARK_COLOR.get(state, ""),
                            (self._ASCII if self.ascii else self._GLYPH).get(state, "."))
      
          def hdr(self, text):
              return self.c("cyan", f"=== {text} ===")
      
      
      TERM = Term(sys.stderr)
      
      EXIT_OK = 0
      EXIT_USAGE = 2
      EXIT_NOT_FOUND = 3
      EXIT_VALIDATION = 4
      EXIT_MISSING_DEP = 5
      EXIT_UNAVAILABLE = 7
      EXIT_DRIFT = 10
      
      SCHEMA = "claude-mods.claude-api-ops.model-table/v1"
      MODELS_API = "https://api.anthropic.com/v1/models?limit=1000"
      ANTHROPIC_VERSION = "2023-06-01"
      
      # --- Cache-economics constants (offline cross-file tripwire) -----------------
      # These numbers are stated in more than one file, so a future edit can desync
      # them silently. This table is the single source of truth: each entry names the
      # canonical value, a regex that must match in EVERY listed file, and why.
      # Changing a real-world value means changing it here AND in every listed file --
      # that is the point: a one-file edit trips this check instead of rotting.
      #
      # Deliberately offline-only. A --live parse of the docs page would be brittle
      # (prose/format churn -> spurious exit 10), and SKILL-RESOURCE-PROTOCOL.md §7 is
      # explicit that a gate which cries wolf is a gate everyone learns to ignore.
      # The human cross-check stays the "Live Documentation" link in SKILL.md.
      CACHE_CONSTANTS = [
          {
              "key": "cache_read_multiplier",
              "value": "0.1x base input",
              "pattern": r"0\.1\s*[x×]",
              "files": ["SKILL.md", "references/caching-and-cost.md",
                        "references/context-engineering.md", "references/compaction.md"],
          },
          {
              "key": "cache_write_5m_multiplier",
              "value": "1.25x base input",
              "pattern": r"1\.25\s*[x×]",
              "files": ["references/caching-and-cost.md", "references/compaction.md"],
          },
          {
              "key": "cache_write_1h_multiplier",
              "value": "2x base input",
              # Anchored to the TTL in either order so a bare "2x" elsewhere can't
              # satisfy it. Note "2×" (U+00D7) has no trailing \b -- don't add one.
              "pattern": (r"(?<![\d.])2\s*[x×][^\n]{0,40}1-hour"
                          r"|1-hour[^\n]{0,40}(?<![\d.])2\s*[x×]"),
              "files": ["references/caching-and-cost.md", "references/compaction.md"],
          },
          {
              "key": "max_breakpoints",
              "value": "4 cache_control breakpoints per request",
              "pattern": r"\*\*4\*\*\s*`?cache_control`?\s*breakpoints",
              "files": ["references/caching-and-cost.md"],
          },
          {
              "key": "lookback_blocks",
              "value": "20-block backward lookback per breakpoint",
              "pattern": r"\b20\b[^.\n]{0,48}\bblock",
              "files": ["references/caching-and-cost.md",
                        "references/context-engineering.md"],
          },
          {
              # context-budget.py encodes the same multipliers as executable constants.
              # Prose and code drifting apart is the exact failure this guards: a doc
              # that says 0.1x beside a calculator that bills 0.15x.
              "key": "cache_multipliers_in_calculator",
              "value": "context-budget.py constants match the documented multipliers",
              "pattern": (r"CACHE_READ_MULTIPLIER\s*=\s*0\.1\b[\s\S]{0,200}?"
                          r"CACHE_WRITE_5M\s*=\s*1\.25\b[\s\S]{0,200}?"
                          r"CACHE_WRITE_1H\s*=\s*2\.0\b"),
              "files": ["scripts/context-budget.py"],
          },
          {
              "key": "context_management_beta",
              "value": "context-management-2025-06-27 beta header",
              "pattern": r"context-management-2025-06-27",
              "files": ["SKILL.md", "references/compaction.md"],
          },
      ]
      
      # Docs whose facts were verified on a date must SAY so, so staleness is visible
      # to a reader rather than only to this script.
      DATE_STAMPED_DOCS = ["references/context-engineering.md", "references/compaction.md"]
      DATE_STAMP_RE = re.compile(r"verified[^\n]{0,80}?(\d{4}-\d{2}-\d{2})", re.I)
      
      # A well-formed alias id: claude-<word>-<digit>... and NO date suffix.
      # Accepts claude-opus-5, claude-fable-5, claude-sonnet-5, claude-haiku-4-5.
      ID_RE = re.compile(r"^claude-[a-z]+-\d+(?:-\d+)?$")
      # A date suffix looks like an 8-digit run (e.g. -20251114).
      DATE_SUFFIX_RE = re.compile(r"-\d{8}$")
      
      # Any Claude model id appearing in the skill body, dated snapshots included.
      ANY_ID_RE = re.compile(r"\bclaude-[a-z]+-\d+(?:-\d+)?(?:-\d{8})?\b")
      # A model id being ASSIGNED - what a reader copies. Covers model="x", model: "x",
      # "model": "x", MODEL = "x", and --model x.
      ASSIGN_RE = re.compile(
          r'(?:"?\bmodel"?\s*[:=]\s*["\']([a-z0-9.\-]+)["\'])'
          r'|(?:--model[=\s]+([a-z0-9.\-]+))',
          re.IGNORECASE,
      )
      # Escape hatch: a line carrying this marker may name a legacy id deliberately
      # (e.g. a migration before/after snippet).
      LEGACY_OK = "legacy-ok"
      # tests/ is fixture material - it deliberately writes malformed ids to prove the
      # verifier rejects them, so scanning it would be self-defeating.
      SCAN_SKIP_DIRS = {"tests", "__pycache__", ".git"}
      SCAN_SUFFIXES = {".md", ".py", ".json", ".sh", ".ts", ".js", ".yaml", ".yml", ".txt"}
      
      
      def note(msg: str, quiet: bool) -> None:
          if not quiet:
              print(msg, file=sys.stderr)
      
      
      def fail_validation(message: str, details: dict, json_mode: bool) -> NoReturn:
          if json_mode:
              print(json.dumps({"error": {"code": "VALIDATION", "message": message,
                                          "details": details}}))
          print(f"{TERM.mark('bad')} ERROR: {message}", file=sys.stderr)
          for k, v in details.items():
              print(f"  {k}: {v}", file=sys.stderr)
          sys.exit(EXIT_VALIDATION)
      
      
      # ---------------------------------------------------------------------------
      # Parsing
      # ---------------------------------------------------------------------------
      
      def split_row(line: str) -> list[str]:
          """Split a markdown table row into trimmed cells (drops outer pipes)."""
          cells = [c.strip() for c in line.strip().strip("|").split("|")]
          return cells
      
      
      def is_separator(cells: list[str]) -> bool:
          return all(re.fullmatch(r":?-{2,}:?", c or "") for c in cells) and bool(cells)
      
      
      def parse_model_table(text: str) -> tuple[list[dict], list[str]]:
          """Parse the SKILL.md 'Current Models' table.
      
          Columns: Model | ID | Context | Max Output | Input $/MTok | Output $/MTok
          Returns one dict per data row.
          """
          lines = text.splitlines()
          # Locate the header row that contains the ID column and a price column.
          start = None
          for i, line in enumerate(lines):
              low = line.lower()
              if line.lstrip().startswith("|") and "id" in low and "context" in low and "output" in low:
                  start = i
                  break
          if start is None:
              return [], []
      
          header = split_row(lines[start])
          rows: list[dict] = []
          # Expect a separator row next, then data rows until a non-table line.
          j = start + 1
          if j < len(lines) and is_separator(split_row(lines[j])):
              j += 1
          while j < len(lines):
              line = lines[j]
              if not line.lstrip().startswith("|"):
                  break
              cells = split_row(line)
              if is_separator(cells):
                  j += 1
                  continue
              if len(cells) >= 6:
                  rows.append({
                      "name": cells[0],
                      "id_cell": cells[1],
                      "context": cells[2],
                      "max_output": cells[3],
                      "input_price": cells[4],
                      "output_price": cells[5],
                  })
              j += 1
          return rows, header
      
      
      def parse_legacy_ids(text: str) -> list[str]:
          """Extract the backticked ids from SKILL.md's 'Legacy (still available...)'
          paragraph. Parsed rather than hard-coded so retiring a model is a one-line
          doc edit, not a code change."""
          m = re.search(r"\*\*Legacy \(still available[^*]*\)\:\*\*(.+?)(?:\n\n|\Z)",
                        text, re.S)
          if not m:
              return []
          return [i for i in re.findall(r"`([^`]+)`", m.group(1)) if ID_RE.match(i)]
      
      
      def scan_body_ids(skill_dir: Path, current: set[str], legacy: set[str]) -> list[dict]:
          """Flag model ids used in the skill body that a reader must not copy.
      
          Two distinct faults, because they have different severities in practice:
            unknown  - an id in neither the current table nor the legacy list. A typo
                       or a hallucinated model; nothing resolves it.
            retired  - a LEGACY id sitting in an assignable position (model=..., MODEL
                       = ..., "model": ..., --model ...). Naming a legacy model in
                       prose is legitimate; shipping one as the value a reader copies
                       is not. Append the `legacy-ok` marker to that line to allow it.
          """
          findings: list[dict] = []
          for path in sorted(skill_dir.rglob("*")):
              if not path.is_file() or path.suffix.lower() not in SCAN_SUFFIXES:
                  continue
              if SCAN_SKIP_DIRS & set(p.name for p in path.relative_to(skill_dir).parents):
                  continue
              rel = path.relative_to(skill_dir).as_posix()
              try:
                  lines = path.read_text(encoding="utf-8").splitlines()
              except UnicodeDecodeError:
                  continue
              for n, line in enumerate(lines, 1):
                  if LEGACY_OK in line:
                      continue
                  # Strip a dated snapshot to its alias before membership testing: a
                  # dated id is a real, resolvable snapshot of a known model, and the
                  # docs discuss several deliberately.
                  for raw in ANY_ID_RE.findall(line):
                      base = DATE_SUFFIX_RE.sub("", raw)
                      if base not in current and base not in legacy:
                          findings.append({"file": rel, "line": n, "id": raw,
                                           "fault": "unknown"})
                  for a, b in ASSIGN_RE.findall(line):
                      val = DATE_SUFFIX_RE.sub("", (a or b).strip())
                      if val in legacy:
                          findings.append({"file": rel, "line": n, "id": val,
                                           "fault": "retired"})
          return findings
      
      
      def parse_cache_min_table(text: str) -> list[dict]:
          """Parse the caching-and-cost.md 'Minimum prefix tokens' table.
      
          Columns: Model | Minimum prefix tokens. The Model cell holds friendly names
          (possibly several comma-separated), not ids.
          """
          lines = text.splitlines()
          start = None
          for i, line in enumerate(lines):
              low = line.lower()
              if line.lstrip().startswith("|") and "model" in low and "minimum" in low and "prefix" in low:
                  start = i
                  break
          if start is None:
              return []
          rows: list[dict] = []
          j = start + 1
          if j < len(lines) and is_separator(split_row(lines[j])):
              j += 1
          while j < len(lines):
              line = lines[j]
              if not line.lstrip().startswith("|"):
                  break
              cells = split_row(line)
              if is_separator(cells):
                  j += 1
                  continue
              if len(cells) >= 2:
                  rows.append({"names": cells[0], "min_tokens": cells[1]})
              j += 1
          return rows
      
      
      # ---------------------------------------------------------------------------
      # Offline validation
      # ---------------------------------------------------------------------------
      
      PRICE_RE = re.compile(r"^\$\d+(?:\.\d+)?$")
      SIZE_RE = re.compile(r"^\d+(?:\.\d+)?[KM]$")
      
      
      def clean_id(id_cell: str) -> str:
          """Strip backtick code fences from an ID cell."""
          return id_cell.strip().strip("`").strip()
      
      
      def validate_cache_constants(skill_dir: Path, json_mode: bool, quiet: bool) -> list[dict]:
          """Assert every file that states a cache-economics constant states the same one.
      
          Presence-based on purpose: if an edit changes 0.1x to 0.15x in one file, that
          file stops matching and this trips. Returns one result row per constant.
          """
          rows: list[dict] = []
          for const in CACHE_CONSTANTS:
              rx = re.compile(const["pattern"])
              missing: list[str] = []
              for rel in const["files"]:
                  path = skill_dir / rel
                  if not path.is_file():
                      fail_validation("file referenced by a cache constant is missing",
                                      {"constant": const["key"], "file": rel}, json_mode)
                  if not rx.search(path.read_text(encoding="utf-8")):
                      missing.append(rel)
              if missing:
                  fail_validation(
                      f"cache constant {const['key']!r} not stated consistently across files",
                      {"expected": const["value"],
                       "pattern": const["pattern"],
                       "files_missing_it": ", ".join(missing),
                       "hint": ("either the doc drifted, or the real value changed and "
                                "CACHE_CONSTANTS in this script needs updating too")},
                      json_mode)
              rows.append({"key": const["key"], "value": const["value"],
                           "files": const["files"], "consistent": True})
          note(f"  {len(rows)} cache constants consistent across files", quiet)
          return rows
      
      
      def validate_date_stamps(skill_dir: Path, json_mode: bool, quiet: bool) -> list[dict]:
          """Every fact-heavy doctrine reference must carry a 'verified <ISO date>' stamp."""
          rows: list[dict] = []
          for rel in DATE_STAMPED_DOCS:
              path = skill_dir / rel
              if not path.is_file():
                  fail_validation("date-stamped doc is missing", {"file": rel}, json_mode)
              m = DATE_STAMP_RE.search(path.read_text(encoding="utf-8"))
              if not m:
                  fail_validation(
                      "doc encodes fast-moving facts but carries no verification date",
                      {"file": rel,
                       "hint": "add a line like 'Facts verified against ... YYYY-MM-DD'"},
                      json_mode)
              rows.append({"file": rel, "verified": m.group(1)})
          note(f"  {len(rows)} doctrine docs carry a verification date", quiet)
          return rows
      
      
      def validate_cited_references(skill_dir: Path, json_mode: bool, quiet: bool) -> list[str]:
          """Every references/*.md on disk must be cited from SKILL.md, and vice versa.
      
          SKILL-RESOURCE-PROTOCOL.md §1: an uncited reference is dead weight the router
          never finds; a cited-but-absent one is a broken link.
          """
          skill_text = (skill_dir / "SKILL.md").read_text(encoding="utf-8")
          on_disk = sorted(p.name for p in (skill_dir / "references").glob("*.md"))
          cited = set(re.findall(r"\(references/([A-Za-z0-9._-]+\.md)\)", skill_text))
      
          uncited = [n for n in on_disk if n not in cited]
          ghosts = sorted(n for n in cited if n not in on_disk)
          if uncited or ghosts:
              fail_validation(
                  "SKILL.md and references/ disagree",
                  {"on_disk_but_uncited": ", ".join(uncited) or "(none)",
                   "cited_but_missing": ", ".join(ghosts) or "(none)"},
                  json_mode)
      
          # Same rule for scripts/ and assets/ (SKILL-RESOURCE-PROTOCOL.md §2.8: every
          # resource is cited from SKILL.md). Those are referenced as backticked paths
          # or worked invocations rather than markdown links, so match on the basename.
          shipped: list[str] = []
          for sub in ("scripts", "assets"):
              for p in sorted((skill_dir / sub).glob("*")):
                  if not p.is_file() or p.name.startswith(".") or p.suffix == ".pyc":
                      continue
                  if p.name not in skill_text:
                      fail_validation(
                          f"{sub}/ file is not cited from SKILL.md",
                          {"file": f"{sub}/{p.name}",
                           "hint": "an uncited resource is dead weight the router never "
                                   "finds - cite it with a worked invocation"},
                          json_mode)
                  shipped.append(f"{sub}/{p.name}")
      
          note(f"  {len(on_disk)} reference files + {len(shipped)} scripts/assets, "
               "all cited from SKILL.md", quiet)
          return on_disk + shipped
      
      
      def validate_offline(skill_dir: Path, json_mode: bool, quiet: bool) -> dict:
          skill_md = skill_dir / "SKILL.md"
          cache_md = skill_dir / "references" / "caching-and-cost.md"
          for p in (skill_md, cache_md):
              if not p.is_file():
                  if json_mode:
                      print(json.dumps({"error": {"code": "NOT_FOUND",
                                                  "message": f"missing file: {p}",
                                                  "details": {}}}))
                  print(f"ERROR: required file not found: {p}", file=sys.stderr)
                  sys.exit(EXIT_NOT_FOUND)
      
          note(TERM.hdr("offline model-table consistency check"), quiet)
      
          model_rows, _ = parse_model_table(skill_md.read_text(encoding="utf-8"))
          if not model_rows:
              fail_validation("could not locate a non-empty Current Models table in SKILL.md",
                              {"file": str(skill_md)}, json_mode)
      
          documented_ids: list[str] = []
          models_out: list[dict] = []
          for row in model_rows:
              mid = clean_id(row["id_cell"])
              problems = []
              if not ID_RE.match(mid):
                  problems.append("id does not match claude-[a-z]+-<digits>")
              if DATE_SUFFIX_RE.search(mid):
                  problems.append("id carries a date suffix (should be a bare alias)")
              if not PRICE_RE.match(row["input_price"]):
                  problems.append(f"input price not numeric: {row['input_price']!r}")
              if not PRICE_RE.match(row["output_price"]):
                  problems.append(f"output price not numeric: {row['output_price']!r}")
              if not SIZE_RE.match(row["context"]):
                  problems.append(f"context not a size (e.g. 1M/200K): {row['context']!r}")
              if not SIZE_RE.match(row["max_output"]):
                  problems.append(f"max output not a size: {row['max_output']!r}")
              if problems:
                  fail_validation(f"malformed model row: {row['name']!r}",
                                  {"id": mid, "problems": "; ".join(problems)}, json_mode)
              documented_ids.append(mid)
              models_out.append({
                  "name": row["name"], "id": mid, "context": row["context"],
                  "max_output": row["max_output"],
                  "input_price": row["input_price"], "output_price": row["output_price"],
              })
      
          # No duplicate ids.
          dupes = {x for x in documented_ids if documented_ids.count(x) > 1}
          if dupes:
              fail_validation("duplicate model ids in the table",
                              {"ids": ", ".join(sorted(dupes))}, json_mode)
      
          # Cache-minimum table.
          cache_rows = parse_cache_min_table(cache_md.read_text(encoding="utf-8"))
          if not cache_rows:
              fail_validation("could not locate the cache-minimum table in caching-and-cost.md",
                              {"file": str(cache_md)}, json_mode)
          for crow in cache_rows:
              if not re.fullmatch(r"\d+", crow["min_tokens"]):
                  fail_validation("cache-minimum value is not an integer",
                                  {"row": crow["names"], "value": crow["min_tokens"]},
                                  json_mode)
      
          # Cross-file consistency: every model NAME (e.g. "Opus 4.8", "Fable 5",
          # "Sonnet 4.6", "Haiku 4.5") in the model table must appear in the cache
          # table's name set, so the two files agree on the model lineup.
          cache_blob = " ".join(c["names"] for c in cache_rows).lower()
          missing_in_cache: list[str] = []
          for m in models_out:
              # Derive the short family+version token, e.g. "Claude Opus 4.8" -> "opus 4.8".
              short = re.sub(r"^claude\s+", "", m["name"], flags=re.I).strip().lower()
              if short not in cache_blob:
                  missing_in_cache.append(m["name"])
          if missing_in_cache:
              fail_validation(
                  "model(s) in SKILL.md absent from the cache-minimum table — files contradict",
                  {"missing": ", ".join(missing_in_cache),
                   "hint": "every documented model needs a prompt-cache minimum row"},
                  json_mode)
      
          # Body scan: the tables can be perfectly consistent while a code sample
          # ships a retired model. Guard what readers copy, not just what they read.
          legacy_ids = set(parse_legacy_ids(skill_md.read_text(encoding="utf-8")))
          body = scan_body_ids(skill_dir, set(documented_ids), legacy_ids)
          if body:
              detail = {}
              for f in body[:12]:
                  detail[f"{f['file']}:{f['line']}"] = f"{f['fault']}: {f['id']}"
              if len(body) > 12:
                  detail["..."] = f"{len(body) - 12} more"
              fail_validation(
                  f"{len(body)} model-id problem(s) in the skill body",
                  {**detail,
                   "hint": "'retired' = a legacy id in an assignable position; retarget "
                           "it at a current model or append the 'legacy-ok' marker. "
                           "'unknown' = an id in neither the model table nor the legacy list."},
                  json_mode)
      
          note(f"  {len(models_out)} model rows, all well-formed", quiet)
          note(f"  {len(cache_rows)} cache-minimum rows, all integer", quiet)
          note("  cross-file model lineup consistent", quiet)
          note(f"  {len(legacy_ids)} legacy id(s) tracked; no retired/unknown id in the body", quiet)
      
          # Context-engineering layer: constants stated in several files, verification
          # date stamps, and SKILL.md <-> references/ citation integrity.
          constants = validate_cache_constants(skill_dir, json_mode, quiet)
          stamps = validate_date_stamps(skill_dir, json_mode, quiet)
          refs = validate_cited_references(skill_dir, json_mode, quiet)
      
          note(f"{TERM.mark('ok')} OK: tables internally consistent.", quiet)
      
          return {
              "mode": "offline",
              "models": models_out,
              "documented_ids": documented_ids,
              "cache_min_rows": cache_rows,
              "cache_constants": constants,
              "date_stamps": stamps,
              "reference_files": refs,
              "legacy_ids": sorted(legacy_ids),
              "consistent": True,
          }
      
      
      # ---------------------------------------------------------------------------
      # Live validation
      # ---------------------------------------------------------------------------
      
      def fetch_live_ids(quiet: bool) -> list[str] | None:
          """Return the live model-id list, or None if unavailable (advisory)."""
          key = os.environ.get("ANTHROPIC_API_KEY", "").strip()
          if not key:
              note("NOTE: ANTHROPIC_API_KEY is unset - skipping live check (advisory).",
                   quiet)
              return None
          cmd = [
              "curl", "-fsS", "--max-time", "20",
              "-H", f"x-api-key: {key}",
              "-H", f"anthropic-version: {ANTHROPIC_VERSION}",
              MODELS_API,
          ]
          try:
              proc = subprocess.run(cmd, capture_output=True, text=True, timeout=30)
          except (subprocess.TimeoutExpired, OSError) as exc:
              note(f"NOTE: Models API call failed ({exc}) — advisory, not a failure.",
                   quiet)
              return None
          if proc.returncode != 0:
              note(f"NOTE: Models API unreachable (curl exit {proc.returncode}) — advisory.",
                   quiet)
              if proc.stderr.strip():
                  note(f"  {proc.stderr.strip().splitlines()[-1]}", quiet)
              return None
          try:
              payload = json.loads(proc.stdout)
          except json.JSONDecodeError:
              note("NOTE: Models API returned non-JSON — advisory, not a failure.", quiet)
              return None
          data = payload.get("data")
          if not isinstance(data, list):
              note("NOTE: Models API JSON missing 'data' list — advisory.", quiet)
              return None
          return [m.get("id", "") for m in data if isinstance(m, dict) and m.get("id")]
      
      
      def validate_live(skill_dir: Path, json_mode: bool, quiet: bool) -> dict:
          if not _have("curl"):
              if json_mode:
                  print(json.dumps({"error": {"code": "PRECONDITION",
                                               "message": "curl required for --live",
                                               "details": {}}}))
              print("ERROR: curl is required for --live", file=sys.stderr)
              sys.exit(EXIT_MISSING_DEP)
      
          # Reuse offline parse for the documented id set (also validates well-formedness).
          note(TERM.hdr("live model-id coverage check"), quiet)
          skill_md = skill_dir / "SKILL.md"
          if not skill_md.is_file():
              print(f"ERROR: required file not found: {skill_md}", file=sys.stderr)
              sys.exit(EXIT_NOT_FOUND)
          parsed = parse_model_table(skill_md.read_text(encoding="utf-8"))
          if not parsed or not parsed[0]:
              fail_validation("could not parse the model table for live comparison",
                              {"file": str(skill_md)}, json_mode)
          documented = [clean_id(r["id_cell"]) for r in parsed[0]]
      
          live = fetch_live_ids(quiet)
          if live is None:
              # Advisory: not a failure. Exit 7.
              if json_mode:
                  print(json.dumps({"data": {"mode": "live", "status": "unavailable",
                                             "documented_ids": documented, "live_ids": None},
                                    "meta": {"schema": SCHEMA, "status": "unavailable"}}))
              sys.exit(EXIT_UNAVAILABLE)
      
          live_set = set(live)
          doc_set = set(documented)
      
          # A documented id absent from the live list = drift (retired/typo).
          missing = sorted(doc_set - live_set)
          # A live id NEWER than anything documented = drift (table lacks a new model).
          # Restrict "newer" to well-formed alias ids so we ignore date-suffixed and
          # snapshot variants the docs intentionally don't list.
          live_alias = {m for m in live_set if ID_RE.match(m) and not DATE_SUFFIX_RE.search(m)}
          new_models = sorted(live_alias - doc_set)
      
          drift = bool(missing or new_models)
          result = {
              "mode": "live",
              "status": "drift" if drift else "ok",
              "documented_ids": documented,
              "live_ids": sorted(live_set),
              "missing_from_live": missing,
              "new_in_live": new_models,
          }
      
          if drift:
              if missing:
                  note(f"{TERM.mark('bad')} {TERM.c('red', 'DRIFT: documented id(s) absent from live Models API:')}", quiet)
                  for m in missing:
                      note(f"  {TERM.c('red', '-')} {m}", quiet)
              if new_models:
                  note(f"{TERM.mark('bad')} {TERM.c('red', 'DRIFT: live Models API has alias id(s) the table lacks:')}", quiet)
                  for m in new_models:
                      note(f"  {TERM.c('green', '+')} {m}", quiet)
              if json_mode:
                  print(json.dumps({"data": result, "meta": {"schema": SCHEMA,
                                                              "status": "drift"}}))
              else:
                  print("DRIFT: model-id table disagrees with the live Models API "
                        f"(missing={missing}, new={new_models})")
              sys.exit(EXIT_DRIFT)
      
          note("OK: every documented id exists live; no newer alias id missing from the table.",
               quiet)
          return result
      
      
      def _have(tool: str) -> bool:
          from shutil import which
          return which(tool) is not None
      
      
      # ---------------------------------------------------------------------------
      # Main
      # ---------------------------------------------------------------------------
      
      def main(argv: list[str]) -> int:
          parser = argparse.ArgumentParser(
              prog="check-model-table.py", add_help=True,
              description="Staleness verifier for the claude-api-ops model + cache tables.",
              epilog=(
                  "EXAMPLES:\n"
                  "  check-model-table.py --offline\n"
                  "  check-model-table.py --offline --json | python -m json.tool\n"
                  "  ANTHROPIC_API_KEY=sk-... check-model-table.py --live\n"
                  "  check-model-table.py --live   # exits 7 (advisory) when key unset\n"
              ),
              formatter_class=argparse.RawDescriptionHelpFormatter,
          )
          mode = parser.add_mutually_exclusive_group()
          mode.add_argument("--offline", action="store_true",
                            help="parse + assert internal consistency, no network (default)")
          mode.add_argument("--live", action="store_true",
                            help="compare documented ids against the live Models API (advisory)")
          parser.add_argument("--json", action="store_true",
                              help="emit the JSON envelope on stdout")
          parser.add_argument("--skill-dir", default=None,
                              help="skill root (default: parent of this script's dir)")
          parser.add_argument("-q", "--quiet", action="store_true",
                              help="suppress stderr progress/notes")
          args = parser.parse_args(argv)
      
          if args.skill_dir:
              skill_dir = Path(args.skill_dir).resolve()
          else:
              skill_dir = Path(__file__).resolve().parent.parent
          if not skill_dir.is_dir():
              print(f"ERROR: skill dir not found: {skill_dir}", file=sys.stderr)
              return EXIT_NOT_FOUND
      
          if args.live:
              result = validate_live(skill_dir, args.json, args.quiet)
          else:
              result = validate_offline(skill_dir, args.json, args.quiet)
      
          if args.json:
              print(json.dumps({"data": result,
                                "meta": {"schema": SCHEMA, "status": "ok"}}))
          return EXIT_OK
      
      
      if __name__ == "__main__":
          try:
              sys.exit(main(sys.argv[1:]))
          except KeyboardInterrupt:
              sys.exit(EXIT_USAGE)
      
    • context-budget.py 17.9 KB
      #!/usr/bin/env python3
      """Append-vs-compact calculator for a cached multi-turn Claude conversation.
      
      Answers the one question the context-engineering doctrine says to ask before
      compacting: over the turns you actually have left, is rewriting the history
      cheaper than carrying it? Under prompt caching the answer is usually no, and
      the arithmetic is short enough that people skip it and guess wrong.
      
      Models both paths in dollars (see references/compaction.md §3):
        append  : history rides the cache at CACHE_READ_MULTIPLIER of base input,
                  every remaining turn, growing by --growth-per-turn.
        compact : summarisation call(s) (cached read in, summary out) + a cache
                  WRITE of the new prefix each time + the remaining turns on the
                  summary, and every token cached before a rewrite is forfeited.
      
                  Compaction RECURS when the history grows. With --growth-per-turn
                  set, the summary climbs back toward the original size and a real
                  system compacts again; this models that cadence rather than
                  charging a single one-shot rewrite. Assuming one compaction
                  understated its cost by ~2.2x on a 60-turn, 4k-per-turn session.
      
      Also checks the hard constraint first: if the projected history overflows the
      context window, cost is moot and compaction (or offloading) is forced.
      
      Two things to understand before trusting the verdict:
      
        * The break-even turn count is SCALE-INVARIANT. History size and price both
          cancel out of fixed_cost / per_turn_saving -- it is driven only by the
          summary ratio, the output/input price ratio, and the write multiplier. So
          "compact" vs "append" is almost entirely a question of how many turns you
          have left, not how big or expensive the conversation is.
        * RECALL IS NOT PRICED. This models dollars only, which is one of the three
          constraints in references/compaction.md. The measured recall cost of
          summarisation (92-100% -> 38-58% on a planted-fact probe) does not appear
          anywhere in these numbers. A "compact" verdict means compaction is cheaper,
          NOT that it is right. Probe your own recall before acting on it --
          assets/recall-probe.py does exactly that.
      
      Usage:   context-budget.py --history-tokens N --turns-remaining N [OPTIONS]
      Input:   argv only; no stdin, no network, no files read or written.
      Output:  stdout = data only (JSON envelope under --json, else a plain summary)
      Stderr:  headers, workings, notes
      Exit:    0 append wins (keep everything), 2 usage, 4 validation (bad numbers),
               10 compaction indicated (cost crossover or context ceiling)
      
      Examples:
        # Short session -> append wins (exit 0)
        context-budget.py --history-tokens 25000 --turns-remaining 5 --base-rate 0.30
      
        # 40 turns left on a 120K history -> cost favours compaction (exit 10)
        context-budget.py --history-tokens 120000 --turns-remaining 40 --base-rate 2.00
      
        # Premium tier, long session, history still growing each turn
        context-budget.py --history-tokens 300000 --turns-remaining 60 \
            --base-rate 10.00 --growth-per-turn 4000 --ttl 1h
      
        # Machine-readable, for an agent deciding mid-run
        context-budget.py --history-tokens 900000 --turns-remaining 10 \
            --base-rate 5.00 --json | python -m json.tool
      """
      from __future__ import annotations
      
      import argparse
      import json
      import os
      import sys
      
      # Windows consoles default to cp1252; force UTF-8 so the section glyphs in the
      # human framing don't raise UnicodeEncodeError (matches check-model-table.py).
      for _stream in (sys.stdout, sys.stderr):
          try:
              _stream.reconfigure(encoding="utf-8")  # type: ignore[attr-defined]
          except (AttributeError, ValueError):
              pass
      
      
      class Term:
          """Tiny ANSI helper mirroring skills/_lib/term.sh (term.sh is bash-only; per
          TERMINAL-DESIGN.md §9 the Python port is inline with matching keys/glyphs).
          Honors FORCE_COLOR / NO_COLOR / TERM_ASCII; color tracks the bound stream's TTY,
          and glyphs fall back to ASCII on TERM_ASCII or a non-UTF stream encoding."""
      
          _C = {"green": "\033[32m", "yellow": "\033[33m", "orange": "\033[38;5;208m",
                "red": "\033[31m", "cyan": "\033[36m", "dim": "\033[2m", "off": "\033[0m"}
          _GLYPH = {"ok": "✓", "bad": "✗", "warn": "▲", "skip": "—", "na": "—", "unknown": "?"}
          _ASCII = {"ok": "+", "bad": "x", "warn": "!", "skip": "-", "na": "-", "unknown": "?"}
          _MARK_COLOR = {"ok": "green", "bad": "red", "warn": "orange", "skip": "dim",
                         "na": "dim", "unknown": "yellow"}
      
          def __init__(self, stream=sys.stderr):
              enc = (getattr(stream, "encoding", "") or "").lower()
              self.ascii = (os.environ.get("TERM_ASCII") == "1"
                            or os.environ.get("FLEET_ASCII") == "1" or "utf" not in enc)
              if os.environ.get("FORCE_COLOR"):
                  self.color = True
              elif (os.environ.get("NO_COLOR") is not None or os.environ.get("TERM") == "dumb"
                    or not getattr(stream, "isatty", lambda: False)()):
                  self.color = False
              else:
                  self.color = True
      
          def c(self, name, text):
              return f"{self._C.get(name, '')}{text}{self._C['off']}" if self.color else text
      
          def mark(self, state):
              return self.c(self._MARK_COLOR.get(state, ""),
                            (self._ASCII if self.ascii else self._GLYPH).get(state, "."))
      
          def hdr(self, text):
              return self.c("cyan", f"=== {text} ===")
      
      
      TERM = Term(sys.stderr)
      
      EXIT_OK = 0
      EXIT_USAGE = 2
      EXIT_VALIDATION = 4
      EXIT_COMPACT = 10
      
      SCHEMA = "claude-mods.claude-api-ops.context-budget/v1"
      
      # --- Cache economics (verified against platform.claude.com 2026-08-30) -------
      # Guarded by check-model-table.py --offline: these literals are cross-checked
      # against the prose in SKILL.md and references/*.md, so a one-file edit trips
      # CI instead of leaving the docs and this calculator quietly disagreeing.
      CACHE_READ_MULTIPLIER = 0.1     # cache read, every model, flat
      CACHE_WRITE_5M = 1.25           # cache write, 5-minute TTL
      CACHE_WRITE_1H = 2.0            # cache write, 1-hour TTL
      
      # Every model in the current lineup prices output at exactly 5x its input rate
      # (10/50, 5/25, 2/10, 1/5), so --output-rate defaults to 5x --base-rate rather
      # than forcing the caller to look it up. Override it if that ever stops holding.
      OUTPUT_RATE_RATIO = 5.0
      
      PER_MTOK = 1_000_000.0
      
      
      def note(msg: str, quiet: bool) -> None:
          if not quiet:
              print(msg, file=sys.stderr)
      
      
      def fail(code: str, message: str, details: dict, json_mode: bool, exit_code: int):
          if json_mode:
              print(json.dumps({"error": {"code": code, "message": message,
                                          "details": details}}))
          print(f"{TERM.mark('bad')} ERROR: {message}", file=sys.stderr)
          for k, v in details.items():
              print(f"  {k}: {v}", file=sys.stderr)
          sys.exit(exit_code)
      
      
      def cached_carry_cost(start_tokens: float, turns: int, growth: float,
                            rate: float) -> float:
          """Cost of re-sending a growing history at cache-read price for `turns` turns."""
          total_tokens = sum(start_tokens + growth * t for t in range(turns))
          return total_tokens * CACHE_READ_MULTIPLIER * rate / PER_MTOK
      
      
      def main(argv: list[str]) -> int:
          parser = argparse.ArgumentParser(
              prog="context-budget.py", add_help=True,
              description="Append-vs-compact calculator for a cached Claude conversation.",
              epilog=(
                  "EXAMPLES:\n"
                  "  context-budget.py --history-tokens 25000 --turns-remaining 5 "
                  "--base-rate 0.30      # append wins (exit 0)\n"
                  "  context-budget.py --history-tokens 120000 --turns-remaining 40 "
                  "--base-rate 2.00      # cost favours compaction (exit 10)\n"
                  "  context-budget.py --history-tokens 300000 --turns-remaining 60 "
                  "--base-rate 10.00 --growth-per-turn 4000 --ttl 1h\n"
                  "  context-budget.py --history-tokens 900000 --turns-remaining 10 "
                  "--base-rate 5.00 --json\n"
                  "\nEXIT: 0 append wins, 2 usage, 4 bad numbers, 10 compaction indicated\n"
              ),
              formatter_class=argparse.RawDescriptionHelpFormatter,
          )
          parser.add_argument("--history-tokens", type=float, required=True,
                              help="current conversation/prefix size in tokens")
          parser.add_argument("--turns-remaining", type=int, required=True,
                              help="turns you expect AFTER this decision (a compaction "
                                   "just before the end is pure loss)")
          parser.add_argument("--base-rate", type=float, default=5.00,
                              help="base INPUT price in $/MTok (default: 5.00)")
          parser.add_argument("--output-rate", type=float, default=None,
                              help=f"output price in $/MTok (default: {OUTPUT_RATE_RATIO}x "
                                   "--base-rate, which holds across the current lineup)")
          parser.add_argument("--summary-tokens", type=float, default=None,
                              help="expected size of the summary (default: 10%% of history)")
          parser.add_argument("--growth-per-turn", type=float, default=0.0,
                              help="tokens the history grows each turn (default: 0)")
          parser.add_argument("--ttl", choices=["5m", "1h"], default="5m",
                              help="cache TTL, sets the write multiplier (default: 5m)")
          parser.add_argument("--context-window", type=float, default=1_000_000,
                              help="model context window in tokens (default: 1000000)")
          parser.add_argument("--json", action="store_true",
                              help="emit the JSON envelope on stdout")
          parser.add_argument("-q", "--quiet", action="store_true",
                              help="suppress stderr framing/workings")
          args = parser.parse_args(argv)
      
          jm, quiet = args.json, args.quiet
      
          # --- validate (agents fabricate plausible inputs; §6 of the resource protocol)
          if args.history_tokens <= 0:
              fail("VALIDATION", "--history-tokens must be positive",
                   {"got": args.history_tokens}, jm, EXIT_VALIDATION)
          if args.turns_remaining < 0:
              fail("VALIDATION", "--turns-remaining cannot be negative",
                   {"got": args.turns_remaining}, jm, EXIT_VALIDATION)
          if args.base_rate <= 0:
              fail("VALIDATION", "--base-rate must be positive",
                   {"got": args.base_rate}, jm, EXIT_VALIDATION)
          if args.growth_per_turn < 0:
              fail("VALIDATION", "--growth-per-turn cannot be negative",
                   {"got": args.growth_per_turn}, jm, EXIT_VALIDATION)
          if args.context_window <= 0:
              fail("VALIDATION", "--context-window must be positive",
                   {"got": args.context_window}, jm, EXIT_VALIDATION)
      
          rate = args.base_rate
          out_rate = args.output_rate if args.output_rate is not None else rate * OUTPUT_RATE_RATIO
          if out_rate <= 0:
              fail("VALIDATION", "--output-rate must be positive", {"got": out_rate},
                   jm, EXIT_VALIDATION)
          summary = (args.summary_tokens if args.summary_tokens is not None
                     else args.history_tokens * 0.10)
          if summary <= 0:
              fail("VALIDATION", "--summary-tokens must be positive", {"got": summary},
                   jm, EXIT_VALIDATION)
          if summary >= args.history_tokens:
              fail("VALIDATION", "--summary-tokens must be smaller than the history "
                                 "(a 'summary' that isn't smaller saves nothing)",
                   {"summary": summary, "history": args.history_tokens},
                   jm, EXIT_VALIDATION)
      
          write_mult = CACHE_WRITE_1H if args.ttl == "1h" else CACHE_WRITE_5M
          turns = args.turns_remaining
      
          note(TERM.hdr("append vs compact"), quiet)
      
          # --- Constraint A: does it even fit? Checked first; cost is moot if not.
          projected = args.history_tokens + args.growth_per_turn * turns
          overflows = projected > args.context_window
      
          # --- Path 1: append. Carry the growing history at cache-read price.
          cost_append = cached_carry_cost(args.history_tokens, turns,
                                          args.growth_per_turn, rate)
      
          # --- Path 2: compact. Summarise now, then carry the summary.
          #   a) one summarisation call: history read from cache, summary generated
          per_summarise = (args.history_tokens * CACHE_READ_MULTIPLIER * rate / PER_MTOK
                           + summary * out_rate / PER_MTOK)
          #   b) writing the new (shorter) prefix into the cache, once per compaction
          per_rewrite = summary * write_mult * rate / PER_MTOK
          #   c) HOW MANY compactions. With growth, the summary climbs back to the
          #      original size after (history - summary)/growth turns and a real
          #      system compacts again. Charging a single one-shot rewrite flatters
          #      compaction badly on long agentic sessions (~2.2x on 60 turns at
          #      4k/turn), which is the regime where people actually reach for it.
          if args.growth_per_turn > 0:
              regrow_turns = (args.history_tokens - summary) / args.growth_per_turn
              compactions = max(1, 1 + int(turns / regrow_turns)) if turns > 0 else 0
          else:
              compactions = 1 if turns > 0 else 0
          cost_summarise = per_summarise * compactions
          cost_rewrite = per_rewrite * compactions
          #   d) carrying the summary for the remaining turns, growing as before
          cost_carry = cached_carry_cost(summary, turns, args.growth_per_turn, rate)
          cost_compact = cost_summarise + cost_rewrite + cost_carry
      
          delta = cost_append - cost_compact          # >0 means compaction is cheaper
          # Break-even: turns needed to repay ONE compaction's fixed cost. With
          # growth the cycle repeats, so this is a per-cycle repayment period, not a
          # whole-session verdict -- the verdict below compares the full modelled
          # costs including every compaction. Per-turn saving is the token difference
          # carried at cache-read price.
          per_turn_saving = ((args.history_tokens - summary)
                             * CACHE_READ_MULTIPLIER * rate / PER_MTOK)
          fixed_cost = per_summarise + per_rewrite   # one compaction's fixed cost
          breakeven = (fixed_cost / per_turn_saving) if per_turn_saving > 0 else float("inf")
      
          if overflows:
              verdict, reason = "compact", "context ceiling: projected history overflows the window"
          elif delta > 0:
              verdict, reason = "compact", "cost ceiling: compaction is cheaper over the remaining turns"
          else:
              verdict, reason = "append", "keep everything: appending is cheaper over the remaining turns"
      
          note(f"  projected history at turn {turns}: {projected:,.0f} tokens "
               f"(window {args.context_window:,.0f})", quiet)
          note(f"  compactions modelled: {compactions}"
               + (" (history regrows at --growth-per-turn)" if args.growth_per_turn > 0
                  else " (no growth given -> one-shot)"), quiet)
          note(f"  append  ${cost_append:.4f}   compact ${cost_compact:.4f}"
               f"   (summarise ${cost_summarise:.4f} + rewrite ${cost_rewrite:.4f}"
               f" + carry ${cost_carry:.4f})", quiet)
          note(f"  break-even at ~{breakeven:.1f} turns per compaction cycle "
               f"(you have {turns} turns, {compactions} compaction(s) modelled)", quiet)
          note("  note: break-even is scale-invariant - size and price cancel out; it "
               "tracks the summary ratio, not how big or costly the conversation is.",
               quiet)
          if verdict == "append":
              note(f"{TERM.mark('ok')} APPEND. {reason}.", quiet)
              note("  Cheaper levers before compaction: cap tool output at the boundary, "
                   "move payloads to files, verify the cache is actually hitting.", quiet)
          else:
              note(f"{TERM.mark('warn')} COMPACT. {reason}.", quiet)
              note("  Still try the cache-preserving levers first (tool-output caps, "
                   "payloads to files) - they do not rewrite the prefix.", quiet)
              note("  RECALL IS NOT PRICED HERE. Cheaper is not the same as better: "
                   "summarisation has measured 92-100% -> 38-58% on planted-fact "
                   "recall. Probe yours (assets/recall-probe.py) before acting.", quiet)
      
          data = {
              "verdict": verdict,
              "reason": reason,
              "context_ceiling_hit": overflows,
              "inputs": {
                  "history_tokens": args.history_tokens,
                  "turns_remaining": turns,
                  "base_rate_per_mtok": rate,
                  "output_rate_per_mtok": out_rate,
                  "summary_tokens": summary,
                  "growth_per_turn": args.growth_per_turn,
                  "ttl": args.ttl,
                  "context_window": args.context_window,
              },
              "multipliers": {
                  "cache_read": CACHE_READ_MULTIPLIER,
                  "cache_write": write_mult,
              },
              "cost_usd": {
                  "append": round(cost_append, 6),
                  "compact": round(cost_compact, 6),
                  "compact_summarise_call": round(cost_summarise, 6),
                  "compact_cache_rewrite": round(cost_rewrite, 6),
                  "compact_carry": round(cost_carry, 6),
                  "delta_append_minus_compact": round(delta, 6),
              },
              "breakeven_turns_per_compaction": (
                  round(breakeven, 2) if breakeven != float("inf") else None),
              "compactions_modelled": compactions,
              "projected_history_tokens": projected,
              "caveats": [
                  "models cost only - recall loss from summarisation is not priced",
                  "break-even turns is scale-invariant: driven by the summary ratio, "
                  "the output/input price ratio and the write multiplier - not by "
                  "history size or price level",
                  "break-even is the repayment period for ONE compaction; the verdict "
                  "compares full modelled costs across all modelled compactions",
              ],
          }
      
          if jm:
              print(json.dumps({"data": data,
                                "meta": {"schema": SCHEMA, "status": verdict}}))
          else:
              print(f"{verdict}\tappend=${cost_append:.4f}\tcompact=${cost_compact:.4f}"
                    f"\tbreakeven_turns_per_compaction={breakeven:.1f}"
                    f"  compactions={compactions}")
      
          return EXIT_COMPACT if verdict == "compact" else EXIT_OK
      
      
      if __name__ == "__main__":
          try:
              sys.exit(main(sys.argv[1:]))
          except KeyboardInterrupt:
              sys.exit(EXIT_USAGE)
      
  • tests
    • run.sh 23.6 KB
      #!/usr/bin/env bash
      # Self-test for claude-api-ops — fully offline: no network, no Anthropic API.
      #
      # Wraps the skill's §7 staleness verifier (scripts/check-model-table.py), which
      # guards the two fast-moving fact tables (SKILL.md "Current Models" and
      # references/caching-and-cost.md cache-minimums) against silent drift. Contract
      # (py_compile + --help), offline happy path against the shipped skill, the
      # --json §7 envelope, and a NEGATIVE proving the verifier actually rejects a bad
      # model id (a date-suffixed alias — exactly what SKILL.md forbids). --live is
      # NEVER invoked: it hits the Models API, and a network blip must never fail a PR.
      #
      # Usage:   bash tests/run.sh
      # Exit:    0 all pass, 1 one or more failures
      
      set -uo pipefail
      
      HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
      SKILL="$(dirname "$HERE")"
      V="$SKILL/scripts/check-model-table.py"
      
      # Pick a python that actually executes (the Windows Store python3 stub exists
      # on PATH but exits non-zero; probe by running it).
      PYTHON=""
      for c in python python3 py; do
        if command -v "$c" >/dev/null 2>&1 && "$c" -c "" >/dev/null 2>&1; then PYTHON="$c"; break; fi
      done
      
      SB="$(mktemp -d)"; trap 'rm -rf "$SB"' EXIT
      PASS=0; FAIL=0
      ok() { PASS=$((PASS+1)); printf '  PASS  %s\n' "$1"; }
      no() { FAIL=$((FAIL+1)); printf '  FAIL  %s\n' "$1"; }
      expect_exit() { [[ "$2" == "$3" ]] && ok "$1 (exit $3)" || no "$1 (want $2 got $3)"; }
      expect_has()  { case "$3" in *"$2"*) ok "$1";; *) no "$1 (missing '$2')";; esac; }
      
      echo "=== claude-api-ops self-test ==="
      
      if [[ -z "$PYTHON" ]]; then
        echo "  SKIP  no working python (verifier is python) — cannot test"
        [[ "$FAIL" -eq 0 ]] || exit 1
        exit 0
      fi
      
      # ── contract ──────────────────────────────────────────────────────────────────
      echo "-- contract --"
      "$PYTHON" -m py_compile "$V" 2>/dev/null && ok "py_compile check-model-table.py" || no "py_compile check-model-table.py"
      "$PYTHON" "$V" --help >/dev/null 2>&1; expect_exit "--help exits 0" 0 $?
      out="$("$PYTHON" "$V" --help 2>&1)"
      expect_has "--help has EXAMPLES" "EXAMPLES" "$out"
      "$PYTHON" "$V" --bogus >/dev/null 2>&1; expect_exit "unknown flag -> 2" 2 $?
      
      # ── offline structural mode (§7 seam: --offline default, --live advisory) ─────
      echo "-- offline structural --"
      "$PYTHON" "$V" --offline >/dev/null 2>&1; expect_exit "--offline clean on shipped skill" 0 $?
      out="$("$PYTHON" "$V" --offline --json 2>/dev/null)"
      expect_has "--offline --json envelope schema" '"schema": "claude-mods.claude-api-ops.model-table/v1"' "$out"
      expect_has "--offline --json consistent" '"consistent": true' "$out"
      
      # ── negative: a date-suffixed alias must be rejected (exit 4 VALIDATION) ──────
      # Run the verifier from a doctored copy; never mutate the shipped skill.
      echo "-- negative --"
      cp -r "$SKILL" "$SB/copy"
      # Append the date suffix SKILL.md explicitly forbids ("Never append date
      # suffixes"). The verifier's DATE_SUFFIX_RE must flag it as VALIDATION drift.
      # Target the table cell uniquely (the prose never pairs "Opus 5 |" with the
      # backticked id) so the edit is surgical. If the lineup changes and this string
      # stops matching, the copy stays clean and the exit-4 assertion below fails loudly
      # rather than silently passing on an unmodified file.
      "$PYTHON" - "$SB/copy/SKILL.md" <<'PY'
      import pathlib, sys
      p = pathlib.Path(sys.argv[1])
      t = p.read_text(encoding="utf-8")
      src = "Opus 5 | `claude-opus-5`"
      assert src in t, "negative-test fixture no longer matches SKILL.md model table"
      t = t.replace(src, "Opus 5 | `claude-opus-5-20260724`")
      p.write_text(t, encoding="utf-8")
      PY
      "$PYTHON" "$SB/copy/scripts/check-model-table.py" --offline >"$SB/neg.out" 2>&1
      expect_exit "--offline flags date-suffixed id -> 4" 4 $?
      expect_has "finding names the date suffix" "date suffix" "$(cat "$SB/neg.out")"
      
      # ── context-engineering layer (offline cross-file tripwires) ─────────────────
      # The verifier also guards the cache-economics constants that are now stated in
      # more than one file, the verification date stamps on the doctrine references,
      # and SKILL.md <-> references/ citation integrity. Each gets a NEGATIVE below:
      # a passing check proves nothing unless it can be made to fail.
      echo "-- context-engineering checks --"
      out="$("$PYTHON" "$V" --offline --json -q 2>/dev/null)"
      expect_has "--json reports cache_constants" '"cache_constants"' "$out"
      expect_has "--json reports date_stamps" '"date_stamps"' "$out"
      expect_has "--json reports reference_files" '"reference_files"' "$out"
      expect_has "--json names the context-management beta" 'context-management-2025-06-27' "$out"
      
      # Fresh sandbox copy per negative: validate_offline runs the model-table checks
      # BEFORE these, so a copy poisoned by an earlier negative would exit 4 on the
      # wrong finding and the assertion would pass for the wrong reason.
      n=0
      fresh_copy() { n=$((n+1)); rm -rf "$SB/c$n"; cp -r "$SKILL" "$SB/c$n"; echo "$SB/c$n"; }
      
      # NEGATIVE 1: desync a cache constant in exactly one file (0.1x -> 0.15x in
      # compaction.md). The real drift this guards: someone updates a multiplier in
      # the doc they happen to be editing and leaves the other three stale.
      C="$(fresh_copy)"
      "$PYTHON" - "$C/references/compaction.md" <<'PY'
      import pathlib, re, sys
      p = pathlib.Path(sys.argv[1]); t = p.read_text(encoding="utf-8")
      # Cover every spelling the verifier's pattern accepts: "0.1x", "0.1×", "0.1 ×".
      p.write_text(re.sub(r"0\.1(\s*)([x×])", r"0.15\1\2", t), encoding="utf-8")
      PY
      "$PYTHON" "$C/scripts/check-model-table.py" --offline >"$SB/n1.out" 2>&1
      expect_exit "desynced cache multiplier -> 4" 4 $?
      expect_has "finding names the constant" "cache_read_multiplier" "$(cat "$SB/n1.out")"
      
      # NEGATIVE 2: strip the verification date stamp from a doctrine reference.
      C="$(fresh_copy)"
      "$PYTHON" - "$C/references/context-engineering.md" <<'PY'
      import pathlib, re, sys
      p = pathlib.Path(sys.argv[1]); t = p.read_text(encoding="utf-8")
      p.write_text(re.sub(r"(?i)verified", "checked", t), encoding="utf-8")
      PY
      "$PYTHON" "$C/scripts/check-model-table.py" --offline >"$SB/n2.out" 2>&1
      expect_exit "missing verification date stamp -> 4" 4 $?
      expect_has "finding names the undated file" "context-engineering.md" "$(cat "$SB/n2.out")"
      
      # NEGATIVE 3: an uncited reference file (dead weight the router never finds --
      # SKILL-RESOURCE-PROTOCOL.md §1).
      C="$(fresh_copy)"
      printf '# orphan\n' > "$C/references/orphan-doc.md"
      "$PYTHON" "$C/scripts/check-model-table.py" --offline >"$SB/n3.out" 2>&1
      expect_exit "uncited reference file -> 4" 4 $?
      expect_has "finding names the orphan" "orphan-doc.md" "$(cat "$SB/n3.out")"
      
      # NEGATIVE 4: a cited-but-missing reference (broken link in SKILL.md).
      C="$(fresh_copy)"
      rm -f "$C/references/compaction.md"
      "$PYTHON" "$C/scripts/check-model-table.py" --offline >"$SB/n4.out" 2>&1
      expect_exit "cited-but-missing reference -> 4" 4 $?
      
      # NEGATIVE 5: an uncited script/asset (same protocol rule as references, but
      # those are cited by basename in prose rather than by markdown link).
      C="$(fresh_copy)"
      printf '# orphan\n' > "$C/assets/orphan-asset.py"
      "$PYTHON" "$C/scripts/check-model-table.py" --offline >"$SB/n5.out" 2>&1
      expect_exit "uncited asset -> 4" 4 $?
      expect_has "finding names the uncited asset" "orphan-asset.py" "$(cat "$SB/n5.out")"
      
      # ── context-budget calculator ────────────────────────────────────────────────
      # The doctrine's break-even arithmetic, executable. Its VERDICT is the contract:
      # exit 0 = append wins, exit 10 = cost favours compaction. Both directions are
      # asserted, because a calculator that can only say one thing is not a calculator.
      echo "-- context-budget --"
      CB="$SKILL/scripts/context-budget.py"
      "$PYTHON" -m py_compile "$CB" 2>/dev/null && ok "py_compile context-budget.py" || no "py_compile context-budget.py"
      "$PYTHON" "$CB" --help >/dev/null 2>&1; expect_exit "context-budget --help exits 0" 0 $?
      expect_has "context-budget --help has EXAMPLES" "EXAMPLES" "$("$PYTHON" "$CB" --help 2>&1)"
      "$PYTHON" "$CB" --bogus >/dev/null 2>&1; expect_exit "context-budget unknown flag -> 2" 2 $?
      
      # Short session: the doctrine's default answer. Must be exit 0 (append).
      "$PYTHON" "$CB" --history-tokens 25000 --turns-remaining 5 --base-rate 0.30 -q >/dev/null 2>&1
      expect_exit "short session -> append (0)" 0 $?
      # Deep session: enough remaining turns to repay the rewrite. Must be exit 10.
      "$PYTHON" "$CB" --history-tokens 120000 --turns-remaining 40 --base-rate 2.00 -q >/dev/null 2>&1
      expect_exit "deep session -> compaction indicated (10)" 10 $?
      # Context ceiling binds regardless of cost.
      "$PYTHON" "$CB" --history-tokens 900000 --turns-remaining 10 --growth-per-turn 50000 \
        --context-window 1000000 -q >/dev/null 2>&1
      expect_exit "context ceiling -> 10" 10 $?
      
      out="$("$PYTHON" "$CB" --history-tokens 25000 --turns-remaining 5 --base-rate 0.30 --json -q 2>/dev/null)"
      expect_has "context-budget --json envelope schema" '"schema": "claude-mods.claude-api-ops.context-budget/v1"' "$out"
      expect_has "context-budget --json verdict" '"verdict": "append"' "$out"
      # The recall caveat must ride along in the machine-readable output: an agent
      # acting on the verdict alone would otherwise treat "cheaper" as "better".
      expect_has "context-budget --json carries the recall caveat" 'recall loss' "$out"
      
      # REGRESSION: the calculator originally charged ONE compaction. With growth the
      # summary regrows and a real system compacts again -- assuming one-shot
      # understated compaction's fixed cost ~2.2x on a 60-turn, 4k/turn session, i.e.
      # it was biased TOWARD compaction, the opposite of the doctrine's default.
      out="$("$PYTHON" "$CB" --history-tokens 120000 --turns-remaining 60 --base-rate 2.00         --growth-per-turn 4000 --json -q 2>/dev/null)"
      expect_has "growing session models >1 compaction" '"compactions_modelled": 3' "$out"
      out="$("$PYTHON" "$CB" --history-tokens 120000 --turns-remaining 60 --base-rate 2.00 --json -q 2>/dev/null)"
      expect_has "no growth -> one-shot compaction" '"compactions_modelled": 1' "$out"
      # Break-even is a per-compaction repayment period, not a session verdict; the
      # key name has to say so or readers compare it against total turns and misread.
      expect_has "break-even key is scoped per compaction" '"breakeven_turns_per_compaction"' "$out"
      
      # Input validation (resource protocol §6 - agents fabricate plausible inputs).
      for bad in "--history-tokens -5 --turns-remaining 10" \
                 "--history-tokens 1000 --turns-remaining -1" \
                 "--history-tokens 1000 --turns-remaining 5 --base-rate 0" \
                 "--history-tokens 1000 --turns-remaining 5 --summary-tokens 5000"; do
        "$PYTHON" "$CB" $bad >/dev/null 2>&1
        expect_exit "rejects bad input ($bad)" 4 $?
      done
      
      # ── cache-correct loop asset ─────────────────────────────────────────────────
      # The asset's only real logic is breakpoint placement, and getting it wrong is
      # silent (a missed cache costs money and raises no error). Exercise it directly
      # with the anthropic SDK stubbed out - no network, no SDK install needed.
      echo "-- cached-agent-loop --"
      "$PYTHON" -m py_compile "$SKILL/assets/cached-agent-loop.py" 2>/dev/null \
        && ok "py_compile cached-agent-loop.py" || no "py_compile cached-agent-loop.py"
      "$PYTHON" -m py_compile "$SKILL/assets/recall-probe.py" 2>/dev/null \
        && ok "py_compile recall-probe.py" || no "py_compile recall-probe.py"
      
      "$PYTHON" - "$SKILL/assets/cached-agent-loop.py" >"$SB/bp.out" 2>&1 <<'PY'
      import sys, types, importlib.util
      stub = types.ModuleType("anthropic"); stub.Anthropic = lambda *a, **k: None
      sys.modules["anthropic"] = stub
      spec = importlib.util.spec_from_file_location("loop", sys.argv[1])
      m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m)
      
      def marks(msgs):
          return [(i, j) for i, msg in enumerate(msgs)
                  for j, b in enumerate(msg["content"])
                  if isinstance(b, dict) and "cache_control" in b]
      
      # The newest turn must always carry a breakpoint, or hits never accrue.
      msgs = [{"role": "user", "content": [{"type": "text", "text": "hi"}]}]
      m.place_message_breakpoints(msgs)
      assert marks(msgs) == [(0, 0)], f"short: {marks(msgs)}"
      
      # A tool-heavy turn appending 40 blocks must not exceed the API's 4-breakpoint
      # limit (one is spent on the system block) and must keep consecutive anchors
      # inside the 20-block backward search, or the lookback silently misses.
      msgs = [{"role": "user", "content": [{"type": "text", "text": f"b{i}"} for i in range(40)]}]
      m.place_message_breakpoints(msgs)
      got = marks(msgs)
      assert len(got) <= m.MAX_BREAKPOINTS - 1, f"too many breakpoints: {got}"
      assert got[-1] == (0, 39), f"newest block unmarked: {got}"
      gaps = [got[i + 1][1] - got[i][1] for i in range(len(got) - 1)]
      assert all(g <= m.LOOKBACK_BLOCKS for g in gaps), f"anchor gap exceeds lookback: {gaps}"
      assert m.BREAKPOINT_EVERY < m.LOOKBACK_BLOCKS, "anchor spacing must fit the window"
      
      # Idempotent: called every turn, it must not accumulate stale markers.
      before = marks(msgs); m.place_message_breakpoints(msgs)
      assert marks(msgs) == before, "not idempotent"
      
      # Tool output is capped at the boundary (the cache-preserving lever).
      assert len(m.capped("x" * 99999)) < 99999, "capped() did not truncate"
      assert m.capped("short") == "short", "capped() mangled a short result"
      
      # === REGRESSION: the shipped loop once placed ZERO breakpoints ===============
      # In a real loop, assistant turns are `response.content` -- SDK block OBJECTS,
      # not dicts -- and cache_control can only be set on a dict. The first version
      # of this asset skipped non-dict blocks silently, so every anchor landed on an
      # assistant block and was dropped: no error, no warning, no caching, full price
      # forever. The original test missed it because it only ever fed dicts. These
      # cases feed the shapes a real loop actually produces.
      class SDKBlock:                       # stands in for anthropic.types.TextBlock
          def __init__(self, t): self.type = "text"; self.text = t
          def model_dump(self, exclude_none=False): return {"type": "text", "text": self.text}
      
      def mixed(n_turns):
          msgs = []
          for k in range(n_turns):
              msgs.append({"role": "user",
                           "content": [{"type": "text", "text": f"u{k}-{j}"} for j in range(4)]})
              msgs.append({"role": "assistant", "content": [SDKBlock(f"a{k}-{j}") for j in range(4)]})
          return msgs
      
      # Un-normalised SDK objects: must still place markers (walks back to a dict).
      msgs = mixed(6)
      placed = m.place_message_breakpoints(msgs)
      assert placed > 0, "REGRESSION: zero breakpoints placed on an SDK-object conversation"
      assert placed <= m.MAX_BREAKPOINTS - 1, f"too many breakpoints: {placed}"
      
      # Normalised via to_blocks() -- the path main() takes. The newest block must
      # carry a marker, and anchors must stay inside the 20-block lookback.
      msgs = [{"role": m0["role"], "content": m.to_blocks(m0["content"])} for m0 in mixed(6)]
      placed = m.place_message_breakpoints(msgs)
      got = marks(msgs)
      assert placed == len(got), f"reported {placed} but set {len(got)}"
      assert got[-1] == (len(msgs) - 1, len(msgs[-1]["content"]) - 1), f"newest block unmarked: {got}"
      flat_pos = []
      for mi, msg in enumerate(msgs):
          for bi in range(len(msg["content"])):
              flat_pos.append((mi, bi))
      idxs = [flat_pos.index(g) for g in got]
      gaps = [idxs[i + 1] - idxs[i] for i in range(len(idxs) - 1)]
      assert all(g <= m.LOOKBACK_BLOCKS for g in gaps), f"anchor gap exceeds lookback: {gaps}"
      
      # String-shorthand content has no block to attach a marker to: warn, don't crash.
      assert m.place_message_breakpoints([{"role": "user", "content": "plain string"}]) == 0
      
      # to_blocks() normalises all three shapes a caller may hand it.
      assert m.to_blocks("hi") == [{"type": "text", "text": "hi"}]
      assert m.to_blocks([{"type": "text", "text": "hi"}]) == [{"type": "text", "text": "hi"}]
      assert m.to_blocks([SDKBlock("hi")]) == [{"type": "text", "text": "hi"}]
      print("OK")
      PY
      expect_has "breakpoint placement, capping and idempotence" "OK" "$(cat "$SB/bp.out")"
      
      # ── negative: a retired model id in a code sample (exit 4 VALIDATION) ─────────
      # The regression this check exists for: on 2026-08-30 the tables were retargeted
      # at the Claude 5 lineup and went green, then a sibling branch landed an asset
      # still pinned to claude-opus-4-8. Tables were consistent; the sample was wrong.
      echo "-- negative: retired id in a sample --"
      rm -rf "$SB/copy2"; cp -r "$SKILL" "$SB/copy2"
      "$PYTHON" - "$SB/copy2" <<'PY'
      import pathlib, sys
      p = pathlib.Path(sys.argv[1]) / "assets/cached-agent-loop.py"
      t = p.read_text(encoding="utf-8")
      assert 'MODEL = "claude-opus-5"' in t, "fixture no longer matches cached-agent-loop.py"
      p.write_text(t.replace('MODEL = "claude-opus-5"', 'MODEL = "claude-opus-4-8"'), encoding="utf-8")
      PY
      "$PYTHON" "$SB/copy2/scripts/check-model-table.py" --offline >"$SB/neg2.out" 2>&1
      expect_exit "retired id in a sample -> 4" 4 $?
      expect_has "finding names the file and fault" "cached-agent-loop.py" "$(cat "$SB/neg2.out")"
      expect_has "finding classifies it retired" "retired" "$(cat "$SB/neg2.out")"
      
      # The escape hatch must actually release the gate, or authors will delete it.
      "$PYTHON" - "$SB/copy2" <<'PY'
      import pathlib, sys
      p = pathlib.Path(sys.argv[1]) / "assets/cached-agent-loop.py"
      p.write_text(p.read_text(encoding="utf-8").replace(
          'MODEL = "claude-opus-4-8"', 'MODEL = "claude-opus-4-8"  # legacy-ok'), encoding="utf-8")
      PY
      "$PYTHON" "$SB/copy2/scripts/check-model-table.py" --offline >/dev/null 2>&1
      expect_exit "legacy-ok marker releases the gate" 0 $?
      
      # An id in neither the table nor the legacy list is a typo or a hallucination.
      rm -rf "$SB/copy3"; cp -r "$SKILL" "$SB/copy3"
      printf '\nUse claude-opus-7 for this.\n' >> "$SB/copy3/references/tool-use.md"
      "$PYTHON" "$SB/copy3/scripts/check-model-table.py" --offline >"$SB/neg3.out" 2>&1
      expect_exit "unknown model id -> 4" 4 $?
      expect_has "finding classifies it unknown" "unknown" "$(cat "$SB/neg3.out")"
      
      # Runtime warnings must be ASCII: this asset prints to a console that is cp1252
      # by default on Windows, where an em-dash renders as a replacement character.
      "$PYTHON" - "$SKILL/assets/cached-agent-loop.py" "$SKILL/assets/recall-probe.py" >"$SB/ascii.out" 2>&1 <<'PY'
      import ast, pathlib, sys
      bad = []
      for f in sys.argv[1:]:
          tree = ast.parse(pathlib.Path(f).read_text(encoding="utf-8"))
          for node in ast.walk(tree):
              if isinstance(node, ast.Call) and getattr(node.func, "id", "") == "print":
                  for a in ast.walk(node):
                      if isinstance(a, ast.Constant) and isinstance(a.value, str) \
                              and any(ord(c) > 127 for c in a.value):
                          bad.append((pathlib.Path(f).name, a.value[:40]))
      print("CLEAN" if not bad else f"NON-ASCII IN PRINT: {bad}")
      PY
      expect_has "asset runtime output is ASCII-safe" "CLEAN" "$(cat "$SB/ascii.out")"
      
      # ── recall-probe fairness ────────────────────────────────────────────────────
      # REGRESSION: the harness originally compacted on EVERY turn once the history
      # exceeded keep_recent, so the compact arm paid a summarisation call per turn
      # (21 API calls vs append's 12 over 10 turns). That inflates the arm under test
      # and would "confirm" the keep-everything result whatever the data said. The
      # cadence knob is the fix; assert it exists and is actually consulted.
      echo "-- recall-probe fairness --"
      RP="$SKILL/assets/recall-probe.py"
      grep -q 'compact_every' "$RP" && ok "recall-probe has a compaction cadence" \
        || no "recall-probe has a compaction cadence"
      grep -q 'turns_since_compaction >= compact_every' "$RP" \
        && ok "cadence actually gates compaction" || no "cadence actually gates compaction"
      grep -q '"--compact-every"' "$RP" && ok "cadence is user-tunable" || no "cadence is user-tunable"
      # Default must not be 1 -- that is the rigged configuration.
      "$PYTHON" - "$RP" >"$SB/rp.out" 2>&1 <<'PY'
      import ast, pathlib, sys
      tree = ast.parse(pathlib.Path(sys.argv[1]).read_text(encoding="utf-8"))
      for node in ast.walk(tree):
          if isinstance(node, ast.Call) and getattr(node.func, "attr", "") == "add_argument":
              if node.args and getattr(node.args[0], "value", "") == "--compact-every":
                  d = [k.value.value for k in node.keywords if k.arg == "default"]
                  print("DEFAULT", d[0] if d else "none")
      PY
      out="$(cat "$SB/rp.out")"
      case "$out" in
        "DEFAULT 1"|"DEFAULT none"|"") no "compact cadence default is unbiased (got '$out')" ;;
        *) ok "compact cadence default is unbiased ($out)" ;;
      esac
      
      # ── SKILL.md sanity ───────────────────────────────────────────────────────────
      echo "-- SKILL.md --"
      # CONTRACT (frontmatter shape): this suite asserts that SKILL.md's frontmatter
      # keeps `name: claude-api-ops` and a `when_to_use:` field, and that
      # len(description) + len(when_to_use) stays within the repo's 1000-char per-skill
      # cap enforced by tests/validate.sh. A description-trim or frontmatter cleanup
      # lane that removes `when_to_use` from this skill WILL break CI here -- that is
      # deliberate, and stated here so the edit site is not the first place you find out.
      grep -q '^name: claude-api-ops$' "$SKILL/SKILL.md" && ok "frontmatter name" || no "frontmatter name"
      grep -q '^when_to_use: ' "$SKILL/SKILL.md" && ok "frontmatter when_to_use present" || no "frontmatter when_to_use present"
      grep -q 'check-model-table.py' "$SKILL/SKILL.md" && ok "verifier cited from SKILL.md" || no "verifier cited from SKILL.md"
      
      # Description budget: mirrors tests/validate.sh's hard cap so this skill fails
      # in its own suite rather than only in the catalog-wide gate.
      combined="$("$PYTHON" - "$SKILL/SKILL.md" <<'PY'
      import pathlib, re, sys
      t = pathlib.Path(sys.argv[1]).read_text(encoding="utf-8")
      fm = t.split("---")[1]
      def field(k):
          m = re.search(r'^%s: "(.*)"$' % k, fm, re.M)
          return m.group(1) if m else ""
      print(len(field("description")) + len(field("when_to_use")))
      PY
      )"
      if [[ "$combined" -gt 0 && "$combined" -le 1000 ]]; then
        ok "description + when_to_use within 1000-char cap ($combined)"
      else
        no "description + when_to_use out of range (got '$combined', cap 1000)"
      fi
      
      # Body-size limit: SKILL-CREATION-PROTOCOL.md Step 3 caps the body at 500 lines;
      # depth belongs in references/*.md.
      lines="$(wc -l < "$SKILL/SKILL.md" | tr -d ' ')"
      [[ "$lines" -lt 500 ]] && ok "SKILL.md body under 500 lines ($lines)" \
                             || no "SKILL.md body is $lines lines (limit 500)"
      
      # The context-engineering content itself is cited and reachable.
      echo "-- context-engineering content --"
      grep -q '^## Context Engineering$' "$SKILL/SKILL.md" \
        && ok "Context Engineering section present" || no "Context Engineering section present"
      for r in context-engineering compaction; do
        [[ -f "$SKILL/references/$r.md" ]] && ok "references/$r.md exists" || no "references/$r.md exists"
        grep -q "(references/$r.md)" "$SKILL/SKILL.md" \
          && ok "references/$r.md cited from SKILL.md" || no "references/$r.md cited from SKILL.md"
      done
      # The load-bearing, counter-intuitive claim must survive edits: compaction is a
      # response to a NAMED constraint, not a default.
      grep -qi 'named constraint' "$SKILL/references/compaction.md" \
        && ok "compaction doctrine states the named-constraint rule" \
        || no "compaction doctrine states the named-constraint rule"
      grep -q 'clear_at_least' "$SKILL/references/compaction.md" \
        && ok "context_management params documented" || no "context_management params documented"
      # The three tiers are the spine of the doctrine reference.
      grep -q 'Tier 1' "$SKILL/references/context-engineering.md" \
        && ok "three-tier model documented" || no "three-tier model documented"
      
      echo ""
      echo "=== $PASS passed, $FAIL failed ==="
      [[ "$FAIL" -eq 0 ]] || exit 1
      exit 0
      
  • SKILL.md 26.4 KB
    ---
    name: claude-api-ops
    description: "Building applications ON Claude - the Anthropic API and Claude Agent SDK. Use for: anthropic api, claude api, messages api, tool use, function calling, prompt caching, agent sdk, claude-agent-sdk, structured output, json schema output, batches api, extended thinking, adaptive thinking, model selection, claude pricing, build claude agent, anthropic sdk, stop_reason handling, streaming claude, token counting, cache_control, output_config, tool_choice, agentic loop, rate limits anthropic, context engineering, context window budget, compaction, context editing, context_management, clear_tool_uses, memory tool, context rot, tool result bloat, subagent context isolation."
    when_to_use: "Use when building applications on the Anthropic API or Claude Agent SDK — e.g. 'add tool use to my Claude app', 'set up prompt caching', 'which Claude model should I use', 'handle stop_reason / streaming', 'should I compact this agent context'."
    license: MIT
    allowed-tools: "Read Write Bash WebFetch"
    metadata:
      author: claude-mods
      related-skills: mcp-ops
    ---
    
    # Claude API Operations
    
    Building applications and agents on Anthropic's API: the Messages API, tool use,
    prompt caching, structured outputs, batches, thinking/effort, and the Claude
    Agent SDK. For developers writing apps *against* the API — not for using Claude
    Code itself.
    
    **API surfaces move fast.** Model IDs, parameters, and betas in this skill were
    verified against platform.claude.com (2026-08). When in doubt — especially for
    "latest model" or pricing questions — verify with WebFetch against
    `https://platform.claude.com/docs/en/about-claude/models/overview.md` or query
    the Models API (`client.models.list()`).
    
    ## Current Models (verified 2026-08)
    
    | Model | ID (exact, no date suffix) | Context | Max Output | Input $/MTok | Output $/MTok |
    |---|---|---|---|---|---|
    | Claude Fable 5 | `claude-fable-5` | 1M | 128K | $10.00 | $50.00 |
    | Claude Opus 5 | `claude-opus-5` | 1M | 128K | $5.00 | $25.00 |
    | Claude Sonnet 5 | `claude-sonnet-5` | 1M | 128K | $2.00 | $10.00 |
    | Claude Haiku 4.5 | `claude-haiku-4-5` | 200K | 64K | $1.00 | $5.00 |
    
    Use these alias IDs verbatim. **Never append date suffixes** (`claude-sonnet-5-20260630`
    is wrong → 404). Haiku 4.5 is the one current model with a *dated* snapshot id
    (`claude-haiku-4-5-20251001`) behind its alias; from the 4.6 generation on, the dateless
    id **is** the pinned snapshot.
    
    **Legacy (still available, no longer current):** `claude-opus-4-8`, `claude-opus-4-7`,
    `claude-opus-4-6`, `claude-opus-4-5`, `claude-sonnet-4-6`, `claude-sonnet-4-5`. Migrating
    off one: `https://platform.claude.com/docs/en/models/opus-5/migration-guide.md` (or run
    `/claude-api migrate` in Claude Code). Live capability lookup:
    `client.models.retrieve("claude-opus-5")` → `.max_input_tokens`, `.max_tokens`,
    `.capabilities` dict.
    
    ## Model Selection Decision Tree
    
    ```
    What is the workload?
    │
    ├─ Hardest problems, long-horizon agents, deep research, ceiling intelligence
    │  └─ claude-fable-5 (premium ceiling) or claude-opus-5 (default flagship)
    │
    ├─ Agentic coding, tool-heavy workflows, production assistants
    │  └─ claude-opus-5 (quality) or claude-sonnet-5 (speed/cost balance)
    │
    ├─ High-volume production: summarization, RAG answers, extraction
    │  └─ claude-sonnet-5
    │
    ├─ Classification, routing, simple Q&A, latency-critical
    │  └─ claude-haiku-4-5
    │
    └─ Subagents inside a larger system
       └─ One tier below the orchestrator (Opus loop → Sonnet/Haiku workers)
    ```
    
    Tiering rule: route by task difficulty, not by uniform default. An Opus
    orchestrator dispatching Haiku classifiers is routinely 5-10x cheaper than
    Opus-everywhere with no quality loss on the simple legs.
    
    ## Which Surface? (API vs Agent SDK vs Batches)
    
    | Need | Use | Why |
    |---|---|---|
    | One request → one response (classify, summarize, extract, Q&A) | **Messages API** | Simplest; full control |
    | Multi-step pipeline, your code controls the logic | **Messages API + tool use** | You own the loop |
    | Custom agent with your own tools, your infra | **Messages API + tool use** (manual loop or SDK tool runner) | Max flexibility |
    | Agent that reads/edits files, runs commands, searches — without building tools | **Claude Agent SDK** | Claude Code's tools + agent loop as a library |
    | CI/CD automation, coding agents, production agent apps | **Claude Agent SDK** | Built-in tools, hooks, sessions, MCP |
    | Large non-urgent workloads (eval runs, backfills, bulk extraction) | **Batches API** | 50% discount, ≤24h turnaround |
    | Hosted agent, Anthropic runs loop + sandbox | **Managed Agents** (beta) | No infra; see official docs |
    
    Rule of thumb: start at the simplest tier. Reach for an agent only when the
    task is genuinely open-ended (multi-step, hard to fully specify, errors
    recoverable, value justifies cost).
    
    ## Messages API Quick Start
    
    Everything goes through `POST /v1/messages`. Headers: `x-api-key`,
    `anthropic-version: 2023-06-01`, `content-type: application/json`.
    
    ```python
    # pip install anthropic
    import anthropic
    
    client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY
    
    response = client.messages.create(
        model="claude-opus-5",
        max_tokens=16000,
        system="You are a concise technical assistant.",
        messages=[{"role": "user", "content": "Explain CRDTs in one paragraph."}],
    )
    for block in response.content:        # content is a list of typed blocks
        if block.type == "text":          # always check .type before .text
            print(block.text)
    print(response.stop_reason, response.usage.input_tokens, response.usage.output_tokens)
    ```
    
    ```typescript
    // npm install @anthropic-ai/sdk
    import Anthropic from "@anthropic-ai/sdk";
    
    const client = new Anthropic();
    
    const response = await client.messages.create({
      model: "claude-opus-5",
      max_tokens: 16000,
      messages: [{ role: "user", content: "Explain CRDTs in one paragraph." }],
    });
    for (const block of response.content) {
      if (block.type === "text") console.log(block.text);  // narrow the union first
    }
    ```
    
    Streaming (default to it for long outputs — non-streaming above ~16K
    `max_tokens` risks SDK HTTP timeouts):
    
    ```python
    with client.messages.stream(model="claude-opus-5", max_tokens=64000,
                                messages=[{"role": "user", "content": "Write a long report"}]) as stream:
        for text in stream.text_stream:
            print(text, end="", flush=True)
        final = stream.get_final_message()   # full Message after streaming
    ```
    
    Full params, response shape, stop reasons, errors, retries, rate limits:
    [references/messages-api.md](references/messages-api.md)
    
    ## Thinking & Effort (quick reference)
    
    - **Adaptive thinking is ON BY DEFAULT on Fable 5 / Opus 5 / Sonnet 5** — send no
      `thinking` field and you still get (and pay for) thinking. On the legacy
      4.6–4.8 models it stays off until you set `thinking: {"type": "adaptive"}`.
    - **Manual budgets are gone.** `{"type": "enabled", "budget_tokens": N}` returns a
      **400 on Opus 4.7 and every later model** (Opus 5, Sonnet 5, Fable 5 included);
      deprecated on Opus 4.6 / Sonnet 4.6. Control depth with `effort`, not tokens.
    - **Turning thinking off:** Sonnet 5 accepts `{"type": "disabled"}`. Opus 5 accepts
      it only at effort `high` or below — pairing it with `xhigh`/`max` is a **400**.
      Fable 5 **rejects it outright**; thinking there is unconditional, so budget for it.
    - **Effort (GA):** `output_config: {"effort": "low" | "medium" | "high" | "xhigh" | "max"}`
      — nested in `output_config`, not top-level. Default `high` (identical to omitting
      it). `xhigh`: Fable 5, Opus 5, Opus 4.8/4.7, **Sonnet 5**. `max`: those plus Opus 4.6
      and Sonnet 4.6. Haiku 4.5 does not support `effort` at all.
    - **Sampling params removed on Opus 4.7 and later** (so Opus 5, Sonnet 5, Fable 5):
      `temperature`, `top_p`, `top_k` all return 400 — and the Python SDK v1.0+ doesn't
      define them, so passing them raises `TypeError`. Steer with prompting + effort.
    - **Forced tool_choice is fine with adaptive thinking.** The auto/none-only
      restriction applies to *manual* extended thinking (`{"type": "enabled"}`) only;
      adaptive mode — including the models where it's on by default — accepts
      `{"type": "any"}` and `{"type": "tool", ...}`.
    - Thinking text is **omitted by default** on Fable 5 / Opus 5 / Sonnet 5 / Opus 4.8 /
      4.7 — opt in with `thinking: {"type": "adaptive", "display": "summarized"}` if you
      surface reasoning to users. Either way the blocks are billed, and must be echoed
      back **unmodified** (empty `thinking` field included) in a tool-use loop, or the
      next request 400s.
    
    Details and gotchas: [references/structured-outputs.md](references/structured-outputs.md)
    (thinking interplay) and [references/messages-api.md](references/messages-api.md).
    
    ## Tool Use (quick reference)
    
    ```python
    tools = [{
        "name": "get_weather",
        "description": "Get current weather. Call when the user asks about weather conditions.",
        "input_schema": {
            "type": "object",
            "properties": {"location": {"type": "string", "description": "City, e.g. Paris"}},
            "required": ["location"],
        },
    }]
    response = client.messages.create(model="claude-opus-5", max_tokens=16000,
                                      tools=tools, messages=messages)
    if response.stop_reason == "tool_use":
        ...  # execute, send tool_result back, loop
    ```
    
    `tool_choice`: `{"type": "auto"}` (default) | `{"type": "any"}` | `{"type":
    "tool", "name": "..."}` | `{"type": "none"}`. Add
    `"disable_parallel_tool_use": true` to force at most one call per response.
    
    The agentic loop, parallel tool results, `pause_turn`, `is_error`, server-side
    tools, and SDK tool runners: [references/tool-use.md](references/tool-use.md)
    
    ## Cost Optimization Checklist
    
    Work top-down; each item is independent:
    
    - [ ] **Right-size the model.** Haiku for classification/routing, Sonnet for
          volume work, Opus/Fable for the hard 10%. Largest single lever.
    - [ ] **Prompt caching** on stable prefixes (system prompt, tool defs, big docs):
          `cache_control: {"type": "ephemeral"}`. Reads cost ~0.1x; up to 90% savings.
          Verify with `usage.cache_read_input_tokens > 0` — zero means a silent
          invalidator (timestamp in system prompt, unsorted JSON, varying tools).
    - [ ] **Batches API** for anything that can wait ≤24h: flat 50% off all tokens,
          stacks with caching.
    - [ ] **Cap output**: set `max_tokens` to what you need (256 for classification);
          stream + generous cap for long generation.
    - [ ] **Tune effort down** where quality allows: `medium` is often the sweet
          spot; `low` for subagents and simple tasks.
    - [ ] **Count before sending**: `client.messages.count_tokens(...)` (never
          tiktoken — it's OpenAI's tokenizer and undercounts Claude by 15-20%).
    - [ ] **Keep prefixes stable**: order requests `tools` → `system` → `messages`,
          volatile content last; don't swap tool sets or models mid-conversation.
    
    Mechanics, breakpoints, TTLs, batch lifecycle, tiering math:
    [references/caching-and-cost.md](references/caching-and-cost.md)
    
    ## Context Engineering
    
    Prompt engineering asks what to write in the prompt. **Context engineering asks what
    earns a place in the window on *this* call** — including everything that lands there
    without you typing it: tool definitions, tool results, retrieved documents, prior
    turns, thinking blocks. It is iterative (every inference) where prompt engineering is
    discrete (written once). Target: the smallest set of high-signal tokens that gets the
    outcome.
    
    The budget is real because attention degrades with length (**context rot** — n²
    pairwise relationships), not just because tokens cost money. A 1M window is a
    capacity, not a target.
    
    ### The three tiers
    
    Every candidate fact lives in exactly one place. Choosing deliberately is most of the job.
    
    | Tier | Where | Cost | Use when |
    |---|---|---|---|
    | **1 — In context** | `tools` / `system` / `messages`, every call | Paid every turn (≈0.1× cached) | It steers *most* turns |
    | **2 — On disk, read on demand** | A file the agent can read; only the **path** stays in context | Paid only when read | The agent can tell from a *name* that it needs this |
    | **3 — Retrieved** | Index / search tool behind a query | Paid only on a hit, plus a relevance gamble | The corpus is too large to enumerate |
    
    When a prompt is too big, **demote before you delete** — a path is ~10 tokens; the
    file it names may be 10,000.
    
    **This repo already runs on the tier-1/tier-2 split.** A skill's `description` is
    always resident (tier 1, so it must carry the routing signal); `SKILL.md` loads on a
    match; `references/*.md` load only when cited and needed. "Description is the
    trigger", "body under 500 lines", "one concept per reference", "every reference must
    be cited" are context-engineering rules wearing authoring clothes.
    
    ### Cache-aware prompt architecture
    
    Requests render `tools` → `system` → `messages`, and the cache is a **prefix match**.
    So **static prefix first, volatile content last** — put the `cache_control` breakpoint
    at the end of the stable part and let per-request content fall after it.
    
    Reordering a prompt destroys the cache **silently**: no error, just a different prefix
    hash, `cache_read_input_tokens: 0`, and a 1.25–2× bill where you expected 0.1×. The
    usage block is the only symptom, which is why asserting `cache_read_input_tokens > 0`
    in staging is a real test.
    
    ### The compaction decision
    
    **Under modern prompt caching, keeping the full history has been measured to beat
    summarisation on cost, latency AND recall at the same time.** A 2026 production-tutor
    evaluation (660 turns, 11 configurations) put keep-everything at 92–100% fact recall,
    $0.11/turn and 17 s TTFT, against 38–58% recall, $0.24/turn and 21 s for its
    clear-plus-summarise preset. Summarising rewrites the cached prefix and forfeits the
    0.1× discount — the cheap move is usually to **append**. (That study ran on a
    non-Claude model; what transfers is the *mechanism*, and Claude's flat 0.1× cache
    read makes it stronger, not weaker. Full caveats in
    [references/compaction.md](references/compaction.md).)
    
    So: **compact only as a deliberate response to a named constraint.**
    
    | Constraint | Diagnose | Try first |
    |---|---|---|
    | **Context ceiling** — it will not fit | Projected tokens > window | Cap tool output → payloads to files → server-side clearing |
    | **Cost ceiling** — the bill is unacceptable | Compare against *cached* cost, not uncached | **Verify the cache is hitting** → tier down → cap tool output |
    | **Latency target** — TTFT too slow at depth | Confirm growth is in the prefix | Cap tool output → lower `effort` → stream |
    
    Capping tool output at the tool boundary is the underrated lever: it shrinks context
    **without rewriting the cached prefix** (the same study measured −38% cost/turn with
    no recall loss). Clearing and summarising both break the cache; they are what people
    reach for first and should reach for last.
    
    First-party clearing is `context_management` (beta `context-management-2025-06-27`):
    `clear_tool_uses_20250919` and `clear_thinking_20251015`, applied server-side. Always
    set `clear_at_least` — it stops a trigger paying a full cache re-write to save a
    handful of tokens. Pair with the memory tool so durable conclusions are written out
    before raw material is cleared.
    
    ### Agentic specifics
    
    - **Tool results are the growth term**, not the system prompt. Design tools to return
      decisions, not dumps.
    - **Summarise vs write-to-file:** needed later *in full* → write to a file, return the
      path. Only the *conclusion* matters → summarise **at the tool boundary** (free of
      cache cost, unlike rewriting history after the fact).
    - **Sub-agents are context isolation**, not just parallelism: 80K tokens of
      exploration are billed once inside the child and discarded; the parent sees a
      ~1–2K-token distillation. Costs: cold cache in the child, a lossy hand-off. Skip it
      when the subtask needs most of the parent's context to make sense.
    
    Full doctrine — tiers, progressive disclosure, instrumentation:
    [references/context-engineering.md](references/context-engineering.md).
    Compaction economics, `context_management` parameters, memory tool:
    [references/compaction.md](references/compaction.md).
    For Claude Code's own context surface see the `claude-code-ops` skill; for
    prompts re-sent on a cadence, `loop-ops`; for cross-provider fan-out, `fleetflow`.
    
    ## Claude Agent SDK (quick reference)
    
    ```python
    # pip install claude-agent-sdk   (Python >= 3.10)
    import asyncio
    from claude_agent_sdk import query, ClaudeAgentOptions
    
    async def main():
        async for message in query(
            prompt="Find and fix the bug in auth.py",
            options=ClaudeAgentOptions(allowed_tools=["Read", "Edit", "Bash"]),
        ):
            if hasattr(message, "result"):
                print(message.result)
    
    asyncio.run(main())
    ```
    
    ```typescript
    // npm install @anthropic-ai/claude-agent-sdk
    import { query } from "@anthropic-ai/claude-agent-sdk";
    
    for await (const message of query({
      prompt: "Find and fix the bug in auth.ts",
      options: { allowedTools: ["Read", "Edit", "Bash"] },
    })) {
      if ("result" in message) console.log(message.result);
    }
    ```
    
    Built-in tools (Read/Write/Edit/Bash/Glob/Grep/WebSearch/WebFetch/...), hooks
    (`PreToolUse`, `PostToolUse`, ...), subagents, MCP servers, sessions
    (resume/fork), permission modes, and the SDK-vs-raw-API decision:
    [references/agent-sdk.md](references/agent-sdk.md)
    
    ## Common Pitfalls
    
    | Pitfall | Symptom | Fix |
    |---|---|---|
    | Date-suffixed or guessed model ID | 404 `not_found_error` | Use exact alias IDs from the table above |
    | `budget_tokens` on Opus 4.7+ (incl. Opus 5 / Sonnet 5 / Fable 5) | 400 | `thinking: {"type": "adaptive"}` + `effort` |
    | Assuming thinking is opt-in on Fable 5 / Opus 5 / Sonnet 5 | Unexpected thinking tokens billed | Adaptive thinking is on by default there; Fable 5 can't be disabled at all |
    | `thinking: {"type": "disabled"}` at `xhigh`/`max` on Opus 5 | 400 | Drop effort to `high` or below, or leave thinking on |
    | `temperature`/`top_p`/`top_k` on Opus 4.7+ | 400 (or `TypeError` on Python SDK v1.0+) | Remove; steer via prompt + `effort` |
    | `effort` on Haiku 4.5 | 400 | Haiku 4.5 doesn't support the parameter |
    | Rebuilding assistant turns in a tool loop (dropping empty `thinking` blocks) | 400 "thinking blocks cannot be modified" | Echo the content list back exactly as received |
    | Assistant-turn prefill on Opus 4.7+ models | 400 | `output_config.format` or system-prompt instruction |
    | Cache marker on <minimum prefix | Silent no-cache (`cache_creation_input_tokens: 0`) | Min 512-4096 tokens depending on model (see caching ref) |
    | Not handling `stop_reason: "tool_use"` | Agent "stops" after first tool call | Loop: execute tools, append `tool_result`, re-request |
    | Missing `tool_result` for a `tool_use` id | 400 on follow-up | One `tool_result` per `tool_use` block, ids matching |
    | Non-streaming with `max_tokens` > ~16K | SDK timeout / `ValueError` | Stream + `get_final_message()` / `finalMessage()` |
    | `output_format` top-level param | Deprecated | `output_config: {"format": {...}}` |
    | tiktoken for Claude token counts | 15-20%+ undercount | `messages.count_tokens` endpoint |
    | String-matching error messages | Fragile retries | Typed exceptions: `anthropic.RateLimitError` etc. |
    | Raw string-matching tool `input` | Breaks on escaping changes | Always `json.loads()` / use parsed `block.input` |
    | Compacting by reflex on a long conversation | Higher cost, worse recall than doing nothing | Name the constraint first; under caching, appending usually wins (see Context Engineering) |
    | `clear_tool_uses` without `clear_at_least` | A full cache re-write to reclaim a few hundred tokens | Set `clear_at_least` so each cache break is worth taking |
    
    ## Resources & Verification
    
    This skill ships a staleness verifier and two copy-and-adapt starter assets. The
    model table and pricing above are the facts most likely to drift — run the
    verifier when you suspect they're stale.
    
    **`scripts/check-model-table.py`** — guards the Current Models table (this file)
    and the per-model prompt-cache minimum table
    ([references/caching-and-cost.md](references/caching-and-cost.md)) against drift.
    Two modes per the [resource protocol §7](../../docs/SKILL-RESOURCE-PROTOCOL.md):
    
    ```bash
    # Structural (default, no network): every row well-formed, ids carry no date
    # suffix, prices numeric, the two files agree on the model lineup. It also
    # guards the cache-economics constants that are stated in more than one file
    # (0.1x read, 1.25x/2x writes, 4 breakpoints, 20-block lookback, the
    # context-management beta id), asserts each doctrine reference carries a
    # "verified <ISO date>" stamp, and checks SKILL.md <-> references/ citation
    # integrity in both directions. It then scans every file in the skill for model
    # ids: an id in neither the table nor the Legacy list is flagged "unknown", and
    # a LEGACY id sitting where a reader would copy it (model=..., "model": ...,
    # --model ...) is flagged "retired" - append a `legacy-ok` comment to that line
    # for a deliberate migration example. Exit 4 on any contradiction.
    python skills/claude-api-ops/scripts/check-model-table.py --offline
    python skills/claude-api-ops/scripts/check-model-table.py --offline --json | python -m json.tool
    
    # Live (advisory, needs ANTHROPIC_API_KEY): curls the Models API and compares
    # its id set against the documented ids. Exit 10 if a documented id is gone or a
    # newer alias id is missing from the table; exit 7 (not a failure) if the key is
    # unset or the API is unreachable. Live mode checks model-ID coverage ONLY — the
    # API returns no pricing, so pricing/context drift stays an --offline + docs concern.
    ANTHROPIC_API_KEY=sk-... python skills/claude-api-ops/scripts/check-model-table.py --live
    ```
    
    **`scripts/context-budget.py`** — append-vs-compact calculator. Models both paths
    in dollars over the turns you actually have left, checks the context ceiling
    first, and exits **10** when cost favours compaction, **0** when appending wins:
    
    ```bash
    # Short session — appending is cheaper (exit 0)
    python skills/claude-api-ops/scripts/context-budget.py \
        --history-tokens 25000 --turns-remaining 5 --base-rate 0.30
    
    # Deep session — cost favours compaction (exit 10)
    python skills/claude-api-ops/scripts/context-budget.py \
        --history-tokens 120000 --turns-remaining 40 --base-rate 2.00 --json
    ```
    
    Two results worth knowing before you trust it: the break-even turn count is
    **scale-invariant** (history size and price cancel out — it tracks the summary
    ratio, not how big or costly the conversation is), and it prices **cost only**.
    Recall loss is not in the model, so "compact" means cheaper, not better.
    
    **`assets/cached-agent-loop.py`** — the cache-aware sibling of the minimal loop
    below, and the executable form of the Context Engineering section: breakpoint at
    the end of the static prefix, a rolling breakpoint on the newest turn, an
    intermediate anchor every ~15 blocks so long tool-heavy turns don't jump the
    20-block lookback, tool output capped at the boundary, and a per-turn
    `cache_read_input_tokens` check that warns when the prefix silently changed.
    Copy it when the agent is long-running; copy `agentic-loop.py` when it isn't.
    
    The footgun it encodes: `cache_control` is a key on a content block, so it can
    only be set on a **dict**. Appending `response.content` verbatim (SDK block
    objects) or using the `"content": "a string"` shorthand leaves nowhere to put a
    marker — every breakpoint aimed at those turns is discarded with no error and no
    warning. Normalise content to dict blocks before placing breakpoints.
    
    **`assets/recall-probe.py`** — the "measure it on your workload" harness:
    plants a fact, buries it under N turns, probes for it, and reports recall, cost
    per turn and TTFT for **append** vs **compact**. Makes real API calls, so start
    small (`--turns 6 --trials 1`). Replace the synthetic filler turns with traffic
    from your own logs — that is the point of running it.
    
    **`assets/agentic-loop.py`** — a minimal, runnable tool-use loop (define a tool,
    call `messages.create`, loop while `stop_reason == "tool_use"`, append
    `tool_result`, re-request until `end_turn`). Copy it as the starting point when
    building a manual agent loop; the `>>> ADAPT` marks show what to change.
    
    **`assets/output-schema.json`** — a known-good structured-outputs request body in
    the canonical `output_config.format` shape (with `additionalProperties: false`
    and a `required` array). Copy and reshape `schema.properties` when adding JSON
    outputs; see [references/structured-outputs.md](references/structured-outputs.md)
    for the rules. (Supported on every current model — Fable 5, Opus 5, Sonnet 5,
    Haiku 4.5 — and the legacy 4.5–4.8 line.)
    
    ## Reference Files
    
    | File | Covers |
    |---|---|
    | [references/messages-api.md](references/messages-api.md) | Params, response shape, streaming events, stop reasons, error handling, retries, rate limits |
    | [references/tool-use.md](references/tool-use.md) | Tool definitions, tool_choice, parallel tools, agentic loop, tool results, server tools, tool runners |
    | [references/caching-and-cost.md](references/caching-and-cost.md) | Prompt caching mechanics, Batches API, token counting, model tiering economics |
    | [references/structured-outputs.md](references/structured-outputs.md) | output_config.format, schema rules/limits, strict tools, parse() helpers, thinking interplay |
    | [references/agent-sdk.md](references/agent-sdk.md) | Python + TS Agent SDK, ClaudeAgentOptions, hooks, MCP, sessions, SDK vs raw API |
    | [references/context-engineering.md](references/context-engineering.md) | Context budget, the three tiers, progressive disclosure, cache-aware ordering, tool-result bloat, sub-agents as isolation, instrumentation |
    | [references/compaction.md](references/compaction.md) | When compaction is justified, break-even arithmetic, context_management edits, memory tool, how to compact well |
    
    ## Live Documentation
    
    When cached facts may be stale, WebFetch (append `.md` for clean markdown):
    
    - Models/pricing: `https://platform.claude.com/docs/en/about-claude/models/overview.md`
    - Messages API: `https://platform.claude.com/docs/en/api/messages`
    - Tool use: `https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview.md`
    - Prompt caching: `https://platform.claude.com/docs/en/build-with-claude/prompt-caching.md`
    - Structured outputs: `https://platform.claude.com/docs/en/build-with-claude/structured-outputs.md`
    - Batches: `https://platform.claude.com/docs/en/build-with-claude/batch-processing.md`
    - Agent SDK: `https://code.claude.com/docs/en/agent-sdk/overview`
    - Context editing: `https://platform.claude.com/docs/en/build-with-claude/context-editing`
    - Context engineering: `https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents`
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related