Claude Skill

agent-routing

Decide which model, effort level, and cascade shape each subagent gets, and how to keep improvement loops safe (evaluator-as-selector, stop on regression). Routes on measured cost-per-completed-task rather than per-token price, because a tier's token count varies more by task sha

LLM Mart · 0 points · 1 views 1 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oaustegard-claude-skills-plugins_ai-and-reasoning_skills_agent-routing-e39c726.zip · 14 KB
Part of oaustegard/claude-skills — 39 skills

Install

skills CLI npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/ai-and-reasoning/skills/agent-routing
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
Git git clone https://github.com/oaustegard/claude-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.

README

agent-routing

Decide which model, effort level, and cascade shape each subagent gets, and how to keep improvement loops safe (evaluator-as-selector, stop on regression). Routes on measured cost-per-completed-task rather than per-token price, because a tier's token count varies more by task shape than price varies across tiers. Covers per-model effort semantics, the concision lever, cascade preconditions, context handoff, and watching a subagent fan-out live. Use when spawning subagents via the Agent or Workflow tools, when fanning out more than a handful of agents, or when asked which model or effort a task should get. Grounded in measured calibration (references/calibration-2026-07-15.md) plus a 2026-08 coding-cost study; Managed Agents API specifics are operational, not calibrated.

Skill manifest

Agent Routing — model, effort, and cascade selection

The rule that decides everything

Cost is output tokens × output price. Prices span ~5× across tiers. Token counts span up to 7× within a single tier depending on task shape. The shape therefore decides more than the tier does, and routing on the per-token discount gets the answer backwards.

Measured 2026-08-17, 14 spec-dense Python modules graded by hidden tests, all tiers at equal quality where noted:

arm tok/task pass $/task vs opus
haiku-solo 20,051 14/14 $0.1007 1.30×
haiku + concision 13,342 12/14 $0.0672 1.01×
opus-solo 3,001 14/14 $0.0774 1.00×
sonnet-base 4,687 14/14 $0.0478 0.62×
sonnet + concision 2,951 13/14 $0.0305 0.42×
sonnet cascade (below) — 14/14 $0.0315 0.41×

Haiku is 5× cheaper per token and cost 30% more per solved task than Opus, because it emitted 6.7× the tokens. Prices: Haiku 4.5 $1/$5, Sonnet 5 $2/$10, Opus 5 $5/$25 per MTok.

Two questions before spawning

  1. Is the output short or long? Short = a schema instance, a label, an answer, a small patch. Long = a module, a document, a plan, a review.
  2. Is it mechanically checkable, or does it need judgment?
short output long output
checkable haiku + verifier sonnet @ medium + concision + verifier
judgment sonnet @ medium sonnet/opus @ high

Output length is the discriminator because it is what the verbosity multiplier multiplies. Haiku's premium is invisible on a 200-token JSON object and ruinous on a 700-token module that costs it 13,000 tokens of thinking to produce.

Routing table

Task shape Model Effort Verify with
Extraction, classification, format transforms, schema-bound output haiku n/a schema / spot-check
Closed-form computation, state tracking, multi-hop lookup haiku n/a deterministic check
Constraint-bound generation (exact counts, required tokens, lipograms) haiku n/a mechanical checker
Bulk scans/greps, per-file summaries, fan-out reads haiku n/a sample audit
Code generation from a spec; any long structured artifact sonnet medium run the tests
Code edits with tests available sonnet medium run the tests
Judging / scoring another model's output sonnet+ medium — (judge ≠ worker)
Ambiguity resolution, novel synthesis, architecture, taste sonnet/opus high human or panel
Long-horizon multi-step agentic work, cross-file reasoning sonnet/opus high/xhigh milestone checks

Haiku holds the top four rows on merit: 240/240 measured across nested modular arithmetic, 30-hop chains, 25-operation state tracking, trap-laden word math, and 5-constraint generation, some with CoT suppressed (references/calibration-2026-07-15.md). Those runs requested effort: low; Haiku ignores effort, so the column says n/a (see below). Do not up-tier short checkable work "to be safe"; there is no measured benefit and it costs 3–5×. The burden of proof is on routing up.

Haiku loses the generation rows on cost alone, not capability — it scored 14/14 on the same suite Opus swept.

Effort is model-specific — verify per model before reusing a level

Measured 2026-08-17 via per-message output_tokens_details.thinking_tokens, thinking as a share of output on identical prompts:

model low medium
Sonnet 5 2.9% 47.7% (61.7% without concision)
Haiku 4.5 88–91% 88–91%

low was a near kill-switch on Sonnet 5: it dropped 14/14 → 10/14. Sonnet 5.5 is different. On the 14-repo seeded-bug battery (2026-09-29, two replicates, temporal-routing-headroom):

arm solved (r1, r2) output tokens, 14 tasks $/completed
Sonnet 5 @ low (2026-09-03) 9/14 16,779 —
Sonnet 5.5 @ low 10, 11 11,466 $0.0104
Sonnet 5.5 @ medium 11, 12 10,853 $0.0090
Sonnet 5.5 @ high 12, 12 14,488 $0.0121
Opus 5.5 @ high 11, 13 18,464 $0.0284
Opus 5 @ high (2026-09-03) 10/14 55,674 $0.1392
  • low and medium cost the same on Sonnet 5.5 for this work, and low missed only trap tasks. Claude Code's <reasoning_effort> tag agrees: 4 at low, 5 at medium.
  • high bought the paired traps: lru_ttl and wrap_text went from 0 of 2 at low to 2 of 2 and 1 of 2.
  • Sonnet 5.5 @ high matched Opus 5.5 @ high at 24 of 28 each, for 0.43× the cost per completed task. Seeded repair does not separate tiers, so this says the two are close on this family and nothing about harder work.
  • The token counts are one replicate (budget deltas from r2); thinking share could not be measured, because Claude Code omits thinking text.

Effort does not reach Haiku 4.5. The Messages API rejects effort on it; Claude Code's Workflow agent() accepts the option and drops it (0 of 12 Haiku probes carried an effort tag or a transcript effort field). The identical 88–91% thinking share at low and medium above is that fact seen from the token side, and the ~26% token drop once attributed to low sits inside the 23% run-to-run gap measured below. Tune Haiku with the prompt; the routing table lists its effort as n/a.

  • Tune Sonnet with the prompt as well as the knob. On Sonnet 5, medium was the working floor and low overshot into thinking-off; on Sonnet 5.5 low is a usable rung 1 and high is the rung that caught the traps.
  • Buy depth only for judgment-heavy roles; drop triage and formatting roles to low without touching the expensive role's budget.

Setting effort — which channel reaches which agent

Measured in Claude Code on 2026-09-28 (muninn.austegard.com/blog/effort-levels-in-claude-code-subagents.html):

channel reaches notes
Workflow agent(prompt, {effort}) one subagent 24/24 Sonnet and Opus probes ran at the requested level while the session sat elsewhere. The only way a parent sets a subagent's effort.
Agent tool, SendMessage nothing No effort parameter. The subagent runs at the session's level when it starts or resumes.
Session setting (/effort) main loop, and every subagent at its next turn A resumed subagent picks up the new level. The model cannot change this itself; it recommends a level and the human sets it.
Managed Agents agent config Set on the agent, not per session: an effort inside a per-session model override is silently ignored. If the create response echoes effort: None, the org's beta header (managed-agents-2026-04-01) doesn't carry the feature and the field was dropped.

A Workflow run needs the user's opt-in (ultracode, an explicit request, or a user-invoked skill that names it). Without one, a subagent spawned through the Agent tool gets the session's level, and the routing table's effort column is advisory.

The concision lever, and its limit

Adding one instruction — this is routine work; do not deliberate at length, do not enumerate test cases or weigh alternative designs; write it directly — cut output 37% on Sonnet and 27% on Haiku, at no quality cost. It composes with effort. Use it on every long-output generation spawn.

It does not reach work whose output is small. On agentic bug repair — a patch plus a paragraph — the same instruction cut Sonnet output 2.9%, at no quality change either way. The lever acts on deliberation the model would have written down, so a task that emits little has little to cut. Measure before carrying it to a new task family; "every long-output generation spawn" is the scope, and repair work is not in it.

Then stop. Thinking below a model's natural level is load-bearing, and cutting into it buys tokens with correctness:

  • An engineered suppression prompt (positive framing, bounded budget, n-shot exemplar) cut Haiku 35% and halved its pass rate, 8/9 → 4/9. Within that arm, passing runs thought 1.9× more than failing runs.
  • Sonnet at low (2.9% thinking) fell 14/14 → 10/14.
  • Priced per passing result the suppressed arms were more expensive: 22,143 tokens vs 17,126 for the un-engineered prompt.

A targeted checklist ("enumerate the spec's rejection rules first") helps only when it names the actual failure mode: it took one validation-heavy task from 15,220 to 9,634 tokens at equal quality, and took a semantics-heavy task from 3/3 to 0/3. Misnaming the failure mode is worse than not intervening.

Cascade

Precondition, checked first: is the cheap tier actually cheaper per task? The first rung is never free, so a cascade pays only when the cheap tier's measured cost per completed task is below the destination's. Verbosity can erase a price discount outright — Haiku at $0.067/task against Sonnet's $0.031 made haiku → sonnet worse than Sonnet alone regardless of p_fail: the attempt cost 2× the destination's entire job. Compute this before designing the ladder.

Second precondition: no verifier ⇒ no cascade. Route by the table instead; silent cheap-tier errors compound with nothing to catch them.

The verifier's holder makes the escalation call. Never the worker. A subagent asked whether it finished says yes: across 58 graded runs carrying an explicit "did you finish" field, 58 said yes and 44 had passed the held-out suite. Every one of the 14 failures self-reported success. The workers were not lying — they had passed the tests they could see, and those tests stay satisfiable while the task is unfinished. "Try it, and ask for help if you fail" therefore fails on exactly the tasks that need escalation. Put the decision wherever the stronger check lives; in a fan-out that is the orchestrator.

The shape that worked (measured, 14/14 at 0.41× Opus):

result = sonnet(task, effort=low, concise)          # rung 1: 10/14, $0.0155
if verify(result) fails:
    result = sonnet(task, effort=medium, concise,   # rung 2: fixed 12/12
                    prior=result, failure=test_output)

Rung 2 is the same model one effort step up. A tier jump is the exception you justify. Measured twice. On a second battery (14 seeded-bug repos, 2026-09-03) rung 2 ran from an identical failed attempt at both settings: sonnet @ medium and opus @ high rescued the same 4 of 5 tasks and both missed the same fifth, at 11,691 against 32,504 output tokens. Composed over the same rung 1, the same-model cascade cost 0.31× always-opus and the tier jump 0.76×. The tier jump costs 2.5× and buys nothing.

Those rungs are Sonnet 5. On Sonnet 5.5, low and medium solved the same tasks at the same cost, so the step that changes anything is low → high: high is where the paired traps started getting caught (see the effort table above). The cascade itself has not been run on 5.5.

A cascade can beat the frontier solo arm on correctness, not only on cost. In that run the sonnet→sonnet cascade solved 13/14 where always-opus solved 10/14. opus starting from the issue text fell into the same stop-early trap as sonnet on three tasks; opus starting from the failed patch and the failing assertions fixed all three.

How to run rung 2 in Claude Code. For a subagent, launch a fresh Workflow agent() one effort step up and pass it the diff and the test output; Agent and SendMessage cannot change a subagent's effort. For the main loop, recommend the next level and let the human set /effort; the session keeps its cache and carries on.

Caching pushes the same way. Caches are model-scoped, so a tier jump discards rung 1's prefix. An effort change does not: in Claude Code on 2026-09-28, five subagent resumes across a level change (Low→Medium, High→Extra, Extra→Max) all read their full prefix from cache, and the parent read 97,839 tokens from cache on its first request after a change. On the API the per-message effort message ({"role": "system", "content": [], "output_config": {"effort": …}}, beta mid-conversation-output-config-2026-07-01) keeps the cache on Fable 5.1, Mythos 5.1, Opus 5, Opus 5.5 and Sonnet 5.5 with thinking on; a change to the top-level effort still invalidates the messages cache.

The TTL is what misses. Every miss in that experiment followed a gap longer than five minutes. Subagents write only to the 5-minute tier and the parent to the 1-hour tier, so a subagent resumed after five idle minutes rebuilds its whole prefix (about 54k tokens for a bare general-purpose agent). Resume a subagent for a quick retry; after a longer gap, start a fresh one, since same-model agents share a cached prefix (32 of 36 fresh probes read from cache on their first request).

Every figure in this skill prices output tokens only; input and cache effects sit outside its cost model.

Carry the prior attempt and the raw failure output into the retry. Informed retry fixed 12/12; a blind re-attempt fixed 9/12 and failed one task identically across all three replicates — a systematic blind spot re-rolling never escapes. The extra input averaged 866 tokens, 5.9% of the retry's cost. Input is 1/5 the price of output, so context is nearly free relative to thinking.

The artifacts, not the prior model's account of itself. Adding rung 1's stated diagnosis on top of the patch and the assertions did nothing: 13/15 against 12/15 over three replicates, 1% fewer output tokens, and the whole difference was one replicate of one unstable task. It does not help and it does not anchor — SWE-Router (arXiv 2607.00053) restarts its strong model from the task description to avoid an anchoring effect that is not there. Pass the diff and the test output; skip the rationale.

Don't pay a frontier model to write guidance. An Opus diagnosis added zero over raw test output in two independent tests, at ~$0.15/task. The failing test already says what the orchestrator would say.

Verify content, not envelope. Strip fences, preambles, and trailing commentary before checking; hard-fail only on semantic content and log envelope deviations as soft. Two Haiku runs returned 7/7 and 6/6 correct fields while both wrapping output in a markdown fence the prompt forbade — a verifier keying on raw.startswith('{') would have escalated both for zero content error. Spurious escalation is a cascade failure mode, not a safety margin.

Judgment tasks fail in a shape checkers miss. Asked to rebut a stakeholder's "spend is down 66%" off a partial-month extract, Haiku killed the bad conclusion but normalized per calendar day across a 40%-weekend window and missed a model-mix confound — while passing every mechanical check available (word count, prose form, internal arithmetic consistency). The cheap tier fails as right headline, missed confound. This is why judgment rows route up rather than cascade.

Context handoff — routing picks the tier; the prompt carries the context

Subagents inherit nothing: not the conversation, not loaded skills, not the existence of artifacts already on disk. Every index, scan output, artifact path, or tool recipe must be serialized into the prompt (or a file the prompt points at). Otherwise the agent falls back to blind rediscovery and the tier premium is spent on crawling. A Sonnet with no handoff wastes more than a Haiku with a good procedure.

Per spawn: (1) artifact paths + how to query them, (2) tool commands verbatim, interpreter path included — subagents don't know your venv, (3) explicit anti-patterns ("no ls/glob discovery"), (4) an output spec.

Evidence: 2026-07-16, four Sonnet Explore agents launched onto a 2,300-file repo without the handoff opened with ls crawls despite a full tree-sitter symbol index sitting on disk; relaunched with per-agent index slices, the verbatim command, and anti-crawl rules, discovery cost dropped to ~zero.

To convert a judgment-shaped task into a cheap-tier-executable one (explicit procedures, n-shot examples), use the sibling down-skilling skill. This skill decides the routing; that one engineers the prompt.

Shared-prefix caching cuts the fan-out multiplier (unmeasured, conditional). When N subagents share a byte-stable prefix — the fixed handoff, not the per-agent slices — prefix caching can pull that portion toward a read-discount rate where the orchestration surface exposes it. Keep per-agent content at the tail. Verify your surface caches subagent prefixes before relying on it.

Loop discipline

Never blind-loop. Re-applying a prompt to a model's own output is the identity at best — an LLM call already unrolls its reasoning internally — and regression-then-freeze at worst: a re-looped haiku broke its own middle line on iteration 2 and froze on the broken text for every iteration after.

  1. Loop only with an out-of-band evaluator — ground truth, mechanical checker, or an up-tier judge scoring every iteration.
  2. Select, don't trust the last: final = argmax_r eval(answer_r).
  3. Stop on first regression. If eval(r) < eval(r-1), stop; loops froze on degraded output rather than recovering.
  4. Loop for diversity, not depth. Vary the angle per iteration; identical re-application converges instantly.
  5. "Improve this" with no headroom is the danger zone. It pressures the model to change something; without a selector, that change ships.

Judge rules

  • Judge model ≠ worker model; judge at least one tier up. Same-model self-assessment is untested.
  • Prefer mechanical checkers wherever a spec can be executed (counts, schemas, tests, regex): free, deterministic, zero judge tokens.
  • Judges are for rubric quality, not arithmetic — don't ask a model to verify a sum a Python one-liner can check.

Escalation triggers (route up despite the table)

  • The verifier fails twice at the same tier. Route up for capability, not for thoroughness — opus @ high fell into the same stop-early trap as sonnet @ low on three of four tasks built to reward a second look. A verifier catches that; a bigger model does not.
  • The task requires weighing trade-offs with no checkable ground truth.
  • Output ships verbatim to a human without review.
  • The subagent must plan its own multi-step tool strategy over many turns.
  • The task spans multiple sources that may disagree and must be reconciled.

Observing the fan-out — you can't govern what you can't watch

Stop-on-regression and "verifier failed twice" assume you can see a subagent's work while it runs. By default you can't: the session stream previews only the primary thread, and a subagent's output lands only after its whole turn buffers.

Attach one stream per thread. Read the session stream for the coordinator; on every session.thread_created (carrying session_thread_id and agent_name), attach a watcher to GET /v1/sessions/{id}/threads/{thread_id}/stream with event_deltas.

  • Preview is a scratch buffer; the buffered event is the record. Deltas are best-effort and shed under load, so concatenated deltas are a prefix of the final text. Reconcile by a single replace when the buffered agent.message arrives; the SDK's accumulate_managed_agents_event folds start/delta/record into one snapshot. One accumulator per connection. (The same trap appears offline: per-message usage records in transcripts include streaming partials — take the max per message id, or you undercount tokens ~2×.)
  • No replay. A stream opened after a request started gets no deltas for it, and reconnects never replay — attach on thread_created or miss the first response.
  • Coordination events live on the primary thread — session.thread_created, agent.thread_message_sent, agent.thread_message_received. Child tool calls cross-posted to the primary carry session_thread_id; skip them.
  • Terminate cleanly. Watchers exit on session.thread_status_idle; the main loop on session.status_idle — print the stop reason when it isn't end_turn, and break on terminated-status events.

Operational, not calibrated. Source: Anthropic Managed Agents notebook CMA_watch_subagents_live (beta managed-agents-2026-04-01); contract in events and streaming.

Measure before trusting this

Everything above is measured on three batteries: a 300-call deterministic calibration (references/calibration-2026-07-15.md), a 14-task hidden-test coding suite (2026-08-17, ~190 subagent runs), and a 14-repo seeded-bug agentic battery (2026-09-03, ~120 subagent runs, oaustegard/experiments → temporal-routing-headroom) that measured the cascade rungs, the escalation signal, and the tier gap against each other. Re-measure when:

  • A model or price revision lands. Both the verbosity multipliers and the cost table above invert on either. Sonnet 5.5 (2026-09-29) is such a revision. Its effort levels are measured above on the seeded-bug battery only; the other Sonnet figures in this skill are Sonnet 5 data, and the verbosity multipliers and the generation-suite costs have not been re-run on the 5.5 generation.
  • The task family is off all three batteries. No deterministic task has made Haiku fail on correctness yet, so the capability cliff is past what's been probed. Seeded-bug repair in a small module is now measured as not tier-separating: three probe shapes aimed at thoroughness, at ambiguity the tests underdetermine, and at a repo with no test suite at all, and sonnet @ low solved all six cells against opus @ high.
  • Output length differs materially from what was measured. The whole cost model keys on token volume; a 10× longer artifact re-opens the tier question.
  • You need pass-rate differences of 1–2 tasks. Run-to-run variance swamps them: two runs of the same model on the same 14 tasks produced disjoint failure sets and a 23% token gap. Token deltas are trustworthy; small pass-rate deltas are not.
Files (claude-skills)
  • references
    • calibration-2026-07-15.md 4.6 KB
      # Calibration evidence — 2026-07-15
      
      Data behind the routing heuristics in SKILL.md. Collected in one CCotw session
      (Fable 5 orchestrator, Workflow tool, session `bf6d02ee`), ~12.3M subagent
      tokens total across four workflow runs, zero agent errors. All numeric grading
      was deterministic (Python scorers against generated ground truth); the only
      judge tokens were Sonnet scoring open-ended text in Experiment 1.
      
      ## Experiment 1 — Looped Haiku subagents (Looped-Mamba analog, arXiv 2607.10110)
      
      Design: `answer_r = Haiku(task, answer_{r-1})`, R=5, out-of-band eval per loop.
      
      | Task family | Instances × loops | Result |
      |---|---|---|
      | Nested mod-23 arithmetic, depth-3 trees | 5 × 5 | 100% at r=1, flat, 0 churn |
      | p-hop function chains (n=7, 8–12 hops) | 5 × 5 | 100% at r=1, flat, 0 churn |
      | Harder: depth-4 trees (16 leaves) | 8 × 5 | 100% at r=1, flat, 0 churn |
      | Harder: n=12 chains, 20–30 hops | 8 × 5 | 100% at r=1, flat, 0 churn |
      | Answer-only control (CoT suppressed, effort=low), depth-4 trees | 8 × 5 | still 100% at r=1, 0 churn |
      | Open-ended (Sonnet judge /10): prime one-liner | 1 × 5 | [10,10,10,10,10] |
      | Open-ended: sky-blue explanation | 1 × 5 | [9,9,9,9,9] |
      | Open-ended: 5-7-5 haiku | 1 × 5 | **[9,3,3,3,3]** — loop 2 broke the middle line to 8 syllables, loops 3–5 froze on identical broken text |
      
      Interpretation: a single LLM call spends variable internal compute (CoT), so
      `answer_1` is already a fixed point on checkable tasks — the paper's
      fixed-compute-per-step premise doesn't hold for LLM subagents. The exit-gate
      idea *does* transfer: the evaluator-as-selector recovers the loop-1 haiku and
      skips wasted loops elsewhere. Answer churn across all deterministic loop-pairs:
      0/96 flips.
      
      ## Experiment 2 — Single-shot routing calibration (60 agents)
      
      20 tasks × 3 configs, single shot, deterministic local scoring:
      
      Tasks: 9 constraint-stack sentences (K=3/4/5 simultaneous constraints: exact
      word count, begin-word, include-word, end-word, no letter 'e'), 4 trap-laden
      word-math problems (distractor numbers), 3 state-tracking problems (3 boxes,
      25 operations), 4 constraint-preserving revisions (exact N words, fixed
      first/last word, ≥2 changes).
      
      | Config | stack | trap | state | revision | total |
      |---|---|---|---|---|---|
      | haiku, effort=low | 9/9 | 4/4 | 3/3 | 4/4 | **20/20** |
      | haiku, effort=high | 9/9 | 4/4 | 3/3 | 4/4 | **20/20** |
      | sonnet, effort=low | 6/9 | 4/4 | 3/3 | 4/4 | **17/20** |
      
      Sonnet-low failures (hand-verified, real):
      - `stack-K3-1`: 15 words where exactly 14 required
      - `stack-K4-0`: 13 words where exactly 12 required
      - `stack-K5-0`: used "quiet" in a no-letter-'e' sentence
      
      Notes:
      - n=20/config: 17 vs 20 is not statistically significant. The defensible claim
        is "no evidence of up-tier benefit on mechanical tasks," which is enough to
        invert the default (burden of proof on routing up).
      - Haiku's K5 lipogram outputs were flawless, e.g. "Our distant stars hang
        bright and high throughout dark night" (10 words, begins "our", includes
        "stars", ends "night", zero e's).
      - Effort had no measurable effect on Haiku here (both 100%) — consistent with
        Experiment 1's answer-only/effort-low control also scoring 100%.
      
      ## Pricing basis (2026-07, per MTok in/out)
      
      Haiku 4.5 $1/$5 · Sonnet 5 **$2/$10** · Opus 5 $5/$25.
      
      **Superseded 2026-08-17:** this section previously read Sonnet 5 at $3/$15 with
      $2/$10 as an intro rate expiring 2026-08-31. $2/$10 is now the standing price
      (user-confirmed). Any analysis computed at $3/$15 understates Sonnet by ~1/3.
      
      The `p_fail`-only break-even formerly stated here (Haiku-first beats Sonnet-direct
      while `p_fail(Haiku) < 1 − c_H/c_S ≈ 2/3`) is **wrong in the general case**. It
      assumes `c_H < c_S` per task, which holds only when outputs are short. On
      long-output work Haiku's verbosity inverts it: measured 2026-08-17, Haiku cost
      $0.067/task against Sonnet's $0.031, so Haiku-first loses at *every* `p_fail`,
      including zero. Check cost-per-task first; see the cascade precondition in
      SKILL.md. The verifier-is-near-free assumption still holds — all Experiment 2
      scoring was local Python.
      
      ## What has NOT been measured
      
      - A deterministic task family where Haiku actually fails (the cliff).
      - Multi-turn agentic tool-use quality per tier.
      - Same-model self-judging reliability.
      - Haiku at higher K (>5 simultaneous constraints) or longer state chains.
      - Multi-source consistency/reconciliation (do two extracts of the same period
        agree?). One anecdote favors up-tiering; untested.
      
      Re-run: generators + scorer live in the session scratchpad pattern
      (`gen.py`/`gen2.py`/`gen3.py`, `score.py`); regenerate with new seeds and a
      current model rev before trusting the table across model versions.
      
  • CHANGELOG.md 3.7 KB
    # agent-routing - Changelog
    
    All notable changes to the `agent-routing` skill are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
    
    ## [2.3.0] - 2026-09-29
    
    ### Other
    
    - agent-routing 2.3.0: Sonnet 5.5 effort measured
    
    ## [2.3.0] - 2026-09-29
    
    ### Added
    
    - Sonnet 5.5 effort measured on the seeded-bug battery, two replicates: `low` 10 and
      11 of 14, `medium` 11 and 12, `high` 12 and 12, against Sonnet 5 `low` at 9/14. `low`
      no longer switches thinking off. Sonnet 5.5 @ `high` matched Opus 5.5 @ `high` (24 of
      28 each) at 0.43x the cost per completed task. Opus 5.5 emitted a third of Opus 5's
      output on the same tasks.
    
    ## [2.2.0] - 2026-09-29
    
    ### Other
    
    - agent-routing 2.2.0: effort channels, cache TTL, Sonnet 5.5 caveat
    
    ## [2.2.0] - 2026-09-29
    
    From the 2026-09-28 Claude Code effort experiments
    (muninn.austegard.com/blog/effort-levels-in-claude-code-subagents.html) and the Sonnet 5.5
    release.
    
    ### Added
    
    - Which channel sets whose effort: Workflow `agent({effort})` is the only way a parent sets
      a subagent's effort; Agent and SendMessage have none; a resumed subagent picks up the
      session's `/effort`.
    - How to run rung 2 in Claude Code: a fresh Workflow agent one effort step up for
      subagents, a recommended `/effort` change for the main loop.
    - The cache TTL, not the effort change, is what misses: subagents cache on the 5-minute
      tier, the parent on the 1-hour tier.
    
    ### Changed
    
    - Caching paragraph: an effort change keeps the cache in Claude Code (five resumes across a
      level change all hit), and the API's per-message effort message now covers Opus 5.5 and
      Sonnet 5.5. Replaces "an effort change invalidates the messages cache on every model".
    - Haiku's effort column reads n/a: the API rejects `effort` on Haiku 4.5 and Claude Code
      drops it.
    - Every Sonnet figure is labelled as Sonnet 5 data pending a re-measure on Sonnet 5.5,
      whose effort levels were recalibrated.
    
    ## [2.1.0] - 2026-09-04
    
    ### Other
    
    - agent-routing 2.1.0: measured cascade rungs, escalation signal, tier gap (#785)
    
    ## [2.1.0] - 2026-09-03
    
    Measured against a 14-repo seeded-bug agentic battery (~120 subagent runs;
    `oaustegard/experiments` -> `temporal-routing-headroom`). Nothing was retracted; the
    cascade section gained the numbers it was missing.
    
    ### Added
    
    - The escalation call belongs to whoever holds the verifier, never the worker. 58 of 58
      graded runs self-reported success; 44 had passed. Every failure claimed to be done.
    - Rung 2 is the same model one effort step up; a tier jump is the exception. From an
      identical failed attempt, `sonnet` @ `medium` and `opus` @ `high` rescued the same 4 of
      5 tasks at 11,691 vs 32,504 output tokens (0.31x vs 0.76x always-`opus` composed).
    - A cascade can beat the frontier solo arm on correctness: 13/14 vs 10/14.
    - Caching in the cascade: caches are model-scoped with no escape hatch, so a tier jump
      discards rung 1's prefix; an `effort` change invalidates the messages cache on every
      model, and the per-message effort hatch is Opus 5 / Fable 5.1 / Mythos 5.1 only.
    - Informed retry means the artifacts, not the prior model's narrative: adding rung 1's
      stated diagnosis moved 12/15 to 13/15 on one replicate of one unstable task.
    - Concision does not reach small-output work: 2.9% on agentic repair vs 37% on generation.
    - Route up for capability, not thoroughness: `opus` @ `high` fell into the same
      stop-early trap as `sonnet` @ `low` on three of four tasks.
    - Seeded-bug repair in a small module is measured as not tier-separating across three
      probe shapes.
    
    ### Changed
    
    - Cost-model caveat made explicit: every figure prices output tokens only.
    
    ## [2.0.0] - 2026-08-18
    
    ### Added
    
    - Add/Update skill: agent-routing (#767)
  • README.md 798 B
    # agent-routing
    
    Decide which model, effort level, and cascade shape each subagent gets, and how to keep improvement loops safe (evaluator-as-selector, stop on regression). Routes on measured cost-per-completed-task rather than per-token price, because a tier's token count varies more by task shape than price varies across tiers. Covers per-model effort semantics, the concision lever, cascade preconditions, context handoff, and watching a subagent fan-out live. Use when spawning subagents via the Agent or Workflow tools, when fanning out more than a handful of agents, or when asked which model or effort a task should get. Grounded in measured calibration (references/calibration-2026-07-15.md) plus a 2026-08 coding-cost study; Managed Agents API specifics are operational, not calibrated.
    
  • SKILL.md 23 KB
    ---
    name: agent-routing
    description: Decide which model, effort level, and cascade shape each subagent gets, and how to keep improvement loops safe (evaluator-as-selector, stop on regression). Routes on measured cost-per-completed-task rather than per-token price, because a tier's token count varies more by task shape than price varies across tiers. Covers per-model effort semantics, the concision lever, cascade preconditions, context handoff, and watching a subagent fan-out live. Use when spawning subagents via the Agent or Workflow tools, when choosing how to escalate a failed attempt, when fanning out more than a handful of agents, or when asked which model or effort a task should get. Grounded in measured calibration (references/calibration-2026-07-15.md), a 2026-08 coding-cost study, and a 2026-09 agentic-repair battery that measured the cascade rungs directly; Managed Agents API specifics are operational, not calibrated.
    compatibility: Designed for Claude Code / Claude Code on the Web — assumes an orchestrator with Agent/Workflow subagent tools. Only the Workflow tool sets a subagent's effort; the Agent tool sets its model. Not applicable to claude.ai chat use.
    metadata:
      author: Oskar Austegard and Claude
      version: "2.3.0"
    ---
    
    # Agent Routing — model, effort, and cascade selection
    
    ## The rule that decides everything
    
    **Cost is output tokens × output price.** Prices span ~5× across tiers. Token
    counts span up to **7× within a single tier** depending on task shape. The shape
    therefore decides more than the tier does, and *routing on the per-token discount
    gets the answer backwards*.
    
    Measured 2026-08-17, 14 spec-dense Python modules graded by hidden tests, all
    tiers at equal quality where noted:
    
    | arm | tok/task | pass | $/task | vs opus |
    |---|---|---|---|---|
    | haiku-solo | 20,051 | 14/14 | $0.1007 | **1.30×** |
    | haiku + concision | 13,342 | 12/14 | $0.0672 | 1.01× |
    | opus-solo | 3,001 | 14/14 | $0.0774 | 1.00× |
    | sonnet-base | 4,687 | 14/14 | $0.0478 | 0.62× |
    | sonnet + concision | 2,951 | 13/14 | $0.0305 | 0.42× |
    | **sonnet cascade** (below) | — | **14/14** | **$0.0315** | **0.41×** |
    
    Haiku is 5× cheaper per token and cost **30% more per solved task** than Opus,
    because it emitted 6.7× the tokens. Prices: Haiku 4.5 $1/$5, Sonnet 5 $2/$10,
    Opus 5 $5/$25 per MTok.
    
    ## Two questions before spawning
    
    1. **Is the output short or long?** Short = a schema instance, a label, an answer,
       a small patch. Long = a module, a document, a plan, a review.
    2. **Is it mechanically checkable, or does it need judgment?**
    
    | | short output | long output |
    |---|---|---|
    | **checkable** | `haiku` + verifier | `sonnet` @ `medium` + concision + verifier |
    | **judgment** | `sonnet` @ `medium` | `sonnet`/`opus` @ `high` |
    
    Output length is the discriminator because it is what the verbosity multiplier
    multiplies. Haiku's premium is invisible on a 200-token JSON object and ruinous on
    a 700-token module that costs it 13,000 tokens of thinking to produce.
    
    ## Routing table
    
    | Task shape | Model | Effort | Verify with |
    |---|---|---|---|
    | Extraction, classification, format transforms, schema-bound output | `haiku` | n/a | schema / spot-check |
    | Closed-form computation, state tracking, multi-hop lookup | `haiku` | n/a | deterministic check |
    | Constraint-bound generation (exact counts, required tokens, lipograms) | `haiku` | n/a | mechanical checker |
    | Bulk scans/greps, per-file summaries, fan-out reads | `haiku` | n/a | sample audit |
    | **Code generation from a spec; any long structured artifact** | **`sonnet`** | **`medium`** | run the tests |
    | Code edits with tests available | `sonnet` | `medium` | run the tests |
    | Judging / scoring another model's output | `sonnet`+ | `medium` | — (judge ≠ worker) |
    | Ambiguity resolution, novel synthesis, architecture, taste | `sonnet`/`opus` | `high` | human or panel |
    | Long-horizon multi-step agentic work, cross-file reasoning | `sonnet`/`opus` | `high`/`xhigh` | milestone checks |
    
    Haiku holds the top four rows on merit: 240/240 measured across nested modular
    arithmetic, 30-hop chains, 25-operation state tracking, trap-laden word math, and
    5-constraint generation, some with CoT suppressed
    (references/calibration-2026-07-15.md). Those runs requested `effort: low`; Haiku ignores
    effort, so the column says n/a (see below). **Do not up-tier short checkable work "to be
    safe"**; there is no measured benefit and it costs 3–5×. The burden of proof is on
    routing up.
    
    Haiku loses the generation rows on cost alone, not capability — it scored 14/14 on
    the same suite Opus swept.
    
    ## Effort is model-specific — verify per model before reusing a level
    
    Measured 2026-08-17 via per-message `output_tokens_details.thinking_tokens`,
    thinking as a share of output on identical prompts:
    
    | model | `low` | `medium` |
    |---|---|---|
    | Sonnet 5 | **2.9%** | 47.7% (61.7% without concision) |
    | Haiku 4.5 | **88–91%** | 88–91% |
    
    `low` was a near kill-switch on Sonnet 5: it dropped 14/14 → 10/14. **Sonnet 5.5 is
    different.** On the 14-repo seeded-bug battery (2026-09-29, two replicates,
    `temporal-routing-headroom`):
    
    | arm | solved (r1, r2) | output tokens, 14 tasks | $/completed |
    |---|---|---|---|
    | Sonnet 5 @ `low` (2026-09-03) | 9/14 | 16,779 | — |
    | Sonnet 5.5 @ `low` | 10, 11 | 11,466 | $0.0104 |
    | Sonnet 5.5 @ `medium` | 11, 12 | 10,853 | $0.0090 |
    | Sonnet 5.5 @ `high` | 12, 12 | 14,488 | $0.0121 |
    | Opus 5.5 @ `high` | 11, 13 | 18,464 | $0.0284 |
    | Opus 5 @ `high` (2026-09-03) | 10/14 | 55,674 | $0.1392 |
    
    - `low` and `medium` cost the same on Sonnet 5.5 for this work, and `low` missed only
      trap tasks. Claude Code's `<reasoning_effort>` tag agrees: 4 at `low`, 5 at `medium`.
    - `high` bought the paired traps: `lru_ttl` and `wrap_text` went from 0 of 2 at `low` to
      2 of 2 and 1 of 2.
    - Sonnet 5.5 @ `high` matched Opus 5.5 @ `high` at 24 of 28 each, for **0.43×** the cost
      per completed task. Seeded repair does not separate tiers, so this says the two are
      close on this family and nothing about harder work.
    - The token counts are one replicate (budget deltas from r2); thinking share could not
      be measured, because Claude Code omits thinking text.
    
    **Effort does not reach Haiku 4.5.** The Messages API rejects `effort` on it; Claude
    Code's Workflow `agent()` accepts the option and drops it (0 of 12 Haiku probes carried an
    effort tag or a transcript effort field). The identical 88–91% thinking share at `low` and
    `medium` above is that fact seen from the token side, and the ~26% token drop once
    attributed to `low` sits inside the 23% run-to-run gap measured below. Tune Haiku with the
    prompt; the routing table lists its effort as n/a.
    
    - **Tune Sonnet with the prompt as well as the knob.** On Sonnet 5, `medium` was the
      working floor and `low` overshot into thinking-off; on Sonnet 5.5 `low` is a usable
      rung 1 and `high` is the rung that caught the traps.
    - Buy depth only for judgment-heavy roles; drop triage and formatting roles to `low`
      without touching the expensive role's budget.
    
    ### Setting effort — which channel reaches which agent
    
    Measured in Claude Code on 2026-09-28 (muninn.austegard.com/blog/effort-levels-in-claude-code-subagents.html):
    
    | channel | reaches | notes |
    |---|---|---|
    | Workflow `agent(prompt, {effort})` | one subagent | 24/24 Sonnet and Opus probes ran at the requested level while the session sat elsewhere. **The only way a parent sets a subagent's effort.** |
    | Agent tool, SendMessage | nothing | No `effort` parameter. The subagent runs at the session's level when it starts or resumes. |
    | Session setting (`/effort`) | main loop, and every subagent at its next turn | A resumed subagent picks up the new level. The model cannot change this itself; it recommends a level and the human sets it. |
    | Managed Agents | agent config | Set on the agent, not per session: an `effort` inside a per-session `model` override is silently ignored. If the create response echoes `effort: None`, the org's beta header (`managed-agents-2026-04-01`) doesn't carry the feature and the field was dropped. |
    
    A Workflow run needs the user's opt-in (ultracode, an explicit request, or a user-invoked
    skill that names it). Without one, a subagent spawned through the Agent tool gets the
    session's level, and the routing table's effort column is advisory.
    
    ## The concision lever, and its limit
    
    Adding one instruction — *this is routine work; do not deliberate at length, do not
    enumerate test cases or weigh alternative designs; write it directly* — cut output
    **37% on Sonnet** and **27% on Haiku**, at no quality cost. It composes with effort.
    Use it on every long-output generation spawn.
    
    **It does not reach work whose output is small.** On agentic bug repair — a patch plus a
    paragraph — the same instruction cut Sonnet output **2.9%**, at no quality change either
    way. The lever acts on deliberation the model would have written down, so a task that
    emits little has little to cut. Measure before carrying it to a new task family; "every
    long-output generation spawn" is the scope, and repair work is not in it.
    
    **Then stop.** Thinking below a model's natural level is load-bearing, and cutting
    into it buys tokens with correctness:
    
    - An engineered suppression prompt (positive framing, bounded budget, n-shot
      exemplar) cut Haiku 35% and **halved** its pass rate, 8/9 → 4/9. Within that arm,
      passing runs thought **1.9×** more than failing runs.
    - Sonnet at `low` (2.9% thinking) fell 14/14 → 10/14.
    - Priced per *passing* result the suppressed arms were **more** expensive: 22,143
      tokens vs 17,126 for the un-engineered prompt.
    
    A targeted checklist ("enumerate the spec's rejection rules first") helps only when
    it names the actual failure mode: it took one validation-heavy task from 15,220 to
    9,634 tokens at equal quality, and took a semantics-heavy task from 3/3 to **0/3**.
    Misnaming the failure mode is worse than not intervening.
    
    ## Cascade
    
    **Precondition, checked first: is the cheap tier actually cheaper per task?** The
    first rung is never free, so a cascade pays only when the cheap tier's *measured*
    cost per completed task is below the destination's. Verbosity can erase a price
    discount outright — Haiku at $0.067/task against Sonnet's $0.031 made
    `haiku → sonnet` worse than Sonnet alone **regardless of `p_fail`**: the attempt
    cost 2× the destination's entire job. Compute this before designing the ladder.
    
    **Second precondition: no verifier ⇒ no cascade.** Route by the table instead;
    silent cheap-tier errors compound with nothing to catch them.
    
    **The verifier's holder makes the escalation call. Never the worker.** A subagent asked
    whether it finished says yes: across 58 graded runs carrying an explicit "did you finish"
    field, 58 said yes and 44 had passed the held-out suite. Every one of the 14 failures
    self-reported success. The workers were not lying — they had passed the tests they could
    see, and those tests stay satisfiable while the task is unfinished. "Try it, and ask for
    help if you fail" therefore fails on exactly the tasks that need escalation. Put the
    decision wherever the stronger check lives; in a fan-out that is the orchestrator.
    
    The shape that worked (measured, 14/14 at 0.41× Opus):
    
    ```
    result = sonnet(task, effort=low, concise)          # rung 1: 10/14, $0.0155
    if verify(result) fails:
        result = sonnet(task, effort=medium, concise,   # rung 2: fixed 12/12
                        prior=result, failure=test_output)
    ```
    
    **Rung 2 is the same model one effort step up. A tier jump is the exception you justify.**
    Measured twice. On a second battery (14 seeded-bug repos, 2026-09-03) rung 2 ran from an
    identical failed attempt at both settings: `sonnet` @ `medium` and `opus` @ `high` rescued
    the same 4 of 5 tasks and both missed the same fifth, at 11,691 against 32,504 output
    tokens. Composed over the same rung 1, the same-model cascade cost **0.31×** always-`opus`
    and the tier jump **0.76×**. The tier jump costs 2.5× and buys nothing.
    
    Those rungs are Sonnet 5. On Sonnet 5.5, `low` and `medium` solved the same tasks at the
    same cost, so the step that changes anything is `low` → `high`: `high` is where the
    paired traps started getting caught (see the effort table above). The cascade itself has
    not been run on 5.5.
    
    **A cascade can beat the frontier solo arm on correctness, not only on cost.** In that run
    the `sonnet`→`sonnet` cascade solved 13/14 where always-`opus` solved 10/14. `opus`
    starting from the issue text fell into the same stop-early trap as `sonnet` on three
    tasks; `opus` starting from the failed patch and the failing assertions fixed all three.
    
    **How to run rung 2 in Claude Code.** For a subagent, launch a fresh Workflow `agent()`
    one effort step up and pass it the diff and the test output; Agent and SendMessage cannot
    change a subagent's effort. For the main loop, recommend the next level and let the human
    set `/effort`; the session keeps its cache and carries on.
    
    **Caching pushes the same way.** Caches are model-scoped, so a tier jump discards rung 1's
    prefix. An effort change does not: in Claude Code on 2026-09-28, five subagent resumes
    across a level change (Low→Medium, High→Extra, Extra→Max) all read their full prefix from
    cache, and the parent read 97,839 tokens from cache on its first request after a change.
    On the API the per-message effort message (`{"role": "system", "content": [],
    "output_config": {"effort": …}}`, beta `mid-conversation-output-config-2026-07-01`) keeps
    the cache on Fable 5.1, Mythos 5.1, Opus 5, Opus 5.5 and Sonnet 5.5 with thinking on; a
    change to the top-level `effort` still invalidates the messages cache.
    
    **The TTL is what misses.** Every miss in that experiment followed a gap longer than five
    minutes. Subagents write only to the 5-minute tier and the parent to the 1-hour tier, so a
    subagent resumed after five idle minutes rebuilds its whole prefix (about 54k tokens for a
    bare general-purpose agent). Resume a subagent for a quick retry; after a longer gap, start
    a fresh one, since same-model agents share a cached prefix (32 of 36 fresh probes read from
    cache on their first request).
    
    Every figure in this skill prices output tokens only; input and cache effects sit outside
    its cost model.
    
    **Carry the prior attempt and the raw failure output into the retry.** Informed retry
    fixed **12/12**; a blind re-attempt fixed **9/12** and failed one task *identically
    across all three replicates* — a systematic blind spot re-rolling never escapes. The
    extra input averaged 866 tokens, **5.9%** of the retry's cost. Input is 1/5 the price
    of output, so context is nearly free relative to thinking.
    
    **The artifacts, not the prior model's account of itself.** Adding rung 1's stated
    diagnosis on top of the patch and the assertions did nothing: 13/15 against 12/15 over
    three replicates, 1% fewer output tokens, and the whole difference was one replicate of
    one unstable task. It does not help and it does not anchor — SWE-Router (arXiv
    2607.00053) restarts its strong model from the task description to avoid an anchoring
    effect that is not there. Pass the diff and the test output; skip the rationale.
    
    **Don't pay a frontier model to write guidance.** An Opus diagnosis added zero over
    raw test output in two independent tests, at ~$0.15/task. The failing test already
    says what the orchestrator would say.
    
    **Verify content, not envelope.** Strip fences, preambles, and trailing commentary
    before checking; hard-fail only on semantic content and log envelope deviations as
    soft. Two Haiku runs returned 7/7 and 6/6 correct fields while both wrapping output
    in a markdown fence the prompt forbade — a verifier keying on `raw.startswith('{')`
    would have escalated both for zero content error. Spurious escalation is a cascade
    failure mode, not a safety margin.
    
    **Judgment tasks fail in a shape checkers miss.** Asked to rebut a stakeholder's
    "spend is down 66%" off a partial-month extract, Haiku killed the bad conclusion but
    normalized per calendar day across a 40%-weekend window and missed a model-mix
    confound — while passing every mechanical check available (word count, prose form,
    internal arithmetic consistency). The cheap tier fails as *right headline, missed
    confound*. This is why judgment rows route up rather than cascade.
    
    ## Context handoff — routing picks the tier; the prompt carries the context
    
    Subagents inherit nothing: not the conversation, not loaded skills, not the existence
    of artifacts already on disk. Every index, scan output, artifact path, or tool recipe
    must be serialized into the prompt (or a file the prompt points at). Otherwise the
    agent falls back to blind rediscovery and the tier premium is spent on crawling. **A
    Sonnet with no handoff wastes more than a Haiku with a good procedure.**
    
    Per spawn: (1) artifact paths + how to query them, (2) tool commands verbatim,
    interpreter path included — subagents don't know your venv, (3) explicit
    anti-patterns ("no `ls`/glob discovery"), (4) an output spec.
    
    Evidence: 2026-07-16, four Sonnet Explore agents launched onto a 2,300-file repo
    without the handoff opened with `ls` crawls despite a full tree-sitter symbol index
    sitting on disk; relaunched with per-agent index slices, the verbatim command, and
    anti-crawl rules, discovery cost dropped to ~zero.
    
    To convert a judgment-shaped task into a cheap-tier-executable one (explicit
    procedures, n-shot examples), use the sibling `down-skilling` skill. This skill
    decides the routing; that one engineers the prompt.
    
    **Shared-prefix caching cuts the fan-out multiplier** (unmeasured, conditional). When
    N subagents share a byte-stable prefix — the *fixed* handoff, not the per-agent
    slices — prefix caching can pull that portion toward a read-discount rate *where the
    orchestration surface exposes it*. Keep per-agent content at the tail. Verify your
    surface caches subagent prefixes before relying on it.
    
    ## Loop discipline
    
    Never blind-loop. Re-applying a prompt to a model's own output is the identity at
    best — an LLM call already unrolls its reasoning internally — and
    regression-then-freeze at worst: a re-looped haiku broke its own middle line on
    iteration 2 and froze on the broken text for every iteration after.
    
    1. **Loop only with an out-of-band evaluator** — ground truth, mechanical checker, or
       an up-tier judge scoring every iteration.
    2. **Select, don't trust the last:** `final = argmax_r eval(answer_r)`.
    3. **Stop on first regression.** If `eval(r) < eval(r-1)`, stop; loops froze on
       degraded output rather than recovering.
    4. **Loop for diversity, not depth.** Vary the angle per iteration; identical
       re-application converges instantly.
    5. **"Improve this" with no headroom is the danger zone.** It pressures the model to
       change something; without a selector, that change ships.
    
    ## Judge rules
    
    - Judge model ≠ worker model; judge at least one tier up. Same-model self-assessment
      is untested.
    - Prefer mechanical checkers wherever a spec can be executed (counts, schemas, tests,
      regex): free, deterministic, zero judge tokens.
    - Judges are for rubric quality, not arithmetic — don't ask a model to verify a sum a
      Python one-liner can check.
    
    ## Escalation triggers (route up despite the table)
    
    - The verifier fails twice at the same tier. **Route up for capability, not for
      thoroughness** — `opus` @ `high` fell into the same stop-early trap as `sonnet` @ `low`
      on three of four tasks built to reward a second look. A verifier catches that; a bigger
      model does not.
    - The task requires weighing trade-offs with no checkable ground truth.
    - Output ships verbatim to a human without review.
    - The subagent must plan its own multi-step tool strategy over many turns.
    - The task spans multiple sources that may disagree and must be reconciled.
    
    ## Observing the fan-out — you can't govern what you can't watch
    
    Stop-on-regression and "verifier failed twice" assume you can see a subagent's work
    *while it runs*. By default you can't: the session stream previews only the primary
    thread, and a subagent's output lands only after its whole turn buffers.
    
    Attach one stream per thread. Read the session stream for the coordinator; on every
    `session.thread_created` (carrying `session_thread_id` and `agent_name`), attach a
    watcher to `GET /v1/sessions/{id}/threads/{thread_id}/stream` with `event_deltas`.
    
    - **Preview is a scratch buffer; the buffered event is the record.** Deltas are
      best-effort and shed under load, so concatenated deltas are a *prefix* of the final
      text. Reconcile by a single replace when the buffered `agent.message` arrives; the
      SDK's `accumulate_managed_agents_event` folds start/delta/record into one snapshot.
      One accumulator per connection. (The same trap appears offline: per-message usage
      records in transcripts include streaming partials — take the **max** per message
      id, or you undercount tokens ~2×.)
    - **No replay.** A stream opened after a request started gets no deltas for it, and
      reconnects never replay — attach on `thread_created` or miss the first response.
    - **Coordination events live on the primary thread** — `session.thread_created`,
      `agent.thread_message_sent`, `agent.thread_message_received`. Child tool calls
      cross-posted to the primary carry `session_thread_id`; skip them.
    - **Terminate cleanly.** Watchers exit on `session.thread_status_idle`; the main loop
      on `session.status_idle` — print the stop reason when it isn't `end_turn`, and
      break on terminated-status events.
    
    Operational, not calibrated. Source: Anthropic Managed Agents notebook
    `CMA_watch_subagents_live` (beta `managed-agents-2026-04-01`); contract in
    [events and streaming](https://platform.claude.com/docs/en/managed-agents/events-and-streaming#event-deltas).
    
    ## Measure before trusting this
    
    Everything above is measured on three batteries: a 300-call deterministic calibration
    (references/calibration-2026-07-15.md), a 14-task hidden-test coding suite (2026-08-17,
    ~190 subagent runs), and a 14-repo seeded-bug agentic battery (2026-09-03, ~120 subagent
    runs, `oaustegard/experiments` → `temporal-routing-headroom`) that measured the cascade
    rungs, the escalation signal, and the tier gap against each other. Re-measure when:
    
    - **A model or price revision lands.** Both the verbosity multipliers and the
      cost table above invert on either. Sonnet 5.5 (2026-09-29) is such a revision. Its
      effort levels are measured above on the seeded-bug battery only; the other Sonnet
      figures in this skill are Sonnet 5 data, and the verbosity multipliers and the
      generation-suite costs have not been re-run on the 5.5 generation.
    - **The task family is off all three batteries.** No deterministic task has made Haiku
      fail on correctness yet, so the capability cliff is past what's been probed. Seeded-bug
      repair in a small module is now measured as *not* tier-separating: three probe shapes
      aimed at thoroughness, at ambiguity the tests underdetermine, and at a repo with no test
      suite at all, and `sonnet` @ `low` solved all six cells against `opus` @ `high`.
    - **Output length differs materially** from what was measured. The whole cost model
      keys on token volume; a 10× longer artifact re-opens the tier question.
    - **You need pass-rate differences of 1–2 tasks.** Run-to-run variance swamps them:
      two runs of the same model on the same 14 tasks produced disjoint failure sets and
      a 23% token gap. Token deltas are trustworthy; small pass-rate deltas are not.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related