{"slug":"escalation-governance","title":"escalation-governance","summary":"Assess whether to escalate models. Use when evaluating reasoning depth.","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-11T17:35:04.988466Z","repo":{"url":"https://github.com/athola/claude-night-market","stars":340,"forks":37,"license":"MIT","updatedAt":"2026-09-30T04:53:30Z"},"bodyHtml":"<hr>\n<p>name: escalation-governance\ndescription: 'Assess whether to escalate models. Use when evaluating reasoning depth.'\nalwaysApply: false\ncategory: agent-workflow\ntags:</p>\n<ul>\n<li>escalation</li>\n<li>model-selection</li>\n<li>governance</li>\n<li>agents</li>\n<li>orchestration\ndependencies: []\nestimated_tokens: 800\nmodel_hint: standard</li>\n</ul>\n<hr>\n<h2>Table of Contents</h2>\n<ul>\n<li><a href=\"#overview\">Overview</a></li>\n<li><a href=\"#the-iron-law\">The Iron Law</a></li>\n<li><a href=\"#when-to-escalate\">When to Escalate</a></li>\n<li><a href=\"#when-not-to-escalate\">When NOT to Escalate</a></li>\n<li><a href=\"#decision-framework\">Decision Framework</a></li>\n<li><a href=\"#1-have-i-understood-the-problem\">1. Have I understood the problem?</a></li>\n<li><a href=\"#2-have-i-investigated-systematically\">2. Have I investigated systematically?</a></li>\n<li><a href=\"#3-is-escalation-the-right-solution\">3. Is escalation the right solution?</a></li>\n<li><a href=\"#4-can-i-justify-the-trade-off\">4. Can I justify the trade-off?</a></li>\n<li><a href=\"#escalation-protocol\">Escalation Protocol</a></li>\n<li><a href=\"#common-rationalizations\">Common Rationalizations</a></li>\n<li><a href=\"#agent-schema\">Agent Schema</a></li>\n<li><a href=\"#orchestrator-authority\">Orchestrator Authority</a></li>\n<li><a href=\"#red-flags-stop-and-investigate\">Red Flags - STOP and Investigate</a></li>\n<li><a href=\"#integration-with-agent-workflow\">Integration with Agent Workflow</a></li>\n<li><a href=\"#quick-reference\">Quick Reference</a></li>\n</ul>\n<h1>Escalation Governance</h1>\n<h2>Overview</h2>\n<p>Model escalation (haiku→sonnet→opus) trades speed/cost for reasoning capability. This trade-off must be justified.</p>\n<p><strong>Core principle:</strong> Escalation is for tasks that genuinely require deeper reasoning, not for \"maybe a smarter model will figure it out.\"</p>\n<h2>The Iron Law</h2>\n<pre><code>NO ESCALATION WITHOUT INVESTIGATION FIRST\n</code></pre>\n<p><strong>Verification:</strong> Run the command with <code>--help</code> flag to verify availability.</p>\n<p>Escalation is never a shortcut. If you haven't understood why the current model is insufficient, escalation is premature.</p>\n<h2>When to Escalate</h2>\n<p><strong>Legitimate escalation triggers:</strong></p>\n<table>\n<thead>\n<tr>\n<th>Trigger</th>\n<th>Description</th>\n<th>Example</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Genuine complexity</td>\n<td>Task inherently requires nuanced judgment</td>\n<td>Security policy trade-offs</td>\n</tr>\n<tr>\n<td>Reasoning depth</td>\n<td>Multiple inference steps with uncertainty</td>\n<td>Architecture decisions</td>\n</tr>\n<tr>\n<td>Novel patterns</td>\n<td>No existing patterns apply</td>\n<td>First-of-kind implementation</td>\n</tr>\n<tr>\n<td>High stakes</td>\n<td>Error cost justifies capability investment</td>\n<td>Production deployment</td>\n</tr>\n<tr>\n<td>Ambiguity resolution</td>\n<td>Multiple valid interpretations need weighing</td>\n<td>Spec clarification</td>\n</tr>\n</tbody>\n</table>\n<h2>When NOT to Escalate</h2>\n<p><strong>Illegitimate escalation triggers:</strong></p>\n<table>\n<thead>\n<tr>\n<th>Anti-Pattern</th>\n<th>Why It's Wrong</th>\n<th>What to Do Instead</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>\"Maybe smarter model will figure it out\"</td>\n<td>This is thrashing</td>\n<td>Investigate root cause</td>\n</tr>\n<tr>\n<td>Multiple failed attempts</td>\n<td>Suggests wrong approach, not insufficient capability</td>\n<td>Question your assumptions</td>\n</tr>\n<tr>\n<td>Time pressure</td>\n<td>Urgency doesn't change task complexity</td>\n<td>Systematic investigation is faster</td>\n</tr>\n<tr>\n<td>Uncertainty without investigation</td>\n<td>You haven't tried to understand yet</td>\n<td>Gather evidence first</td>\n</tr>\n<tr>\n<td>\"Just to be safe\"</td>\n<td>False safety - wastes resources</td>\n<td>Assess actual complexity</td>\n</tr>\n</tbody>\n</table>\n<h2>Decision Framework</h2>\n<p>Before escalating, answer these questions:</p>\n<h3>1. Have I understood the problem?</h3>\n<ul>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Can I articulate why the current model is insufficient?</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Have I identified what specific reasoning capability is missing?</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Is this a capability gap or a knowledge gap?</li>\n</ul>\n<p><strong>If knowledge gap:</strong> Gather more information, don't escalate.</p>\n<h3>2. Have I investigated systematically?</h3>\n<ul>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Did I read error messages/outputs carefully?</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Did I check for similar solved problems?</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Did I form and test a hypothesis?</li>\n</ul>\n<p><strong>If not investigated:</strong> Complete investigation first.</p>\n<h3>3. Is escalation the right solution?</h3>\n<ul>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Would a different approach work at current model level?</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Is the task inherently complex, or am I making it complex?</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Would breaking the task into smaller pieces help?</li>\n</ul>\n<p><strong>If decomposable:</strong> Break down, don't escalate.</p>\n<h3>4. Can I justify the trade-off?</h3>\n<ul>\n<li><input disabled=\"disabled\" type=\"checkbox\"> What's the cost (latency, tokens, money) of escalation?</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> What's the benefit (accuracy, safety, completeness)?</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Is the benefit proportional to the cost?</li>\n</ul>\n<p><strong>If not proportional:</strong> Don't escalate.</p>\n<h2>Escalation Protocol</h2>\n<p>When escalation IS justified:</p>\n<ol>\n<li><strong>Document the reason</strong> - State why current model is insufficient</li>\n<li><strong>Specify the scope</strong> - What specific subtask needs higher capability?</li>\n<li><strong>Define success</strong> - How will you know the escalated task succeeded?</li>\n<li><strong>Return promptly</strong> - Drop back to efficient model after reasoning task</li>\n</ol>\n<h2>Common Rationalizations</h2>\n<table>\n<thead>\n<tr>\n<th>Excuse</th>\n<th>Reality</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>\"This is complex\"</td>\n<td>Complex for whom? Have you tried?</td>\n</tr>\n<tr>\n<td>\"Better safe than sorry\"</td>\n<td>Safety theater wastes resources</td>\n</tr>\n<tr>\n<td>\"I tried and failed\"</td>\n<td>How many times? Did you investigate why?</td>\n</tr>\n<tr>\n<td>\"The user expects quality\"</td>\n<td>Quality comes from process, not model size</td>\n</tr>\n<tr>\n<td>\"Just this once\"</td>\n<td>Exceptions become habits</td>\n</tr>\n<tr>\n<td>\"Time is money\"</td>\n<td>Systematic approach is faster than thrashing</td>\n</tr>\n</tbody>\n</table>\n<h2>Agent Schema</h2>\n<p>Agents can declare escalation hints in frontmatter:</p>\n<pre><code>model: haiku\nescalation:\n  to: sonnet                 # Suggested escalation target\n  hints:                     # Advisory triggers (orchestrator may override)\n    - security_sensitive     # Touches auth, secrets, permissions\n    - ambiguous_input        # Multiple valid interpretations\n    - novel_pattern          # No existing patterns apply\n    - high_stakes            # Error would be costly\n</code></pre>\n<p><strong>Verification:</strong> Run the command with <code>--help</code> flag to verify availability.</p>\n<p><strong>Key points:</strong></p>\n<ul>\n<li>Hints are advisory, not mandatory</li>\n<li>Orchestrator has final authority</li>\n<li>Orchestrator can escalate without hints (broader context)</li>\n<li>Orchestrator can ignore hints (task is actually simple)</li>\n</ul>\n<h2>Orchestrator Authority</h2>\n<p>The orchestrator (typically Opus) makes final escalation decisions:</p>\n<p><strong>Can follow hints:</strong> When hint matches observed conditions\n<strong>Can override to escalate:</strong> When context demands it (even without hints)\n<strong>Can override to stay:</strong> When task is simpler than hints suggest\n<strong>Can escalate beyond hint:</strong> Go to opus even if hint says sonnet</p>\n<p>The orchestrator's judgment, informed by conversation context, supersedes static hints.</p>\n<h2>Red Flags - STOP and Investigate</h2>\n<p>If you catch yourself thinking:</p>\n<ul>\n<li>\"Let me try with a better model\"</li>\n<li>\"This should be simple but isn't working\"</li>\n<li>\"I've tried everything\" (but haven't investigated why)</li>\n<li>\"The smarter model will know what to do\"</li>\n<li>\"I don't understand why this isn't working\"</li>\n</ul>\n<p><strong>ALL of these mean: STOP. Investigate first.</strong></p>\n<h2>Integration with Agent Workflow</h2>\n<pre><code>**Verification:** Run the command with `--help` flag to verify availability.\nAgent starts task at assigned model\n├── Task succeeds → Complete\n└── Task struggles →\n    ├── Investigate systematically\n    │   ├── Root cause found → Fix at current model\n    │   └── Genuine capability gap → Escalate with justification\n    └── Don't investigate → WRONG PATH\n        └── \"Maybe escalate?\" → NO. Investigate first.\n</code></pre>\n<p><strong>Verification:</strong> Run the command with <code>--help</code> flag to verify availability.</p>\n<h2>Quick Reference</h2>\n<table>\n<thead>\n<tr>\n<th>Situation</th>\n<th>Action</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Task inherently requires nuanced reasoning</td>\n<td>Escalate</td>\n</tr>\n<tr>\n<td>Agent uncertain but hasn't investigated</td>\n<td>Investigate first</td>\n</tr>\n<tr>\n<td>Multiple attempts failed</td>\n<td>Question approach, not model</td>\n</tr>\n<tr>\n<td>Security/high-stakes decision</td>\n<td>Escalate</td>\n</tr>\n<tr>\n<td>\"Maybe smarter model knows\"</td>\n<td>Never escalate on this basis</td>\n</tr>\n<tr>\n<td>Hint fires, task is actually simple</td>\n<td>Override, stay at current model</td>\n</tr>\n<tr>\n<td>No hint fires, task is actually complex</td>\n<td>Override, escalate</td>\n</tr>\n</tbody>\n</table>\n<h2>Model Capability Notes</h2>\n<p><strong>MCP Tool Search (Claude Code 2.1.7+)</strong>: Haiku models do not support MCP tool search. If a workflow uses many MCP tools (descriptions exceeding 10% of context), those tools load upfront on haiku instead of being deferred. This can consume significant context. Consider escalating to sonnet for MCP-heavy workflows or ensure haiku agents use only native tools (Read, Write, Bash, etc.).</p>\n<p><strong>Claude.ai MCP Connectors (Claude Code 2.1.46+)</strong>: Users with claude.ai connectors configured may have additional MCP tools auto-loaded, increasing the total tool description footprint. This makes it more likely that haiku agents will exceed the 10% tool search threshold. When escalation decisions involve MCP-heavy workflows, factor in claude.ai connector tool count via <code>/mcp</code>.</p>\n<p><strong>Effort Controls as Escalation Alternative (Opus 4.6 / Claude Code 2.1.32+)</strong>: Opus 4.6 introduces adaptive thinking with effort levels (<code>low</code>, <code>medium</code>, <code>high</code>). The <code>max</code> level was removed in 2.1.72 for Opus 4.6, and <code>high</code> became the ceiling on that model. Claude Code 2.1.111 reintroduced <code>max</code> and added <code>xhigh</code> (between <code>high</code> and <code>max</code>) for Opus 4.7 only; on other models <code>xhigh</code> falls back to <code>high</code>. Symbols: ○ (low) ◐ (medium) ● (high) ◉ (xhigh) ★ (max). Use <code>/effort</code> (interactive slider since 2.1.111) or <code>/effort auto</code> to reset. Before escalating between models, consider whether adjusting effort on the current model would suffice:</p>\n<table>\n<thead>\n<tr>\n<th>Instead of...</th>\n<th>Consider...</th>\n<th>When</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Haiku → Sonnet</td>\n<td>Stay on Haiku</td>\n<td>Task is still deterministic, just needs more context</td>\n</tr>\n<tr>\n<td>Sonnet → Opus</td>\n<td>Opus@medium</td>\n<td>Moderate reasoning, not deep architectural analysis</td>\n</tr>\n<tr>\n<td>Opus@medium → \"maybe try again\"</td>\n<td>Opus@high or \"ultrathink\"</td>\n<td>Genuine complexity needing deeper reasoning</td>\n</tr>\n<tr>\n<td>Opus 4.7@high → escalate</td>\n<td>Opus 4.7@xhigh or @max</td>\n<td>Deep architectural analysis on Opus 4.7 specifically</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Default effort change (2.1.68+)</strong>: Opus 4.6 now\ndefaults to <strong>medium effort</strong> for Max and Team\nsubscribers. Use <code>/model</code> to change effort level, or\ntype \"ultrathink\" in your prompt to enable high effort\nfor the next turn.</p>\n<p><strong>Opus 4/4.1 removed (2.1.68+)</strong>: Opus 4 and 4.1 are\nno longer available on the first-party API. Users with\nthese models pinned are automatically migrated to\nOpus 4.6. No action needed for agents using <code>model</code>\nfrontmatter, as the migration is transparent.</p>\n<p><strong>Sonnet 4.5 → 4.6 migration (2.1.69+)</strong>: Sonnet 4.5\nusers on Pro/Max/Team Premium are automatically migrated\nto Sonnet 4.6. Agent model frontmatter referencing\nSonnet resolves transparently. The <code>--model</code> flags for\n<code>claude-opus-4-0</code> and <code>claude-opus-4-1</code> now correctly\nresolve to Opus 4.6 instead of deprecated versions.</p>\n<p><strong>Effort parameter fix (2.1.70+)</strong>: Fixed API 400 error\n<code>This model does not support the effort parameter</code> when\nusing custom Bedrock inference profiles or non-standard\nClaude model identifiers. Effort controls now work\nreliably across all deployment configurations.</p>\n<p><strong>Default Opus 4.6 on providers (2.1.73+)</strong>: Bedrock,\nVertex, and Microsoft Foundry now default to Opus 4.6\n(was Opus 4.1). Subagent <code>model: opus</code>/<code>sonnet</code>/<code>haiku</code>\naliases now resolve to the current version on all\nproviders; previously they were silently downgraded to\nolder versions (e.g., Opus 4.1 instead of 4.6). This\nfix means agent dispatch workflows on third-party\nproviders now match first-party API behavior.</p>\n<p><strong><code>modelOverrides</code> setting (2.1.73+)</strong>: Maps model\npicker entries to provider-specific IDs (Bedrock\ninference profile ARNs, Vertex version names, Foundry\ndeployment names). Use when routing model selections to\nspecific inference profiles. See the model optimization\nguide for configuration details.</p>\n<p><strong><code>/output-style</code> deprecated (2.1.73+)</strong>: Use <code>/config</code>\ninstead. Output style is now fixed at session start for\nbetter prompt caching.</p>\n<p><strong>Full model IDs in agent frontmatter (2.1.74+)</strong>: Agent\n<code>model:</code> fields now accept full model IDs (e.g.,\n<code>claude-opus-4-6</code>) in addition to aliases (<code>opus</code>,\n<code>sonnet</code>, <code>haiku</code>). Previously, full IDs were silently\nignored. Agents now accept the same values as <code>--model</code>.</p>\n<p>Effort controls do NOT replace the escalation governance\nframework: they provide an additional axis. The Iron Law\nstill applies: investigate before changing either model\nor effort level.</p>\n<h2>Exit Criteria</h2>\n<ul>\n<li><input disabled=\"disabled\" type=\"checkbox\"> A decision (escalate / stay) is stated with a named trigger from the \"When to Escalate\" or\n\"When NOT to Escalate\" tables, not a vague claim of complexity.</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> If escalation is recommended, the specific target model and the subtask scope are documented\nbefore the escalation occurs.</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> Any instance of escalation triggered by \"maybe a smarter model will figure it out\" is flagged\nas an Iron Law violation and blocked.</li>\n<li><input disabled=\"disabled\" type=\"checkbox\"> If the task is decomposable into smaller pieces that each fit the current model, that\ndecomposition is proposed instead of escalation.</li>\n</ul>\n","files":[{"path":"SKILL.md","sizeBytes":11579,"isText":true},{"path":"test-authority.md","sizeBytes":3402,"isText":true},{"path":"test-convenience.md","sizeBytes":2826,"isText":true},{"path":"test-false-complexity.md","sizeBytes":3071,"isText":true},{"path":"test-thrashing.md","sizeBytes":3489,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"notes-only","suspicious":0,"notes":1,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-24T06:49:46.197049Z","sha256":"F0C0397BD9E7FCE85E34B3B8F5CD714451B0F04EAC906453874903ED0C3B4F23","sizeBytes":11400},"review":null,"source":{"repositoryUrl":"https://github.com/athola/claude-night-market","path":"plugins/abstract/skills/escalation-governance","license":"MIT","commit":"904583125527ac9ac25c0604db68d3d19b836a8d","subtreeSha":"6423E23B68DDCD172045BAB68CA09B9ED31C89DD9A6F6D4C695F992E1DFDC176","lastSyncedAt":"2026-10-01T15:24:25.628638Z"},"reviewedAt":"2026-09-24T06:51:18.835198Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/athola/claude-night-market/tree/master/plugins/abstract/skills/escalation-governance"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install athola-claude-night-market@llmmart"},{"target":"git","command":"git clone https://github.com/athola/claude-night-market.git"}]}