Claude Skill

loop-ops

Design and safely run OUTER loops - scheduled discover-triage-implement-verify-escalate agent loops. Native primitives schedule; loop-ops governs. Risk-tier ladder (L1 report -> L3 unattended), STATE/run-log/budget spine, kill switch, pattern catalog. Triggers: outer loop, schedu

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download 0xdarkmatter-claude-mods-skills_loop-ops-3dfaf0b.zip · 92 KB
Part of 0xdarkmatter/claude-mods — 94 skills

Install

skills CLI npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/loop-ops
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
Git git clone https://github.com/0xDarkMatter/claude-mods.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Loop Ops — outer-loop design discipline

A loop is not a prompt. Turn-by-turn prompting puts you in the loop forever. Loop engineering inverts it: you design a recurring process with memory, verification, and boundaries that discovers work, hands it to agents, verifies the result, and decides — on a schedule or until a goal is met — whether to land it or escalate to a human.

"You shouldn't be prompting coding agents anymore. You should be designing the loops that prompt your agents." — Peter Steinberger

This skill is the outer loop: the orchestration layer above a single agent run. It is the twin of iterate — iterate is the inner loop (one metric, one session, git-as-memory); loop-ops is the design discipline for the loop that schedules and gates inner runs. It does not reimplement spawning or landing; it composes what this repo already ships.

Native primitives schedule; loop-ops governs

Claude Code now ships the cadence half natively. Do not hand-roll a scheduler — pick a native host, declare it as host: in the config, and spend the discipline where the primitives leave a hole. Verified surface, parameters and limits (2026-08-30): references/native-scheduling.md.

Native primitive What it gives you What it does NOT give you
/loop — a bundled skill (docs), driving CronCreate/CronList/CronDelete; ScheduleWakeup for its self-paced mode Fixed-cron or Claude-paced ticks (delay clamped 60 s–1 h), a built-in maintenance prompt, .claude/loop.md to override it, Esc to stop Session-scoped and in-memory — fires only while the session is idle, dies with the conversation, and every recurring job self-deletes after 7 days. No state spine, no budget, no gate. L1-supervised only.
Desktop scheduled tasks — the scheduled-tasks MCP server (docs) Durable local ticks (≥1 min) with local files, a fresh session per run, a per-task permission mode with saved approvals, a task folder, run history, an Active/Paused toggle The worktree toggle is off by default (runs against uncommitted changes); one catch-up only for a missed window; a Manual-mode task stalls on an unapproved tool. No STATE spine, no token budget, no verify gate.
Cloud routines — /schedule (docs) Machine-off ticks (≥1 h), plus native event triggers: an API /fire endpoint and GitHub pull_request/release events with filters. A real push guard on non-claude/ branches No permission mode at all and every connector attaches by default; no local files (fresh clone); green run status ≠ task success. The boundary must come from repos + environment + connectors.
/goal (docs) A native completion gate — keep going until a fast model confirms the condition Not a cadence, and not an audit trail.

What none of them provide — and what this skill is for: a state spine that survives ticks, a token budget, a verify gate you can trust, an escalation rule, and the risk-tier ladder that decides whether the loop has earned the autonomy you're about to grant it. The plumbing moved into the harness; the judgement did not.

When the native primitive is enough — stop here

Don't scaffold a loop for work the harness already does. Use the primitive raw when all of these hold:

  • it writes nothing you'd have to undo (watch a deploy, poll a build, remind you), or the only writes are ones you'll review anyway;
  • it is supervised or short-lived — you're watching, or it stops in a session;
  • nothing needs to be remembered between ticks beyond what's in the repo;
  • and you'd shrug if a tick silently didn't fire.

/loop 5m check if the deploy finished is a complete, correct answer. Wrapping it in a loop.config.yaml adds ceremony and no safety.

Reach for loop-ops the moment any one of those flips: the loop starts changing things, runs unattended, needs to know what it did last time, or its silence would cost you. That is the whole trigger — everything below is what to do once it fires.


The six primitives → what owns each here

Every durable loop rests on six primitives. The discipline is wiring them; the parts already exist:

Primitive What it is Owned in claude-mods by
Schedule fire the loop on a cadence or an event native-first, declared as host: — session-cron (/loop+CronCreate, L1 only), desktop-task (scheduled-tasks MCP: local + durable), cloud-routine (/schedule: machine-off, plus API/GitHub event triggers), /goal for completion. external (cron/Task Scheduler + loop-run.sh) only for non-Claude-Code control
Worktree isolated, discardable execution context git-ops worktrees, fleet-worker (per-task worktree)
Skills persistent project knowledge the run loads this repo's skill layer + your CLAUDE.md
Sub-agents maker/checker separation Agent/Task; dispatching skills (review, testgen)
Connectors reach tickets / CI / chat MCP tools, gh, github-ops
+ State a durable spine outside the conversation STATE.md + run-log + budget (this skill)

The inner improvement loop is iterate; cheap parallel makers are fleet-worker; the test-gated merge queue is fleet-ops; inter-loop signalling is pigeon. loop-ops is the doctrine that connects them.

Loop anatomy

   ┌──────────────────────────────────────────────────────────────┐
   │  SCHEDULE (cadence)                                           │
   │     └─▶ TRIAGE      read STATE.md → pick the next unit of work │
   │           └─▶ WORKTREE   isolate (git worktree)               │
   │                 └─▶ MAKER     implementer run (or fleet-worker)│
   │                       └─▶ CHECKER  verify gate + guard (tests) │
   │                             └─▶ GATE  safe & allowlisted?      │
   │                                   ├─ yes → LAND  (commit/PR)   │
   │                                   └─ no  → ESCALATE (+context) │
   │     └─▶ write STATE.md, append run-log, decrement budget ──────┘

The gate is the load-bearing decision. Everything before it is mechanical; the gate is where a loop earns the right to run unattended — or doesn't.

The risk-tier ladder (the heart of the discipline)

Never start a loop unattended. Graduate it. Each tier maps to a concrete Claude Code permission mode — full mapping, the headless-profile table, and the enumerate vs isolate fork in references/risk-tiers.md.

Tier Posture Permission mode May do Lands by
L1 Report read-only discovery + triage plan / dontAsk+read allowlist scan, summarize, propose — writes nothing a human reads the report
L2 Assisted suggest changes, human gates the merge dontAsk+narrow allowlist, or auto edit in a worktree, run tests, open a PR a human approves the PR (or fleet-ops)
L3 Unattended autonomous land within a denylist bypassPermissions in an isolated container only commit/merge allowlisted classes the loop itself, inside its boundary

The host is part of the tier. session-cron (/loop + CronCreate) cannot host L2+ at all: it needs an open idle session and every recurring job expires after 7 days. cloud-routine has no permission mode, so its tier is expressed as repos + environment network policy + connectors instead — which means a routine is effectively autonomous the moment it is created, and the L1 posture has to come from the prompt being read-only. loop-doctor enforces both. Details: references/native-scheduling.md.

The cardinal rule, straight from Claude Code's own gate model: an unattended loop is a scheduler/script that invokes claude -p, not a Claude session that spawns ungated children. A session in auto mode that tries to launch a --permission-mode bypassPermissions child is blocked as Create Unsafe Agents — by design. See references/risk-tiers.md and the repo's auto-mode-classifier reference.

The escalation gate

What a loop may land vs what it must escalate is not a vibe — it mirrors Claude Code's classifier tiers. Bake these into the config's escalation: field:

  • Always escalate (never auto-land): force-push, push to main, production deploys or migrations, mass deletion, granting IAM/repo permissions, anything destroying pre-session files, editing .claude//settings (self-modification), curl | bash.
  • Safe to auto-land at L2/L3 (when allowlisted): a green PR on a feature branch, a lockfile patch bump that passes the guard, a generated changelog draft, a label/ triage classification, a comment.
  • The test: would a careful human let this happen unattended in this repo? If the action's blast radius exceeds the loop's stated purpose, it escalates. A general goal ("keep CI green") is not authorization for a specific high-blast action it implies.
  • Scope the tools, not just the mode. Allowlist exactly the tools/MCP connectors the job needs (read-only at L1); keep gh pr merge out and land_via: fleet-ops in. Full connector/MCP-scope discipline + the auto-merge guard: references/risk-tiers.md. On a cloud routine this is the whole gate — there is no permission mode, and every connected connector is attached by default with full write access. Prune them.
  • A task that reschedules itself is self-modification. The scheduled-tasks MCP lets a running task call update_scheduled_task on its own schedule or prompt. Useful, and on the always-escalate list unless adaptive cadence is the loop's stated purpose — a loop that can rewrite its own trigger has left the boundary you audited.

The state spine

A loop's memory lives outside the conversation, in three files (schemas + read/write contract in references/state-spine.md):

  • STATE.md — the triage snapshot: priority / watch / noise + a readiness line. Read at the top of every run, rewritten at the end.
  • run-log.md — append one line per run (timestamp, action, outcome, tokens). The audit trail that answers "what has this loop been doing?"
  • loop.config.yaml — the loop's definition (goal, tier, cadence, host, scope, gate, budget, escalation). Scaffolded by loop-scaffold, scored by loop-check.

A native host gives you a place for this spine (a Desktop task's folder) but never the spine itself: no host writes STATE.md, enforces a token budget, or records what the loop decided. Run history says a tick happened; the run-log says what it did and cost.

Pattern catalog (a morphology, not a fixed list)

Patterns are compositions of three axes — trigger (cadence / event via a Channel / goal) × posture (L1/L2/L3) × locus (connector→cloud routine / local→Desktop task). The named patterns are well-trodden points in that space; compose your own from the axes. Full recipes + the morphology in references/pattern-catalog.md:

Pattern Trigger · Locus Tier One-line job
daily-scan cadence · local L1 discover + prioritize, report only
pr-watch event|cadence · connector L1 watch review state, surface stuck PRs
ci-watch event · local L2 triage build failures, propose a fix
dep-bump cadence · local L2 patch-only bumps behind cooldown + guard
changelog-gen event(tag)|cadence · local L1 draft release notes for approval
merge-hygiene cadence · local L1 dead branches, stale flags
issue-sort cadence · connector L1 classify + label, propose only
metric-chase goal · local L2 drive a metric (coverage/latency/eval) via iterate
regression-watch cadence|event · local L1 run a benchmark/eval, flag a regression
digest cadence · connector L1 summarize email/Asana/news (cloud routine)
backfill goal · local L2 drain a migration/queue to completion
monitor event · local L1 error/deploy webhook → triage + page
freshness cadence · local L1 re-check docs/data/deps vs reality

Start any pattern at L1. Graduate to L2 only after the L1 reports prove its judgment. Prefer event over cadence where a webhook exists (cheaper, faster than polling).

Multi-loop coordination & the kill switch

Running several loops? Two non-negotiables (detail in references/state-spine.md):

  • Priority order prevents collisions: CI Watch → PR Watch → Dependency Bump → Merge-Hygiene/Changelog → Daily Scan (off-peak). A higher-priority loop's worktree wins; lowers defer. Loops signal each other via pigeon.
  • A kill switch every loop honors. A single stop signal — a PAUSED sentinel file or a loop-pause label — that every loop checks at the top of its run and exits on. No loop ships without one. Put it in kill_switch: and check it first.

Composition map — don't rebuild what exists

You need to… Use Not
improve one metric in one session iterate a hand-rolled inner loop
spawn cheap parallel makers fleet-worker bespoke claude -p plumbing
route models across a fan-out (cheap finders, Opus judges) fleet-worker model-routing every agent on the orchestrator's model
test-gate + land winning branches fleet-ops a manual merge step
fire on a cadence or an event a native host: — /loop, Desktop scheduled task, cloud routine (schedule/API/GitHub triggers); /goal for completion a custom cron in this skill
trust the verify gate's judgement the evals-ops skill — a gate is an eval (golden set, judge bias, pass^k, blocking vs advisory) eyeballing a few runs and calling it proven
reason about per-tick prompt-cache cost claude-api-ops caching-and-cost a TTL number memorised from a blog post
commit / PR / release git-ops, github-ops raw git push
signal between loops pigeon a shared scratch file

loop-ops is the design layer; these are the execution layers.


Tools

Six scripts, all following the Skill Resource Protocol (stdout = data, semantic exit codes, --help with EXAMPLES, --json envelopes): init scaffolds the loop, audit scores whether the config is well-formed, doctor preflights whether it will actually run (host-aware), cost estimates spend (caching-aware), and two drift guards — check-pricing-sync for the pricing table and check-native-facts for the native-scheduling limits. The discipline before scheduling is init → fill → cost → audit → doctor --live.

scripts/loop-scaffold.sh — scaffold a loop's state spine

Writes <dir>/<name>/ with five files from the bundled templates: loop.config.yaml (assets/loop.config.template.yaml), STATE.md (assets/STATE.template.md), run-log.md, run.md (the headless run prompt, assets/run.template.md), and an executable loop-run.sh (assets/run.sh.template) — the runner-agnostic tick wrapper any scheduler invokes (cron / Windows Task Scheduler / systemd / by hand), no GitHub Actions required. Pass a known --pattern (pr-watch, ci-watch, dep-bump, …) and the config is seeded with that pattern's scope/goal/escalation — and, at L2+, its gate — so you get a near-ready config to review, not blank placeholders (it audits clean immediately). Doctrine holds: it still scaffolds at L1 by default with a graduation block.

--host records where ticks will execute (local default, or session-cron / desktop-task / cloud-routine / external) so loop-doctor checks that host's real constraints instead of assuming a local claude -p.

# Create .loops/pr-watch/ with config + STATE.md + run-log.md + run.md from templates:
bash scripts/loop-scaffold.sh --name pr-watch --pattern pr-watch --tier L1

# A connector-driven loop bound for a cloud routine (>=1h floor, no permission mode):
bash scripts/loop-scaffold.sh --name digest --pattern digest --host cloud-routine --cadence 1h

# Custom dir + cadence, preview without writing:
bash scripts/loop-scaffold.sh --name dep-bump --pattern dep-bump \
  --tier L2 --cadence 1d --dir .loops --dry-run

Refuses to overwrite a populated <dir>/<name>/ (exit 5) unless --force. Atomic writes. --dry-run prints what it would create and writes nothing. stdout = the created config path.

scripts/loop-check.sh — readiness scorer (run before you schedule)

The question this answers: is this loop safe to turn on at its declared tier? It scores a loop.config.yaml against the readiness rubric — gate present, scope bounded, escalation defined, guard + worktree at L2+, budget + kill switch set, permission mode consistent with tier — and refuses a green light if any critical gap exists.

bash scripts/loop-check.sh .loops/pr-watch/loop.config.yaml   # exit 0 ready, 10 not ready
bash scripts/loop-check.sh --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.severity=="error")'
bash scripts/loop-check.sh --min 80 .loops/ci-watch/loop.config.yaml   # raise the score bar

Exit 0 = ready (no errors, score ≥ --min), 10 = not ready (findings on stdout), 2 usage, 3 config not found, 4 config unparseable. --strict counts warnings toward the not-ready signal.

scripts/loop-doctor.sh — live preflight (will it actually run?)

loop-check proves the config is well-formed; loop-doctor proves the loop will execute — catching the "blocked at 3am" failures audit can't see. --offline (CI-safe): the budget fits a tick's estimated tokens, the permission mode is achievable (not interactive), an L3 bypass declares an isolation boundary. --live adds runtime preflight: the verify/guard gate's leading binary resolves on PATH, claude/git are present, the kill-switch sentinel's parent dir exists.

It is host-aware. host: changes what "will it run" even means, so the doctor checks against the declared surface: a cloud-routine faster than its 1-hour floor is rejected at creation; a routine with no named repos/environment/connector boundary has no gate at all (it has no permission mode either, so demanding one there would be a false finding); a session-cron host at L2+ can't run unattended and is called a predicted failure; and --live is skipped, not passed, for a cloud routine — this machine's PATH says nothing about a fresh cloud clone, and a green check there would be false confidence.

bash scripts/loop-doctor.sh --offline .loops/pr-watch/loop.config.yaml   # CI gate
bash scripts/loop-doctor.sh --live .loops/ci-watch/loop.config.yaml          # before scheduling
bash scripts/loop-doctor.sh --live --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.state=="bad")'
bash scripts/loop-doctor.sh --offline .loops/digest/loop.config.yaml   # host: cloud-routine -> floor + boundary

Exit 0 = will run, 10 = a check predicts a runtime failure (gate binary missing, bypass on host without isolation, budget too small for a tick), 2 usage, 3 not found, 4 unparseable, 5 missing core dep. Run it after loop-check and before scheduling.

scripts/loop-estimate.py — token/$ estimate by pattern × cadence × model (caching-aware)

Estimate spend before committing to a cadence — the cost of an outer loop is runs/day × tokens/run × price, and sub-agents multiply it. It also models prompt caching: a loop re-sends the same run.md+system prefix every tick (the Ralph property), so the prefix should be cache-written once then read (~0.1×) — but only if the tick interval fits the cache TTL. The TTL is a choice, not a constant: 5 minutes by default (1.25× write) or 1 hour with "ttl": "1h" (2× write), so the daemon window is ~4.5 min or ~55 min — not a fixed 270 s. The estimator picks the cheapest TTL that stays warm at your cadence and names it; past 1 h nothing caches at all. Mechanics and break-even: claude-api-ops caching-and-cost. The estimate itself is host-agnostic — tokens are tokens wherever the tick fires; the host-dependent limit is the minimum cadence, which loop-doctor enforces. Pricing reads from assets/model-pricing.json (date-stamped; claude-api-ops is the source of truth — run its check-model-table.py if you suspect drift).

python scripts/loop-estimate.py --pattern pr-watch --cadence 10m --model claude-haiku-4-5
python scripts/loop-estimate.py --pattern ci-watch --cadence 15m --model claude-sonnet-5 --days 30 --json
python scripts/loop-estimate.py --list-models      # the pricing table + its as-of date

Exit 0 ok, 2 usage, 3 pricing file missing, 4 bad cadence/model. Output names every assumption (runs/day, tokens/run, sub-agent multiplier) — it's an estimate, and it says so.

scripts/check-pricing-sync.py — offline drift guard (CI)

model-pricing.json is a copy of claude-api-ops's authoritative model table, and a copy drifts silently. This offline verifier asserts every model in assets/model-pricing.json matches claude-api-ops's "Current Models" table (prices included). Both files are in-repo, so it's network-free and gates PR CI via tests/check-resources.sh; live model-id drift is owned by claude-api-ops's check-model-table.py.

python scripts/check-pricing-sync.py --offline   # exit 0 in sync, 10 drift, 3 a file missing

scripts/check-native-facts.py — native-scheduling staleness guard

references/native-scheduling.md encodes a fast-moving external surface, and loop-doctor refuses configs on those numbers — so a silently stale limit becomes a wrong refusal. --offline (PR CI) proves internal consistency: the host vocabulary is one set across the config template, loop-scaffold --host, loop-doctor's case arm and the reference; the reference still carries its Verified <date> stamp; and every limit the doctor enforces is still stated in the prose that justifies it. --live (scheduled, never a PR gate) fetches the three published docs pages and checks our numbers still appear in them.

python scripts/check-native-facts.py --offline   # exit 0 in sync, 10 drift, 3 file missing
python scripts/check-native-facts.py --live      # exit 7 = docs unreachable (advisory)

End-to-end workflow

  1. Pick a pattern from the catalog (or custom), and pick the host — does the tick need local files? must it run with the machine off? is it supervised? Start at L1.
  2. Scaffold: bash scripts/loop-scaffold.sh --name <n> --pattern <p> --tier L1 --host <h>.
  3. Fill loop.config.yaml — the real goal, scope (bounded globs, never *), verify gate, escalation rule, budget_tokens, kill_switch. On a cloud-routine, name the boundary that replaces the absent permission mode: repos, environment network policy, and the connectors you kept.
  4. Cost it: python scripts/loop-estimate.py --pattern <p> --cadence <c> --model <m> — sanity-check the monthly spend against the value.
  5. Audit it: bash scripts/loop-check.sh .loops/<n>/loop.config.yaml — fix every error before scheduling. Don't schedule a loop that fails its own audit.
  6. Doctor it: bash scripts/loop-doctor.sh --live .loops/<n>/loop.config.yaml — prove it will actually run (gate binary on PATH, budget fits a tick). Audit = well-formed; doctor = will-run.
  7. Schedule the L1 run on the declared host — the recipe selector in references/claude-code-loops.md prescribes which, because they're not interchangeable: connector-driven (email/Asana, no local code) → cloud routine; touches local code → Desktop scheduled task; sustained & token-sensitive → a cache-warm daemon (claude -p inside the cache TTL you paid for), not /loop (which grows a session and chews tokens); fixed-criteria long task → /goal; quick supervised polling → /loop. Per-primitive limits: references/native-scheduling.md. (L1 is read-only — it just writes STATE.md + a report.)
  8. Read the reports. Only after the loop's judgment is proven do you graduate it to L2 (worktree + guard + fleet-ops landing), change host: if the proving host was session-cron, and re-audit at the higher tier. If the gate's verdict is a judgement rather than a green test run, harden it with the evals-ops discipline before you let it decide unattended.

Worked example

A complete, audit + doctor-clean L1 loop ships at assets/examples/pr-watch/: a filled loop.config.yaml, a populated STATE.md, the run.md run prompt, a sample run-log.md, the runner-agnostic loop-run.sh (the tick wrapper, with the kill-switch gate and dontAsk + allowlist baked in — point cron / Task Scheduler at it), and an optional github-actions.yml for repos already on GitHub. Copy the dir, adjust scope/cadence, run loop-check + loop-doctor --live, then wire loop-run.sh to your scheduler. The other patterns don't ship as static dirs that rot — loop-scaffold --pattern <name> generates the same, seeded and gate-clean, for any pattern at any tier. CI runs loop-check + loop-doctor on this example every build, so it can't drift out of validity.

Anti-patterns (these are detected and wrong)

The incident-shaped catalog — symptom → mechanism → the control that catches each — is references/failure-modes.md (runaway budget, the 3am-dead loop, cache-cold, force-push, ungated-child spawn, colliding loops, silent-stop, gate reward-hacking, and the native-host trio — the expired 7-day loop, the over-connected routine, the stalled/skipped Desktop task). The headline ones:

  • Routing around the gate. Wrapping claude -p --permission-mode bypassPermissions in a script to dodge the classifier is Auto-Mode Bypass — a hard_deny nothing clears. If an outcome is blocked, authorize it (a narrow allow rule, or run the scheduler outside the auto-mode session), never disguise it.
  • The orchestrator session spawning ungated children. A session in auto mode is the wrong place to launch the loop. The scheduler/cron/Task-Scheduler/CI runner that invokes claude -p is the authorizer. See references/risk-tiers.md §"enumerate vs isolate".
  • No gate. A loop whose verify: is empty is not a loop, it's an unsupervised typer. loop-check errors on it. Nor is a green run status a gate: on a cloud routine green means "the session started and exited without an infrastructure error", never that the task succeeded. Grade the work, not the process.
  • Assuming the native host gave you a boundary. It gave you a cadence. A Desktop task's worktree toggle is off by default; a cloud routine has no permission mode and attaches every connector; /loop evaporates after 7 days. Each is a default that reads as safe and isn't.
  • Unbounded scope. scope: "*" means "may touch anything" — the audit rejects it.
  • No kill switch / no budget. A loop you can't stop, or whose spend you didn't bound, will eventually surprise you. Both are audit findings.
  • Skipping L1. Starting a fresh loop at L3 is how comprehension debt and incidents compound. The ladder exists precisely so trust is earned before it's granted.

See also

Files (claude-mods)
  • assets
    • examples
      • pr-watch
        • github-actions.yml 2.7 KB
          # OPTIONAL scheduler — use this ONLY if your repo already lives on GitHub. The portable,
          # runner-agnostic path is loop-run.sh (cron / Windows Task Scheduler / systemd / by hand) —
          # no GitHub Actions dependency. This file just wraps the same loop-run.sh idea in Actions.
          # Copy to .github/workflows/pr-watch.yml, PIN the action/CLI versions, add the
          # ANTHROPIC_API_KEY secret. The SCHEDULER is the authorizer (no auto-mode session in the
          # loop), and the child runs gated (--permission-mode dontAsk + a narrow allowlist), never
          # bypassPermissions on a shared runner. See references/claude-code-loops.md.
          name: pr-watch
          on:
            schedule:
              - cron: "*/10 * * * *"   # every 10 min (matches loop.config.yaml cadence: 10m)
            workflow_dispatch: {}
          
          permissions:
            contents: write            # commit STATE.md / run-log.md back
            pull-requests: write       # post the at-most-one summary comment (L1 stays report-only)
          
          concurrency:
            group: pr-watch       # never overlap two ticks
            cancel-in-progress: false
          
          jobs:
            tick:
              runs-on: ubuntu-latest
              steps:
                - uses: actions/checkout@v4   # <-- pin to a SHA in production
          
                # Kill switch: a 'loop-pause' label on the repo, or a committed PAUSED sentinel.
                - name: Honor the kill switch
                  id: gate
                  env: { GH_TOKEN: "${{ github.token }}" }
                  run: |
                    if [ -f .loops/pr-watch/PAUSED ]; then echo "paused=1" >> "$GITHUB_OUTPUT"; fi
                    if gh label list --limit 100 | grep -qi '^loop-pause'; then echo "paused=1" >> "$GITHUB_OUTPUT"; fi
          
                - name: Install Claude Code
                  if: steps.gate.outputs.paused != '1'
                  run: npm i -g @anthropic-ai/claude-code   # <-- pin a version
          
                # The run: same prompt every tick (cache-friendly), gated with dontAsk + an
                # allowlist scoped to exactly what an L1 report loop needs (read-only + gh + STATE writes).
                - name: Run one tick
                  if: steps.gate.outputs.paused != '1'
                  env:
                    ANTHROPIC_API_KEY: "${{ secrets.ANTHROPIC_API_KEY }}"
                  run: |
                    cd .loops/pr-watch
                    claude -p "$(cat run.md)" \
                      --permission-mode dontAsk \
                      --append-system-prompt "$(cat STATE.md)" \
                      --allowedTools 'Bash(gh pr list:*)' 'Bash(gh pr view:*)' 'Bash(gh pr comment:*)' 'Read' 'Write(STATE.md)' 'Write(run-log.md)' \
                      --max-turns 30
          
                - name: Persist STATE + run-log
                  if: steps.gate.outputs.paused != '1'
                  run: |
                    git config user.name  "pr-watch-loop"
                    git config user.email "loop@users.noreply.github.com"
                    git add .loops/pr-watch/STATE.md .loops/pr-watch/run-log.md
                    git diff --cached --quiet || git commit -m "chore(loop): pr-watch tick $(date -u +%FT%TZ)"
                    git push
          
        • loop-run.sh 1.6 KB
          #!/usr/bin/env bash
          # loop-run.sh - one tick of the pr-watch loop. RUNNER-AGNOSTIC: point any scheduler
          # at it. No GitHub Actions required.
          #   cron:                 */10 * * * *  /path/.loops/pr-watch/loop-run.sh >> tick.log 2>&1
          #   Windows Task Scheduler: schtasks /Create /SC MINUTE /MO 10 /TN pr-watch \
          #                             /TR "bash -lc '/path/.loops/pr-watch/loop-run.sh'"
          #   by hand:              bash loop-run.sh
          # The scheduler is the authorizer; this runs a gated `claude -p` (dontAsk + an allowlist),
          # never bypassPermissions on a shared host. (github-actions.yml is one OPTIONAL scheduler.)
          set -uo pipefail
          HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
          cd "$HERE"
          
          # 1. Kill switch first.
          if [ -f PAUSED ]; then echo "pr-watch: paused (PAUSED sentinel) - skipping tick" >&2; exit 0; fi
          command -v claude >/dev/null 2>&1 || { echo "pr-watch: 'claude' not on PATH" >&2; exit 5; }
          
          # 2. One tick. SAME prompt every time (cache-friendly). Allowlist = exactly what an L1
          #    report loop needs: read-only gh + Read + the STATE/run-log writes. No 'gh pr merge'.
          claude -p "$(cat run.md)" \
            --permission-mode dontAsk \
            --append-system-prompt "$(cat STATE.md)" \
            --allowedTools 'Bash(gh pr list:*)' 'Bash(gh pr view:*)' 'Bash(gh pr comment:*)' 'Read' 'Write(STATE.md)' 'Write(run-log.md)' \
            --max-turns 30
          
          # 3. Persist STATE + run-log if this dir lives in a git repo.
          if git rev-parse --is-inside-work-tree >/dev/null 2>&1; then
            git add STATE.md run-log.md 2>/dev/null || true
            git diff --cached --quiet 2>/dev/null || git commit -q -m "chore(loop): pr-watch tick" || true
          fi
          
        • loop.config.yaml 1.1 KB
          # WORKED EXAMPLE — a complete, audit-clean L1 loop. Copy the dir, adjust scope/cadence,
          # run loop-check + loop-doctor --live, then wire the scheduler (github-actions.yml).
          # The other patterns: `loop-scaffold --pattern <name>` generates the same, seeded.
          name: pr-watch
          pattern: pr-watch
          tier: L1
          permission_mode: dontAsk
          cadence: 10m
          # Ticks run on this machine via the wrapper below, so the local host applies. A 10m
          # cadence rules out cloud-routine (1h floor); session-cron would work while you watch
          # it, but expires after 7 days. See references/native-scheduling.md.
          host: external
          goal: "Watch open PRs; flag stuck (no review > 4h), failing checks, and merge conflicts in STATE.md; post at most one summary comment per PR; never merge."
          scope:
            - "src/**"
          escalation: "a human reviews and merges; never merge to main; never push; never close a PR"
          budget_tokens: 60000
          kill_switch: ".loops/pr-watch/PAUSED exists, OR the 'loop-pause' label is on the repo"
          
          # ── graduate to L2 (assisted: open a fix-the-PR-description / rebase branch) ──
          # tier: L2
          # verify: "npm test"
          # guard: "npm run typecheck"
          # worktree: true
          # land_via: fleet-ops
          
        • run-log.md 436 B
          # pr-watch — run log (append-only; one line per run)
          # format: <ISO-Z>  run#N  action=<reported|none>  pr=<n|->  outcome=<…>  tokens=<N>
          2026-06-22T14:05:00Z  run#142  action=reported  pr=412  outcome=commented-flaky-ci   tokens=14820
          2026-06-22T13:55:00Z  run#141  action=reported  pr=408  outcome=flagged-conflict      tokens=12110
          2026-06-22T13:45:00Z  run#140  action=none      pr=-    outcome=quiet                 tokens=2090
          
        • run.md 1.7 KB
          <!--
          run.md — fed to `claude -p` each tick. SAME every run (fresh context; state lives in
          STATE.md + git, not the conversation). Keep it BYTE-IDENTICAL so the prompt cache hits.
          Wired by github-actions.yml:
            claude -p "$(cat run.md)" --permission-mode dontAsk --append-system-prompt "$(cat STATE.md)"
          -->
          
          # Run: pr-watch  (tier L1, report-only)
          
          You are one tick of a scheduled loop. Goal: **watch open PRs and report; never merge, push, or close.**
          
          ## Do these in order
          1. **Kill switch first.** If `.loops/pr-watch/PAUSED` exists or the repo has the `loop-pause` label, STOP — do nothing.
          2. **Read `STATE.md`** (in your system prompt): the Priority / Watch / Noise lists from last run.
          3. **List open PRs:** `gh pr list --state open --json number,title,reviewDecision,statusCheckRollup,mergeStateStatus,updatedAt`.
          4. **Classify each** PR: failing checks · merge conflict · awaiting-review past 4h · draft · healthy.
          5. **Report only.** You may post **at most one** summary comment per PR that newly needs attention (preview the text in the run log; never spam). You do **not** merge, push, rebase, or close — that escalates to a human.
          6. **Respect the budget:** stop if you approach 60000 output tokens.
          7. **Rewrite `STATE.md`:** move PRs across Priority / Watch / Noise; bump `_Updated_` + run number + readiness.
          8. **Append one line to `run-log.md`:** `<ISO-Z>  run#N  action=reported  pr=<n|->  outcome=<…>  tokens=<N>`.
          
          ## Hard rules
          - Never merge/push/close/rebase — those are the escalation cases. A general goal is not authorization for them.
          - Stay within scope (`src/**`); never touch another session's `.claude/worktrees/`.
          - Leave the repo clean every tick.
          
        • STATE.md 963 B
          # pr-watch — STATE
          _Updated: 2026-06-22T14:05:00Z · run #142 · readiness 100/100_
          
          <!-- Read first every run: check the kill switch, then act on the Priority list.
               This is a realistic populated snapshot — yours starts from STATE.template.md. -->
          
          ## Priority   (act on these next)
          - [P1] PR #412 failing CI 3h10m — `build` job red, looks like a flaky e2e; owner @dana pinged 14:02
          - [P1] PR #408 merge conflict with main (settings.ts) — author notified, awaiting rebase
          - [P2] PR #415 awaiting review 5h — past the 4h threshold; nudged the reviewers group
          
          ## Watch      (not yet actionable)
          - PR #417 awaiting review 1h12m — under threshold, recheck next run
          - PR #410 draft — skip until marked ready
          
          ## Noise      (seen + dismissed this run)
          - PR #419 opened 6m ago, checks still running — too early
          - PR #402 already merged since last run — drop from tracking
          
          ---
          _Source: .github/workflows/pr-watch.yml · config: loop.config.yaml_
          
    • loop.config.template.yaml 4.2 KB
      # loop.config.yaml — one OUTER-loop definition.
      # Scaffolded by loop-scaffold.sh, scored by loop-check.sh. Flat YAML on purpose: every
      # key sits at column 0 so the audit parses it without a yq dependency.
      # Full field semantics: skills/loop-ops/references/state-spine.md
      #
      # >>> ADAPT every <PLACEHOLDER> below. The audit errors on unbounded scope, a
      #     missing gate, an undefined escalation rule, or a tier/permission mismatch.
      
      # ── identity ────────────────────────────────────────────────────────────────
      name: <loop-name>            # matches the .loops/<name>/ directory
      pattern: <pattern-key>       # a catalog key (pr-watch, ci-watch, …) or "custom"
      
      # ── autonomy ────────────────────────────────────────────────────────────────
      tier: L1                     # L1 report-only | L2 assisted (worktree+gate) | L3 unattended
      permission_mode: dontAsk     # plan | dontAsk | auto | acceptEdits | bypassPermissions
                                   #   L1 → plan or dontAsk · L2 → dontAsk/auto · L3 → bypassPermissions (container only)
      
      # ── cadence & host ──────────────────────────────────────────────────────────
      cadence: 1h                  # 10m | 1h | 6h | 1d, or a cron string ("*/10 * * * *")
      host: local                  # where ticks execute — decides which constraints apply:
                                   #   local          generic local run (default; no host checks)
                                   #   session-cron   /loop + CronCreate — L1 only: needs an open
                                   #                  idle session, 7-day expiry, ≥1 min
                                   #   desktop-task   scheduled-tasks MCP — durable, local files,
                                   #                  per-task permission mode, ≥1 min
                                   #   cloud-routine  /schedule Routines — no local files, NO
                                   #                  permission mode, ≥1 hour
                                   #   external       cron / Task Scheduler / systemd → loop-run.sh
                                   # Full semantics: references/native-scheduling.md
      
      # ── purpose & bounds ────────────────────────────────────────────────────────
      goal: "<one sentence: what this loop does AND what it must never do>"
      scope:                       # bounded globs the loop may touch — NEVER "*" or "**"
        - "<src/**>"
      
      # ── the gate (required at L2+) ──────────────────────────────────────────────
      verify: "<command that decides pass/fail, e.g. npm test>"   # a loop with no gate is invalid at L2+
      guard: "<must-always-pass, e.g. npm run typecheck>"          # required at L2+
      
      # ── isolation & landing (required at L2+) ───────────────────────────────────
      worktree: true               # isolate code changes in a git worktree (required L2+)
      land_via: fleet-ops          # who test-gates + lands winning branches (L2+)
      
      # ── the escalation rule (required) ──────────────────────────────────────────
      # What the loop ESCALATES instead of doing. Mirror the never-auto-land classes:
      # force-push, push to main, prod deploy/migration, mass delete, IAM grants, .claude edits.
      escalation: "<e.g. open a PR with context; never merge to main; never deploy>"
      
      # ── safety rails ────────────────────────────────────────────────────────────
      budget_tokens: 200000        # per-run output-token ceiling (stop the run when reached)
      kill_switch: ".loops/<loop-name>/PAUSED exists, OR the loop-pause label is set"
      
    • model-pricing.json 2.1 KB
      {
        "_comment": "Per-model USD pricing per million tokens, read by loop-estimate.py. SOURCE OF TRUTH is skills/claude-api-ops/SKILL.md 'Current Models' table — run skills/claude-api-ops/scripts/check-model-table.py --live if you suspect drift. Date-stamp this file when updating.",
        "_as_of": "2026-08",
        "_schema": "claude-mods.loop-ops.pricing/v1",
        "models": {
          "claude-fable-5":    { "input_per_mtok": 10.0, "output_per_mtok": 50.0 },
          "claude-opus-5":     { "input_per_mtok": 5.0,  "output_per_mtok": 25.0 },
          "claude-sonnet-5":   { "input_per_mtok": 2.0,  "output_per_mtok": 10.0 },
          "claude-haiku-4-5":  { "input_per_mtok": 1.0,  "output_per_mtok": 5.0 }
        },
        "_pattern_defaults": {
          "_comment": "Rough per-run token estimates by pattern. input = context the run reads (STATE, diffs, tool results); output = tokens the model generates; subagents multiplies tokens when the pattern fans out makers/checkers. Estimates, not guarantees — reconcile against run-log.md actuals.",
          "daily-scan":        { "input": 40000,  "output": 6000,  "subagents": 1 },
          "pr-watch":       { "input": 15000,  "output": 3000,  "subagents": 1 },
          "ci-watch":          { "input": 60000,  "output": 12000, "subagents": 2 },
          "dep-bump":  { "input": 30000,  "output": 8000,  "subagents": 1 },
          "changelog-gen":   { "input": 25000,  "output": 9000,  "subagents": 1 },
          "merge-hygiene":  { "input": 20000,  "output": 4000,  "subagents": 1 },
          "issue-sort":        { "input": 18000,  "output": 4000,  "subagents": 1 },
          "metric-chase":      { "input": 50000,  "output": 20000, "subagents": 3 },
          "regression-watch":  { "input": 45000,  "output": 8000,  "subagents": 1 },
          "digest":            { "input": 30000,  "output": 7000,  "subagents": 1 },
          "backfill":          { "input": 40000,  "output": 15000, "subagents": 2 },
          "monitor":           { "input": 12000,  "output": 3000,  "subagents": 1 },
          "freshness":         { "input": 35000,  "output": 6000,  "subagents": 1 },
          "custom":              { "input": 30000,  "output": 8000,  "subagents": 1 }
        }
      }
      
    • run.sh.template 1.4 KB · in bundle
    • run.template.md 2.2 KB
      <!--
      run.md — the prompt a scheduler feeds to `claude -p` each tick. It is the SAME
      every run (fresh context each time; state lives in STATE.md + the codebase + git,
      not the conversation — the Ralph property). Fill the <PLACEHOLDERS>; keep it short.
      
      Wire it up (see references/claude-code-loops.md):
        claude -p "$(cat .loops/<loop-name>/run.md)" \
          --permission-mode <permission_mode-from-config> \
          --append-system-prompt "$(cat .loops/<loop-name>/STATE.md)"
      The SCHEDULER (cron / Task Scheduler / CI), not a Claude session, invokes this.
      -->
      
      # Run: <loop-name>  (tier <L1|L2|L3>)
      
      You are one tick of a scheduled loop. Goal: **<one sentence — what to do AND what never to do>**.
      
      ## Do these in order
      
      1. **Check the kill switch FIRST.** If <kill_switch — e.g. `.loops/<loop-name>/PAUSED` exists, or the `loop-pause` label is set>, STOP immediately and do nothing else.
      2. **Read `STATE.md`** (appended to your system prompt). It is your memory of prior runs: the Priority / Watch / Noise lists.
      3. **Pick the next unit of work** from the Priority list. Stay strictly within scope: `<scope globs — never *>`.
      4. **Do the tier-appropriate action:**
         - **L1 (report-only):** investigate and summarize. Write NOTHING but `STATE.md`.
         - **L2 (assisted):** make the change in a **git worktree**; run the gate `<verify>` and guard `<guard>`; if both pass, hand the branch to `<land_via — e.g. fleet-ops>`; otherwise discard.
      5. **Apply the escalation rule.** If the action would <escalation — e.g. force-push / push to main / deploy / delete pre-existing files / edit .claude>, do NOT do it — escalate to a human with context instead.
      6. **Respect the budget.** Stop this run if you approach `<budget_tokens>` output tokens.
      7. **Rewrite `STATE.md`:** promote/demote items across Priority / Watch / Noise, bump the `_Updated_` line + run number + readiness score.
      8. **Append one line to `run-log.md`:** `<ISO-Z>  run#N  action=<…>  outcome=<…>  tokens=<N>`.
      
      ## Hard rules
      - A general goal is NOT authorization for a specific high-blast action it implies — when in doubt, escalate.
      - Never act outside `scope`. Never touch another session's `.claude/worktrees/`.
      - Leave the repo in a clean, reviewable state every tick.
      
    • STATE.template.md 880 B
      # <loop-name> — STATE
      _Updated: <ISO-8601 Z> · run #0 · readiness 100/100_
      
      <!--
      The triage snapshot. The loop READS this at the top of every run and REWRITES it at
      the end. Not a database — a lightweight snapshot of: what to act on, what to watch,
      what was seen-and-dismissed. Read/write contract: references/state-spine.md.
      First action of every run: check the kill switch, then read the Priority list.
      -->
      
      ## Priority   (act on these next)
      <!-- the next units of work, highest first. e.g. "[P1] PR #412 failing CI 3h" -->
      - (none yet)
      
      ## Watch      (not yet actionable)
      <!-- things being tracked that aren't ready to action -->
      - (none yet)
      
      ## Noise      (seen + dismissed this run)
      <!-- items deliberately skipped, so the next run doesn't re-surface them -->
      - (none yet)
      
      ---
      _Source: <scheduler, e.g. .github/workflows/<loop-name>.yml> · config: loop.config.yaml_
      
  • references
    • claude-code-loops.md 17.2 KB
      # Where Loops Actually Live in Claude Code
      
      The outer loop is a *cadence + a headless run*. This file is the mechanics: the concrete
      ways to fire a loop in Claude Code, when to use each, and how they compose with the tier
      model. The doctrine — *a scheduler invokes `claude -p`, not a session that spawns ungated
      children* — is in [risk-tiers.md](risk-tiers.md); this is the how. The primitives
      themselves — every parameter, limit and failure semantic, verified and date-stamped — are
      in [native-scheduling.md](native-scheduling.md); read that before trusting a number here.
      
      ---
      
      A loop's **trigger** answers *when a tick fires* — a **cadence** (poll on a clock) or an
      **event** (something pushed in from outside) — and its **completion** rule answers *when
      the work stops*. Claude Code has native answers to all three. **Prefer the native
      mechanisms — zero/low-infra, no GitHub Actions.** Reach for an external scheduler only for
      non-Claude-Code control.
      
      ## Cadence — when a tick fires
      
      | Mechanism | `host:` | Runs on | Local files? | Open session? | Min interval | Best for |
      |---|---|---|---|---|---|---|
      | **`/loop`** (bundled skill) + `CronCreate` | `session-cron` | your machine | ✅ | **yes, idle** | 1 min | supervised, in-session polling (**L1 only** — 7-day expiry) |
      | **`ScheduleWakeup`** — `/loop`'s dynamic mode | `session-cron` | your machine | ✅ | yes | 60 s–1 h clamp | self-pacing one task; Claude picks each delay |
      | **Desktop scheduled task** (`scheduled-tasks` MCP) | `desktop-task` | your machine | ✅ | no (app open) | 1 min | **the local-first unattended default** — loops that touch the repo/build/tools |
      | **Cloud routine** (`/schedule` → [Routines](https://code.claude.com/docs/en/routines)) | `cloud-routine` | **Anthropic cloud** | ❌ **fresh clone** | no | **1 hour** | unattended loops needing **no** local state (GitHub PRs, web, connectors) |
      | external scheduler + `loop-run.sh` | `external` | your machine | ✅ | no | your call | non-Claude-Code control: cron / Task Scheduler / systemd / process-compose / CI |
      | **GitHub Actions** | `external` | GH runner | fresh clone | no | — | *optional* — only if the repo already lives on GitHub |
      
      Declare the choice as `host:` in `loop.config.yaml`; `loop-doctor` then enforces that
      host's real constraints instead of assuming a local `claude -p`.
      
      > **Three load-bearing caveats, all verified 2026-08-30:**
      >
      > 1. **Cloud routines run on a fresh clone with no access to your local files.** A loop
      >    that touches a local repo, build, model dir, or tool **cannot** be a cloud routine —
      >    use a Desktop scheduled task or `/loop`. They also have **no permission mode at all**:
      >    the boundary is repos + environment network policy + connectors, and *every* connected
      >    connector attaches by default.
      > 2. **`/loop` and `CronCreate` are session-scoped and expire.** In-memory, gone on a new
      >    conversation (`--resume` restores unexpired ones), fire only while the session is
      >    **idle**, and every recurring job **self-deletes 7 days after creation**. That makes
      >    `session-cron` an L1-supervised host — never the home of an unattended loop.
      > 3. **A Desktop task's worktree toggle is OFF by default**, so a run works against your
      >    working directory *including uncommitted changes*. At L2+ turn it on, or the loop's
      >    "isolation" is imaginary.
      
      The unattended options (Desktop task, cloud routine, external scheduler, Actions) are the
      human-configured **authorizer** — no parent auto-mode session, so nothing blocks the
      headless child. Many loop frameworks are CI/Actions-centric; loop-ops is
      runner-agnostic and **native-first** on purpose.
      
      ## Event — when something happens (routine triggers, Channels)
      
      Polling burns tokens while nothing changes and lags the thing it watches. There are now
      **three** ways to fire on an event instead of a timer, and they differ in whether a session
      must stay alive:
      
      | Event source | Needs a live session? | Fires |
      |---|---|---|
      | **Routine API trigger** — `POST /fire` + bearer token | **no** | your alerting system, deploy pipeline or internal tool starts a cloud run |
      | **Routine GitHub trigger** — `pull_request` / `release` + filters | **no** | a repo event starts a cloud run |
      | **Channel** — an MCP plugin pushing into a session | **yes** | anything you can build a receiver for |
      
      The routine triggers are the important addition: **a native event loop no longer has to be
      a kept-alive background session.** An alert-triage or deploy-verification loop is an API
      trigger; a PR-review loop is a GitHub trigger with filters. Both carry the cloud-routine
      constraints above. `text` sent to `/fire` arrives wrapped in a `<routine-fire-payload>`
      block **labelled untrusted** — the prompt must explicitly opt in to acting on it, which is
      what stops a leaked bearer token from becoming instruction injection.
      
      A [**Channel**](https://code.claude.com/docs/en/channels) (v2.1.80+, research preview) is an
      MCP plugin that **pushes** an external event — a CI failure, an error-tracker alert, a
      deploy webhook, a chat message — straight into a running session, so the tick fires *on the
      event* instead of on a timer.
      
      - **Cheaper + faster than polling** — no idle ticks; the loop reacts the instant the event
        lands. The right trigger for `ci-watch`, `pr-watch`, `monitor`.
      - **The trade-off:** an event arrives only while a session is open, so an unattended
        event-loop is a **persistent background session** (`claude --channels plugin:<name> …`,
        or `-p` for non-interactive) kept alive — not a fully-detached cron. Detachment traded
        for responsiveness.
      - **Setup:** install a channel plugin (Telegram/Discord/iMessage ship in the preview; build
        a [webhook receiver](https://code.claude.com/docs/en/channels-reference) for CI/error/
        deploy), launch with `--channels`, lock the sender allowlist. Anthropic-auth only (not
        Bedrock/Vertex/Foundry).
      - **Still gated** — an event-driven tick runs under the same permission mode + allowlist as
        any other; a webhook firing the loop never widens what it may do.
      
      ## Completion — when the work stops: `/goal`
      
      [`/goal <condition>`](https://code.claude.com/docs/en/goal) (v2.1.139+) keeps the session
      working turn-after-turn until a small fast model confirms the condition holds, then
      auto-clears — the **native inner-loop gate**. It's the native expression of a loop's
      `verify`/Until rule: *"keep going until the acceptance criteria hold."* Bound it with
      `or stop after N turns`. It's a session-scoped **prompt-based Stop hook**, and it pairs
      with auto mode (auto removes per-*tool* prompts; `/goal` removes per-*turn* prompts).
      Headless, one tick to completion:
      
      ```bash
      claude -p "/goal all tests in test/auth pass and lint is clean, or stop after 20 turns"
      ```
      
      `/loop`'s **dynamic mode** carries its own completion rule: Claude ends the loop itself by
      calling `ScheduleWakeup` with `stop: true` once the task is done, and an iteration that
      neither reschedules nor stops gets one ~20-minute fallback wakeup before the loop ends.
      That is a *self-judged* stop, so it is weaker than `/goal`'s explicit condition — use it
      for exploratory watching, not as a loop's `verify` gate.
      
      **The fully-native, zero-external-infra loop** = a **Desktop scheduled task** (local, has
      files, no open session) that runs `claude -p "/goal <tick condition>"` against the STATE
      spine. No cron, no Task Scheduler, no Actions.
      
      ---
      
      ## Which mechanism? — the recipe selector
      
      These mechanisms are **not interchangeable** — each has a load-bearing trade-off. Pick by
      answering: does it need **local code**, is it **connector-driven**, is it **recurring** or
      **run-to-completion**, and **does token cost matter**?
      
      | Your situation | Prescribed recipe | The trade-off that decides it |
      |---|---|---|
      | **Connector work, no local code** — triage email, Asana, Slack, calendar, issues via your claude.ai connectors | **Cloud routine** (`/schedule`) | Runs unattended in the cloud and **keeps all your claude.ai connectors** — email/Asana/tools work with your machine *off*. The fresh-clone/no-local-files limit doesn't bite because the work isn't in your repo. (≥1-hour cadence.) |
      | **Touches local code / build / tools**, unattended | **Desktop scheduled task**, or a **background daemon** running `claude -p` | Both have local files and need no open session. The daemon adds fresh context per tick + deterministic, tunable cost (next row). |
      | **Sustained / heavy cadence where tokens matter** | a **deterministic daemon** (or cron) firing `claude -p` — **not** `/loop` | `/loop` runs in one *growing* session: context accumulates, tokens climb, quality drifts past ~150k. A daemon fires a **fresh** `claude -p` each tick — bounded cost, no drift — and is deterministic. **Wake it inside the cache TTL you're paying for** so the static `run.md`+system prefix stays warm and each tick reads it at ~0.1×: ~240–270 s for the default 5-minute TTL, or up to ~55 min if the prefix is written with `"ttl": "1h"`. Fresh context *and* cache reads — the cheap sustained-loop recipe. |
      | **Supervised, light, you're watching** | **`/loop`** | Quickest to start, in-session — perfect for a short burst ("watch this deploy"). But it's **token-hungry if left running heavy**; graduate to a daemon for anything sustained. |
      | **Long task with a fixed, verifiable end state** — "migrate until tests pass", "split until each file < N lines", "drain the labeled backlog" | **`/goal`** (+ auto mode) | Runs turn-after-turn until a fast model confirms the criteria, then stops — a *completion gate*, not a cadence. Auto mode makes each turn unattended; bound with `or stop after N turns`. |
      
      **Cadence × completion compose.** A recurring loop whose every tick should run *to
      completion* = a cadence mechanism driving `claude -p "/goal <tick condition>"`. E.g. a
      Desktop task (or daemon) every morning running `/goal` over the issue backlog.
      
      ### The economics (why the daemon beats `/loop` at scale)
      
      Cadence is the top cost lever, **caching is the next** ([state-spine.md](state-spine.md),
      [loop-estimate](../scripts/loop-estimate.py)). The two interact:
      
      - **`/loop`** keeps one session alive; its input grows every iteration (accumulating
        transcript), so cost climbs and the cache helps less. Great for short supervised runs.
      - **A daemon/cron `claude -p`** starts fresh each tick (the Ralph property → flat per-tick
        cost) and, fired **inside the cache TTL**, keeps the static prefix warm (~0.1× reads).
        **The TTL is a choice, not a constant:** 5 minutes by default (1.25× write), or 1 hour
        with `"ttl": "1h"` (2× write) — so the practical daemon window is ~4.5 min *or* ~55 min.
        `loop-estimate` picks the cheapest TTL that stays warm at your cadence and says which;
        past 1 h nothing caches. Break-even and the multipliers:
        [claude-api-ops caching-and-cost](../../claude-api-ops/references/caching-and-cost.md).
        The same reasoning is why an in-session `/loop` gains less: its input *grows*, so the
        cached prefix is a shrinking share of each tick — the fresh-context daemon is what keeps
        the cacheable part dominant.
      
      A minimal local daemon (no scheduler infra) — wake under the cache window, fresh context each tick:
      
      ```bash
      # fires loop-run.sh every ~4.5 min: fresh `claude -p`, prefix stays cache-warm (5m TTL)
      while true; do .loops/<name>/loop-run.sh; sleep 270; done
      # with a 1h-TTL cache write on the prefix, ~55 min still reads warm: sleep 3300
      # or run it under process-compose / a systemd timer / nohup for boot persistence
      ```
      
      ---
      
      ## The external-scheduler shape (when you're not using a native mechanism)
      
      Native paths (Desktop task, cloud routine, `/loop`) run the tick prompt — or
      `claude -p "/goal …"` — **directly**, so they need no wrapper. When you instead drive the
      loop from an **external** scheduler (cron / Task Scheduler / systemd / process-compose /
      CI — e.g. for sub-minute cadence or to fit existing infra), `loop-scaffold` scaffolds a
      **`loop-run.sh`** in the loop dir as the runner-agnostic glue. No GitHub Actions required.
      
      ```
         any scheduler ──▶ .loops/<name>/loop-run.sh
         (the authorizer)      ├─ kill switch first (PAUSED sentinel) → exit if set
                               ├─ claude -p "$(cat run.md)" --permission-mode dontAsk \
                               │     --append-system-prompt "$(cat STATE.md)" --allowedTools …
                               └─ git add/commit STATE.md + run-log.md (if in a repo)
      ```
      
      Wire it with whatever you already run — **no cloud dependency**:
      
      ```bash
      # cron (Linux/macOS):
      */10 * * * *  /path/.loops/pr-watch/loop-run.sh >> /path/.loops/pr-watch/tick.log 2>&1
      
      # Windows Task Scheduler (every 10 min; S4U logon, see windows-ops for the hardened form):
      schtasks /Create /SC MINUTE /MO 10 /TN pr-watch \
        /TR "bash -lc '/c/path/.loops/pr-watch/loop-run.sh'"
      
      # process-compose / systemd timer / a while-sleep loop — all work; loop-run.sh is just a script.
      ```
      
      - The **scheduler** (not a Claude session) invokes `loop-run.sh`. It is the
        human-configured authorizer; nothing upstream gates the run.
      - `--permission-mode dontAsk` + a curated allowlist = a **gated** worker that runs
        anywhere. (For L3 arbitrary-execution jobs, swap to a container + `bypassPermissions` —
        see the enumerate-vs-isolate fork in [risk-tiers.md](risk-tiers.md).)
      - The run prompt (`run.md`) is the same every tick — fresh context each time (the Ralph
        property). State survives in `STATE.md` + the codebase + git, not the conversation.
      - **GitHub Actions** is one option, not a requirement — the worked example ships an
        optional `github-actions.yml` for repos already on GitHub; everyone else uses the local
        schedulers above.
      
      ### Why not "a Claude session that launches the loop"?
      
      Because an `auto`-mode session that spawns a detached `claude -p --permission-mode
      bypassPermissions` child is blocked as **Create Unsafe Agents** — an ungated autonomous
      agent with no human gate. The fix is structural, not a workaround: move the launch to the
      scheduler. Trying to wrap the bypass flag in a script to dodge the gate is **Auto-Mode
      Bypass**, a `hard_deny` (see [risk-tiers.md](risk-tiers.md) and the
      [classifier reference](../../../docs/AUTO-MODE-CLASSIFIER.md)).
      
      ---
      
      ## Hooks — the loop's reflexes
      
      Hooks fire shell commands at points in the agent's lifecycle. Useful loop wiring:
      
      | Hook | Loop use |
      |---|---|
      | `PreToolUse` | enforce scope/kill-switch before a tool runs (deterministic gate 1) |
      | `PermissionDenied` | react to a classifier denial — log it, signal a retry, escalate |
      | `Stop` | write the run-log line + rewrite `STATE.md` as the run ends |
      | `SessionStart` | load `STATE.md` into context at the top of a run |
      
      A `PreToolUse` hook that checks `.loops/<name>/PAUSED` is the cheapest possible kill
      switch — it blocks every tool the instant the sentinel appears, no matter where the run
      is. See [`claude-code-ops`](../../claude-code-ops/SKILL.md) for the full 30-event hook
      catalog and the stdin/stdout JSON contracts.
      
      ---
      
      ## Composing with the execution layers
      
      The cadence fires; the work is done by the layers this repo already ships:
      
      ```
      /schedule (cadence)
         └─▶ claude -p  (the run; dontAsk + allowlist)
               ├─▶ iterate          # inner improvement loop, if the unit of work is "improve metric X"
               ├─▶ fleet-worker     # spawn cheap parallel makers in worktrees
               └─▶ fleet-ops        # test-gate + land the winning branch
         └─▶ Stop hook → rewrite STATE.md + append run-log
      ```
      
      - **`iterate`** when the unit of work is "drive metric X to target in this session".
      - **`fleet-worker`** when one tick should fan out several maker attempts cheaply.
      - **`fleet-ops`** as the `land_via` — the sequential, test-gated merge queue that turns a
        worker's green branch into a landed change (or escalates it).
      - **`pigeon`** to coordinate across concurrent loops (the priority-order standoff).
      
      ---
      
      ## A worked L1 → L2 graduation
      
      1. **L1, supervised:** `/loop 15m` in a session (`host: session-cron`), running a
         read-only "report PR state to STATE.md" prompt. You watch it; it writes nothing but the
         snapshot. Permission mode `plan`. Remember the 7-day expiry — this host is for the
         proving period, not the destination.
      2. **Prove judgment:** read a week of `STATE.md` snapshots + the run-log. Is its triage
         right? Does readiness hold?
      3. **L2, unattended:** move the host — `desktop-task` if the loop touches local code,
         `cloud-routine` if it doesn't, `external` for sub-minute cadence — and update `host:`
         so `loop-doctor` checks the right constraints. Switch the
         run prompt to "open a fix PR in a worktree" with `--permission-mode dontAsk` + a narrow
         allowlist (`Bash(npm test)`, `Bash(git …)`). Add a `guard`, set `land_via: fleet-ops`,
         write the `escalation` rule. Re-run `loop-check` at L2 — fix every error — then enable.
      
      The point of the ladder: the cadence mechanism *changes* (session `/loop` → scheduled
      `claude -p`) exactly when the autonomy does, and the audit gates the transition.
      
      ## See also
      
      - [native-scheduling.md](native-scheduling.md) — the primitives themselves: verified parameters, limits and failure semantics per host.
      - [risk-tiers.md](risk-tiers.md) — the permission-mode mapping + scheduler-not-session rule.
      - [state-spine.md](state-spine.md) — the STATE.md the run reads and rewrites.
      - [../../claude-code-ops/SKILL.md](../../claude-code-ops/SKILL.md) — the full hook catalog, `claude -p` flags, headless reference.
      
    • failure-modes.md 10.9 KB
      # Failure Modes — how loops actually break, and what catches each
      
      Incident-shaped scar tissue. Every entry is a real way an outer loop goes wrong: the
      **symptom** you'd observe, the **mechanism** underneath, and the **catch** — the
      specific `loop-ops` control (or Claude Code gate) that prevents or surfaces it. Read this
      before you schedule anything unattended; most of these only bite once you're not watching.
      
      The meta-lesson (Addy Osmani): *"build the loop like someone who intends to stay the
      engineer."* These failures are what happens when the loop is given more autonomy than its
      judgment has earned.
      
      ---
      
      ## 1. The runaway-budget loop
      
      - **Symptom:** a day's token spend gone in an hour; the bill is 5–10× the estimate.
      - **Mechanism:** cadence too tight, or scope crept so each tick reads/does far more than
        scoped (a "report PRs" loop that started crawling diffs). Sub-agents multiply it.
      - **Catch:** set `budget_tokens` (a per-run ceiling). Estimate with `loop-estimate` *before*
        scheduling; `loop-doctor` fails the loop if `budget_tokens` < estimated tokens/run.
        Reconcile the `loop-estimate` estimate against `run-log.md` actuals periodically — a tick
        that used to cost 2k now costing 40k means scope crept.
      
      ## 2. The 3am-dead loop
      
      - **Symptom:** every tick aborts immediately; `run-log.md` shows nothing but failures.
      - **Mechanism:** the `verify`/`guard` gate command's binary isn't on PATH in the
        *scheduler's* environment (works on your laptop, absent on the CI runner), or `claude`
        itself isn't installed there. In non-interactive `-p`, a hard denial **aborts the
        session** — no human to prompt.
      - **Catch:** `loop-doctor --live` resolves the gate's leading binary and checks
        `claude`/`git` are on PATH *before* you schedule. Run it in the target environment.
      
      ## 3. The cache-cold loop
      
      - **Symptom:** cost far higher than `loop-estimate --cached` projected; `cache_read_input_tokens`
        stays 0.
      - **Mechanism:** the run prompt isn't byte-identical every tick — a `datetime.now()`, a
        per-run UUID, or unsorted JSON in the prefix invalidates the cache. Or the cadence is
        slower than the cache TTL (a 6h loop can't keep a 1h entry warm), so every tick is a
        cold write.
      - **Catch:** keep `run.md` byte-identical (the template enforces this — fresh context,
        same prompt). `loop-estimate` tells you whether the cadence can cache at all and which TTL;
        if it can't, don't pay the write multiplier — run uncached.
      
      ## 4. The force-push / push-to-main loop
      
      - **Symptom:** the loop force-pushed, pushed to `main`, or ran a production migration —
        "to fix the thing."
      - **Mechanism:** a *general* goal ("keep CI green", "clean up the repo") was taken as
        authorization for a *specific* high-blast action it merely implied.
      - **Catch:** the escalation gate — these classes (force-push, push to `main`, prod
        deploy/migration, mass delete, IAM grants, deleting pre-session files, `.claude` edits)
        are **always** escalated, declared in `escalation:`. Claude Code's auto-mode classifier
        also hard/soft-denies them independently: a general goal is *not* explicit intent.
      
      ## 5. The ungated-child spawn
      
      - **Symptom:** the orchestrator session dies with *Create Unsafe Agents* / *Auto-Mode
        Bypass*; the loop never starts.
      - **Mechanism:** a session in `auto` mode tried to launch a detached `claude -p
        --permission-mode bypassPermissions` child (an ungated autonomous agent). Wrapping the
        flag in a script to dodge the classifier is a `hard_deny` nothing clears.
      - **Catch:** the cardinal rule — **a scheduler invokes `claude -p`, not a session that
        spawns ungated children.** Move the launch to cron/Actions/Task Scheduler (the human
        authorizer), and give the child gates (`dontAsk` + allowlist), not bypass — unless it's
        in an isolated container. (`rules/loop-engineering.md` directive #2.)
      
      ## 6. The colliding loops
      
      - **Symptom:** two loops fight over the same branch/worktree; one clobbers the other's
        work; merge churn.
      - **Mechanism:** several loops run against one repo with no coordination.
      - **Catch:** the multi-loop **priority order** (CI > PR > deps > cleanup > triage) — the
        higher-priority loop wins worktree contention, lowers defer to their next tick. Each
        loop isolates in its **own** worktree; they announce what they hold via `pigeon` so a
        peer can stand off. Never touch another session's `.claude/worktrees/`.
      
      ## 7. The silent-stop loop
      
      - **Symptom:** nobody noticed the loop stopped running for a week; stale `STATE.md`.
      - **Mechanism:** the schedule quietly stopped firing — a disabled workflow, a cron typo,
        a paused runner, an expired token. Loops fail *open* into silence, not error.
      - **Catch:** treat `STATE.md`'s `_Updated_` timestamp + the `run-log.md` tail as a
        heartbeat — if the latest run is older than ~2× the cadence, the loop is down. A
        separate cheap monitor (or a `daily-scan` loop) that flags stale loop heartbeats
        closes this; the kill switch is for stopping, the heartbeat is for noticing it stopped.
      
      ## 8. The test-deleting "fix" (gate reward-hacking)
      
      - **Symptom:** CI is green again — because the loop deleted or `skip`-ped the failing
        test, not because it fixed the bug.
      - **Mechanism:** the loop optimized the literal gate (`verify` passes) rather than the
        intent. A loop, like any optimizer, games a weak metric.
      - **Catch:** make the gate hard to hack — a `guard` that runs the **full** suite +
        typecheck, a `scope` that **excludes** test files and CI config, and a human review at
        L2 (the PR gate). Never let an L2/L3 loop modify the very tests that gate it.
      
      ## 9. The unbounded-scope loop
      
      - **Symptom:** the loop edited files far outside its job.
      - **Mechanism:** `scope: "*"` (or `**`, or empty) — "may touch anything."
      - **Catch:** `loop-check` **rejects** an unbounded or placeholder scope (exit 10). Scope
        is bounded globs, always.
      
      ## 10. The no-kill-switch loop
      
      - **Symptom:** the loop is misbehaving and there's no fast way to stop it.
      - **Mechanism:** no stop signal was designed in; stopping means disabling the workflow by
        hand mid-tick.
      - **Catch:** `kill_switch` is mandatory (`loop-check` errors without one) and checked
        **first** every run. The cheapest implementation is a `PreToolUse` hook that blocks
        every tool the instant a `PAUSED` sentinel appears — an instant breaker.
      
      ## 11. The comprehension-debt loop
      
      - **Symptom:** the codebase works but no one on the team understands the changes the loop
        shipped; onboarding slows, incidents take longer.
      - **Mechanism:** an unattended loop shipped correct-but-unreviewed changes for weeks;
        comprehension debt compounded silently.
      - **Catch:** the tier ladder is the antidote — **start at L1 (report-only)** and *read
        the reports*; graduate to L2 only once you trust its judgment, and keep the human in the
        PR loop. Autonomy is earned with evidence, not granted up front. Build the loop like you
        intend to stay the engineer.
      
      ---
      
      ## 12. The expired loop (native-host)
      
      - **Symptom:** a `/loop`-scheduled watch that ran fine for a week is simply gone. No
        error, no final report — and `CronList` shows nothing.
      - **Mechanism:** `CronCreate` recurring jobs **self-delete 7 days after creation** (they
        fire one final time first), and the whole session store is in-memory: a new conversation
        clears it, and `durable` is documented as having no effect. A loop hosted on
        `session-cron` has a hard, silent lifetime ceiling. This is silent-stop (#7) with a
        cause you cannot fix by watching the scheduler.
      - **Catch:** don't host anything unattended on `session-cron`. Declare
        `host: session-cron` and `loop-doctor` refuses it at L2+; at L1 it warns. For anything
        that must outlive a week, use `desktop-task`, `cloud-routine` or `external`.
      
      ## 13. The over-connected routine (native-host)
      
      - **Symptom:** a read-only "summarise my inbox" routine sent an email / closed an issue /
        posted to Slack.
      - **Mechanism:** cloud routines have **no permission mode and no approval prompts**, and
        **every connected connector is attached by default** with full write access. Nothing
        between the prompt and the tools — the "read-only" was only ever a wording in the prompt.
        A leaked `/fire` bearer token compounds it (fire text is at least wrapped as untrusted,
        so it can't issue instructions the prompt didn't opt into).
      - **Catch:** prune connectors to the minimum on every routine; scope the environment's
        network access; select only the repos the work needs. `loop-doctor` refuses a
        `cloud-routine` config that doesn't name that boundary, precisely because there is no
        permission mode to fall back on.
      
      ## 14. The stalled task (native-host)
      
      - **Symptom:** a Desktop scheduled task shows a session open in the sidebar and no output
        for hours, or a task "runs" daily but the run history is mostly *skipped*.
      - **Mechanism:** three separate defaults. A task in Manual permission mode that needs an
        unapproved tool **stalls waiting for a human** rather than failing (as does any MCP tool
        marked `requiresUserInteraction`, on every call). Tasks fire only while the app is open
        and the machine is awake; a sleeping machine skips the run. And a machine that was
        asleep all day gets **exactly one** catch-up for the most recent missed window, so a 9am
        task can execute at 11pm.
      - **Catch:** click *Run now* once after creating a task and always-allow each tool it
        needs; then put **time guardrails in the run prompt itself** ("only review today's
        commits; if it's after 5pm, skip and summarise what was missed") — the loop cannot
        assume it is running when it was scheduled to.
      
      ## At a glance — symptom → control
      
      | Failure | Primary control |
      |---|---|
      | Runaway budget | `budget_tokens` + `loop-estimate` + `loop-doctor` budget check |
      | 3am-dead | `loop-doctor --live` (gate binary + PATH) |
      | Cache-cold | byte-identical `run.md` + `loop-estimate` TTL guidance |
      | Force-push / prod | escalation gate + auto-mode classifier |
      | Ungated-child spawn | scheduler-invokes-`claude -p` (rule #2) |
      | Colliding loops | priority order + per-loop worktree + `pigeon` |
      | Silent-stop | `STATE.md`/run-log heartbeat staleness |
      | Gate reward-hacking | full-suite `guard` + scope excludes tests + human PR gate |
      | Unbounded scope | `loop-check` rejects `*` |
      | No kill switch | mandatory `kill_switch` + PreToolUse PAUSED hook |
      | Comprehension debt | L1-first graduation; read the reports |
      | Expired loop (7-day) | never host unattended on `session-cron`; `loop-doctor` host check |
      | Over-connected routine | prune connectors + scope environment; `loop-doctor` boundary check |
      | Stalled / skipped task | always-allow the tools once; time guardrails in `run.md` |
      
      ## See also
      
      - [native-scheduling.md](native-scheduling.md) — the per-host limits behind #12–#14.
      - [risk-tiers.md](risk-tiers.md) — the graduated-autonomy ladder behind #11.
      - [state-spine.md](state-spine.md) — budget + heartbeat + multi-loop coordination.
      - [../../../rules/loop-engineering.md](../../../rules/loop-engineering.md) — the directives that prevent #4/#5/#9/#10.
      
    • native-scheduling.md 17.1 KB
      # Native Scheduling Primitives — what the harness now ships, and what it doesn't
      
      **Verified 2026-08-30** against the live tool schemas in-session (`CronCreate`/`CronList`/
      `CronDelete`, the `scheduled-tasks` MCP server, `ScheduleWakeup`) and the current docs:
      [scheduled-tasks](https://code.claude.com/docs/en/scheduled-tasks),
      [desktop-scheduled-tasks](https://code.claude.com/docs/en/desktop-scheduled-tasks),
      [routines](https://code.claude.com/docs/en/routines). These are fast-moving surfaces —
      **re-verify before trusting a number here.** Where the docs and a tool description
      disagree, this file says so rather than picking a winner.
      
      [claude-code-loops.md](claude-code-loops.md) owns *which mechanism to pick and how to wire
      it*. This file owns *what each primitive actually is*: parameters, limits, and the
      failure semantics that decide whether a loop survives a night.
      
      ---
      
      ## The four hosts
      
      A loop's **host** is where its ticks execute. It is not a style preference — it changes
      what the loop can reach, what can stop it, and what "it didn't run" means.
      
      | Host | Primitive | Executes on | Local files | Needs a session | Min cadence | Survives restart |
      |---|---|---|---|---|---|---|
      | `session-cron` | `CronCreate` / `/loop` | your machine | ✅ | **yes, open + idle** | 1 min | only via `--resume`, if unexpired |
      | `desktop-task` | `scheduled-tasks` MCP / Routines→**Local** | your machine | ✅ | no (app open) | 1 min | ✅ on disk |
      | `cloud-routine` | `/schedule` → Routines→**Cloud** | Anthropic cloud | ❌ fresh clone | no | **1 hour** | ✅ |
      | `external` | cron / Task Scheduler / systemd → `loop-run.sh` | your machine | ✅ | no | yours | ✅ |
      | `local` | *undeclared* — a generic local run | your machine | ✅ | — | — | — |
      
      `local` is the template default and means "not yet decided": `loop-doctor` applies only
      the host-agnostic checks. It is fine while scaffolding and while a loop is still L1, but
      **pick a real host before scheduling** — a loop whose execution surface nobody named is a
      loop whose limits nobody checked.
      
      The doctrine is unchanged and now has a native shape: **the authorizer is the scheduler,
      never a session that spawns ungated children** ([risk-tiers.md](risk-tiers.md)). Every host
      above except `session-cron` is a human-configured authorizer.
      
      ---
      
      ## `session-cron` — `CronCreate` / `CronList` / `CronDelete`, and `/loop`
      
      The in-session scheduler. `/loop` is a **bundled skill** that drives these tools; you do
      not author it.
      
      **Verified surface.** `CronCreate` takes a standard 5-field cron expression evaluated in
      **local time** (`minute hour day-of-month month day-of-week`), a `prompt`, and
      `recurring` (default `true`; `false` = fire once then auto-delete). It returns an
      8-character job ID for `CronDelete`. `CronList` lists them. Wildcards, steps, ranges and
      lists are supported; extended syntax (`L`, `W`, `?`, `MON`/`JAN` aliases) is not. When
      day-of-month and day-of-week are both constrained, a date matches if **either** does
      (vixie-cron semantics).
      
      **The limits that decide whether you may use it for a real loop:**
      
      - **Session-scoped and in-memory.** The `durable` parameter is present but documented as
        having **no effect** — "durable persistence is not available". A new conversation clears
        every task; `--resume` / `--continue` restores unexpired ones.
        *One caveat worth knowing before you trust that flatly:* the docs describe a narrow
        case where a task you ask to keep across sessions **is** written to the project's
        `.claude` directory — when feature-flag fetching is off (and it errors if that path is a
        symlink). So "no effect" is what the tool reports in a normal session, not a universal
        law. Either way it is not a foundation for an unattended loop; don't design around it.
      - **Seven-day expiry.** A recurring task fires one final time 7 days after creation, then
        deletes itself. This is a hard ceiling on unattended lifetime.
      - **Fires only while the REPL is idle.** Not mid-response. If Claude is busy when a task
        comes due, it waits for the turn to end.
      - **No catch-up.** A missed window fires **once** on return to idle, never once per
        missed interval.
      - **Jitter, and the sources disagree.** The docs say recurring tasks fire up to **30 min**
        late (or up to half the interval for sub-hourly jobs); the in-session tool description
        says up to **10% of the period, max 15 min**. Both agree one-shots at `:00`/`:30` can
        fire up to 90 s *early*, and that the offset is derived from the task ID (so it is
        stable per task). **Do not build a loop that depends on exact fire times.** Picking a
        minute that is not `:00`/`:30` avoids the one-shot jitter and spreads API load.
      - **50 tasks per session.**
      - `CLAUDE_CODE_DISABLE_CRON=1` disables the scheduler entirely — cron tools *and* `/loop`.
      
      **Verdict for loop-ops:** `session-cron` is an **L1 supervised** host only. It cannot host
      an unattended L2/L3 loop: it needs an open idle session and evaporates after 7 days.
      `loop-doctor` treats `host: session-cron` at L2+ as a predicted runtime failure.
      
      ### `/loop` — the bundled skill, and its dynamic mode
      
      `/loop` is not something you author; it ships with the harness and reads its argument in
      three shapes:
      
      | You type | Behaviour |
      |---|---|
      | `/loop 5m <prompt>` | fixed cron cadence. `s`/`m`/`h`/`d`; seconds round up to a minute; awkward steps (`7m`, `90m`) round to the nearest clean cron step and Claude says what it picked |
      | `/loop <prompt>` | **dynamic (self-paced)** — Claude picks each delay itself |
      | `/loop` (bare) | the built-in maintenance prompt, self-paced |
      
      A skill can be the prompt (`/loop 20m /review-pr 1234`), but a scheduled fire only runs
      skills Claude may invoke on its own — built-ins (`/permissions`, `/model`), skills marked
      `disable-model-invocation: true`, skill deny-rules and MCP prompts arrive as plain text
      instead of executing. **A loop whose tick is a slash command must check that the command
      is model-invocable**, or every tick silently no-ops.
      
      **Dynamic mode** is driven by the `ScheduleWakeup` tool. Each iteration Claude calls it
      with a `delaySeconds` **clamped to [60, 3600]**, a `reason` shown back to the user, and a
      `noop` flag (`true` = nothing changed; consecutive noop ticks collapse in the transcript).
      `stop: true` ends the loop immediately. If an iteration neither reschedules nor stops, one
      fallback wakeup fires ~20 minutes later and the loop ends if that one doesn't reschedule
      either. `Esc` clears a pending self-paced wakeup.
      
      **The two sentinels.** An autonomous `/loop` with no user prompt passes a literal sentinel
      back as its `prompt` so the runtime can re-resolve the instructions at fire time. They are
      **not interchangeable**:
      
      - `<<autonomous-loop-dynamic>>` — for `ScheduleWakeup` (self-paced mode)
      - `<<autonomous-loop>>` — for the `CronCreate` fixed-cadence mode
      
      Passing the wrong one wires an autonomous loop to the wrong pacing engine.
      
      **Customising the default.** `.claude/loop.md` (project, wins) or `~/.claude/loop.md`
      (user) replaces the built-in maintenance prompt for a bare `/loop`. It is ignored whenever
      you supply a prompt. Edits take effect on the next iteration — you can refine a running
      loop's instructions in place. Content past **25,000 bytes is truncated**.
      
      **Polling vs pushing.** The docs are explicit that where the `Monitor` tool is available,
      streaming a background script's output beats re-running a prompt on an interval — cheaper
      and more responsive. Reach for a cadence only when there is nothing to stream.
      
      ---
      
      ## `desktop-task` — the `scheduled-tasks` MCP server
      
      The durable local host, and the closest native analogue to a loop-ops loop.
      
      **Verified surface.** Four tools: `create_scheduled_task` (`taskId` kebab-case, `prompt`,
      `description`, plus **at most one** of `cronExpression` (recurring, local time) or
      `fireAt` (ISO-8601 with offset, one-time, auto-disables after firing) — omit both for an
      ad-hoc task that only runs manually; `notifyOnCompletion` defaults true),
      `list_scheduled_tasks` (returns `taskId`, schedule, `enabled`, `nextRunAt`, `lastRunAt`
      and a `path` to the task's `SKILL.md`), `update_scheduled_task` (partial; `enabled: false`
      pauses), `delete_scheduled_task` (leaves the `SKILL.md` on disk so the prompt is
      recoverable).
      
      **On-disk shape.** Each task is `<config-dir>/scheduled-tasks/<task-id>/SKILL.md` —
      YAML frontmatter carrying `name` and `description`, body = the prompt. Schedule, folder,
      model and enabled state live **outside** that file (edit them through the app or by
      asking). The config dir is `~/.claude` unless `CLAUDE_CONFIG_DIR` overrides it.
      
      **What this gives a loop for free — and the caveats:**
      
      | Native feature | What it replaces | The caveat that still bites |
      |---|---|---|
      | The task folder | a place for `STATE.md` / `run-log.md` | nothing is written for you; the *prompt* must read and rewrite them |
      | Fresh session per run | the Ralph property | **no memory of the creating conversation** — the prompt must be fully self-contained |
      | Per-task permission mode + saved always-allow approvals | `--permission-mode` on a wrapper | a task in Manual mode that hits an unapproved tool **stalls** until you answer — the classic 3am-dead loop, natively |
      | Worktree toggle | `worktree: true` + manual setup | **off by default** — a task runs against your working dir *including uncommitted changes* |
      | Active/Paused status toggle | the kill switch | pausing is out-of-band; an in-prompt sentinel check still stops a run *mid-tick* |
      | Run history incl. skipped runs + reasons | part of the run-log | it records *that* a run happened, not what the loop decided |
      
      **Failure semantics you must design around:**
      
      - **Only runs while the app is open and the machine is awake.** Sleep through a window and
        the run is skipped.
      - **Exactly one catch-up.** On launch or wake, Desktop starts one catch-up run for the
        *most recently* missed time within the last 7 days and discards everything older. A
        daily task that missed six days runs **once**. The docs' own advice is the right advice:
        put time guardrails in the prompt ("only review today's commits; if it's after 5pm, skip
        and post a summary of what was missed").
      - **Deterministic stagger** of a few minutes after the scheduled time.
      - MCP tools marked `requiresUserInteraction` prompt every call and stall the run each time.
      - A task can call `update_scheduled_task` on **itself** to change its own schedule or
        prompt. That is genuinely useful (reschedule earlier when a release branch appears) and
        it is also **self-modification** — put it on the escalation list unless the loop's stated
        purpose is adaptive cadence.
      
      ---
      
      ## `cloud-routine` — Routines (`/schedule`)
      
      Research preview. Runs on Anthropic-managed cloud infrastructure, so it keeps working with
      the machine off — at the cost of the local filesystem.
      
      **Triggers are no longer cadence-only.** A routine may carry any combination of:
      
      - **Schedule** — presets (hourly/daily/weekdays/weekly) or a cron set via `/schedule
        update`; **minimum interval one hour, faster expressions are rejected**. Also one-off
        runs at a timestamp, which auto-disable after firing and do **not** count against the
        daily run cap.
      - **API** — a per-routine `/fire` endpoint. `POST` with a bearer token starts a run and
        returns a session URL. An optional `text` field carries run-specific context.
      - **GitHub** — `pull_request` and `release` events, with filters (author, title, body,
        base/head branch, labels, draft, merged) combined by equals / contains / starts-with /
        one-of / regex. `matches regex` tests the **whole** field: use `.*hotfix.*`, not
        `hotfix`.
      
      The API trigger matters for loop-ops: it is a **native event trigger that needs no
      persistent session**, which is the thing [Channels](claude-code-loops.md) could not offer.
      An `event`-triggered `monitor` or `ci-watch` loop no longer has to be a kept-alive
      session — it can be an alerting system POSTing to `/fire`.
      
      **Two security properties worth encoding in the loop's design:**
      
      - **Fire text is untrusted by construction.** The `text` payload arrives wrapped in a
        `<routine-fire-payload>` block labelled as untrusted data. A routine's prompt must
        *opt in* by referencing the payload explicitly, or the text is inert context. Anyone
        holding the bearer token can send it, so this wrapper is the control that keeps a leaked
        token from becoming instruction injection. Treat it exactly as
        [prompt-injection-defense](../../prompt-injection-defense/SKILL.md) treats any ingested
        content. The token is shown **once** at generation; rotate via Regenerate/Revoke.
      - **A native escalation gate on pushes.** Claude pushes to `claude/`-prefixed branches
        freely; a push to any other branch is **rejected** if the branch is protected, someone
        else has an open PR from it, or it carries commits authored by someone else. That is
        close to loop-ops' "never push to main" rule, enforced by the platform.
      
      **No permission mode at all.** Routines "run autonomously as full Claude Code cloud
      sessions: there is no permission-mode picker and no approval prompts during a run." The
      boundary therefore *cannot* come from a permission mode — it comes from three other
      places, and scoping them is the entire safety story:
      
      1. **Repositories** selected (each cloned fresh from its default branch)
      2. **Environment** network policy — the Default environment is *Trusted*, allowing only
         the default allowlist; off-list hosts fail `403 x-deny-reason: host_not_allowed`
      3. **Connectors** — **all connected connectors are attached by default**, and Claude may
         use every tool from an included connector, writes included, without asking. Remove
         everything the routine does not need. This is the single most common over-grant.
      
      **Other limits:** research preview (surface may change); the `/fire` endpoint ships behind
      a dated beta header; GitHub webhook events have per-routine and per-account hourly caps
      and events beyond them are **dropped**; there is a daily cap on runs started per account
      (one-off runs exempt); routines belong to an individual account and act as that identity;
      Team/Enterprise Owners can disable them org-wide.
      
      **And the trap that looks like success:** a green status in the run list "means the session
      started and exited without an infrastructure error. It does not mean the task in your
      prompt succeeded." A loop that grades itself on run status is grading the wrong thing —
      which is exactly why the loop's own `verify` gate stays load-bearing.
      
      ---
      
      ## What the native primitives replaced — and what they did not
      
      The plumbing is theirs now. The discipline is still yours.
      
      | Loop-ops primitive | Native answer (2026-08-30) | Still yours to build |
      |---|---|---|
      | **Schedule** | ✅ all four hosts | picking the host against the constraint, not the habit |
      | **Fresh context per tick** | ✅ desktop-task, cloud-routine | writing a genuinely self-contained prompt |
      | **Isolation** | ~ desktop-task worktree toggle (**off by default**) | worktree at L2+, and verifying it is on |
      | **Kill switch** | ~ Paused toggle, `Esc`, `CronDelete` | an **in-prompt** sentinel check — the out-of-band ones can't stop a tick already running |
      | **Run log** | ~ run history / skipped-run reasons | what the loop *decided* and what it cost |
      | **State between ticks** | ❌ (the task folder is a place, not a spine) | `STATE.md` — [state-spine.md](state-spine.md) |
      | **Budget** | ❌ (a daily run cap is not a token budget) | `budget_tokens`, enforced in the run prompt |
      | **The verify gate** | ❌ (green status ≠ success) | `verify:` — and it is an eval, see below |
      | **The escalation rule** | ~ routines' branch-push guard only | the full never-auto-land list |
      | **The tier ladder** | ❌ | L1 → L2 → L3, earned |
      
      **A loop's `verify` gate is an eval.** Everything the eval discipline says about scoring —
      outcome vs step vs trajectory, `pass^k` over `pass@k` for anything non-deterministic,
      judge bias, and gates that are blocking rather than advisory — applies to the gate that
      decides land-vs-escalate. Invoke the **`evals-ops`** skill when the gate is a judgement
      call rather than a green test run; a gate you cannot trust is a loop you cannot graduate.
      
      ---
      
      ## Choosing a host — the short version
      
      ```
      Needs local files / build / tools?
        ├─ no  → cloud-routine        (machine off; ≥1h; scope repos+env+connectors, no perm mode)
        └─ yes → unattended?
                  ├─ no  → session-cron (/loop)   L1 supervised only; 7-day expiry
                  └─ yes → desktop-task           (durable, per-task perms, worktree toggle)
                           └─ need sub-minute cadence, or non-Claude-Code control?
                              → external + loop-run.sh
      ```
      
      Declare the answer as `host:` in `loop.config.yaml`. `loop-doctor` checks the loop against
      its host's real constraints — a cadence faster than the host allows, a local PATH check
      that proves nothing about a cloud run, an unattended tier on a session-scoped host.
      
      ## See also
      
      - [claude-code-loops.md](claude-code-loops.md) — which mechanism, and how to wire it.
      - [risk-tiers.md](risk-tiers.md) — L1/L2/L3 ↔ permission modes; the scheduler-not-session rule.
      - [state-spine.md](state-spine.md) — the STATE/run-log/budget spine none of these hosts provide.
      - [failure-modes.md](failure-modes.md) — the incident catalog these limits produce.
      
    • pattern-catalog.md 8.8 KB
      # Pattern Catalog — a morphology of loop shapes
      
      Loops aren't a fixed list of recipes — they're **compositions of three orthogonal axes**.
      Name the axes and the patterns fall out; you can also compose ones not named here. The
      named patterns below are the well-trodden *points* in this space. `loop-scaffold` seeds a
      `loop.config.yaml` keyed by `--pattern <name>` (the canonical keys); the rest of the space
      you compose by hand. **Start every pattern at L1** and graduate only once its reports prove
      its judgment.
      
      ## The three axes
      
      **1. Trigger — what starts a tick.**
      
      | Trigger | Fires when | Mechanism | Best for |
      |---|---|---|---|
      | `cadence` | a clock interval elapses | `/loop` (supervised), Desktop task, cloud routine, or a daemon | steady polling — backlog, PRs, deps |
      | `event` | an external thing happens (CI fail, error, deploy, message) | a cloud routine's **API `/fire`** or **GitHub** trigger (no session needed), or a **Channel** (MCP receiver) pushing into a live session | responsiveness + low cost — no idle polling |
      | `goal` | runs continuously **until a condition holds**, then stops | `/goal` (+ auto mode) | run-to-completion — migrations, metric targets |
      
      > **Event beats poll when you can get it.** A CI webhook firing the tick is cheaper and
      > faster than a 10-min poll — and it no longer has to cost you detachment. A **routine API
      > trigger** (`POST /fire` with a bearer token) or a **GitHub trigger** starts a cloud run
      > with no session alive at all; only a **Channel** needs a persistent session
      > (`claude --channels …` in a background process, research-preview, Anthropic-auth only).
      > Pick the routine trigger when the work can run in the cloud, the Channel when it must
      > touch local state. See [claude-code-loops.md](claude-code-loops.md).
      
      **2. Posture — how much autonomy** (the [risk tier](risk-tiers.md)): `L1` report · `L2`
      propose-and-human-gates · `L3` autonomous-in-a-denylist.
      
      **3. Locus — where it runs / what it can touch.**
      
      | Locus | Mechanism | Can touch | Use when |
      |---|---|---|---|
      | `connector` | **cloud routine** (`/schedule`) | your claude.ai connectors (email, Asana, Slack, issues) — **no local files** | the work lives in services, not your repo |
      | `local` | Desktop task / daemon / `/loop` | the repo, build, models, local tools | the work touches local state |
      
      **Locus is `host:`.** The axis is not decorative — write the resolved answer into the
      config's `host:` field (`cloud-routine` for `connector`; `desktop-task`, `session-cron` or
      `external` for `local`) so `loop-doctor` enforces that surface's real limits instead of
      assuming a local `claude -p`. Per-host limits: [native-scheduling.md](native-scheduling.md).
      
      The recipe-selector in [claude-code-loops.md](claude-code-loops.md) is just these axes
      resolved to a mechanism. A loop = **(trigger × posture × locus) + the [state spine](state-spine.md)**.
      
      ---
      
      ## The catalog
      
      Each row: the axes, the recommended native mechanism, the job (gate → what it escalates),
      and the **failure mode to watch** ([failure-modes.md](failure-modes.md)).
      
      | Pattern | Trigger · Locus | Start tier | Mechanism | Job → escalates | Watch |
      |---|---|---|---|---|---|
      | `daily-scan` | cadence · local | L1 | Desktop task (off-peak) | sweep backlog/alerts, write `STATE.md` → all to a human | silent-stop |
      | `pr-watch` | event\|cadence · connector | L1 | cloud routine on a **GitHub `pull_request` trigger** (no session needed), or a Channel | flag stuck/failing/conflicted PRs → never merges | runaway tokens if polled tight |
      | `ci-watch` | **event** · local | L2 | Channel (CI webhook) → fix in a worktree | failing test passes + full guard → flaky/deploy/secrets | gate reward-hacking |
      | `dep-bump` | cadence · local | L2 | Desktop task/daemon | patch-only behind cooldown + guard → minor/major, advisories | supply-chain |
      | `changelog-gen` | event(on tag)\|cadence · local | L1 | tag-event or Desktop task | draft `RELEASE_NOTES_DRAFT.md` → human publishes | — |
      | `merge-hygiene` | cadence · local | L1 | Desktop task (off-peak) | dead branches / stale flags → ambiguous deletes | worktree-boundaries |
      | `issue-sort` | cadence\|event · connector | L1 | cloud routine | classify + suggest labels → priority/dupe-close | — |
      | **`metric-chase`** | **goal** · local | L2 | `/goal` driving [`iterate`](../../iterate/SKILL.md) | drive coverage/latency/bundle/**eval-score** to target → unreachable / guard fails | gate reward-hacking · **high cost** |
      | **`regression-watch`** | cadence\|event(on release) · local | L1→L2 | Desktop task/daemon | run a benchmark/eval, diff vs baseline → a real regression | flaky bench = false alarm · **high/run** |
      | **`digest`** | cadence · **connector** | L1 | **cloud routine** | summarize email/Asana/calendar/news → nothing (read-only) | over-scoped connector |
      | **`backfill`** | **goal** · local | L2/L3 | `/goal` (+ worktree/container) | drain a migration/queue **to completion** → an item needing judgment | runaway budget · **long** |
      | **`monitor`** | **event** · local | L1 | **Channel** (error/log/deploy webhook) | triage the event → page a human on anomaly | alert fatigue · needs a live session |
      | **`freshness`** | cadence · local | L1 | Desktop task (daily/weekly) | re-check docs/data/deps vs reality → confirmed drift | transient failure ≠ drift |
      
      ---
      
      ## Notes on the patterns that need them
      
      - **`ci-watch` / `pr-watch` — prefer event over poll.** A polled `pr-watch` at 5 min costs
        ~3× a 15-min one for marginal freshness; the event-driven version costs ~nothing while
        quiet. Two ways to get the event now: a cloud routine's **GitHub trigger** (`pr-watch`)
        or **API `/fire`** from your CI (`ci-watch`) — neither needs a session alive — or a
        **Channel** when the tick must touch local state. At L2, `ci-watch` opens a fix in a
        worktree and hands the branch to `fleet-ops`; never auto-merges `main`.
      - **`metric-chase` is the bridge to [`iterate`](../../iterate/SKILL.md).** The loop's
        *trigger* is a `/goal` ("coverage ≥ 90, or stop after N turns"); the *work* each turn is
        an `iterate` step (modify → measure → keep/discard). Use it for any measurable target —
        including an **eval score** (this is the GLM/Opus-bench shape). Highest cost class; bound it.
      - **`digest` is the canonical cloud-routine pattern.** It needs *connectors, not code*, so
        it's the one archetype where the fresh-clone cloud routine is exactly right — it keeps
        your claude.ai connectors and runs with the machine off. Read-only: no write scopes.
      - **`backfill` is run-to-completion, not recurring.** A `/goal` drains the queue/migration;
        when the condition holds it stops and clears. Bound it (`or stop after N`, a token budget)
        — it's the runaway-budget risk made flesh. For arbitrary execution, run it in a container.
      - **`monitor` is the purest event loop.** An error-tracker/deploy webhook → a Channel → a
        persistent background session that triages and pages on anomaly. No polling at all. The
        trade-off is keeping that session alive.
      - **`regression-watch`** runs a real suite each tick (expensive), so cadence it slowly or
        trigger it on a release event. Treat a transient/flaky failure as advisory (don't page on
        one red run) — the same exit-7-vs-exit-10 discipline our staleness verifiers use.
      
      ---
      
      ## Composing a pattern not in the catalog
      
      Pick a point in the space the named patterns don't cover. Examples:
      
      - *event · connector · L1* — a Slack message (Channel) triggers a read-only lookup against
        a connector. (A "support-triage" loop.)
      - *goal · connector · L2* — drain an Asana backlog to empty via `/goal`, updating tasks
        through the connector.
      - *cadence · local · L3* — a nightly autonomous refactor in an isolated container.
      
      The discipline is identical regardless of the point: bounded scope, a gate, an escalation
      rule, a kill switch, a budget — and **start at L1**.
      
      ## Choosing — the short version
      
      1. **Locus first:** does it touch local code? → `local` (Desktop task/daemon). Pure
         connector work? → `connector` (cloud routine).
      2. **Trigger next:** is there an event to react to? → `event` (Channel) — cheaper + faster.
         A clear finish line? → `goal`. Otherwise → `cadence`, slowest that still catches the work.
      3. **Posture:** start **L1**. Graduate to L2 (with a guard, worktree, escalation, `land_via`)
         only once the reports earn it; re-run `loop-check` + `loop-doctor --live` at the new tier.
      
      ## See also
      
      - [risk-tiers.md](risk-tiers.md) — the posture axis (permission-mode mapping).
      - [claude-code-loops.md](claude-code-loops.md) — the trigger/locus axes resolved to mechanisms + the recipe selector.
      - [failure-modes.md](failure-modes.md) — the "watch" column, in depth.
      - [state-spine.md](state-spine.md) — the multi-loop priority order these share.
      - [../assets/loop.config.template.yaml](../assets/loop.config.template.yaml) — the config every pattern fills in.
      
    • risk-tiers.md 9.4 KB
      # Risk Tiers ↔ Claude Code's permission model
      
      The single best idea in loop engineering is **graduated autonomy**: a loop earns the
      right to act unattended, it isn't granted it. This file maps the L1→L2→L3 ladder onto
      Claude Code's *actual* permission machinery — which is what makes this skill more than a
      generic-agent methodology. The authority for the gate behaviour is the repo's
      [auto-mode-classifier reference](../../../docs/AUTO-MODE-CLASSIFIER.md); read it for the
      full two-gate model. This file is the loop-specific projection.
      
      ---
      
      ## The ladder
      
      ```
      L1 Report ───────► L2 Assisted ───────► L3 Unattended
      read-only          suggest + human-gate    autonomous within a denylist
      (plan/dontAsk)     (dontAsk/auto)          (bypassPermissions, ISOLATED only)
      ```
      
      **Never skip a rung.** A fresh loop starts at L1. It graduates only after its reports
      prove its judgment over real runs. Each rung adds exactly one new power and one new
      guardrail.
      
      | | L1 Report | L2 Assisted | L3 Unattended |
      |---|---|---|---|
      | **Posture** | discovery + triage | propose changes | autonomous land |
      | **Writes?** | no — report only | yes, in a worktree | yes, allowlisted classes |
      | **Permission mode** | `plan` or `dontAsk` + read allowlist | `dontAsk` + narrow allowlist, or `auto` | `bypassPermissions` **in a container** |
      | **Required guardrails** | bounded scope, kill switch | + guard command, + worktree, + escalation | + denylist, + isolation boundary, + budget cap |
      | **Lands by** | a human reads the report | a human approves the PR (or `fleet-ops`) | the loop, inside its boundary |
      | **Blast radius** | zero (no writes) | one PR, reviewable | bounded by the denylist + container |
      
      ---
      
      ## How each tier maps to a permission mode
      
      Claude Code has six permission modes. Loops use four of them:
      
      | Mode | Behaviour | Loop tier |
      |---|---|---|
      | `plan` | read/explore only; cannot edit | L1 (strictest) |
      | `dontAsk` | auto-**denies** anything not pre-approved; read-only Bash always allowed; fully non-interactive | L1 / L2 (**recommended default for workers**) |
      | `auto` | a classifier model gates each unresolved action; "trust the direction" autonomy | L2 (long runs) |
      | `acceptEdits` | in-scope edits + common fs commands auto-approved; other Bash needs an allow rule | L2 (edit-heavy, known command set) |
      | `bypassPermissions` | no gates at all | L3 — **only** inside an isolated container/VM without internet |
      
      `default` (prompt each action) is interactive — not for unattended loops.
      `acceptEdits` is the middle option when the command set is known.
      
      ### Why `dontAsk` is the workhorse for L1/L2 workers
      
      `dontAsk` is fully non-interactive (it never prompts; it auto-denies the unknown), so it
      runs anywhere — no container required — and read-only Bash is always allowed. Pair it
      with a **narrow** `permissions.allow` list (`Bash(npm test)`, `Bash(git status)`) and you
      get a worker that can do exactly its job and nothing else. This is the safe default for
      headless loop workers.
      
      ---
      
      ## The headless-profile table (what a `claude -p` worker should use)
      
      The loop's *maker* runs are headless `claude -p` sessions. Pick the least privilege that
      still lets the job run:
      
      | Profile | Behaviour | Use for |
      |---|---|---|
      | `--permission-mode dontAsk` + curated `permissions.allow` | auto-denies anything not pre-approved; read-only Bash allowed; non-interactive | **locked-down workers (recommended default)** |
      | `--permission-mode auto` | classifier-gated; configure `autoMode.environment` for your infra. In `-p`, repeated blocks abort the session | long "trust-the-direction" runs |
      | `--permission-mode acceptEdits` + allow rules | edits + common fs auto-approved; other Bash needs an allow rule | edit-heavy tasks, known command set |
      | `--dangerously-skip-permissions` (= `bypassPermissions`) | no gates; refuses root/sudo; `rm -rf /`\|`~` still circuit-break | **only** in an isolated container/VM/devcontainer without internet |
      
      In **non-interactive `-p` mode** a hard denial **aborts the session** (there's no human
      to prompt). So an `auto`-mode worker that hits a wall dies; a `dontAsk` worker with a
      correct allowlist never hits one. This is why enumerating permissions beats relying on
      the classifier for batch workers.
      
      ---
      
      ## The cardinal rule: scheduler invokes `claude -p`, not session-spawns-loop
      
      This is the one thing a generic-agent methodology can't tell you because it isn't grounded
      in Claude Code's gate. **An unattended loop must be a scheduler/script that invokes
      `claude -p` — not a Claude session that tries to launch the loop.**
      
      Why: the auto-mode classifier evaluates tool calls *inside* an auto-mode session. A
      session that tries to spawn a detached `claude -p --permission-mode bypassPermissions`
      child is blocked as **Create Unsafe Agents** (an ungated autonomous agent with no human
      gate). Two independent fixes, combine for best result:
      
      1. **Move the launch outside the auto-mode session.** A human — or a human-configured
         Task Scheduler / cron / CI runner / plain script — running `claude -p …` is the
         authorizer, with no parent classifier in the loop. Don't run the *orchestrator*
         session itself in auto mode if its job is spawning agents.
      2. **Give the child gates instead of bypass.** The denial is about the *ungated*
         property, not headless-ness. A `dontAsk`+allowlist child is gated and runs fine.
      
      > **Subagents can't escalate.** Agent/Task subagents inherit the parent's mode; the
      > classifier uses the parent mode and ignores `permissionMode` in subagent frontmatter.
      > A full-bypass worker fleet must be the isolated-container path launched *outside* the
      > auto-mode session — never an in-session subagent.
      
      ---
      
      ## The real fork: enumerate vs isolate
      
      When a loop needs real power, there are exactly two legitimate shapes. Reaching for
      `bypassPermissions` on the host *to avoid enumerating permissions* is precisely the
      pattern the classifier blocks.
      
      | | **Enumerate** | **Isolate** |
      |---|---|---|
      | Shape | `dontAsk` + a curated allowlist | container/VM + `bypassPermissions` |
      | Runs | anywhere (host, CI, laptop) | only inside the sandbox |
      | Safety | bounded by the allowlist | bounded by the container |
      | Cost | you list the commands once | you stand up isolation |
      | Best for | most loops; CI/PR/dep workers | heavy autonomous refactors, untrusted-input runs |
      
      **Default to enumerate.** Reach for isolate only when the job genuinely needs arbitrary
      execution *and* you have a real sandbox (no internet, can't damage the host).
      
      ---
      
      ## Connector & MCP scopes — least privilege for the loop's tools
      
      A loop is only as safe as the tools it can reach. The permission *mode* gates Bash + file
      edits; the *tool surface* gates everything else — MCP connectors (Slack, GitHub, Jira, a
      DB), `WebFetch`, the `Agent` tool. Scope them per tier:
      
      - **Allowlist, don't blanket.** Headless, name exactly what the job needs:
        `--allowedTools 'Bash(gh pr list:*)' 'Bash(gh pr view:*)' 'Read' 'mcp__github__*'`. Use
        `--disallowedTools` to subtract a dangerous one (block `WebFetch` on a loop that
        shouldn't read the web; block a Slack `post_message` on a read-only triage loop).
      - **Read-scoped connectors at L1.** An L1 report loop gets read-only MCP scopes
        (list/get/search), never write (post/create/delete/merge). Scope the connector *itself*
        least-privilege — don't hand a triage loop a write-capable GitHub token "just in case".
      - **The auto-merge guard.** Never give a loop a path to merge `main`: keep `gh pr merge`
        out of the allowlist, set `land_via: fleet-ops` (test-gated, human-or-queue), and list
        main-push in `escalation`. A green PR on a feature branch is the *most* a loop
        auto-produces.
      - **MCP tool descriptions are instructions.** A poisoned connector description is
        prompt-injection straight into the loop's context — vet a connector (and prefer read
        scopes) before a loop uses it. See [`prompt-injection-defense`](../../prompt-injection-defense/SKILL.md).
      
      ## Why Claude Code-specific (not a multi-tool matrix)
      
      `loop-ops` is deliberately scoped to **Claude Code**, not a cross-tool primitives matrix
      (Grok / Codex / …). The whole edge is grounding in Claude Code's *actual* gate model — the
      permission modes, the auto-mode classifier, `claude -p`, the hook events. A generic
      multi-agent matrix would dilute exactly that. The *doctrine* ports to any agent (the tier
      ladder, the gate, the kill switch, the STATE spine, the escalation classes); the
      permission-mode **mapping** is Claude Code's, and that specificity is the point.
      
      ## Tier checklist (what `loop-check` enforces)
      
      - **L1:** bounded `scope` (never `*`), a `kill_switch`, `permission_mode` ∈ {plan,
        dontAsk}, **no** `verify` that writes. Report-only.
      - **L2:** all of L1, plus a `verify` gate **and** a `guard` (must-always-pass),
        `worktree: true`, a concrete `escalation:` rule, and a `land_via` (e.g. `fleet-ops`).
      - **L3:** all of L2, plus `permission_mode: bypassPermissions` **with** an isolation note
        in `escalation`/scope, a denylist of never-auto-land classes, and a `budget_tokens`
        cap. The audit warns hard if L3 is declared without an isolation boundary.
      
      ## See also
      
      - [../../../docs/AUTO-MODE-CLASSIFIER.md](../../../docs/AUTO-MODE-CLASSIFIER.md) — the full two-gate model, classifier categories, legitimate-authorization decision tree.
      - [claude-code-loops.md](claude-code-loops.md) — the scheduler/`claude -p` mechanics this tier model runs on.
      - [pattern-catalog.md](pattern-catalog.md) — each pattern's recommended starting tier.
      
    • state-spine.md 6.9 KB
      # The State Spine — memory outside the conversation
      
      A loop's durability comes from state that lives **outside** the conversation window. The
      conversation is ephemeral and degrades as it fills (the Ralph insight: quality drops past
      ~100–150k tokens). The spine is three files the loop reads at the start of every run and
      writes at the end. This is the loop's working memory, audit trail, and definition.
      
      ```
      .loops/<name>/
      ├── loop.config.yaml    # the definition (immutable-ish; edited by a human)
      ├── STATE.md            # the triage snapshot (rewritten every run)
      └── run-log.md          # append-only audit trail (one line per run)
      ```
      
      `loop-scaffold` scaffolds all three. The config is human-owned; `STATE.md` and `run-log.md`
      are loop-owned.
      
      ---
      
      ## `loop.config.yaml` — the definition
      
      Flat YAML so it's trivially parseable (no `yq` dependency). Full annotated template:
      [../assets/loop.config.template.yaml](../assets/loop.config.template.yaml). Fields:
      
      | Field | Required | Meaning |
      |---|---|---|
      | `name` | yes | the loop's identifier; matches the directory |
      | `pattern` | yes | a catalog key (`pr-watch`, …) or `custom` |
      | `tier` | yes | `L1` / `L2` / `L3` — the autonomy rung |
      | `cadence` | yes | `10m` / `1h` / `6h` / `1d`, or a cron string |
      | `host` | rec | where ticks execute: `local` (default) / `session-cron` / `desktop-task` / `cloud-routine` / `external`. Selects which hard limits `loop-doctor` enforces — see [native-scheduling.md](native-scheduling.md) |
      | `goal` | yes | one sentence: what this loop does and what it must NOT do |
      | `scope` | yes | bounded globs the loop may touch — **never `*`** |
      | `verify` | L2+ | the gate command (the metric/check); a loop with no gate is invalid |
      | `guard` | L2+ | a must-always-pass command (full suite / typecheck) |
      | `permission_mode` | yes | `plan` / `dontAsk` / `auto` / `acceptEdits` / `bypassPermissions`. **Ignored when `host: cloud-routine`** — routines have no permission picker; the boundary is repos + environment + connectors |
      | `worktree` | L2+ | `true` to isolate code changes in a git worktree |
      | `escalation` | yes | what the loop escalates instead of doing (the gate rule) |
      | `budget_tokens` | rec | per-run output-token ceiling |
      | `kill_switch` | yes | the stop signal every run checks first |
      | `land_via` | L2+ | who gates + lands winning branches (e.g. `fleet-ops`) |
      
      `loop-check` reads this file and scores it against the tier's requirements.
      
      ---
      
      ## `STATE.md` — the triage snapshot
      
      Rewritten at the end of every run; read at the top of the next. It is **not** a database
      — it's a lightweight snapshot of what the loop needs, what it's watching, and what it
      ignored. Template: [../assets/STATE.template.md](../assets/STATE.template.md). Shape:
      
      ```markdown
      # <loop-name> — STATE
      _Updated: 2026-06-22T14:05:00Z · run #142 · readiness 100/100_
      
      ## Priority   (act on these next)
      - [P1] PR #412 failing CI 3h — owner pinged
      - [P2] dep `axios` patch 1.14.0→1.14.1 available, cooldown clears 2026-06-25
      
      ## Watch     (not yet actionable)
      - PR #408 awaiting review 1h
      - flag `new-checkout` at 100% rollout 6d — cleanup candidate
      
      ## Noise      (seen + dismissed this run)
      - PR #410 draft — skip until ready
      - dep `left-pad` major bump — escalates, not auto
      
      ---
      _Source: .github/workflows/<loop>.yml · config: loop.config.yaml_
      ```
      
      **The read/write contract:**
      1. **Read** `STATE.md` first thing — it's the loop's memory of the last run.
      2. **Check the kill switch** (`kill_switch:` from config) — exit immediately if set.
      3. Do the run's work, drawing the next unit from the Priority list.
      4. **Rewrite** `STATE.md` — promote/demote items across Priority/Watch/Noise, bump the
         `_Updated_` line + run number + readiness.
      
      `readiness` is the loop's self-assessment (0–100): is its config still coherent, its
      gate still passing, its scope still valid? A dropping readiness is an early signal to
      re-audit.
      
      ---
      
      ## `run-log.md` — the append-only audit trail
      
      One line per run, appended, never rewritten. Answers "what has this loop been doing, and
      what did it cost?"
      
      ```
      2026-06-22T14:05:00Z  run#142  action=reported  pr=412  outcome=escalated  tokens=18420
      2026-06-22T13:55:00Z  run#141  action=none       -       outcome=quiet      tokens=2110
      2026-06-22T13:45:00Z  run#140  action=proposed   pr=409  outcome=pr-opened  tokens=44380
      ```
      
      The `tokens` column feeds back into the budget. Tail it to see drift: a loop that used to
      cost 2k/run quietly now costing 40k/run is doing more than it was scoped to.
      
      ---
      
      ## Budget control
      
      A loop's cost is `runs/day × tokens/run × price`, and sub-agents multiply tokens/run.
      Two controls:
      
      - **`budget_tokens`** in the config — a per-run output ceiling. The loop stops the run
        when it's reached (the same discipline as a dynamic `/loop` watching `budget.remaining()`).
      - **The run-log** — the actual spend, line by line. Reconcile estimate (`loop-estimate`)
        against actual periodically; if they diverge, the loop's scope crept.
      
      Estimate before you schedule: [../scripts/loop-estimate.py](../scripts/loop-estimate.py). The
      cheapest lever is **cadence** — halving the frequency halves the cost. The next is
      **model** — a Haiku triage loop costs a fifth of an Opus one; put the cheap model on the
      maker and reserve the expensive one for the gate decision.
      
      ---
      
      ## Multi-loop coordination
      
      Running several loops against one repo, two rules prevent them tripping over each other:
      
      ### Priority order (collision avoidance)
      
      ```
      CI Watch  ►  PR Watch  ►  Dependency Bump  ►  Post-Merge/Changelog  ►  Daily Scan
       (highest)                                                                        (off-peak)
      ```
      
      A red build blocks everyone, so the CI watch wins any worktree contention; daily scan
      yields to all. When two loops want the same worktree/branch, the higher-priority one
      proceeds and the lower defers to its next cadence tick. Loops announce what they're
      touching via [`pigeon`](../../pigeon/SKILL.md) so a peer can see "ci-watch holds a
      worktree on PR #412" and stand off.
      
      ### The kill switch (every loop honors it)
      
      One stop signal, checked at the top of **every** run, that halts **every** loop:
      
      - a **sentinel file** — `.loops/PAUSED` (global) or `.loops/<name>/PAUSED` (one loop), or
      - a **label** — `loop-pause` on the repo/issue, checked via `gh`.
      
      No loop ships without one. It's the difference between "the loops are misbehaving, give me
      a minute" and "the loops are misbehaving, where's the breaker?". Put the exact mechanism
      in `kill_switch:` and make checking it the first action of every run, before the work.
      
      ## See also
      
      - [risk-tiers.md](risk-tiers.md) — the autonomy ladder the config's `tier` selects.
      - [pattern-catalog.md](pattern-catalog.md) — each pattern's place in the priority order.
      - [claude-code-loops.md](claude-code-loops.md) — how the cadence actually fires.
      - [native-scheduling.md](native-scheduling.md) — the `host:` values and the limits each one imposes.
      
  • scripts
    • check-native-facts.py 9.7 KB
      #!/usr/bin/env python3
      """Staleness verifier for loop-ops' native-scheduling facts.
      
      references/native-scheduling.md encodes a fast-moving external surface: the native
      scheduling primitives (CronCreate / the scheduled-tasks MCP / cloud routines) and
      their hard limits. Those limits are load-bearing - loop-doctor refuses a config on
      them - and they are exactly the kind of fact that rots invisibly
      (SKILL-RESOURCE-PROTOCOL.md §7). Two modes guard it:
      
        --offline (default, safe for PR CI): internal consistency, no network.
          * the host vocabulary is ONE set across all four places it appears:
            assets/loop.config.template.yaml, scripts/loop-scaffold.sh (--host),
            scripts/loop-doctor.sh (its case arms), references/native-scheduling.md
          * native-scheduling.md carries its "Verified <date>" stamp
          * every load-bearing limit loop-doctor enforces is still stated in the prose
            (the 1-hour cloud floor, the 7-day session-cron expiry)
        --live (scheduled freshness.yml, never a PR gate): fetch the three upstream doc
          pages and check the numbers we encode still appear in them. A changed number
          upstream is real drift; an unreachable docs host is advisory, not a failure.
      
      Usage:   check-native-facts.py [--offline | --live] [--skill DIR] [--json] [--timeout S]
      Input:   argv flags only (no stdin).
      Output:  stdout = findings (plain rows, or a --json envelope). Data only.
      Stderr:  the verdict line, notices, errors.
      Exit:    0 in sync, 2 usage, 3 a required file is missing, 4 unparseable,
               7 docs unreachable (live, advisory - never a real failure),
               10 drift found
      
      Examples:
        check-native-facts.py --offline              # PR CI: host vocabulary + limits are one set
        check-native-facts.py --live                 # weekly: our numbers vs the published docs
        check-native-facts.py --offline --json | jq '.data[]'
      """
      from __future__ import annotations
      
      import argparse
      import json
      import re
      import sys
      import urllib.error
      import urllib.request
      from pathlib import Path
      
      EX_OK = 0
      EX_USAGE = 2
      EX_NOTFOUND = 3
      EX_UNPARSEABLE = 4
      EX_UNREACHABLE = 7
      EX_DRIFT = 10
      
      SCHEMA = "claude-mods.loop-ops.native-facts/v1"
      
      # The canonical host vocabulary. Every file below must agree with exactly this set -
      # a host added in one place and forgotten in another is the drift this catches.
      HOSTS = {"local", "session-cron", "desktop-task", "cloud-routine", "external"}
      
      # Limits loop-doctor actually enforces, so the prose that justifies them must state
      # them. (needle, where it must appear, why it matters)
      LIMITS = [
          ("1 hour", "cloud-routine minimum interval"),
          ("7 days", "session-cron recurring-task expiry"),
      ]
      
      # Live checks: (url, [(needle, label)]). Needles are the published numbers we encode.
      LIVE_PAGES = [
          ("https://code.claude.com/docs/en/routines", [
              ("minimum interval is one hour", "cloud routine >=1h floor"),
          ]),
          ("https://code.claude.com/docs/en/scheduled-tasks", [
              ("expire 7 days", "session-cron 7-day expiry"),
              ("50 scheduled tasks", "50-task-per-session cap"),
          ]),
          ("https://code.claude.com/docs/en/desktop-scheduled-tasks", [
              ("scheduled-tasks", "desktop task on-disk location"),
          ]),
      ]
      
      
      class Term:
          """Minimal stderr styling; honours TERM_ASCII=1 and a non-tty stderr."""
      
          def __init__(self) -> None:
              import os
      
              self.plain = os.environ.get("TERM_ASCII") == "1" or not sys.stderr.isatty()
      
          def say(self, msg: str) -> None:
              print(msg, file=sys.stderr)
      
      
      class Finding:
          def __init__(self, state: str, check: str, detail: str) -> None:
              self.state, self.check, self.detail = state, check, detail
      
          def as_dict(self) -> dict:
              return {"state": self.state, "check": self.check, "detail": self.detail}
      
      
      def read(path: Path) -> str:
          try:
              return path.read_text(encoding="utf-8", errors="replace")
          except OSError as exc:
              raise FileNotFoundError(str(exc)) from exc
      
      
      def hosts_in_template(text: str) -> set:
          """Hosts named in the `host:` block's inline comments."""
          block = re.search(r"^host:.*?(?=\n[a-z_]+:)", text, re.M | re.S)
          if not block:
              return set()
          return {h for h in HOSTS if re.search(r"\b" + re.escape(h) + r"\b", block.group(0))}
      
      
      def hosts_in_scaffold(text: str) -> set:
          """Hosts accepted by loop-scaffold's --host validation case arm."""
          arm = re.search(r"case \"\$HOST\" in\n\s*([a-z|\-]+)\)", text)
          if not arm:
              return set()
          return set(arm.group(1).split("|"))
      
      
      def hosts_in_doctor(text: str) -> set:
          """Hosts loop-doctor recognises in its host-coherence case arm."""
          arm = re.search(r"case \"\$HOST\" in\n\s*([a-z|\-]+)\) row ok \"host\"", text)
          if not arm:
              return set()
          return set(arm.group(1).split("|"))
      
      
      def hosts_in_reference(text: str) -> set:
          return {h for h in HOSTS if "`" + h + "`" in text}
      
      
      def check_offline(skill: Path) -> list:
          findings = []
          tpl = skill / "assets" / "loop.config.template.yaml"
          scaffold = skill / "scripts" / "loop-scaffold.sh"
          doctor = skill / "scripts" / "loop-doctor.sh"
          ref = skill / "references" / "native-scheduling.md"
          for p in (tpl, scaffold, doctor, ref):
              if not p.is_file():
                  raise FileNotFoundError(str(p))
      
          sources = {
              "loop.config.template.yaml": hosts_in_template(read(tpl)),
              "loop-scaffold.sh --host": hosts_in_scaffold(read(scaffold)),
              "loop-doctor.sh host arm": hosts_in_doctor(read(doctor)),
              "native-scheduling.md": hosts_in_reference(read(ref)),
          }
          for where, found in sources.items():
              if not found:
                  findings.append(Finding("bad", "hosts", f"{where}: no host vocabulary found (parser or format changed)"))
              elif found != HOSTS:
                  missing = sorted(HOSTS - found)
                  extra = sorted(found - HOSTS)
                  detail = f"{where}: " + ", ".join(
                      filter(None, [f"missing {missing}" if missing else "", f"unknown {extra}" if extra else ""])
                  )
                  findings.append(Finding("bad", "hosts", detail))
              else:
                  findings.append(Finding("ok", "hosts", f"{where}: all {len(HOSTS)} hosts"))
      
          ref_text = read(ref)
          stamp = re.search(r"\*\*Verified (\d{4}-\d{2}-\d{2})\*\*", ref_text)
          if stamp:
              findings.append(Finding("ok", "date-stamp", f"native-scheduling.md verified {stamp.group(1)}"))
          else:
              findings.append(Finding("bad", "date-stamp", "native-scheduling.md has no '**Verified YYYY-MM-DD**' stamp"))
      
          for needle, why in LIMITS:
              if needle in ref_text:
                  findings.append(Finding("ok", "limit", f"{why}: '{needle}' documented"))
              else:
                  findings.append(Finding("bad", "limit", f"{why}: '{needle}' not stated - loop-doctor enforces it unexplained"))
          return findings
      
      
      def check_live(timeout: float) -> list:
          findings = []
          unreachable = 0
          for url, needles in LIVE_PAGES:
              try:
                  req = urllib.request.Request(url, headers={"User-Agent": "claude-mods-loop-ops-verifier"})
                  with urllib.request.urlopen(req, timeout=timeout) as resp:
                      body = resp.read().decode("utf-8", errors="replace")
              except (urllib.error.URLError, OSError, ValueError) as exc:
                  unreachable += 1
                  findings.append(Finding("skip", "fetch", f"{url}: unreachable ({exc.__class__.__name__})"))
                  continue
              for needle, label in needles:
                  if needle.lower() in body.lower():
                      findings.append(Finding("ok", "live", f"{label}: still published"))
                  else:
                      findings.append(Finding("bad", "live", f"{label}: '{needle}' no longer in {url} - re-verify native-scheduling.md"))
          if unreachable == len(LIVE_PAGES):
              findings.append(Finding("skip", "live", "all docs pages unreachable - live check advisory only"))
          return findings
      
      
      def main() -> int:
          ap = argparse.ArgumentParser(add_help=False)
          ap.add_argument("--offline", action="store_true")
          ap.add_argument("--live", action="store_true")
          ap.add_argument("--skill", default=str(Path(__file__).resolve().parent.parent))
          ap.add_argument("--json", action="store_true")
          ap.add_argument("--timeout", type=float, default=15.0)
          ap.add_argument("-h", "--help", action="store_true")
          try:
              args = ap.parse_args()
          except SystemExit:
              return EX_USAGE
          if args.help:
              print(__doc__)
              return EX_OK
          if args.offline and args.live:
              print("error: --offline and --live are mutually exclusive", file=sys.stderr)
              return EX_USAGE
      
          term = Term()
          skill = Path(args.skill)
          mode = "live" if args.live else "offline"
      
          try:
              findings = check_offline(skill)
          except FileNotFoundError as exc:
              print(f"error: required file missing: {exc}", file=sys.stderr)
              return EX_NOTFOUND
          except re.error as exc:
              print(f"error: could not parse a source file: {exc}", file=sys.stderr)
              return EX_UNPARSEABLE
      
          if args.live:
              findings += check_live(args.timeout)
      
          bad = [f for f in findings if f.state == "bad"]
          skipped = [f for f in findings if f.state == "skip"]
      
          if args.json:
              print(json.dumps({
                  "schema": SCHEMA,
                  "mode": mode,
                  "in_sync": not bad,
                  "data": [f.as_dict() for f in findings],
              }, indent=2))
          else:
              for f in findings:
                  print(f"{f.state:<5} {f.check:<12} {f.detail}")
      
          if bad:
              term.say(f"native-facts: {len(bad)} drift finding(s) - re-verify references/native-scheduling.md")
              return EX_DRIFT
          if args.live and len(skipped) >= len(LIVE_PAGES):
              term.say("native-facts: docs unreachable - live check skipped (advisory)")
              return EX_UNREACHABLE
          term.say(f"native-facts: in sync ({mode})")
          return EX_OK
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • check-pricing-sync.py 7.1 KB
      #!/usr/bin/env python3
      """Offline verifier: loop-ops pricing must match claude-api-ops's model table.
      
      loop-estimate.py reads assets/model-pricing.json. That table is a *copy* of the
      authoritative "Current Models" table in skills/claude-api-ops/SKILL.md — and a
      copy drifts silently (the exact §7 failure mode). This asserts every model in
      loop-ops' pricing exists in the claude-api-ops table with matching input/output
      prices. Both files are in-repo, so this is a pure OFFLINE consistency check and
      safe to gate PR CI (no network). Live model-id drift is owned by
      claude-api-ops/scripts/check-model-table.py.
      
      Usage:   check-pricing-sync.py [--offline] [--pricing FILE] [--table FILE] [--json]
      Input:   argv flags only (no stdin).
      Output:  stdout = drift findings (plain rows, or --json envelope). Data only.
      Stderr:  the verdict panel, notices, errors.
      Exit:    0 in sync, 2 usage, 3 a file missing, 4 unparseable, 10 drift found
      
      --offline is the default and only mode (accepted for parity with the other §7
      verifiers invoked by tests/check-resources.sh).
      
      Examples:
        check-pricing-sync.py --offline
        check-pricing-sync.py --json | jq '.data[]'
      """
      from __future__ import annotations
      
      import argparse
      import json
      import os
      import re
      import sys
      from pathlib import Path
      
      EX_OK = 0
      EX_USAGE = 2
      EX_NOTFOUND = 3
      EX_UNPARSEABLE = 4
      EX_DRIFT = 10
      
      HERE = Path(__file__).resolve().parent
      DEFAULT_PRICING = HERE.parent / "assets" / "model-pricing.json"
      DEFAULT_TABLE = HERE.parent.parent / "claude-api-ops" / "SKILL.md"
      
      PRICE_RE = re.compile(r"\$?\s*([0-9]+(?:\.[0-9]+)?)")
      
      
      class Term:
          """Minimal ANSI helper (term.sh is bash-only; per TERMINAL-DESIGN.md §9 the
          Python port is inline). Honors FORCE_COLOR / NO_COLOR / TERM_ASCII and the
          bound stream's TTY + encoding so piped data stays plain ASCII."""
      
          _C = {"green": "\033[32m", "red": "\033[31m", "cyan": "\033[36m",
                "dim": "\033[2m", "off": "\033[0m"}
      
          def __init__(self, stream=sys.stderr):
              enc = (getattr(stream, "encoding", "") or "").lower()
              self.ascii = os.environ.get("TERM_ASCII") == "1" or "utf" not in enc
              if os.environ.get("FORCE_COLOR"):
                  self.color = True
              elif (os.environ.get("NO_COLOR") is not None
                    or os.environ.get("TERM") == "dumb"
                    or not getattr(stream, "isatty", lambda: False)()):
                  self.color = False
              else:
                  self.color = True
      
          def c(self, name, text):
              return f"{self._C.get(name,'')}{text}{self._C['off']}" if self.color else text
      
          def mark(self, ok):
              g = ("+" if self.ascii else "✓") if ok else ("x" if self.ascii else "✗")
              return self.c("green" if ok else "red", g)
      
      
      def parse_price(cell: str) -> float | None:
          m = PRICE_RE.search(cell)
          return float(m.group(1)) if m else None
      
      
      def load_pricing(path: Path) -> dict:
          """{model_id: (input_per_mtok, output_per_mtok)} from loop-ops' JSON."""
          if not path.is_file():
              print(f"error: pricing file not found: {path}", file=sys.stderr)
              raise SystemExit(EX_NOTFOUND)
          try:
              data = json.loads(path.read_text(encoding="utf-8"))
              out = {}
              for mid, pr in data.get("models", {}).items():
                  out[mid] = (float(pr["input_per_mtok"]), float(pr["output_per_mtok"]))
              if not out:
                  print(f"error: no models in {path}", file=sys.stderr)
                  raise SystemExit(EX_UNPARSEABLE)
              return out
          except (json.JSONDecodeError, KeyError, TypeError, ValueError) as exc:
              print(f"error: could not parse pricing file: {exc}", file=sys.stderr)
              raise SystemExit(EX_UNPARSEABLE)
      
      
      def load_table(path: Path) -> dict:
          """{model_id: (input_price, output_price)} from the claude-api-ops markdown
          'Current Models' table. Columns: Model | ID | Context | Max Output | Input | Output."""
          if not path.is_file():
              print(f"error: claude-api-ops table not found: {path}", file=sys.stderr)
              raise SystemExit(EX_NOTFOUND)
          table: dict = {}
          in_table = False
          for line in path.read_text(encoding="utf-8").splitlines():
              s = line.strip()
              low = s.lower()
              if s.startswith("|") and "id" in low and "context" in low and "output" in low:
                  in_table = True
                  continue
              if in_table:
                  if not s.startswith("|"):
                      if table:  # table ended
                          break
                      continue
                  if set(s) <= set("|-: "):  # separator row
                      continue
                  cells = [c.strip() for c in s.strip("|").split("|")]
                  if len(cells) < 6:
                      continue
                  mid = cells[1].strip("`").strip()
                  if not mid.startswith("claude-"):
                      continue
                  ip, op = parse_price(cells[4]), parse_price(cells[5])
                  if ip is not None and op is not None:
                      table[mid] = (ip, op)
          if not table:
              print(f"error: no model rows parsed from {path}", file=sys.stderr)
              raise SystemExit(EX_UNPARSEABLE)
          return table
      
      
      def main(argv: list[str]) -> int:
          p = argparse.ArgumentParser(
              prog="check-pricing-sync.py",
              description="Verify loop-ops pricing matches claude-api-ops's model table (offline).",
          )
          p.add_argument("--offline", action="store_true", help="offline consistency check (default/only mode)")
          p.add_argument("--pricing", default=str(DEFAULT_PRICING), help="loop-ops model-pricing.json")
          p.add_argument("--table", default=str(DEFAULT_TABLE), help="claude-api-ops SKILL.md with the model table")
          p.add_argument("--json", action="store_true", help="emit a JSON envelope")
          try:
              args = p.parse_args(argv)
          except SystemExit as exc:
              return EX_USAGE if exc.code not in (0, None) else (exc.code or EX_OK)
      
          pricing = load_pricing(Path(args.pricing))
          table = load_table(Path(args.table))
      
          findings = []
          for mid, (ip, op) in sorted(pricing.items()):
              if mid not in table:
                  findings.append({"model": mid, "issue": "absent from claude-api-ops table",
                                   "loop_ops": [ip, op], "authoritative": None})
                  continue
              tip, top = table[mid]
              if abs(ip - tip) > 1e-9 or abs(op - top) > 1e-9:
                  findings.append({"model": mid, "issue": "price mismatch",
                                   "loop_ops": [ip, op], "authoritative": [tip, top]})
      
          if args.json:
              print(json.dumps({
                  "data": findings,
                  "meta": {"count": len(findings), "models_checked": len(pricing),
                           "in_sync": not findings, "schema": "claude-mods.loop-ops.pricing-sync/v1"},
              }, indent=2))
          else:
              for f in findings:
                  auth = f"authoritative {f['authoritative']}" if f["authoritative"] else "not in table"
                  print(f"DRIFT  {f['model']}: {f['issue']} (loop-ops {f['loop_ops']} vs {auth})")
              t = Term(sys.stderr)
              ok = not findings
              print(f"{t.mark(ok)} pricing-sync: {len(pricing)} model(s) checked, "
                    f"{len(findings)} drift "
                    f"{t.c('dim', '(authoritative: claude-api-ops/SKILL.md)')}", file=sys.stderr)
      
          return EX_DRIFT if findings else EX_OK
      
      
      if __name__ == "__main__":
          sys.exit(main(sys.argv[1:]))
      
    • loop-check.sh 12 KB
      #!/usr/bin/env bash
      # Score an outer-loop config for readiness before it is scheduled.
      #
      # Usage:   loop-check.sh [OPTIONS] <loop.config.yaml>
      # Input:   argv flags + a config path (no stdin).
      # Output:  stdout = findings (plain `SEVERITY  message` rows, or --json envelope).
      #          Data only.
      # Stderr:  the readiness panel (score + verdict), notices, errors.
      # Exit:    0 ready (no errors, score >= --min), 2 usage, 3 config not found,
      #          4 config unparseable, 10 NOT ready (findings present)
      #
      # Scores a flat loop.config.yaml against the tier's requirements: a bounded scope,
      # a defined escalation rule + kill switch, and — at L2+ — a verify gate, a guard, a
      # worktree, and a landing path. The config is parsed without a yq dependency.
      # Pair with loop-scaffold.sh (scaffold) and references/risk-tiers.md (the rubric).
      #
      # Examples:
      #   loop-check.sh .loops/pr-watch/loop.config.yaml
      #   loop-check.sh --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.severity=="error")'
      #   loop-check.sh --min 80 --strict .loops/ci-watch/loop.config.yaml
      set -uo pipefail
      
      readonly EX_OK=0 EX_USAGE=2 EX_NOTFOUND=3 EX_UNPARSEABLE=4 EX_FINDINGS=10
      
      # Terminal design system. stdout = findings (data); the score panel frames on stderr.
      __lib="$(cd "$(dirname "${BASH_SOURCE[0]}")/../../_lib" 2>/dev/null && pwd || true)"
      if [ -n "${__lib:-}" ] && [ -f "$__lib/term.sh" ]; then . "$__lib/term.sh"; term_init 2
      else
        term_panel_open() { :; }; term_panel_close() { :; }; term_panel_vert() { :; }
        term_status_row() { shift; printf '  - %s %s\n' "$1" "${2:-}"; }
        term_pip_bar() { printf '%s/%s' "$2" "$3"; }
        term_color() { shift; printf '%s' "$*"; }; TERM_DOT="|"
      fi
      
      CFG=""
      MIN=70
      STRICT=0
      JSON=0
      
      usage() {
        cat <<'EOF'
      loop-check.sh — score an outer-loop config for readiness.
      
      Usage:
        loop-check.sh [OPTIONS] <loop.config.yaml>
      
      Options:
        --min N        readiness score (0-100) required for a "ready" verdict (default: 70).
        --strict       count warnings toward the NOT-ready signal (exit 10).
        --json         emit a JSON envelope instead of plain rows.
        -h, --help     show this help and exit 0.
      
      Exit codes:
        0 ready    2 usage    3 config not found    4 unparseable    10 NOT ready (findings)
      
      Examples:
        loop-check.sh .loops/pr-watch/loop.config.yaml
        loop-check.sh --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.severity=="error")'
        loop-check.sh --min 80 --strict .loops/ci-watch/loop.config.yaml
      EOF
      }
      
      die_usage() { printf 'error: %s\n' "$1" >&2; echo >&2; usage >&2; exit "$EX_USAGE"; }
      
      # ── parse args ──────────────────────────────────────────────────────────────
      while [[ $# -gt 0 ]]; do
        case "$1" in
          --min)     [[ $# -ge 2 ]] || die_usage "--min needs a value"; MIN="$2"; shift 2 ;;
          --strict)  STRICT=1; shift ;;
          --json)    JSON=1; shift ;;
          -h|--help) usage; exit "$EX_OK" ;;
          -*)        die_usage "unknown flag: $1" ;;
          *)         [[ -z "$CFG" ]] || die_usage "unexpected extra argument: $1"; CFG="$1"; shift ;;
        esac
      done
      
      [[ -n "$CFG" ]] || die_usage "a loop.config.yaml path is required"
      [[ "$MIN" =~ ^[0-9]+$ ]] || die_usage "--min must be an integer (got '$MIN')"
      [[ -f "$CFG" ]] || { printf 'error: config not found: %s\n' "$CFG" >&2; exit "$EX_NOTFOUND"; }
      
      # Normalize Windows-authored configs: strip a leading UTF-8 BOM (line 1) and CR
      # line-endings so a CRLF/BOM file parses identically to a clean LF one (octal BOM +
      # gsub \r are portable across gawk/mawk/BSD awk). Falls back to the original on failure.
      __NORM="$(mktemp 2>/dev/null)" && awk 'NR==1{sub(/^\357\273\277/,"")} {gsub(/\r/,""); print}' "$CFG" > "$__NORM" 2>/dev/null && CFG="$__NORM" && trap 'rm -f "$__NORM"' EXIT
      
      # Unparseable: no top-level `key:` lines at all.
      if ! grep -Eq '^[a-z_]+:' "$CFG"; then
        printf 'error: no parseable top-level keys in %s\n' "$CFG" >&2
        exit "$EX_UNPARSEABLE"
      fi
      
      # ── flat-YAML readers (no yq) ───────────────────────────────────────────────
      cfg_scalar() { # inline scalar value for `^KEY:`; empty if absent or block-list
        awk -v k="$1" -v q="'" '
          $0 ~ "^"k":" {
            sub("^"k":[ \t]*","")
            sub(/[ \t]*#.*$/,"")
            gsub(/^[ \t]+|[ \t]+$/,"")
            gsub(/^"|"$/,""); gsub("^"q"|"q"$","")
            print; exit
          }' "$CFG"
      }
      cfg_has_key() { grep -Eq "^$1:" "$CFG"; }
      cfg_list_items() { # `  - item` lines under `^KEY:`, until the next top-level key
        awk -v k="$1" -v q="'" '
          $0 ~ "^"k":" { inlist=1; next }
          inlist==1 {
            if ($0 ~ /^[ \t]*-[ \t]+/) {
              line=$0
              sub(/^[ \t]*-[ \t]+/,"",line); sub(/[ \t]*#.*$/,"",line)
              gsub(/^[ \t]+|[ \t]+$/,"",line); gsub(/^"|"$/,"",line); gsub("^"q"|"q"$","",line)
              if (line != "") print line
            } else if ($0 ~ /^[^ \t#]/) { inlist=0 }
          }' "$CFG"
      }
      is_placeholder() { [[ "$1" == *"<"*">"* ]]; }   # an unfilled <PLACEHOLDER>
      
      # ── findings + scoring ──────────────────────────────────────────────────────
      FIND_SEV=(); FIND_MSG=()
      CHECKS_TOTAL=0; CHECKS_PASS=0
      add() { FIND_SEV+=("$1"); FIND_MSG+=("$2"); }
      pass() { CHECKS_TOTAL=$((CHECKS_TOTAL+1)); CHECKS_PASS=$((CHECKS_PASS+1)); }
      fail() { CHECKS_TOTAL=$((CHECKS_TOTAL+1)); add "$1" "$2"; }    # $1=severity $2=message
      
      # require <severity> <ok?> <message-on-fail>  — a present+valid scalar check.
      require() { if [[ "$2" -eq 1 ]]; then pass; else fail "$1" "$3"; fi; }
      
      TIER="$(cfg_scalar tier)"
      PMODE="$(cfg_scalar permission_mode)"
      NAME="$(cfg_scalar name)"
      GOAL="$(cfg_scalar goal)"
      ESCAL="$(cfg_scalar escalation)"
      KILL="$(cfg_scalar kill_switch)"
      BUDGET="$(cfg_scalar budget_tokens)"
      VERIFY="$(cfg_scalar verify)"
      GUARD="$(cfg_scalar guard)"
      WORKTREE="$(cfg_scalar worktree)"
      LANDVIA="$(cfg_scalar land_via)"
      CADENCE="$(cfg_scalar cadence)"
      PATTERN="$(cfg_scalar pattern)"
      
      is_l2plus=0; [[ "$TIER" == "L2" || "$TIER" == "L3" ]] && is_l2plus=1
      
      # present-and-not-placeholder predicate
      filled() { [[ -n "$1" ]] && ! is_placeholder "$1"; }
      
      # ── always-applicable checks ────────────────────────────────────────────────
      require error  "$(filled "$NAME" && echo 1 || echo 0)"  "name: missing or placeholder"
      require warning "$(filled "$PATTERN" && echo 1 || echo 0)" "pattern: missing"
      case "$TIER" in L1|L2|L3) pass ;; *) fail error "tier: must be L1|L2|L3 (got '${TIER:-empty}')" ;; esac
      require warning "$(filled "$CADENCE" && echo 1 || echo 0)" "cadence: missing"
      require error  "$(filled "$GOAL" && echo 1 || echo 0)"  "goal: missing or placeholder"
      require error  "$(filled "$ESCAL" && echo 1 || echo 0)" "escalation: undefined — every loop must declare what it escalates"
      require error  "$(filled "$KILL" && echo 1 || echo 0)"  "kill_switch: undefined — no loop ships without a stop signal"
      
      # budget present + numeric
      if [[ -n "$BUDGET" && "$BUDGET" =~ ^[0-9]+$ ]]; then pass; else fail warning "budget_tokens: missing or non-numeric — bound the per-run spend"; fi
      
      # scope present + bounded + not placeholder
      mapfile -t SCOPE_ITEMS < <(cfg_list_items scope)
      SCOPE_INLINE="$(cfg_scalar scope)"
      [[ -n "$SCOPE_INLINE" ]] && SCOPE_ITEMS+=("$SCOPE_INLINE")
      if ! cfg_has_key scope || [[ ${#SCOPE_ITEMS[@]} -eq 0 ]]; then
        fail error "scope: missing — bound what the loop may touch"
      else
        scope_bad=0
        for it in "${SCOPE_ITEMS[@]}"; do
          if is_placeholder "$it"; then fail error "scope: unfilled placeholder ('$it')"; scope_bad=1; break; fi
          case "$it" in '*'|'**'|'.'|'./'|'/'|'') fail error "scope: unbounded ('$it') — a loop that may touch anything is not bounded"; scope_bad=1; break ;; esac
        done
        [[ "$scope_bad" -eq 0 ]] && pass
      fi
      
      # permission_mode present + valid
      case "$PMODE" in
        plan|dontAsk|auto|acceptEdits|bypassPermissions) pass ;;
        "") fail error "permission_mode: missing" ;;
        *)  fail error "permission_mode: invalid ('$PMODE')" ;;
      esac
      
      # permission_mode consistent with tier (warning)
      case "$TIER" in
        L1) case "$PMODE" in plan|dontAsk) pass ;; *) fail warning "permission_mode '$PMODE' is broad for L1 (report-only) — prefer plan or dontAsk" ;; esac ;;
        L2) case "$PMODE" in dontAsk|auto|acceptEdits) pass ;; *) fail warning "permission_mode '$PMODE' fits L2 poorly — prefer dontAsk/auto/acceptEdits" ;; esac ;;
        L3) case "$PMODE" in bypassPermissions) pass ;; *) fail warning "L3 unattended usually needs bypassPermissions in a container (got '$PMODE')" ;; esac ;;
        *)  : ;;
      esac
      
      # ── L2+ checks (code-changing tiers) ────────────────────────────────────────
      if [[ "$is_l2plus" -eq 1 ]]; then
        require error "$(filled "$VERIFY" && echo 1 || echo 0)" "verify: no gate command — a code-changing loop with no gate is invalid"
        require error "$(filled "$GUARD" && echo 1 || echo 0)"  "guard: no must-always-pass command at $TIER"
        if [[ "$WORKTREE" == "true" ]]; then pass; else fail error "worktree: must be true at $TIER — isolate code changes"; fi
        require warning "$(filled "$LANDVIA" && echo 1 || echo 0)" "land_via: undefined — name who gates+lands (e.g. fleet-ops)"
      fi
      
      # ── L3-specific isolation check ─────────────────────────────────────────────
      if [[ "$TIER" == "L3" ]]; then
        if printf '%s %s' "$ESCAL" "${SCOPE_ITEMS[*]:-}" | grep -Eqi 'container|isolat|sandbox|devcontainer'; then
          pass
        else
          fail warning "L3 declares no isolation boundary — bypassPermissions is only safe in a container/VM; note it in escalation"
        fi
      fi
      
      # ── verdict ─────────────────────────────────────────────────────────────────
      ERRORS=0; WARNINGS=0
      for s in "${FIND_SEV[@]:-}"; do
        [[ "$s" == "error" ]] && ERRORS=$((ERRORS+1))
        [[ "$s" == "warning" ]] && WARNINGS=$((WARNINGS+1))
      done
      SCORE=0
      [[ "$CHECKS_TOTAL" -gt 0 ]] && SCORE=$(( CHECKS_PASS * 100 / CHECKS_TOTAL ))
      
      READY=1
      [[ "$ERRORS" -gt 0 ]] && READY=0
      [[ "$SCORE" -lt "$MIN" ]] && READY=0
      [[ "$STRICT" -eq 1 && "$WARNINGS" -gt 0 ]] && READY=0
      
      # ── output ──────────────────────────────────────────────────────────────────
      if [[ "$JSON" -eq 1 ]]; then
        printf '{\n  "data": [\n'
        for i in "${!FIND_SEV[@]}"; do
          msg="${FIND_MSG[$i]//\\/\\\\}"; msg="${msg//\"/\\\"}"
          sep=","; [[ "$i" -eq $(( ${#FIND_SEV[@]} - 1 )) ]] && sep=""
          printf '    {"severity": "%s", "message": "%s"}%s\n' "${FIND_SEV[$i]}" "$msg" "$sep"
        done
        printf '  ],\n  "meta": {"count": %d, "errors": %d, "warnings": %d, "score": %d, "min": %d, "ready": %s, "tier": "%s", "schema": "claude-mods.loop-ops.check/v1"}\n}\n' \
          "${#FIND_SEV[@]}" "$ERRORS" "$WARNINGS" "$SCORE" "$MIN" "$([[ "$READY" -eq 1 ]] && echo true || echo false)" "${TIER:-unknown}"
      else
        if [[ ${#FIND_SEV[@]} -gt 0 ]]; then
          for i in "${!FIND_SEV[@]}"; do
            printf '%-7s %s\n' "$(printf '%s' "${FIND_SEV[$i]}" | tr '[:lower:]' '[:upper:]')" "${FIND_MSG[$i]}"
          done
        fi
        verdict="$([[ "$READY" -eq 1 ]] && echo READY || echo "NOT READY")"
        vstate="$([[ "$READY" -eq 1 ]] && echo ok || echo bad)"
        {
          term_panel_open loop "loop ${TERM_DOT} audit" "${NAME:-$(basename "$(dirname "$CFG")")}"
          term_panel_vert
          term_status_row "$vstate" "$verdict  $(term_pip_bar score "$SCORE" 100)" "score $SCORE/100 ${TERM_DOT} tier ${TIER:-?}"
          term_status_row "$([[ "$ERRORS" -eq 0 ]] && echo ok || echo bad)" "$ERRORS error(s)" "must be 0 to be ready"
          term_status_row "$([[ "$WARNINGS" -eq 0 ]] && echo ok || echo warn)" "$WARNINGS warning(s)" "$([[ "$STRICT" -eq 1 ]] && echo 'block under --strict' || echo advisory)"
          term_panel_vert
          term_panel_close "min $MIN ${TERM_DOT} fix errors before scheduling" ""
        } >&2
      fi
      
      [[ "$READY" -eq 1 ]] && exit "$EX_OK" || exit "$EX_FINDINGS"
      
    • loop-doctor.sh 17 KB
      #!/usr/bin/env bash
      # Preflight a loop config - will this loop actually RUN, or die at 3am?
      #
      # loop-check checks the config is well-formed; loop-doctor checks the loop will
      # execute: the gate command's binary resolves, claude/git are on PATH, the budget
      # can fit a tick, and the permission mode is achievable from where it launches.
      #
      # HOST-AWARE. Since native scheduling landed, "where it launches" is a real
      # variable, so the config's optional `host:` selects which constraints apply - a
      # cloud routine has NO permission mode and a >=1h floor, and this machine's PATH
      # says nothing about it; a session-cron host cannot run unattended at all.
      # Verified surface + limits: references/native-scheduling.md (2026-08-30).
      # Modeled on fleet-worker/scripts/fleet-doctor.sh.
      #
      # Usage:   loop-doctor.sh [--offline|--live] [--json] [-q] <loop.config.yaml>
      # Input:   argv flags + a config path (no stdin).
      # Output:  stdout = check rows (TSV: state<TAB>check<TAB>detail), or a --json envelope.
      # Stderr:  the preflight panel, notices, errors.
      # Exit:    0 ok, 2 usage, 3 config not found, 4 unparseable, 5 missing core dep,
      #          10 a check predicts a runtime failure (a gate binary missing, bypass on
      #          host without isolation, budget too small for a tick)
      #
      #   --offline (default): no PATH/exec - config-shape + budget-vs-cost + permission/
      #                        isolation coherence. Safe for PR CI.
      #   --live:              adds runtime preflight - claude/git on PATH, the verify/guard
      #                        leading binary resolvable, the kill-switch path's parent exists.
      #                        Skipped (not failed) when host: cloud-routine - the tick does
      #                        not run on this machine, so this machine's PATH is irrelevant.
      #
      # Examples:
      #   loop-doctor.sh --offline .loops/pr-watch/loop.config.yaml
      #   loop-doctor.sh --live .loops/ci-watch/loop.config.yaml
      #   loop-doctor.sh --live --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.state=="bad")'
      set -uo pipefail
      
      readonly EX_OK=0 EX_USAGE=2 EX_NOTFOUND=3 EX_UNPARSEABLE=4 EX_MISSING_DEP=5 EX_FINDINGS=10
      
      __lib="$(cd "$(dirname "${BASH_SOURCE[0]}")/../../_lib" 2>/dev/null && pwd || true)"
      if [ -n "${__lib:-}" ] && [ -f "$__lib/term.sh" ]; then . "$__lib/term.sh"; term_init 2
      else
        term_panel_open() { :; }; term_panel_close() { :; }; term_panel_vert() { :; }
        term_status_row() { shift; printf '  - %s %s\n' "$1" "${2:-}"; }
        term_color() { shift; printf '%s' "$*"; }; TERM_DOT="|"
      fi
      
      HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
      PRICING="$HERE/../assets/model-pricing.json"
      
      CFG=""; MODE="offline"; JSON=0; QUIET=0
      
      usage() {
        cat <<'EOF'
      loop-doctor.sh - preflight a loop config (will it actually run?).
      
      Usage:
        loop-doctor.sh [--offline|--live] [--json] [-q] <loop.config.yaml>
      
      Options:
        --offline      config-shape + budget-vs-cost + permission coherence (default; no PATH/exec).
        --live         adds runtime preflight: claude/git on PATH, verify/guard binary resolvable
                       (skipped for host: cloud-routine - ticks do not run on this machine).
        --json         emit a JSON envelope.
        -q, --quiet    suppress the stderr panel.
        -h, --help     show this help and exit 0.
      
      Exit codes:
        0 ok   2 usage   3 not found   4 unparseable   5 missing dep   10 predicted runtime failure
      
      Examples:
        loop-doctor.sh --offline .loops/pr-watch/loop.config.yaml
        loop-doctor.sh --live .loops/ci-watch/loop.config.yaml
        loop-doctor.sh --live --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.state=="bad")'
      EOF
      }
      die_usage() { printf 'error: %s\n' "$1" >&2; echo >&2; usage >&2; exit "$EX_USAGE"; }
      
      while [[ $# -gt 0 ]]; do
        case "$1" in
          --offline) MODE="offline"; shift ;;
          --live)    MODE="live"; shift ;;
          --json)    JSON=1; shift ;;
          -q|--quiet) QUIET=1; shift ;;
          -h|--help) usage; exit "$EX_OK" ;;
          -*)        die_usage "unknown flag: $1" ;;
          *)         [[ -z "$CFG" ]] || die_usage "unexpected extra argument: $1"; CFG="$1"; shift ;;
        esac
      done
      
      command -v awk  >/dev/null 2>&1 || { echo "loop-doctor: awk required" >&2; exit "$EX_MISSING_DEP"; }
      command -v grep >/dev/null 2>&1 || { echo "loop-doctor: grep required" >&2; exit "$EX_MISSING_DEP"; }
      
      [[ -n "$CFG" ]] || die_usage "a loop.config.yaml path is required"
      [[ -f "$CFG" ]] || { printf 'error: config not found: %s\n' "$CFG" >&2; exit "$EX_NOTFOUND"; }
      # Normalize Windows-authored configs: strip a leading UTF-8 BOM + CR line-endings so a
      # CRLF/BOM file parses like a clean LF one (portable octal BOM + gsub \r).
      __NORM="$(mktemp 2>/dev/null)" && awk 'NR==1{sub(/^\357\273\277/,"")} {gsub(/\r/,""); print}' "$CFG" > "$__NORM" 2>/dev/null && CFG="$__NORM" && trap 'rm -f "$__NORM"' EXIT
      grep -Eq '^[a-z_]+:' "$CFG" || { printf 'error: no parseable keys in %s\n' "$CFG" >&2; exit "$EX_UNPARSEABLE"; }
      
      # Pick a working python for the budget-vs-cost check (skipped gracefully if none).
      PY=""
      for c in python python3 py; do
        if command -v "$c" >/dev/null 2>&1 && "$c" -c "" >/dev/null 2>&1; then PY="$c"; break; fi
      done
      
      # ── flat-YAML readers (no yq), same contract as loop-check.sh ────────────────
      cfg_scalar() {
        awk -v k="$1" -v q="'" '
          $0 ~ "^"k":" { sub("^"k":[ \t]*",""); sub(/[ \t]*#.*$/,""); gsub(/^[ \t]+|[ \t]+$/,"");
            gsub(/^"|"$/,""); gsub("^"q"|"q"$",""); print; exit }' "$CFG"
      }
      cfg_list_items() {
        awk -v k="$1" -v q="'" '
          $0 ~ "^"k":" { inlist=1; next }
          inlist==1 { if ($0 ~ /^[ \t]*-[ \t]+/) { line=$0; sub(/^[ \t]*-[ \t]+/,"",line); sub(/[ \t]*#.*$/,"",line);
              gsub(/^[ \t]+|[ \t]+$/,"",line); gsub(/^"|"$/,"",line); gsub("^"q"|"q"$","",line); if (line!="") print line }
            else if ($0 ~ /^[^ \t#]/) { inlist=0 } }' "$CFG"
      }
      
      TIER="$(cfg_scalar tier)"; PMODE="$(cfg_scalar permission_mode)"; PATTERN="$(cfg_scalar pattern)"
      VERIFY="$(cfg_scalar verify)"; GUARD="$(cfg_scalar guard)"; BUDGET="$(cfg_scalar budget_tokens)"
      KILL="$(cfg_scalar kill_switch)"; ESCAL="$(cfg_scalar escalation)"
      CADENCE="$(cfg_scalar cadence)"; HOST="$(cfg_scalar host)"; [[ -z "$HOST" ]] && HOST="local"
      WORKTREE="$(cfg_scalar worktree)"
      is_l2plus=0; [[ "$TIER" == "L2" || "$TIER" == "L3" ]] && is_l2plus=1
      
      # ── findings ─────────────────────────────────────────────────────────────
      ROWS=()       # "state\tcheck\tdetail"
      FINDING=0
      row() { ROWS+=("$1"$'\t'"$2"$'\t'"$3"); [[ "$1" == "bad" ]] && FINDING=1; }
      
      # leading binary of a command string (first whitespace token; strips a leading VAR= prefix)
      lead_bin() { awk '{ for(i=1;i<=NF;i++){ if($i !~ /=/){print $i; exit} } }' <<<"$1"; }
      
      # ── OFFLINE checks ───────────────────────────────────────────────────────
      # Cadence in minutes, for the host floor checks. Nm/Nh/Nd, or "*/N * * * *" -> N.
      # Anything richer returns empty and the floor check is SKIPPED rather than guessed:
      # a wrong floor finding is worse than no finding.
      #
      # The digits-only guard is load-bearing, not defensive padding: `1a2m` matches the
      # *[0-9]m glob, and ${1%m} would hand `1a2` to arithmetic - which errors to stderr and
      # leaves a garbage CAD_MIN that later trips `[[ -lt ]]`. Malformed cadence is loop-check's
      # finding to report; here it must simply yield "unknown" and skip the floor check.
      cadence_minutes() {
        local n=""
        case "$1" in
          *[0-9]m) n="${1%m}"; [[ "$n" =~ ^[0-9]+$ ]] && printf '%s' "$n" ;;
          *[0-9]h) n="${1%h}"; [[ "$n" =~ ^[0-9]+$ ]] && printf '%s' "$(( n * 60 ))" ;;
          *[0-9]d) n="${1%d}"; [[ "$n" =~ ^[0-9]+$ ]] && printf '%s' "$(( n * 1440 ))" ;;
          */[0-9]*\ *) awk '{ n=$1; sub(/^\*\//,"",n); if (n ~ /^[0-9]+$/ && $2=="*") print n }' <<<"$1" ;;
          *) printf '' ;;
        esac
      }
      CAD_MIN="$(cadence_minutes "$CADENCE" 2>/dev/null)"
      
      # Host coherence. `host:` names where ticks execute; each surface has different hard
      # limits (references/native-scheduling.md, verified 2026-08-30).
      case "$HOST" in
        local|external|desktop-task|cloud-routine|session-cron) row ok "host" "$HOST" ;;
        *) row bad "host" "unknown host '$HOST' - use local|session-cron|desktop-task|cloud-routine|external" ;;
      esac
      
      case "$HOST" in
        session-cron)
          # /loop + CronCreate are session-scoped: they need an open, idle session and every
          # recurring job self-deletes 7 days after creation. Fine for L1 supervised polling;
          # it cannot host an unattended loop, which is what L2+ means.
          if [[ "$is_l2plus" -eq 1 ]]; then
            row bad "host/tier" "session-cron can't run unattended ($TIER) - needs an open idle session and expires after 7 days; use desktop-task or external"
          else
            row warn "host/tier" "session-cron is supervised-only: open idle session, 7-day expiry, no catch-up for missed fires"
          fi
          ;;
        desktop-task)
          # One catch-up run for the most recently missed window; older ones are discarded,
          # so a slow tick can land at any hour. The prompt needs its own time guardrails.
          if [[ -n "$CAD_MIN" ]] && [[ "$CAD_MIN" -ge 720 ]]; then
            row warn "catch-up" "desktop-task runs ONE catch-up for the latest missed window - a $CADENCE tick may fire hours late; put time guardrails in run.md"
          fi
          # `worktree: true` in this config is a DECLARATION, not the switch. The real toggle
          # lives on the task itself and is OFF by default, so a task can satisfy the config
          # while running against the working dir including uncommitted changes. Nothing on
          # disk lets us verify it - so say so rather than implying the config settled it.
          if [[ "$WORKTREE" == "true" ]]; then
            row warn "worktree" "config declares worktree: true - confirm the TASK's worktree toggle is on (it is off by default); this file cannot enforce it"
          fi
          ;;
        cloud-routine)
          # Routines run autonomously in the cloud: no permission-mode picker, >=1h floor,
          # fresh clone with no local files. The boundary is repos + environment + connectors.
          if [[ -n "$CAD_MIN" ]] && [[ "$CAD_MIN" -lt 60 ]]; then
            row bad "cadence" "cloud-routine minimum interval is 1 hour - '$CADENCE' is rejected at creation"
          fi
          if printf '%s %s' "$ESCAL" "$(cfg_list_items scope | tr '\n' ' ')" | grep -Eqi 'connectors?|environment|network access|repositor'; then
            row ok "boundary" "cloud-routine boundary names repos/environment/connectors"
          else
            row bad "boundary" "cloud-routine has NO permission mode - the boundary must be repos + environment network policy + connectors (ALL connectors attach by default); name it in scope/escalation"
          fi
          ;;
      esac
      
      # Permission mode achievability. A cloud routine has no permission-mode picker at all,
      # so requiring one there would be a false finding - the boundary check above replaces it.
      if [[ "$HOST" == "cloud-routine" ]]; then
        if [[ -n "$PMODE" ]]; then
          row warn "permission_mode" "'$PMODE' is ignored by cloud routines (they run autonomously, no approval prompts)"
        else
          row ok "permission_mode" "n/a for cloud-routine"
        fi
      else
        case "$PMODE" in
          default) row bad "permission_mode" "default is interactive - a headless 'claude -p' tick can't answer prompts; use dontAsk/auto/bypassPermissions" ;;
          "")      row bad "permission_mode" "missing" ;;
          *)       row ok  "permission_mode" "$PMODE" ;;
        esac
      fi
      # L3 bypass needs an isolation boundary.
      #
      # A cloud routine already IS one: it runs in Anthropic-managed cloud infrastructure on a
      # fresh clone, and its permission_mode is ignored entirely. Demanding a "container" note
      # there is a false finding that teaches people to write a bogus note to satisfy the tool -
      # so the cloud host reports its real boundary (environment + connectors, checked above)
      # instead of the host-isolation one.
      if [[ "$TIER" == "L3" && "$HOST" == "cloud-routine" ]]; then
        row ok "isolation" "cloud-routine runs in Anthropic-managed cloud infra - the boundary is environment + connectors, not a local container"
      elif [[ "$TIER" == "L3" && "$PMODE" == "bypassPermissions" ]]; then
        if printf '%s %s' "$ESCAL" "$(cfg_list_items scope | tr '\n' ' ')" | grep -Eqi 'container|isolat|sandbox|devcontainer'; then
          row ok "isolation" "L3 bypass declares an isolation boundary"
        else
          row bad "isolation" "L3 + bypassPermissions with no container/sandbox note - only safe in an isolated VM/container"
        fi
      fi
      # Budget vs estimated tokens/run.
      if [[ -n "$BUDGET" && "$BUDGET" =~ ^[0-9]+$ && -n "$PY" && -n "$PATTERN" && -f "$PRICING" ]]; then
        TPR="$(PR="$PRICING" PAT="$PATTERN" "$PY" -c "import json,os
      try:
       d=json.load(open(os.environ['PR']))['_pattern_defaults'].get(os.environ['PAT'])
       print((int(d['input'])+int(d['output']))*int(d.get('subagents',1)) if d else '')
      except Exception: print('')" 2>/dev/null)"
        if [[ -n "$TPR" && "$TPR" =~ ^[0-9]+$ ]]; then
          if [[ "$BUDGET" -lt "$TPR" ]]; then
            row bad "budget" "budget_tokens $BUDGET < ~$TPR est. tokens/run for $PATTERN - a tick can't complete"
          else
            row ok "budget" "budget_tokens $BUDGET >= ~$TPR est. tokens/run"
          fi
        fi
      fi
      
      # ── LIVE checks ──────────────────────────────────────────────────────────
      # Skipped wholesale for cloud-routine: the tick runs on a fresh cloud clone, so this
      # machine's PATH, git and gate binaries say nothing about whether it will run. A pass
      # here would be false confidence - worse than no check.
      if [[ "$MODE" == "live" && "$HOST" == "cloud-routine" ]]; then
        row warn "live" "skipped - host cloud-routine runs on a fresh cloud clone; verify the gate in the routine's environment setup script instead"
      elif [[ "$MODE" == "live" ]]; then
        if command -v claude >/dev/null 2>&1; then row ok "claude" "on PATH"; else row warn "claude" "not on PATH - the scheduler that runs 'claude -p' must have it"; fi
        if command -v git >/dev/null 2>&1; then
          row ok "git" "on PATH"
          if [[ "$is_l2plus" -eq 1 ]] && ! git worktree list >/dev/null 2>&1; then
            row warn "worktree" "'git worktree' unavailable here - L2+ isolates changes in a worktree"
          fi
        elif [[ "$is_l2plus" -eq 1 ]]; then
          row bad "git" "git not on PATH - L2+ needs it for worktree isolation + landing"
        else
          row warn "git" "git not on PATH"
        fi
        # verify / guard leading binary resolvable
        for pair in "verify:$VERIFY" "guard:$GUARD"; do
          label="${pair%%:*}"; cmd="${pair#*:}"
          [[ -z "$cmd" ]] && continue
          case "$cmd" in *"<"*">"*) continue ;; esac   # unfilled placeholder - audit's job
          bin="$(lead_bin "$cmd")"
          [[ -z "$bin" ]] && continue
          if [[ "$bin" == */* ]]; then
            [[ -x "$bin" ]] && row ok "$label" "$bin executable" || row bad "$label" "$bin not executable - the gate can't run"
          elif command -v "$bin" >/dev/null 2>&1; then
            row ok "$label" "$bin resolves"
          else
            row bad "$label" "'$bin' not on PATH - the gate command can't run at tick time"
          fi
        done
        # kill-switch path parent exists (only when it clearly names a path)
        ks_path="$(grep -oE '[^ "'"'"']*/[^ "'"'"']*' <<<"$KILL" | head -1)"
        if [[ -n "$ks_path" ]]; then
          parent="$(dirname "$ks_path")"
          [[ -d "$parent" || "$parent" == "." ]] && row ok "kill_switch" "sentinel path parent exists ($parent)" \
            || row warn "kill_switch" "sentinel parent dir missing ($parent) - create it so the switch works"
        fi
      fi
      
      # ── output ───────────────────────────────────────────────────────────────
      n_bad=0; n_warn=0; n_ok=0
      for r in "${ROWS[@]:-}"; do
        case "${r%%$'\t'*}" in bad) n_bad=$((n_bad+1));; warn) n_warn=$((n_warn+1));; ok) n_ok=$((n_ok+1));; esac
      done
      
      if [[ "$JSON" -eq 1 ]]; then
        printf '{\n  "data": [\n'
        if [[ ${#ROWS[@]} -gt 0 ]]; then
         for i in "${!ROWS[@]}"; do
          IFS=$'\t' read -r st ck dt <<<"${ROWS[$i]}"
          dt="${dt//\\/\\\\}"; dt="${dt//\"/\\\"}"
          sep=","; [[ "$i" -eq $(( ${#ROWS[@]} - 1 )) ]] && sep=""
          printf '    {"state": "%s", "check": "%s", "detail": "%s"}%s\n' "$st" "$ck" "$dt" "$sep"
         done
        fi
        printf '  ],\n  "meta": {"mode": "%s", "ok": %d, "warn": %d, "bad": %d, "will_run": %s, "tier": "%s", "schema": "claude-mods.loop-ops.doctor/v1"}\n}\n' \
          "$MODE" "$n_ok" "$n_warn" "$n_bad" "$([[ "$FINDING" -eq 0 ]] && echo true || echo false)" "${TIER:-unknown}"
      else
        if [[ ${#ROWS[@]} -gt 0 ]]; then
          for r in "${ROWS[@]}"; do
            IFS=$'\t' read -r st ck dt <<<"$r"
            printf '%-5s %-14s %s\n' "$st" "$ck" "$dt"
          done
        fi
        if [[ "$QUIET" -eq 0 ]]; then
          verdict="$([[ "$FINDING" -eq 0 ]] && echo "WILL RUN" || echo "WILL FAIL")"
          vstate="$([[ "$FINDING" -eq 0 ]] && echo ok || echo bad)"
          {
            term_panel_open loop "loop ${TERM_DOT} doctor ($MODE)" "$(basename "$(dirname "$CFG")")"
            term_panel_vert
            term_status_row "$vstate" "$verdict" "$n_bad blocking ${TERM_DOT} $n_warn advisory ${TERM_DOT} $n_ok ok"
            [[ "$MODE" == "offline" ]] && term_status_row skip "run --live before scheduling" "checks gate binaries + PATH"
            term_panel_vert
            term_panel_close "audit = well-formed ${TERM_DOT} doctor = will-run" ""
          } >&2
        fi
      fi
      
      [[ "$FINDING" -eq 0 ]] && exit "$EX_OK" || exit "$EX_FINDINGS"
      
    • loop-estimate.py 14.7 KB
      #!/usr/bin/env python3
      """Estimate the token/$ cost of an outer loop by pattern × cadence × model.
      
      A loop's cost is runs/day × tokens/run × price, and sub-agents multiply tokens/run.
      This computes that - and, crucially, models **prompt caching**: a loop re-sends the
      SAME run.md + system prefix every tick (the Ralph property), which is the textbook
      caching case. Whether caching helps depends on cadence vs cache TTL, so this picks the
      TTL and reports the cached projection alongside the naive one.
      
      Pricing reads from assets/model-pricing.json (date-stamped; skills/claude-api-ops is
      the source of truth - run its check-model-table.py if you suspect drift).
      
      Usage:   loop-estimate.py --pattern P --cadence C --model M [OPTIONS]
      Input:   argv flags only (no stdin).
      Output:  stdout = the cost breakdown (plain rows, or --json envelope). Data only.
      Stderr:  the assumptions + caching note, errors.
      Exit:    0 ok, 2 usage, 3 pricing file missing, 4 bad cadence/model/pattern
      
      Estimates, not guarantees - reconcile against the loop's run-log.md actuals. Levers in
      order of impact: cadence (halving frequency halves cost), prompt caching (model below),
      model tier.
      
      Examples:
        loop-estimate.py --pattern pr-watch --cadence 10m --model claude-haiku-4-5
        loop-estimate.py --pattern ci-watch --cadence 15m --model claude-sonnet-5 --days 30 --json
        loop-estimate.py --pattern daily-scan --cadence 6h --model claude-opus-5   # too slow to cache
        loop-estimate.py --list-models
      """
      from __future__ import annotations
      
      import argparse
      import json
      import os
      import re
      import sys
      from pathlib import Path
      
      EX_OK = 0
      EX_USAGE = 2
      EX_NOTFOUND = 3
      EX_VALIDATION = 4
      
      DEFAULT_PRICING = Path(__file__).resolve().parent.parent / "assets" / "model-pricing.json"
      
      # Prompt-caching multipliers vs base input price (claude-api-ops/references/caching-and-cost.md).
      CACHE_WRITE_5M = 1.25   # write a 5-minute-TTL entry
      CACHE_WRITE_1H = 2.0    # write a 1-hour-TTL entry
      CACHE_READ = 0.1        # read any cached entry
      
      # Minimum cacheable prefix (tokens) - below this the cache_control marker is silently
      # ignored (caching-and-cost.md). A loop whose static prefix is smaller can't cache.
      MIN_PREFIX = {
          "claude-fable-5": 512,
          "claude-opus-5": 512,
          "claude-sonnet-5": 1024,
          "claude-haiku-4-5": 4096,
      }
      DEFAULT_MIN_PREFIX = 1024
      
      
      class Term:
          """Minimal ANSI helper (term.sh is bash-only; per TERMINAL-DESIGN.md §9 the Python
          port is inline). Honors FORCE_COLOR / NO_COLOR / TERM_ASCII and the bound stream's
          TTY + encoding, so piped data stays plain ASCII."""
      
          _C = {"green": "\033[32m", "cyan": "\033[36m", "dim": "\033[2m", "off": "\033[0m"}
      
          def __init__(self, stream=sys.stderr):
              enc = (getattr(stream, "encoding", "") or "").lower()
              self.ascii = os.environ.get("TERM_ASCII") == "1" or "utf" not in enc
              if os.environ.get("FORCE_COLOR"):
                  self.color = True
              elif (os.environ.get("NO_COLOR") is not None
                    or os.environ.get("TERM") == "dumb"
                    or not getattr(stream, "isatty", lambda: False)()):
                  self.color = False
              else:
                  self.color = True
      
          def c(self, name, text):
              return f"{self._C.get(name,'')}{text}{self._C['off']}" if self.color else text
      
      
      def load_pricing(path: Path) -> dict:
          if not path.is_file():
              print(f"error: pricing file not found: {path}", file=sys.stderr)
              raise SystemExit(EX_NOTFOUND)
          try:
              return json.loads(path.read_text(encoding="utf-8"))
          except (json.JSONDecodeError, OSError) as exc:
              print(f"error: could not read pricing file: {exc}", file=sys.stderr)
              raise SystemExit(EX_VALIDATION)
      
      
      def runs_per_day(cadence: str, override: float | None) -> float:
          """Translate a cadence into runs/day. Supports Nm/Nh/Nd and the common cron
          forms `*/N * * * *` and `N * * * *`. --runs-per-day overrides everything."""
          if override is not None:
              if override <= 0:
                  print("error: --runs-per-day must be positive", file=sys.stderr)
                  raise SystemExit(EX_VALIDATION)
              return float(override)
      
          s = cadence.strip()
          m = re.fullmatch(r"(\d+)([mhd])", s)
          if m:
              n = int(m.group(1))
              if n <= 0:
                  print(f"error: cadence value must be positive (got '{cadence}')", file=sys.stderr)
                  raise SystemExit(EX_VALIDATION)
              return {"m": 1440.0, "h": 24.0, "d": 1.0}[m.group(2)] / n
          cron_min = re.fullmatch(r"\*/(\d+) \* \* \* \*", s)
          if cron_min:
              n = int(cron_min.group(1))
              return 1440.0 / n if n > 0 else 1440.0
          if re.fullmatch(r"\d+ \* \* \* \*", s):
              return 24.0
          print(
              f"error: cannot derive runs/day from cadence '{cadence}' - "
              "use Nm/Nh/Nd, `*/N * * * *`, or pass --runs-per-day",
              file=sys.stderr,
          )
          raise SystemExit(EX_VALIDATION)
      
      
      def caching_projection(in_tok, out_tok, sub, in_price, out_price, rpd, model,
                             prefix_frac, ttl_choice):
          """Model prompt-caching of the static run-prompt prefix across ticks.
      
          Returns a dict: ttl, beneficial, reason, cost_per_run/day, prefix_tokens.
          The cache stays warm only when the tick interval is <= the TTL (reads refresh it);
          a loop slower than the 1h max TTL writes a cold entry every tick - caching can't help.
          """
          interval_min = 1440.0 / rpd if rpd > 0 else 1e9
          prefix_tokens = int(round(in_tok * prefix_frac))
          variable_in = in_tok - prefix_tokens
          min_prefix = MIN_PREFIX.get(model, DEFAULT_MIN_PREFIX)
      
          # Pick TTL: smallest that stays warm at this cadence.
          if ttl_choice == "5m":
              ttl, warm = "5m", interval_min <= 5
          elif ttl_choice == "1h":
              ttl, warm = "1h", interval_min <= 60
          else:  # auto
              if interval_min <= 5:
                  ttl, warm = "5m", True
              elif interval_min <= 60:
                  ttl, warm = "1h", True
              else:
                  ttl, warm = None, False
      
          out_cost_day = out_tok / 1e6 * out_price * rpd
      
          if prefix_tokens < min_prefix:
              return {"ttl": ttl, "beneficial": False,
                      "reason": f"static prefix ~{prefix_tokens} tok < {model} minimum {min_prefix} tok "
                                "- cache marker silently ignored; enlarge the run prompt/system or skip caching",
                      "prefix_tokens": prefix_tokens, "cost_per_day": None, "cost_per_run": None}
          if not warm or ttl is None:
              return {"ttl": ttl, "beneficial": False,
                      "reason": f"tick interval ~{interval_min:.0f} min exceeds the cache TTL "
                                "- the entry expires between ticks, so every tick is a cold write; caching won't help",
                      "prefix_tokens": prefix_tokens, "cost_per_day": None, "cost_per_run": None}
      
          write_mult = CACHE_WRITE_5M if ttl == "5m" else CACHE_WRITE_1H
          # Per day, warm: ~1 cache write of the prefix + (rpd-1) reads; variable input + output full price.
          prefix_day = prefix_tokens / 1e6 * in_price * (write_mult + max(rpd - 1, 0) * CACHE_READ)
          variable_day = variable_in / 1e6 * in_price * rpd
          cost_day = (prefix_day + variable_day + out_cost_day) * sub
          return {"ttl": ttl, "beneficial": True, "reason": "",
                  "prefix_tokens": prefix_tokens, "write_mult": write_mult,
                  "cost_per_day": cost_day, "cost_per_run": cost_day / rpd if rpd else cost_day}
      
      
      def fmt_money(x: float) -> str:
          if x < 1:
              return f"${x:.4f}"
          return f"${x:,.2f}"
      
      
      def main(argv: list[str]) -> int:
          p = argparse.ArgumentParser(
              prog="loop-estimate.py",
              description="Estimate outer-loop cost by pattern × cadence × model, with prompt caching.",
          )
          p.add_argument("--pattern", default="custom", help="catalog pattern key (default: custom)")
          p.add_argument("--cadence", default="1h", help="10m | 1h | 6h | 1d, or a cron string (default: 1h)")
          p.add_argument("--model", default="claude-haiku-4-5", help="model id (default: claude-haiku-4-5)")
          p.add_argument("--days", type=int, default=30, help="horizon in days for the total (default: 30)")
          p.add_argument("--runs-per-day", type=float, default=None, help="override the cadence-derived runs/day")
          p.add_argument("--input-tokens", type=int, default=None, help="override per-run input tokens")
          p.add_argument("--output-tokens", type=int, default=None, help="override per-run output tokens")
          p.add_argument("--subagents", type=int, default=None, help="override the sub-agent fan-out multiplier")
          p.add_argument("--cache-prefix-frac", type=float, default=0.6,
                         help="fraction of input that is the static, cacheable run-prompt prefix (default: 0.6)")
          p.add_argument("--cache-ttl", choices=["auto", "5m", "1h"], default="auto",
                         help="cache TTL to model (default: auto - pick by cadence)")
          p.add_argument("--no-cache", action="store_true", help="report the uncached cost only")
          p.add_argument("--pricing", default=str(DEFAULT_PRICING), help="path to model-pricing.json")
          p.add_argument("--list-models", action="store_true", help="print the pricing table + as-of date, exit 0")
          p.add_argument("--json", action="store_true", help="emit a JSON envelope")
          try:
              args = p.parse_args(argv)
          except SystemExit as exc:
              return EX_USAGE if exc.code not in (0, None) else (exc.code or EX_OK)
      
          pricing = load_pricing(Path(args.pricing))
          models = pricing.get("models", {})
          as_of = pricing.get("_as_of", "unknown")
          pattern_defaults = pricing.get("_pattern_defaults", {})
      
          if args.list_models:
              if args.json:
                  print(json.dumps({"data": models, "meta": {"as_of": as_of, "schema": "claude-mods.loop-ops.pricing/v1"}}, indent=2))
              else:
                  print(f"{'model':<22}{'input $/MTok':>14}{'output $/MTok':>16}")
                  for mid, pr in models.items():
                      print(f"{mid:<22}{pr.get('input_per_mtok', 0):>14.2f}{pr.get('output_per_mtok', 0):>16.2f}")
                  print(f"\n(as of {as_of}; source of truth: claude-api-ops)", file=sys.stderr)
              return EX_OK
      
          if args.days <= 0:
              print("error: --days must be positive", file=sys.stderr)
              return EX_VALIDATION
          if not (0.0 <= args.cache_prefix_frac <= 1.0):
              print("error: --cache-prefix-frac must be between 0 and 1", file=sys.stderr)
              return EX_VALIDATION
      
          if args.model not in models:
              print(f"error: unknown model '{args.model}' - known: {', '.join(models) or '(none)'}", file=sys.stderr)
              return EX_VALIDATION
          in_price = float(models[args.model]["input_per_mtok"])
          out_price = float(models[args.model]["output_per_mtok"])
      
          if args.input_tokens is not None and args.output_tokens is not None:
              in_tok, out_tok = args.input_tokens, args.output_tokens
              sub = args.subagents if args.subagents is not None else 1
          elif args.pattern in pattern_defaults and not args.pattern.startswith("_"):
              d = pattern_defaults[args.pattern]
              in_tok = args.input_tokens if args.input_tokens is not None else int(d["input"])
              out_tok = args.output_tokens if args.output_tokens is not None else int(d["output"])
              sub = args.subagents if args.subagents is not None else int(d.get("subagents", 1))
          else:
              print(
                  f"error: unknown pattern '{args.pattern}' - pass --input-tokens and "
                  f"--output-tokens, or use one of: {', '.join(k for k in pattern_defaults if not k.startswith('_'))}",
                  file=sys.stderr,
              )
              return EX_VALIDATION
      
          if min(in_tok, out_tok, sub) < 0:
              print("error: token counts and --subagents must be non-negative", file=sys.stderr)
              return EX_VALIDATION
      
          rpd = runs_per_day(args.cadence, args.runs_per_day)
      
          # ── uncached (naive) ──
          cost_in = in_tok / 1_000_000 * in_price
          cost_out = out_tok / 1_000_000 * out_price
          cost_run = (cost_in + cost_out) * sub
          tokens_run = (in_tok + out_tok) * sub
          cost_day = cost_run * rpd
          cost_horizon = cost_day * args.days
      
          # ── cached projection ──
          cache = None
          if not args.no_cache:
              cache = caching_projection(in_tok, out_tok, sub, in_price, out_price, rpd,
                                         args.model, args.cache_prefix_frac, args.cache_ttl)
      
          if args.json:
              data = {
                  "pattern": args.pattern, "model": args.model, "cadence": args.cadence,
                  "runs_per_day": round(rpd, 3), "tokens_per_run": tokens_run,
                  "input_tokens": in_tok, "output_tokens": out_tok, "subagents": sub,
                  "cost_per_run": round(cost_run, 6), "cost_per_day": round(cost_day, 4),
                  "days": args.days, "cost_per_horizon": round(cost_horizon, 2),
              }
              if cache is not None:
                  if cache["beneficial"]:
                      cd = cache["cost_per_day"]
                      data["caching"] = {
                          "beneficial": True, "ttl": cache["ttl"], "prefix_tokens": cache["prefix_tokens"],
                          "cost_per_day": round(cd, 4), "cost_per_horizon": round(cd * args.days, 2),
                          "savings_pct": round((cost_day - cd) / cost_day * 100, 1) if cost_day else 0.0,
                      }
                  else:
                      data["caching"] = {"beneficial": False, "reason": cache["reason"],
                                         "prefix_tokens": cache["prefix_tokens"]}
              print(json.dumps({"data": data, "meta": {"as_of": as_of, "schema": "claude-mods.loop-ops.estimate/v1"}}, indent=2))
              return EX_OK
      
          t = Term(sys.stderr)
          print(f"{'pattern:':<16}{args.pattern}")
          print(f"{'model:':<16}{args.model}")
          print(f"{'cadence:':<16}{args.cadence}  ->  {rpd:g} runs/day")
          print(f"{'tokens/run:':<16}{tokens_run:,} ({in_tok:,} in + {out_tok:,} out) x {sub} subagent(s)")
          print(f"{'cost/run:':<16}{fmt_money(cost_run)}")
          print(f"{'cost/day:':<16}{fmt_money(cost_day)}")
          print(f"{'cost/'+str(args.days)+'d:':<16}{fmt_money(cost_horizon)}  (uncached)")
          if cache is not None:
              if cache["beneficial"]:
                  cd, ch = cache["cost_per_day"], cache["cost_per_day"] * args.days
                  save = (cost_day - cd) / cost_day * 100 if cost_day else 0.0
                  print(f"{'cached/'+str(args.days)+'d:':<16}{t.c('cyan', fmt_money(ch))}  "
                        f"({t.c('green', f'-{save:.0f}%')}, TTL {cache['ttl']}, prefix ~{cache['prefix_tokens']:,} tok)")
                  print(f"recommendation: cache the static run.md+system prefix at TTL {cache['ttl']} "
                        f"-> ~-{save:.0f}%/mo. Keep run.md BYTE-IDENTICAL every tick or the cache never hits.",
                        file=sys.stderr)
              else:
                  print(f"caching: not beneficial here", file=sys.stderr)
                  print(f"  why: {cache['reason']}", file=sys.stderr)
          print(f"estimate (as of {as_of} pricing) - reconcile against run-log.md actuals; "
                "cadence is the biggest lever, then caching, then model tier", file=sys.stderr)
          return EX_OK
      
      
      if __name__ == "__main__":
          sys.exit(main(sys.argv[1:]))
      
    • loop-scaffold.sh 16 KB
      #!/usr/bin/env bash
      # Scaffold an outer-loop state spine (loop.config.yaml + STATE.md + run-log.md).
      #
      # Usage:   loop-scaffold.sh --name NAME [OPTIONS]
      # Input:   argv flags only (no stdin).
      # Output:  stdout = the created loop.config.yaml path (data). Under --dry-run, the
      #          path then the rendered config. Data only.
      # Stderr:  the creation panel, reminders, warnings, errors.
      # Exit:    0 created (or dry-run rendered), 2 usage, 3 template/dir not found,
      #          5 precondition (target dir already populated, no --force)
      #
      # Creates <dir>/<name>/ from the bundled templates, substituting name/pattern/tier/
      # cadence/permission_mode. Never clobbers a populated loop dir. Atomic writes.
      # Next step: fill the config, then `loop-check.sh <dir>/<name>/loop.config.yaml`.
      #
      # Examples:
      #   loop-scaffold.sh --name pr-watch --pattern pr-watch --tier L1
      #   loop-scaffold.sh --name dep-bump --pattern dep-bump --tier L2 --cadence 1d
      #   loop-scaffold.sh --name nightly --cadence "0 3 * * *" --dry-run
      #   loop-scaffold.sh --name digest --pattern digest --host cloud-routine --cadence 1h
      set -uo pipefail
      
      readonly EX_OK=0 EX_USAGE=2 EX_NOTFOUND=3 EX_PRECOND=5
      
      # Terminal design system (skills/_lib/term.sh). stdout = the created path (data);
      # the creation panel frames on stderr, so detect color on fd 2. Degrade to plain
      # stderr lines if the shared lib is unreachable.
      __lib="$(cd "$(dirname "${BASH_SOURCE[0]}")/../../_lib" 2>/dev/null && pwd || true)"
      if [ -n "${__lib:-}" ] && [ -f "$__lib/term.sh" ]; then . "$__lib/term.sh"; term_init 2
      else
        term_panel_open() { :; }; term_panel_close() { :; }; term_panel_vert() { :; }
        term_status_row() { shift; printf '  - %s %s\n' "$1" "${2:-}"; }
        term_alert() { shift; printf '  ! %s\n' "$*"; }
        term_color() { shift; printf '%s' "$*"; }; TERM_DOT="|"
      fi
      
      HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
      ASSETS="$HERE/../assets"
      CFG_TPL="$ASSETS/loop.config.template.yaml"
      STATE_TPL="$ASSETS/STATE.template.md"
      RUN_TPL="$ASSETS/run.template.md"
      RUN_SH_TPL="$ASSETS/run.sh.template"
      
      # ── defaults ────────────────────────────────────────────────────────────────
      NAME=""
      PATTERN="custom"
      TIER="L1"
      CADENCE="1h"
      HOST="local"
      DIR=".loops"
      DRY_RUN=0
      FORCE=0
      
      usage() {
        cat <<'EOF'
      loop-scaffold.sh — scaffold an outer-loop state spine.
      
      Usage:
        loop-scaffold.sh --name NAME [OPTIONS]
      
      Options:
        --name NAME        loop identifier, kebab-case (required). Names the directory.
        --pattern KEY      catalog key (pr-watch, ci-watch, dep-bump,
                           changelog-gen, merge-hygiene, issue-sort,
                           daily-scan) or "custom" (default: custom).
        --tier L1|L2|L3    starting autonomy tier (default: L1).
        --cadence STR      10m | 1h | 6h | 1d, or a cron string (default: 1h).
        --host HOST        where ticks execute: local (default) | session-cron |
                           desktop-task | cloud-routine | external. Decides which
                           constraints loop-doctor enforces (references/native-scheduling.md).
        --dir DIR          parent directory for the loop (default: .loops).
        --dry-run          print the target path + rendered config; write nothing.
        --force            overwrite an already-populated <dir>/<name>/ directory.
        -h, --help         show this help and exit 0.
      
      Exit codes:
        0 created (or dry-run)   2 usage   3 template/dir not found   5 dir populated
      
      Examples:
        loop-scaffold.sh --name pr-watch --pattern pr-watch --tier L1
        loop-scaffold.sh --name dep-bump --pattern dep-bump --tier L2 --cadence 1d
        loop-scaffold.sh --name nightly --cadence "0 3 * * *" --dry-run
        loop-scaffold.sh --name digest --pattern digest --host cloud-routine --cadence 1h
      EOF
      }
      
      die_usage() { printf 'error: %s\n' "$1" >&2; echo >&2; usage >&2; exit "$EX_USAGE"; }
      
      # ── parse args ──────────────────────────────────────────────────────────────
      while [[ $# -gt 0 ]]; do
        case "$1" in
          --name)    [[ $# -ge 2 ]] || die_usage "--name needs a value"; NAME="$2"; shift 2 ;;
          --pattern) [[ $# -ge 2 ]] || die_usage "--pattern needs a value"; PATTERN="$2"; shift 2 ;;
          --tier)    [[ $# -ge 2 ]] || die_usage "--tier needs a value"; TIER="$2"; shift 2 ;;
          --cadence) [[ $# -ge 2 ]] || die_usage "--cadence needs a value"; CADENCE="$2"; shift 2 ;;
          --host)    [[ $# -ge 2 ]] || die_usage "--host needs a value"; HOST="$2"; shift 2 ;;
          --dir)     [[ $# -ge 2 ]] || die_usage "--dir needs a value"; DIR="$2"; shift 2 ;;
          --dry-run) DRY_RUN=1; shift ;;
          --force)   FORCE=1; shift ;;
          -h|--help) usage; exit "$EX_OK" ;;
          -*)        die_usage "unknown flag: $1" ;;
          *)         die_usage "unexpected positional argument: $1" ;;
        esac
      done
      
      # ── validate ────────────────────────────────────────────────────────────────
      [[ -n "$NAME" ]] || die_usage "--name is required"
      [[ "$NAME" =~ ^[a-z0-9]+(-[a-z0-9]+)*$ ]] || die_usage "--name must be kebab-case (got '$NAME')"
      [[ "$PATTERN" =~ ^[a-z0-9]+(-[a-z0-9]+)*$ ]] || die_usage "--pattern must be kebab-case (got '$PATTERN')"
      case "$TIER" in L1|L2|L3) ;; *) die_usage "--tier must be L1|L2|L3 (got '$TIER')" ;; esac
      # cadence: Nm/Nh/Nd OR a cron-ish string (digits, spaces, * / , -)
      [[ "$CADENCE" =~ ^[0-9]+[mhd]$ || "$CADENCE" =~ ^[-0-9*/,\ ]+$ ]] \
        || die_usage "--cadence must be like 10m/1h/1d or a cron string (got '$CADENCE')"
      # host: the execution surface. Its constraints are enforced by loop-doctor, not here -
      # scaffolding a not-yet-valid combination is fine; scheduling one is not.
      case "$HOST" in
        local|session-cron|desktop-task|cloud-routine|external) ;;
        *) die_usage "--host must be local|session-cron|desktop-task|cloud-routine|external (got '$HOST')" ;;
      esac
      
      [[ -f "$CFG_TPL" ]]   || { printf 'error: config template not found at %s\n' "$CFG_TPL" >&2; exit "$EX_NOTFOUND"; }
      [[ -f "$STATE_TPL" ]] || { printf 'error: STATE template not found at %s\n' "$STATE_TPL" >&2; exit "$EX_NOTFOUND"; }
      [[ -f "$RUN_TPL" ]]   || { printf 'error: run template not found at %s\n' "$RUN_TPL" >&2; exit "$EX_NOTFOUND"; }
      [[ -f "$RUN_SH_TPL" ]] || { printf 'error: run.sh template not found at %s\n' "$RUN_SH_TPL" >&2; exit "$EX_NOTFOUND"; }
      
      # Default permission_mode from tier (the workhorse mapping; see references/risk-tiers.md).
      case "$TIER" in
        L1|L2) PMODE="dontAsk" ;;
        L3)    PMODE="bypassPermissions" ;;
      esac
      
      # ── pattern presets ─────────────────────────────────────────────────────────
      # Seed a near-ready config for a known --pattern (the user reviews, doesn't start
      # from blank placeholders). Doctrine: always scaffold at the chosen tier; report/
      # propose/draft patterns carry no gate (VERIFY_SEED empty), code-changing ones do.
      SEEDED=0; SCOPE_SEED=""; GOAL_SEED=""; ESCAL_SEED=""; VERIFY_SEED=""; GUARD_SEED=""; BUDGET_SEED=""
      case "$PATTERN" in
        daily-scan) SEEDED=1
          SCOPE_SEED="src/**"
          GOAL_SEED="Sweep the backlog/issues/alerts and write the day's STATE.md priority list; report only."
          ESCAL_SEED="everything - a human decides what to action; this loop never changes code" ;;
        pr-watch) SEEDED=1
          SCOPE_SEED="src/**"
          GOAL_SEED="Watch open PRs; flag stuck/failing/conflicted; post a summary comment at most; never merge."
          ESCAL_SEED="a human reviews and merges; never merge to main" ;;
        ci-watch) SEEDED=1
          SCOPE_SEED="src/**"
          GOAL_SEED="Detect red CI; classify the failure; at L2 propose a fix in a worktree; never auto-merge to main."
          ESCAL_SEED="flaky/infra failures, anything touching deploy/secrets, ambiguous root cause"
          VERIFY_SEED="npm test"; GUARD_SEED="npm run typecheck" ;;
        dep-bump) SEEDED=1
          SCOPE_SEED="package.json"
          GOAL_SEED="Patch-only dependency bumps behind the release cooldown + guard; open a PR; never minor/major."
          ESCAL_SEED="minor/major bumps, guard failures, any flagged advisory"
          VERIFY_SEED="npm test"; GUARD_SEED="npm run build && npm test" ;;
        changelog-gen) SEEDED=1
          SCOPE_SEED="CHANGELOG.md"
          GOAL_SEED="Summarize merged PRs since the last tag into RELEASE_NOTES_DRAFT.md; never publish a release."
          ESCAL_SEED="the human edits and publishes; never run gh release create" ;;
        merge-hygiene) SEEDED=1
          SCOPE_SEED="src/**"
          GOAL_SEED="Find merged-deletable branches / stale flags / orphaned artifacts; report; never delete unmerged work."
          ESCAL_SEED="anything ambiguous; never delete a branch with unmerged commits" ;;
        issue-sort) SEEDED=1
          SCOPE_SEED="src/**"
          GOAL_SEED="Classify new issues and suggest labels + priority; propose only; never close or set priority unattended."
          ESCAL_SEED="priority calls, dupe-closing, anything needing product judgment" ;;
        metric-chase) SEEDED=1
          SCOPE_SEED="src/**"
          GOAL_SEED="Drive a measurable target (coverage/latency/bundle/eval score) to goal via iterate; keep gains, discard regressions."
          ESCAL_SEED="target unreachable after the budget, guard failures, any change to the gate/test itself"
          VERIFY_SEED="npm test -- --coverage"; GUARD_SEED="npm run typecheck"
          BUDGET_SEED=400000 ;;   # iterate fan-out is the most expensive tick — fit it
        regression-watch) SEEDED=1
          SCOPE_SEED="bench/**"
          GOAL_SEED="Run the benchmark/eval suite, diff against the recorded baseline; report a regression; never edit the suite."
          ESCAL_SEED="a confirmed regression (a human triages); a single flaky run is advisory, not a page" ;;
        digest) SEEDED=1
          SCOPE_SEED="reports/**"
          GOAL_SEED="Summarize email/Asana/calendar/news via connectors into a morning report; read-only, never act."
          ESCAL_SEED="anything requiring a reply or an action; this loop only summarizes" ;;
        backfill) SEEDED=1
          SCOPE_SEED="src/**"
          GOAL_SEED="Drain a migration/queue to completion via /goal; one item per step, verify each; stop when empty or after the bound."
          ESCAL_SEED="any item needing a judgment call; never exceed the stop-after-N / token bound"
          VERIFY_SEED="npm test"; GUARD_SEED="npm run typecheck" ;;
        monitor) SEEDED=1
          SCOPE_SEED="src/**"
          GOAL_SEED="React to an error/log/deploy event (via a Channel); triage it and page a human on a real anomaly; never auto-remediate prod."
          ESCAL_SEED="any anomaly worth a human; production remediation; anything destructive" ;;
        freshness) SEEDED=1
          SCOPE_SEED="docs/**"
          GOAL_SEED="Re-check docs/data/deps/links against reality on a cadence; report confirmed drift; never auto-edit on a transient failure."
          ESCAL_SEED="confirmed drift a human should fix; a transient/network failure is advisory only" ;;
      esac
      
      TARGET_DIR="$DIR/$NAME"
      CFG_OUT="$TARGET_DIR/loop.config.yaml"
      STATE_OUT="$TARGET_DIR/STATE.md"
      LOG_OUT="$TARGET_DIR/run-log.md"
      RUN_OUT="$TARGET_DIR/run.md"
      RUN_SH_OUT="$TARGET_DIR/loop-run.sh"
      
      # Refuse a populated target unless --force.
      if [[ -d "$TARGET_DIR" ]] && [[ -n "$(ls -A "$TARGET_DIR" 2>/dev/null)" ]] && [[ "$FORCE" -ne 1 ]]; then
        printf 'error: loop directory already populated: %s (use --force to overwrite)\n' "$TARGET_DIR" >&2
        exit "$EX_PRECOND"
      fi
      
      NOW="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
      
      # ── render config from template ─────────────────────────────────────────────
      # Line-anchored sed substitutions: identity placeholders globally, the three
      # tunable scalar lines by their default value. Kill-switch path carries <loop-name>.
      render_config() {
        sed -E \
          -e "s|<loop-name>|$NAME|g" \
          -e "s|<pattern-key>|$PATTERN|" \
          -e "s|^tier: L1|tier: $TIER|" \
          -e "s|^cadence: 1h|cadence: $CADENCE|" \
          -e "s|^host: local|host: $HOST|" \
          -e "s|^permission_mode: dontAsk|permission_mode: $PMODE|" \
          "$CFG_TPL"
      }
      
      render_state() {
        sed -E \
          -e "s|<loop-name>|$NAME|g" \
          -e "s|<ISO-8601 Z>|$NOW|" \
          "$STATE_TPL"
      }
      
      render_log() {
        cat <<EOF
      # $NAME — run log (append-only; one line per run)
      # format: <ISO-Z>  run#N  action=<reported|proposed|none>  <key=val…>  outcome=<…>  tokens=<N>
      EOF
      }
      
      render_run() {
        sed -E \
          -e "s|<loop-name>|$NAME|g" \
          -e "s|tier <L1\\|L2\\|L3>|tier $TIER|g" \
          "$RUN_TPL"
      }
      
      # The runner-agnostic tick wrapper any scheduler invokes (cron / Task Scheduler /
      # systemd / process-compose / by hand) — no GitHub Actions required.
      render_run_sh() {
        sed -E \
          -e "s|<loop-name>|$NAME|g" \
          -e "s|<permission-mode>|$PMODE|g" \
          "$RUN_SH_TPL"
      }
      
      # Seeded config for a known pattern. L1 stays report-only (gate fields are a
      # commented graduation block); L2/L3 emit verify/guard/worktree/land_via — using
      # the pattern's gate if it has one, else a <fill:…> placeholder the audit will flag.
      render_seeded_config() {
        cat <<EOF
      # loop.config.yaml - $PATTERN (seeded by loop-scaffold at $TIER; REVIEW before scheduling)
      # Full field semantics: skills/loop-ops/references/state-spine.md
      name: $NAME
      pattern: $PATTERN
      tier: $TIER
      permission_mode: $PMODE
      cadence: $CADENCE
      host: $HOST
      goal: "$GOAL_SEED"
      scope:
        - "$SCOPE_SEED"
      escalation: "$ESCAL_SEED"
      budget_tokens: ${BUDGET_SEED:-200000}
      kill_switch: ".loops/$NAME/PAUSED exists, OR the loop-pause label is set"
      EOF
        if [[ "$TIER" == "L1" ]]; then
          cat <<EOF
      
      # ── graduate to L2 (assisted): set tier: L2, uncomment + fill, re-run loop-check + loop-doctor --live ──
      # verify: "${VERIFY_SEED:-<fill: the gate command, e.g. npm test>}"
      # guard: "${GUARD_SEED:-<fill: a must-always-pass command>}"
      # worktree: true
      # land_via: fleet-ops
      EOF
        else
          cat <<EOF
      verify: "${VERIFY_SEED:-<fill: the gate command for this loop>}"
      guard: "${GUARD_SEED:-<fill: a must-always-pass command>}"
      worktree: true
      land_via: fleet-ops
      EOF
        fi
      }
      
      # Pick the seeded renderer for a known pattern, else the generic template.
      emit_config() { if [[ "$SEEDED" -eq 1 ]]; then render_seeded_config; else render_config; fi; }
      
      # ── dry-run: print and stop ─────────────────────────────────────────────────
      if [[ "$DRY_RUN" -eq 1 ]]; then
        printf '%s\n' "$CFG_OUT"
        {
          term_panel_open loop "loop ${TERM_DOT} init (dry-run)" "$NAME"
          term_panel_vert
          term_status_row skip "would create  $TARGET_DIR/" "tier $TIER ${TERM_DOT} $PATTERN ${TERM_DOT} $CADENCE"
          term_status_row skip "  loop.config.yaml" "permission_mode: $PMODE"
          term_status_row skip "  STATE.md / run-log.md / run.md / loop-run.sh" ""
          term_panel_vert
          term_panel_close "nothing written" ""
        } >&2
        emit_config
        exit "$EX_OK"
      fi
      
      # ── atomic writes ───────────────────────────────────────────────────────────
      mkdir -p "$TARGET_DIR" || { printf 'error: could not create %s\n' "$TARGET_DIR" >&2; exit 1; }
      
      write_atomic() {  # write_atomic <dest> <content>
        local dest="$1" content="$2" tmp
        tmp="$dest.tmp.$$"
        printf '%s\n' "$content" > "$tmp" || { printf 'error: failed to write %s\n' "$tmp" >&2; exit 1; }
        mv -f "$tmp" "$dest" || { rm -f "$tmp"; printf 'error: failed to move into place: %s\n' "$dest" >&2; exit 1; }
      }
      
      write_atomic "$CFG_OUT"   "$(emit_config)"
      write_atomic "$STATE_OUT" "$(render_state)"
      write_atomic "$LOG_OUT"   "$(render_log)"
      write_atomic "$RUN_OUT"   "$(render_run)"
      write_atomic "$RUN_SH_OUT" "$(render_run_sh)"
      chmod +x "$RUN_SH_OUT" 2>/dev/null || true
      
      printf '%s\n' "$CFG_OUT"
      
      {
        term_panel_open loop "loop ${TERM_DOT} init" "$NAME"
        term_panel_vert
        term_status_row ok "created  $TARGET_DIR/" "tier $TIER ${TERM_DOT} $PATTERN ${TERM_DOT} $CADENCE"
        term_status_row ok "  loop.config.yaml" "permission_mode: $PMODE"
        term_status_row ok "  STATE.md / run-log.md / run.md / loop-run.sh" ""
        if [[ "$TIER" != "L1" ]]; then
          term_alert warning "tier $TIER needs a verify gate, guard, worktree, escalation + land_via — fill them before auditing"
        fi
        term_panel_vert
        term_panel_close "then: fill the config ${TERM_DOT} loop-check.sh $CFG_OUT" ""
      } >&2
      
      exit "$EX_OK"
      
  • tests
    • run.sh 32 KB
      #!/usr/bin/env bash
      # Self-test for loop-ops scripts (loop-scaffold.sh, loop-check.sh, loop-estimate.py).
      #
      # Offline-deterministic (no network). Scaffolds throwaway loop fixtures, asserts the
      # documented exit codes + key output of each script, then cleans up. Resolves paths
      # relative to itself so it works both in the repo and installed to ~/.claude/.
      #
      # Usage:   bash tests/run.sh
      # Exit:    0 all pass, 1 one or more failures
      
      set -uo pipefail
      
      HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
      SKILL="$(dirname "$HERE")"
      SCRIPTS="$SKILL/scripts"
      INIT="$SCRIPTS/loop-scaffold.sh"
      AUDIT="$SCRIPTS/loop-check.sh"
      COST="$SCRIPTS/loop-estimate.py"
      SYNC="$SCRIPTS/check-pricing-sync.py"
      DOCTOR="$SCRIPTS/loop-doctor.sh"
      
      # Pick a python that actually executes — skips the Windows Store python3 stub.
      PYTHON=""
      for c in python python3 py; do
        if command -v "$c" >/dev/null 2>&1 && "$c" -c "" >/dev/null 2>&1; then PYTHON="$c"; break; fi
      done
      [[ -z "$PYTHON" ]] && { echo "no working python found — skipping" >&2; exit 0; }
      
      SB="$(mktemp -d)"; trap 'rm -rf "$SB"' EXIT
      
      PASS=0; FAIL=0
      ok() { PASS=$((PASS+1)); printf '  PASS  %s\n' "$1"; }
      no() { FAIL=$((FAIL+1)); printf '  FAIL  %s\n' "$1"; }
      expect_exit() { [[ "$2" == "$3" ]] && ok "$1 (exit $3)" || no "$1 (want $2 got $3)"; }
      expect_has()  { case "$3" in *"$2"*) ok "$1";; *) no "$1 (missing '$2')";; esac; }
      
      # Write a filled, READY L1 report-only config.
      good_l1() { cat > "$1" <<'EOF'
      name: test-l1
      pattern: pr-watch
      tier: L1
      permission_mode: dontAsk
      cadence: 10m
      goal: "Watch open PRs and report; never merge."
      scope:
        - "src/**"
      escalation: "comment on the PR; never merge to main"
      budget_tokens: 200000
      kill_switch: ".loops/test-l1/PAUSED exists or loop-pause label"
      EOF
      }
      
      # Write a filled, READY L2 assisted config.
      good_l2() { cat > "$1" <<'EOF'
      name: dep-bump
      pattern: dep-bump
      tier: L2
      permission_mode: dontAsk
      cadence: 1d
      goal: "Patch-only dependency bumps behind cooldown; open a PR."
      scope:
        - "package.json"
        - "package-lock.json"
      verify: "npm test"
      guard: "npm run typecheck"
      worktree: true
      land_via: fleet-ops
      escalation: "minor/major bumps escalate; never merge to main"
      budget_tokens: 300000
      kill_switch: ".loops/dep-bump/PAUSED"
      EOF
      }
      
      echo "=== loop-ops self-test (python: $PYTHON) ==="
      
      # ── --help contracts (exit 0) ──────────────────────────────────────────────
      echo "-- --help --"
      bash "$INIT"  --help >/dev/null 2>&1; expect_exit "loop-scaffold --help" 0 $?
      bash "$AUDIT" --help >/dev/null 2>&1; expect_exit "loop-check --help" 0 $?
      "$PYTHON" "$COST" --help >/dev/null 2>&1; expect_exit "loop-estimate --help" 0 $?
      
      # ── loop-scaffold: scaffolds dir + 3 files, substitutes fields ─────────────────
      echo "-- loop-scaffold --"
      out="$(bash "$INIT" --name pr-watch --pattern pr-watch --tier L1 --cadence 5m --dir "$SB/loops" 2>/dev/null)"; rc=$?
      expect_exit "loop-scaffold -> 0" 0 "$rc"
      expect_has  "prints the config path" "pr-watch/loop.config.yaml" "$out"
      [[ -f "$SB/loops/pr-watch/loop.config.yaml" ]] && ok "wrote loop.config.yaml" || no "no loop.config.yaml"
      [[ -f "$SB/loops/pr-watch/STATE.md" ]] && ok "wrote STATE.md" || no "no STATE.md"
      [[ -f "$SB/loops/pr-watch/run-log.md" ]] && ok "wrote run-log.md" || no "no run-log.md"
      [[ -f "$SB/loops/pr-watch/run.md" ]] && ok "wrote run.md" || no "no run.md"
      runmd="$(cat "$SB/loops/pr-watch/run.md")"
      expect_has "run.md substitutes loop name" "Run: pr-watch" "$runmd"
      expect_has "run.md substitutes tier" "tier L1)" "$runmd"
      # runner-agnostic wrapper: emitted, executable, fully substituted, no GH Actions dep
      [[ -f "$SB/loops/pr-watch/loop-run.sh" ]] && ok "wrote loop-run.sh" || no "no loop-run.sh"
      runsh="$(cat "$SB/loops/pr-watch/loop-run.sh")"
      case "$runsh" in *"<loop-name>"*|*"<permission-mode>"*) no "loop-run.sh left a placeholder";; *) ok "loop-run.sh fully substituted";; esac
      expect_has "loop-run.sh wires the gated mode" "--permission-mode dontAsk" "$runsh"
      cfg="$(cat "$SB/loops/pr-watch/loop.config.yaml")"
      expect_has "substituted name" "name: pr-watch" "$cfg"
      expect_has "substituted tier" "tier: L1" "$cfg"
      expect_has "substituted cadence" "cadence: 5m" "$cfg"
      expect_has "L1 default permission_mode" "permission_mode: dontAsk" "$cfg"
      # L3 default permission_mode is bypassPermissions
      bash "$INIT" --name big-job --tier L3 --dir "$SB/loops" >/dev/null 2>&1
      expect_has "L3 default permission_mode" "permission_mode: bypassPermissions" "$(cat "$SB/loops/big-job/loop.config.yaml")"
      
      # ── loop-scaffold: refuses a populated dir -> 5, --force overwrites ─────────────
      bash "$INIT" --name pr-watch --dir "$SB/loops" >/dev/null 2>&1; expect_exit "refuse populated dir -> 5" 5 $?
      bash "$INIT" --name pr-watch --dir "$SB/loops" --force >/dev/null 2>&1; expect_exit "--force overwrites -> 0" 0 $?
      
      # ── loop-scaffold: --dry-run writes nothing ────────────────────────────────────
      out="$(bash "$INIT" --name ghost --dir "$SB/dryloops" --dry-run 2>/dev/null)"; rc=$?
      expect_exit "dry-run -> 0" 0 "$rc"
      [[ -e "$SB/dryloops" ]] && no "dry-run created files" || ok "dry-run wrote nothing"
      expect_has "dry-run prints config path" "ghost/loop.config.yaml" "$out"
      
      # ── loop-scaffold: usage errors ────────────────────────────────────────────────
      bash "$INIT" --dir "$SB/loops" >/dev/null 2>&1; expect_exit "missing --name -> 2" 2 $?
      bash "$INIT" --name BadName --dir "$SB/loops" >/dev/null 2>&1; expect_exit "non-kebab name -> 2" 2 $?
      bash "$INIT" --name x --tier L9 --dir "$SB/loops" >/dev/null 2>&1; expect_exit "bad tier -> 2" 2 $?
      
      # pattern-seeding: a known pattern seeds a near-ready, audit-clean config
      bash "$INIT" --name seed-l1 --pattern ci-watch --tier L1 --cadence 15m --dir "$SB/seed" >/dev/null 2>&1
      seedcfg="$(cat "$SB/seed/seed-l1/loop.config.yaml")"
      expect_has "seeded config carries the pattern goal" "Detect red CI" "$seedcfg"
      expect_has "seeded L1 leaves a graduation block" "graduate to L2" "$seedcfg"
      bash "$AUDIT" "$SB/seed/seed-l1/loop.config.yaml" >/dev/null 2>&1; expect_exit "seeded L1 audits clean -> 0" 0 $?
      # at L2 the pattern's gate is filled (not commented) and audits clean
      bash "$INIT" --name seed-l2 --pattern ci-watch --tier L2 --cadence 15m --dir "$SB/seed" >/dev/null 2>&1
      l2cfg="$(cat "$SB/seed/seed-l2/loop.config.yaml")"
      case "$l2cfg" in *$'\nverify: "npm test"'*) ok "seeded L2 fills the gate";; *) no "seeded L2 did not fill the gate";; esac
      bash "$AUDIT" "$SB/seed/seed-l2/loop.config.yaml" >/dev/null 2>&1; expect_exit "seeded L2 audits clean -> 0" 0 $?
      # an unknown pattern falls back to the generic placeholder template (not ready)
      bash "$INIT" --name seed-x --pattern custom --tier L1 --dir "$SB/seed" >/dev/null 2>&1
      case "$(cat "$SB/seed/seed-x/loop.config.yaml")" in *"<one sentence"*) ok "unknown pattern uses generic template";; *) no "unknown pattern did not use template";; esac
      # v2 archetypes: scaffold must be audit-clean AND doctor-clean at L1 + known to the cost model
      # (doctor-clean catches budget < tokens/run — the metric-chase trap: it seeds a bigger budget)
      for p in metric-chase regression-watch digest backfill monitor freshness; do
        bash "$INIT" --name "a-$p" --pattern "$p" --tier L1 --dir "$SB/arch" >/dev/null 2>&1
        bash "$AUDIT" "$SB/arch/a-$p/loop.config.yaml" >/dev/null 2>&1; expect_exit "archetype $p seeds audit-clean (L1)" 0 $?
        bash "$DOCTOR" --offline "$SB/arch/a-$p/loop.config.yaml" >/dev/null 2>&1; expect_exit "archetype $p doctors clean (L1)" 0 $?
        "$PYTHON" "$COST" --pattern "$p" --cadence 1h --model claude-haiku-4-5 >/dev/null 2>&1; expect_exit "cost model knows $p" 0 $?
      done
      # the most expensive archetype at L2: gate filled, budget fits the tick (audit + doctor clean)
      bash "$INIT" --name a-mc --pattern metric-chase --tier L2 --cadence 1h --dir "$SB/arch" >/dev/null 2>&1
      bash "$AUDIT" "$SB/arch/a-mc/loop.config.yaml" >/dev/null 2>&1; expect_exit "metric-chase L2 audits clean -> 0" 0 $?
      bash "$DOCTOR" --offline "$SB/arch/a-mc/loop.config.yaml" >/dev/null 2>&1; expect_exit "metric-chase L2 doctors clean (budget fits) -> 0" 0 $?
      
      # ── loop-scaffold: --host records the execution surface ────────────────────────
      # `host:` selects which constraints loop-doctor enforces (references/native-scheduling.md).
      echo "-- loop-scaffold --host --"
      bash "$INIT" --name h-default --dir "$SB/hosts" >/dev/null 2>&1
      expect_has "host defaults to local" "host: local" "$(cat "$SB/hosts/h-default/loop.config.yaml")"
      bash "$INIT" --name h-cloud --host cloud-routine --dir "$SB/hosts" >/dev/null 2>&1
      expect_has "template render carries --host" "host: cloud-routine" "$(cat "$SB/hosts/h-cloud/loop.config.yaml")"
      # the seeded (known --pattern) path is a separate renderer - it must carry host too
      bash "$INIT" --name h-seed --pattern digest --host desktop-task --dir "$SB/hosts" >/dev/null 2>&1
      expect_has "seeded render carries --host" "host: desktop-task" "$(cat "$SB/hosts/h-seed/loop.config.yaml")"
      bash "$INIT" --name h-bad --host nonsense --dir "$SB/hosts" >/dev/null 2>&1; expect_exit "unknown --host -> 2" 2 $?
      
      # ── loop-check: a freshly-init'd config is NOT ready (placeholders) -> 10 ───
      echo "-- loop-check --"
      bash "$INIT" --name raw --pattern custom --tier L1 --dir "$SB/loops" >/dev/null 2>&1
      out="$(bash "$AUDIT" "$SB/loops/raw/loop.config.yaml" 2>/dev/null)"; rc=$?
      expect_exit "raw scaffold not ready -> 10" 10 "$rc"
      expect_has  "flags the goal placeholder" "goal:" "$out"
      
      # ── loop-check: filled L1 config is READY -> 0 ─────────────────────────────
      good_l1 "$SB/l1.yaml"
      out="$(bash "$AUDIT" "$SB/l1.yaml" 2>/dev/null)"; rc=$?
      expect_exit "filled L1 ready -> 0" 0 "$rc"
      
      # ── loop-check: filled L2 config is READY -> 0 ─────────────────────────────
      good_l2 "$SB/l2.yaml"
      bash "$AUDIT" "$SB/l2.yaml" >/dev/null 2>&1; expect_exit "filled L2 ready -> 0" 0 $?
      
      # ── loop-check: L2 missing the gate -> 10, names verify ────────────────────
      grep -v '^verify:' "$SB/l2.yaml" > "$SB/l2-nogate.yaml"
      out="$(bash "$AUDIT" "$SB/l2-nogate.yaml" 2>/dev/null)"; rc=$?
      expect_exit "L2 missing gate -> 10" 10 "$rc"
      expect_has  "names the missing gate" "verify:" "$out"
      
      # ── loop-check: unbounded scope -> 10 ──────────────────────────────────────
      sed 's|  - "src/\*\*"|  - "*"|' "$SB/l1.yaml" > "$SB/l1-unbounded.yaml"
      out="$(bash "$AUDIT" "$SB/l1-unbounded.yaml" 2>/dev/null)"; rc=$?
      expect_exit "unbounded scope -> 10" 10 "$rc"
      expect_has  "names unbounded scope" "unbounded" "$out"
      
      # ── loop-check: missing escalation -> 10 ───────────────────────────────────
      grep -v '^escalation:' "$SB/l1.yaml" > "$SB/l1-noescal.yaml"
      out="$(bash "$AUDIT" "$SB/l1-noescal.yaml" 2>/dev/null)"; rc=$?
      expect_exit "missing escalation -> 10" 10 "$rc"
      expect_has  "names escalation" "escalation:" "$out"
      
      # ── loop-check: missing file -> 3, unparseable -> 4, bad --min -> 2 ────────
      bash "$AUDIT" "$SB/no-such.yaml" >/dev/null 2>&1; expect_exit "missing config -> 3" 3 $?
      printf 'just some prose, no keys\n' > "$SB/garbage.yaml"
      bash "$AUDIT" "$SB/garbage.yaml" >/dev/null 2>&1; expect_exit "unparseable -> 4" 4 $?
      bash "$AUDIT" --min abc "$SB/l1.yaml" >/dev/null 2>&1; expect_exit "bad --min -> 2" 2 $?
      
      # ── loop-check: --json envelope schema + ready flag ────────────────────────
      out="$(bash "$AUDIT" --json "$SB/l1.yaml" 2>/dev/null)"
      expect_has "audit json schema" "claude-mods.loop-ops.check/v1" "$out"
      expect_has "audit json ready true" '"ready": true' "$out"
      out="$(bash "$AUDIT" --json "$SB/l2-nogate.yaml" 2>/dev/null)"
      expect_has "audit json ready false" '"ready": false' "$out"
      
      # ── loop-check: --strict turns a warning into NOT ready ────────────────────
      # An L1 with permission_mode: auto is consistent-enough to pass errors but warns
      # (broad for L1). Normally ready; --strict flips it.
      sed 's|permission_mode: dontAsk|permission_mode: auto|' "$SB/l1.yaml" > "$SB/l1-warn.yaml"
      bash "$AUDIT" "$SB/l1-warn.yaml" >/dev/null 2>&1; expect_exit "warning, normally ready -> 0" 0 $?
      bash "$AUDIT" --strict "$SB/l1-warn.yaml" >/dev/null 2>&1; expect_exit "warning, --strict not ready -> 10" 10 $?
      
      # ── loop-estimate: basic run, --json, --list-models, cadence forms ─────────────
      echo "-- loop-estimate --"
      out="$("$PYTHON" "$COST" --pattern pr-watch --cadence 10m --model claude-haiku-4-5 2>/dev/null)"; rc=$?
      expect_exit "loop-estimate -> 0" 0 "$rc"
      expect_has  "prints a daily cost" "cost/day:" "$out"
      expect_has  "derives runs/day from 10m" "144 runs/day" "$out"
      out="$("$PYTHON" "$COST" --pattern ci-watch --cadence 15m --model claude-sonnet-5 --json 2>/dev/null)"
      expect_has "cost json schema" "claude-mods.loop-ops.estimate/v1" "$out"
      expect_has "cost json carries runs_per_day" "runs_per_day" "$out"
      out="$("$PYTHON" "$COST" --list-models 2>/dev/null)"; rc=$?
      expect_exit "list-models -> 0" 0 "$rc"
      expect_has  "list-models shows a model" "claude-opus-5" "$out"
      # cron cadence parses
      "$PYTHON" "$COST" --pattern daily-scan --cadence '*/10 * * * *' --model claude-haiku-4-5 >/dev/null 2>&1
      expect_exit "cron cadence -> 0" 0 $?
      # --runs-per-day override
      out="$("$PYTHON" "$COST" --pattern custom --cadence weird --runs-per-day 5 --model claude-haiku-4-5 2>/dev/null)"; rc=$?
      expect_exit "runs-per-day override -> 0" 0 "$rc"
      expect_has  "uses the override" "5 runs/day" "$out"
      # caching: a fast loop (10m -> 1h TTL) projects a cached saving
      out="$("$PYTHON" "$COST" --pattern ci-watch --cadence 10m --model claude-sonnet-5 2>&1)"
      expect_has "fast loop shows a cached projection" "cached/" "$out"
      # caching: a slow loop (6h > 1h TTL) is not cache-beneficial
      out="$("$PYTHON" "$COST" --pattern daily-scan --cadence 6h --model claude-opus-5 2>&1)"
      expect_has "slow loop: caching not beneficial" "not beneficial" "$out"
      # --no-cache suppresses the cached projection
      out="$("$PYTHON" "$COST" --pattern ci-watch --cadence 10m --model claude-sonnet-5 --no-cache 2>&1)"
      case "$out" in *"cached/"*) no "--no-cache still showed caching";; *) ok "--no-cache suppresses caching";; esac
      # json caching block present for a cacheable loop
      out="$("$PYTHON" "$COST" --pattern ci-watch --cadence 5m --model claude-sonnet-5 --json 2>/dev/null)"
      expect_has "cost json carries caching block" '"caching"' "$out"
      
      # ── loop-doctor: preflight (offline budget, live binary), json ─────────────
      echo "-- loop-doctor --"
      bash "$DOCTOR" --help >/dev/null 2>&1; expect_exit "loop-doctor --help -> 0" 0 $?
      bash "$DOCTOR" --offline "$SB/l1.yaml" >/dev/null 2>&1; expect_exit "doctor offline healthy L1 -> 0" 0 $?
      bash "$DOCTOR" --live "$SB/l1.yaml" >/dev/null 2>&1; expect_exit "doctor live healthy L1 -> 0" 0 $?
      # budget too small for the pattern -> bad -> 10
      sed 's/^budget_tokens: 300000/budget_tokens: 100/' "$SB/l2.yaml" > "$SB/l2-poor.yaml"
      out="$(bash "$DOCTOR" --offline "$SB/l2-poor.yaml" 2>/dev/null)"; rc=$?
      expect_exit "doctor budget-too-small -> 10" 10 "$rc"
      expect_has  "doctor names the budget gap" "tokens/run" "$out"
      # live: a verify gate whose binary is missing -> bad -> 10
      sed 's/^verify: "npm test"/verify: "totally-missing-binary-zzz run"/' "$SB/l2.yaml" > "$SB/l2-nobin.yaml"
      bash "$DOCTOR" --live "$SB/l2-nobin.yaml" >/dev/null 2>&1; expect_exit "doctor missing gate binary -> 10" 10 $?
      # missing config -> 3, json schema
      bash "$DOCTOR" --offline "$SB/no-such.yaml" >/dev/null 2>&1; expect_exit "doctor missing config -> 3" 3 $?
      out="$(bash "$DOCTOR" --offline --json "$SB/l1.yaml" 2>/dev/null)"
      expect_has "doctor json schema" "claude-mods.loop-ops.doctor/v1" "$out"
      
      # ── loop-doctor: host-aware checks (native scheduling) ─────────────────────────
      # Each native host has different hard limits, so "will it run" depends on `host:`.
      # Facts asserted here are verified against the live tool schemas + docs (2026-08-30)
      # and written up in references/native-scheduling.md - if a limit changes upstream,
      # these are the assertions that should fail first.
      echo "-- loop-doctor (host-aware) --"
      
      # A config with NO host: must behave exactly as before (host defaults to local).
      grep -v '^host:' "$SB/l1.yaml" > "$SB/l1-nohost.yaml" 2>/dev/null || cp "$SB/l1.yaml" "$SB/l1-nohost.yaml"
      bash "$DOCTOR" --offline "$SB/l1-nohost.yaml" >/dev/null 2>&1; expect_exit "no host: still healthy (defaults local) -> 0" 0 $?
      
      # cloud-routine: minimum interval is 1 hour; faster expressions are rejected at creation.
      sed 's|^kill_switch:.*|kill_switch: "pause the routine (Repeats toggle)"|' "$SB/l1.yaml" > "$SB/cloud.yaml"
      cat >> "$SB/cloud.yaml" <<'EOF'
      host: cloud-routine
      EOF
      sed -i 's|^cadence: 10m|cadence: 10m|' "$SB/cloud.yaml"
      out="$(bash "$DOCTOR" --offline "$SB/cloud.yaml" 2>/dev/null)"; rc=$?
      expect_exit "cloud-routine sub-hour cadence -> 10" 10 "$rc"
      expect_has  "names the 1-hour floor" "1 hour" "$out"
      
      # cloud-routine has NO permission mode, so the boundary must be named instead
      # (repos + environment network policy + connectors). Absent -> a finding.
      sed 's|^cadence: 10m|cadence: 1h|' "$SB/cloud.yaml" > "$SB/cloud-1h.yaml"
      out="$(bash "$DOCTOR" --offline "$SB/cloud-1h.yaml" 2>/dev/null)"; rc=$?
      expect_exit "cloud-routine without a named boundary -> 10" 10 "$rc"
      expect_has  "names the missing boundary" "boundary" "$out"
      
      # With the boundary named it passes, and permission_mode is reported as ignored, not required.
      sed 's|^escalation:.*|escalation: "connectors pruned to GitHub read; environment network Trusted; never merge"|' \
        "$SB/cloud-1h.yaml" > "$SB/cloud-ok.yaml"
      bash "$DOCTOR" --offline "$SB/cloud-ok.yaml" >/dev/null 2>&1; expect_exit "cloud-routine with boundary -> 0" 0 $?
      out="$(bash "$DOCTOR" --offline "$SB/cloud-ok.yaml" 2>/dev/null)"
      expect_has "permission_mode reported as ignored on cloud" "ignored by cloud routines" "$out"
      
      # --live must be SKIPPED (not passed) for a cloud routine: the tick runs on a fresh
      # cloud clone, so this machine's PATH proves nothing. A missing gate binary here must
      # NOT fail, and must not be silently reported as ok either.
      sed 's|^escalation:|verify: "totally-missing-binary-zzz run"\nescalation:|' "$SB/cloud-ok.yaml" > "$SB/cloud-live.yaml"
      out="$(bash "$DOCTOR" --live "$SB/cloud-live.yaml" 2>/dev/null)"; rc=$?
      expect_exit "cloud-routine --live does not fail on local PATH -> 0" 0 "$rc"
      expect_has  "cloud-routine --live is explicitly skipped" "skipped" "$out"
      
      # session-cron (/loop + CronCreate) is session-scoped with a 7-day expiry: L1 only.
      sed 's|^host: cloud-routine|host: session-cron|' "$SB/cloud-ok.yaml" > "$SB/sess-l1.yaml"
      bash "$DOCTOR" --offline "$SB/sess-l1.yaml" >/dev/null 2>&1; expect_exit "session-cron at L1 -> 0 (warns only)" 0 $?
      out="$(bash "$DOCTOR" --offline "$SB/sess-l1.yaml" 2>/dev/null)"
      expect_has "session-cron L1 warns about the 7-day expiry" "7-day expiry" "$out"
      # same host at L2 = unattended, which it cannot do
      { sed 's|^tier: L1|tier: L2|' "$SB/sess-l1.yaml"; printf 'verify: "true"\nguard: "true"\nworktree: true\nland_via: fleet-ops\n'; } > "$SB/sess-l2.yaml"
      out="$(bash "$DOCTOR" --offline "$SB/sess-l2.yaml" 2>/dev/null)"; rc=$?
      expect_exit "session-cron at L2 (unattended) -> 10" 10 "$rc"
      expect_has  "names the unattended mismatch" "can't run unattended" "$out"
      
      # An unknown host is a finding, not a silent pass.
      sed 's|^host: session-cron|host: made-up|' "$SB/sess-l1.yaml" > "$SB/host-bad.yaml"
      bash "$DOCTOR" --offline "$SB/host-bad.yaml" >/dev/null 2>&1; expect_exit "unknown host -> 10" 10 $?
      
      # ── loop-doctor: regressions found by adversarial review ───────────────────
      # Each of these shipped broken once. They are cheap to re-break and silent when broken.
      echo "-- loop-doctor (adversarial regressions) --"
      
      # R1. A cron-string cadence must still hit the cloud floor - the Nm/Nh/Nd path is not
      # the only way to express "every 10 minutes".
      sed 's|^cadence: 1h|cadence: "*/10 * * * *"|' "$SB/cloud-ok.yaml" > "$SB/cloud-cron.yaml"
      out="$(bash "$DOCTOR" --offline "$SB/cloud-cron.yaml" 2>/dev/null)"; rc=$?
      expect_exit "cron-string cadence hits the cloud floor -> 10" 10 "$rc"
      # ...and an hourly cron must NOT be flagged (1h is legal on cloud).
      sed 's|^cadence: 1h|cadence: "0 * * * *"|' "$SB/cloud-ok.yaml" > "$SB/cloud-hourly.yaml"
      bash "$DOCTOR" --offline "$SB/cloud-hourly.yaml" >/dev/null 2>&1; expect_exit "hourly cron is legal on cloud -> 0" 0 $?
      
      # R2. A malformed cadence must not leak a bash arithmetic error. '1a2m' matches the
      # *[0-9]m glob, so without a digits-only guard ${1%m} hands '1a2' to $(( )).
      sed 's|^cadence: 1h|cadence: 1a2m|' "$SB/cloud-ok.yaml" > "$SB/cad-junk.yaml"
      allout="$(bash "$DOCTOR" --offline "$SB/cad-junk.yaml" 2>&1)"
      case "$allout" in
        *"value too great for base"*|*"syntax error"*) no "malformed cadence leaks an arithmetic error" ;;
        *) ok "malformed cadence degrades quietly (no arithmetic leak)" ;;
      esac
      
      # R3. L3 on a cloud routine must NOT demand a local container note - the cloud infra IS
      # the isolation, and its permission_mode is ignored. A false finding here teaches people
      # to write a bogus "container" note to satisfy the tool.
      { sed -e 's|^tier: L1|tier: L3|' -e 's|^permission_mode: dontAsk|permission_mode: bypassPermissions|' "$SB/cloud-ok.yaml"
        printf 'verify: "true"\nguard: "true"\nworktree: true\nland_via: fleet-ops\n'; } > "$SB/cloud-l3.yaml"
      out="$(bash "$DOCTOR" --offline "$SB/cloud-l3.yaml" 2>/dev/null)"
      case "$out" in
        *"bad"*"isolation"*) no "L3 cloud-routine wrongly demands a container note" ;;
        *) ok "L3 cloud-routine does not demand a local container" ;;
      esac
      # The same check must still bite for a LOCAL L3 bypass - the fix must not disarm it.
      { sed -e 's|^tier: L1|tier: L3|' -e 's|^host: cloud-routine|host: local|' \
            -e 's|^permission_mode: dontAsk|permission_mode: bypassPermissions|' "$SB/cloud-ok.yaml"
        printf 'verify: "true"\nguard: "true"\nworktree: true\nland_via: fleet-ops\n'; } > "$SB/local-l3.yaml"
      out="$(bash "$DOCTOR" --offline "$SB/local-l3.yaml" 2>/dev/null)"
      expect_has "local L3 bypass still demands isolation" "isolated VM/container" "$out"
      
      # R4. `worktree: true` on a desktop-task is a declaration, not the switch: the real
      # toggle is per-task and OFF by default, so the doctor must say it cannot enforce it.
      { sed -e 's|^host: cloud-routine|host: desktop-task|' -e 's|^tier: L1|tier: L2|' "$SB/cloud-ok.yaml"
        printf 'verify: "true"\nguard: "true"\nworktree: true\nland_via: fleet-ops\n'; } > "$SB/desk-l2.yaml"
      out="$(bash "$DOCTOR" --offline "$SB/desk-l2.yaml" 2>/dev/null)"
      expect_has "desktop-task flags the out-of-band worktree toggle" "toggle is on" "$out"
      
      # ── loop-estimate: validation errors ───────────────────────────────────────────
      "$PYTHON" "$COST" --pattern pr-watch --cadence 10m --model claude-nope >/dev/null 2>&1; expect_exit "unknown model -> 4" 4 $?
      "$PYTHON" "$COST" --pattern not-a-pattern --cadence 10m --model claude-haiku-4-5 >/dev/null 2>&1; expect_exit "unknown pattern -> 4" 4 $?
      "$PYTHON" "$COST" --pattern pr-watch --cadence "garbage cron" --model claude-haiku-4-5 >/dev/null 2>&1; expect_exit "bad cadence -> 4" 4 $?
      "$PYTHON" "$COST" --pricing "$SB/no-pricing.json" --pattern custom --cadence 1h --input-tokens 1 --output-tokens 1 --model x >/dev/null 2>&1; expect_exit "missing pricing file -> 3" 3 $?
      
      # ── check-pricing-sync: offline clean -> 0, drift -> 10, --json ────────────
      echo "-- check-pricing-sync --"
      "$PYTHON" "$SYNC" --help >/dev/null 2>&1; expect_exit "pricing-sync --help -> 0" 0 $?
      "$PYTHON" "$SYNC" --offline >/dev/null 2>&1; expect_exit "pricing-sync offline in sync -> 0" 0 $?
      # Tamper a copy: opus input price 5.0 -> 999.0 (sed; argv path is MSYS-converted for python).
      sed 's/"input_per_mtok": 5\.0/"input_per_mtok": 999.0/' "$SKILL/assets/model-pricing.json" > "$SB/badprice.json"
      "$PYTHON" "$SYNC" --pricing "$SB/badprice.json" >/dev/null 2>&1; expect_exit "pricing-sync drift -> 10" 10 $?
      "$PYTHON" "$SYNC" --pricing "$SB/no-such.json" >/dev/null 2>&1; expect_exit "pricing-sync missing file -> 3" 3 $?
      out="$("$PYTHON" "$SYNC" --json 2>/dev/null)"
      expect_has "pricing-sync json schema" "claude-mods.loop-ops.pricing-sync/v1" "$out"
      expect_has "pricing-sync json in_sync" '"in_sync": true' "$out"
      
      # ── Windows-authored configs: CRLF + UTF-8 BOM must parse like clean LF ─────
      echo "-- windows-authored configs (CRLF / BOM) --"
      good_l1 "$SB/win.yaml"
      sed 's/$/\r/' "$SB/win.yaml" > "$SB/win-crlf.yaml"                       # LF -> CRLF
      bash "$AUDIT"  "$SB/win-crlf.yaml" >/dev/null 2>&1; expect_exit "CRLF config audits clean -> 0" 0 $?
      bash "$DOCTOR" --offline "$SB/win-crlf.yaml" >/dev/null 2>&1; expect_exit "CRLF config doctors clean -> 0" 0 $?
      printf '\xEF\xBB\xBF' > "$SB/win-bom.yaml"; cat "$SB/win.yaml" >> "$SB/win-bom.yaml"  # prepend BOM
      bash "$AUDIT"  "$SB/win-bom.yaml" >/dev/null 2>&1; expect_exit "BOM config audits clean -> 0" 0 $?
      bash "$DOCTOR" --offline "$SB/win-bom.yaml" >/dev/null 2>&1; expect_exit "BOM config doctors clean -> 0" 0 $?
      
      # ── worked example: the shipped example stays gate-clean ───────────────────
      echo "-- worked example --"
      EX="$SKILL/assets/examples/pr-watch/loop.config.yaml"
      [[ -f "$EX" ]] && ok "worked example present" || no "worked example missing"
      bash "$AUDIT" "$EX" >/dev/null 2>&1; expect_exit "shipped example audits clean -> 0" 0 $?
      bash "$DOCTOR" --offline "$EX" >/dev/null 2>&1; expect_exit "shipped example doctors clean -> 0" 0 $?
      [[ -f "$SKILL/assets/examples/pr-watch/loop-run.sh" ]] && ok "example ships loop-run.sh (runner-agnostic)" || no "example missing loop-run.sh"
      [[ -f "$SKILL/assets/examples/pr-watch/github-actions.yml" ]] && ok "example ships an optional GH Actions scheduler" || no "example missing GH Actions option"
      [[ -f "$SKILL/assets/examples/pr-watch/run.md" ]] && ok "example ships a run prompt" || no "example missing run.md"
      # The flagship example is what people copy, so it must model the doctrine it teaches:
      # a loop declares where it runs before it is scheduled.
      expect_has "example declares a host" "host:" "$(cat "$EX")"
      
      # ── terminal design system ─────────────────────────────────────────────────
      echo "-- terminal design system --"
      for s in "$INIT" "$AUDIT" "$DOCTOR"; do
        b="$(basename "$s")"
        grep -q '_lib/term.sh' "$s" && ok "$b sources _lib/term.sh" || no "$b does not source _lib/term.sh"
      done
      grep -q 'class Term' "$COST" && ok "loop-estimate carries inline Term helper" || no "loop-estimate missing inline Term helper"
      grep -q 'class Term' "$SYNC" && ok "check-pricing-sync carries inline Term helper" || no "check-pricing-sync missing inline Term helper"
      grep -q 'BRAND::loop' "$SKILL/../_lib/term.sh" && ok "term.sh registers the loop brand glyph" || no "term.sh missing loop brand glyph"
      # Piped audit findings stay plain (no ANSI in the data stream).
      po="$(bash "$AUDIT" "$SB/l2-nogate.yaml" 2>/dev/null)"
      case "$po" in *$'\033'*) no "piped audit leaked ANSI into data";; *) ok "piped audit stays plain data";; esac
      
      # ── check-native-facts: the native-scheduling staleness guard ──────────────
      # Asserts the guard actually detects drift, not just that it exits 0 today: a
      # verifier that can only pass is decoration.
      echo "-- check-native-facts --"
      NATIVE_CHK="$SCRIPTS/check-native-facts.py"
      "$PYTHON" "$NATIVE_CHK" --help >/dev/null 2>&1; expect_exit "native-facts --help -> 0" 0 $?
      "$PYTHON" "$NATIVE_CHK" --offline >/dev/null 2>&1; expect_exit "native-facts offline in sync -> 0" 0 $?
      "$PYTHON" "$NATIVE_CHK" --offline --live >/dev/null 2>&1; expect_exit "mutually exclusive modes -> 2" 2 $?
      "$PYTHON" "$NATIVE_CHK" --offline --skill "$SB/not-a-skill" >/dev/null 2>&1; expect_exit "missing skill dir -> 3" 3 $?
      out="$("$PYTHON" "$NATIVE_CHK" --offline --json 2>/dev/null)"
      expect_has "native-facts json schema" "claude-mods.loop-ops.native-facts/v1" "$out"
      expect_has "native-facts json in_sync" '"in_sync": true' "$out"
      # Drift detection: copy the skill, drop a host from the reference -> must be caught.
      mkdir -p "$SB/drift"
      cp -r "$SKILL/assets" "$SKILL/scripts" "$SKILL/references" "$SB/drift/" 2>/dev/null
      "$PYTHON" "$NATIVE_CHK" --offline --skill "$SB/drift" >/dev/null 2>&1
      expect_exit "unmodified copy still in sync -> 0" 0 $?
      sed 's|`cloud-routine`|`clown-routine`|g' "$SKILL/references/native-scheduling.md" > "$SB/drift/references/native-scheduling.md"
      out="$("$PYTHON" "$NATIVE_CHK" --offline --skill "$SB/drift" 2>/dev/null)"; rc=$?
      expect_exit "host vocabulary drift -> 10" 10 "$rc"
      expect_has  "names the drifted source" "native-scheduling.md" "$out"
      # A reference that loses its date stamp is drift too - undated is how a doc rots.
      sed 's|\*\*Verified 2026-08-30\*\*|Verified recently|' "$SKILL/references/native-scheduling.md" > "$SB/drift/references/native-scheduling.md"
      "$PYTHON" "$NATIVE_CHK" --offline --skill "$SB/drift" >/dev/null 2>&1; expect_exit "missing date stamp -> 10" 10 $?
      
      # ── docs: the native-scheduling reference must exist AND be cited ──────────
      # A reference SKILL.md never links is dead weight the router can't find
      # (docs/SKILL-CREATION-PROTOCOL.md step 4), so both halves are asserted.
      echo "-- native-scheduling reference --"
      NATIVE="$SKILL/references/native-scheduling.md"
      [[ -f "$NATIVE" ]] && ok "native-scheduling.md present" || no "native-scheduling.md missing"
      skillmd="$(cat "$SKILL/SKILL.md")"
      expect_has "SKILL.md cites native-scheduling.md" "references/native-scheduling.md" "$skillmd"
      expect_has "claude-code-loops.md cites native-scheduling.md" "native-scheduling.md" "$(cat "$SKILL/references/claude-code-loops.md")"
      # The reference is a claim about a fast-moving external surface, so it must carry the
      # date it was verified - an undated table is how a scheduling doc rots invisibly.
      expect_has "native-scheduling.md is date-stamped" "Verified 2026-08-30" "$(cat "$NATIVE")"
      # Each host named by the config template must be documented in the reference.
      for h in session-cron desktop-task cloud-routine; do
        expect_has "reference documents host '$h'" "$h" "$(cat "$NATIVE")"
      done
      # The gate is an eval; loop-ops links the discipline rather than restating it.
      expect_has "SKILL.md routes gate-judgement work to evals-ops" "evals-ops" "$skillmd"
      # A skill that repositions itself against a native primitive must say when NOT to use
      # itself, or it just re-sells ceremony for work the harness already does.
      expect_has "SKILL.md says when the native primitive is enough" "native primitive is enough" "$skillmd"
      
      # FRONTMATTER CONTRACT (stated here because a later description-trim lane edits
      # frontmatter without reading this suite - see SKILL-CREATION-PROTOCOL.md step 5):
      #   1. `description` must keep the repositioning clause "Native primitives schedule;
      #      loop-ops governs" - it is what stops a reader reaching for this skill to build a
      #      scheduler the harness already ships. Trimming it changes what the skill IS.
      #   2. `description` must keep the native trigger words, or the router never fires on
      #      "cron", "scheduled task" or "cloud routine".
      #   3. Nothing here requires `when_to_use`; this skill does not use that field.
      fm="$(grep -m1 '^description:' "$SKILL/SKILL.md")"
      expect_has "description keeps the repositioning clause" "Native primitives schedule; loop-ops governs" "$fm"
      expect_has "description keeps the native scheduling triggers" "cloud routine" "$fm"
      # Progressive disclosure: the body stays under the 500-line cap (protocol step 3).
      sk_lines="$(wc -l < "$SKILL/SKILL.md" | tr -d ' ')"
      [[ "$sk_lines" -lt 500 ]] && ok "SKILL.md under the 500-line cap ($sk_lines)" || no "SKILL.md is $sk_lines lines (cap 500)"
      
      # ── summary ────────────────────────────────────────────────────────────────
      echo "=== $PASS passed, $FAIL failed ==="
      [[ "$FAIL" -eq 0 ]] || exit 1
      
  • SKILL.md 30.5 KB
    ---
    name: loop-ops
    description: "Design and safely run OUTER loops - scheduled discover-triage-implement-verify-escalate agent loops. Native primitives schedule; loop-ops governs. Risk-tier ladder (L1 report -> L3 unattended), STATE/run-log/budget spine, kill switch, pattern catalog. Triggers: outer loop, scheduled/autonomous agent loop, PR watch, CI watch, dep-bump loop, run on a schedule, kill switch, risk tier, CronCreate, scheduled task, cloud routine, /loop."
    license: MIT
    allowed-tools: "Read Write Edit Bash Glob Grep"
    metadata:
      author: claude-mods
      related-skills: "iterate, fleet-ops, fleet-worker, pigeon, git-ops, ci-cd-ops"
    ---
    
    # Loop Ops — outer-loop design discipline
    
    **A loop is not a prompt.** Turn-by-turn prompting puts you in the loop forever. *Loop
    engineering* inverts it: you design a **recurring process with memory, verification, and
    boundaries** that discovers work, hands it to agents, verifies the result, and decides —
    on a schedule or until a goal is met — whether to **land it or escalate to a human**.
    
    > "You shouldn't be prompting coding agents anymore. You should be designing the loops
    > that prompt your agents." — Peter Steinberger
    
    This skill is the **outer loop**: the orchestration layer *above* a single agent run. It
    is the twin of [`iterate`](../iterate/SKILL.md) — `iterate` is the *inner* loop (one
    metric, one session, git-as-memory); `loop-ops` is the design discipline for the loop
    that *schedules and gates* inner runs. It does not reimplement spawning or landing; it
    **composes** what this repo already ships.
    
    ## Native primitives schedule; loop-ops governs
    
    Claude Code now ships the *cadence* half natively. **Do not hand-roll a scheduler** — pick
    a native host, declare it as `host:` in the config, and spend the discipline where the
    primitives leave a hole. Verified surface, parameters and limits (2026-08-30):
    [references/native-scheduling.md](references/native-scheduling.md).
    
    | Native primitive | What it gives you | What it does NOT give you |
    |---|---|---|
    | **`/loop`** — a *bundled* skill ([docs](https://code.claude.com/docs/en/scheduled-tasks)), driving `CronCreate`/`CronList`/`CronDelete`; `ScheduleWakeup` for its self-paced mode | Fixed-cron or Claude-paced ticks (delay clamped 60 s–1 h), a built-in maintenance prompt, `.claude/loop.md` to override it, `Esc` to stop | **Session-scoped and in-memory** — fires only while the session is idle, dies with the conversation, and every recurring job **self-deletes after 7 days**. No state spine, no budget, no gate. L1-supervised only. |
    | **Desktop scheduled tasks** — the `scheduled-tasks` MCP server ([docs](https://code.claude.com/docs/en/desktop-scheduled-tasks)) | Durable local ticks (≥1 min) with local files, a **fresh session per run**, a per-task permission mode with saved approvals, a task folder, run history, an Active/Paused toggle | The worktree toggle is **off by default** (runs against uncommitted changes); one catch-up only for a missed window; a Manual-mode task **stalls** on an unapproved tool. No STATE spine, no token budget, no verify gate. |
    | **Cloud routines** — `/schedule` ([docs](https://code.claude.com/docs/en/routines)) | Machine-off ticks (≥1 h), plus **native event triggers**: an API `/fire` endpoint and GitHub `pull_request`/`release` events with filters. A real push guard on non-`claude/` branches | **No permission mode at all** and **every connector attaches by default**; no local files (fresh clone); green run status ≠ task success. The boundary must come from repos + environment + connectors. |
    | **`/goal`** ([docs](https://code.claude.com/docs/en/goal)) | A native *completion* gate — keep going until a fast model confirms the condition | Not a cadence, and not an audit trail. |
    
    What **none** of them provide — and what this skill is for: a **state spine** that survives
    ticks, a **token budget**, a **verify gate** you can trust, an **escalation rule**, and the
    **risk-tier ladder** that decides whether the loop has earned the autonomy you're about to
    grant it. The plumbing moved into the harness; the judgement did not.
    
    ### When the native primitive is enough — stop here
    
    Don't scaffold a loop for work the harness already does. **Use the primitive raw** when
    *all* of these hold:
    
    - it **writes nothing** you'd have to undo (watch a deploy, poll a build, remind you), or
      the only writes are ones you'll review anyway;
    - it is **supervised or short-lived** — you're watching, or it stops in a session;
    - **nothing needs to be remembered between ticks** beyond what's in the repo;
    - and you'd **shrug if a tick silently didn't fire**.
    
    `/loop 5m check if the deploy finished` is a complete, correct answer. Wrapping it in a
    `loop.config.yaml` adds ceremony and no safety.
    
    Reach for loop-ops the moment **any one** of those flips: the loop starts *changing*
    things, runs *unattended*, needs to know what it did *last time*, or its silence would
    cost you. That is the whole trigger — everything below is what to do once it fires.
    
    ---
    
    ## The six primitives → what owns each here
    
    Every durable loop rests on six primitives. The discipline is wiring them; the parts
    already exist:
    
    | Primitive | What it is | Owned in claude-mods by |
    |---|---|---|
    | **Schedule** | fire the loop on a cadence *or an event* | native-first, declared as `host:` — `session-cron` (`/loop`+`CronCreate`, L1 only), `desktop-task` (`scheduled-tasks` MCP: local + durable), `cloud-routine` (`/schedule`: machine-off, plus API/GitHub event triggers), `/goal` for completion. `external` (cron/Task Scheduler + `loop-run.sh`) only for non-Claude-Code control |
    | **Worktree** | isolated, discardable execution context | `git-ops` worktrees, `fleet-worker` (per-task worktree) |
    | **Skills** | persistent project knowledge the run loads | this repo's skill layer + your `CLAUDE.md` |
    | **Sub-agents** | maker/checker separation | `Agent`/`Task`; dispatching skills (`review`, `testgen`) |
    | **Connectors** | reach tickets / CI / chat | MCP tools, `gh`, `github-ops` |
    | **+ State** | a durable spine *outside* the conversation | `STATE.md` + run-log + budget (this skill) |
    
    The inner improvement loop is `iterate`; cheap parallel makers are `fleet-worker`; the
    test-gated merge queue is `fleet-ops`; inter-loop signalling is `pigeon`. `loop-ops` is
    the doctrine that connects them.
    
    ## Loop anatomy
    
    ```
       ┌──────────────────────────────────────────────────────────────┐
       │  SCHEDULE (cadence)                                           │
       │     └─▶ TRIAGE      read STATE.md → pick the next unit of work │
       │           └─▶ WORKTREE   isolate (git worktree)               │
       │                 └─▶ MAKER     implementer run (or fleet-worker)│
       │                       └─▶ CHECKER  verify gate + guard (tests) │
       │                             └─▶ GATE  safe & allowlisted?      │
       │                                   ├─ yes → LAND  (commit/PR)   │
       │                                   └─ no  → ESCALATE (+context) │
       │     └─▶ write STATE.md, append run-log, decrement budget ──────┘
    ```
    
    The **gate** is the load-bearing decision. Everything before it is mechanical; the gate
    is where a loop earns the right to run unattended — or doesn't.
    
    ## The risk-tier ladder (the heart of the discipline)
    
    Never start a loop unattended. Graduate it. Each tier maps to a concrete Claude Code
    **permission mode** — full mapping, the headless-profile table, and the *enumerate vs
    isolate* fork in [references/risk-tiers.md](references/risk-tiers.md).
    
    | Tier | Posture | Permission mode | May do | Lands by |
    |---|---|---|---|---|
    | **L1 Report** | read-only discovery + triage | `plan` / `dontAsk`+read allowlist | scan, summarize, propose — **writes nothing** | a human reads the report |
    | **L2 Assisted** | suggest changes, human gates the merge | `dontAsk`+narrow allowlist, or `auto` | edit in a **worktree**, run tests, open a PR | a human approves the PR (or `fleet-ops`) |
    | **L3 Unattended** | autonomous land within a denylist | `bypassPermissions` **in an isolated container only** | commit/merge allowlisted classes | the loop itself, inside its boundary |
    
    **The host is part of the tier.** `session-cron` (`/loop` + `CronCreate`) cannot host L2+
    at all: it needs an open idle session and every recurring job expires after 7 days.
    `cloud-routine` has *no permission mode*, so its tier is expressed as repos + environment
    network policy + connectors instead — which means a routine is effectively autonomous the
    moment it is created, and the L1 posture has to come from the prompt being read-only.
    `loop-doctor` enforces both. Details: [references/native-scheduling.md](references/native-scheduling.md).
    
    The cardinal rule, straight from Claude Code's own gate model: **an unattended loop is a
    *scheduler/script that invokes `claude -p`*, not a Claude session that spawns ungated
    children.** A session in `auto` mode that tries to launch a `--permission-mode
    bypassPermissions` child is blocked as *Create Unsafe Agents* — by design. See
    [references/risk-tiers.md](references/risk-tiers.md) and the repo's
    [auto-mode-classifier reference](../../docs/AUTO-MODE-CLASSIFIER.md).
    
    ## The escalation gate
    
    What a loop may **land** vs what it must **escalate** is not a vibe — it mirrors Claude
    Code's classifier tiers. Bake these into the config's `escalation:` field:
    
    - **Always escalate (never auto-land):** force-push, push to `main`, production deploys
      or migrations, mass deletion, granting IAM/repo permissions, anything destroying
      pre-session files, editing `.claude/`/settings (self-modification), `curl | bash`.
    - **Safe to auto-land at L2/L3 (when allowlisted):** a green PR on a feature branch,
      a lockfile patch bump that passes the guard, a generated changelog draft, a label/
      triage classification, a comment.
    - **The test:** *would a careful human let this happen unattended in this repo?* If the
      action's blast radius exceeds the loop's stated purpose, it escalates. A general goal
      ("keep CI green") is **not** authorization for a specific high-blast action it implies.
    - **Scope the tools, not just the mode.** Allowlist exactly the tools/MCP connectors the
      job needs (read-only at L1); keep `gh pr merge` out and `land_via: fleet-ops` in. Full
      connector/MCP-scope discipline + the auto-merge guard: [references/risk-tiers.md](references/risk-tiers.md).
      On a **cloud routine this is the whole gate** — there is no permission mode, and every
      connected connector is attached by default with full write access. Prune them.
    - **A task that reschedules itself is self-modification.** The `scheduled-tasks` MCP lets a
      running task call `update_scheduled_task` on its own schedule or prompt. Useful, and on
      the always-escalate list unless adaptive cadence is the loop's *stated* purpose — a loop
      that can rewrite its own trigger has left the boundary you audited.
    
    ## The state spine
    
    A loop's memory lives **outside** the conversation, in three files (schemas +
    read/write contract in [references/state-spine.md](references/state-spine.md)):
    
    - **`STATE.md`** — the triage snapshot: priority / watch / noise + a readiness line.
      Read at the top of every run, rewritten at the end.
    - **`run-log.md`** — append one line per run (timestamp, action, outcome, tokens). The
      audit trail that answers "what has this loop been doing?"
    - **`loop.config.yaml`** — the loop's definition (goal, tier, cadence, **host**, scope,
      gate, budget, escalation). Scaffolded by `loop-scaffold`, scored by `loop-check`.
    
    A native host gives you a *place* for this spine (a Desktop task's folder) but never the
    spine itself: no host writes `STATE.md`, enforces a token budget, or records what the loop
    *decided*. Run history says a tick happened; the run-log says what it did and cost.
    
    ## Pattern catalog (a morphology, not a fixed list)
    
    Patterns are **compositions of three axes** — `trigger` (cadence / **event** via a Channel
    / `goal`) × `posture` (L1/L2/L3) × `locus` (connector→cloud routine / local→Desktop task).
    The named patterns are well-trodden points in that space; compose your own from the axes.
    Full recipes + the morphology in [references/pattern-catalog.md](references/pattern-catalog.md):
    
    | Pattern | Trigger · Locus | Tier | One-line job |
    |---|---|---|---|
    | `daily-scan` | cadence · local | L1 | discover + prioritize, report only |
    | `pr-watch` | event\|cadence · connector | L1 | watch review state, surface stuck PRs |
    | `ci-watch` | **event** · local | L2 | triage build failures, propose a fix |
    | `dep-bump` | cadence · local | L2 | patch-only bumps behind cooldown + guard |
    | `changelog-gen` | event(tag)\|cadence · local | L1 | draft release notes for approval |
    | `merge-hygiene` | cadence · local | L1 | dead branches, stale flags |
    | `issue-sort` | cadence · connector | L1 | classify + label, propose only |
    | `metric-chase` | **goal** · local | L2 | drive a metric (coverage/latency/eval) via `iterate` |
    | `regression-watch` | cadence\|event · local | L1 | run a benchmark/eval, flag a regression |
    | `digest` | cadence · **connector** | L1 | summarize email/Asana/news (cloud routine) |
    | `backfill` | **goal** · local | L2 | drain a migration/queue **to completion** |
    | `monitor` | **event** · local | L1 | error/deploy webhook → triage + page |
    | `freshness` | cadence · local | L1 | re-check docs/data/deps vs reality |
    
    Start any pattern at L1. Graduate to L2 only after the L1 reports prove its judgment.
    **Prefer `event` over `cadence`** where a webhook exists (cheaper, faster than polling).
    
    ## Multi-loop coordination & the kill switch
    
    Running several loops? Two non-negotiables (detail in
    [references/state-spine.md](references/state-spine.md)):
    
    - **Priority order** prevents collisions: `CI Watch → PR Watch → Dependency Bump →
      Merge-Hygiene/Changelog → Daily Scan (off-peak)`. A higher-priority loop's
      worktree wins; lowers defer. Loops signal each other via [`pigeon`](../pigeon/SKILL.md).
    - **A kill switch every loop honors.** A single stop signal — a `PAUSED` sentinel file
      or a `loop-pause` label — that every loop checks at the top of its run and exits on.
      No loop ships without one. Put it in `kill_switch:` and check it first.
    
    ## Composition map — don't rebuild what exists
    
    | You need to… | Use | Not |
    |---|---|---|
    | improve one metric in one session | [`iterate`](../iterate/SKILL.md) | a hand-rolled inner loop |
    | spawn cheap parallel makers | [`fleet-worker`](../fleet-worker/SKILL.md) | bespoke `claude -p` plumbing |
    | route models across a fan-out (cheap finders, Opus judges) | [`fleet-worker` model-routing](../fleet-worker/references/model-routing.md) | every agent on the orchestrator's model |
    | test-gate + land winning branches | [`fleet-ops`](../fleet-ops/SKILL.md) | a manual merge step |
    | fire on a cadence or an event | a native `host:` — `/loop`, Desktop scheduled task, cloud routine (schedule/API/GitHub triggers); `/goal` for completion | a custom cron in this skill |
    | trust the `verify` gate's judgement | the **`evals-ops`** skill — a gate is an eval (golden set, judge bias, `pass^k`, blocking vs advisory) | eyeballing a few runs and calling it proven |
    | reason about per-tick prompt-cache cost | [`claude-api-ops` caching-and-cost](../claude-api-ops/references/caching-and-cost.md) | a TTL number memorised from a blog post |
    | commit / PR / release | [`git-ops`](../git-ops/SKILL.md), [`github-ops`](../github-ops/SKILL.md) | raw `git push` |
    | signal between loops | [`pigeon`](../pigeon/SKILL.md) | a shared scratch file |
    
    `loop-ops` is the **design layer**; these are the **execution layers**.
    
    ---
    
    ## Tools
    
    Six scripts, all following the [Skill Resource Protocol](../../docs/SKILL-RESOURCE-PROTOCOL.md)
    (stdout = data, semantic exit codes, `--help` with EXAMPLES, `--json` envelopes): **init**
    scaffolds the loop, **audit** scores whether the config is *well-formed*, **doctor**
    preflights whether it will actually *run* (host-aware), **cost** estimates spend
    (caching-aware), and two drift guards — **check-pricing-sync** for the pricing table and
    **check-native-facts** for the native-scheduling limits. The discipline before scheduling
    is `init → fill → cost → audit → doctor --live`.
    
    ### `scripts/loop-scaffold.sh` — scaffold a loop's state spine
    
    Writes `<dir>/<name>/` with five files from the bundled templates:
    `loop.config.yaml` ([assets/loop.config.template.yaml](assets/loop.config.template.yaml)),
    `STATE.md` ([assets/STATE.template.md](assets/STATE.template.md)), `run-log.md`, `run.md`
    (the headless run prompt, [assets/run.template.md](assets/run.template.md)), and an
    executable **`loop-run.sh`** ([assets/run.sh.template](assets/run.sh.template)) — the
    runner-agnostic tick wrapper any scheduler invokes (cron / Windows Task Scheduler /
    systemd / by hand), **no GitHub Actions required**. Pass a known `--pattern`
    (pr-watch, ci-watch, dep-bump, …) and the config is **seeded** with that
    pattern's scope/goal/escalation — and, at L2+, its gate — so you get a near-ready config to
    review, not blank placeholders (it audits clean immediately). Doctrine holds: it still
    scaffolds at L1 by default with a graduation block.
    
    `--host` records where ticks will execute (`local` default, or `session-cron` /
    `desktop-task` / `cloud-routine` / `external`) so `loop-doctor` checks that host's real
    constraints instead of assuming a local `claude -p`.
    
    ```bash
    # Create .loops/pr-watch/ with config + STATE.md + run-log.md + run.md from templates:
    bash scripts/loop-scaffold.sh --name pr-watch --pattern pr-watch --tier L1
    
    # A connector-driven loop bound for a cloud routine (>=1h floor, no permission mode):
    bash scripts/loop-scaffold.sh --name digest --pattern digest --host cloud-routine --cadence 1h
    
    # Custom dir + cadence, preview without writing:
    bash scripts/loop-scaffold.sh --name dep-bump --pattern dep-bump \
      --tier L2 --cadence 1d --dir .loops --dry-run
    ```
    
    Refuses to overwrite a populated `<dir>/<name>/` (exit 5) unless `--force`. Atomic
    writes. `--dry-run` prints what it would create and writes nothing. stdout = the created
    config path.
    
    ### `scripts/loop-check.sh` — readiness scorer (run before you schedule)
    
    The question this answers: *is this loop safe to turn on at its declared tier?* It scores
    a `loop.config.yaml` against the readiness rubric — gate present, scope bounded,
    escalation defined, guard + worktree at L2+, budget + kill switch set, permission mode
    consistent with tier — and refuses a green light if any **critical** gap exists.
    
    ```bash
    bash scripts/loop-check.sh .loops/pr-watch/loop.config.yaml   # exit 0 ready, 10 not ready
    bash scripts/loop-check.sh --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.severity=="error")'
    bash scripts/loop-check.sh --min 80 .loops/ci-watch/loop.config.yaml   # raise the score bar
    ```
    
    Exit **0** = ready (no errors, score ≥ `--min`), **10** = not ready (findings on stdout),
    `2` usage, `3` config not found, `4` config unparseable. `--strict` counts warnings
    toward the not-ready signal.
    
    ### `scripts/loop-doctor.sh` — live preflight (will it actually run?)
    
    `loop-check` proves the config is *well-formed*; `loop-doctor` proves the loop will
    *execute* — catching the "blocked at 3am" failures audit can't see. `--offline` (CI-safe):
    the budget fits a tick's estimated tokens, the permission mode is achievable (not
    interactive), an L3 bypass declares an isolation boundary. `--live` adds runtime preflight:
    the `verify`/`guard` gate's leading binary resolves on PATH, `claude`/`git` are present,
    the kill-switch sentinel's parent dir exists.
    
    **It is host-aware.** `host:` changes what "will it run" even means, so the doctor checks
    against the declared surface: a `cloud-routine` faster than its 1-hour floor is rejected at
    creation; a routine with no named repos/environment/connector boundary has no gate at all
    (it has no permission mode either, so demanding one there would be a false finding); a
    `session-cron` host at L2+ can't run unattended and is called a predicted failure; and
    `--live` is **skipped, not passed**, for a cloud routine — this machine's PATH says nothing
    about a fresh cloud clone, and a green check there would be false confidence.
    
    ```bash
    bash scripts/loop-doctor.sh --offline .loops/pr-watch/loop.config.yaml   # CI gate
    bash scripts/loop-doctor.sh --live .loops/ci-watch/loop.config.yaml          # before scheduling
    bash scripts/loop-doctor.sh --live --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.state=="bad")'
    bash scripts/loop-doctor.sh --offline .loops/digest/loop.config.yaml   # host: cloud-routine -> floor + boundary
    ```
    
    Exit **0** = will run, **10** = a check predicts a runtime failure (gate binary missing,
    bypass on host without isolation, budget too small for a tick), `2` usage, `3` not found,
    `4` unparseable, `5` missing core dep. Run it **after** `loop-check` and before scheduling.
    
    ### `scripts/loop-estimate.py` — token/$ estimate by pattern × cadence × model (caching-aware)
    
    Estimate spend **before** committing to a cadence — the cost of an outer loop is
    runs/day × tokens/run × price, and sub-agents multiply it. It also models **prompt
    caching**: a loop re-sends the same `run.md`+system prefix every tick (the Ralph
    property), so the prefix should be cache-written once then read (~0.1×) — *but only if the
    tick interval fits the cache TTL*. **The TTL is a choice, not a constant:** 5 minutes by
    default (1.25× write) or 1 hour with `"ttl": "1h"` (2× write), so the daemon window is
    ~4.5 min *or* ~55 min — not a fixed 270 s. The estimator picks the cheapest TTL that stays
    warm at your cadence and names it; past 1 h nothing caches at all. Mechanics and
    break-even: [`claude-api-ops` caching-and-cost](../claude-api-ops/references/caching-and-cost.md).
    The estimate itself is **host-agnostic** — tokens are tokens wherever the tick fires; the
    host-dependent limit is the *minimum cadence*, which `loop-doctor` enforces. Pricing reads from
    `assets/model-pricing.json` (date-stamped; [`claude-api-ops`](../claude-api-ops/SKILL.md)
    is the source of truth — run its `check-model-table.py` if you suspect drift).
    
    ```bash
    python scripts/loop-estimate.py --pattern pr-watch --cadence 10m --model claude-haiku-4-5
    python scripts/loop-estimate.py --pattern ci-watch --cadence 15m --model claude-sonnet-5 --days 30 --json
    python scripts/loop-estimate.py --list-models      # the pricing table + its as-of date
    ```
    
    Exit `0` ok, `2` usage, `3` pricing file missing, `4` bad cadence/model. Output names
    every assumption (runs/day, tokens/run, sub-agent multiplier) — it's an estimate, and it
    says so.
    
    ### `scripts/check-pricing-sync.py` — offline drift guard (CI)
    
    `model-pricing.json` is a *copy* of claude-api-ops's authoritative model table, and a copy
    drifts silently. This offline verifier asserts every model in
    [assets/model-pricing.json](assets/model-pricing.json) matches claude-api-ops's "Current
    Models" table (prices included). Both files are in-repo, so it's network-free and gates PR
    CI via `tests/check-resources.sh`; live model-id drift is owned by claude-api-ops's
    `check-model-table.py`.
    
    ```bash
    python scripts/check-pricing-sync.py --offline   # exit 0 in sync, 10 drift, 3 a file missing
    ```
    
    ### `scripts/check-native-facts.py` — native-scheduling staleness guard
    
    [references/native-scheduling.md](references/native-scheduling.md) encodes a **fast-moving
    external surface**, and `loop-doctor` refuses configs on those numbers — so a silently
    stale limit becomes a wrong refusal. `--offline` (PR CI) proves internal consistency: the
    host vocabulary is *one* set across the config template, `loop-scaffold --host`,
    `loop-doctor`'s case arm and the reference; the reference still carries its `Verified
    <date>` stamp; and every limit the doctor enforces is still stated in the prose that
    justifies it. `--live` (scheduled, never a PR gate) fetches the three published docs pages
    and checks our numbers still appear in them.
    
    ```bash
    python scripts/check-native-facts.py --offline   # exit 0 in sync, 10 drift, 3 file missing
    python scripts/check-native-facts.py --live      # exit 7 = docs unreachable (advisory)
    ```
    
    ---
    
    ## End-to-end workflow
    
    1. **Pick a pattern** from the catalog (or `custom`), and **pick the host** — does the tick
       need local files? must it run with the machine off? is it supervised? Start at **L1**.
    2. **Scaffold:** `bash scripts/loop-scaffold.sh --name <n> --pattern <p> --tier L1
       --host <h>`.
    3. **Fill `loop.config.yaml`** — the real `goal`, `scope` (bounded globs, never `*`),
       `verify` gate, `escalation` rule, `budget_tokens`, `kill_switch`. On a `cloud-routine`,
       name the boundary that replaces the absent permission mode: repos, environment network
       policy, and the connectors you kept.
    4. **Cost it:** `python scripts/loop-estimate.py --pattern <p> --cadence <c> --model <m>` —
       sanity-check the monthly spend against the value.
    5. **Audit it:** `bash scripts/loop-check.sh .loops/<n>/loop.config.yaml` — fix every
       error before scheduling. Don't schedule a loop that fails its own audit.
    6. **Doctor it:** `bash scripts/loop-doctor.sh --live .loops/<n>/loop.config.yaml` — prove
       it will actually *run* (gate binary on PATH, budget fits a tick). Audit = well-formed;
       doctor = will-run.
    7. **Schedule** the L1 run on the declared host — the **recipe selector** in
       [references/claude-code-loops.md](references/claude-code-loops.md) prescribes which,
       because they're not interchangeable: connector-driven (email/Asana, no local code) →
       **cloud routine**; touches local code → **Desktop scheduled task**; sustained &
       token-sensitive → a **cache-warm daemon** (`claude -p` inside the cache TTL you paid
       for), *not* `/loop` (which grows a session and chews tokens); fixed-criteria long task →
       **`/goal`**; quick supervised polling → `/loop`. Per-primitive limits:
       [references/native-scheduling.md](references/native-scheduling.md). (L1 is read-only —
       it just writes `STATE.md` + a report.)
    8. **Read the reports.** Only after the loop's judgment is proven do you graduate it to
       **L2** (worktree + guard + `fleet-ops` landing), change `host:` if the proving host was
       `session-cron`, and re-audit at the higher tier. If the gate's verdict is a judgement
       rather than a green test run, harden it with the **`evals-ops`** discipline before you
       let it decide unattended.
    
    ## Worked example
    
    A complete, **audit + doctor-clean** L1 loop ships at
    [assets/examples/pr-watch/](assets/examples/pr-watch/): a filled
    `loop.config.yaml`, a *populated* `STATE.md`, the `run.md` run prompt, a sample
    `run-log.md`, the runner-agnostic **`loop-run.sh`** (the tick wrapper, with the
    kill-switch gate and `dontAsk` + allowlist baked in — point cron / Task Scheduler at it),
    and an *optional* `github-actions.yml` for repos already on GitHub. Copy the dir, adjust
    scope/cadence, run `loop-check` + `loop-doctor --live`, then wire `loop-run.sh` to your
    scheduler. The other patterns don't ship as
    static dirs that rot — `loop-scaffold --pattern <name>` *generates* the same, seeded and
    gate-clean, for any pattern at any tier. CI runs `loop-check` + `loop-doctor` on this
    example every build, so it can't drift out of validity.
    
    ## Anti-patterns (these are detected and wrong)
    
    The incident-shaped catalog — symptom → mechanism → the control that catches each — is
    [references/failure-modes.md](references/failure-modes.md) (runaway budget, the 3am-dead
    loop, cache-cold, force-push, ungated-child spawn, colliding loops, silent-stop,
    gate reward-hacking, and the native-host trio — the **expired** 7-day loop, the
    **over-connected** routine, the **stalled/skipped** Desktop task). The headline ones:
    
    - **Routing around the gate.** Wrapping `claude -p --permission-mode bypassPermissions`
      in a script to dodge the classifier is *Auto-Mode Bypass* — a `hard_deny` nothing
      clears. If an outcome is blocked, **authorize it** (a narrow allow rule, or run the
      scheduler outside the auto-mode session), never **disguise it**.
    - **The orchestrator session spawning ungated children.** A session in `auto` mode is
      the wrong place to launch the loop. The scheduler/cron/Task-Scheduler/CI runner that
      invokes `claude -p` is the authorizer. See [references/risk-tiers.md](references/risk-tiers.md) §"enumerate vs isolate".
    - **No gate.** A loop whose `verify:` is empty is not a loop, it's an unsupervised typer.
      `loop-check` errors on it. Nor is a *green run status* a gate: on a cloud routine green
      means "the session started and exited without an infrastructure error", never that the
      task succeeded. Grade the work, not the process.
    - **Assuming the native host gave you a boundary.** It gave you a cadence. A Desktop task's
      worktree toggle is off by default; a cloud routine has no permission mode and attaches
      every connector; `/loop` evaporates after 7 days. Each is a default that reads as safe
      and isn't.
    - **Unbounded scope.** `scope: "*"` means "may touch anything" — the audit rejects it.
    - **No kill switch / no budget.** A loop you can't stop, or whose spend you didn't
      bound, will eventually surprise you. Both are audit findings.
    - **Skipping L1.** Starting a fresh loop at L3 is how comprehension debt and incidents
      compound. The ladder exists precisely so trust is *earned* before it's *granted*.
    
    ## See also
    
    - [references/risk-tiers.md](references/risk-tiers.md) — L1/L2/L3 ↔ permission modes, headless profiles, enumerate-vs-isolate.
    - [references/pattern-catalog.md](references/pattern-catalog.md) — the seven patterns, full skeletons + escalation rules.
    - [references/state-spine.md](references/state-spine.md) — STATE.md / run-log / budget schemas, multi-loop coordination.
    - [references/native-scheduling.md](references/native-scheduling.md) — the native primitives themselves (verified 2026-08-30): `CronCreate`/`/loop` + its dynamic mode, the `scheduled-tasks` MCP, cloud routines — parameters, limits, failure semantics, and what each still doesn't give you.
    - [references/claude-code-loops.md](references/claude-code-loops.md) — which mechanism and how to wire it: the recipe selector, event triggers, hooks, the external-scheduler shape.
    - [references/failure-modes.md](references/failure-modes.md) — how loops break (incident-shaped) and the control that catches each.
    - [assets/loop.config.template.yaml](assets/loop.config.template.yaml) — the loop definition starter; [assets/STATE.template.md](assets/STATE.template.md) — the state-spine starter; [assets/run.template.md](assets/run.template.md) — the headless run prompt.
    - Lineage (public sources): the [Ralph loop](https://ghuntley.com/ralph/) (fresh-context inner brute-force) and the broader *loop engineering* discipline framed by Peter Steinberger and Addy Osmani.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related