loop-ops
Design and safely run OUTER loops - scheduled discover-triage-implement-verify-escalate agent loops. Native primitives schedule; loop-ops governs. Risk-tier ladder (L1 report -> L3 unattended), STATE/run-log/budget spine, kill switch, pattern catalog. Triggers: outer loop, schedu
Install
npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/loop-ops
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
git clone https://github.com/0xDarkMatter/claude-mods.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Loop Ops — outer-loop design discipline
A loop is not a prompt. Turn-by-turn prompting puts you in the loop forever. Loop engineering inverts it: you design a recurring process with memory, verification, and boundaries that discovers work, hands it to agents, verifies the result, and decides — on a schedule or until a goal is met — whether to land it or escalate to a human.
"You shouldn't be prompting coding agents anymore. You should be designing the loops that prompt your agents." — Peter Steinberger
This skill is the outer loop: the orchestration layer above a single agent run. It
is the twin of iterate — iterate is the inner loop (one
metric, one session, git-as-memory); loop-ops is the design discipline for the loop
that schedules and gates inner runs. It does not reimplement spawning or landing; it
composes what this repo already ships.
Native primitives schedule; loop-ops governs
Claude Code now ships the cadence half natively. Do not hand-roll a scheduler — pick
a native host, declare it as host: in the config, and spend the discipline where the
primitives leave a hole. Verified surface, parameters and limits (2026-08-30):
references/native-scheduling.md.
| Native primitive | What it gives you | What it does NOT give you |
|---|---|---|
/loop — a bundled skill (docs), driving CronCreate/CronList/CronDelete; ScheduleWakeup for its self-paced mode |
Fixed-cron or Claude-paced ticks (delay clamped 60 s–1 h), a built-in maintenance prompt, .claude/loop.md to override it, Esc to stop |
Session-scoped and in-memory — fires only while the session is idle, dies with the conversation, and every recurring job self-deletes after 7 days. No state spine, no budget, no gate. L1-supervised only. |
Desktop scheduled tasks — the scheduled-tasks MCP server (docs) |
Durable local ticks (≥1 min) with local files, a fresh session per run, a per-task permission mode with saved approvals, a task folder, run history, an Active/Paused toggle | The worktree toggle is off by default (runs against uncommitted changes); one catch-up only for a missed window; a Manual-mode task stalls on an unapproved tool. No STATE spine, no token budget, no verify gate. |
Cloud routines — /schedule (docs) |
Machine-off ticks (≥1 h), plus native event triggers: an API /fire endpoint and GitHub pull_request/release events with filters. A real push guard on non-claude/ branches |
No permission mode at all and every connector attaches by default; no local files (fresh clone); green run status ≠ task success. The boundary must come from repos + environment + connectors. |
/goal (docs) |
A native completion gate — keep going until a fast model confirms the condition | Not a cadence, and not an audit trail. |
What none of them provide — and what this skill is for: a state spine that survives ticks, a token budget, a verify gate you can trust, an escalation rule, and the risk-tier ladder that decides whether the loop has earned the autonomy you're about to grant it. The plumbing moved into the harness; the judgement did not.
When the native primitive is enough — stop here
Don't scaffold a loop for work the harness already does. Use the primitive raw when all of these hold:
- it writes nothing you'd have to undo (watch a deploy, poll a build, remind you), or the only writes are ones you'll review anyway;
- it is supervised or short-lived — you're watching, or it stops in a session;
- nothing needs to be remembered between ticks beyond what's in the repo;
- and you'd shrug if a tick silently didn't fire.
/loop 5m check if the deploy finished is a complete, correct answer. Wrapping it in a
loop.config.yaml adds ceremony and no safety.
Reach for loop-ops the moment any one of those flips: the loop starts changing things, runs unattended, needs to know what it did last time, or its silence would cost you. That is the whole trigger — everything below is what to do once it fires.
The six primitives → what owns each here
Every durable loop rests on six primitives. The discipline is wiring them; the parts already exist:
| Primitive | What it is | Owned in claude-mods by |
|---|---|---|
| Schedule | fire the loop on a cadence or an event | native-first, declared as host: — session-cron (/loop+CronCreate, L1 only), desktop-task (scheduled-tasks MCP: local + durable), cloud-routine (/schedule: machine-off, plus API/GitHub event triggers), /goal for completion. external (cron/Task Scheduler + loop-run.sh) only for non-Claude-Code control |
| Worktree | isolated, discardable execution context | git-ops worktrees, fleet-worker (per-task worktree) |
| Skills | persistent project knowledge the run loads | this repo's skill layer + your CLAUDE.md |
| Sub-agents | maker/checker separation | Agent/Task; dispatching skills (review, testgen) |
| Connectors | reach tickets / CI / chat | MCP tools, gh, github-ops |
| + State | a durable spine outside the conversation | STATE.md + run-log + budget (this skill) |
The inner improvement loop is iterate; cheap parallel makers are fleet-worker; the
test-gated merge queue is fleet-ops; inter-loop signalling is pigeon. loop-ops is
the doctrine that connects them.
Loop anatomy
┌──────────────────────────────────────────────────────────────┐
│ SCHEDULE (cadence) │
│ └─▶ TRIAGE read STATE.md → pick the next unit of work │
│ └─▶ WORKTREE isolate (git worktree) │
│ └─▶ MAKER implementer run (or fleet-worker)│
│ └─▶ CHECKER verify gate + guard (tests) │
│ └─▶ GATE safe & allowlisted? │
│ ├─ yes → LAND (commit/PR) │
│ └─ no → ESCALATE (+context) │
│ └─▶ write STATE.md, append run-log, decrement budget ──────┘
The gate is the load-bearing decision. Everything before it is mechanical; the gate is where a loop earns the right to run unattended — or doesn't.
The risk-tier ladder (the heart of the discipline)
Never start a loop unattended. Graduate it. Each tier maps to a concrete Claude Code permission mode — full mapping, the headless-profile table, and the enumerate vs isolate fork in references/risk-tiers.md.
| Tier | Posture | Permission mode | May do | Lands by |
|---|---|---|---|---|
| L1 Report | read-only discovery + triage | plan / dontAsk+read allowlist |
scan, summarize, propose — writes nothing | a human reads the report |
| L2 Assisted | suggest changes, human gates the merge | dontAsk+narrow allowlist, or auto |
edit in a worktree, run tests, open a PR | a human approves the PR (or fleet-ops) |
| L3 Unattended | autonomous land within a denylist | bypassPermissions in an isolated container only |
commit/merge allowlisted classes | the loop itself, inside its boundary |
The host is part of the tier. session-cron (/loop + CronCreate) cannot host L2+
at all: it needs an open idle session and every recurring job expires after 7 days.
cloud-routine has no permission mode, so its tier is expressed as repos + environment
network policy + connectors instead — which means a routine is effectively autonomous the
moment it is created, and the L1 posture has to come from the prompt being read-only.
loop-doctor enforces both. Details: references/native-scheduling.md.
The cardinal rule, straight from Claude Code's own gate model: an unattended loop is a
scheduler/script that invokes claude -p, not a Claude session that spawns ungated
children. A session in auto mode that tries to launch a --permission-mode bypassPermissions child is blocked as Create Unsafe Agents — by design. See
references/risk-tiers.md and the repo's
auto-mode-classifier reference.
The escalation gate
What a loop may land vs what it must escalate is not a vibe — it mirrors Claude
Code's classifier tiers. Bake these into the config's escalation: field:
- Always escalate (never auto-land): force-push, push to
main, production deploys or migrations, mass deletion, granting IAM/repo permissions, anything destroying pre-session files, editing.claude//settings (self-modification),curl | bash. - Safe to auto-land at L2/L3 (when allowlisted): a green PR on a feature branch, a lockfile patch bump that passes the guard, a generated changelog draft, a label/ triage classification, a comment.
- The test: would a careful human let this happen unattended in this repo? If the action's blast radius exceeds the loop's stated purpose, it escalates. A general goal ("keep CI green") is not authorization for a specific high-blast action it implies.
- Scope the tools, not just the mode. Allowlist exactly the tools/MCP connectors the
job needs (read-only at L1); keep
gh pr mergeout andland_via: fleet-opsin. Full connector/MCP-scope discipline + the auto-merge guard: references/risk-tiers.md. On a cloud routine this is the whole gate — there is no permission mode, and every connected connector is attached by default with full write access. Prune them. - A task that reschedules itself is self-modification. The
scheduled-tasksMCP lets a running task callupdate_scheduled_taskon its own schedule or prompt. Useful, and on the always-escalate list unless adaptive cadence is the loop's stated purpose — a loop that can rewrite its own trigger has left the boundary you audited.
The state spine
A loop's memory lives outside the conversation, in three files (schemas + read/write contract in references/state-spine.md):
STATE.md— the triage snapshot: priority / watch / noise + a readiness line. Read at the top of every run, rewritten at the end.run-log.md— append one line per run (timestamp, action, outcome, tokens). The audit trail that answers "what has this loop been doing?"loop.config.yaml— the loop's definition (goal, tier, cadence, host, scope, gate, budget, escalation). Scaffolded byloop-scaffold, scored byloop-check.
A native host gives you a place for this spine (a Desktop task's folder) but never the
spine itself: no host writes STATE.md, enforces a token budget, or records what the loop
decided. Run history says a tick happened; the run-log says what it did and cost.
Pattern catalog (a morphology, not a fixed list)
Patterns are compositions of three axes — trigger (cadence / event via a Channel
/ goal) × posture (L1/L2/L3) × locus (connector→cloud routine / local→Desktop task).
The named patterns are well-trodden points in that space; compose your own from the axes.
Full recipes + the morphology in references/pattern-catalog.md:
| Pattern | Trigger · Locus | Tier | One-line job |
|---|---|---|---|
daily-scan |
cadence · local | L1 | discover + prioritize, report only |
pr-watch |
event|cadence · connector | L1 | watch review state, surface stuck PRs |
ci-watch |
event · local | L2 | triage build failures, propose a fix |
dep-bump |
cadence · local | L2 | patch-only bumps behind cooldown + guard |
changelog-gen |
event(tag)|cadence · local | L1 | draft release notes for approval |
merge-hygiene |
cadence · local | L1 | dead branches, stale flags |
issue-sort |
cadence · connector | L1 | classify + label, propose only |
metric-chase |
goal · local | L2 | drive a metric (coverage/latency/eval) via iterate |
regression-watch |
cadence|event · local | L1 | run a benchmark/eval, flag a regression |
digest |
cadence · connector | L1 | summarize email/Asana/news (cloud routine) |
backfill |
goal · local | L2 | drain a migration/queue to completion |
monitor |
event · local | L1 | error/deploy webhook → triage + page |
freshness |
cadence · local | L1 | re-check docs/data/deps vs reality |
Start any pattern at L1. Graduate to L2 only after the L1 reports prove its judgment.
Prefer event over cadence where a webhook exists (cheaper, faster than polling).
Multi-loop coordination & the kill switch
Running several loops? Two non-negotiables (detail in references/state-spine.md):
- Priority order prevents collisions:
CI Watch → PR Watch → Dependency Bump → Merge-Hygiene/Changelog → Daily Scan (off-peak). A higher-priority loop's worktree wins; lowers defer. Loops signal each other viapigeon. - A kill switch every loop honors. A single stop signal — a
PAUSEDsentinel file or aloop-pauselabel — that every loop checks at the top of its run and exits on. No loop ships without one. Put it inkill_switch:and check it first.
Composition map — don't rebuild what exists
| You need to… | Use | Not |
|---|---|---|
| improve one metric in one session | iterate |
a hand-rolled inner loop |
| spawn cheap parallel makers | fleet-worker |
bespoke claude -p plumbing |
| route models across a fan-out (cheap finders, Opus judges) | fleet-worker model-routing |
every agent on the orchestrator's model |
| test-gate + land winning branches | fleet-ops |
a manual merge step |
| fire on a cadence or an event | a native host: — /loop, Desktop scheduled task, cloud routine (schedule/API/GitHub triggers); /goal for completion |
a custom cron in this skill |
trust the verify gate's judgement |
the evals-ops skill — a gate is an eval (golden set, judge bias, pass^k, blocking vs advisory) |
eyeballing a few runs and calling it proven |
| reason about per-tick prompt-cache cost | claude-api-ops caching-and-cost |
a TTL number memorised from a blog post |
| commit / PR / release | git-ops, github-ops |
raw git push |
| signal between loops | pigeon |
a shared scratch file |
loop-ops is the design layer; these are the execution layers.
Tools
Six scripts, all following the Skill Resource Protocol
(stdout = data, semantic exit codes, --help with EXAMPLES, --json envelopes): init
scaffolds the loop, audit scores whether the config is well-formed, doctor
preflights whether it will actually run (host-aware), cost estimates spend
(caching-aware), and two drift guards — check-pricing-sync for the pricing table and
check-native-facts for the native-scheduling limits. The discipline before scheduling
is init → fill → cost → audit → doctor --live.
scripts/loop-scaffold.sh — scaffold a loop's state spine
Writes <dir>/<name>/ with five files from the bundled templates:
loop.config.yaml (assets/loop.config.template.yaml),
STATE.md (assets/STATE.template.md), run-log.md, run.md
(the headless run prompt, assets/run.template.md), and an
executable loop-run.sh (assets/run.sh.template) — the
runner-agnostic tick wrapper any scheduler invokes (cron / Windows Task Scheduler /
systemd / by hand), no GitHub Actions required. Pass a known --pattern
(pr-watch, ci-watch, dep-bump, …) and the config is seeded with that
pattern's scope/goal/escalation — and, at L2+, its gate — so you get a near-ready config to
review, not blank placeholders (it audits clean immediately). Doctrine holds: it still
scaffolds at L1 by default with a graduation block.
--host records where ticks will execute (local default, or session-cron /
desktop-task / cloud-routine / external) so loop-doctor checks that host's real
constraints instead of assuming a local claude -p.
# Create .loops/pr-watch/ with config + STATE.md + run-log.md + run.md from templates:
bash scripts/loop-scaffold.sh --name pr-watch --pattern pr-watch --tier L1
# A connector-driven loop bound for a cloud routine (>=1h floor, no permission mode):
bash scripts/loop-scaffold.sh --name digest --pattern digest --host cloud-routine --cadence 1h
# Custom dir + cadence, preview without writing:
bash scripts/loop-scaffold.sh --name dep-bump --pattern dep-bump \
--tier L2 --cadence 1d --dir .loops --dry-run
Refuses to overwrite a populated <dir>/<name>/ (exit 5) unless --force. Atomic
writes. --dry-run prints what it would create and writes nothing. stdout = the created
config path.
scripts/loop-check.sh — readiness scorer (run before you schedule)
The question this answers: is this loop safe to turn on at its declared tier? It scores
a loop.config.yaml against the readiness rubric — gate present, scope bounded,
escalation defined, guard + worktree at L2+, budget + kill switch set, permission mode
consistent with tier — and refuses a green light if any critical gap exists.
bash scripts/loop-check.sh .loops/pr-watch/loop.config.yaml # exit 0 ready, 10 not ready
bash scripts/loop-check.sh --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.severity=="error")'
bash scripts/loop-check.sh --min 80 .loops/ci-watch/loop.config.yaml # raise the score bar
Exit 0 = ready (no errors, score ≥ --min), 10 = not ready (findings on stdout),
2 usage, 3 config not found, 4 config unparseable. --strict counts warnings
toward the not-ready signal.
scripts/loop-doctor.sh — live preflight (will it actually run?)
loop-check proves the config is well-formed; loop-doctor proves the loop will
execute — catching the "blocked at 3am" failures audit can't see. --offline (CI-safe):
the budget fits a tick's estimated tokens, the permission mode is achievable (not
interactive), an L3 bypass declares an isolation boundary. --live adds runtime preflight:
the verify/guard gate's leading binary resolves on PATH, claude/git are present,
the kill-switch sentinel's parent dir exists.
It is host-aware. host: changes what "will it run" even means, so the doctor checks
against the declared surface: a cloud-routine faster than its 1-hour floor is rejected at
creation; a routine with no named repos/environment/connector boundary has no gate at all
(it has no permission mode either, so demanding one there would be a false finding); a
session-cron host at L2+ can't run unattended and is called a predicted failure; and
--live is skipped, not passed, for a cloud routine — this machine's PATH says nothing
about a fresh cloud clone, and a green check there would be false confidence.
bash scripts/loop-doctor.sh --offline .loops/pr-watch/loop.config.yaml # CI gate
bash scripts/loop-doctor.sh --live .loops/ci-watch/loop.config.yaml # before scheduling
bash scripts/loop-doctor.sh --live --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.state=="bad")'
bash scripts/loop-doctor.sh --offline .loops/digest/loop.config.yaml # host: cloud-routine -> floor + boundary
Exit 0 = will run, 10 = a check predicts a runtime failure (gate binary missing,
bypass on host without isolation, budget too small for a tick), 2 usage, 3 not found,
4 unparseable, 5 missing core dep. Run it after loop-check and before scheduling.
scripts/loop-estimate.py — token/$ estimate by pattern × cadence × model (caching-aware)
Estimate spend before committing to a cadence — the cost of an outer loop is
runs/day × tokens/run × price, and sub-agents multiply it. It also models prompt
caching: a loop re-sends the same run.md+system prefix every tick (the Ralph
property), so the prefix should be cache-written once then read (~0.1×) — but only if the
tick interval fits the cache TTL. The TTL is a choice, not a constant: 5 minutes by
default (1.25× write) or 1 hour with "ttl": "1h" (2× write), so the daemon window is
~4.5 min or ~55 min — not a fixed 270 s. The estimator picks the cheapest TTL that stays
warm at your cadence and names it; past 1 h nothing caches at all. Mechanics and
break-even: claude-api-ops caching-and-cost.
The estimate itself is host-agnostic — tokens are tokens wherever the tick fires; the
host-dependent limit is the minimum cadence, which loop-doctor enforces. Pricing reads from
assets/model-pricing.json (date-stamped; claude-api-ops
is the source of truth — run its check-model-table.py if you suspect drift).
python scripts/loop-estimate.py --pattern pr-watch --cadence 10m --model claude-haiku-4-5
python scripts/loop-estimate.py --pattern ci-watch --cadence 15m --model claude-sonnet-5 --days 30 --json
python scripts/loop-estimate.py --list-models # the pricing table + its as-of date
Exit 0 ok, 2 usage, 3 pricing file missing, 4 bad cadence/model. Output names
every assumption (runs/day, tokens/run, sub-agent multiplier) — it's an estimate, and it
says so.
scripts/check-pricing-sync.py — offline drift guard (CI)
model-pricing.json is a copy of claude-api-ops's authoritative model table, and a copy
drifts silently. This offline verifier asserts every model in
assets/model-pricing.json matches claude-api-ops's "Current
Models" table (prices included). Both files are in-repo, so it's network-free and gates PR
CI via tests/check-resources.sh; live model-id drift is owned by claude-api-ops's
check-model-table.py.
python scripts/check-pricing-sync.py --offline # exit 0 in sync, 10 drift, 3 a file missing
scripts/check-native-facts.py — native-scheduling staleness guard
references/native-scheduling.md encodes a fast-moving
external surface, and loop-doctor refuses configs on those numbers — so a silently
stale limit becomes a wrong refusal. --offline (PR CI) proves internal consistency: the
host vocabulary is one set across the config template, loop-scaffold --host,
loop-doctor's case arm and the reference; the reference still carries its Verified <date> stamp; and every limit the doctor enforces is still stated in the prose that
justifies it. --live (scheduled, never a PR gate) fetches the three published docs pages
and checks our numbers still appear in them.
python scripts/check-native-facts.py --offline # exit 0 in sync, 10 drift, 3 file missing
python scripts/check-native-facts.py --live # exit 7 = docs unreachable (advisory)
End-to-end workflow
- Pick a pattern from the catalog (or
custom), and pick the host — does the tick need local files? must it run with the machine off? is it supervised? Start at L1. - Scaffold:
bash scripts/loop-scaffold.sh --name <n> --pattern <p> --tier L1 --host <h>. - Fill
loop.config.yaml— the realgoal,scope(bounded globs, never*),verifygate,escalationrule,budget_tokens,kill_switch. On acloud-routine, name the boundary that replaces the absent permission mode: repos, environment network policy, and the connectors you kept. - Cost it:
python scripts/loop-estimate.py --pattern <p> --cadence <c> --model <m>— sanity-check the monthly spend against the value. - Audit it:
bash scripts/loop-check.sh .loops/<n>/loop.config.yaml— fix every error before scheduling. Don't schedule a loop that fails its own audit. - Doctor it:
bash scripts/loop-doctor.sh --live .loops/<n>/loop.config.yaml— prove it will actually run (gate binary on PATH, budget fits a tick). Audit = well-formed; doctor = will-run. - Schedule the L1 run on the declared host — the recipe selector in
references/claude-code-loops.md prescribes which,
because they're not interchangeable: connector-driven (email/Asana, no local code) →
cloud routine; touches local code → Desktop scheduled task; sustained &
token-sensitive → a cache-warm daemon (
claude -pinside the cache TTL you paid for), not/loop(which grows a session and chews tokens); fixed-criteria long task →/goal; quick supervised polling →/loop. Per-primitive limits: references/native-scheduling.md. (L1 is read-only — it just writesSTATE.md+ a report.) - Read the reports. Only after the loop's judgment is proven do you graduate it to
L2 (worktree + guard +
fleet-opslanding), changehost:if the proving host wassession-cron, and re-audit at the higher tier. If the gate's verdict is a judgement rather than a green test run, harden it with theevals-opsdiscipline before you let it decide unattended.
Worked example
A complete, audit + doctor-clean L1 loop ships at
assets/examples/pr-watch/: a filled
loop.config.yaml, a populated STATE.md, the run.md run prompt, a sample
run-log.md, the runner-agnostic loop-run.sh (the tick wrapper, with the
kill-switch gate and dontAsk + allowlist baked in — point cron / Task Scheduler at it),
and an optional github-actions.yml for repos already on GitHub. Copy the dir, adjust
scope/cadence, run loop-check + loop-doctor --live, then wire loop-run.sh to your
scheduler. The other patterns don't ship as
static dirs that rot — loop-scaffold --pattern <name> generates the same, seeded and
gate-clean, for any pattern at any tier. CI runs loop-check + loop-doctor on this
example every build, so it can't drift out of validity.
Anti-patterns (these are detected and wrong)
The incident-shaped catalog — symptom → mechanism → the control that catches each — is references/failure-modes.md (runaway budget, the 3am-dead loop, cache-cold, force-push, ungated-child spawn, colliding loops, silent-stop, gate reward-hacking, and the native-host trio — the expired 7-day loop, the over-connected routine, the stalled/skipped Desktop task). The headline ones:
- Routing around the gate. Wrapping
claude -p --permission-mode bypassPermissionsin a script to dodge the classifier is Auto-Mode Bypass — ahard_denynothing clears. If an outcome is blocked, authorize it (a narrow allow rule, or run the scheduler outside the auto-mode session), never disguise it. - The orchestrator session spawning ungated children. A session in
automode is the wrong place to launch the loop. The scheduler/cron/Task-Scheduler/CI runner that invokesclaude -pis the authorizer. See references/risk-tiers.md §"enumerate vs isolate". - No gate. A loop whose
verify:is empty is not a loop, it's an unsupervised typer.loop-checkerrors on it. Nor is a green run status a gate: on a cloud routine green means "the session started and exited without an infrastructure error", never that the task succeeded. Grade the work, not the process. - Assuming the native host gave you a boundary. It gave you a cadence. A Desktop task's
worktree toggle is off by default; a cloud routine has no permission mode and attaches
every connector;
/loopevaporates after 7 days. Each is a default that reads as safe and isn't. - Unbounded scope.
scope: "*"means "may touch anything" — the audit rejects it. - No kill switch / no budget. A loop you can't stop, or whose spend you didn't bound, will eventually surprise you. Both are audit findings.
- Skipping L1. Starting a fresh loop at L3 is how comprehension debt and incidents compound. The ladder exists precisely so trust is earned before it's granted.
See also
- references/risk-tiers.md — L1/L2/L3 ↔ permission modes, headless profiles, enumerate-vs-isolate.
- references/pattern-catalog.md — the seven patterns, full skeletons + escalation rules.
- references/state-spine.md — STATE.md / run-log / budget schemas, multi-loop coordination.
- references/native-scheduling.md — the native primitives themselves (verified 2026-08-30):
CronCreate//loop+ its dynamic mode, thescheduled-tasksMCP, cloud routines — parameters, limits, failure semantics, and what each still doesn't give you. - references/claude-code-loops.md — which mechanism and how to wire it: the recipe selector, event triggers, hooks, the external-scheduler shape.
- references/failure-modes.md — how loops break (incident-shaped) and the control that catches each.
- assets/loop.config.template.yaml — the loop definition starter; assets/STATE.template.md — the state-spine starter; assets/run.template.md — the headless run prompt.
- Lineage (public sources): the Ralph loop (fresh-context inner brute-force) and the broader loop engineering discipline framed by Peter Steinberger and Addy Osmani.
Files (claude-mods)
-
assets
-
examples
-
pr-watch
-
github-actions.yml 2.7 KB
# OPTIONAL scheduler — use this ONLY if your repo already lives on GitHub. The portable, # runner-agnostic path is loop-run.sh (cron / Windows Task Scheduler / systemd / by hand) — # no GitHub Actions dependency. This file just wraps the same loop-run.sh idea in Actions. # Copy to .github/workflows/pr-watch.yml, PIN the action/CLI versions, add the # ANTHROPIC_API_KEY secret. The SCHEDULER is the authorizer (no auto-mode session in the # loop), and the child runs gated (--permission-mode dontAsk + a narrow allowlist), never # bypassPermissions on a shared runner. See references/claude-code-loops.md. name: pr-watch on: schedule: - cron: "*/10 * * * *" # every 10 min (matches loop.config.yaml cadence: 10m) workflow_dispatch: {} permissions: contents: write # commit STATE.md / run-log.md back pull-requests: write # post the at-most-one summary comment (L1 stays report-only) concurrency: group: pr-watch # never overlap two ticks cancel-in-progress: false jobs: tick: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 # <-- pin to a SHA in production # Kill switch: a 'loop-pause' label on the repo, or a committed PAUSED sentinel. - name: Honor the kill switch id: gate env: { GH_TOKEN: "${{ github.token }}" } run: | if [ -f .loops/pr-watch/PAUSED ]; then echo "paused=1" >> "$GITHUB_OUTPUT"; fi if gh label list --limit 100 | grep -qi '^loop-pause'; then echo "paused=1" >> "$GITHUB_OUTPUT"; fi - name: Install Claude Code if: steps.gate.outputs.paused != '1' run: npm i -g @anthropic-ai/claude-code # <-- pin a version # The run: same prompt every tick (cache-friendly), gated with dontAsk + an # allowlist scoped to exactly what an L1 report loop needs (read-only + gh + STATE writes). - name: Run one tick if: steps.gate.outputs.paused != '1' env: ANTHROPIC_API_KEY: "${{ secrets.ANTHROPIC_API_KEY }}" run: | cd .loops/pr-watch claude -p "$(cat run.md)" \ --permission-mode dontAsk \ --append-system-prompt "$(cat STATE.md)" \ --allowedTools 'Bash(gh pr list:*)' 'Bash(gh pr view:*)' 'Bash(gh pr comment:*)' 'Read' 'Write(STATE.md)' 'Write(run-log.md)' \ --max-turns 30 - name: Persist STATE + run-log if: steps.gate.outputs.paused != '1' run: | git config user.name "pr-watch-loop" git config user.email "loop@users.noreply.github.com" git add .loops/pr-watch/STATE.md .loops/pr-watch/run-log.md git diff --cached --quiet || git commit -m "chore(loop): pr-watch tick $(date -u +%FT%TZ)" git push -
loop-run.sh 1.6 KB
#!/usr/bin/env bash # loop-run.sh - one tick of the pr-watch loop. RUNNER-AGNOSTIC: point any scheduler # at it. No GitHub Actions required. # cron: */10 * * * * /path/.loops/pr-watch/loop-run.sh >> tick.log 2>&1 # Windows Task Scheduler: schtasks /Create /SC MINUTE /MO 10 /TN pr-watch \ # /TR "bash -lc '/path/.loops/pr-watch/loop-run.sh'" # by hand: bash loop-run.sh # The scheduler is the authorizer; this runs a gated `claude -p` (dontAsk + an allowlist), # never bypassPermissions on a shared host. (github-actions.yml is one OPTIONAL scheduler.) set -uo pipefail HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" cd "$HERE" # 1. Kill switch first. if [ -f PAUSED ]; then echo "pr-watch: paused (PAUSED sentinel) - skipping tick" >&2; exit 0; fi command -v claude >/dev/null 2>&1 || { echo "pr-watch: 'claude' not on PATH" >&2; exit 5; } # 2. One tick. SAME prompt every time (cache-friendly). Allowlist = exactly what an L1 # report loop needs: read-only gh + Read + the STATE/run-log writes. No 'gh pr merge'. claude -p "$(cat run.md)" \ --permission-mode dontAsk \ --append-system-prompt "$(cat STATE.md)" \ --allowedTools 'Bash(gh pr list:*)' 'Bash(gh pr view:*)' 'Bash(gh pr comment:*)' 'Read' 'Write(STATE.md)' 'Write(run-log.md)' \ --max-turns 30 # 3. Persist STATE + run-log if this dir lives in a git repo. if git rev-parse --is-inside-work-tree >/dev/null 2>&1; then git add STATE.md run-log.md 2>/dev/null || true git diff --cached --quiet 2>/dev/null || git commit -q -m "chore(loop): pr-watch tick" || true fi -
loop.config.yaml 1.1 KB
# WORKED EXAMPLE — a complete, audit-clean L1 loop. Copy the dir, adjust scope/cadence, # run loop-check + loop-doctor --live, then wire the scheduler (github-actions.yml). # The other patterns: `loop-scaffold --pattern <name>` generates the same, seeded. name: pr-watch pattern: pr-watch tier: L1 permission_mode: dontAsk cadence: 10m # Ticks run on this machine via the wrapper below, so the local host applies. A 10m # cadence rules out cloud-routine (1h floor); session-cron would work while you watch # it, but expires after 7 days. See references/native-scheduling.md. host: external goal: "Watch open PRs; flag stuck (no review > 4h), failing checks, and merge conflicts in STATE.md; post at most one summary comment per PR; never merge." scope: - "src/**" escalation: "a human reviews and merges; never merge to main; never push; never close a PR" budget_tokens: 60000 kill_switch: ".loops/pr-watch/PAUSED exists, OR the 'loop-pause' label is on the repo" # ── graduate to L2 (assisted: open a fix-the-PR-description / rebase branch) ── # tier: L2 # verify: "npm test" # guard: "npm run typecheck" # worktree: true # land_via: fleet-ops -
run-log.md 436 B
# pr-watch — run log (append-only; one line per run) # format: <ISO-Z> run#N action=<reported|none> pr=<n|-> outcome=<…> tokens=<N> 2026-06-22T14:05:00Z run#142 action=reported pr=412 outcome=commented-flaky-ci tokens=14820 2026-06-22T13:55:00Z run#141 action=reported pr=408 outcome=flagged-conflict tokens=12110 2026-06-22T13:45:00Z run#140 action=none pr=- outcome=quiet tokens=2090 -
run.md 1.7 KB
<!-- run.md — fed to `claude -p` each tick. SAME every run (fresh context; state lives in STATE.md + git, not the conversation). Keep it BYTE-IDENTICAL so the prompt cache hits. Wired by github-actions.yml: claude -p "$(cat run.md)" --permission-mode dontAsk --append-system-prompt "$(cat STATE.md)" --> # Run: pr-watch (tier L1, report-only) You are one tick of a scheduled loop. Goal: **watch open PRs and report; never merge, push, or close.** ## Do these in order 1. **Kill switch first.** If `.loops/pr-watch/PAUSED` exists or the repo has the `loop-pause` label, STOP — do nothing. 2. **Read `STATE.md`** (in your system prompt): the Priority / Watch / Noise lists from last run. 3. **List open PRs:** `gh pr list --state open --json number,title,reviewDecision,statusCheckRollup,mergeStateStatus,updatedAt`. 4. **Classify each** PR: failing checks · merge conflict · awaiting-review past 4h · draft · healthy. 5. **Report only.** You may post **at most one** summary comment per PR that newly needs attention (preview the text in the run log; never spam). You do **not** merge, push, rebase, or close — that escalates to a human. 6. **Respect the budget:** stop if you approach 60000 output tokens. 7. **Rewrite `STATE.md`:** move PRs across Priority / Watch / Noise; bump `_Updated_` + run number + readiness. 8. **Append one line to `run-log.md`:** `<ISO-Z> run#N action=reported pr=<n|-> outcome=<…> tokens=<N>`. ## Hard rules - Never merge/push/close/rebase — those are the escalation cases. A general goal is not authorization for them. - Stay within scope (`src/**`); never touch another session's `.claude/worktrees/`. - Leave the repo clean every tick. -
STATE.md 963 B
# pr-watch — STATE _Updated: 2026-06-22T14:05:00Z · run #142 · readiness 100/100_ <!-- Read first every run: check the kill switch, then act on the Priority list. This is a realistic populated snapshot — yours starts from STATE.template.md. --> ## Priority (act on these next) - [P1] PR #412 failing CI 3h10m — `build` job red, looks like a flaky e2e; owner @dana pinged 14:02 - [P1] PR #408 merge conflict with main (settings.ts) — author notified, awaiting rebase - [P2] PR #415 awaiting review 5h — past the 4h threshold; nudged the reviewers group ## Watch (not yet actionable) - PR #417 awaiting review 1h12m — under threshold, recheck next run - PR #410 draft — skip until marked ready ## Noise (seen + dismissed this run) - PR #419 opened 6m ago, checks still running — too early - PR #402 already merged since last run — drop from tracking --- _Source: .github/workflows/pr-watch.yml · config: loop.config.yaml_
-
-
-
loop.config.template.yaml 4.2 KB
# loop.config.yaml — one OUTER-loop definition. # Scaffolded by loop-scaffold.sh, scored by loop-check.sh. Flat YAML on purpose: every # key sits at column 0 so the audit parses it without a yq dependency. # Full field semantics: skills/loop-ops/references/state-spine.md # # >>> ADAPT every <PLACEHOLDER> below. The audit errors on unbounded scope, a # missing gate, an undefined escalation rule, or a tier/permission mismatch. # ── identity ──────────────────────────────────────────────────────────────── name: <loop-name> # matches the .loops/<name>/ directory pattern: <pattern-key> # a catalog key (pr-watch, ci-watch, …) or "custom" # ── autonomy ──────────────────────────────────────────────────────────────── tier: L1 # L1 report-only | L2 assisted (worktree+gate) | L3 unattended permission_mode: dontAsk # plan | dontAsk | auto | acceptEdits | bypassPermissions # L1 → plan or dontAsk · L2 → dontAsk/auto · L3 → bypassPermissions (container only) # ── cadence & host ────────────────────────────────────────────────────────── cadence: 1h # 10m | 1h | 6h | 1d, or a cron string ("*/10 * * * *") host: local # where ticks execute — decides which constraints apply: # local generic local run (default; no host checks) # session-cron /loop + CronCreate — L1 only: needs an open # idle session, 7-day expiry, ≥1 min # desktop-task scheduled-tasks MCP — durable, local files, # per-task permission mode, ≥1 min # cloud-routine /schedule Routines — no local files, NO # permission mode, ≥1 hour # external cron / Task Scheduler / systemd → loop-run.sh # Full semantics: references/native-scheduling.md # ── purpose & bounds ──────────────────────────────────────────────────────── goal: "<one sentence: what this loop does AND what it must never do>" scope: # bounded globs the loop may touch — NEVER "*" or "**" - "<src/**>" # ── the gate (required at L2+) ────────────────────────────────────────────── verify: "<command that decides pass/fail, e.g. npm test>" # a loop with no gate is invalid at L2+ guard: "<must-always-pass, e.g. npm run typecheck>" # required at L2+ # ── isolation & landing (required at L2+) ─────────────────────────────────── worktree: true # isolate code changes in a git worktree (required L2+) land_via: fleet-ops # who test-gates + lands winning branches (L2+) # ── the escalation rule (required) ────────────────────────────────────────── # What the loop ESCALATES instead of doing. Mirror the never-auto-land classes: # force-push, push to main, prod deploy/migration, mass delete, IAM grants, .claude edits. escalation: "<e.g. open a PR with context; never merge to main; never deploy>" # ── safety rails ──────────────────────────────────────────────────────────── budget_tokens: 200000 # per-run output-token ceiling (stop the run when reached) kill_switch: ".loops/<loop-name>/PAUSED exists, OR the loop-pause label is set" -
model-pricing.json 2.1 KB
{ "_comment": "Per-model USD pricing per million tokens, read by loop-estimate.py. SOURCE OF TRUTH is skills/claude-api-ops/SKILL.md 'Current Models' table — run skills/claude-api-ops/scripts/check-model-table.py --live if you suspect drift. Date-stamp this file when updating.", "_as_of": "2026-08", "_schema": "claude-mods.loop-ops.pricing/v1", "models": { "claude-fable-5": { "input_per_mtok": 10.0, "output_per_mtok": 50.0 }, "claude-opus-5": { "input_per_mtok": 5.0, "output_per_mtok": 25.0 }, "claude-sonnet-5": { "input_per_mtok": 2.0, "output_per_mtok": 10.0 }, "claude-haiku-4-5": { "input_per_mtok": 1.0, "output_per_mtok": 5.0 } }, "_pattern_defaults": { "_comment": "Rough per-run token estimates by pattern. input = context the run reads (STATE, diffs, tool results); output = tokens the model generates; subagents multiplies tokens when the pattern fans out makers/checkers. Estimates, not guarantees — reconcile against run-log.md actuals.", "daily-scan": { "input": 40000, "output": 6000, "subagents": 1 }, "pr-watch": { "input": 15000, "output": 3000, "subagents": 1 }, "ci-watch": { "input": 60000, "output": 12000, "subagents": 2 }, "dep-bump": { "input": 30000, "output": 8000, "subagents": 1 }, "changelog-gen": { "input": 25000, "output": 9000, "subagents": 1 }, "merge-hygiene": { "input": 20000, "output": 4000, "subagents": 1 }, "issue-sort": { "input": 18000, "output": 4000, "subagents": 1 }, "metric-chase": { "input": 50000, "output": 20000, "subagents": 3 }, "regression-watch": { "input": 45000, "output": 8000, "subagents": 1 }, "digest": { "input": 30000, "output": 7000, "subagents": 1 }, "backfill": { "input": 40000, "output": 15000, "subagents": 2 }, "monitor": { "input": 12000, "output": 3000, "subagents": 1 }, "freshness": { "input": 35000, "output": 6000, "subagents": 1 }, "custom": { "input": 30000, "output": 8000, "subagents": 1 } } } -
run.sh.template 1.4 KB · in bundle
-
run.template.md 2.2 KB
<!-- run.md — the prompt a scheduler feeds to `claude -p` each tick. It is the SAME every run (fresh context each time; state lives in STATE.md + the codebase + git, not the conversation — the Ralph property). Fill the <PLACEHOLDERS>; keep it short. Wire it up (see references/claude-code-loops.md): claude -p "$(cat .loops/<loop-name>/run.md)" \ --permission-mode <permission_mode-from-config> \ --append-system-prompt "$(cat .loops/<loop-name>/STATE.md)" The SCHEDULER (cron / Task Scheduler / CI), not a Claude session, invokes this. --> # Run: <loop-name> (tier <L1|L2|L3>) You are one tick of a scheduled loop. Goal: **<one sentence — what to do AND what never to do>**. ## Do these in order 1. **Check the kill switch FIRST.** If <kill_switch — e.g. `.loops/<loop-name>/PAUSED` exists, or the `loop-pause` label is set>, STOP immediately and do nothing else. 2. **Read `STATE.md`** (appended to your system prompt). It is your memory of prior runs: the Priority / Watch / Noise lists. 3. **Pick the next unit of work** from the Priority list. Stay strictly within scope: `<scope globs — never *>`. 4. **Do the tier-appropriate action:** - **L1 (report-only):** investigate and summarize. Write NOTHING but `STATE.md`. - **L2 (assisted):** make the change in a **git worktree**; run the gate `<verify>` and guard `<guard>`; if both pass, hand the branch to `<land_via — e.g. fleet-ops>`; otherwise discard. 5. **Apply the escalation rule.** If the action would <escalation — e.g. force-push / push to main / deploy / delete pre-existing files / edit .claude>, do NOT do it — escalate to a human with context instead. 6. **Respect the budget.** Stop this run if you approach `<budget_tokens>` output tokens. 7. **Rewrite `STATE.md`:** promote/demote items across Priority / Watch / Noise, bump the `_Updated_` line + run number + readiness score. 8. **Append one line to `run-log.md`:** `<ISO-Z> run#N action=<…> outcome=<…> tokens=<N>`. ## Hard rules - A general goal is NOT authorization for a specific high-blast action it implies — when in doubt, escalate. - Never act outside `scope`. Never touch another session's `.claude/worktrees/`. - Leave the repo in a clean, reviewable state every tick. -
STATE.template.md 880 B
# <loop-name> — STATE _Updated: <ISO-8601 Z> · run #0 · readiness 100/100_ <!-- The triage snapshot. The loop READS this at the top of every run and REWRITES it at the end. Not a database — a lightweight snapshot of: what to act on, what to watch, what was seen-and-dismissed. Read/write contract: references/state-spine.md. First action of every run: check the kill switch, then read the Priority list. --> ## Priority (act on these next) <!-- the next units of work, highest first. e.g. "[P1] PR #412 failing CI 3h" --> - (none yet) ## Watch (not yet actionable) <!-- things being tracked that aren't ready to action --> - (none yet) ## Noise (seen + dismissed this run) <!-- items deliberately skipped, so the next run doesn't re-surface them --> - (none yet) --- _Source: <scheduler, e.g. .github/workflows/<loop-name>.yml> · config: loop.config.yaml_
-
-
references
-
claude-code-loops.md 17.2 KB
# Where Loops Actually Live in Claude Code The outer loop is a *cadence + a headless run*. This file is the mechanics: the concrete ways to fire a loop in Claude Code, when to use each, and how they compose with the tier model. The doctrine — *a scheduler invokes `claude -p`, not a session that spawns ungated children* — is in [risk-tiers.md](risk-tiers.md); this is the how. The primitives themselves — every parameter, limit and failure semantic, verified and date-stamped — are in [native-scheduling.md](native-scheduling.md); read that before trusting a number here. --- A loop's **trigger** answers *when a tick fires* — a **cadence** (poll on a clock) or an **event** (something pushed in from outside) — and its **completion** rule answers *when the work stops*. Claude Code has native answers to all three. **Prefer the native mechanisms — zero/low-infra, no GitHub Actions.** Reach for an external scheduler only for non-Claude-Code control. ## Cadence — when a tick fires | Mechanism | `host:` | Runs on | Local files? | Open session? | Min interval | Best for | |---|---|---|---|---|---|---| | **`/loop`** (bundled skill) + `CronCreate` | `session-cron` | your machine | ✅ | **yes, idle** | 1 min | supervised, in-session polling (**L1 only** — 7-day expiry) | | **`ScheduleWakeup`** — `/loop`'s dynamic mode | `session-cron` | your machine | ✅ | yes | 60 s–1 h clamp | self-pacing one task; Claude picks each delay | | **Desktop scheduled task** (`scheduled-tasks` MCP) | `desktop-task` | your machine | ✅ | no (app open) | 1 min | **the local-first unattended default** — loops that touch the repo/build/tools | | **Cloud routine** (`/schedule` → [Routines](https://code.claude.com/docs/en/routines)) | `cloud-routine` | **Anthropic cloud** | ❌ **fresh clone** | no | **1 hour** | unattended loops needing **no** local state (GitHub PRs, web, connectors) | | external scheduler + `loop-run.sh` | `external` | your machine | ✅ | no | your call | non-Claude-Code control: cron / Task Scheduler / systemd / process-compose / CI | | **GitHub Actions** | `external` | GH runner | fresh clone | no | — | *optional* — only if the repo already lives on GitHub | Declare the choice as `host:` in `loop.config.yaml`; `loop-doctor` then enforces that host's real constraints instead of assuming a local `claude -p`. > **Three load-bearing caveats, all verified 2026-08-30:** > > 1. **Cloud routines run on a fresh clone with no access to your local files.** A loop > that touches a local repo, build, model dir, or tool **cannot** be a cloud routine — > use a Desktop scheduled task or `/loop`. They also have **no permission mode at all**: > the boundary is repos + environment network policy + connectors, and *every* connected > connector attaches by default. > 2. **`/loop` and `CronCreate` are session-scoped and expire.** In-memory, gone on a new > conversation (`--resume` restores unexpired ones), fire only while the session is > **idle**, and every recurring job **self-deletes 7 days after creation**. That makes > `session-cron` an L1-supervised host — never the home of an unattended loop. > 3. **A Desktop task's worktree toggle is OFF by default**, so a run works against your > working directory *including uncommitted changes*. At L2+ turn it on, or the loop's > "isolation" is imaginary. The unattended options (Desktop task, cloud routine, external scheduler, Actions) are the human-configured **authorizer** — no parent auto-mode session, so nothing blocks the headless child. Many loop frameworks are CI/Actions-centric; loop-ops is runner-agnostic and **native-first** on purpose. ## Event — when something happens (routine triggers, Channels) Polling burns tokens while nothing changes and lags the thing it watches. There are now **three** ways to fire on an event instead of a timer, and they differ in whether a session must stay alive: | Event source | Needs a live session? | Fires | |---|---|---| | **Routine API trigger** — `POST /fire` + bearer token | **no** | your alerting system, deploy pipeline or internal tool starts a cloud run | | **Routine GitHub trigger** — `pull_request` / `release` + filters | **no** | a repo event starts a cloud run | | **Channel** — an MCP plugin pushing into a session | **yes** | anything you can build a receiver for | The routine triggers are the important addition: **a native event loop no longer has to be a kept-alive background session.** An alert-triage or deploy-verification loop is an API trigger; a PR-review loop is a GitHub trigger with filters. Both carry the cloud-routine constraints above. `text` sent to `/fire` arrives wrapped in a `<routine-fire-payload>` block **labelled untrusted** — the prompt must explicitly opt in to acting on it, which is what stops a leaked bearer token from becoming instruction injection. A [**Channel**](https://code.claude.com/docs/en/channels) (v2.1.80+, research preview) is an MCP plugin that **pushes** an external event — a CI failure, an error-tracker alert, a deploy webhook, a chat message — straight into a running session, so the tick fires *on the event* instead of on a timer. - **Cheaper + faster than polling** — no idle ticks; the loop reacts the instant the event lands. The right trigger for `ci-watch`, `pr-watch`, `monitor`. - **The trade-off:** an event arrives only while a session is open, so an unattended event-loop is a **persistent background session** (`claude --channels plugin:<name> …`, or `-p` for non-interactive) kept alive — not a fully-detached cron. Detachment traded for responsiveness. - **Setup:** install a channel plugin (Telegram/Discord/iMessage ship in the preview; build a [webhook receiver](https://code.claude.com/docs/en/channels-reference) for CI/error/ deploy), launch with `--channels`, lock the sender allowlist. Anthropic-auth only (not Bedrock/Vertex/Foundry). - **Still gated** — an event-driven tick runs under the same permission mode + allowlist as any other; a webhook firing the loop never widens what it may do. ## Completion — when the work stops: `/goal` [`/goal <condition>`](https://code.claude.com/docs/en/goal) (v2.1.139+) keeps the session working turn-after-turn until a small fast model confirms the condition holds, then auto-clears — the **native inner-loop gate**. It's the native expression of a loop's `verify`/Until rule: *"keep going until the acceptance criteria hold."* Bound it with `or stop after N turns`. It's a session-scoped **prompt-based Stop hook**, and it pairs with auto mode (auto removes per-*tool* prompts; `/goal` removes per-*turn* prompts). Headless, one tick to completion: ```bash claude -p "/goal all tests in test/auth pass and lint is clean, or stop after 20 turns" ``` `/loop`'s **dynamic mode** carries its own completion rule: Claude ends the loop itself by calling `ScheduleWakeup` with `stop: true` once the task is done, and an iteration that neither reschedules nor stops gets one ~20-minute fallback wakeup before the loop ends. That is a *self-judged* stop, so it is weaker than `/goal`'s explicit condition — use it for exploratory watching, not as a loop's `verify` gate. **The fully-native, zero-external-infra loop** = a **Desktop scheduled task** (local, has files, no open session) that runs `claude -p "/goal <tick condition>"` against the STATE spine. No cron, no Task Scheduler, no Actions. --- ## Which mechanism? — the recipe selector These mechanisms are **not interchangeable** — each has a load-bearing trade-off. Pick by answering: does it need **local code**, is it **connector-driven**, is it **recurring** or **run-to-completion**, and **does token cost matter**? | Your situation | Prescribed recipe | The trade-off that decides it | |---|---|---| | **Connector work, no local code** — triage email, Asana, Slack, calendar, issues via your claude.ai connectors | **Cloud routine** (`/schedule`) | Runs unattended in the cloud and **keeps all your claude.ai connectors** — email/Asana/tools work with your machine *off*. The fresh-clone/no-local-files limit doesn't bite because the work isn't in your repo. (≥1-hour cadence.) | | **Touches local code / build / tools**, unattended | **Desktop scheduled task**, or a **background daemon** running `claude -p` | Both have local files and need no open session. The daemon adds fresh context per tick + deterministic, tunable cost (next row). | | **Sustained / heavy cadence where tokens matter** | a **deterministic daemon** (or cron) firing `claude -p` — **not** `/loop` | `/loop` runs in one *growing* session: context accumulates, tokens climb, quality drifts past ~150k. A daemon fires a **fresh** `claude -p` each tick — bounded cost, no drift — and is deterministic. **Wake it inside the cache TTL you're paying for** so the static `run.md`+system prefix stays warm and each tick reads it at ~0.1×: ~240–270 s for the default 5-minute TTL, or up to ~55 min if the prefix is written with `"ttl": "1h"`. Fresh context *and* cache reads — the cheap sustained-loop recipe. | | **Supervised, light, you're watching** | **`/loop`** | Quickest to start, in-session — perfect for a short burst ("watch this deploy"). But it's **token-hungry if left running heavy**; graduate to a daemon for anything sustained. | | **Long task with a fixed, verifiable end state** — "migrate until tests pass", "split until each file < N lines", "drain the labeled backlog" | **`/goal`** (+ auto mode) | Runs turn-after-turn until a fast model confirms the criteria, then stops — a *completion gate*, not a cadence. Auto mode makes each turn unattended; bound with `or stop after N turns`. | **Cadence × completion compose.** A recurring loop whose every tick should run *to completion* = a cadence mechanism driving `claude -p "/goal <tick condition>"`. E.g. a Desktop task (or daemon) every morning running `/goal` over the issue backlog. ### The economics (why the daemon beats `/loop` at scale) Cadence is the top cost lever, **caching is the next** ([state-spine.md](state-spine.md), [loop-estimate](../scripts/loop-estimate.py)). The two interact: - **`/loop`** keeps one session alive; its input grows every iteration (accumulating transcript), so cost climbs and the cache helps less. Great for short supervised runs. - **A daemon/cron `claude -p`** starts fresh each tick (the Ralph property → flat per-tick cost) and, fired **inside the cache TTL**, keeps the static prefix warm (~0.1× reads). **The TTL is a choice, not a constant:** 5 minutes by default (1.25× write), or 1 hour with `"ttl": "1h"` (2× write) — so the practical daemon window is ~4.5 min *or* ~55 min. `loop-estimate` picks the cheapest TTL that stays warm at your cadence and says which; past 1 h nothing caches. Break-even and the multipliers: [claude-api-ops caching-and-cost](../../claude-api-ops/references/caching-and-cost.md). The same reasoning is why an in-session `/loop` gains less: its input *grows*, so the cached prefix is a shrinking share of each tick — the fresh-context daemon is what keeps the cacheable part dominant. A minimal local daemon (no scheduler infra) — wake under the cache window, fresh context each tick: ```bash # fires loop-run.sh every ~4.5 min: fresh `claude -p`, prefix stays cache-warm (5m TTL) while true; do .loops/<name>/loop-run.sh; sleep 270; done # with a 1h-TTL cache write on the prefix, ~55 min still reads warm: sleep 3300 # or run it under process-compose / a systemd timer / nohup for boot persistence ``` --- ## The external-scheduler shape (when you're not using a native mechanism) Native paths (Desktop task, cloud routine, `/loop`) run the tick prompt — or `claude -p "/goal …"` — **directly**, so they need no wrapper. When you instead drive the loop from an **external** scheduler (cron / Task Scheduler / systemd / process-compose / CI — e.g. for sub-minute cadence or to fit existing infra), `loop-scaffold` scaffolds a **`loop-run.sh`** in the loop dir as the runner-agnostic glue. No GitHub Actions required. ``` any scheduler ──▶ .loops/<name>/loop-run.sh (the authorizer) ├─ kill switch first (PAUSED sentinel) → exit if set ├─ claude -p "$(cat run.md)" --permission-mode dontAsk \ │ --append-system-prompt "$(cat STATE.md)" --allowedTools … └─ git add/commit STATE.md + run-log.md (if in a repo) ``` Wire it with whatever you already run — **no cloud dependency**: ```bash # cron (Linux/macOS): */10 * * * * /path/.loops/pr-watch/loop-run.sh >> /path/.loops/pr-watch/tick.log 2>&1 # Windows Task Scheduler (every 10 min; S4U logon, see windows-ops for the hardened form): schtasks /Create /SC MINUTE /MO 10 /TN pr-watch \ /TR "bash -lc '/c/path/.loops/pr-watch/loop-run.sh'" # process-compose / systemd timer / a while-sleep loop — all work; loop-run.sh is just a script. ``` - The **scheduler** (not a Claude session) invokes `loop-run.sh`. It is the human-configured authorizer; nothing upstream gates the run. - `--permission-mode dontAsk` + a curated allowlist = a **gated** worker that runs anywhere. (For L3 arbitrary-execution jobs, swap to a container + `bypassPermissions` — see the enumerate-vs-isolate fork in [risk-tiers.md](risk-tiers.md).) - The run prompt (`run.md`) is the same every tick — fresh context each time (the Ralph property). State survives in `STATE.md` + the codebase + git, not the conversation. - **GitHub Actions** is one option, not a requirement — the worked example ships an optional `github-actions.yml` for repos already on GitHub; everyone else uses the local schedulers above. ### Why not "a Claude session that launches the loop"? Because an `auto`-mode session that spawns a detached `claude -p --permission-mode bypassPermissions` child is blocked as **Create Unsafe Agents** — an ungated autonomous agent with no human gate. The fix is structural, not a workaround: move the launch to the scheduler. Trying to wrap the bypass flag in a script to dodge the gate is **Auto-Mode Bypass**, a `hard_deny` (see [risk-tiers.md](risk-tiers.md) and the [classifier reference](../../../docs/AUTO-MODE-CLASSIFIER.md)). --- ## Hooks — the loop's reflexes Hooks fire shell commands at points in the agent's lifecycle. Useful loop wiring: | Hook | Loop use | |---|---| | `PreToolUse` | enforce scope/kill-switch before a tool runs (deterministic gate 1) | | `PermissionDenied` | react to a classifier denial — log it, signal a retry, escalate | | `Stop` | write the run-log line + rewrite `STATE.md` as the run ends | | `SessionStart` | load `STATE.md` into context at the top of a run | A `PreToolUse` hook that checks `.loops/<name>/PAUSED` is the cheapest possible kill switch — it blocks every tool the instant the sentinel appears, no matter where the run is. See [`claude-code-ops`](../../claude-code-ops/SKILL.md) for the full 30-event hook catalog and the stdin/stdout JSON contracts. --- ## Composing with the execution layers The cadence fires; the work is done by the layers this repo already ships: ``` /schedule (cadence) └─▶ claude -p (the run; dontAsk + allowlist) ├─▶ iterate # inner improvement loop, if the unit of work is "improve metric X" ├─▶ fleet-worker # spawn cheap parallel makers in worktrees └─▶ fleet-ops # test-gate + land the winning branch └─▶ Stop hook → rewrite STATE.md + append run-log ``` - **`iterate`** when the unit of work is "drive metric X to target in this session". - **`fleet-worker`** when one tick should fan out several maker attempts cheaply. - **`fleet-ops`** as the `land_via` — the sequential, test-gated merge queue that turns a worker's green branch into a landed change (or escalates it). - **`pigeon`** to coordinate across concurrent loops (the priority-order standoff). --- ## A worked L1 → L2 graduation 1. **L1, supervised:** `/loop 15m` in a session (`host: session-cron`), running a read-only "report PR state to STATE.md" prompt. You watch it; it writes nothing but the snapshot. Permission mode `plan`. Remember the 7-day expiry — this host is for the proving period, not the destination. 2. **Prove judgment:** read a week of `STATE.md` snapshots + the run-log. Is its triage right? Does readiness hold? 3. **L2, unattended:** move the host — `desktop-task` if the loop touches local code, `cloud-routine` if it doesn't, `external` for sub-minute cadence — and update `host:` so `loop-doctor` checks the right constraints. Switch the run prompt to "open a fix PR in a worktree" with `--permission-mode dontAsk` + a narrow allowlist (`Bash(npm test)`, `Bash(git …)`). Add a `guard`, set `land_via: fleet-ops`, write the `escalation` rule. Re-run `loop-check` at L2 — fix every error — then enable. The point of the ladder: the cadence mechanism *changes* (session `/loop` → scheduled `claude -p`) exactly when the autonomy does, and the audit gates the transition. ## See also - [native-scheduling.md](native-scheduling.md) — the primitives themselves: verified parameters, limits and failure semantics per host. - [risk-tiers.md](risk-tiers.md) — the permission-mode mapping + scheduler-not-session rule. - [state-spine.md](state-spine.md) — the STATE.md the run reads and rewrites. - [../../claude-code-ops/SKILL.md](../../claude-code-ops/SKILL.md) — the full hook catalog, `claude -p` flags, headless reference. -
failure-modes.md 10.9 KB
# Failure Modes — how loops actually break, and what catches each Incident-shaped scar tissue. Every entry is a real way an outer loop goes wrong: the **symptom** you'd observe, the **mechanism** underneath, and the **catch** — the specific `loop-ops` control (or Claude Code gate) that prevents or surfaces it. Read this before you schedule anything unattended; most of these only bite once you're not watching. The meta-lesson (Addy Osmani): *"build the loop like someone who intends to stay the engineer."* These failures are what happens when the loop is given more autonomy than its judgment has earned. --- ## 1. The runaway-budget loop - **Symptom:** a day's token spend gone in an hour; the bill is 5–10× the estimate. - **Mechanism:** cadence too tight, or scope crept so each tick reads/does far more than scoped (a "report PRs" loop that started crawling diffs). Sub-agents multiply it. - **Catch:** set `budget_tokens` (a per-run ceiling). Estimate with `loop-estimate` *before* scheduling; `loop-doctor` fails the loop if `budget_tokens` < estimated tokens/run. Reconcile the `loop-estimate` estimate against `run-log.md` actuals periodically — a tick that used to cost 2k now costing 40k means scope crept. ## 2. The 3am-dead loop - **Symptom:** every tick aborts immediately; `run-log.md` shows nothing but failures. - **Mechanism:** the `verify`/`guard` gate command's binary isn't on PATH in the *scheduler's* environment (works on your laptop, absent on the CI runner), or `claude` itself isn't installed there. In non-interactive `-p`, a hard denial **aborts the session** — no human to prompt. - **Catch:** `loop-doctor --live` resolves the gate's leading binary and checks `claude`/`git` are on PATH *before* you schedule. Run it in the target environment. ## 3. The cache-cold loop - **Symptom:** cost far higher than `loop-estimate --cached` projected; `cache_read_input_tokens` stays 0. - **Mechanism:** the run prompt isn't byte-identical every tick — a `datetime.now()`, a per-run UUID, or unsorted JSON in the prefix invalidates the cache. Or the cadence is slower than the cache TTL (a 6h loop can't keep a 1h entry warm), so every tick is a cold write. - **Catch:** keep `run.md` byte-identical (the template enforces this — fresh context, same prompt). `loop-estimate` tells you whether the cadence can cache at all and which TTL; if it can't, don't pay the write multiplier — run uncached. ## 4. The force-push / push-to-main loop - **Symptom:** the loop force-pushed, pushed to `main`, or ran a production migration — "to fix the thing." - **Mechanism:** a *general* goal ("keep CI green", "clean up the repo") was taken as authorization for a *specific* high-blast action it merely implied. - **Catch:** the escalation gate — these classes (force-push, push to `main`, prod deploy/migration, mass delete, IAM grants, deleting pre-session files, `.claude` edits) are **always** escalated, declared in `escalation:`. Claude Code's auto-mode classifier also hard/soft-denies them independently: a general goal is *not* explicit intent. ## 5. The ungated-child spawn - **Symptom:** the orchestrator session dies with *Create Unsafe Agents* / *Auto-Mode Bypass*; the loop never starts. - **Mechanism:** a session in `auto` mode tried to launch a detached `claude -p --permission-mode bypassPermissions` child (an ungated autonomous agent). Wrapping the flag in a script to dodge the classifier is a `hard_deny` nothing clears. - **Catch:** the cardinal rule — **a scheduler invokes `claude -p`, not a session that spawns ungated children.** Move the launch to cron/Actions/Task Scheduler (the human authorizer), and give the child gates (`dontAsk` + allowlist), not bypass — unless it's in an isolated container. (`rules/loop-engineering.md` directive #2.) ## 6. The colliding loops - **Symptom:** two loops fight over the same branch/worktree; one clobbers the other's work; merge churn. - **Mechanism:** several loops run against one repo with no coordination. - **Catch:** the multi-loop **priority order** (CI > PR > deps > cleanup > triage) — the higher-priority loop wins worktree contention, lowers defer to their next tick. Each loop isolates in its **own** worktree; they announce what they hold via `pigeon` so a peer can stand off. Never touch another session's `.claude/worktrees/`. ## 7. The silent-stop loop - **Symptom:** nobody noticed the loop stopped running for a week; stale `STATE.md`. - **Mechanism:** the schedule quietly stopped firing — a disabled workflow, a cron typo, a paused runner, an expired token. Loops fail *open* into silence, not error. - **Catch:** treat `STATE.md`'s `_Updated_` timestamp + the `run-log.md` tail as a heartbeat — if the latest run is older than ~2× the cadence, the loop is down. A separate cheap monitor (or a `daily-scan` loop) that flags stale loop heartbeats closes this; the kill switch is for stopping, the heartbeat is for noticing it stopped. ## 8. The test-deleting "fix" (gate reward-hacking) - **Symptom:** CI is green again — because the loop deleted or `skip`-ped the failing test, not because it fixed the bug. - **Mechanism:** the loop optimized the literal gate (`verify` passes) rather than the intent. A loop, like any optimizer, games a weak metric. - **Catch:** make the gate hard to hack — a `guard` that runs the **full** suite + typecheck, a `scope` that **excludes** test files and CI config, and a human review at L2 (the PR gate). Never let an L2/L3 loop modify the very tests that gate it. ## 9. The unbounded-scope loop - **Symptom:** the loop edited files far outside its job. - **Mechanism:** `scope: "*"` (or `**`, or empty) — "may touch anything." - **Catch:** `loop-check` **rejects** an unbounded or placeholder scope (exit 10). Scope is bounded globs, always. ## 10. The no-kill-switch loop - **Symptom:** the loop is misbehaving and there's no fast way to stop it. - **Mechanism:** no stop signal was designed in; stopping means disabling the workflow by hand mid-tick. - **Catch:** `kill_switch` is mandatory (`loop-check` errors without one) and checked **first** every run. The cheapest implementation is a `PreToolUse` hook that blocks every tool the instant a `PAUSED` sentinel appears — an instant breaker. ## 11. The comprehension-debt loop - **Symptom:** the codebase works but no one on the team understands the changes the loop shipped; onboarding slows, incidents take longer. - **Mechanism:** an unattended loop shipped correct-but-unreviewed changes for weeks; comprehension debt compounded silently. - **Catch:** the tier ladder is the antidote — **start at L1 (report-only)** and *read the reports*; graduate to L2 only once you trust its judgment, and keep the human in the PR loop. Autonomy is earned with evidence, not granted up front. Build the loop like you intend to stay the engineer. --- ## 12. The expired loop (native-host) - **Symptom:** a `/loop`-scheduled watch that ran fine for a week is simply gone. No error, no final report — and `CronList` shows nothing. - **Mechanism:** `CronCreate` recurring jobs **self-delete 7 days after creation** (they fire one final time first), and the whole session store is in-memory: a new conversation clears it, and `durable` is documented as having no effect. A loop hosted on `session-cron` has a hard, silent lifetime ceiling. This is silent-stop (#7) with a cause you cannot fix by watching the scheduler. - **Catch:** don't host anything unattended on `session-cron`. Declare `host: session-cron` and `loop-doctor` refuses it at L2+; at L1 it warns. For anything that must outlive a week, use `desktop-task`, `cloud-routine` or `external`. ## 13. The over-connected routine (native-host) - **Symptom:** a read-only "summarise my inbox" routine sent an email / closed an issue / posted to Slack. - **Mechanism:** cloud routines have **no permission mode and no approval prompts**, and **every connected connector is attached by default** with full write access. Nothing between the prompt and the tools — the "read-only" was only ever a wording in the prompt. A leaked `/fire` bearer token compounds it (fire text is at least wrapped as untrusted, so it can't issue instructions the prompt didn't opt into). - **Catch:** prune connectors to the minimum on every routine; scope the environment's network access; select only the repos the work needs. `loop-doctor` refuses a `cloud-routine` config that doesn't name that boundary, precisely because there is no permission mode to fall back on. ## 14. The stalled task (native-host) - **Symptom:** a Desktop scheduled task shows a session open in the sidebar and no output for hours, or a task "runs" daily but the run history is mostly *skipped*. - **Mechanism:** three separate defaults. A task in Manual permission mode that needs an unapproved tool **stalls waiting for a human** rather than failing (as does any MCP tool marked `requiresUserInteraction`, on every call). Tasks fire only while the app is open and the machine is awake; a sleeping machine skips the run. And a machine that was asleep all day gets **exactly one** catch-up for the most recent missed window, so a 9am task can execute at 11pm. - **Catch:** click *Run now* once after creating a task and always-allow each tool it needs; then put **time guardrails in the run prompt itself** ("only review today's commits; if it's after 5pm, skip and summarise what was missed") — the loop cannot assume it is running when it was scheduled to. ## At a glance — symptom → control | Failure | Primary control | |---|---| | Runaway budget | `budget_tokens` + `loop-estimate` + `loop-doctor` budget check | | 3am-dead | `loop-doctor --live` (gate binary + PATH) | | Cache-cold | byte-identical `run.md` + `loop-estimate` TTL guidance | | Force-push / prod | escalation gate + auto-mode classifier | | Ungated-child spawn | scheduler-invokes-`claude -p` (rule #2) | | Colliding loops | priority order + per-loop worktree + `pigeon` | | Silent-stop | `STATE.md`/run-log heartbeat staleness | | Gate reward-hacking | full-suite `guard` + scope excludes tests + human PR gate | | Unbounded scope | `loop-check` rejects `*` | | No kill switch | mandatory `kill_switch` + PreToolUse PAUSED hook | | Comprehension debt | L1-first graduation; read the reports | | Expired loop (7-day) | never host unattended on `session-cron`; `loop-doctor` host check | | Over-connected routine | prune connectors + scope environment; `loop-doctor` boundary check | | Stalled / skipped task | always-allow the tools once; time guardrails in `run.md` | ## See also - [native-scheduling.md](native-scheduling.md) — the per-host limits behind #12–#14. - [risk-tiers.md](risk-tiers.md) — the graduated-autonomy ladder behind #11. - [state-spine.md](state-spine.md) — budget + heartbeat + multi-loop coordination. - [../../../rules/loop-engineering.md](../../../rules/loop-engineering.md) — the directives that prevent #4/#5/#9/#10. -
native-scheduling.md 17.1 KB
# Native Scheduling Primitives — what the harness now ships, and what it doesn't **Verified 2026-08-30** against the live tool schemas in-session (`CronCreate`/`CronList`/ `CronDelete`, the `scheduled-tasks` MCP server, `ScheduleWakeup`) and the current docs: [scheduled-tasks](https://code.claude.com/docs/en/scheduled-tasks), [desktop-scheduled-tasks](https://code.claude.com/docs/en/desktop-scheduled-tasks), [routines](https://code.claude.com/docs/en/routines). These are fast-moving surfaces — **re-verify before trusting a number here.** Where the docs and a tool description disagree, this file says so rather than picking a winner. [claude-code-loops.md](claude-code-loops.md) owns *which mechanism to pick and how to wire it*. This file owns *what each primitive actually is*: parameters, limits, and the failure semantics that decide whether a loop survives a night. --- ## The four hosts A loop's **host** is where its ticks execute. It is not a style preference — it changes what the loop can reach, what can stop it, and what "it didn't run" means. | Host | Primitive | Executes on | Local files | Needs a session | Min cadence | Survives restart | |---|---|---|---|---|---|---| | `session-cron` | `CronCreate` / `/loop` | your machine | ✅ | **yes, open + idle** | 1 min | only via `--resume`, if unexpired | | `desktop-task` | `scheduled-tasks` MCP / Routines→**Local** | your machine | ✅ | no (app open) | 1 min | ✅ on disk | | `cloud-routine` | `/schedule` → Routines→**Cloud** | Anthropic cloud | ❌ fresh clone | no | **1 hour** | ✅ | | `external` | cron / Task Scheduler / systemd → `loop-run.sh` | your machine | ✅ | no | yours | ✅ | | `local` | *undeclared* — a generic local run | your machine | ✅ | — | — | — | `local` is the template default and means "not yet decided": `loop-doctor` applies only the host-agnostic checks. It is fine while scaffolding and while a loop is still L1, but **pick a real host before scheduling** — a loop whose execution surface nobody named is a loop whose limits nobody checked. The doctrine is unchanged and now has a native shape: **the authorizer is the scheduler, never a session that spawns ungated children** ([risk-tiers.md](risk-tiers.md)). Every host above except `session-cron` is a human-configured authorizer. --- ## `session-cron` — `CronCreate` / `CronList` / `CronDelete`, and `/loop` The in-session scheduler. `/loop` is a **bundled skill** that drives these tools; you do not author it. **Verified surface.** `CronCreate` takes a standard 5-field cron expression evaluated in **local time** (`minute hour day-of-month month day-of-week`), a `prompt`, and `recurring` (default `true`; `false` = fire once then auto-delete). It returns an 8-character job ID for `CronDelete`. `CronList` lists them. Wildcards, steps, ranges and lists are supported; extended syntax (`L`, `W`, `?`, `MON`/`JAN` aliases) is not. When day-of-month and day-of-week are both constrained, a date matches if **either** does (vixie-cron semantics). **The limits that decide whether you may use it for a real loop:** - **Session-scoped and in-memory.** The `durable` parameter is present but documented as having **no effect** — "durable persistence is not available". A new conversation clears every task; `--resume` / `--continue` restores unexpired ones. *One caveat worth knowing before you trust that flatly:* the docs describe a narrow case where a task you ask to keep across sessions **is** written to the project's `.claude` directory — when feature-flag fetching is off (and it errors if that path is a symlink). So "no effect" is what the tool reports in a normal session, not a universal law. Either way it is not a foundation for an unattended loop; don't design around it. - **Seven-day expiry.** A recurring task fires one final time 7 days after creation, then deletes itself. This is a hard ceiling on unattended lifetime. - **Fires only while the REPL is idle.** Not mid-response. If Claude is busy when a task comes due, it waits for the turn to end. - **No catch-up.** A missed window fires **once** on return to idle, never once per missed interval. - **Jitter, and the sources disagree.** The docs say recurring tasks fire up to **30 min** late (or up to half the interval for sub-hourly jobs); the in-session tool description says up to **10% of the period, max 15 min**. Both agree one-shots at `:00`/`:30` can fire up to 90 s *early*, and that the offset is derived from the task ID (so it is stable per task). **Do not build a loop that depends on exact fire times.** Picking a minute that is not `:00`/`:30` avoids the one-shot jitter and spreads API load. - **50 tasks per session.** - `CLAUDE_CODE_DISABLE_CRON=1` disables the scheduler entirely — cron tools *and* `/loop`. **Verdict for loop-ops:** `session-cron` is an **L1 supervised** host only. It cannot host an unattended L2/L3 loop: it needs an open idle session and evaporates after 7 days. `loop-doctor` treats `host: session-cron` at L2+ as a predicted runtime failure. ### `/loop` — the bundled skill, and its dynamic mode `/loop` is not something you author; it ships with the harness and reads its argument in three shapes: | You type | Behaviour | |---|---| | `/loop 5m <prompt>` | fixed cron cadence. `s`/`m`/`h`/`d`; seconds round up to a minute; awkward steps (`7m`, `90m`) round to the nearest clean cron step and Claude says what it picked | | `/loop <prompt>` | **dynamic (self-paced)** — Claude picks each delay itself | | `/loop` (bare) | the built-in maintenance prompt, self-paced | A skill can be the prompt (`/loop 20m /review-pr 1234`), but a scheduled fire only runs skills Claude may invoke on its own — built-ins (`/permissions`, `/model`), skills marked `disable-model-invocation: true`, skill deny-rules and MCP prompts arrive as plain text instead of executing. **A loop whose tick is a slash command must check that the command is model-invocable**, or every tick silently no-ops. **Dynamic mode** is driven by the `ScheduleWakeup` tool. Each iteration Claude calls it with a `delaySeconds` **clamped to [60, 3600]**, a `reason` shown back to the user, and a `noop` flag (`true` = nothing changed; consecutive noop ticks collapse in the transcript). `stop: true` ends the loop immediately. If an iteration neither reschedules nor stops, one fallback wakeup fires ~20 minutes later and the loop ends if that one doesn't reschedule either. `Esc` clears a pending self-paced wakeup. **The two sentinels.** An autonomous `/loop` with no user prompt passes a literal sentinel back as its `prompt` so the runtime can re-resolve the instructions at fire time. They are **not interchangeable**: - `<<autonomous-loop-dynamic>>` — for `ScheduleWakeup` (self-paced mode) - `<<autonomous-loop>>` — for the `CronCreate` fixed-cadence mode Passing the wrong one wires an autonomous loop to the wrong pacing engine. **Customising the default.** `.claude/loop.md` (project, wins) or `~/.claude/loop.md` (user) replaces the built-in maintenance prompt for a bare `/loop`. It is ignored whenever you supply a prompt. Edits take effect on the next iteration — you can refine a running loop's instructions in place. Content past **25,000 bytes is truncated**. **Polling vs pushing.** The docs are explicit that where the `Monitor` tool is available, streaming a background script's output beats re-running a prompt on an interval — cheaper and more responsive. Reach for a cadence only when there is nothing to stream. --- ## `desktop-task` — the `scheduled-tasks` MCP server The durable local host, and the closest native analogue to a loop-ops loop. **Verified surface.** Four tools: `create_scheduled_task` (`taskId` kebab-case, `prompt`, `description`, plus **at most one** of `cronExpression` (recurring, local time) or `fireAt` (ISO-8601 with offset, one-time, auto-disables after firing) — omit both for an ad-hoc task that only runs manually; `notifyOnCompletion` defaults true), `list_scheduled_tasks` (returns `taskId`, schedule, `enabled`, `nextRunAt`, `lastRunAt` and a `path` to the task's `SKILL.md`), `update_scheduled_task` (partial; `enabled: false` pauses), `delete_scheduled_task` (leaves the `SKILL.md` on disk so the prompt is recoverable). **On-disk shape.** Each task is `<config-dir>/scheduled-tasks/<task-id>/SKILL.md` — YAML frontmatter carrying `name` and `description`, body = the prompt. Schedule, folder, model and enabled state live **outside** that file (edit them through the app or by asking). The config dir is `~/.claude` unless `CLAUDE_CONFIG_DIR` overrides it. **What this gives a loop for free — and the caveats:** | Native feature | What it replaces | The caveat that still bites | |---|---|---| | The task folder | a place for `STATE.md` / `run-log.md` | nothing is written for you; the *prompt* must read and rewrite them | | Fresh session per run | the Ralph property | **no memory of the creating conversation** — the prompt must be fully self-contained | | Per-task permission mode + saved always-allow approvals | `--permission-mode` on a wrapper | a task in Manual mode that hits an unapproved tool **stalls** until you answer — the classic 3am-dead loop, natively | | Worktree toggle | `worktree: true` + manual setup | **off by default** — a task runs against your working dir *including uncommitted changes* | | Active/Paused status toggle | the kill switch | pausing is out-of-band; an in-prompt sentinel check still stops a run *mid-tick* | | Run history incl. skipped runs + reasons | part of the run-log | it records *that* a run happened, not what the loop decided | **Failure semantics you must design around:** - **Only runs while the app is open and the machine is awake.** Sleep through a window and the run is skipped. - **Exactly one catch-up.** On launch or wake, Desktop starts one catch-up run for the *most recently* missed time within the last 7 days and discards everything older. A daily task that missed six days runs **once**. The docs' own advice is the right advice: put time guardrails in the prompt ("only review today's commits; if it's after 5pm, skip and post a summary of what was missed"). - **Deterministic stagger** of a few minutes after the scheduled time. - MCP tools marked `requiresUserInteraction` prompt every call and stall the run each time. - A task can call `update_scheduled_task` on **itself** to change its own schedule or prompt. That is genuinely useful (reschedule earlier when a release branch appears) and it is also **self-modification** — put it on the escalation list unless the loop's stated purpose is adaptive cadence. --- ## `cloud-routine` — Routines (`/schedule`) Research preview. Runs on Anthropic-managed cloud infrastructure, so it keeps working with the machine off — at the cost of the local filesystem. **Triggers are no longer cadence-only.** A routine may carry any combination of: - **Schedule** — presets (hourly/daily/weekdays/weekly) or a cron set via `/schedule update`; **minimum interval one hour, faster expressions are rejected**. Also one-off runs at a timestamp, which auto-disable after firing and do **not** count against the daily run cap. - **API** — a per-routine `/fire` endpoint. `POST` with a bearer token starts a run and returns a session URL. An optional `text` field carries run-specific context. - **GitHub** — `pull_request` and `release` events, with filters (author, title, body, base/head branch, labels, draft, merged) combined by equals / contains / starts-with / one-of / regex. `matches regex` tests the **whole** field: use `.*hotfix.*`, not `hotfix`. The API trigger matters for loop-ops: it is a **native event trigger that needs no persistent session**, which is the thing [Channels](claude-code-loops.md) could not offer. An `event`-triggered `monitor` or `ci-watch` loop no longer has to be a kept-alive session — it can be an alerting system POSTing to `/fire`. **Two security properties worth encoding in the loop's design:** - **Fire text is untrusted by construction.** The `text` payload arrives wrapped in a `<routine-fire-payload>` block labelled as untrusted data. A routine's prompt must *opt in* by referencing the payload explicitly, or the text is inert context. Anyone holding the bearer token can send it, so this wrapper is the control that keeps a leaked token from becoming instruction injection. Treat it exactly as [prompt-injection-defense](../../prompt-injection-defense/SKILL.md) treats any ingested content. The token is shown **once** at generation; rotate via Regenerate/Revoke. - **A native escalation gate on pushes.** Claude pushes to `claude/`-prefixed branches freely; a push to any other branch is **rejected** if the branch is protected, someone else has an open PR from it, or it carries commits authored by someone else. That is close to loop-ops' "never push to main" rule, enforced by the platform. **No permission mode at all.** Routines "run autonomously as full Claude Code cloud sessions: there is no permission-mode picker and no approval prompts during a run." The boundary therefore *cannot* come from a permission mode — it comes from three other places, and scoping them is the entire safety story: 1. **Repositories** selected (each cloned fresh from its default branch) 2. **Environment** network policy — the Default environment is *Trusted*, allowing only the default allowlist; off-list hosts fail `403 x-deny-reason: host_not_allowed` 3. **Connectors** — **all connected connectors are attached by default**, and Claude may use every tool from an included connector, writes included, without asking. Remove everything the routine does not need. This is the single most common over-grant. **Other limits:** research preview (surface may change); the `/fire` endpoint ships behind a dated beta header; GitHub webhook events have per-routine and per-account hourly caps and events beyond them are **dropped**; there is a daily cap on runs started per account (one-off runs exempt); routines belong to an individual account and act as that identity; Team/Enterprise Owners can disable them org-wide. **And the trap that looks like success:** a green status in the run list "means the session started and exited without an infrastructure error. It does not mean the task in your prompt succeeded." A loop that grades itself on run status is grading the wrong thing — which is exactly why the loop's own `verify` gate stays load-bearing. --- ## What the native primitives replaced — and what they did not The plumbing is theirs now. The discipline is still yours. | Loop-ops primitive | Native answer (2026-08-30) | Still yours to build | |---|---|---| | **Schedule** | ✅ all four hosts | picking the host against the constraint, not the habit | | **Fresh context per tick** | ✅ desktop-task, cloud-routine | writing a genuinely self-contained prompt | | **Isolation** | ~ desktop-task worktree toggle (**off by default**) | worktree at L2+, and verifying it is on | | **Kill switch** | ~ Paused toggle, `Esc`, `CronDelete` | an **in-prompt** sentinel check — the out-of-band ones can't stop a tick already running | | **Run log** | ~ run history / skipped-run reasons | what the loop *decided* and what it cost | | **State between ticks** | ❌ (the task folder is a place, not a spine) | `STATE.md` — [state-spine.md](state-spine.md) | | **Budget** | ❌ (a daily run cap is not a token budget) | `budget_tokens`, enforced in the run prompt | | **The verify gate** | ❌ (green status ≠ success) | `verify:` — and it is an eval, see below | | **The escalation rule** | ~ routines' branch-push guard only | the full never-auto-land list | | **The tier ladder** | ❌ | L1 → L2 → L3, earned | **A loop's `verify` gate is an eval.** Everything the eval discipline says about scoring — outcome vs step vs trajectory, `pass^k` over `pass@k` for anything non-deterministic, judge bias, and gates that are blocking rather than advisory — applies to the gate that decides land-vs-escalate. Invoke the **`evals-ops`** skill when the gate is a judgement call rather than a green test run; a gate you cannot trust is a loop you cannot graduate. --- ## Choosing a host — the short version ``` Needs local files / build / tools? ├─ no → cloud-routine (machine off; ≥1h; scope repos+env+connectors, no perm mode) └─ yes → unattended? ├─ no → session-cron (/loop) L1 supervised only; 7-day expiry └─ yes → desktop-task (durable, per-task perms, worktree toggle) └─ need sub-minute cadence, or non-Claude-Code control? → external + loop-run.sh ``` Declare the answer as `host:` in `loop.config.yaml`. `loop-doctor` checks the loop against its host's real constraints — a cadence faster than the host allows, a local PATH check that proves nothing about a cloud run, an unattended tier on a session-scoped host. ## See also - [claude-code-loops.md](claude-code-loops.md) — which mechanism, and how to wire it. - [risk-tiers.md](risk-tiers.md) — L1/L2/L3 ↔ permission modes; the scheduler-not-session rule. - [state-spine.md](state-spine.md) — the STATE/run-log/budget spine none of these hosts provide. - [failure-modes.md](failure-modes.md) — the incident catalog these limits produce. -
pattern-catalog.md 8.8 KB
# Pattern Catalog — a morphology of loop shapes Loops aren't a fixed list of recipes — they're **compositions of three orthogonal axes**. Name the axes and the patterns fall out; you can also compose ones not named here. The named patterns below are the well-trodden *points* in this space. `loop-scaffold` seeds a `loop.config.yaml` keyed by `--pattern <name>` (the canonical keys); the rest of the space you compose by hand. **Start every pattern at L1** and graduate only once its reports prove its judgment. ## The three axes **1. Trigger — what starts a tick.** | Trigger | Fires when | Mechanism | Best for | |---|---|---|---| | `cadence` | a clock interval elapses | `/loop` (supervised), Desktop task, cloud routine, or a daemon | steady polling — backlog, PRs, deps | | `event` | an external thing happens (CI fail, error, deploy, message) | a cloud routine's **API `/fire`** or **GitHub** trigger (no session needed), or a **Channel** (MCP receiver) pushing into a live session | responsiveness + low cost — no idle polling | | `goal` | runs continuously **until a condition holds**, then stops | `/goal` (+ auto mode) | run-to-completion — migrations, metric targets | > **Event beats poll when you can get it.** A CI webhook firing the tick is cheaper and > faster than a 10-min poll — and it no longer has to cost you detachment. A **routine API > trigger** (`POST /fire` with a bearer token) or a **GitHub trigger** starts a cloud run > with no session alive at all; only a **Channel** needs a persistent session > (`claude --channels …` in a background process, research-preview, Anthropic-auth only). > Pick the routine trigger when the work can run in the cloud, the Channel when it must > touch local state. See [claude-code-loops.md](claude-code-loops.md). **2. Posture — how much autonomy** (the [risk tier](risk-tiers.md)): `L1` report · `L2` propose-and-human-gates · `L3` autonomous-in-a-denylist. **3. Locus — where it runs / what it can touch.** | Locus | Mechanism | Can touch | Use when | |---|---|---|---| | `connector` | **cloud routine** (`/schedule`) | your claude.ai connectors (email, Asana, Slack, issues) — **no local files** | the work lives in services, not your repo | | `local` | Desktop task / daemon / `/loop` | the repo, build, models, local tools | the work touches local state | **Locus is `host:`.** The axis is not decorative — write the resolved answer into the config's `host:` field (`cloud-routine` for `connector`; `desktop-task`, `session-cron` or `external` for `local`) so `loop-doctor` enforces that surface's real limits instead of assuming a local `claude -p`. Per-host limits: [native-scheduling.md](native-scheduling.md). The recipe-selector in [claude-code-loops.md](claude-code-loops.md) is just these axes resolved to a mechanism. A loop = **(trigger × posture × locus) + the [state spine](state-spine.md)**. --- ## The catalog Each row: the axes, the recommended native mechanism, the job (gate → what it escalates), and the **failure mode to watch** ([failure-modes.md](failure-modes.md)). | Pattern | Trigger · Locus | Start tier | Mechanism | Job → escalates | Watch | |---|---|---|---|---|---| | `daily-scan` | cadence · local | L1 | Desktop task (off-peak) | sweep backlog/alerts, write `STATE.md` → all to a human | silent-stop | | `pr-watch` | event\|cadence · connector | L1 | cloud routine on a **GitHub `pull_request` trigger** (no session needed), or a Channel | flag stuck/failing/conflicted PRs → never merges | runaway tokens if polled tight | | `ci-watch` | **event** · local | L2 | Channel (CI webhook) → fix in a worktree | failing test passes + full guard → flaky/deploy/secrets | gate reward-hacking | | `dep-bump` | cadence · local | L2 | Desktop task/daemon | patch-only behind cooldown + guard → minor/major, advisories | supply-chain | | `changelog-gen` | event(on tag)\|cadence · local | L1 | tag-event or Desktop task | draft `RELEASE_NOTES_DRAFT.md` → human publishes | — | | `merge-hygiene` | cadence · local | L1 | Desktop task (off-peak) | dead branches / stale flags → ambiguous deletes | worktree-boundaries | | `issue-sort` | cadence\|event · connector | L1 | cloud routine | classify + suggest labels → priority/dupe-close | — | | **`metric-chase`** | **goal** · local | L2 | `/goal` driving [`iterate`](../../iterate/SKILL.md) | drive coverage/latency/bundle/**eval-score** to target → unreachable / guard fails | gate reward-hacking · **high cost** | | **`regression-watch`** | cadence\|event(on release) · local | L1→L2 | Desktop task/daemon | run a benchmark/eval, diff vs baseline → a real regression | flaky bench = false alarm · **high/run** | | **`digest`** | cadence · **connector** | L1 | **cloud routine** | summarize email/Asana/calendar/news → nothing (read-only) | over-scoped connector | | **`backfill`** | **goal** · local | L2/L3 | `/goal` (+ worktree/container) | drain a migration/queue **to completion** → an item needing judgment | runaway budget · **long** | | **`monitor`** | **event** · local | L1 | **Channel** (error/log/deploy webhook) | triage the event → page a human on anomaly | alert fatigue · needs a live session | | **`freshness`** | cadence · local | L1 | Desktop task (daily/weekly) | re-check docs/data/deps vs reality → confirmed drift | transient failure ≠ drift | --- ## Notes on the patterns that need them - **`ci-watch` / `pr-watch` — prefer event over poll.** A polled `pr-watch` at 5 min costs ~3× a 15-min one for marginal freshness; the event-driven version costs ~nothing while quiet. Two ways to get the event now: a cloud routine's **GitHub trigger** (`pr-watch`) or **API `/fire`** from your CI (`ci-watch`) — neither needs a session alive — or a **Channel** when the tick must touch local state. At L2, `ci-watch` opens a fix in a worktree and hands the branch to `fleet-ops`; never auto-merges `main`. - **`metric-chase` is the bridge to [`iterate`](../../iterate/SKILL.md).** The loop's *trigger* is a `/goal` ("coverage ≥ 90, or stop after N turns"); the *work* each turn is an `iterate` step (modify → measure → keep/discard). Use it for any measurable target — including an **eval score** (this is the GLM/Opus-bench shape). Highest cost class; bound it. - **`digest` is the canonical cloud-routine pattern.** It needs *connectors, not code*, so it's the one archetype where the fresh-clone cloud routine is exactly right — it keeps your claude.ai connectors and runs with the machine off. Read-only: no write scopes. - **`backfill` is run-to-completion, not recurring.** A `/goal` drains the queue/migration; when the condition holds it stops and clears. Bound it (`or stop after N`, a token budget) — it's the runaway-budget risk made flesh. For arbitrary execution, run it in a container. - **`monitor` is the purest event loop.** An error-tracker/deploy webhook → a Channel → a persistent background session that triages and pages on anomaly. No polling at all. The trade-off is keeping that session alive. - **`regression-watch`** runs a real suite each tick (expensive), so cadence it slowly or trigger it on a release event. Treat a transient/flaky failure as advisory (don't page on one red run) — the same exit-7-vs-exit-10 discipline our staleness verifiers use. --- ## Composing a pattern not in the catalog Pick a point in the space the named patterns don't cover. Examples: - *event · connector · L1* — a Slack message (Channel) triggers a read-only lookup against a connector. (A "support-triage" loop.) - *goal · connector · L2* — drain an Asana backlog to empty via `/goal`, updating tasks through the connector. - *cadence · local · L3* — a nightly autonomous refactor in an isolated container. The discipline is identical regardless of the point: bounded scope, a gate, an escalation rule, a kill switch, a budget — and **start at L1**. ## Choosing — the short version 1. **Locus first:** does it touch local code? → `local` (Desktop task/daemon). Pure connector work? → `connector` (cloud routine). 2. **Trigger next:** is there an event to react to? → `event` (Channel) — cheaper + faster. A clear finish line? → `goal`. Otherwise → `cadence`, slowest that still catches the work. 3. **Posture:** start **L1**. Graduate to L2 (with a guard, worktree, escalation, `land_via`) only once the reports earn it; re-run `loop-check` + `loop-doctor --live` at the new tier. ## See also - [risk-tiers.md](risk-tiers.md) — the posture axis (permission-mode mapping). - [claude-code-loops.md](claude-code-loops.md) — the trigger/locus axes resolved to mechanisms + the recipe selector. - [failure-modes.md](failure-modes.md) — the "watch" column, in depth. - [state-spine.md](state-spine.md) — the multi-loop priority order these share. - [../assets/loop.config.template.yaml](../assets/loop.config.template.yaml) — the config every pattern fills in. -
risk-tiers.md 9.4 KB
# Risk Tiers ↔ Claude Code's permission model The single best idea in loop engineering is **graduated autonomy**: a loop earns the right to act unattended, it isn't granted it. This file maps the L1→L2→L3 ladder onto Claude Code's *actual* permission machinery — which is what makes this skill more than a generic-agent methodology. The authority for the gate behaviour is the repo's [auto-mode-classifier reference](../../../docs/AUTO-MODE-CLASSIFIER.md); read it for the full two-gate model. This file is the loop-specific projection. --- ## The ladder ``` L1 Report ───────► L2 Assisted ───────► L3 Unattended read-only suggest + human-gate autonomous within a denylist (plan/dontAsk) (dontAsk/auto) (bypassPermissions, ISOLATED only) ``` **Never skip a rung.** A fresh loop starts at L1. It graduates only after its reports prove its judgment over real runs. Each rung adds exactly one new power and one new guardrail. | | L1 Report | L2 Assisted | L3 Unattended | |---|---|---|---| | **Posture** | discovery + triage | propose changes | autonomous land | | **Writes?** | no — report only | yes, in a worktree | yes, allowlisted classes | | **Permission mode** | `plan` or `dontAsk` + read allowlist | `dontAsk` + narrow allowlist, or `auto` | `bypassPermissions` **in a container** | | **Required guardrails** | bounded scope, kill switch | + guard command, + worktree, + escalation | + denylist, + isolation boundary, + budget cap | | **Lands by** | a human reads the report | a human approves the PR (or `fleet-ops`) | the loop, inside its boundary | | **Blast radius** | zero (no writes) | one PR, reviewable | bounded by the denylist + container | --- ## How each tier maps to a permission mode Claude Code has six permission modes. Loops use four of them: | Mode | Behaviour | Loop tier | |---|---|---| | `plan` | read/explore only; cannot edit | L1 (strictest) | | `dontAsk` | auto-**denies** anything not pre-approved; read-only Bash always allowed; fully non-interactive | L1 / L2 (**recommended default for workers**) | | `auto` | a classifier model gates each unresolved action; "trust the direction" autonomy | L2 (long runs) | | `acceptEdits` | in-scope edits + common fs commands auto-approved; other Bash needs an allow rule | L2 (edit-heavy, known command set) | | `bypassPermissions` | no gates at all | L3 — **only** inside an isolated container/VM without internet | `default` (prompt each action) is interactive — not for unattended loops. `acceptEdits` is the middle option when the command set is known. ### Why `dontAsk` is the workhorse for L1/L2 workers `dontAsk` is fully non-interactive (it never prompts; it auto-denies the unknown), so it runs anywhere — no container required — and read-only Bash is always allowed. Pair it with a **narrow** `permissions.allow` list (`Bash(npm test)`, `Bash(git status)`) and you get a worker that can do exactly its job and nothing else. This is the safe default for headless loop workers. --- ## The headless-profile table (what a `claude -p` worker should use) The loop's *maker* runs are headless `claude -p` sessions. Pick the least privilege that still lets the job run: | Profile | Behaviour | Use for | |---|---|---| | `--permission-mode dontAsk` + curated `permissions.allow` | auto-denies anything not pre-approved; read-only Bash allowed; non-interactive | **locked-down workers (recommended default)** | | `--permission-mode auto` | classifier-gated; configure `autoMode.environment` for your infra. In `-p`, repeated blocks abort the session | long "trust-the-direction" runs | | `--permission-mode acceptEdits` + allow rules | edits + common fs auto-approved; other Bash needs an allow rule | edit-heavy tasks, known command set | | `--dangerously-skip-permissions` (= `bypassPermissions`) | no gates; refuses root/sudo; `rm -rf /`\|`~` still circuit-break | **only** in an isolated container/VM/devcontainer without internet | In **non-interactive `-p` mode** a hard denial **aborts the session** (there's no human to prompt). So an `auto`-mode worker that hits a wall dies; a `dontAsk` worker with a correct allowlist never hits one. This is why enumerating permissions beats relying on the classifier for batch workers. --- ## The cardinal rule: scheduler invokes `claude -p`, not session-spawns-loop This is the one thing a generic-agent methodology can't tell you because it isn't grounded in Claude Code's gate. **An unattended loop must be a scheduler/script that invokes `claude -p` — not a Claude session that tries to launch the loop.** Why: the auto-mode classifier evaluates tool calls *inside* an auto-mode session. A session that tries to spawn a detached `claude -p --permission-mode bypassPermissions` child is blocked as **Create Unsafe Agents** (an ungated autonomous agent with no human gate). Two independent fixes, combine for best result: 1. **Move the launch outside the auto-mode session.** A human — or a human-configured Task Scheduler / cron / CI runner / plain script — running `claude -p …` is the authorizer, with no parent classifier in the loop. Don't run the *orchestrator* session itself in auto mode if its job is spawning agents. 2. **Give the child gates instead of bypass.** The denial is about the *ungated* property, not headless-ness. A `dontAsk`+allowlist child is gated and runs fine. > **Subagents can't escalate.** Agent/Task subagents inherit the parent's mode; the > classifier uses the parent mode and ignores `permissionMode` in subagent frontmatter. > A full-bypass worker fleet must be the isolated-container path launched *outside* the > auto-mode session — never an in-session subagent. --- ## The real fork: enumerate vs isolate When a loop needs real power, there are exactly two legitimate shapes. Reaching for `bypassPermissions` on the host *to avoid enumerating permissions* is precisely the pattern the classifier blocks. | | **Enumerate** | **Isolate** | |---|---|---| | Shape | `dontAsk` + a curated allowlist | container/VM + `bypassPermissions` | | Runs | anywhere (host, CI, laptop) | only inside the sandbox | | Safety | bounded by the allowlist | bounded by the container | | Cost | you list the commands once | you stand up isolation | | Best for | most loops; CI/PR/dep workers | heavy autonomous refactors, untrusted-input runs | **Default to enumerate.** Reach for isolate only when the job genuinely needs arbitrary execution *and* you have a real sandbox (no internet, can't damage the host). --- ## Connector & MCP scopes — least privilege for the loop's tools A loop is only as safe as the tools it can reach. The permission *mode* gates Bash + file edits; the *tool surface* gates everything else — MCP connectors (Slack, GitHub, Jira, a DB), `WebFetch`, the `Agent` tool. Scope them per tier: - **Allowlist, don't blanket.** Headless, name exactly what the job needs: `--allowedTools 'Bash(gh pr list:*)' 'Bash(gh pr view:*)' 'Read' 'mcp__github__*'`. Use `--disallowedTools` to subtract a dangerous one (block `WebFetch` on a loop that shouldn't read the web; block a Slack `post_message` on a read-only triage loop). - **Read-scoped connectors at L1.** An L1 report loop gets read-only MCP scopes (list/get/search), never write (post/create/delete/merge). Scope the connector *itself* least-privilege — don't hand a triage loop a write-capable GitHub token "just in case". - **The auto-merge guard.** Never give a loop a path to merge `main`: keep `gh pr merge` out of the allowlist, set `land_via: fleet-ops` (test-gated, human-or-queue), and list main-push in `escalation`. A green PR on a feature branch is the *most* a loop auto-produces. - **MCP tool descriptions are instructions.** A poisoned connector description is prompt-injection straight into the loop's context — vet a connector (and prefer read scopes) before a loop uses it. See [`prompt-injection-defense`](../../prompt-injection-defense/SKILL.md). ## Why Claude Code-specific (not a multi-tool matrix) `loop-ops` is deliberately scoped to **Claude Code**, not a cross-tool primitives matrix (Grok / Codex / …). The whole edge is grounding in Claude Code's *actual* gate model — the permission modes, the auto-mode classifier, `claude -p`, the hook events. A generic multi-agent matrix would dilute exactly that. The *doctrine* ports to any agent (the tier ladder, the gate, the kill switch, the STATE spine, the escalation classes); the permission-mode **mapping** is Claude Code's, and that specificity is the point. ## Tier checklist (what `loop-check` enforces) - **L1:** bounded `scope` (never `*`), a `kill_switch`, `permission_mode` ∈ {plan, dontAsk}, **no** `verify` that writes. Report-only. - **L2:** all of L1, plus a `verify` gate **and** a `guard` (must-always-pass), `worktree: true`, a concrete `escalation:` rule, and a `land_via` (e.g. `fleet-ops`). - **L3:** all of L2, plus `permission_mode: bypassPermissions` **with** an isolation note in `escalation`/scope, a denylist of never-auto-land classes, and a `budget_tokens` cap. The audit warns hard if L3 is declared without an isolation boundary. ## See also - [../../../docs/AUTO-MODE-CLASSIFIER.md](../../../docs/AUTO-MODE-CLASSIFIER.md) — the full two-gate model, classifier categories, legitimate-authorization decision tree. - [claude-code-loops.md](claude-code-loops.md) — the scheduler/`claude -p` mechanics this tier model runs on. - [pattern-catalog.md](pattern-catalog.md) — each pattern's recommended starting tier. -
state-spine.md 6.9 KB
# The State Spine — memory outside the conversation A loop's durability comes from state that lives **outside** the conversation window. The conversation is ephemeral and degrades as it fills (the Ralph insight: quality drops past ~100–150k tokens). The spine is three files the loop reads at the start of every run and writes at the end. This is the loop's working memory, audit trail, and definition. ``` .loops/<name>/ ├── loop.config.yaml # the definition (immutable-ish; edited by a human) ├── STATE.md # the triage snapshot (rewritten every run) └── run-log.md # append-only audit trail (one line per run) ``` `loop-scaffold` scaffolds all three. The config is human-owned; `STATE.md` and `run-log.md` are loop-owned. --- ## `loop.config.yaml` — the definition Flat YAML so it's trivially parseable (no `yq` dependency). Full annotated template: [../assets/loop.config.template.yaml](../assets/loop.config.template.yaml). Fields: | Field | Required | Meaning | |---|---|---| | `name` | yes | the loop's identifier; matches the directory | | `pattern` | yes | a catalog key (`pr-watch`, …) or `custom` | | `tier` | yes | `L1` / `L2` / `L3` — the autonomy rung | | `cadence` | yes | `10m` / `1h` / `6h` / `1d`, or a cron string | | `host` | rec | where ticks execute: `local` (default) / `session-cron` / `desktop-task` / `cloud-routine` / `external`. Selects which hard limits `loop-doctor` enforces — see [native-scheduling.md](native-scheduling.md) | | `goal` | yes | one sentence: what this loop does and what it must NOT do | | `scope` | yes | bounded globs the loop may touch — **never `*`** | | `verify` | L2+ | the gate command (the metric/check); a loop with no gate is invalid | | `guard` | L2+ | a must-always-pass command (full suite / typecheck) | | `permission_mode` | yes | `plan` / `dontAsk` / `auto` / `acceptEdits` / `bypassPermissions`. **Ignored when `host: cloud-routine`** — routines have no permission picker; the boundary is repos + environment + connectors | | `worktree` | L2+ | `true` to isolate code changes in a git worktree | | `escalation` | yes | what the loop escalates instead of doing (the gate rule) | | `budget_tokens` | rec | per-run output-token ceiling | | `kill_switch` | yes | the stop signal every run checks first | | `land_via` | L2+ | who gates + lands winning branches (e.g. `fleet-ops`) | `loop-check` reads this file and scores it against the tier's requirements. --- ## `STATE.md` — the triage snapshot Rewritten at the end of every run; read at the top of the next. It is **not** a database — it's a lightweight snapshot of what the loop needs, what it's watching, and what it ignored. Template: [../assets/STATE.template.md](../assets/STATE.template.md). Shape: ```markdown # <loop-name> — STATE _Updated: 2026-06-22T14:05:00Z · run #142 · readiness 100/100_ ## Priority (act on these next) - [P1] PR #412 failing CI 3h — owner pinged - [P2] dep `axios` patch 1.14.0→1.14.1 available, cooldown clears 2026-06-25 ## Watch (not yet actionable) - PR #408 awaiting review 1h - flag `new-checkout` at 100% rollout 6d — cleanup candidate ## Noise (seen + dismissed this run) - PR #410 draft — skip until ready - dep `left-pad` major bump — escalates, not auto --- _Source: .github/workflows/<loop>.yml · config: loop.config.yaml_ ``` **The read/write contract:** 1. **Read** `STATE.md` first thing — it's the loop's memory of the last run. 2. **Check the kill switch** (`kill_switch:` from config) — exit immediately if set. 3. Do the run's work, drawing the next unit from the Priority list. 4. **Rewrite** `STATE.md` — promote/demote items across Priority/Watch/Noise, bump the `_Updated_` line + run number + readiness. `readiness` is the loop's self-assessment (0–100): is its config still coherent, its gate still passing, its scope still valid? A dropping readiness is an early signal to re-audit. --- ## `run-log.md` — the append-only audit trail One line per run, appended, never rewritten. Answers "what has this loop been doing, and what did it cost?" ``` 2026-06-22T14:05:00Z run#142 action=reported pr=412 outcome=escalated tokens=18420 2026-06-22T13:55:00Z run#141 action=none - outcome=quiet tokens=2110 2026-06-22T13:45:00Z run#140 action=proposed pr=409 outcome=pr-opened tokens=44380 ``` The `tokens` column feeds back into the budget. Tail it to see drift: a loop that used to cost 2k/run quietly now costing 40k/run is doing more than it was scoped to. --- ## Budget control A loop's cost is `runs/day × tokens/run × price`, and sub-agents multiply tokens/run. Two controls: - **`budget_tokens`** in the config — a per-run output ceiling. The loop stops the run when it's reached (the same discipline as a dynamic `/loop` watching `budget.remaining()`). - **The run-log** — the actual spend, line by line. Reconcile estimate (`loop-estimate`) against actual periodically; if they diverge, the loop's scope crept. Estimate before you schedule: [../scripts/loop-estimate.py](../scripts/loop-estimate.py). The cheapest lever is **cadence** — halving the frequency halves the cost. The next is **model** — a Haiku triage loop costs a fifth of an Opus one; put the cheap model on the maker and reserve the expensive one for the gate decision. --- ## Multi-loop coordination Running several loops against one repo, two rules prevent them tripping over each other: ### Priority order (collision avoidance) ``` CI Watch ► PR Watch ► Dependency Bump ► Post-Merge/Changelog ► Daily Scan (highest) (off-peak) ``` A red build blocks everyone, so the CI watch wins any worktree contention; daily scan yields to all. When two loops want the same worktree/branch, the higher-priority one proceeds and the lower defers to its next cadence tick. Loops announce what they're touching via [`pigeon`](../../pigeon/SKILL.md) so a peer can see "ci-watch holds a worktree on PR #412" and stand off. ### The kill switch (every loop honors it) One stop signal, checked at the top of **every** run, that halts **every** loop: - a **sentinel file** — `.loops/PAUSED` (global) or `.loops/<name>/PAUSED` (one loop), or - a **label** — `loop-pause` on the repo/issue, checked via `gh`. No loop ships without one. It's the difference between "the loops are misbehaving, give me a minute" and "the loops are misbehaving, where's the breaker?". Put the exact mechanism in `kill_switch:` and make checking it the first action of every run, before the work. ## See also - [risk-tiers.md](risk-tiers.md) — the autonomy ladder the config's `tier` selects. - [pattern-catalog.md](pattern-catalog.md) — each pattern's place in the priority order. - [claude-code-loops.md](claude-code-loops.md) — how the cadence actually fires. - [native-scheduling.md](native-scheduling.md) — the `host:` values and the limits each one imposes.
-
-
scripts
-
check-native-facts.py 9.7 KB
#!/usr/bin/env python3 """Staleness verifier for loop-ops' native-scheduling facts. references/native-scheduling.md encodes a fast-moving external surface: the native scheduling primitives (CronCreate / the scheduled-tasks MCP / cloud routines) and their hard limits. Those limits are load-bearing - loop-doctor refuses a config on them - and they are exactly the kind of fact that rots invisibly (SKILL-RESOURCE-PROTOCOL.md §7). Two modes guard it: --offline (default, safe for PR CI): internal consistency, no network. * the host vocabulary is ONE set across all four places it appears: assets/loop.config.template.yaml, scripts/loop-scaffold.sh (--host), scripts/loop-doctor.sh (its case arms), references/native-scheduling.md * native-scheduling.md carries its "Verified <date>" stamp * every load-bearing limit loop-doctor enforces is still stated in the prose (the 1-hour cloud floor, the 7-day session-cron expiry) --live (scheduled freshness.yml, never a PR gate): fetch the three upstream doc pages and check the numbers we encode still appear in them. A changed number upstream is real drift; an unreachable docs host is advisory, not a failure. Usage: check-native-facts.py [--offline | --live] [--skill DIR] [--json] [--timeout S] Input: argv flags only (no stdin). Output: stdout = findings (plain rows, or a --json envelope). Data only. Stderr: the verdict line, notices, errors. Exit: 0 in sync, 2 usage, 3 a required file is missing, 4 unparseable, 7 docs unreachable (live, advisory - never a real failure), 10 drift found Examples: check-native-facts.py --offline # PR CI: host vocabulary + limits are one set check-native-facts.py --live # weekly: our numbers vs the published docs check-native-facts.py --offline --json | jq '.data[]' """ from __future__ import annotations import argparse import json import re import sys import urllib.error import urllib.request from pathlib import Path EX_OK = 0 EX_USAGE = 2 EX_NOTFOUND = 3 EX_UNPARSEABLE = 4 EX_UNREACHABLE = 7 EX_DRIFT = 10 SCHEMA = "claude-mods.loop-ops.native-facts/v1" # The canonical host vocabulary. Every file below must agree with exactly this set - # a host added in one place and forgotten in another is the drift this catches. HOSTS = {"local", "session-cron", "desktop-task", "cloud-routine", "external"} # Limits loop-doctor actually enforces, so the prose that justifies them must state # them. (needle, where it must appear, why it matters) LIMITS = [ ("1 hour", "cloud-routine minimum interval"), ("7 days", "session-cron recurring-task expiry"), ] # Live checks: (url, [(needle, label)]). Needles are the published numbers we encode. LIVE_PAGES = [ ("https://code.claude.com/docs/en/routines", [ ("minimum interval is one hour", "cloud routine >=1h floor"), ]), ("https://code.claude.com/docs/en/scheduled-tasks", [ ("expire 7 days", "session-cron 7-day expiry"), ("50 scheduled tasks", "50-task-per-session cap"), ]), ("https://code.claude.com/docs/en/desktop-scheduled-tasks", [ ("scheduled-tasks", "desktop task on-disk location"), ]), ] class Term: """Minimal stderr styling; honours TERM_ASCII=1 and a non-tty stderr.""" def __init__(self) -> None: import os self.plain = os.environ.get("TERM_ASCII") == "1" or not sys.stderr.isatty() def say(self, msg: str) -> None: print(msg, file=sys.stderr) class Finding: def __init__(self, state: str, check: str, detail: str) -> None: self.state, self.check, self.detail = state, check, detail def as_dict(self) -> dict: return {"state": self.state, "check": self.check, "detail": self.detail} def read(path: Path) -> str: try: return path.read_text(encoding="utf-8", errors="replace") except OSError as exc: raise FileNotFoundError(str(exc)) from exc def hosts_in_template(text: str) -> set: """Hosts named in the `host:` block's inline comments.""" block = re.search(r"^host:.*?(?=\n[a-z_]+:)", text, re.M | re.S) if not block: return set() return {h for h in HOSTS if re.search(r"\b" + re.escape(h) + r"\b", block.group(0))} def hosts_in_scaffold(text: str) -> set: """Hosts accepted by loop-scaffold's --host validation case arm.""" arm = re.search(r"case \"\$HOST\" in\n\s*([a-z|\-]+)\)", text) if not arm: return set() return set(arm.group(1).split("|")) def hosts_in_doctor(text: str) -> set: """Hosts loop-doctor recognises in its host-coherence case arm.""" arm = re.search(r"case \"\$HOST\" in\n\s*([a-z|\-]+)\) row ok \"host\"", text) if not arm: return set() return set(arm.group(1).split("|")) def hosts_in_reference(text: str) -> set: return {h for h in HOSTS if "`" + h + "`" in text} def check_offline(skill: Path) -> list: findings = [] tpl = skill / "assets" / "loop.config.template.yaml" scaffold = skill / "scripts" / "loop-scaffold.sh" doctor = skill / "scripts" / "loop-doctor.sh" ref = skill / "references" / "native-scheduling.md" for p in (tpl, scaffold, doctor, ref): if not p.is_file(): raise FileNotFoundError(str(p)) sources = { "loop.config.template.yaml": hosts_in_template(read(tpl)), "loop-scaffold.sh --host": hosts_in_scaffold(read(scaffold)), "loop-doctor.sh host arm": hosts_in_doctor(read(doctor)), "native-scheduling.md": hosts_in_reference(read(ref)), } for where, found in sources.items(): if not found: findings.append(Finding("bad", "hosts", f"{where}: no host vocabulary found (parser or format changed)")) elif found != HOSTS: missing = sorted(HOSTS - found) extra = sorted(found - HOSTS) detail = f"{where}: " + ", ".join( filter(None, [f"missing {missing}" if missing else "", f"unknown {extra}" if extra else ""]) ) findings.append(Finding("bad", "hosts", detail)) else: findings.append(Finding("ok", "hosts", f"{where}: all {len(HOSTS)} hosts")) ref_text = read(ref) stamp = re.search(r"\*\*Verified (\d{4}-\d{2}-\d{2})\*\*", ref_text) if stamp: findings.append(Finding("ok", "date-stamp", f"native-scheduling.md verified {stamp.group(1)}")) else: findings.append(Finding("bad", "date-stamp", "native-scheduling.md has no '**Verified YYYY-MM-DD**' stamp")) for needle, why in LIMITS: if needle in ref_text: findings.append(Finding("ok", "limit", f"{why}: '{needle}' documented")) else: findings.append(Finding("bad", "limit", f"{why}: '{needle}' not stated - loop-doctor enforces it unexplained")) return findings def check_live(timeout: float) -> list: findings = [] unreachable = 0 for url, needles in LIVE_PAGES: try: req = urllib.request.Request(url, headers={"User-Agent": "claude-mods-loop-ops-verifier"}) with urllib.request.urlopen(req, timeout=timeout) as resp: body = resp.read().decode("utf-8", errors="replace") except (urllib.error.URLError, OSError, ValueError) as exc: unreachable += 1 findings.append(Finding("skip", "fetch", f"{url}: unreachable ({exc.__class__.__name__})")) continue for needle, label in needles: if needle.lower() in body.lower(): findings.append(Finding("ok", "live", f"{label}: still published")) else: findings.append(Finding("bad", "live", f"{label}: '{needle}' no longer in {url} - re-verify native-scheduling.md")) if unreachable == len(LIVE_PAGES): findings.append(Finding("skip", "live", "all docs pages unreachable - live check advisory only")) return findings def main() -> int: ap = argparse.ArgumentParser(add_help=False) ap.add_argument("--offline", action="store_true") ap.add_argument("--live", action="store_true") ap.add_argument("--skill", default=str(Path(__file__).resolve().parent.parent)) ap.add_argument("--json", action="store_true") ap.add_argument("--timeout", type=float, default=15.0) ap.add_argument("-h", "--help", action="store_true") try: args = ap.parse_args() except SystemExit: return EX_USAGE if args.help: print(__doc__) return EX_OK if args.offline and args.live: print("error: --offline and --live are mutually exclusive", file=sys.stderr) return EX_USAGE term = Term() skill = Path(args.skill) mode = "live" if args.live else "offline" try: findings = check_offline(skill) except FileNotFoundError as exc: print(f"error: required file missing: {exc}", file=sys.stderr) return EX_NOTFOUND except re.error as exc: print(f"error: could not parse a source file: {exc}", file=sys.stderr) return EX_UNPARSEABLE if args.live: findings += check_live(args.timeout) bad = [f for f in findings if f.state == "bad"] skipped = [f for f in findings if f.state == "skip"] if args.json: print(json.dumps({ "schema": SCHEMA, "mode": mode, "in_sync": not bad, "data": [f.as_dict() for f in findings], }, indent=2)) else: for f in findings: print(f"{f.state:<5} {f.check:<12} {f.detail}") if bad: term.say(f"native-facts: {len(bad)} drift finding(s) - re-verify references/native-scheduling.md") return EX_DRIFT if args.live and len(skipped) >= len(LIVE_PAGES): term.say("native-facts: docs unreachable - live check skipped (advisory)") return EX_UNREACHABLE term.say(f"native-facts: in sync ({mode})") return EX_OK if __name__ == "__main__": sys.exit(main()) -
check-pricing-sync.py 7.1 KB
#!/usr/bin/env python3 """Offline verifier: loop-ops pricing must match claude-api-ops's model table. loop-estimate.py reads assets/model-pricing.json. That table is a *copy* of the authoritative "Current Models" table in skills/claude-api-ops/SKILL.md — and a copy drifts silently (the exact §7 failure mode). This asserts every model in loop-ops' pricing exists in the claude-api-ops table with matching input/output prices. Both files are in-repo, so this is a pure OFFLINE consistency check and safe to gate PR CI (no network). Live model-id drift is owned by claude-api-ops/scripts/check-model-table.py. Usage: check-pricing-sync.py [--offline] [--pricing FILE] [--table FILE] [--json] Input: argv flags only (no stdin). Output: stdout = drift findings (plain rows, or --json envelope). Data only. Stderr: the verdict panel, notices, errors. Exit: 0 in sync, 2 usage, 3 a file missing, 4 unparseable, 10 drift found --offline is the default and only mode (accepted for parity with the other §7 verifiers invoked by tests/check-resources.sh). Examples: check-pricing-sync.py --offline check-pricing-sync.py --json | jq '.data[]' """ from __future__ import annotations import argparse import json import os import re import sys from pathlib import Path EX_OK = 0 EX_USAGE = 2 EX_NOTFOUND = 3 EX_UNPARSEABLE = 4 EX_DRIFT = 10 HERE = Path(__file__).resolve().parent DEFAULT_PRICING = HERE.parent / "assets" / "model-pricing.json" DEFAULT_TABLE = HERE.parent.parent / "claude-api-ops" / "SKILL.md" PRICE_RE = re.compile(r"\$?\s*([0-9]+(?:\.[0-9]+)?)") class Term: """Minimal ANSI helper (term.sh is bash-only; per TERMINAL-DESIGN.md §9 the Python port is inline). Honors FORCE_COLOR / NO_COLOR / TERM_ASCII and the bound stream's TTY + encoding so piped data stays plain ASCII.""" _C = {"green": "\033[32m", "red": "\033[31m", "cyan": "\033[36m", "dim": "\033[2m", "off": "\033[0m"} def __init__(self, stream=sys.stderr): enc = (getattr(stream, "encoding", "") or "").lower() self.ascii = os.environ.get("TERM_ASCII") == "1" or "utf" not in enc if os.environ.get("FORCE_COLOR"): self.color = True elif (os.environ.get("NO_COLOR") is not None or os.environ.get("TERM") == "dumb" or not getattr(stream, "isatty", lambda: False)()): self.color = False else: self.color = True def c(self, name, text): return f"{self._C.get(name,'')}{text}{self._C['off']}" if self.color else text def mark(self, ok): g = ("+" if self.ascii else "✓") if ok else ("x" if self.ascii else "✗") return self.c("green" if ok else "red", g) def parse_price(cell: str) -> float | None: m = PRICE_RE.search(cell) return float(m.group(1)) if m else None def load_pricing(path: Path) -> dict: """{model_id: (input_per_mtok, output_per_mtok)} from loop-ops' JSON.""" if not path.is_file(): print(f"error: pricing file not found: {path}", file=sys.stderr) raise SystemExit(EX_NOTFOUND) try: data = json.loads(path.read_text(encoding="utf-8")) out = {} for mid, pr in data.get("models", {}).items(): out[mid] = (float(pr["input_per_mtok"]), float(pr["output_per_mtok"])) if not out: print(f"error: no models in {path}", file=sys.stderr) raise SystemExit(EX_UNPARSEABLE) return out except (json.JSONDecodeError, KeyError, TypeError, ValueError) as exc: print(f"error: could not parse pricing file: {exc}", file=sys.stderr) raise SystemExit(EX_UNPARSEABLE) def load_table(path: Path) -> dict: """{model_id: (input_price, output_price)} from the claude-api-ops markdown 'Current Models' table. Columns: Model | ID | Context | Max Output | Input | Output.""" if not path.is_file(): print(f"error: claude-api-ops table not found: {path}", file=sys.stderr) raise SystemExit(EX_NOTFOUND) table: dict = {} in_table = False for line in path.read_text(encoding="utf-8").splitlines(): s = line.strip() low = s.lower() if s.startswith("|") and "id" in low and "context" in low and "output" in low: in_table = True continue if in_table: if not s.startswith("|"): if table: # table ended break continue if set(s) <= set("|-: "): # separator row continue cells = [c.strip() for c in s.strip("|").split("|")] if len(cells) < 6: continue mid = cells[1].strip("`").strip() if not mid.startswith("claude-"): continue ip, op = parse_price(cells[4]), parse_price(cells[5]) if ip is not None and op is not None: table[mid] = (ip, op) if not table: print(f"error: no model rows parsed from {path}", file=sys.stderr) raise SystemExit(EX_UNPARSEABLE) return table def main(argv: list[str]) -> int: p = argparse.ArgumentParser( prog="check-pricing-sync.py", description="Verify loop-ops pricing matches claude-api-ops's model table (offline).", ) p.add_argument("--offline", action="store_true", help="offline consistency check (default/only mode)") p.add_argument("--pricing", default=str(DEFAULT_PRICING), help="loop-ops model-pricing.json") p.add_argument("--table", default=str(DEFAULT_TABLE), help="claude-api-ops SKILL.md with the model table") p.add_argument("--json", action="store_true", help="emit a JSON envelope") try: args = p.parse_args(argv) except SystemExit as exc: return EX_USAGE if exc.code not in (0, None) else (exc.code or EX_OK) pricing = load_pricing(Path(args.pricing)) table = load_table(Path(args.table)) findings = [] for mid, (ip, op) in sorted(pricing.items()): if mid not in table: findings.append({"model": mid, "issue": "absent from claude-api-ops table", "loop_ops": [ip, op], "authoritative": None}) continue tip, top = table[mid] if abs(ip - tip) > 1e-9 or abs(op - top) > 1e-9: findings.append({"model": mid, "issue": "price mismatch", "loop_ops": [ip, op], "authoritative": [tip, top]}) if args.json: print(json.dumps({ "data": findings, "meta": {"count": len(findings), "models_checked": len(pricing), "in_sync": not findings, "schema": "claude-mods.loop-ops.pricing-sync/v1"}, }, indent=2)) else: for f in findings: auth = f"authoritative {f['authoritative']}" if f["authoritative"] else "not in table" print(f"DRIFT {f['model']}: {f['issue']} (loop-ops {f['loop_ops']} vs {auth})") t = Term(sys.stderr) ok = not findings print(f"{t.mark(ok)} pricing-sync: {len(pricing)} model(s) checked, " f"{len(findings)} drift " f"{t.c('dim', '(authoritative: claude-api-ops/SKILL.md)')}", file=sys.stderr) return EX_DRIFT if findings else EX_OK if __name__ == "__main__": sys.exit(main(sys.argv[1:])) -
loop-check.sh 12 KB
#!/usr/bin/env bash # Score an outer-loop config for readiness before it is scheduled. # # Usage: loop-check.sh [OPTIONS] <loop.config.yaml> # Input: argv flags + a config path (no stdin). # Output: stdout = findings (plain `SEVERITY message` rows, or --json envelope). # Data only. # Stderr: the readiness panel (score + verdict), notices, errors. # Exit: 0 ready (no errors, score >= --min), 2 usage, 3 config not found, # 4 config unparseable, 10 NOT ready (findings present) # # Scores a flat loop.config.yaml against the tier's requirements: a bounded scope, # a defined escalation rule + kill switch, and — at L2+ — a verify gate, a guard, a # worktree, and a landing path. The config is parsed without a yq dependency. # Pair with loop-scaffold.sh (scaffold) and references/risk-tiers.md (the rubric). # # Examples: # loop-check.sh .loops/pr-watch/loop.config.yaml # loop-check.sh --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.severity=="error")' # loop-check.sh --min 80 --strict .loops/ci-watch/loop.config.yaml set -uo pipefail readonly EX_OK=0 EX_USAGE=2 EX_NOTFOUND=3 EX_UNPARSEABLE=4 EX_FINDINGS=10 # Terminal design system. stdout = findings (data); the score panel frames on stderr. __lib="$(cd "$(dirname "${BASH_SOURCE[0]}")/../../_lib" 2>/dev/null && pwd || true)" if [ -n "${__lib:-}" ] && [ -f "$__lib/term.sh" ]; then . "$__lib/term.sh"; term_init 2 else term_panel_open() { :; }; term_panel_close() { :; }; term_panel_vert() { :; } term_status_row() { shift; printf ' - %s %s\n' "$1" "${2:-}"; } term_pip_bar() { printf '%s/%s' "$2" "$3"; } term_color() { shift; printf '%s' "$*"; }; TERM_DOT="|" fi CFG="" MIN=70 STRICT=0 JSON=0 usage() { cat <<'EOF' loop-check.sh — score an outer-loop config for readiness. Usage: loop-check.sh [OPTIONS] <loop.config.yaml> Options: --min N readiness score (0-100) required for a "ready" verdict (default: 70). --strict count warnings toward the NOT-ready signal (exit 10). --json emit a JSON envelope instead of plain rows. -h, --help show this help and exit 0. Exit codes: 0 ready 2 usage 3 config not found 4 unparseable 10 NOT ready (findings) Examples: loop-check.sh .loops/pr-watch/loop.config.yaml loop-check.sh --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.severity=="error")' loop-check.sh --min 80 --strict .loops/ci-watch/loop.config.yaml EOF } die_usage() { printf 'error: %s\n' "$1" >&2; echo >&2; usage >&2; exit "$EX_USAGE"; } # ── parse args ────────────────────────────────────────────────────────────── while [[ $# -gt 0 ]]; do case "$1" in --min) [[ $# -ge 2 ]] || die_usage "--min needs a value"; MIN="$2"; shift 2 ;; --strict) STRICT=1; shift ;; --json) JSON=1; shift ;; -h|--help) usage; exit "$EX_OK" ;; -*) die_usage "unknown flag: $1" ;; *) [[ -z "$CFG" ]] || die_usage "unexpected extra argument: $1"; CFG="$1"; shift ;; esac done [[ -n "$CFG" ]] || die_usage "a loop.config.yaml path is required" [[ "$MIN" =~ ^[0-9]+$ ]] || die_usage "--min must be an integer (got '$MIN')" [[ -f "$CFG" ]] || { printf 'error: config not found: %s\n' "$CFG" >&2; exit "$EX_NOTFOUND"; } # Normalize Windows-authored configs: strip a leading UTF-8 BOM (line 1) and CR # line-endings so a CRLF/BOM file parses identically to a clean LF one (octal BOM + # gsub \r are portable across gawk/mawk/BSD awk). Falls back to the original on failure. __NORM="$(mktemp 2>/dev/null)" && awk 'NR==1{sub(/^\357\273\277/,"")} {gsub(/\r/,""); print}' "$CFG" > "$__NORM" 2>/dev/null && CFG="$__NORM" && trap 'rm -f "$__NORM"' EXIT # Unparseable: no top-level `key:` lines at all. if ! grep -Eq '^[a-z_]+:' "$CFG"; then printf 'error: no parseable top-level keys in %s\n' "$CFG" >&2 exit "$EX_UNPARSEABLE" fi # ── flat-YAML readers (no yq) ─────────────────────────────────────────────── cfg_scalar() { # inline scalar value for `^KEY:`; empty if absent or block-list awk -v k="$1" -v q="'" ' $0 ~ "^"k":" { sub("^"k":[ \t]*","") sub(/[ \t]*#.*$/,"") gsub(/^[ \t]+|[ \t]+$/,"") gsub(/^"|"$/,""); gsub("^"q"|"q"$","") print; exit }' "$CFG" } cfg_has_key() { grep -Eq "^$1:" "$CFG"; } cfg_list_items() { # ` - item` lines under `^KEY:`, until the next top-level key awk -v k="$1" -v q="'" ' $0 ~ "^"k":" { inlist=1; next } inlist==1 { if ($0 ~ /^[ \t]*-[ \t]+/) { line=$0 sub(/^[ \t]*-[ \t]+/,"",line); sub(/[ \t]*#.*$/,"",line) gsub(/^[ \t]+|[ \t]+$/,"",line); gsub(/^"|"$/,"",line); gsub("^"q"|"q"$","",line) if (line != "") print line } else if ($0 ~ /^[^ \t#]/) { inlist=0 } }' "$CFG" } is_placeholder() { [[ "$1" == *"<"*">"* ]]; } # an unfilled <PLACEHOLDER> # ── findings + scoring ────────────────────────────────────────────────────── FIND_SEV=(); FIND_MSG=() CHECKS_TOTAL=0; CHECKS_PASS=0 add() { FIND_SEV+=("$1"); FIND_MSG+=("$2"); } pass() { CHECKS_TOTAL=$((CHECKS_TOTAL+1)); CHECKS_PASS=$((CHECKS_PASS+1)); } fail() { CHECKS_TOTAL=$((CHECKS_TOTAL+1)); add "$1" "$2"; } # $1=severity $2=message # require <severity> <ok?> <message-on-fail> — a present+valid scalar check. require() { if [[ "$2" -eq 1 ]]; then pass; else fail "$1" "$3"; fi; } TIER="$(cfg_scalar tier)" PMODE="$(cfg_scalar permission_mode)" NAME="$(cfg_scalar name)" GOAL="$(cfg_scalar goal)" ESCAL="$(cfg_scalar escalation)" KILL="$(cfg_scalar kill_switch)" BUDGET="$(cfg_scalar budget_tokens)" VERIFY="$(cfg_scalar verify)" GUARD="$(cfg_scalar guard)" WORKTREE="$(cfg_scalar worktree)" LANDVIA="$(cfg_scalar land_via)" CADENCE="$(cfg_scalar cadence)" PATTERN="$(cfg_scalar pattern)" is_l2plus=0; [[ "$TIER" == "L2" || "$TIER" == "L3" ]] && is_l2plus=1 # present-and-not-placeholder predicate filled() { [[ -n "$1" ]] && ! is_placeholder "$1"; } # ── always-applicable checks ──────────────────────────────────────────────── require error "$(filled "$NAME" && echo 1 || echo 0)" "name: missing or placeholder" require warning "$(filled "$PATTERN" && echo 1 || echo 0)" "pattern: missing" case "$TIER" in L1|L2|L3) pass ;; *) fail error "tier: must be L1|L2|L3 (got '${TIER:-empty}')" ;; esac require warning "$(filled "$CADENCE" && echo 1 || echo 0)" "cadence: missing" require error "$(filled "$GOAL" && echo 1 || echo 0)" "goal: missing or placeholder" require error "$(filled "$ESCAL" && echo 1 || echo 0)" "escalation: undefined — every loop must declare what it escalates" require error "$(filled "$KILL" && echo 1 || echo 0)" "kill_switch: undefined — no loop ships without a stop signal" # budget present + numeric if [[ -n "$BUDGET" && "$BUDGET" =~ ^[0-9]+$ ]]; then pass; else fail warning "budget_tokens: missing or non-numeric — bound the per-run spend"; fi # scope present + bounded + not placeholder mapfile -t SCOPE_ITEMS < <(cfg_list_items scope) SCOPE_INLINE="$(cfg_scalar scope)" [[ -n "$SCOPE_INLINE" ]] && SCOPE_ITEMS+=("$SCOPE_INLINE") if ! cfg_has_key scope || [[ ${#SCOPE_ITEMS[@]} -eq 0 ]]; then fail error "scope: missing — bound what the loop may touch" else scope_bad=0 for it in "${SCOPE_ITEMS[@]}"; do if is_placeholder "$it"; then fail error "scope: unfilled placeholder ('$it')"; scope_bad=1; break; fi case "$it" in '*'|'**'|'.'|'./'|'/'|'') fail error "scope: unbounded ('$it') — a loop that may touch anything is not bounded"; scope_bad=1; break ;; esac done [[ "$scope_bad" -eq 0 ]] && pass fi # permission_mode present + valid case "$PMODE" in plan|dontAsk|auto|acceptEdits|bypassPermissions) pass ;; "") fail error "permission_mode: missing" ;; *) fail error "permission_mode: invalid ('$PMODE')" ;; esac # permission_mode consistent with tier (warning) case "$TIER" in L1) case "$PMODE" in plan|dontAsk) pass ;; *) fail warning "permission_mode '$PMODE' is broad for L1 (report-only) — prefer plan or dontAsk" ;; esac ;; L2) case "$PMODE" in dontAsk|auto|acceptEdits) pass ;; *) fail warning "permission_mode '$PMODE' fits L2 poorly — prefer dontAsk/auto/acceptEdits" ;; esac ;; L3) case "$PMODE" in bypassPermissions) pass ;; *) fail warning "L3 unattended usually needs bypassPermissions in a container (got '$PMODE')" ;; esac ;; *) : ;; esac # ── L2+ checks (code-changing tiers) ──────────────────────────────────────── if [[ "$is_l2plus" -eq 1 ]]; then require error "$(filled "$VERIFY" && echo 1 || echo 0)" "verify: no gate command — a code-changing loop with no gate is invalid" require error "$(filled "$GUARD" && echo 1 || echo 0)" "guard: no must-always-pass command at $TIER" if [[ "$WORKTREE" == "true" ]]; then pass; else fail error "worktree: must be true at $TIER — isolate code changes"; fi require warning "$(filled "$LANDVIA" && echo 1 || echo 0)" "land_via: undefined — name who gates+lands (e.g. fleet-ops)" fi # ── L3-specific isolation check ───────────────────────────────────────────── if [[ "$TIER" == "L3" ]]; then if printf '%s %s' "$ESCAL" "${SCOPE_ITEMS[*]:-}" | grep -Eqi 'container|isolat|sandbox|devcontainer'; then pass else fail warning "L3 declares no isolation boundary — bypassPermissions is only safe in a container/VM; note it in escalation" fi fi # ── verdict ───────────────────────────────────────────────────────────────── ERRORS=0; WARNINGS=0 for s in "${FIND_SEV[@]:-}"; do [[ "$s" == "error" ]] && ERRORS=$((ERRORS+1)) [[ "$s" == "warning" ]] && WARNINGS=$((WARNINGS+1)) done SCORE=0 [[ "$CHECKS_TOTAL" -gt 0 ]] && SCORE=$(( CHECKS_PASS * 100 / CHECKS_TOTAL )) READY=1 [[ "$ERRORS" -gt 0 ]] && READY=0 [[ "$SCORE" -lt "$MIN" ]] && READY=0 [[ "$STRICT" -eq 1 && "$WARNINGS" -gt 0 ]] && READY=0 # ── output ────────────────────────────────────────────────────────────────── if [[ "$JSON" -eq 1 ]]; then printf '{\n "data": [\n' for i in "${!FIND_SEV[@]}"; do msg="${FIND_MSG[$i]//\\/\\\\}"; msg="${msg//\"/\\\"}" sep=","; [[ "$i" -eq $(( ${#FIND_SEV[@]} - 1 )) ]] && sep="" printf ' {"severity": "%s", "message": "%s"}%s\n' "${FIND_SEV[$i]}" "$msg" "$sep" done printf ' ],\n "meta": {"count": %d, "errors": %d, "warnings": %d, "score": %d, "min": %d, "ready": %s, "tier": "%s", "schema": "claude-mods.loop-ops.check/v1"}\n}\n' \ "${#FIND_SEV[@]}" "$ERRORS" "$WARNINGS" "$SCORE" "$MIN" "$([[ "$READY" -eq 1 ]] && echo true || echo false)" "${TIER:-unknown}" else if [[ ${#FIND_SEV[@]} -gt 0 ]]; then for i in "${!FIND_SEV[@]}"; do printf '%-7s %s\n' "$(printf '%s' "${FIND_SEV[$i]}" | tr '[:lower:]' '[:upper:]')" "${FIND_MSG[$i]}" done fi verdict="$([[ "$READY" -eq 1 ]] && echo READY || echo "NOT READY")" vstate="$([[ "$READY" -eq 1 ]] && echo ok || echo bad)" { term_panel_open loop "loop ${TERM_DOT} audit" "${NAME:-$(basename "$(dirname "$CFG")")}" term_panel_vert term_status_row "$vstate" "$verdict $(term_pip_bar score "$SCORE" 100)" "score $SCORE/100 ${TERM_DOT} tier ${TIER:-?}" term_status_row "$([[ "$ERRORS" -eq 0 ]] && echo ok || echo bad)" "$ERRORS error(s)" "must be 0 to be ready" term_status_row "$([[ "$WARNINGS" -eq 0 ]] && echo ok || echo warn)" "$WARNINGS warning(s)" "$([[ "$STRICT" -eq 1 ]] && echo 'block under --strict' || echo advisory)" term_panel_vert term_panel_close "min $MIN ${TERM_DOT} fix errors before scheduling" "" } >&2 fi [[ "$READY" -eq 1 ]] && exit "$EX_OK" || exit "$EX_FINDINGS" -
loop-doctor.sh 17 KB
#!/usr/bin/env bash # Preflight a loop config - will this loop actually RUN, or die at 3am? # # loop-check checks the config is well-formed; loop-doctor checks the loop will # execute: the gate command's binary resolves, claude/git are on PATH, the budget # can fit a tick, and the permission mode is achievable from where it launches. # # HOST-AWARE. Since native scheduling landed, "where it launches" is a real # variable, so the config's optional `host:` selects which constraints apply - a # cloud routine has NO permission mode and a >=1h floor, and this machine's PATH # says nothing about it; a session-cron host cannot run unattended at all. # Verified surface + limits: references/native-scheduling.md (2026-08-30). # Modeled on fleet-worker/scripts/fleet-doctor.sh. # # Usage: loop-doctor.sh [--offline|--live] [--json] [-q] <loop.config.yaml> # Input: argv flags + a config path (no stdin). # Output: stdout = check rows (TSV: state<TAB>check<TAB>detail), or a --json envelope. # Stderr: the preflight panel, notices, errors. # Exit: 0 ok, 2 usage, 3 config not found, 4 unparseable, 5 missing core dep, # 10 a check predicts a runtime failure (a gate binary missing, bypass on # host without isolation, budget too small for a tick) # # --offline (default): no PATH/exec - config-shape + budget-vs-cost + permission/ # isolation coherence. Safe for PR CI. # --live: adds runtime preflight - claude/git on PATH, the verify/guard # leading binary resolvable, the kill-switch path's parent exists. # Skipped (not failed) when host: cloud-routine - the tick does # not run on this machine, so this machine's PATH is irrelevant. # # Examples: # loop-doctor.sh --offline .loops/pr-watch/loop.config.yaml # loop-doctor.sh --live .loops/ci-watch/loop.config.yaml # loop-doctor.sh --live --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.state=="bad")' set -uo pipefail readonly EX_OK=0 EX_USAGE=2 EX_NOTFOUND=3 EX_UNPARSEABLE=4 EX_MISSING_DEP=5 EX_FINDINGS=10 __lib="$(cd "$(dirname "${BASH_SOURCE[0]}")/../../_lib" 2>/dev/null && pwd || true)" if [ -n "${__lib:-}" ] && [ -f "$__lib/term.sh" ]; then . "$__lib/term.sh"; term_init 2 else term_panel_open() { :; }; term_panel_close() { :; }; term_panel_vert() { :; } term_status_row() { shift; printf ' - %s %s\n' "$1" "${2:-}"; } term_color() { shift; printf '%s' "$*"; }; TERM_DOT="|" fi HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" PRICING="$HERE/../assets/model-pricing.json" CFG=""; MODE="offline"; JSON=0; QUIET=0 usage() { cat <<'EOF' loop-doctor.sh - preflight a loop config (will it actually run?). Usage: loop-doctor.sh [--offline|--live] [--json] [-q] <loop.config.yaml> Options: --offline config-shape + budget-vs-cost + permission coherence (default; no PATH/exec). --live adds runtime preflight: claude/git on PATH, verify/guard binary resolvable (skipped for host: cloud-routine - ticks do not run on this machine). --json emit a JSON envelope. -q, --quiet suppress the stderr panel. -h, --help show this help and exit 0. Exit codes: 0 ok 2 usage 3 not found 4 unparseable 5 missing dep 10 predicted runtime failure Examples: loop-doctor.sh --offline .loops/pr-watch/loop.config.yaml loop-doctor.sh --live .loops/ci-watch/loop.config.yaml loop-doctor.sh --live --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.state=="bad")' EOF } die_usage() { printf 'error: %s\n' "$1" >&2; echo >&2; usage >&2; exit "$EX_USAGE"; } while [[ $# -gt 0 ]]; do case "$1" in --offline) MODE="offline"; shift ;; --live) MODE="live"; shift ;; --json) JSON=1; shift ;; -q|--quiet) QUIET=1; shift ;; -h|--help) usage; exit "$EX_OK" ;; -*) die_usage "unknown flag: $1" ;; *) [[ -z "$CFG" ]] || die_usage "unexpected extra argument: $1"; CFG="$1"; shift ;; esac done command -v awk >/dev/null 2>&1 || { echo "loop-doctor: awk required" >&2; exit "$EX_MISSING_DEP"; } command -v grep >/dev/null 2>&1 || { echo "loop-doctor: grep required" >&2; exit "$EX_MISSING_DEP"; } [[ -n "$CFG" ]] || die_usage "a loop.config.yaml path is required" [[ -f "$CFG" ]] || { printf 'error: config not found: %s\n' "$CFG" >&2; exit "$EX_NOTFOUND"; } # Normalize Windows-authored configs: strip a leading UTF-8 BOM + CR line-endings so a # CRLF/BOM file parses like a clean LF one (portable octal BOM + gsub \r). __NORM="$(mktemp 2>/dev/null)" && awk 'NR==1{sub(/^\357\273\277/,"")} {gsub(/\r/,""); print}' "$CFG" > "$__NORM" 2>/dev/null && CFG="$__NORM" && trap 'rm -f "$__NORM"' EXIT grep -Eq '^[a-z_]+:' "$CFG" || { printf 'error: no parseable keys in %s\n' "$CFG" >&2; exit "$EX_UNPARSEABLE"; } # Pick a working python for the budget-vs-cost check (skipped gracefully if none). PY="" for c in python python3 py; do if command -v "$c" >/dev/null 2>&1 && "$c" -c "" >/dev/null 2>&1; then PY="$c"; break; fi done # ── flat-YAML readers (no yq), same contract as loop-check.sh ──────────────── cfg_scalar() { awk -v k="$1" -v q="'" ' $0 ~ "^"k":" { sub("^"k":[ \t]*",""); sub(/[ \t]*#.*$/,""); gsub(/^[ \t]+|[ \t]+$/,""); gsub(/^"|"$/,""); gsub("^"q"|"q"$",""); print; exit }' "$CFG" } cfg_list_items() { awk -v k="$1" -v q="'" ' $0 ~ "^"k":" { inlist=1; next } inlist==1 { if ($0 ~ /^[ \t]*-[ \t]+/) { line=$0; sub(/^[ \t]*-[ \t]+/,"",line); sub(/[ \t]*#.*$/,"",line); gsub(/^[ \t]+|[ \t]+$/,"",line); gsub(/^"|"$/,"",line); gsub("^"q"|"q"$","",line); if (line!="") print line } else if ($0 ~ /^[^ \t#]/) { inlist=0 } }' "$CFG" } TIER="$(cfg_scalar tier)"; PMODE="$(cfg_scalar permission_mode)"; PATTERN="$(cfg_scalar pattern)" VERIFY="$(cfg_scalar verify)"; GUARD="$(cfg_scalar guard)"; BUDGET="$(cfg_scalar budget_tokens)" KILL="$(cfg_scalar kill_switch)"; ESCAL="$(cfg_scalar escalation)" CADENCE="$(cfg_scalar cadence)"; HOST="$(cfg_scalar host)"; [[ -z "$HOST" ]] && HOST="local" WORKTREE="$(cfg_scalar worktree)" is_l2plus=0; [[ "$TIER" == "L2" || "$TIER" == "L3" ]] && is_l2plus=1 # ── findings ───────────────────────────────────────────────────────────── ROWS=() # "state\tcheck\tdetail" FINDING=0 row() { ROWS+=("$1"$'\t'"$2"$'\t'"$3"); [[ "$1" == "bad" ]] && FINDING=1; } # leading binary of a command string (first whitespace token; strips a leading VAR= prefix) lead_bin() { awk '{ for(i=1;i<=NF;i++){ if($i !~ /=/){print $i; exit} } }' <<<"$1"; } # ── OFFLINE checks ─────────────────────────────────────────────────────── # Cadence in minutes, for the host floor checks. Nm/Nh/Nd, or "*/N * * * *" -> N. # Anything richer returns empty and the floor check is SKIPPED rather than guessed: # a wrong floor finding is worse than no finding. # # The digits-only guard is load-bearing, not defensive padding: `1a2m` matches the # *[0-9]m glob, and ${1%m} would hand `1a2` to arithmetic - which errors to stderr and # leaves a garbage CAD_MIN that later trips `[[ -lt ]]`. Malformed cadence is loop-check's # finding to report; here it must simply yield "unknown" and skip the floor check. cadence_minutes() { local n="" case "$1" in *[0-9]m) n="${1%m}"; [[ "$n" =~ ^[0-9]+$ ]] && printf '%s' "$n" ;; *[0-9]h) n="${1%h}"; [[ "$n" =~ ^[0-9]+$ ]] && printf '%s' "$(( n * 60 ))" ;; *[0-9]d) n="${1%d}"; [[ "$n" =~ ^[0-9]+$ ]] && printf '%s' "$(( n * 1440 ))" ;; */[0-9]*\ *) awk '{ n=$1; sub(/^\*\//,"",n); if (n ~ /^[0-9]+$/ && $2=="*") print n }' <<<"$1" ;; *) printf '' ;; esac } CAD_MIN="$(cadence_minutes "$CADENCE" 2>/dev/null)" # Host coherence. `host:` names where ticks execute; each surface has different hard # limits (references/native-scheduling.md, verified 2026-08-30). case "$HOST" in local|external|desktop-task|cloud-routine|session-cron) row ok "host" "$HOST" ;; *) row bad "host" "unknown host '$HOST' - use local|session-cron|desktop-task|cloud-routine|external" ;; esac case "$HOST" in session-cron) # /loop + CronCreate are session-scoped: they need an open, idle session and every # recurring job self-deletes 7 days after creation. Fine for L1 supervised polling; # it cannot host an unattended loop, which is what L2+ means. if [[ "$is_l2plus" -eq 1 ]]; then row bad "host/tier" "session-cron can't run unattended ($TIER) - needs an open idle session and expires after 7 days; use desktop-task or external" else row warn "host/tier" "session-cron is supervised-only: open idle session, 7-day expiry, no catch-up for missed fires" fi ;; desktop-task) # One catch-up run for the most recently missed window; older ones are discarded, # so a slow tick can land at any hour. The prompt needs its own time guardrails. if [[ -n "$CAD_MIN" ]] && [[ "$CAD_MIN" -ge 720 ]]; then row warn "catch-up" "desktop-task runs ONE catch-up for the latest missed window - a $CADENCE tick may fire hours late; put time guardrails in run.md" fi # `worktree: true` in this config is a DECLARATION, not the switch. The real toggle # lives on the task itself and is OFF by default, so a task can satisfy the config # while running against the working dir including uncommitted changes. Nothing on # disk lets us verify it - so say so rather than implying the config settled it. if [[ "$WORKTREE" == "true" ]]; then row warn "worktree" "config declares worktree: true - confirm the TASK's worktree toggle is on (it is off by default); this file cannot enforce it" fi ;; cloud-routine) # Routines run autonomously in the cloud: no permission-mode picker, >=1h floor, # fresh clone with no local files. The boundary is repos + environment + connectors. if [[ -n "$CAD_MIN" ]] && [[ "$CAD_MIN" -lt 60 ]]; then row bad "cadence" "cloud-routine minimum interval is 1 hour - '$CADENCE' is rejected at creation" fi if printf '%s %s' "$ESCAL" "$(cfg_list_items scope | tr '\n' ' ')" | grep -Eqi 'connectors?|environment|network access|repositor'; then row ok "boundary" "cloud-routine boundary names repos/environment/connectors" else row bad "boundary" "cloud-routine has NO permission mode - the boundary must be repos + environment network policy + connectors (ALL connectors attach by default); name it in scope/escalation" fi ;; esac # Permission mode achievability. A cloud routine has no permission-mode picker at all, # so requiring one there would be a false finding - the boundary check above replaces it. if [[ "$HOST" == "cloud-routine" ]]; then if [[ -n "$PMODE" ]]; then row warn "permission_mode" "'$PMODE' is ignored by cloud routines (they run autonomously, no approval prompts)" else row ok "permission_mode" "n/a for cloud-routine" fi else case "$PMODE" in default) row bad "permission_mode" "default is interactive - a headless 'claude -p' tick can't answer prompts; use dontAsk/auto/bypassPermissions" ;; "") row bad "permission_mode" "missing" ;; *) row ok "permission_mode" "$PMODE" ;; esac fi # L3 bypass needs an isolation boundary. # # A cloud routine already IS one: it runs in Anthropic-managed cloud infrastructure on a # fresh clone, and its permission_mode is ignored entirely. Demanding a "container" note # there is a false finding that teaches people to write a bogus note to satisfy the tool - # so the cloud host reports its real boundary (environment + connectors, checked above) # instead of the host-isolation one. if [[ "$TIER" == "L3" && "$HOST" == "cloud-routine" ]]; then row ok "isolation" "cloud-routine runs in Anthropic-managed cloud infra - the boundary is environment + connectors, not a local container" elif [[ "$TIER" == "L3" && "$PMODE" == "bypassPermissions" ]]; then if printf '%s %s' "$ESCAL" "$(cfg_list_items scope | tr '\n' ' ')" | grep -Eqi 'container|isolat|sandbox|devcontainer'; then row ok "isolation" "L3 bypass declares an isolation boundary" else row bad "isolation" "L3 + bypassPermissions with no container/sandbox note - only safe in an isolated VM/container" fi fi # Budget vs estimated tokens/run. if [[ -n "$BUDGET" && "$BUDGET" =~ ^[0-9]+$ && -n "$PY" && -n "$PATTERN" && -f "$PRICING" ]]; then TPR="$(PR="$PRICING" PAT="$PATTERN" "$PY" -c "import json,os try: d=json.load(open(os.environ['PR']))['_pattern_defaults'].get(os.environ['PAT']) print((int(d['input'])+int(d['output']))*int(d.get('subagents',1)) if d else '') except Exception: print('')" 2>/dev/null)" if [[ -n "$TPR" && "$TPR" =~ ^[0-9]+$ ]]; then if [[ "$BUDGET" -lt "$TPR" ]]; then row bad "budget" "budget_tokens $BUDGET < ~$TPR est. tokens/run for $PATTERN - a tick can't complete" else row ok "budget" "budget_tokens $BUDGET >= ~$TPR est. tokens/run" fi fi fi # ── LIVE checks ────────────────────────────────────────────────────────── # Skipped wholesale for cloud-routine: the tick runs on a fresh cloud clone, so this # machine's PATH, git and gate binaries say nothing about whether it will run. A pass # here would be false confidence - worse than no check. if [[ "$MODE" == "live" && "$HOST" == "cloud-routine" ]]; then row warn "live" "skipped - host cloud-routine runs on a fresh cloud clone; verify the gate in the routine's environment setup script instead" elif [[ "$MODE" == "live" ]]; then if command -v claude >/dev/null 2>&1; then row ok "claude" "on PATH"; else row warn "claude" "not on PATH - the scheduler that runs 'claude -p' must have it"; fi if command -v git >/dev/null 2>&1; then row ok "git" "on PATH" if [[ "$is_l2plus" -eq 1 ]] && ! git worktree list >/dev/null 2>&1; then row warn "worktree" "'git worktree' unavailable here - L2+ isolates changes in a worktree" fi elif [[ "$is_l2plus" -eq 1 ]]; then row bad "git" "git not on PATH - L2+ needs it for worktree isolation + landing" else row warn "git" "git not on PATH" fi # verify / guard leading binary resolvable for pair in "verify:$VERIFY" "guard:$GUARD"; do label="${pair%%:*}"; cmd="${pair#*:}" [[ -z "$cmd" ]] && continue case "$cmd" in *"<"*">"*) continue ;; esac # unfilled placeholder - audit's job bin="$(lead_bin "$cmd")" [[ -z "$bin" ]] && continue if [[ "$bin" == */* ]]; then [[ -x "$bin" ]] && row ok "$label" "$bin executable" || row bad "$label" "$bin not executable - the gate can't run" elif command -v "$bin" >/dev/null 2>&1; then row ok "$label" "$bin resolves" else row bad "$label" "'$bin' not on PATH - the gate command can't run at tick time" fi done # kill-switch path parent exists (only when it clearly names a path) ks_path="$(grep -oE '[^ "'"'"']*/[^ "'"'"']*' <<<"$KILL" | head -1)" if [[ -n "$ks_path" ]]; then parent="$(dirname "$ks_path")" [[ -d "$parent" || "$parent" == "." ]] && row ok "kill_switch" "sentinel path parent exists ($parent)" \ || row warn "kill_switch" "sentinel parent dir missing ($parent) - create it so the switch works" fi fi # ── output ─────────────────────────────────────────────────────────────── n_bad=0; n_warn=0; n_ok=0 for r in "${ROWS[@]:-}"; do case "${r%%$'\t'*}" in bad) n_bad=$((n_bad+1));; warn) n_warn=$((n_warn+1));; ok) n_ok=$((n_ok+1));; esac done if [[ "$JSON" -eq 1 ]]; then printf '{\n "data": [\n' if [[ ${#ROWS[@]} -gt 0 ]]; then for i in "${!ROWS[@]}"; do IFS=$'\t' read -r st ck dt <<<"${ROWS[$i]}" dt="${dt//\\/\\\\}"; dt="${dt//\"/\\\"}" sep=","; [[ "$i" -eq $(( ${#ROWS[@]} - 1 )) ]] && sep="" printf ' {"state": "%s", "check": "%s", "detail": "%s"}%s\n' "$st" "$ck" "$dt" "$sep" done fi printf ' ],\n "meta": {"mode": "%s", "ok": %d, "warn": %d, "bad": %d, "will_run": %s, "tier": "%s", "schema": "claude-mods.loop-ops.doctor/v1"}\n}\n' \ "$MODE" "$n_ok" "$n_warn" "$n_bad" "$([[ "$FINDING" -eq 0 ]] && echo true || echo false)" "${TIER:-unknown}" else if [[ ${#ROWS[@]} -gt 0 ]]; then for r in "${ROWS[@]}"; do IFS=$'\t' read -r st ck dt <<<"$r" printf '%-5s %-14s %s\n' "$st" "$ck" "$dt" done fi if [[ "$QUIET" -eq 0 ]]; then verdict="$([[ "$FINDING" -eq 0 ]] && echo "WILL RUN" || echo "WILL FAIL")" vstate="$([[ "$FINDING" -eq 0 ]] && echo ok || echo bad)" { term_panel_open loop "loop ${TERM_DOT} doctor ($MODE)" "$(basename "$(dirname "$CFG")")" term_panel_vert term_status_row "$vstate" "$verdict" "$n_bad blocking ${TERM_DOT} $n_warn advisory ${TERM_DOT} $n_ok ok" [[ "$MODE" == "offline" ]] && term_status_row skip "run --live before scheduling" "checks gate binaries + PATH" term_panel_vert term_panel_close "audit = well-formed ${TERM_DOT} doctor = will-run" "" } >&2 fi fi [[ "$FINDING" -eq 0 ]] && exit "$EX_OK" || exit "$EX_FINDINGS" -
loop-estimate.py 14.7 KB
#!/usr/bin/env python3 """Estimate the token/$ cost of an outer loop by pattern × cadence × model. A loop's cost is runs/day × tokens/run × price, and sub-agents multiply tokens/run. This computes that - and, crucially, models **prompt caching**: a loop re-sends the SAME run.md + system prefix every tick (the Ralph property), which is the textbook caching case. Whether caching helps depends on cadence vs cache TTL, so this picks the TTL and reports the cached projection alongside the naive one. Pricing reads from assets/model-pricing.json (date-stamped; skills/claude-api-ops is the source of truth - run its check-model-table.py if you suspect drift). Usage: loop-estimate.py --pattern P --cadence C --model M [OPTIONS] Input: argv flags only (no stdin). Output: stdout = the cost breakdown (plain rows, or --json envelope). Data only. Stderr: the assumptions + caching note, errors. Exit: 0 ok, 2 usage, 3 pricing file missing, 4 bad cadence/model/pattern Estimates, not guarantees - reconcile against the loop's run-log.md actuals. Levers in order of impact: cadence (halving frequency halves cost), prompt caching (model below), model tier. Examples: loop-estimate.py --pattern pr-watch --cadence 10m --model claude-haiku-4-5 loop-estimate.py --pattern ci-watch --cadence 15m --model claude-sonnet-5 --days 30 --json loop-estimate.py --pattern daily-scan --cadence 6h --model claude-opus-5 # too slow to cache loop-estimate.py --list-models """ from __future__ import annotations import argparse import json import os import re import sys from pathlib import Path EX_OK = 0 EX_USAGE = 2 EX_NOTFOUND = 3 EX_VALIDATION = 4 DEFAULT_PRICING = Path(__file__).resolve().parent.parent / "assets" / "model-pricing.json" # Prompt-caching multipliers vs base input price (claude-api-ops/references/caching-and-cost.md). CACHE_WRITE_5M = 1.25 # write a 5-minute-TTL entry CACHE_WRITE_1H = 2.0 # write a 1-hour-TTL entry CACHE_READ = 0.1 # read any cached entry # Minimum cacheable prefix (tokens) - below this the cache_control marker is silently # ignored (caching-and-cost.md). A loop whose static prefix is smaller can't cache. MIN_PREFIX = { "claude-fable-5": 512, "claude-opus-5": 512, "claude-sonnet-5": 1024, "claude-haiku-4-5": 4096, } DEFAULT_MIN_PREFIX = 1024 class Term: """Minimal ANSI helper (term.sh is bash-only; per TERMINAL-DESIGN.md §9 the Python port is inline). Honors FORCE_COLOR / NO_COLOR / TERM_ASCII and the bound stream's TTY + encoding, so piped data stays plain ASCII.""" _C = {"green": "\033[32m", "cyan": "\033[36m", "dim": "\033[2m", "off": "\033[0m"} def __init__(self, stream=sys.stderr): enc = (getattr(stream, "encoding", "") or "").lower() self.ascii = os.environ.get("TERM_ASCII") == "1" or "utf" not in enc if os.environ.get("FORCE_COLOR"): self.color = True elif (os.environ.get("NO_COLOR") is not None or os.environ.get("TERM") == "dumb" or not getattr(stream, "isatty", lambda: False)()): self.color = False else: self.color = True def c(self, name, text): return f"{self._C.get(name,'')}{text}{self._C['off']}" if self.color else text def load_pricing(path: Path) -> dict: if not path.is_file(): print(f"error: pricing file not found: {path}", file=sys.stderr) raise SystemExit(EX_NOTFOUND) try: return json.loads(path.read_text(encoding="utf-8")) except (json.JSONDecodeError, OSError) as exc: print(f"error: could not read pricing file: {exc}", file=sys.stderr) raise SystemExit(EX_VALIDATION) def runs_per_day(cadence: str, override: float | None) -> float: """Translate a cadence into runs/day. Supports Nm/Nh/Nd and the common cron forms `*/N * * * *` and `N * * * *`. --runs-per-day overrides everything.""" if override is not None: if override <= 0: print("error: --runs-per-day must be positive", file=sys.stderr) raise SystemExit(EX_VALIDATION) return float(override) s = cadence.strip() m = re.fullmatch(r"(\d+)([mhd])", s) if m: n = int(m.group(1)) if n <= 0: print(f"error: cadence value must be positive (got '{cadence}')", file=sys.stderr) raise SystemExit(EX_VALIDATION) return {"m": 1440.0, "h": 24.0, "d": 1.0}[m.group(2)] / n cron_min = re.fullmatch(r"\*/(\d+) \* \* \* \*", s) if cron_min: n = int(cron_min.group(1)) return 1440.0 / n if n > 0 else 1440.0 if re.fullmatch(r"\d+ \* \* \* \*", s): return 24.0 print( f"error: cannot derive runs/day from cadence '{cadence}' - " "use Nm/Nh/Nd, `*/N * * * *`, or pass --runs-per-day", file=sys.stderr, ) raise SystemExit(EX_VALIDATION) def caching_projection(in_tok, out_tok, sub, in_price, out_price, rpd, model, prefix_frac, ttl_choice): """Model prompt-caching of the static run-prompt prefix across ticks. Returns a dict: ttl, beneficial, reason, cost_per_run/day, prefix_tokens. The cache stays warm only when the tick interval is <= the TTL (reads refresh it); a loop slower than the 1h max TTL writes a cold entry every tick - caching can't help. """ interval_min = 1440.0 / rpd if rpd > 0 else 1e9 prefix_tokens = int(round(in_tok * prefix_frac)) variable_in = in_tok - prefix_tokens min_prefix = MIN_PREFIX.get(model, DEFAULT_MIN_PREFIX) # Pick TTL: smallest that stays warm at this cadence. if ttl_choice == "5m": ttl, warm = "5m", interval_min <= 5 elif ttl_choice == "1h": ttl, warm = "1h", interval_min <= 60 else: # auto if interval_min <= 5: ttl, warm = "5m", True elif interval_min <= 60: ttl, warm = "1h", True else: ttl, warm = None, False out_cost_day = out_tok / 1e6 * out_price * rpd if prefix_tokens < min_prefix: return {"ttl": ttl, "beneficial": False, "reason": f"static prefix ~{prefix_tokens} tok < {model} minimum {min_prefix} tok " "- cache marker silently ignored; enlarge the run prompt/system or skip caching", "prefix_tokens": prefix_tokens, "cost_per_day": None, "cost_per_run": None} if not warm or ttl is None: return {"ttl": ttl, "beneficial": False, "reason": f"tick interval ~{interval_min:.0f} min exceeds the cache TTL " "- the entry expires between ticks, so every tick is a cold write; caching won't help", "prefix_tokens": prefix_tokens, "cost_per_day": None, "cost_per_run": None} write_mult = CACHE_WRITE_5M if ttl == "5m" else CACHE_WRITE_1H # Per day, warm: ~1 cache write of the prefix + (rpd-1) reads; variable input + output full price. prefix_day = prefix_tokens / 1e6 * in_price * (write_mult + max(rpd - 1, 0) * CACHE_READ) variable_day = variable_in / 1e6 * in_price * rpd cost_day = (prefix_day + variable_day + out_cost_day) * sub return {"ttl": ttl, "beneficial": True, "reason": "", "prefix_tokens": prefix_tokens, "write_mult": write_mult, "cost_per_day": cost_day, "cost_per_run": cost_day / rpd if rpd else cost_day} def fmt_money(x: float) -> str: if x < 1: return f"${x:.4f}" return f"${x:,.2f}" def main(argv: list[str]) -> int: p = argparse.ArgumentParser( prog="loop-estimate.py", description="Estimate outer-loop cost by pattern × cadence × model, with prompt caching.", ) p.add_argument("--pattern", default="custom", help="catalog pattern key (default: custom)") p.add_argument("--cadence", default="1h", help="10m | 1h | 6h | 1d, or a cron string (default: 1h)") p.add_argument("--model", default="claude-haiku-4-5", help="model id (default: claude-haiku-4-5)") p.add_argument("--days", type=int, default=30, help="horizon in days for the total (default: 30)") p.add_argument("--runs-per-day", type=float, default=None, help="override the cadence-derived runs/day") p.add_argument("--input-tokens", type=int, default=None, help="override per-run input tokens") p.add_argument("--output-tokens", type=int, default=None, help="override per-run output tokens") p.add_argument("--subagents", type=int, default=None, help="override the sub-agent fan-out multiplier") p.add_argument("--cache-prefix-frac", type=float, default=0.6, help="fraction of input that is the static, cacheable run-prompt prefix (default: 0.6)") p.add_argument("--cache-ttl", choices=["auto", "5m", "1h"], default="auto", help="cache TTL to model (default: auto - pick by cadence)") p.add_argument("--no-cache", action="store_true", help="report the uncached cost only") p.add_argument("--pricing", default=str(DEFAULT_PRICING), help="path to model-pricing.json") p.add_argument("--list-models", action="store_true", help="print the pricing table + as-of date, exit 0") p.add_argument("--json", action="store_true", help="emit a JSON envelope") try: args = p.parse_args(argv) except SystemExit as exc: return EX_USAGE if exc.code not in (0, None) else (exc.code or EX_OK) pricing = load_pricing(Path(args.pricing)) models = pricing.get("models", {}) as_of = pricing.get("_as_of", "unknown") pattern_defaults = pricing.get("_pattern_defaults", {}) if args.list_models: if args.json: print(json.dumps({"data": models, "meta": {"as_of": as_of, "schema": "claude-mods.loop-ops.pricing/v1"}}, indent=2)) else: print(f"{'model':<22}{'input $/MTok':>14}{'output $/MTok':>16}") for mid, pr in models.items(): print(f"{mid:<22}{pr.get('input_per_mtok', 0):>14.2f}{pr.get('output_per_mtok', 0):>16.2f}") print(f"\n(as of {as_of}; source of truth: claude-api-ops)", file=sys.stderr) return EX_OK if args.days <= 0: print("error: --days must be positive", file=sys.stderr) return EX_VALIDATION if not (0.0 <= args.cache_prefix_frac <= 1.0): print("error: --cache-prefix-frac must be between 0 and 1", file=sys.stderr) return EX_VALIDATION if args.model not in models: print(f"error: unknown model '{args.model}' - known: {', '.join(models) or '(none)'}", file=sys.stderr) return EX_VALIDATION in_price = float(models[args.model]["input_per_mtok"]) out_price = float(models[args.model]["output_per_mtok"]) if args.input_tokens is not None and args.output_tokens is not None: in_tok, out_tok = args.input_tokens, args.output_tokens sub = args.subagents if args.subagents is not None else 1 elif args.pattern in pattern_defaults and not args.pattern.startswith("_"): d = pattern_defaults[args.pattern] in_tok = args.input_tokens if args.input_tokens is not None else int(d["input"]) out_tok = args.output_tokens if args.output_tokens is not None else int(d["output"]) sub = args.subagents if args.subagents is not None else int(d.get("subagents", 1)) else: print( f"error: unknown pattern '{args.pattern}' - pass --input-tokens and " f"--output-tokens, or use one of: {', '.join(k for k in pattern_defaults if not k.startswith('_'))}", file=sys.stderr, ) return EX_VALIDATION if min(in_tok, out_tok, sub) < 0: print("error: token counts and --subagents must be non-negative", file=sys.stderr) return EX_VALIDATION rpd = runs_per_day(args.cadence, args.runs_per_day) # ── uncached (naive) ── cost_in = in_tok / 1_000_000 * in_price cost_out = out_tok / 1_000_000 * out_price cost_run = (cost_in + cost_out) * sub tokens_run = (in_tok + out_tok) * sub cost_day = cost_run * rpd cost_horizon = cost_day * args.days # ── cached projection ── cache = None if not args.no_cache: cache = caching_projection(in_tok, out_tok, sub, in_price, out_price, rpd, args.model, args.cache_prefix_frac, args.cache_ttl) if args.json: data = { "pattern": args.pattern, "model": args.model, "cadence": args.cadence, "runs_per_day": round(rpd, 3), "tokens_per_run": tokens_run, "input_tokens": in_tok, "output_tokens": out_tok, "subagents": sub, "cost_per_run": round(cost_run, 6), "cost_per_day": round(cost_day, 4), "days": args.days, "cost_per_horizon": round(cost_horizon, 2), } if cache is not None: if cache["beneficial"]: cd = cache["cost_per_day"] data["caching"] = { "beneficial": True, "ttl": cache["ttl"], "prefix_tokens": cache["prefix_tokens"], "cost_per_day": round(cd, 4), "cost_per_horizon": round(cd * args.days, 2), "savings_pct": round((cost_day - cd) / cost_day * 100, 1) if cost_day else 0.0, } else: data["caching"] = {"beneficial": False, "reason": cache["reason"], "prefix_tokens": cache["prefix_tokens"]} print(json.dumps({"data": data, "meta": {"as_of": as_of, "schema": "claude-mods.loop-ops.estimate/v1"}}, indent=2)) return EX_OK t = Term(sys.stderr) print(f"{'pattern:':<16}{args.pattern}") print(f"{'model:':<16}{args.model}") print(f"{'cadence:':<16}{args.cadence} -> {rpd:g} runs/day") print(f"{'tokens/run:':<16}{tokens_run:,} ({in_tok:,} in + {out_tok:,} out) x {sub} subagent(s)") print(f"{'cost/run:':<16}{fmt_money(cost_run)}") print(f"{'cost/day:':<16}{fmt_money(cost_day)}") print(f"{'cost/'+str(args.days)+'d:':<16}{fmt_money(cost_horizon)} (uncached)") if cache is not None: if cache["beneficial"]: cd, ch = cache["cost_per_day"], cache["cost_per_day"] * args.days save = (cost_day - cd) / cost_day * 100 if cost_day else 0.0 print(f"{'cached/'+str(args.days)+'d:':<16}{t.c('cyan', fmt_money(ch))} " f"({t.c('green', f'-{save:.0f}%')}, TTL {cache['ttl']}, prefix ~{cache['prefix_tokens']:,} tok)") print(f"recommendation: cache the static run.md+system prefix at TTL {cache['ttl']} " f"-> ~-{save:.0f}%/mo. Keep run.md BYTE-IDENTICAL every tick or the cache never hits.", file=sys.stderr) else: print(f"caching: not beneficial here", file=sys.stderr) print(f" why: {cache['reason']}", file=sys.stderr) print(f"estimate (as of {as_of} pricing) - reconcile against run-log.md actuals; " "cadence is the biggest lever, then caching, then model tier", file=sys.stderr) return EX_OK if __name__ == "__main__": sys.exit(main(sys.argv[1:])) -
loop-scaffold.sh 16 KB
#!/usr/bin/env bash # Scaffold an outer-loop state spine (loop.config.yaml + STATE.md + run-log.md). # # Usage: loop-scaffold.sh --name NAME [OPTIONS] # Input: argv flags only (no stdin). # Output: stdout = the created loop.config.yaml path (data). Under --dry-run, the # path then the rendered config. Data only. # Stderr: the creation panel, reminders, warnings, errors. # Exit: 0 created (or dry-run rendered), 2 usage, 3 template/dir not found, # 5 precondition (target dir already populated, no --force) # # Creates <dir>/<name>/ from the bundled templates, substituting name/pattern/tier/ # cadence/permission_mode. Never clobbers a populated loop dir. Atomic writes. # Next step: fill the config, then `loop-check.sh <dir>/<name>/loop.config.yaml`. # # Examples: # loop-scaffold.sh --name pr-watch --pattern pr-watch --tier L1 # loop-scaffold.sh --name dep-bump --pattern dep-bump --tier L2 --cadence 1d # loop-scaffold.sh --name nightly --cadence "0 3 * * *" --dry-run # loop-scaffold.sh --name digest --pattern digest --host cloud-routine --cadence 1h set -uo pipefail readonly EX_OK=0 EX_USAGE=2 EX_NOTFOUND=3 EX_PRECOND=5 # Terminal design system (skills/_lib/term.sh). stdout = the created path (data); # the creation panel frames on stderr, so detect color on fd 2. Degrade to plain # stderr lines if the shared lib is unreachable. __lib="$(cd "$(dirname "${BASH_SOURCE[0]}")/../../_lib" 2>/dev/null && pwd || true)" if [ -n "${__lib:-}" ] && [ -f "$__lib/term.sh" ]; then . "$__lib/term.sh"; term_init 2 else term_panel_open() { :; }; term_panel_close() { :; }; term_panel_vert() { :; } term_status_row() { shift; printf ' - %s %s\n' "$1" "${2:-}"; } term_alert() { shift; printf ' ! %s\n' "$*"; } term_color() { shift; printf '%s' "$*"; }; TERM_DOT="|" fi HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" ASSETS="$HERE/../assets" CFG_TPL="$ASSETS/loop.config.template.yaml" STATE_TPL="$ASSETS/STATE.template.md" RUN_TPL="$ASSETS/run.template.md" RUN_SH_TPL="$ASSETS/run.sh.template" # ── defaults ──────────────────────────────────────────────────────────────── NAME="" PATTERN="custom" TIER="L1" CADENCE="1h" HOST="local" DIR=".loops" DRY_RUN=0 FORCE=0 usage() { cat <<'EOF' loop-scaffold.sh — scaffold an outer-loop state spine. Usage: loop-scaffold.sh --name NAME [OPTIONS] Options: --name NAME loop identifier, kebab-case (required). Names the directory. --pattern KEY catalog key (pr-watch, ci-watch, dep-bump, changelog-gen, merge-hygiene, issue-sort, daily-scan) or "custom" (default: custom). --tier L1|L2|L3 starting autonomy tier (default: L1). --cadence STR 10m | 1h | 6h | 1d, or a cron string (default: 1h). --host HOST where ticks execute: local (default) | session-cron | desktop-task | cloud-routine | external. Decides which constraints loop-doctor enforces (references/native-scheduling.md). --dir DIR parent directory for the loop (default: .loops). --dry-run print the target path + rendered config; write nothing. --force overwrite an already-populated <dir>/<name>/ directory. -h, --help show this help and exit 0. Exit codes: 0 created (or dry-run) 2 usage 3 template/dir not found 5 dir populated Examples: loop-scaffold.sh --name pr-watch --pattern pr-watch --tier L1 loop-scaffold.sh --name dep-bump --pattern dep-bump --tier L2 --cadence 1d loop-scaffold.sh --name nightly --cadence "0 3 * * *" --dry-run loop-scaffold.sh --name digest --pattern digest --host cloud-routine --cadence 1h EOF } die_usage() { printf 'error: %s\n' "$1" >&2; echo >&2; usage >&2; exit "$EX_USAGE"; } # ── parse args ────────────────────────────────────────────────────────────── while [[ $# -gt 0 ]]; do case "$1" in --name) [[ $# -ge 2 ]] || die_usage "--name needs a value"; NAME="$2"; shift 2 ;; --pattern) [[ $# -ge 2 ]] || die_usage "--pattern needs a value"; PATTERN="$2"; shift 2 ;; --tier) [[ $# -ge 2 ]] || die_usage "--tier needs a value"; TIER="$2"; shift 2 ;; --cadence) [[ $# -ge 2 ]] || die_usage "--cadence needs a value"; CADENCE="$2"; shift 2 ;; --host) [[ $# -ge 2 ]] || die_usage "--host needs a value"; HOST="$2"; shift 2 ;; --dir) [[ $# -ge 2 ]] || die_usage "--dir needs a value"; DIR="$2"; shift 2 ;; --dry-run) DRY_RUN=1; shift ;; --force) FORCE=1; shift ;; -h|--help) usage; exit "$EX_OK" ;; -*) die_usage "unknown flag: $1" ;; *) die_usage "unexpected positional argument: $1" ;; esac done # ── validate ──────────────────────────────────────────────────────────────── [[ -n "$NAME" ]] || die_usage "--name is required" [[ "$NAME" =~ ^[a-z0-9]+(-[a-z0-9]+)*$ ]] || die_usage "--name must be kebab-case (got '$NAME')" [[ "$PATTERN" =~ ^[a-z0-9]+(-[a-z0-9]+)*$ ]] || die_usage "--pattern must be kebab-case (got '$PATTERN')" case "$TIER" in L1|L2|L3) ;; *) die_usage "--tier must be L1|L2|L3 (got '$TIER')" ;; esac # cadence: Nm/Nh/Nd OR a cron-ish string (digits, spaces, * / , -) [[ "$CADENCE" =~ ^[0-9]+[mhd]$ || "$CADENCE" =~ ^[-0-9*/,\ ]+$ ]] \ || die_usage "--cadence must be like 10m/1h/1d or a cron string (got '$CADENCE')" # host: the execution surface. Its constraints are enforced by loop-doctor, not here - # scaffolding a not-yet-valid combination is fine; scheduling one is not. case "$HOST" in local|session-cron|desktop-task|cloud-routine|external) ;; *) die_usage "--host must be local|session-cron|desktop-task|cloud-routine|external (got '$HOST')" ;; esac [[ -f "$CFG_TPL" ]] || { printf 'error: config template not found at %s\n' "$CFG_TPL" >&2; exit "$EX_NOTFOUND"; } [[ -f "$STATE_TPL" ]] || { printf 'error: STATE template not found at %s\n' "$STATE_TPL" >&2; exit "$EX_NOTFOUND"; } [[ -f "$RUN_TPL" ]] || { printf 'error: run template not found at %s\n' "$RUN_TPL" >&2; exit "$EX_NOTFOUND"; } [[ -f "$RUN_SH_TPL" ]] || { printf 'error: run.sh template not found at %s\n' "$RUN_SH_TPL" >&2; exit "$EX_NOTFOUND"; } # Default permission_mode from tier (the workhorse mapping; see references/risk-tiers.md). case "$TIER" in L1|L2) PMODE="dontAsk" ;; L3) PMODE="bypassPermissions" ;; esac # ── pattern presets ───────────────────────────────────────────────────────── # Seed a near-ready config for a known --pattern (the user reviews, doesn't start # from blank placeholders). Doctrine: always scaffold at the chosen tier; report/ # propose/draft patterns carry no gate (VERIFY_SEED empty), code-changing ones do. SEEDED=0; SCOPE_SEED=""; GOAL_SEED=""; ESCAL_SEED=""; VERIFY_SEED=""; GUARD_SEED=""; BUDGET_SEED="" case "$PATTERN" in daily-scan) SEEDED=1 SCOPE_SEED="src/**" GOAL_SEED="Sweep the backlog/issues/alerts and write the day's STATE.md priority list; report only." ESCAL_SEED="everything - a human decides what to action; this loop never changes code" ;; pr-watch) SEEDED=1 SCOPE_SEED="src/**" GOAL_SEED="Watch open PRs; flag stuck/failing/conflicted; post a summary comment at most; never merge." ESCAL_SEED="a human reviews and merges; never merge to main" ;; ci-watch) SEEDED=1 SCOPE_SEED="src/**" GOAL_SEED="Detect red CI; classify the failure; at L2 propose a fix in a worktree; never auto-merge to main." ESCAL_SEED="flaky/infra failures, anything touching deploy/secrets, ambiguous root cause" VERIFY_SEED="npm test"; GUARD_SEED="npm run typecheck" ;; dep-bump) SEEDED=1 SCOPE_SEED="package.json" GOAL_SEED="Patch-only dependency bumps behind the release cooldown + guard; open a PR; never minor/major." ESCAL_SEED="minor/major bumps, guard failures, any flagged advisory" VERIFY_SEED="npm test"; GUARD_SEED="npm run build && npm test" ;; changelog-gen) SEEDED=1 SCOPE_SEED="CHANGELOG.md" GOAL_SEED="Summarize merged PRs since the last tag into RELEASE_NOTES_DRAFT.md; never publish a release." ESCAL_SEED="the human edits and publishes; never run gh release create" ;; merge-hygiene) SEEDED=1 SCOPE_SEED="src/**" GOAL_SEED="Find merged-deletable branches / stale flags / orphaned artifacts; report; never delete unmerged work." ESCAL_SEED="anything ambiguous; never delete a branch with unmerged commits" ;; issue-sort) SEEDED=1 SCOPE_SEED="src/**" GOAL_SEED="Classify new issues and suggest labels + priority; propose only; never close or set priority unattended." ESCAL_SEED="priority calls, dupe-closing, anything needing product judgment" ;; metric-chase) SEEDED=1 SCOPE_SEED="src/**" GOAL_SEED="Drive a measurable target (coverage/latency/bundle/eval score) to goal via iterate; keep gains, discard regressions." ESCAL_SEED="target unreachable after the budget, guard failures, any change to the gate/test itself" VERIFY_SEED="npm test -- --coverage"; GUARD_SEED="npm run typecheck" BUDGET_SEED=400000 ;; # iterate fan-out is the most expensive tick — fit it regression-watch) SEEDED=1 SCOPE_SEED="bench/**" GOAL_SEED="Run the benchmark/eval suite, diff against the recorded baseline; report a regression; never edit the suite." ESCAL_SEED="a confirmed regression (a human triages); a single flaky run is advisory, not a page" ;; digest) SEEDED=1 SCOPE_SEED="reports/**" GOAL_SEED="Summarize email/Asana/calendar/news via connectors into a morning report; read-only, never act." ESCAL_SEED="anything requiring a reply or an action; this loop only summarizes" ;; backfill) SEEDED=1 SCOPE_SEED="src/**" GOAL_SEED="Drain a migration/queue to completion via /goal; one item per step, verify each; stop when empty or after the bound." ESCAL_SEED="any item needing a judgment call; never exceed the stop-after-N / token bound" VERIFY_SEED="npm test"; GUARD_SEED="npm run typecheck" ;; monitor) SEEDED=1 SCOPE_SEED="src/**" GOAL_SEED="React to an error/log/deploy event (via a Channel); triage it and page a human on a real anomaly; never auto-remediate prod." ESCAL_SEED="any anomaly worth a human; production remediation; anything destructive" ;; freshness) SEEDED=1 SCOPE_SEED="docs/**" GOAL_SEED="Re-check docs/data/deps/links against reality on a cadence; report confirmed drift; never auto-edit on a transient failure." ESCAL_SEED="confirmed drift a human should fix; a transient/network failure is advisory only" ;; esac TARGET_DIR="$DIR/$NAME" CFG_OUT="$TARGET_DIR/loop.config.yaml" STATE_OUT="$TARGET_DIR/STATE.md" LOG_OUT="$TARGET_DIR/run-log.md" RUN_OUT="$TARGET_DIR/run.md" RUN_SH_OUT="$TARGET_DIR/loop-run.sh" # Refuse a populated target unless --force. if [[ -d "$TARGET_DIR" ]] && [[ -n "$(ls -A "$TARGET_DIR" 2>/dev/null)" ]] && [[ "$FORCE" -ne 1 ]]; then printf 'error: loop directory already populated: %s (use --force to overwrite)\n' "$TARGET_DIR" >&2 exit "$EX_PRECOND" fi NOW="$(date -u +%Y-%m-%dT%H:%M:%SZ)" # ── render config from template ───────────────────────────────────────────── # Line-anchored sed substitutions: identity placeholders globally, the three # tunable scalar lines by their default value. Kill-switch path carries <loop-name>. render_config() { sed -E \ -e "s|<loop-name>|$NAME|g" \ -e "s|<pattern-key>|$PATTERN|" \ -e "s|^tier: L1|tier: $TIER|" \ -e "s|^cadence: 1h|cadence: $CADENCE|" \ -e "s|^host: local|host: $HOST|" \ -e "s|^permission_mode: dontAsk|permission_mode: $PMODE|" \ "$CFG_TPL" } render_state() { sed -E \ -e "s|<loop-name>|$NAME|g" \ -e "s|<ISO-8601 Z>|$NOW|" \ "$STATE_TPL" } render_log() { cat <<EOF # $NAME — run log (append-only; one line per run) # format: <ISO-Z> run#N action=<reported|proposed|none> <key=val…> outcome=<…> tokens=<N> EOF } render_run() { sed -E \ -e "s|<loop-name>|$NAME|g" \ -e "s|tier <L1\\|L2\\|L3>|tier $TIER|g" \ "$RUN_TPL" } # The runner-agnostic tick wrapper any scheduler invokes (cron / Task Scheduler / # systemd / process-compose / by hand) — no GitHub Actions required. render_run_sh() { sed -E \ -e "s|<loop-name>|$NAME|g" \ -e "s|<permission-mode>|$PMODE|g" \ "$RUN_SH_TPL" } # Seeded config for a known pattern. L1 stays report-only (gate fields are a # commented graduation block); L2/L3 emit verify/guard/worktree/land_via — using # the pattern's gate if it has one, else a <fill:…> placeholder the audit will flag. render_seeded_config() { cat <<EOF # loop.config.yaml - $PATTERN (seeded by loop-scaffold at $TIER; REVIEW before scheduling) # Full field semantics: skills/loop-ops/references/state-spine.md name: $NAME pattern: $PATTERN tier: $TIER permission_mode: $PMODE cadence: $CADENCE host: $HOST goal: "$GOAL_SEED" scope: - "$SCOPE_SEED" escalation: "$ESCAL_SEED" budget_tokens: ${BUDGET_SEED:-200000} kill_switch: ".loops/$NAME/PAUSED exists, OR the loop-pause label is set" EOF if [[ "$TIER" == "L1" ]]; then cat <<EOF # ── graduate to L2 (assisted): set tier: L2, uncomment + fill, re-run loop-check + loop-doctor --live ── # verify: "${VERIFY_SEED:-<fill: the gate command, e.g. npm test>}" # guard: "${GUARD_SEED:-<fill: a must-always-pass command>}" # worktree: true # land_via: fleet-ops EOF else cat <<EOF verify: "${VERIFY_SEED:-<fill: the gate command for this loop>}" guard: "${GUARD_SEED:-<fill: a must-always-pass command>}" worktree: true land_via: fleet-ops EOF fi } # Pick the seeded renderer for a known pattern, else the generic template. emit_config() { if [[ "$SEEDED" -eq 1 ]]; then render_seeded_config; else render_config; fi; } # ── dry-run: print and stop ───────────────────────────────────────────────── if [[ "$DRY_RUN" -eq 1 ]]; then printf '%s\n' "$CFG_OUT" { term_panel_open loop "loop ${TERM_DOT} init (dry-run)" "$NAME" term_panel_vert term_status_row skip "would create $TARGET_DIR/" "tier $TIER ${TERM_DOT} $PATTERN ${TERM_DOT} $CADENCE" term_status_row skip " loop.config.yaml" "permission_mode: $PMODE" term_status_row skip " STATE.md / run-log.md / run.md / loop-run.sh" "" term_panel_vert term_panel_close "nothing written" "" } >&2 emit_config exit "$EX_OK" fi # ── atomic writes ─────────────────────────────────────────────────────────── mkdir -p "$TARGET_DIR" || { printf 'error: could not create %s\n' "$TARGET_DIR" >&2; exit 1; } write_atomic() { # write_atomic <dest> <content> local dest="$1" content="$2" tmp tmp="$dest.tmp.$$" printf '%s\n' "$content" > "$tmp" || { printf 'error: failed to write %s\n' "$tmp" >&2; exit 1; } mv -f "$tmp" "$dest" || { rm -f "$tmp"; printf 'error: failed to move into place: %s\n' "$dest" >&2; exit 1; } } write_atomic "$CFG_OUT" "$(emit_config)" write_atomic "$STATE_OUT" "$(render_state)" write_atomic "$LOG_OUT" "$(render_log)" write_atomic "$RUN_OUT" "$(render_run)" write_atomic "$RUN_SH_OUT" "$(render_run_sh)" chmod +x "$RUN_SH_OUT" 2>/dev/null || true printf '%s\n' "$CFG_OUT" { term_panel_open loop "loop ${TERM_DOT} init" "$NAME" term_panel_vert term_status_row ok "created $TARGET_DIR/" "tier $TIER ${TERM_DOT} $PATTERN ${TERM_DOT} $CADENCE" term_status_row ok " loop.config.yaml" "permission_mode: $PMODE" term_status_row ok " STATE.md / run-log.md / run.md / loop-run.sh" "" if [[ "$TIER" != "L1" ]]; then term_alert warning "tier $TIER needs a verify gate, guard, worktree, escalation + land_via — fill them before auditing" fi term_panel_vert term_panel_close "then: fill the config ${TERM_DOT} loop-check.sh $CFG_OUT" "" } >&2 exit "$EX_OK"
-
-
tests
-
run.sh 32 KB
#!/usr/bin/env bash # Self-test for loop-ops scripts (loop-scaffold.sh, loop-check.sh, loop-estimate.py). # # Offline-deterministic (no network). Scaffolds throwaway loop fixtures, asserts the # documented exit codes + key output of each script, then cleans up. Resolves paths # relative to itself so it works both in the repo and installed to ~/.claude/. # # Usage: bash tests/run.sh # Exit: 0 all pass, 1 one or more failures set -uo pipefail HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SKILL="$(dirname "$HERE")" SCRIPTS="$SKILL/scripts" INIT="$SCRIPTS/loop-scaffold.sh" AUDIT="$SCRIPTS/loop-check.sh" COST="$SCRIPTS/loop-estimate.py" SYNC="$SCRIPTS/check-pricing-sync.py" DOCTOR="$SCRIPTS/loop-doctor.sh" # Pick a python that actually executes — skips the Windows Store python3 stub. PYTHON="" for c in python python3 py; do if command -v "$c" >/dev/null 2>&1 && "$c" -c "" >/dev/null 2>&1; then PYTHON="$c"; break; fi done [[ -z "$PYTHON" ]] && { echo "no working python found — skipping" >&2; exit 0; } SB="$(mktemp -d)"; trap 'rm -rf "$SB"' EXIT PASS=0; FAIL=0 ok() { PASS=$((PASS+1)); printf ' PASS %s\n' "$1"; } no() { FAIL=$((FAIL+1)); printf ' FAIL %s\n' "$1"; } expect_exit() { [[ "$2" == "$3" ]] && ok "$1 (exit $3)" || no "$1 (want $2 got $3)"; } expect_has() { case "$3" in *"$2"*) ok "$1";; *) no "$1 (missing '$2')";; esac; } # Write a filled, READY L1 report-only config. good_l1() { cat > "$1" <<'EOF' name: test-l1 pattern: pr-watch tier: L1 permission_mode: dontAsk cadence: 10m goal: "Watch open PRs and report; never merge." scope: - "src/**" escalation: "comment on the PR; never merge to main" budget_tokens: 200000 kill_switch: ".loops/test-l1/PAUSED exists or loop-pause label" EOF } # Write a filled, READY L2 assisted config. good_l2() { cat > "$1" <<'EOF' name: dep-bump pattern: dep-bump tier: L2 permission_mode: dontAsk cadence: 1d goal: "Patch-only dependency bumps behind cooldown; open a PR." scope: - "package.json" - "package-lock.json" verify: "npm test" guard: "npm run typecheck" worktree: true land_via: fleet-ops escalation: "minor/major bumps escalate; never merge to main" budget_tokens: 300000 kill_switch: ".loops/dep-bump/PAUSED" EOF } echo "=== loop-ops self-test (python: $PYTHON) ===" # ── --help contracts (exit 0) ────────────────────────────────────────────── echo "-- --help --" bash "$INIT" --help >/dev/null 2>&1; expect_exit "loop-scaffold --help" 0 $? bash "$AUDIT" --help >/dev/null 2>&1; expect_exit "loop-check --help" 0 $? "$PYTHON" "$COST" --help >/dev/null 2>&1; expect_exit "loop-estimate --help" 0 $? # ── loop-scaffold: scaffolds dir + 3 files, substitutes fields ───────────────── echo "-- loop-scaffold --" out="$(bash "$INIT" --name pr-watch --pattern pr-watch --tier L1 --cadence 5m --dir "$SB/loops" 2>/dev/null)"; rc=$? expect_exit "loop-scaffold -> 0" 0 "$rc" expect_has "prints the config path" "pr-watch/loop.config.yaml" "$out" [[ -f "$SB/loops/pr-watch/loop.config.yaml" ]] && ok "wrote loop.config.yaml" || no "no loop.config.yaml" [[ -f "$SB/loops/pr-watch/STATE.md" ]] && ok "wrote STATE.md" || no "no STATE.md" [[ -f "$SB/loops/pr-watch/run-log.md" ]] && ok "wrote run-log.md" || no "no run-log.md" [[ -f "$SB/loops/pr-watch/run.md" ]] && ok "wrote run.md" || no "no run.md" runmd="$(cat "$SB/loops/pr-watch/run.md")" expect_has "run.md substitutes loop name" "Run: pr-watch" "$runmd" expect_has "run.md substitutes tier" "tier L1)" "$runmd" # runner-agnostic wrapper: emitted, executable, fully substituted, no GH Actions dep [[ -f "$SB/loops/pr-watch/loop-run.sh" ]] && ok "wrote loop-run.sh" || no "no loop-run.sh" runsh="$(cat "$SB/loops/pr-watch/loop-run.sh")" case "$runsh" in *"<loop-name>"*|*"<permission-mode>"*) no "loop-run.sh left a placeholder";; *) ok "loop-run.sh fully substituted";; esac expect_has "loop-run.sh wires the gated mode" "--permission-mode dontAsk" "$runsh" cfg="$(cat "$SB/loops/pr-watch/loop.config.yaml")" expect_has "substituted name" "name: pr-watch" "$cfg" expect_has "substituted tier" "tier: L1" "$cfg" expect_has "substituted cadence" "cadence: 5m" "$cfg" expect_has "L1 default permission_mode" "permission_mode: dontAsk" "$cfg" # L3 default permission_mode is bypassPermissions bash "$INIT" --name big-job --tier L3 --dir "$SB/loops" >/dev/null 2>&1 expect_has "L3 default permission_mode" "permission_mode: bypassPermissions" "$(cat "$SB/loops/big-job/loop.config.yaml")" # ── loop-scaffold: refuses a populated dir -> 5, --force overwrites ───────────── bash "$INIT" --name pr-watch --dir "$SB/loops" >/dev/null 2>&1; expect_exit "refuse populated dir -> 5" 5 $? bash "$INIT" --name pr-watch --dir "$SB/loops" --force >/dev/null 2>&1; expect_exit "--force overwrites -> 0" 0 $? # ── loop-scaffold: --dry-run writes nothing ──────────────────────────────────── out="$(bash "$INIT" --name ghost --dir "$SB/dryloops" --dry-run 2>/dev/null)"; rc=$? expect_exit "dry-run -> 0" 0 "$rc" [[ -e "$SB/dryloops" ]] && no "dry-run created files" || ok "dry-run wrote nothing" expect_has "dry-run prints config path" "ghost/loop.config.yaml" "$out" # ── loop-scaffold: usage errors ──────────────────────────────────────────────── bash "$INIT" --dir "$SB/loops" >/dev/null 2>&1; expect_exit "missing --name -> 2" 2 $? bash "$INIT" --name BadName --dir "$SB/loops" >/dev/null 2>&1; expect_exit "non-kebab name -> 2" 2 $? bash "$INIT" --name x --tier L9 --dir "$SB/loops" >/dev/null 2>&1; expect_exit "bad tier -> 2" 2 $? # pattern-seeding: a known pattern seeds a near-ready, audit-clean config bash "$INIT" --name seed-l1 --pattern ci-watch --tier L1 --cadence 15m --dir "$SB/seed" >/dev/null 2>&1 seedcfg="$(cat "$SB/seed/seed-l1/loop.config.yaml")" expect_has "seeded config carries the pattern goal" "Detect red CI" "$seedcfg" expect_has "seeded L1 leaves a graduation block" "graduate to L2" "$seedcfg" bash "$AUDIT" "$SB/seed/seed-l1/loop.config.yaml" >/dev/null 2>&1; expect_exit "seeded L1 audits clean -> 0" 0 $? # at L2 the pattern's gate is filled (not commented) and audits clean bash "$INIT" --name seed-l2 --pattern ci-watch --tier L2 --cadence 15m --dir "$SB/seed" >/dev/null 2>&1 l2cfg="$(cat "$SB/seed/seed-l2/loop.config.yaml")" case "$l2cfg" in *$'\nverify: "npm test"'*) ok "seeded L2 fills the gate";; *) no "seeded L2 did not fill the gate";; esac bash "$AUDIT" "$SB/seed/seed-l2/loop.config.yaml" >/dev/null 2>&1; expect_exit "seeded L2 audits clean -> 0" 0 $? # an unknown pattern falls back to the generic placeholder template (not ready) bash "$INIT" --name seed-x --pattern custom --tier L1 --dir "$SB/seed" >/dev/null 2>&1 case "$(cat "$SB/seed/seed-x/loop.config.yaml")" in *"<one sentence"*) ok "unknown pattern uses generic template";; *) no "unknown pattern did not use template";; esac # v2 archetypes: scaffold must be audit-clean AND doctor-clean at L1 + known to the cost model # (doctor-clean catches budget < tokens/run — the metric-chase trap: it seeds a bigger budget) for p in metric-chase regression-watch digest backfill monitor freshness; do bash "$INIT" --name "a-$p" --pattern "$p" --tier L1 --dir "$SB/arch" >/dev/null 2>&1 bash "$AUDIT" "$SB/arch/a-$p/loop.config.yaml" >/dev/null 2>&1; expect_exit "archetype $p seeds audit-clean (L1)" 0 $? bash "$DOCTOR" --offline "$SB/arch/a-$p/loop.config.yaml" >/dev/null 2>&1; expect_exit "archetype $p doctors clean (L1)" 0 $? "$PYTHON" "$COST" --pattern "$p" --cadence 1h --model claude-haiku-4-5 >/dev/null 2>&1; expect_exit "cost model knows $p" 0 $? done # the most expensive archetype at L2: gate filled, budget fits the tick (audit + doctor clean) bash "$INIT" --name a-mc --pattern metric-chase --tier L2 --cadence 1h --dir "$SB/arch" >/dev/null 2>&1 bash "$AUDIT" "$SB/arch/a-mc/loop.config.yaml" >/dev/null 2>&1; expect_exit "metric-chase L2 audits clean -> 0" 0 $? bash "$DOCTOR" --offline "$SB/arch/a-mc/loop.config.yaml" >/dev/null 2>&1; expect_exit "metric-chase L2 doctors clean (budget fits) -> 0" 0 $? # ── loop-scaffold: --host records the execution surface ──────────────────────── # `host:` selects which constraints loop-doctor enforces (references/native-scheduling.md). echo "-- loop-scaffold --host --" bash "$INIT" --name h-default --dir "$SB/hosts" >/dev/null 2>&1 expect_has "host defaults to local" "host: local" "$(cat "$SB/hosts/h-default/loop.config.yaml")" bash "$INIT" --name h-cloud --host cloud-routine --dir "$SB/hosts" >/dev/null 2>&1 expect_has "template render carries --host" "host: cloud-routine" "$(cat "$SB/hosts/h-cloud/loop.config.yaml")" # the seeded (known --pattern) path is a separate renderer - it must carry host too bash "$INIT" --name h-seed --pattern digest --host desktop-task --dir "$SB/hosts" >/dev/null 2>&1 expect_has "seeded render carries --host" "host: desktop-task" "$(cat "$SB/hosts/h-seed/loop.config.yaml")" bash "$INIT" --name h-bad --host nonsense --dir "$SB/hosts" >/dev/null 2>&1; expect_exit "unknown --host -> 2" 2 $? # ── loop-check: a freshly-init'd config is NOT ready (placeholders) -> 10 ─── echo "-- loop-check --" bash "$INIT" --name raw --pattern custom --tier L1 --dir "$SB/loops" >/dev/null 2>&1 out="$(bash "$AUDIT" "$SB/loops/raw/loop.config.yaml" 2>/dev/null)"; rc=$? expect_exit "raw scaffold not ready -> 10" 10 "$rc" expect_has "flags the goal placeholder" "goal:" "$out" # ── loop-check: filled L1 config is READY -> 0 ───────────────────────────── good_l1 "$SB/l1.yaml" out="$(bash "$AUDIT" "$SB/l1.yaml" 2>/dev/null)"; rc=$? expect_exit "filled L1 ready -> 0" 0 "$rc" # ── loop-check: filled L2 config is READY -> 0 ───────────────────────────── good_l2 "$SB/l2.yaml" bash "$AUDIT" "$SB/l2.yaml" >/dev/null 2>&1; expect_exit "filled L2 ready -> 0" 0 $? # ── loop-check: L2 missing the gate -> 10, names verify ──────────────────── grep -v '^verify:' "$SB/l2.yaml" > "$SB/l2-nogate.yaml" out="$(bash "$AUDIT" "$SB/l2-nogate.yaml" 2>/dev/null)"; rc=$? expect_exit "L2 missing gate -> 10" 10 "$rc" expect_has "names the missing gate" "verify:" "$out" # ── loop-check: unbounded scope -> 10 ────────────────────────────────────── sed 's| - "src/\*\*"| - "*"|' "$SB/l1.yaml" > "$SB/l1-unbounded.yaml" out="$(bash "$AUDIT" "$SB/l1-unbounded.yaml" 2>/dev/null)"; rc=$? expect_exit "unbounded scope -> 10" 10 "$rc" expect_has "names unbounded scope" "unbounded" "$out" # ── loop-check: missing escalation -> 10 ─────────────────────────────────── grep -v '^escalation:' "$SB/l1.yaml" > "$SB/l1-noescal.yaml" out="$(bash "$AUDIT" "$SB/l1-noescal.yaml" 2>/dev/null)"; rc=$? expect_exit "missing escalation -> 10" 10 "$rc" expect_has "names escalation" "escalation:" "$out" # ── loop-check: missing file -> 3, unparseable -> 4, bad --min -> 2 ──────── bash "$AUDIT" "$SB/no-such.yaml" >/dev/null 2>&1; expect_exit "missing config -> 3" 3 $? printf 'just some prose, no keys\n' > "$SB/garbage.yaml" bash "$AUDIT" "$SB/garbage.yaml" >/dev/null 2>&1; expect_exit "unparseable -> 4" 4 $? bash "$AUDIT" --min abc "$SB/l1.yaml" >/dev/null 2>&1; expect_exit "bad --min -> 2" 2 $? # ── loop-check: --json envelope schema + ready flag ──────────────────────── out="$(bash "$AUDIT" --json "$SB/l1.yaml" 2>/dev/null)" expect_has "audit json schema" "claude-mods.loop-ops.check/v1" "$out" expect_has "audit json ready true" '"ready": true' "$out" out="$(bash "$AUDIT" --json "$SB/l2-nogate.yaml" 2>/dev/null)" expect_has "audit json ready false" '"ready": false' "$out" # ── loop-check: --strict turns a warning into NOT ready ──────────────────── # An L1 with permission_mode: auto is consistent-enough to pass errors but warns # (broad for L1). Normally ready; --strict flips it. sed 's|permission_mode: dontAsk|permission_mode: auto|' "$SB/l1.yaml" > "$SB/l1-warn.yaml" bash "$AUDIT" "$SB/l1-warn.yaml" >/dev/null 2>&1; expect_exit "warning, normally ready -> 0" 0 $? bash "$AUDIT" --strict "$SB/l1-warn.yaml" >/dev/null 2>&1; expect_exit "warning, --strict not ready -> 10" 10 $? # ── loop-estimate: basic run, --json, --list-models, cadence forms ───────────── echo "-- loop-estimate --" out="$("$PYTHON" "$COST" --pattern pr-watch --cadence 10m --model claude-haiku-4-5 2>/dev/null)"; rc=$? expect_exit "loop-estimate -> 0" 0 "$rc" expect_has "prints a daily cost" "cost/day:" "$out" expect_has "derives runs/day from 10m" "144 runs/day" "$out" out="$("$PYTHON" "$COST" --pattern ci-watch --cadence 15m --model claude-sonnet-5 --json 2>/dev/null)" expect_has "cost json schema" "claude-mods.loop-ops.estimate/v1" "$out" expect_has "cost json carries runs_per_day" "runs_per_day" "$out" out="$("$PYTHON" "$COST" --list-models 2>/dev/null)"; rc=$? expect_exit "list-models -> 0" 0 "$rc" expect_has "list-models shows a model" "claude-opus-5" "$out" # cron cadence parses "$PYTHON" "$COST" --pattern daily-scan --cadence '*/10 * * * *' --model claude-haiku-4-5 >/dev/null 2>&1 expect_exit "cron cadence -> 0" 0 $? # --runs-per-day override out="$("$PYTHON" "$COST" --pattern custom --cadence weird --runs-per-day 5 --model claude-haiku-4-5 2>/dev/null)"; rc=$? expect_exit "runs-per-day override -> 0" 0 "$rc" expect_has "uses the override" "5 runs/day" "$out" # caching: a fast loop (10m -> 1h TTL) projects a cached saving out="$("$PYTHON" "$COST" --pattern ci-watch --cadence 10m --model claude-sonnet-5 2>&1)" expect_has "fast loop shows a cached projection" "cached/" "$out" # caching: a slow loop (6h > 1h TTL) is not cache-beneficial out="$("$PYTHON" "$COST" --pattern daily-scan --cadence 6h --model claude-opus-5 2>&1)" expect_has "slow loop: caching not beneficial" "not beneficial" "$out" # --no-cache suppresses the cached projection out="$("$PYTHON" "$COST" --pattern ci-watch --cadence 10m --model claude-sonnet-5 --no-cache 2>&1)" case "$out" in *"cached/"*) no "--no-cache still showed caching";; *) ok "--no-cache suppresses caching";; esac # json caching block present for a cacheable loop out="$("$PYTHON" "$COST" --pattern ci-watch --cadence 5m --model claude-sonnet-5 --json 2>/dev/null)" expect_has "cost json carries caching block" '"caching"' "$out" # ── loop-doctor: preflight (offline budget, live binary), json ───────────── echo "-- loop-doctor --" bash "$DOCTOR" --help >/dev/null 2>&1; expect_exit "loop-doctor --help -> 0" 0 $? bash "$DOCTOR" --offline "$SB/l1.yaml" >/dev/null 2>&1; expect_exit "doctor offline healthy L1 -> 0" 0 $? bash "$DOCTOR" --live "$SB/l1.yaml" >/dev/null 2>&1; expect_exit "doctor live healthy L1 -> 0" 0 $? # budget too small for the pattern -> bad -> 10 sed 's/^budget_tokens: 300000/budget_tokens: 100/' "$SB/l2.yaml" > "$SB/l2-poor.yaml" out="$(bash "$DOCTOR" --offline "$SB/l2-poor.yaml" 2>/dev/null)"; rc=$? expect_exit "doctor budget-too-small -> 10" 10 "$rc" expect_has "doctor names the budget gap" "tokens/run" "$out" # live: a verify gate whose binary is missing -> bad -> 10 sed 's/^verify: "npm test"/verify: "totally-missing-binary-zzz run"/' "$SB/l2.yaml" > "$SB/l2-nobin.yaml" bash "$DOCTOR" --live "$SB/l2-nobin.yaml" >/dev/null 2>&1; expect_exit "doctor missing gate binary -> 10" 10 $? # missing config -> 3, json schema bash "$DOCTOR" --offline "$SB/no-such.yaml" >/dev/null 2>&1; expect_exit "doctor missing config -> 3" 3 $? out="$(bash "$DOCTOR" --offline --json "$SB/l1.yaml" 2>/dev/null)" expect_has "doctor json schema" "claude-mods.loop-ops.doctor/v1" "$out" # ── loop-doctor: host-aware checks (native scheduling) ───────────────────────── # Each native host has different hard limits, so "will it run" depends on `host:`. # Facts asserted here are verified against the live tool schemas + docs (2026-08-30) # and written up in references/native-scheduling.md - if a limit changes upstream, # these are the assertions that should fail first. echo "-- loop-doctor (host-aware) --" # A config with NO host: must behave exactly as before (host defaults to local). grep -v '^host:' "$SB/l1.yaml" > "$SB/l1-nohost.yaml" 2>/dev/null || cp "$SB/l1.yaml" "$SB/l1-nohost.yaml" bash "$DOCTOR" --offline "$SB/l1-nohost.yaml" >/dev/null 2>&1; expect_exit "no host: still healthy (defaults local) -> 0" 0 $? # cloud-routine: minimum interval is 1 hour; faster expressions are rejected at creation. sed 's|^kill_switch:.*|kill_switch: "pause the routine (Repeats toggle)"|' "$SB/l1.yaml" > "$SB/cloud.yaml" cat >> "$SB/cloud.yaml" <<'EOF' host: cloud-routine EOF sed -i 's|^cadence: 10m|cadence: 10m|' "$SB/cloud.yaml" out="$(bash "$DOCTOR" --offline "$SB/cloud.yaml" 2>/dev/null)"; rc=$? expect_exit "cloud-routine sub-hour cadence -> 10" 10 "$rc" expect_has "names the 1-hour floor" "1 hour" "$out" # cloud-routine has NO permission mode, so the boundary must be named instead # (repos + environment network policy + connectors). Absent -> a finding. sed 's|^cadence: 10m|cadence: 1h|' "$SB/cloud.yaml" > "$SB/cloud-1h.yaml" out="$(bash "$DOCTOR" --offline "$SB/cloud-1h.yaml" 2>/dev/null)"; rc=$? expect_exit "cloud-routine without a named boundary -> 10" 10 "$rc" expect_has "names the missing boundary" "boundary" "$out" # With the boundary named it passes, and permission_mode is reported as ignored, not required. sed 's|^escalation:.*|escalation: "connectors pruned to GitHub read; environment network Trusted; never merge"|' \ "$SB/cloud-1h.yaml" > "$SB/cloud-ok.yaml" bash "$DOCTOR" --offline "$SB/cloud-ok.yaml" >/dev/null 2>&1; expect_exit "cloud-routine with boundary -> 0" 0 $? out="$(bash "$DOCTOR" --offline "$SB/cloud-ok.yaml" 2>/dev/null)" expect_has "permission_mode reported as ignored on cloud" "ignored by cloud routines" "$out" # --live must be SKIPPED (not passed) for a cloud routine: the tick runs on a fresh # cloud clone, so this machine's PATH proves nothing. A missing gate binary here must # NOT fail, and must not be silently reported as ok either. sed 's|^escalation:|verify: "totally-missing-binary-zzz run"\nescalation:|' "$SB/cloud-ok.yaml" > "$SB/cloud-live.yaml" out="$(bash "$DOCTOR" --live "$SB/cloud-live.yaml" 2>/dev/null)"; rc=$? expect_exit "cloud-routine --live does not fail on local PATH -> 0" 0 "$rc" expect_has "cloud-routine --live is explicitly skipped" "skipped" "$out" # session-cron (/loop + CronCreate) is session-scoped with a 7-day expiry: L1 only. sed 's|^host: cloud-routine|host: session-cron|' "$SB/cloud-ok.yaml" > "$SB/sess-l1.yaml" bash "$DOCTOR" --offline "$SB/sess-l1.yaml" >/dev/null 2>&1; expect_exit "session-cron at L1 -> 0 (warns only)" 0 $? out="$(bash "$DOCTOR" --offline "$SB/sess-l1.yaml" 2>/dev/null)" expect_has "session-cron L1 warns about the 7-day expiry" "7-day expiry" "$out" # same host at L2 = unattended, which it cannot do { sed 's|^tier: L1|tier: L2|' "$SB/sess-l1.yaml"; printf 'verify: "true"\nguard: "true"\nworktree: true\nland_via: fleet-ops\n'; } > "$SB/sess-l2.yaml" out="$(bash "$DOCTOR" --offline "$SB/sess-l2.yaml" 2>/dev/null)"; rc=$? expect_exit "session-cron at L2 (unattended) -> 10" 10 "$rc" expect_has "names the unattended mismatch" "can't run unattended" "$out" # An unknown host is a finding, not a silent pass. sed 's|^host: session-cron|host: made-up|' "$SB/sess-l1.yaml" > "$SB/host-bad.yaml" bash "$DOCTOR" --offline "$SB/host-bad.yaml" >/dev/null 2>&1; expect_exit "unknown host -> 10" 10 $? # ── loop-doctor: regressions found by adversarial review ─────────────────── # Each of these shipped broken once. They are cheap to re-break and silent when broken. echo "-- loop-doctor (adversarial regressions) --" # R1. A cron-string cadence must still hit the cloud floor - the Nm/Nh/Nd path is not # the only way to express "every 10 minutes". sed 's|^cadence: 1h|cadence: "*/10 * * * *"|' "$SB/cloud-ok.yaml" > "$SB/cloud-cron.yaml" out="$(bash "$DOCTOR" --offline "$SB/cloud-cron.yaml" 2>/dev/null)"; rc=$? expect_exit "cron-string cadence hits the cloud floor -> 10" 10 "$rc" # ...and an hourly cron must NOT be flagged (1h is legal on cloud). sed 's|^cadence: 1h|cadence: "0 * * * *"|' "$SB/cloud-ok.yaml" > "$SB/cloud-hourly.yaml" bash "$DOCTOR" --offline "$SB/cloud-hourly.yaml" >/dev/null 2>&1; expect_exit "hourly cron is legal on cloud -> 0" 0 $? # R2. A malformed cadence must not leak a bash arithmetic error. '1a2m' matches the # *[0-9]m glob, so without a digits-only guard ${1%m} hands '1a2' to $(( )). sed 's|^cadence: 1h|cadence: 1a2m|' "$SB/cloud-ok.yaml" > "$SB/cad-junk.yaml" allout="$(bash "$DOCTOR" --offline "$SB/cad-junk.yaml" 2>&1)" case "$allout" in *"value too great for base"*|*"syntax error"*) no "malformed cadence leaks an arithmetic error" ;; *) ok "malformed cadence degrades quietly (no arithmetic leak)" ;; esac # R3. L3 on a cloud routine must NOT demand a local container note - the cloud infra IS # the isolation, and its permission_mode is ignored. A false finding here teaches people # to write a bogus "container" note to satisfy the tool. { sed -e 's|^tier: L1|tier: L3|' -e 's|^permission_mode: dontAsk|permission_mode: bypassPermissions|' "$SB/cloud-ok.yaml" printf 'verify: "true"\nguard: "true"\nworktree: true\nland_via: fleet-ops\n'; } > "$SB/cloud-l3.yaml" out="$(bash "$DOCTOR" --offline "$SB/cloud-l3.yaml" 2>/dev/null)" case "$out" in *"bad"*"isolation"*) no "L3 cloud-routine wrongly demands a container note" ;; *) ok "L3 cloud-routine does not demand a local container" ;; esac # The same check must still bite for a LOCAL L3 bypass - the fix must not disarm it. { sed -e 's|^tier: L1|tier: L3|' -e 's|^host: cloud-routine|host: local|' \ -e 's|^permission_mode: dontAsk|permission_mode: bypassPermissions|' "$SB/cloud-ok.yaml" printf 'verify: "true"\nguard: "true"\nworktree: true\nland_via: fleet-ops\n'; } > "$SB/local-l3.yaml" out="$(bash "$DOCTOR" --offline "$SB/local-l3.yaml" 2>/dev/null)" expect_has "local L3 bypass still demands isolation" "isolated VM/container" "$out" # R4. `worktree: true` on a desktop-task is a declaration, not the switch: the real # toggle is per-task and OFF by default, so the doctor must say it cannot enforce it. { sed -e 's|^host: cloud-routine|host: desktop-task|' -e 's|^tier: L1|tier: L2|' "$SB/cloud-ok.yaml" printf 'verify: "true"\nguard: "true"\nworktree: true\nland_via: fleet-ops\n'; } > "$SB/desk-l2.yaml" out="$(bash "$DOCTOR" --offline "$SB/desk-l2.yaml" 2>/dev/null)" expect_has "desktop-task flags the out-of-band worktree toggle" "toggle is on" "$out" # ── loop-estimate: validation errors ─────────────────────────────────────────── "$PYTHON" "$COST" --pattern pr-watch --cadence 10m --model claude-nope >/dev/null 2>&1; expect_exit "unknown model -> 4" 4 $? "$PYTHON" "$COST" --pattern not-a-pattern --cadence 10m --model claude-haiku-4-5 >/dev/null 2>&1; expect_exit "unknown pattern -> 4" 4 $? "$PYTHON" "$COST" --pattern pr-watch --cadence "garbage cron" --model claude-haiku-4-5 >/dev/null 2>&1; expect_exit "bad cadence -> 4" 4 $? "$PYTHON" "$COST" --pricing "$SB/no-pricing.json" --pattern custom --cadence 1h --input-tokens 1 --output-tokens 1 --model x >/dev/null 2>&1; expect_exit "missing pricing file -> 3" 3 $? # ── check-pricing-sync: offline clean -> 0, drift -> 10, --json ──────────── echo "-- check-pricing-sync --" "$PYTHON" "$SYNC" --help >/dev/null 2>&1; expect_exit "pricing-sync --help -> 0" 0 $? "$PYTHON" "$SYNC" --offline >/dev/null 2>&1; expect_exit "pricing-sync offline in sync -> 0" 0 $? # Tamper a copy: opus input price 5.0 -> 999.0 (sed; argv path is MSYS-converted for python). sed 's/"input_per_mtok": 5\.0/"input_per_mtok": 999.0/' "$SKILL/assets/model-pricing.json" > "$SB/badprice.json" "$PYTHON" "$SYNC" --pricing "$SB/badprice.json" >/dev/null 2>&1; expect_exit "pricing-sync drift -> 10" 10 $? "$PYTHON" "$SYNC" --pricing "$SB/no-such.json" >/dev/null 2>&1; expect_exit "pricing-sync missing file -> 3" 3 $? out="$("$PYTHON" "$SYNC" --json 2>/dev/null)" expect_has "pricing-sync json schema" "claude-mods.loop-ops.pricing-sync/v1" "$out" expect_has "pricing-sync json in_sync" '"in_sync": true' "$out" # ── Windows-authored configs: CRLF + UTF-8 BOM must parse like clean LF ───── echo "-- windows-authored configs (CRLF / BOM) --" good_l1 "$SB/win.yaml" sed 's/$/\r/' "$SB/win.yaml" > "$SB/win-crlf.yaml" # LF -> CRLF bash "$AUDIT" "$SB/win-crlf.yaml" >/dev/null 2>&1; expect_exit "CRLF config audits clean -> 0" 0 $? bash "$DOCTOR" --offline "$SB/win-crlf.yaml" >/dev/null 2>&1; expect_exit "CRLF config doctors clean -> 0" 0 $? printf '\xEF\xBB\xBF' > "$SB/win-bom.yaml"; cat "$SB/win.yaml" >> "$SB/win-bom.yaml" # prepend BOM bash "$AUDIT" "$SB/win-bom.yaml" >/dev/null 2>&1; expect_exit "BOM config audits clean -> 0" 0 $? bash "$DOCTOR" --offline "$SB/win-bom.yaml" >/dev/null 2>&1; expect_exit "BOM config doctors clean -> 0" 0 $? # ── worked example: the shipped example stays gate-clean ─────────────────── echo "-- worked example --" EX="$SKILL/assets/examples/pr-watch/loop.config.yaml" [[ -f "$EX" ]] && ok "worked example present" || no "worked example missing" bash "$AUDIT" "$EX" >/dev/null 2>&1; expect_exit "shipped example audits clean -> 0" 0 $? bash "$DOCTOR" --offline "$EX" >/dev/null 2>&1; expect_exit "shipped example doctors clean -> 0" 0 $? [[ -f "$SKILL/assets/examples/pr-watch/loop-run.sh" ]] && ok "example ships loop-run.sh (runner-agnostic)" || no "example missing loop-run.sh" [[ -f "$SKILL/assets/examples/pr-watch/github-actions.yml" ]] && ok "example ships an optional GH Actions scheduler" || no "example missing GH Actions option" [[ -f "$SKILL/assets/examples/pr-watch/run.md" ]] && ok "example ships a run prompt" || no "example missing run.md" # The flagship example is what people copy, so it must model the doctrine it teaches: # a loop declares where it runs before it is scheduled. expect_has "example declares a host" "host:" "$(cat "$EX")" # ── terminal design system ───────────────────────────────────────────────── echo "-- terminal design system --" for s in "$INIT" "$AUDIT" "$DOCTOR"; do b="$(basename "$s")" grep -q '_lib/term.sh' "$s" && ok "$b sources _lib/term.sh" || no "$b does not source _lib/term.sh" done grep -q 'class Term' "$COST" && ok "loop-estimate carries inline Term helper" || no "loop-estimate missing inline Term helper" grep -q 'class Term' "$SYNC" && ok "check-pricing-sync carries inline Term helper" || no "check-pricing-sync missing inline Term helper" grep -q 'BRAND::loop' "$SKILL/../_lib/term.sh" && ok "term.sh registers the loop brand glyph" || no "term.sh missing loop brand glyph" # Piped audit findings stay plain (no ANSI in the data stream). po="$(bash "$AUDIT" "$SB/l2-nogate.yaml" 2>/dev/null)" case "$po" in *$'\033'*) no "piped audit leaked ANSI into data";; *) ok "piped audit stays plain data";; esac # ── check-native-facts: the native-scheduling staleness guard ────────────── # Asserts the guard actually detects drift, not just that it exits 0 today: a # verifier that can only pass is decoration. echo "-- check-native-facts --" NATIVE_CHK="$SCRIPTS/check-native-facts.py" "$PYTHON" "$NATIVE_CHK" --help >/dev/null 2>&1; expect_exit "native-facts --help -> 0" 0 $? "$PYTHON" "$NATIVE_CHK" --offline >/dev/null 2>&1; expect_exit "native-facts offline in sync -> 0" 0 $? "$PYTHON" "$NATIVE_CHK" --offline --live >/dev/null 2>&1; expect_exit "mutually exclusive modes -> 2" 2 $? "$PYTHON" "$NATIVE_CHK" --offline --skill "$SB/not-a-skill" >/dev/null 2>&1; expect_exit "missing skill dir -> 3" 3 $? out="$("$PYTHON" "$NATIVE_CHK" --offline --json 2>/dev/null)" expect_has "native-facts json schema" "claude-mods.loop-ops.native-facts/v1" "$out" expect_has "native-facts json in_sync" '"in_sync": true' "$out" # Drift detection: copy the skill, drop a host from the reference -> must be caught. mkdir -p "$SB/drift" cp -r "$SKILL/assets" "$SKILL/scripts" "$SKILL/references" "$SB/drift/" 2>/dev/null "$PYTHON" "$NATIVE_CHK" --offline --skill "$SB/drift" >/dev/null 2>&1 expect_exit "unmodified copy still in sync -> 0" 0 $? sed 's|`cloud-routine`|`clown-routine`|g' "$SKILL/references/native-scheduling.md" > "$SB/drift/references/native-scheduling.md" out="$("$PYTHON" "$NATIVE_CHK" --offline --skill "$SB/drift" 2>/dev/null)"; rc=$? expect_exit "host vocabulary drift -> 10" 10 "$rc" expect_has "names the drifted source" "native-scheduling.md" "$out" # A reference that loses its date stamp is drift too - undated is how a doc rots. sed 's|\*\*Verified 2026-08-30\*\*|Verified recently|' "$SKILL/references/native-scheduling.md" > "$SB/drift/references/native-scheduling.md" "$PYTHON" "$NATIVE_CHK" --offline --skill "$SB/drift" >/dev/null 2>&1; expect_exit "missing date stamp -> 10" 10 $? # ── docs: the native-scheduling reference must exist AND be cited ────────── # A reference SKILL.md never links is dead weight the router can't find # (docs/SKILL-CREATION-PROTOCOL.md step 4), so both halves are asserted. echo "-- native-scheduling reference --" NATIVE="$SKILL/references/native-scheduling.md" [[ -f "$NATIVE" ]] && ok "native-scheduling.md present" || no "native-scheduling.md missing" skillmd="$(cat "$SKILL/SKILL.md")" expect_has "SKILL.md cites native-scheduling.md" "references/native-scheduling.md" "$skillmd" expect_has "claude-code-loops.md cites native-scheduling.md" "native-scheduling.md" "$(cat "$SKILL/references/claude-code-loops.md")" # The reference is a claim about a fast-moving external surface, so it must carry the # date it was verified - an undated table is how a scheduling doc rots invisibly. expect_has "native-scheduling.md is date-stamped" "Verified 2026-08-30" "$(cat "$NATIVE")" # Each host named by the config template must be documented in the reference. for h in session-cron desktop-task cloud-routine; do expect_has "reference documents host '$h'" "$h" "$(cat "$NATIVE")" done # The gate is an eval; loop-ops links the discipline rather than restating it. expect_has "SKILL.md routes gate-judgement work to evals-ops" "evals-ops" "$skillmd" # A skill that repositions itself against a native primitive must say when NOT to use # itself, or it just re-sells ceremony for work the harness already does. expect_has "SKILL.md says when the native primitive is enough" "native primitive is enough" "$skillmd" # FRONTMATTER CONTRACT (stated here because a later description-trim lane edits # frontmatter without reading this suite - see SKILL-CREATION-PROTOCOL.md step 5): # 1. `description` must keep the repositioning clause "Native primitives schedule; # loop-ops governs" - it is what stops a reader reaching for this skill to build a # scheduler the harness already ships. Trimming it changes what the skill IS. # 2. `description` must keep the native trigger words, or the router never fires on # "cron", "scheduled task" or "cloud routine". # 3. Nothing here requires `when_to_use`; this skill does not use that field. fm="$(grep -m1 '^description:' "$SKILL/SKILL.md")" expect_has "description keeps the repositioning clause" "Native primitives schedule; loop-ops governs" "$fm" expect_has "description keeps the native scheduling triggers" "cloud routine" "$fm" # Progressive disclosure: the body stays under the 500-line cap (protocol step 3). sk_lines="$(wc -l < "$SKILL/SKILL.md" | tr -d ' ')" [[ "$sk_lines" -lt 500 ]] && ok "SKILL.md under the 500-line cap ($sk_lines)" || no "SKILL.md is $sk_lines lines (cap 500)" # ── summary ──────────────────────────────────────────────────────────────── echo "=== $PASS passed, $FAIL failed ===" [[ "$FAIL" -eq 0 ]] || exit 1
-
-
SKILL.md 30.5 KB
--- name: loop-ops description: "Design and safely run OUTER loops - scheduled discover-triage-implement-verify-escalate agent loops. Native primitives schedule; loop-ops governs. Risk-tier ladder (L1 report -> L3 unattended), STATE/run-log/budget spine, kill switch, pattern catalog. Triggers: outer loop, scheduled/autonomous agent loop, PR watch, CI watch, dep-bump loop, run on a schedule, kill switch, risk tier, CronCreate, scheduled task, cloud routine, /loop." license: MIT allowed-tools: "Read Write Edit Bash Glob Grep" metadata: author: claude-mods related-skills: "iterate, fleet-ops, fleet-worker, pigeon, git-ops, ci-cd-ops" --- # Loop Ops — outer-loop design discipline **A loop is not a prompt.** Turn-by-turn prompting puts you in the loop forever. *Loop engineering* inverts it: you design a **recurring process with memory, verification, and boundaries** that discovers work, hands it to agents, verifies the result, and decides — on a schedule or until a goal is met — whether to **land it or escalate to a human**. > "You shouldn't be prompting coding agents anymore. You should be designing the loops > that prompt your agents." — Peter Steinberger This skill is the **outer loop**: the orchestration layer *above* a single agent run. It is the twin of [`iterate`](../iterate/SKILL.md) — `iterate` is the *inner* loop (one metric, one session, git-as-memory); `loop-ops` is the design discipline for the loop that *schedules and gates* inner runs. It does not reimplement spawning or landing; it **composes** what this repo already ships. ## Native primitives schedule; loop-ops governs Claude Code now ships the *cadence* half natively. **Do not hand-roll a scheduler** — pick a native host, declare it as `host:` in the config, and spend the discipline where the primitives leave a hole. Verified surface, parameters and limits (2026-08-30): [references/native-scheduling.md](references/native-scheduling.md). | Native primitive | What it gives you | What it does NOT give you | |---|---|---| | **`/loop`** — a *bundled* skill ([docs](https://code.claude.com/docs/en/scheduled-tasks)), driving `CronCreate`/`CronList`/`CronDelete`; `ScheduleWakeup` for its self-paced mode | Fixed-cron or Claude-paced ticks (delay clamped 60 s–1 h), a built-in maintenance prompt, `.claude/loop.md` to override it, `Esc` to stop | **Session-scoped and in-memory** — fires only while the session is idle, dies with the conversation, and every recurring job **self-deletes after 7 days**. No state spine, no budget, no gate. L1-supervised only. | | **Desktop scheduled tasks** — the `scheduled-tasks` MCP server ([docs](https://code.claude.com/docs/en/desktop-scheduled-tasks)) | Durable local ticks (≥1 min) with local files, a **fresh session per run**, a per-task permission mode with saved approvals, a task folder, run history, an Active/Paused toggle | The worktree toggle is **off by default** (runs against uncommitted changes); one catch-up only for a missed window; a Manual-mode task **stalls** on an unapproved tool. No STATE spine, no token budget, no verify gate. | | **Cloud routines** — `/schedule` ([docs](https://code.claude.com/docs/en/routines)) | Machine-off ticks (≥1 h), plus **native event triggers**: an API `/fire` endpoint and GitHub `pull_request`/`release` events with filters. A real push guard on non-`claude/` branches | **No permission mode at all** and **every connector attaches by default**; no local files (fresh clone); green run status ≠ task success. The boundary must come from repos + environment + connectors. | | **`/goal`** ([docs](https://code.claude.com/docs/en/goal)) | A native *completion* gate — keep going until a fast model confirms the condition | Not a cadence, and not an audit trail. | What **none** of them provide — and what this skill is for: a **state spine** that survives ticks, a **token budget**, a **verify gate** you can trust, an **escalation rule**, and the **risk-tier ladder** that decides whether the loop has earned the autonomy you're about to grant it. The plumbing moved into the harness; the judgement did not. ### When the native primitive is enough — stop here Don't scaffold a loop for work the harness already does. **Use the primitive raw** when *all* of these hold: - it **writes nothing** you'd have to undo (watch a deploy, poll a build, remind you), or the only writes are ones you'll review anyway; - it is **supervised or short-lived** — you're watching, or it stops in a session; - **nothing needs to be remembered between ticks** beyond what's in the repo; - and you'd **shrug if a tick silently didn't fire**. `/loop 5m check if the deploy finished` is a complete, correct answer. Wrapping it in a `loop.config.yaml` adds ceremony and no safety. Reach for loop-ops the moment **any one** of those flips: the loop starts *changing* things, runs *unattended*, needs to know what it did *last time*, or its silence would cost you. That is the whole trigger — everything below is what to do once it fires. --- ## The six primitives → what owns each here Every durable loop rests on six primitives. The discipline is wiring them; the parts already exist: | Primitive | What it is | Owned in claude-mods by | |---|---|---| | **Schedule** | fire the loop on a cadence *or an event* | native-first, declared as `host:` — `session-cron` (`/loop`+`CronCreate`, L1 only), `desktop-task` (`scheduled-tasks` MCP: local + durable), `cloud-routine` (`/schedule`: machine-off, plus API/GitHub event triggers), `/goal` for completion. `external` (cron/Task Scheduler + `loop-run.sh`) only for non-Claude-Code control | | **Worktree** | isolated, discardable execution context | `git-ops` worktrees, `fleet-worker` (per-task worktree) | | **Skills** | persistent project knowledge the run loads | this repo's skill layer + your `CLAUDE.md` | | **Sub-agents** | maker/checker separation | `Agent`/`Task`; dispatching skills (`review`, `testgen`) | | **Connectors** | reach tickets / CI / chat | MCP tools, `gh`, `github-ops` | | **+ State** | a durable spine *outside* the conversation | `STATE.md` + run-log + budget (this skill) | The inner improvement loop is `iterate`; cheap parallel makers are `fleet-worker`; the test-gated merge queue is `fleet-ops`; inter-loop signalling is `pigeon`. `loop-ops` is the doctrine that connects them. ## Loop anatomy ``` ┌──────────────────────────────────────────────────────────────┐ │ SCHEDULE (cadence) │ │ └─▶ TRIAGE read STATE.md → pick the next unit of work │ │ └─▶ WORKTREE isolate (git worktree) │ │ └─▶ MAKER implementer run (or fleet-worker)│ │ └─▶ CHECKER verify gate + guard (tests) │ │ └─▶ GATE safe & allowlisted? │ │ ├─ yes → LAND (commit/PR) │ │ └─ no → ESCALATE (+context) │ │ └─▶ write STATE.md, append run-log, decrement budget ──────┘ ``` The **gate** is the load-bearing decision. Everything before it is mechanical; the gate is where a loop earns the right to run unattended — or doesn't. ## The risk-tier ladder (the heart of the discipline) Never start a loop unattended. Graduate it. Each tier maps to a concrete Claude Code **permission mode** — full mapping, the headless-profile table, and the *enumerate vs isolate* fork in [references/risk-tiers.md](references/risk-tiers.md). | Tier | Posture | Permission mode | May do | Lands by | |---|---|---|---|---| | **L1 Report** | read-only discovery + triage | `plan` / `dontAsk`+read allowlist | scan, summarize, propose — **writes nothing** | a human reads the report | | **L2 Assisted** | suggest changes, human gates the merge | `dontAsk`+narrow allowlist, or `auto` | edit in a **worktree**, run tests, open a PR | a human approves the PR (or `fleet-ops`) | | **L3 Unattended** | autonomous land within a denylist | `bypassPermissions` **in an isolated container only** | commit/merge allowlisted classes | the loop itself, inside its boundary | **The host is part of the tier.** `session-cron` (`/loop` + `CronCreate`) cannot host L2+ at all: it needs an open idle session and every recurring job expires after 7 days. `cloud-routine` has *no permission mode*, so its tier is expressed as repos + environment network policy + connectors instead — which means a routine is effectively autonomous the moment it is created, and the L1 posture has to come from the prompt being read-only. `loop-doctor` enforces both. Details: [references/native-scheduling.md](references/native-scheduling.md). The cardinal rule, straight from Claude Code's own gate model: **an unattended loop is a *scheduler/script that invokes `claude -p`*, not a Claude session that spawns ungated children.** A session in `auto` mode that tries to launch a `--permission-mode bypassPermissions` child is blocked as *Create Unsafe Agents* — by design. See [references/risk-tiers.md](references/risk-tiers.md) and the repo's [auto-mode-classifier reference](../../docs/AUTO-MODE-CLASSIFIER.md). ## The escalation gate What a loop may **land** vs what it must **escalate** is not a vibe — it mirrors Claude Code's classifier tiers. Bake these into the config's `escalation:` field: - **Always escalate (never auto-land):** force-push, push to `main`, production deploys or migrations, mass deletion, granting IAM/repo permissions, anything destroying pre-session files, editing `.claude/`/settings (self-modification), `curl | bash`. - **Safe to auto-land at L2/L3 (when allowlisted):** a green PR on a feature branch, a lockfile patch bump that passes the guard, a generated changelog draft, a label/ triage classification, a comment. - **The test:** *would a careful human let this happen unattended in this repo?* If the action's blast radius exceeds the loop's stated purpose, it escalates. A general goal ("keep CI green") is **not** authorization for a specific high-blast action it implies. - **Scope the tools, not just the mode.** Allowlist exactly the tools/MCP connectors the job needs (read-only at L1); keep `gh pr merge` out and `land_via: fleet-ops` in. Full connector/MCP-scope discipline + the auto-merge guard: [references/risk-tiers.md](references/risk-tiers.md). On a **cloud routine this is the whole gate** — there is no permission mode, and every connected connector is attached by default with full write access. Prune them. - **A task that reschedules itself is self-modification.** The `scheduled-tasks` MCP lets a running task call `update_scheduled_task` on its own schedule or prompt. Useful, and on the always-escalate list unless adaptive cadence is the loop's *stated* purpose — a loop that can rewrite its own trigger has left the boundary you audited. ## The state spine A loop's memory lives **outside** the conversation, in three files (schemas + read/write contract in [references/state-spine.md](references/state-spine.md)): - **`STATE.md`** — the triage snapshot: priority / watch / noise + a readiness line. Read at the top of every run, rewritten at the end. - **`run-log.md`** — append one line per run (timestamp, action, outcome, tokens). The audit trail that answers "what has this loop been doing?" - **`loop.config.yaml`** — the loop's definition (goal, tier, cadence, **host**, scope, gate, budget, escalation). Scaffolded by `loop-scaffold`, scored by `loop-check`. A native host gives you a *place* for this spine (a Desktop task's folder) but never the spine itself: no host writes `STATE.md`, enforces a token budget, or records what the loop *decided*. Run history says a tick happened; the run-log says what it did and cost. ## Pattern catalog (a morphology, not a fixed list) Patterns are **compositions of three axes** — `trigger` (cadence / **event** via a Channel / `goal`) × `posture` (L1/L2/L3) × `locus` (connector→cloud routine / local→Desktop task). The named patterns are well-trodden points in that space; compose your own from the axes. Full recipes + the morphology in [references/pattern-catalog.md](references/pattern-catalog.md): | Pattern | Trigger · Locus | Tier | One-line job | |---|---|---|---| | `daily-scan` | cadence · local | L1 | discover + prioritize, report only | | `pr-watch` | event\|cadence · connector | L1 | watch review state, surface stuck PRs | | `ci-watch` | **event** · local | L2 | triage build failures, propose a fix | | `dep-bump` | cadence · local | L2 | patch-only bumps behind cooldown + guard | | `changelog-gen` | event(tag)\|cadence · local | L1 | draft release notes for approval | | `merge-hygiene` | cadence · local | L1 | dead branches, stale flags | | `issue-sort` | cadence · connector | L1 | classify + label, propose only | | `metric-chase` | **goal** · local | L2 | drive a metric (coverage/latency/eval) via `iterate` | | `regression-watch` | cadence\|event · local | L1 | run a benchmark/eval, flag a regression | | `digest` | cadence · **connector** | L1 | summarize email/Asana/news (cloud routine) | | `backfill` | **goal** · local | L2 | drain a migration/queue **to completion** | | `monitor` | **event** · local | L1 | error/deploy webhook → triage + page | | `freshness` | cadence · local | L1 | re-check docs/data/deps vs reality | Start any pattern at L1. Graduate to L2 only after the L1 reports prove its judgment. **Prefer `event` over `cadence`** where a webhook exists (cheaper, faster than polling). ## Multi-loop coordination & the kill switch Running several loops? Two non-negotiables (detail in [references/state-spine.md](references/state-spine.md)): - **Priority order** prevents collisions: `CI Watch → PR Watch → Dependency Bump → Merge-Hygiene/Changelog → Daily Scan (off-peak)`. A higher-priority loop's worktree wins; lowers defer. Loops signal each other via [`pigeon`](../pigeon/SKILL.md). - **A kill switch every loop honors.** A single stop signal — a `PAUSED` sentinel file or a `loop-pause` label — that every loop checks at the top of its run and exits on. No loop ships without one. Put it in `kill_switch:` and check it first. ## Composition map — don't rebuild what exists | You need to… | Use | Not | |---|---|---| | improve one metric in one session | [`iterate`](../iterate/SKILL.md) | a hand-rolled inner loop | | spawn cheap parallel makers | [`fleet-worker`](../fleet-worker/SKILL.md) | bespoke `claude -p` plumbing | | route models across a fan-out (cheap finders, Opus judges) | [`fleet-worker` model-routing](../fleet-worker/references/model-routing.md) | every agent on the orchestrator's model | | test-gate + land winning branches | [`fleet-ops`](../fleet-ops/SKILL.md) | a manual merge step | | fire on a cadence or an event | a native `host:` — `/loop`, Desktop scheduled task, cloud routine (schedule/API/GitHub triggers); `/goal` for completion | a custom cron in this skill | | trust the `verify` gate's judgement | the **`evals-ops`** skill — a gate is an eval (golden set, judge bias, `pass^k`, blocking vs advisory) | eyeballing a few runs and calling it proven | | reason about per-tick prompt-cache cost | [`claude-api-ops` caching-and-cost](../claude-api-ops/references/caching-and-cost.md) | a TTL number memorised from a blog post | | commit / PR / release | [`git-ops`](../git-ops/SKILL.md), [`github-ops`](../github-ops/SKILL.md) | raw `git push` | | signal between loops | [`pigeon`](../pigeon/SKILL.md) | a shared scratch file | `loop-ops` is the **design layer**; these are the **execution layers**. --- ## Tools Six scripts, all following the [Skill Resource Protocol](../../docs/SKILL-RESOURCE-PROTOCOL.md) (stdout = data, semantic exit codes, `--help` with EXAMPLES, `--json` envelopes): **init** scaffolds the loop, **audit** scores whether the config is *well-formed*, **doctor** preflights whether it will actually *run* (host-aware), **cost** estimates spend (caching-aware), and two drift guards — **check-pricing-sync** for the pricing table and **check-native-facts** for the native-scheduling limits. The discipline before scheduling is `init → fill → cost → audit → doctor --live`. ### `scripts/loop-scaffold.sh` — scaffold a loop's state spine Writes `<dir>/<name>/` with five files from the bundled templates: `loop.config.yaml` ([assets/loop.config.template.yaml](assets/loop.config.template.yaml)), `STATE.md` ([assets/STATE.template.md](assets/STATE.template.md)), `run-log.md`, `run.md` (the headless run prompt, [assets/run.template.md](assets/run.template.md)), and an executable **`loop-run.sh`** ([assets/run.sh.template](assets/run.sh.template)) — the runner-agnostic tick wrapper any scheduler invokes (cron / Windows Task Scheduler / systemd / by hand), **no GitHub Actions required**. Pass a known `--pattern` (pr-watch, ci-watch, dep-bump, …) and the config is **seeded** with that pattern's scope/goal/escalation — and, at L2+, its gate — so you get a near-ready config to review, not blank placeholders (it audits clean immediately). Doctrine holds: it still scaffolds at L1 by default with a graduation block. `--host` records where ticks will execute (`local` default, or `session-cron` / `desktop-task` / `cloud-routine` / `external`) so `loop-doctor` checks that host's real constraints instead of assuming a local `claude -p`. ```bash # Create .loops/pr-watch/ with config + STATE.md + run-log.md + run.md from templates: bash scripts/loop-scaffold.sh --name pr-watch --pattern pr-watch --tier L1 # A connector-driven loop bound for a cloud routine (>=1h floor, no permission mode): bash scripts/loop-scaffold.sh --name digest --pattern digest --host cloud-routine --cadence 1h # Custom dir + cadence, preview without writing: bash scripts/loop-scaffold.sh --name dep-bump --pattern dep-bump \ --tier L2 --cadence 1d --dir .loops --dry-run ``` Refuses to overwrite a populated `<dir>/<name>/` (exit 5) unless `--force`. Atomic writes. `--dry-run` prints what it would create and writes nothing. stdout = the created config path. ### `scripts/loop-check.sh` — readiness scorer (run before you schedule) The question this answers: *is this loop safe to turn on at its declared tier?* It scores a `loop.config.yaml` against the readiness rubric — gate present, scope bounded, escalation defined, guard + worktree at L2+, budget + kill switch set, permission mode consistent with tier — and refuses a green light if any **critical** gap exists. ```bash bash scripts/loop-check.sh .loops/pr-watch/loop.config.yaml # exit 0 ready, 10 not ready bash scripts/loop-check.sh --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.severity=="error")' bash scripts/loop-check.sh --min 80 .loops/ci-watch/loop.config.yaml # raise the score bar ``` Exit **0** = ready (no errors, score ≥ `--min`), **10** = not ready (findings on stdout), `2` usage, `3` config not found, `4` config unparseable. `--strict` counts warnings toward the not-ready signal. ### `scripts/loop-doctor.sh` — live preflight (will it actually run?) `loop-check` proves the config is *well-formed*; `loop-doctor` proves the loop will *execute* — catching the "blocked at 3am" failures audit can't see. `--offline` (CI-safe): the budget fits a tick's estimated tokens, the permission mode is achievable (not interactive), an L3 bypass declares an isolation boundary. `--live` adds runtime preflight: the `verify`/`guard` gate's leading binary resolves on PATH, `claude`/`git` are present, the kill-switch sentinel's parent dir exists. **It is host-aware.** `host:` changes what "will it run" even means, so the doctor checks against the declared surface: a `cloud-routine` faster than its 1-hour floor is rejected at creation; a routine with no named repos/environment/connector boundary has no gate at all (it has no permission mode either, so demanding one there would be a false finding); a `session-cron` host at L2+ can't run unattended and is called a predicted failure; and `--live` is **skipped, not passed**, for a cloud routine — this machine's PATH says nothing about a fresh cloud clone, and a green check there would be false confidence. ```bash bash scripts/loop-doctor.sh --offline .loops/pr-watch/loop.config.yaml # CI gate bash scripts/loop-doctor.sh --live .loops/ci-watch/loop.config.yaml # before scheduling bash scripts/loop-doctor.sh --live --json .loops/dep-bump/loop.config.yaml | jq '.data[] | select(.state=="bad")' bash scripts/loop-doctor.sh --offline .loops/digest/loop.config.yaml # host: cloud-routine -> floor + boundary ``` Exit **0** = will run, **10** = a check predicts a runtime failure (gate binary missing, bypass on host without isolation, budget too small for a tick), `2` usage, `3` not found, `4` unparseable, `5` missing core dep. Run it **after** `loop-check` and before scheduling. ### `scripts/loop-estimate.py` — token/$ estimate by pattern × cadence × model (caching-aware) Estimate spend **before** committing to a cadence — the cost of an outer loop is runs/day × tokens/run × price, and sub-agents multiply it. It also models **prompt caching**: a loop re-sends the same `run.md`+system prefix every tick (the Ralph property), so the prefix should be cache-written once then read (~0.1×) — *but only if the tick interval fits the cache TTL*. **The TTL is a choice, not a constant:** 5 minutes by default (1.25× write) or 1 hour with `"ttl": "1h"` (2× write), so the daemon window is ~4.5 min *or* ~55 min — not a fixed 270 s. The estimator picks the cheapest TTL that stays warm at your cadence and names it; past 1 h nothing caches at all. Mechanics and break-even: [`claude-api-ops` caching-and-cost](../claude-api-ops/references/caching-and-cost.md). The estimate itself is **host-agnostic** — tokens are tokens wherever the tick fires; the host-dependent limit is the *minimum cadence*, which `loop-doctor` enforces. Pricing reads from `assets/model-pricing.json` (date-stamped; [`claude-api-ops`](../claude-api-ops/SKILL.md) is the source of truth — run its `check-model-table.py` if you suspect drift). ```bash python scripts/loop-estimate.py --pattern pr-watch --cadence 10m --model claude-haiku-4-5 python scripts/loop-estimate.py --pattern ci-watch --cadence 15m --model claude-sonnet-5 --days 30 --json python scripts/loop-estimate.py --list-models # the pricing table + its as-of date ``` Exit `0` ok, `2` usage, `3` pricing file missing, `4` bad cadence/model. Output names every assumption (runs/day, tokens/run, sub-agent multiplier) — it's an estimate, and it says so. ### `scripts/check-pricing-sync.py` — offline drift guard (CI) `model-pricing.json` is a *copy* of claude-api-ops's authoritative model table, and a copy drifts silently. This offline verifier asserts every model in [assets/model-pricing.json](assets/model-pricing.json) matches claude-api-ops's "Current Models" table (prices included). Both files are in-repo, so it's network-free and gates PR CI via `tests/check-resources.sh`; live model-id drift is owned by claude-api-ops's `check-model-table.py`. ```bash python scripts/check-pricing-sync.py --offline # exit 0 in sync, 10 drift, 3 a file missing ``` ### `scripts/check-native-facts.py` — native-scheduling staleness guard [references/native-scheduling.md](references/native-scheduling.md) encodes a **fast-moving external surface**, and `loop-doctor` refuses configs on those numbers — so a silently stale limit becomes a wrong refusal. `--offline` (PR CI) proves internal consistency: the host vocabulary is *one* set across the config template, `loop-scaffold --host`, `loop-doctor`'s case arm and the reference; the reference still carries its `Verified <date>` stamp; and every limit the doctor enforces is still stated in the prose that justifies it. `--live` (scheduled, never a PR gate) fetches the three published docs pages and checks our numbers still appear in them. ```bash python scripts/check-native-facts.py --offline # exit 0 in sync, 10 drift, 3 file missing python scripts/check-native-facts.py --live # exit 7 = docs unreachable (advisory) ``` --- ## End-to-end workflow 1. **Pick a pattern** from the catalog (or `custom`), and **pick the host** — does the tick need local files? must it run with the machine off? is it supervised? Start at **L1**. 2. **Scaffold:** `bash scripts/loop-scaffold.sh --name <n> --pattern <p> --tier L1 --host <h>`. 3. **Fill `loop.config.yaml`** — the real `goal`, `scope` (bounded globs, never `*`), `verify` gate, `escalation` rule, `budget_tokens`, `kill_switch`. On a `cloud-routine`, name the boundary that replaces the absent permission mode: repos, environment network policy, and the connectors you kept. 4. **Cost it:** `python scripts/loop-estimate.py --pattern <p> --cadence <c> --model <m>` — sanity-check the monthly spend against the value. 5. **Audit it:** `bash scripts/loop-check.sh .loops/<n>/loop.config.yaml` — fix every error before scheduling. Don't schedule a loop that fails its own audit. 6. **Doctor it:** `bash scripts/loop-doctor.sh --live .loops/<n>/loop.config.yaml` — prove it will actually *run* (gate binary on PATH, budget fits a tick). Audit = well-formed; doctor = will-run. 7. **Schedule** the L1 run on the declared host — the **recipe selector** in [references/claude-code-loops.md](references/claude-code-loops.md) prescribes which, because they're not interchangeable: connector-driven (email/Asana, no local code) → **cloud routine**; touches local code → **Desktop scheduled task**; sustained & token-sensitive → a **cache-warm daemon** (`claude -p` inside the cache TTL you paid for), *not* `/loop` (which grows a session and chews tokens); fixed-criteria long task → **`/goal`**; quick supervised polling → `/loop`. Per-primitive limits: [references/native-scheduling.md](references/native-scheduling.md). (L1 is read-only — it just writes `STATE.md` + a report.) 8. **Read the reports.** Only after the loop's judgment is proven do you graduate it to **L2** (worktree + guard + `fleet-ops` landing), change `host:` if the proving host was `session-cron`, and re-audit at the higher tier. If the gate's verdict is a judgement rather than a green test run, harden it with the **`evals-ops`** discipline before you let it decide unattended. ## Worked example A complete, **audit + doctor-clean** L1 loop ships at [assets/examples/pr-watch/](assets/examples/pr-watch/): a filled `loop.config.yaml`, a *populated* `STATE.md`, the `run.md` run prompt, a sample `run-log.md`, the runner-agnostic **`loop-run.sh`** (the tick wrapper, with the kill-switch gate and `dontAsk` + allowlist baked in — point cron / Task Scheduler at it), and an *optional* `github-actions.yml` for repos already on GitHub. Copy the dir, adjust scope/cadence, run `loop-check` + `loop-doctor --live`, then wire `loop-run.sh` to your scheduler. The other patterns don't ship as static dirs that rot — `loop-scaffold --pattern <name>` *generates* the same, seeded and gate-clean, for any pattern at any tier. CI runs `loop-check` + `loop-doctor` on this example every build, so it can't drift out of validity. ## Anti-patterns (these are detected and wrong) The incident-shaped catalog — symptom → mechanism → the control that catches each — is [references/failure-modes.md](references/failure-modes.md) (runaway budget, the 3am-dead loop, cache-cold, force-push, ungated-child spawn, colliding loops, silent-stop, gate reward-hacking, and the native-host trio — the **expired** 7-day loop, the **over-connected** routine, the **stalled/skipped** Desktop task). The headline ones: - **Routing around the gate.** Wrapping `claude -p --permission-mode bypassPermissions` in a script to dodge the classifier is *Auto-Mode Bypass* — a `hard_deny` nothing clears. If an outcome is blocked, **authorize it** (a narrow allow rule, or run the scheduler outside the auto-mode session), never **disguise it**. - **The orchestrator session spawning ungated children.** A session in `auto` mode is the wrong place to launch the loop. The scheduler/cron/Task-Scheduler/CI runner that invokes `claude -p` is the authorizer. See [references/risk-tiers.md](references/risk-tiers.md) §"enumerate vs isolate". - **No gate.** A loop whose `verify:` is empty is not a loop, it's an unsupervised typer. `loop-check` errors on it. Nor is a *green run status* a gate: on a cloud routine green means "the session started and exited without an infrastructure error", never that the task succeeded. Grade the work, not the process. - **Assuming the native host gave you a boundary.** It gave you a cadence. A Desktop task's worktree toggle is off by default; a cloud routine has no permission mode and attaches every connector; `/loop` evaporates after 7 days. Each is a default that reads as safe and isn't. - **Unbounded scope.** `scope: "*"` means "may touch anything" — the audit rejects it. - **No kill switch / no budget.** A loop you can't stop, or whose spend you didn't bound, will eventually surprise you. Both are audit findings. - **Skipping L1.** Starting a fresh loop at L3 is how comprehension debt and incidents compound. The ladder exists precisely so trust is *earned* before it's *granted*. ## See also - [references/risk-tiers.md](references/risk-tiers.md) — L1/L2/L3 ↔ permission modes, headless profiles, enumerate-vs-isolate. - [references/pattern-catalog.md](references/pattern-catalog.md) — the seven patterns, full skeletons + escalation rules. - [references/state-spine.md](references/state-spine.md) — STATE.md / run-log / budget schemas, multi-loop coordination. - [references/native-scheduling.md](references/native-scheduling.md) — the native primitives themselves (verified 2026-08-30): `CronCreate`/`/loop` + its dynamic mode, the `scheduled-tasks` MCP, cloud routines — parameters, limits, failure semantics, and what each still doesn't give you. - [references/claude-code-loops.md](references/claude-code-loops.md) — which mechanism and how to wire it: the recipe selector, event triggers, hooks, the external-scheduler shape. - [references/failure-modes.md](references/failure-modes.md) — how loops break (incident-shaped) and the control that catches each. - [assets/loop.config.template.yaml](assets/loop.config.template.yaml) — the loop definition starter; [assets/STATE.template.md](assets/STATE.template.md) — the state-spine starter; [assets/run.template.md](assets/run.template.md) — the headless run prompt. - Lineage (public sources): the [Ralph loop](https://ghuntley.com/ralph/) (fresh-context inner brute-force) and the broader *loop engineering* discipline framed by Peter Steinberger and Addy Osmani.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.