{"slug":"agent-optimize","title":"agent-optimize","summary":"Free-form optimization algorithm for agent orchestration mode: the conversational agent owns the whole search — proposing capability edits itself, screening them cheaply, gating each on full val, and sealing test once. Use when orchestration_mode is agent and algorithm_skill is a","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-01T18:59:07.796322Z","repo":{"url":"https://github.com/skillberry-ai/cap-evolve","stars":61,"forks":16,"license":"Apache-2.0","updatedAt":"2026-09-26T21:16:54Z"},"bodyHtml":"<hr>\n<h2>name: agent-optimize\ndescription: 'Free-form optimization algorithm for agent orchestration mode: the conversational agent owns the whole search — proposing capability edits itself, screening them cheaply, gating each on full val, and sealing test once. Use when orchestration_mode is agent and algorithm_skill is agent-optimize. For a deterministic loop use hill-climb, gepa or skillopt instead.'\ncomponent: algorithm\nargument-hint: \"agent-mode only — set orchestration_mode: agent + algorithm_skill: agent-optimize\"\nallowed-tools: Read, Write, Edit, Bash, Task\nprovides: [candidate]\nneeds: [scores, traces, candidate]</h2>\n<h1>agent-optimize — the free-form loop you own</h1>\n<p>The one algorithm with <strong>no deterministic subprocess</strong> and <strong>no per-iteration optimizer</strong>: you — the agent\nthat ran intake — are the optimizer, the scheduler and the stopping rule. <code>cap-evolve run</code> (with\n<code>orchestration_mode: agent</code>) does check → baseline, prints a handoff with the <code>run_dir</code>, and returns. From\nthere the search is yours, bounded by the invariants core enforces and the free-text <strong><code>stop_condition</code></strong>.\nDrive the <em>existing</em> primitives so the run dir and dashboard stay populated as in a deterministic run.</p>\n<h2>Shell variables used below</h2>\n<pre><code>R=\"&lt;run_dir from the agent-mode handoff&gt;\"      # e.g. .capevolve/run_20250101_120000\nP=\"&lt;project dir&gt;\"                              # the dir holding capevolve.yaml + adapters/\nS=\"${CAPEVOLVE_SKILLS_DIR:?set CAPEVOLVE_SKILLS_DIR to the skills/ dir}\"\nA=\"$S/algorithms/agent-optimize/scripts\"       # this skill's helpers\nmkdir -p \"$R/work\"                             # working copies live here (RunDir does NOT create it)\n</code></pre>\n<p>Every script imports <code>_bootstrap</code> itself (no <code>PYTHONPATH</code>) and prints JSON on stdout.</p>\n<h2>Phase 0 — understand before you optimize</h2>\n<p>Once, before any edit, and <strong>ask the user any blocking question here</strong> so the loop then runs unattended.\nRead <code>PROJECT.md</code>, <code>capevolve.yaml</code>, the adapter and every file under <code>capability_path</code>, and understand what\n<strong>one evaluation</strong> does: what a task is, what <code>run_target</code> produces, what <code>score()</code> rewards, and what the\nper-task <strong>feedback</strong> says — that is your learning signal. Note the val/test sizes, <code>num_trials</code>,\n<code>gate_mode</code>/<code>gate_k_se</code> and the allowed edit surface.</p>\n<p>Then let <code>spend.py</code> parse the free-text <strong><code>stop_condition</code></strong> rather than restating it from memory: it prints\n<code>constraints.predicates</code>, every concrete check it could extract, with its measured actual. <strong>If\n<code>constraints.ambiguous</code> is non-empty, ASK THE USER before the loop starts</strong> — a vague clause is reported,\nnever guessed at, and this is the one moment where asking is cheap.</p>\n<h2>Agent-mode loop</h2>\n<p>Baseline has scored the seed on val and set <code>best_id = seed</code>. Each round:</p>\n<p><strong>0. Check you can afford the round — for the number of candidates you actually intend to run</strong>, with\n<code>--n-siblings N</code> whenever you plan N of them, <em>before</em> spending:</p>\n<pre><code>python \"$A/spend.py\" --run-dir \"$R\" --project \"$P\" --n-siblings 3\n</code></pre>\n<p>Act on the single <code>recommendation</code>: <strong><code>stop</code></strong> (a ceiling breached, <code>budget_exhausted()</code> true, or the score\ngoal met on FULL val) → <strong>Stop &amp; seal</strong>; <strong><code>narrow_scope</code></strong> (≥80% of a ceiling consumed, goal unmet) → ONE\ncheap candidate at tier 1, no fan-out; <strong><code>continue</code></strong> → run the round you planned.</p>\n<p><code>afford.affordable: false</code> (with <code>afford.blockers</code> naming the ceiling) means <strong>do not fan out N</strong> — check\nBEFORE dispatching proposers, since N candidates can blow a budget with room for one.\n<code>afford.runner_spend_metered: false</code> means $0 is <em>unmetered</em>, not free — bound such a run with\n<code>max_metric_calls</code> and report <strong>rollout counts, not dollars</strong>.</p>\n<p><strong>1. Read the signal.</strong> Free — no new evaluation:</p>\n<pre><code>BEST=\"$(python \"$A/spend.py\" --run-dir \"$R\" | python -c 'import json,sys;print(json.load(sys.stdin)[\"best_id\"])')\"\npython \"$S/phases/diagnose/scripts/run.py\" --run-dir \"$R\" --tag \"$BEST\" --split train\npython \"$S/phases/diagnose/scripts/run.py\" --run-dir \"$R\" --tag \"$BEST\" --split val\n</code></pre>\n<p>Read <code>clusters</code> for what to fix and <code>kept_good</code> for what not to break. <strong>With a disjoint train split,\ndiagnose it too and compare its cluster signatures to val's</strong> — free, and it decides whether the round can\nwork at all: if the signatures are disjoint, no train-driven edit can move the val mean, and every candidate\nis rejected for a reason that looks exactly like a null result. Say which, in the report. (Baseline\nscores val only, so pay one <code>evaluate --split train</code> first.)</p>\n<p><strong>Read the per-task pass rate, not the per-task pass/fail.</strong> At <code>num_trials: n</code> a task's reward is\n<code>k/n</code>, and that fraction is what separates defects from noise:</p>\n<table>\n<thead>\n<tr>\n<th>per-task rate</th>\n<th>what it is</th>\n<th>what to do</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>0/n</code> – <code>3/10</code></td>\n<td>a real, reproducible defect</td>\n<td>this is where every edit should aim</td>\n</tr>\n<tr>\n<td><code>4/10</code> – <code>7/10</code></td>\n<td>genuinely unstable behaviour</td>\n<td>fix by <em>removing</em> ambiguity, not adding rules</td>\n</tr>\n<tr>\n<td><code>8/10</code> – <code>9/10</code></td>\n<td>noise around a working path</td>\n<td><strong>leave it alone</strong>; \"fixing\" it is how churn starts</td>\n</tr>\n</tbody>\n</table>\n<p>A task that \"regressed\" from <code>10/10</code> to <code>9/10</code> between rounds is a re-measurement, not damage.</p>\n<p><strong>Audit the MEASUREMENT before you credit a failure</strong>, in round 1 while it is still free (scoring\nre-derives on persisted rollouts): a failing task is a claim by the scorer, and an optimizer that skips\nthis optimizes against its own instrumentation. Does the feedback name the <strong>defect</strong> or only the tool;\ndoes any helper fail <strong>silently</strong>; is <em>silent</em> distinguished from <em>wrong</em>; did the rollout <strong>run</strong>, or is\nthis missing data wearing a 0.0; which components actually <strong>gate</strong>? The failure behind each item:\n<code>references/edit-design-lessons.md</code>.</p>\n<p><strong>After two rejected rounds, read the candidate's TRACE before writing a third</strong> — not \"was the rule\nright\" but \"did the agent follow it at all\". Never exercised ⇒ the <strong>form</strong> is wrong; exercised and\nstill wrong ⇒ the content is.</p>\n<p><strong>2. Propose an edit per candidate — and address EVERY cluster the round can afford</strong>, either as\n<strong>sibling candidates, default N≥3</strong> (one cluster each, gated independently — the safe default) or as <strong>one bold\nmulti-part edit</strong> (higher variance, but the only way a fix needing a prompt change <em>and</em> a tool change\nlands together). Bundle only <em>independent</em> parts — different files, different rules — so a rejected bundle\ncan be resubmitted as its surviving part; read <code>regressed</code> (screen) and <code>regressions</code> (gate) to know which\nto drop. That is also what stops <strong>churn</strong> — a candidate whose mean matches its parent while a <em>different</em>\nset of tasks passes — from reading as a tie.</p>\n<pre><code>TAG=\"cand_1\"                                   # unique per candidate — it IS the rollout tag\ncp -r \"$R/candidates/$BEST\" \"$R/work/$TAG\"\n# …now edit the files under $R/work/$TAG that your capability owns. Example only:\n# read capability_path in Phase 0 for the real layout.\n</code></pre>\n<p>Every edit encodes a <strong>general rule</strong> — never a task's id, gold value, or answer.</p>\n<p><strong>Choose the edit FORM from the failure TYPE — before you write a word.</strong> The form matters more than the\nwording, because the form that repairs one failure type measurably backfires on another:</p>\n<table>\n<thead>\n<tr>\n<th>the failure you observed</th>\n<th>the form that fixes it</th>\n<th>the form that makes it worse</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>the rule is stated and the agent skips it under pressure</td>\n<td>a prohibition plus the symptom that precedes it (\"if you are about to X, you have already failed\")</td>\n<td>restating the rule — a mid-tier model gets <em>less</em> compliant</td>\n</tr>\n<tr>\n<td>the agent complies but the call has the wrong shape</td>\n<td>a <strong>positive recipe</strong>: what the correct call IS, its parts, in order</td>\n<td>a list of things not to do — it produced <em>more</em> unwanted output than no guidance</td>\n</tr>\n<tr>\n<td>a required element is missing</td>\n<td>a <strong>structural REQUIRED slot</strong>, or a code-level precondition</td>\n<td>a prose reminder mid-document</td>\n</tr>\n<tr>\n<td>behaviour should differ by situation</td>\n<td>a conditional on an <strong>observable predicate</strong> the agent can evaluate from tool output</td>\n<td>an unconditional rule plus exemptions</td>\n</tr>\n</tbody>\n</table>\n<p>Then: <strong>no nuance clauses</strong>; <strong>exemption clauses do not scope</strong> (still suppresses X); <strong>prefer an in-code\nguard to a prose rule where the capability owns its tools</strong> — prose when the agent lacks a decision\ncriterion, code when it has one and violates it. Costs, and the guard-closure trap: <code>edit-design-lessons.md</code>.</p>\n<p><strong>Every round evaluates a null control</strong>: a byte-for-byte copy of the current best, evaluated like any\ncandidate — first, not after a surprising result. That eval is the round's own noise floor. And <strong>read\n<code>$R/rejected.jsonl</code> and make each proposal STRUCTURALLY different from what is in it</strong>: a different\nform, surface, or cluster — never a narrower version of a rejected rule.</p>\n<p><strong>3. Cheap SUBSET screen — the promotion ladder.</strong> Do not pay full val to learn an edit is bad:</p>\n<pre><code>python \"$A/screen.py\" --run-dir \"$R\" --project \"$P\" \\\n       --candidate \"$R/work/$TAG\" --tier 1 --k-se 1.0\n</code></pre>\n<p>Only the candidate pays, and only for the subset. <code>decision</code> is <code>kill</code> or <code>promote</code> — <strong>never accept</strong> — and\nit kills only on proven harm. <strong>Check the arithmetic before trusting a screen:</strong>\n<code>savings.breakeven_kill_rate</code> (<code>fired / full_val_rollouts</code>) is the fraction it must kill to pay for itself;\n<code>savings.net_rollouts</code> books what it cost. Screen only when that break-even sits below your observed kill\nrate — on a small val the tier-1 floor makes it unreachable, so pay full val directly — and read a screen as\nevidence about the tasks the edit targeted, never as a gate decision.</p>\n<p><strong>4. Honest gate on FULL val.</strong> Evaluate the whole split (this writes rollouts + results under tag\n<code>$TAG</code> — the evaluate phase tags by the candidate <strong>dir name</strong>), then decide off those rollouts:</p>\n<pre><code>python \"$S/phases/evaluate/scripts/run.py\" --run-dir \"$R\" --project \"$P\" \\\n       --candidate \"$R/work/$TAG\" --split val --n-trials &lt;num_trials&gt;\npython \"$A/gate_check.py\" --run-dir \"$R\" --candidate \"$TAG\" --k-se &lt;gate_k_se&gt;\n</code></pre>\n<p><code>\"verdict\"</code> is evidence, not a command — decide accept/reject yourself, citing the numbers in\n<code>commit.py --note</code> (<code>references/algorithm.md</code>, \"Gate as evidence\"). <code>\"indecisive\"</code> means too little of val\nran, not a rejection. <strong><code>regressions</code> is diagnosis, not a veto</strong> — a per-task drop at <code>n</code> trials is an\nestimate, not proof, and the old no-regression veto rejected byte-identical seed copies often enough to dominate false\nrejects (<code>--veto-regressions</code> restores it, at that rate). <code>phases/gate/scripts/run.py</code> inspects the same\ngate but books no decision.</p>\n<p><strong>5. Commit the decision through the run dir</strong>, so <code>best_id</code>, the stall counter, the dashboard and the\naudit log stay real. <code>--decision reject</code> keeps the old best; either way it snapshots the candidate, logs\nthe event and advances <code>iterations</code> + stall:</p>\n<pre><code>python \"$A/commit.py\" --run-dir \"$R\" --candidate-id \"$TAG\" --from-dir \"$R/work/$TAG\" \\\n       --decision accept --val &lt;cand_mean&gt; --note \"&lt;one line: the general rule you added&gt;\"\n</code></pre>\n<p><strong>On a reject, pass <code>--reject-basis</code></strong> — <code>screen.py</code>'s \"promote\" means \"could not prove harm\", never \"was\nevaluated on full val\", so conflating the two makes the run's artifacts contradict themselves. <code>gate</code> (a\nfull-val paired gate ran and said reject), <code>screen_kill</code> (the screen proved harm), <code>ceiling</code> (arithmetic\nproved no accept reachable, full val never paid), <code>budget</code> (screen evidence plus a budget call, not a\ngate decision), <code>infra</code> (missing data). So <code>screen: promote</code> + <code>reject_basis: ceiling</code> is coherent.</p>\n<p><code>commit.py</code> <strong>refuses a <code>--candidate-id</code> that already carries a decision event</strong> (<code>--force</code> only to\nrepair a record deliberately): two drivers tagging a candidate alike otherwise produce two decision\nevents over ONE set of rollouts. Pass <code>--optimizer-usd/--optimizer-tokens/--optimizer-seconds</code> for\n<strong>your own</strong> proposal cost — the evaluate phase records the runner's, nothing records the proposer's.</p>\n<p><strong>Two decisions that are NOT rejects</strong> (a reject advances <strong>stall</strong>): <code>--decision inconclusive</code> for an\nunresolved round (<code>verdict_stable: false</code>), re-measured under a FRESH tag; <code>--decision provisional</code> for a\nΔ&gt;0 round under the bar (<code>directionally_positive_but_inconclusive</code>), after which <code>grow.py</code> buys trials on\nthe SAME candidate and re-gates at the pooled n, capped at 2 rounds. <code>references/algorithm.md</code>.</p>\n<p><strong>6. Write the handover before ending this round</strong> — append one <code>## Iteration &lt;cid&gt;</code> entry below\n<code>JOURNAL.md</code>'s marker: what you tried, why, what the numbers said. The only thing the NEXT round reads,\nand <code>commit.py</code> folds in only what you wrote (<code>references/algorithm.md</code>).</p>\n<h2>Parallel round (optional)</h2>\n<p><strong>The whole of steps 3–4 for a round is one command.</strong> <code>round.py</code> builds the null control, evaluates\nevery tag in parallel <em>processes</em> (each runs its own adapter <code>apply()</code>, which mutates a process-global\nregistry and must never be shared), gates them serially, and prints one table:</p>\n<pre><code>python \"$A/round.py\" --run-dir \"$R\" --project \"$P\" \\\n       --candidates cand_1,cand_2,cand_3 \\\n       --n-trials &lt;num_trials&gt; --k-se &lt;gate_k_se&gt; --concurrency 8 --max-parallel 2\n</code></pre>\n<p><code>--concurrency</code> is the gate's <em>measurement</em> concurrency and defaults deliberately low; <code>round.py</code>\nrefuses one too hot to resolve its own verdict, so never raise it to buy wall clock. Read\n<code>noise_floor_from_control</code> FIRST — a candidate inside that band is not evidence, whatever its verdict.\n<code>round.py</code> never commits: which part of a bundle to keep is your judgement.</p>\n<p>Four invariants, to state before every fan-out (the reasoning, and where fan-out pays best, are under\n<em>Parallelism</em> in <a href=\"references/algorithm.md\"><code>references/algorithm.md</code></a>):</p>\n<ol>\n<li><strong>Diagnosis fans out freely</strong> — read-only, zero rollouts: one <code>cap-evolve-diagnoser</code> per failure\ncluster or rollout shard, then merge their JSON.</li>\n<li><strong>Proposal fans out across distinct working copies, one <code>cp -r</code> per sibling, tag unique per sibling</strong> —\nrollouts are <code>&lt;task&gt;__&lt;tag&gt;__t&lt;k&gt;.json</code>, so a shared tag interleaves two evals into the same filenames\nand corrupts both scores.</li>\n<li><strong>The gate stays serial</strong> — gate + commit one sibling at a time, and after any accept <strong>re-run\n<code>gate_check.py</code> for every remaining sibling against the new best</strong>. Skipping that re-gate\ndouble-counts a gain and admits an edit that never beat what it now stacks on.</li>\n<li><strong>Never fan out across the test split, and pay before you fan out</strong> — <code>spend.py --n-siblings N</code>\nmust say <code>affordable: true</code> first.</li>\n</ol>\n<p>Concurrency also composes <em>inside</em> one evaluation (<code>screen.py --workers N</code> / <code>CAPEVOLVE_WORKERS=N</code>, pooling\nrollout generation only — numbers stay byte-identical to serial). Opt in only when <code>run_target</code> is\nthread-safe: no shared scratch dir, single live container, or module-global client.</p>\n<h3>Per-task fan-out — the cheap gradient</h3>\n<p>Reach for this only when the baseline's <code>k/n</code> bands show the loss <strong>concentrated in a few named tasks</strong>: one\ntask at <code>n_trials</code> then buys the same bit as a <code>val_n × n_trials</code> full-val round, about a failure that\ndemonstrably exists. Helpers, in order — <code>taskeval.py</code> (run <strong>detached</strong>: a per-task eval can outlive a\nharness timeout while healthy), <code>mechanisms.py</code> (the shared ledger; <code>list</code> BEFORE you diagnose, or two\noptimisers implement one fix and collide at merge with only one measured), <code>integrate.py</code>, <code>funcmerge.py</code>,\n<code>merge_taskopt.py</code> — then gate the artifact once on full val via <code>round.py</code>. Economics, briefing contract,\ncanary selection, every flag: <a href=\"references/per-task-fanout.md\"><code>references/per-task-fanout.md</code></a>. Two rules\ndecide whether the shape is safe at all, so they live here:</p>\n<p><strong>A parallel optimiser's deliverable is a MECHANISM WITH TRACE PROOF, not a rate.</strong> A fan-out is a\nhigh-load regime by construction — where a per-task rate cannot resolve the effect — so ask for\nload-independent evidence (the guard fired, the next action changed), then gate the survivors serially.</p>\n<p><strong>A multi-branch artifact is assembled with <code>integrate.py</code>, never by one merge</strong>, one branch at a time with\na measurement after each: fewer mechanisms routinely beat more, and one number for N simultaneous changes\ncannot tell you that. <code>funcmerge</code> merging cleanly is <strong>not</strong> evidence the branches compose — Clean merge is\na syntactic property; composition is an empirical one.</p>\n<h2>Measurement discipline</h2>\n<p><strong>Measure step 2's null control twice</strong>: the gap between two byte-identical parents is the round's bar, and a\nbar smaller than that is not a gate. <code>round.py</code> does that, and reuses the replicates while <code>best_id</code> is\nunchanged (<code>control_reuse</code>). Two more rules; the rest — ceiling arithmetic, the binomial floor,\nmechanism-vs-artifact designs, gating the sum not each addend, the sign test below the floor — is in\n<a href=\"references/measured-lessons.md\"><code>references/measured-lessons.md</code></a>.</p>\n<ol>\n<li><strong>Explore fast, gate slow, gate ALONE.</strong> The load knob is <em>total in-flight requests</em> (K processes at\nconcurrency C is K·C), not any per-process flag, and oversubscription fails silently as latency, not an\nerror. Pause the fan-out, run both gate arms in one batch alone; if you cannot quiet the machine, say\nso next to the verdict.</li>\n<li><strong>Two independently-seeded blocks, agreeing in sign, before a small effect is a result.</strong> A paired\nrun's SE is over <em>tasks</em>, so it cannot see run-to-run nondeterminism; <code>multirep.py</code> takes the error\nacross whole runs (<code>--base-seed</code> picks the block — raising <code>--n</code> extends the same one, not a\nreplication). Several full runs unaffordable ⇒ \"not resolvable at this budget\" is the honest output.</li>\n</ol>\n<h2>Stop &amp; seal, then MEASURE (once)</h2>\n<p>Spend is not a CLI subcommand: <strong>every 2–3 rounds</strong> (and always before a fan-out) run <code>spend.py</code>.\nEverything it reports is re-read from the run dir, never a total in your head — which keeps a <code>$6.00</code>\ncap from becoming <code>$6.01</code>. (The Stop hook re-nudges you until finalized; a <code>PostToolUse</code> hook\nre-injects the same predicates on a cadence even if you skip <code>spend.py</code> — <code>goal_reminder.py</code>.)\nStop when <code>recommendation</code> is <code>stop</code>, then produce the run's one honest table — seed vs best on\n<strong>val</strong>, on <strong>train</strong> when the spec defines one worth reporting, and on the <strong>sealed test</strong> split\nscored once:</p>\n<pre><code>python \"$A/measure.py\" --run-dir \"$R\" --project \"$P\" --train auto\npython \"$S/phases/report/scripts/run.py\" --run-dir \"$R\"\n</code></pre>\n<p><code>measure.py</code> reads val off the rollouts the gate already used (free), evaluates train only when it adds\ninformation, and seals test through the same <code>harness.finalize</code> the finalize phase calls — so it is\ninterchangeable with <code>phases/finalize/scripts/run.py</code>. Report its four refusals unsoftened: an <strong>empty</strong> split is <code>empty</code>, not 0.0; a <strong>no-holdout</strong> spec is a <strong>FIT metric, not\ngeneralisation</strong>, with the overlap counted; a negative <code>screen_ledger.net_rollouts</code> says screening was\npure overhead; <code>best_id == \"seed\"</code> is a <strong>null result with a diagnosed cause</strong>, not a 0.000 gain.\n(Sealing is that phase script, <strong>not a CLI subcommand</strong>; a second finalize raises <code>TestSealError</code>.)\nNo finalize, no result.</p>\n<h2>Honesty invariants that are yours by hand</h2>\n<p>Core enforces the split seal, the val-only gate and the tamper guard whether you cooperate or not\n(<code>skills/phases/{evaluate,gate,finalize}</code> document them). Two are yours, because no script can do them\nfor you: <strong>never hand a subset result to <code>gate_check.py</code></strong> — its <code>coverage</code> reads 1.0 because its\ndenominator <em>is</em> the subset; and <strong>a round that produced no run-dir artifacts is a bug</strong>, so fix it\nrather than driving around the primitives.</p>\n<h2>References</h2>\n<p>One level deep — each is read on its own, and none points at another.</p>\n<ul>\n<li><a href=\"references/algorithm.md\"><code>references/algorithm.md</code></a> — why free-form, how the honesty invariants\nsurvive full autonomy, the screening break-even (incl. targeted-cluster holdouts), which steps\nparallelise safely, the constraint surface, provisional candidates. <strong>Load</strong> before relying on a\nscreen, growing a candidate, or skipping a rule.</li>\n<li><a href=\"references/measured-lessons.md\"><code>references/measured-lessons.md</code></a> — every measurement rule with the\nnumber that bought it: binomial floor, full val vs a hard subset, the load-vs-noise tables, the sign\ntest, the across-runs estimator. <strong>Load</strong> before your first gate decision on a new benchmark, or\nwhen a result surprises you.</li>\n<li><a href=\"references/per-task-fanout.md\"><code>references/per-task-fanout.md</code></a> — the fan-out's economics, the\nsubagent briefing contract, canary selection, every helper's flags. <strong>Load</strong> when the loss is\nconcentrated in a few named tasks.</li>\n<li><a href=\"references/edit-design-lessons.md\"><code>references/edit-design-lessons.md</code></a> — the scorer audit, guard\nclosure, and the measured backfires behind the edit-form table. <strong>Load</strong> before editing a surface\nfor the first time, or after two rejects.</li>\n</ul>\n","files":[{"path":"meta.yaml","sizeBytes":946,"isText":true},{"path":"references/algorithm.md","sizeBytes":54558,"isText":true},{"path":"references/context-sources.md","sizeBytes":1937,"isText":true},{"path":"references/edit-design-lessons.md","sizeBytes":9474,"isText":true},{"path":"references/measured-lessons.md","sizeBytes":46178,"isText":true},{"path":"references/microcase.md","sizeBytes":3906,"isText":true},{"path":"references/per-task-fanout.md","sizeBytes":8509,"isText":true},{"path":"scripts/abstract.py","sizeBytes":826,"isText":true},{"path":"scripts/_bootstrap.py","sizeBytes":3694,"isText":true},{"path":"scripts/check.py","sizeBytes":73843,"isText":true},{"path":"scripts/commit.py","sizeBytes":38595,"isText":true},{"path":"scripts/funcmerge.py","sizeBytes":26117,"isText":true},{"path":"scripts/gate_check.py","sizeBytes":15547,"isText":true},{"path":"scripts/grow.py","sizeBytes":10147,"isText":true},{"path":"scripts/host.py","sizeBytes":76537,"isText":true},{"path":"scripts/integrate.py","sizeBytes":14816,"isText":true},{"path":"scripts/linkcheck.py","sizeBytes":2325,"isText":true},{"path":"scripts/measure.py","sizeBytes":15922,"isText":true},{"path":"scripts/mechanisms.py","sizeBytes":7542,"isText":true},{"path":"scripts/merge_search.py","sizeBytes":18513,"isText":true},{"path":"scripts/merge_taskopt.py","sizeBytes":7937,"isText":true},{"path":"scripts/microcase.py","sizeBytes":17338,"isText":true},{"path":"scripts/multirep.py","sizeBytes":2821,"isText":true},{"path":"scripts/prepare_candidate.py","sizeBytes":7379,"isText":true},{"path":"scripts/round.py","sizeBytes":57888,"isText":true},{"path":"scripts/run.py","sizeBytes":2280,"isText":true},{"path":"scripts/screen.py","sizeBytes":12625,"isText":true},{"path":"scripts/spend.py","sizeBytes":10422,"isText":true},{"path":"scripts/taskeval.py","sizeBytes":12008,"isText":true},{"path":"scripts/watchdog.py","sizeBytes":6159,"isText":true},{"path":"SKILL.md","sizeBytes":20789,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-26T23:12:18.093322Z","sha256":"991533CDEA09CF9CF24BCE9D133964B07BABE66BCF1CBC3149FA9984A1AE1714","sizeBytes":219049},"review":null,"source":{"repositoryUrl":"https://github.com/skillberry-ai/cap-evolve","path":"skills/algorithms/agent-optimize","license":"Apache-2.0","commit":"da4781c8c51a48f2c2bd7fb0cfc779d3940bf05b","subtreeSha":"C81AE030E8B8CE8E63552C734489078FB8E8E628DC34E6F6C37FB4A6CB5D1224","lastSyncedAt":"2026-09-26T23:11:58.154287Z"},"reviewedAt":"2026-09-26T23:12:47.80709Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/skillberry-ai/cap-evolve/tree/main/skills/algorithms/agent-optimize"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install skillberry-ai-cap-evolve@llmmart"},{"target":"git","command":"git clone https://github.com/skillberry-ai/cap-evolve.git"}]}