{"slug":"gate","title":"gate","summary":"Apply the acceptance decision that keeps optimization honest — always on the val split, by default requiring the improvement to exceed the significance bar (Δ > k·SE) so noise is not mistaken for progress. Use to inspect or reproduce a single accept/reject decision; the algorithm","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-01T18:59:10.644501Z","repo":{"url":"https://github.com/skillberry-ai/cap-evolve","stars":61,"forks":16,"license":"Apache-2.0","updatedAt":"2026-09-26T21:16:54Z"},"bodyHtml":"<hr>\n<h2>name: gate\ndescription: Apply the acceptance decision that keeps optimization honest — always on the val split, by default requiring the improvement to exceed the significance bar (Δ &gt; k·SE) so noise is not mistaken for progress. Use to inspect or reproduce a single accept/reject decision; the algorithms apply it internally every iteration.\ncomponent: phase\nargument-hint: \"--current R --candidate R --mode significant --k-se 1.0\"\nallowed-tools: Bash\nprovides: [decision]\nneeds: [scores]\nsources: []</h2>\n<h1>gate — accept only real improvements, on val</h1>\n<p>The gate is where dishonest optimization is prevented. Search is a noise\namplifier: try enough candidates and some will <em>look</em> better by chance alone\n(the more candidates you screen, the larger the expected best-of-noise). The gate\nis the rule that keeps a lucky draw from being promoted to \"the new best\". It\nrefuses any split but <code>val</code>, and by default accepts a candidate only when its val\nreward beats the current best by <strong>more than <code>k</code> standard errors</strong>.</p>\n<h2>Inputs / outputs (manifest tokens)</h2>\n<ul>\n<li><strong>needs:</strong> <code>scores</code> — the candidate's and current best's val reward <em>and</em>\n<code>stderr</code> (from <code>evaluate</code>). The SE is not optional: significance is meaningless\nwithout it.</li>\n<li><strong>provides:</strong> <code>decision</code> — <code>{accept, reason, delta, threshold}</code>, the audit\nrecord of why a candidate was kept or rejected.</li>\n</ul>\n<h2>The significance rule</h2>\n<pre><code>paired (the default):   accept ⟺ mean(Δ[t]) &gt; k · SE(Δ)     over the SAME val tasks\nsignificant (fallback): accept ⟺ Δ = cand − curr &gt; k · sqrt(cand_se² + curr_se²)\n</code></pre>\n<p>The bar is <code>Δ &gt; k·SE</code> and <strong>not <code>Δ &gt; 0</code></strong> because search is a noise amplifier:\nscreen enough candidates and the best-looking one is best by <em>luck</em>, so <code>Δ &gt; 0</code>\nbanks noise as progress and the val curve climbs while nothing improved. Clearing\n<code>k</code> standard errors of the measurement's own error is what makes an accept mean\nsomething — turn this down and the run's numbers stop being evidence. <code>k=1</code> is\nlenient (~1σ); raise it to be stricter. It is the textual-optimization analogue of\nKoehn's bootstrap significance test for metric differences.</p>\n<p><code>paired</code> is stronger because both sides were scored on the <em>same</em> val tasks, so\nper-task difficulty cancels and only the paired variance counts; <code>significant</code>\ntreats the two means as independent samples and is only correct when they are.</p>\n<p><strong>Single-trial scores report <code>stderr=0</code>, collapsing <code>k·SE</code> to 0</strong> — then\n<code>significant</code> silently degrades to <code>strict</code> and accepts any positive blip. If you\nrun the significance gate, score with multiple trials (see <code>evaluate</code>).</p>\n<h2>Modes</h2>\n<ul>\n<li><code>paired</code> (<strong>the default</strong>): <code>mean(per-task Δ) &gt; k·SE(Δ)</code>. The loop selects it\nwhenever per-task val data exists (<code>harness.py:1524-1526</code>, <code>gepa.py:741-743</code>)\nand <code>capevolve.yaml</code> ships <code>gate_mode: paired</code>.</li>\n<li><code>significant</code>: <code>Δ &gt; k·SE_combined</code> — the <strong>unpaired fallback</strong>, used when the two\nsides aren't aligned per task. <code>decide()</code>'s own <code>mode=</code> parameter defaults here\nfor bare callers with no per-task data; that is not the default of a real run.</li>\n<li><code>threshold</code>: <code>Δ &gt; T</code> — a flat margin (use when you have a domain minimum\nworthwhile gain, e.g. \"don't bother unless +2pp\").</li>\n<li><code>strict</code>: <code>Δ &gt; 0</code> — any improvement. Only safe with a near-zero-variance scorer\n(deterministic, single correct answer).</li>\n</ul>\n<p>Anything else raises. There is no simplicity/size mode: it was unreachable dead\ncode (nothing ever supplied a size) so it silently behaved as <code>strict</code>, and it has\nbeen removed rather than documented.</p>\n<h2>No-regression (the second gate)</h2>\n<p>A mean can rise while previously-passing tasks silently break. Pair the\nsignificance gate with a <strong>no-regression</strong> check: reject a candidate that improves\nthe aggregate but <em>drops</em> any task that the current best passed. This is the same\ndual-gate discipline SWE-bench-style harnesses use (a patch must pass the new\ntests <strong>and</strong> not break the existing ones — FAIL_TO_PASS <em>and</em> PASS_TO_PASS).\n<code>diagnose</code> provides <code>kept_good</code> (the currently-passing tasks) precisely so this\ncheck has something to protect.</p>\n<h2>Dual-mode</h2>\n<p>This phase runs two ways from the <strong>same</strong> SKILL.md: standalone as the slash command <code>/cap-evolve:gate</code> (the <code>argument-hint</code> shows its run.py args), and orchestrator-callable — <code>cap-evolve run</code> / the <code>orchestrate</code> skill invokes the same <code>scripts/run.py</code> headlessly and threads the run dir between phases.</p>\n<h2>How to run</h2>\n<pre><code>python scripts/run.py --current 0.50 --candidate 0.62 \\\n    --mode significant --k-se 1.0 --candidate-stderr 0.03 --current-stderr 0.03\n</code></pre>\n<p>Algorithms call the gate internally every iteration via the harness; this skill\nexists so a human or agent can reproduce and <em>understand</em> a single decision.</p>\n<h2>What good vs bad looks like</h2>\n<ul>\n<li><strong>Good:</strong> <code>paired</code> mode (the default) with real multi-trial SEs; a no-regression check on\ntop; every accept/reject logged with its <code>reason</code>.</li>\n<li><strong>Bad:</strong> gating on <code>train</code> (the tool refuses this — it overfits the optimizer to\nthe data it edits against); <code>strict</code> mode on a noisy agent (accepts noise);\nraising the mean while quietly regressing tasks because no-regression was off.</li>\n</ul>\n<h2>References</h2>\n<ul>\n<li><code>references/concepts.md</code> — the difference-of-means SE, choosing <code>k</code>, the\nmultiple-comparisons motivation, the dual-gate / no-regression rationale, and\nwhy gating on val (never train, never test) is the honest split, with sources.</li>\n</ul>\n","files":[{"path":"meta.yaml","sizeBytes":301,"isText":true},{"path":"references/concepts.md","sizeBytes":5107,"isText":true},{"path":"scripts/abstract.py","sizeBytes":161,"isText":true},{"path":"scripts/_bootstrap.py","sizeBytes":3694,"isText":true},{"path":"scripts/check.py","sizeBytes":1735,"isText":true},{"path":"scripts/run.py","sizeBytes":4851,"isText":true},{"path":"SKILL.md","sizeBytes":5386,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-01T18:59:50.322724Z","sha256":"9CDD469C5DE8058762A42B5D47804309B7BFC1B775B8EA0FABCDCF9737B18DBE","sizeBytes":10256},"review":null,"source":{"repositoryUrl":"https://github.com/skillberry-ai/cap-evolve","path":"skills/phases/gate","license":"Apache-2.0","commit":"da4781c8c51a48f2c2bd7fb0cfc779d3940bf05b","subtreeSha":"E382FC8B5EB2D4050742D63AC68ED0F145548E28FB598F4BFFEB19E2A1E8A2D7","lastSyncedAt":"2026-09-26T23:11:58.154287Z"},"reviewedAt":"2026-09-01T19:01:23.492253Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/skillberry-ai/cap-evolve/tree/main/skills/phases/gate"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install skillberry-ai-cap-evolve@llmmart"},{"target":"git","command":"git clone https://github.com/skillberry-ai/cap-evolve.git"}]}