{"slug":"finalize-3","title":"finalize","summary":"Score the best candidate on the held-out TEST split exactly once and seal the run. Use as the last evaluation step, after optimization stops. The run dir enforces the seal — a second finalize raises an error — so the headline number is produced once on data the optimizer never sa","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-01T18:59:10.496514Z","repo":{"url":"https://github.com/skillberry-ai/cap-evolve","stars":61,"forks":16,"license":"Apache-2.0","updatedAt":"2026-09-26T21:16:54Z"},"bodyHtml":"<hr>\n<h2>name: finalize\ndescription: Score the best candidate on the held-out TEST split exactly once and seal the run. Use as the last evaluation step, after optimization stops. The run dir enforces the seal — a second finalize raises an error — so the headline number is produced once on data the optimizer never saw, the way an honest benchmark result must be.\ncomponent: phase\nargument-hint: \"--run-dir DIR --project DIR\"\nallowed-tools: Read, Bash\nprovides: [report]\nneeds: [candidate]\nsources: [tau2bench]</h2>\n<h1>finalize — the one honest number</h1>\n<p>Optimization hill-climbs on val: every accept decision consumed val as a tuning\nsignal, so by the end of search val is <em>optimistic</em> — it has been selected\nagainst. The number you <strong>report</strong> must come from data nothing was tuned against.\nfinalize scores the run's best candidate on the sealed <code>test</code> split, once, and\nwrites <code>final.json</code>. That file is the run's result.</p>\n<h2>One finalize is two evals, not one</h2>\n<p>A bare test number cannot be defended — a reader cannot tell whether it beat the\ncapability you started with. So one finalize scores test <strong>twice</strong>: the best\ncandidate as tag <code>FINAL</code>, and the untouched <code>seed</code> candidate as <code>FINAL_seed</code>\n(<code>harness.finalize</code>). <code>final.json</code> therefore carries <code>test</code>, <code>test_baseline</code>,\n<code>baseline_id</code>, and <code>test_delta</code> — the held-out <em>improvement</em>, which is the figure\n<code>report</code>, the dashboard, and the event stream all headline. If the best candidate\nIS the seed (nothing was accepted), the second eval is skipped and <code>test_delta</code> is\n0 by construction.</p>\n<p>So budget <code>--n-trials 3</code> as 3 trials × <strong>2 candidates</strong> × |test| rollouts — twice\nwhat the flag looks like it buys on a paid benchmark.</p>\n<p>Both evals sit inside <strong>one</strong> attempt and neither is a selection event: the delta\nis <em>reported</em>, never chosen on. That is why the seal counts attempts, not evals.</p>\n<h2>The seal (why \"exactly once\")</h2>\n<p>The instant test informs <em>any</em> choice — picking between finalists, \"double-\nchecking\" a low number, re-running until it looks better — it stops being held\nout, because each peek is a selection event that pulls the number from an\nunbiased estimate toward an optimistic fit metric (<code>references/concepts.md</code>).</p>\n<p><code>cap_evolve</code> enforces this in three parts (<code>rundir.py:358-407</code>), and the split\nbetween them is the whole design:</p>\n<ul>\n<li><strong>reserve</strong> — every <code>split=\"test\"</code> eval first <em>checks</em> the seal without burning\nit, so no phase other than finalize can reach test at all.</li>\n<li><strong>commit</strong> — the seal burns only after <code>final.json</code> is written, so a finalize\nthat dies <em>before</em> scoring leaves it unused and is honestly retryable. A\ntransient crash must not destroy a run's headline number.</li>\n<li><strong>attempt guard</strong> — seal-on-success alone cannot tell \"crashed before scoring\"\nfrom \"crashed after\". A real run hit the second case: a finalize killed by a\ntimeout had already scored test, the retry scored it again, and the reported\nheadline was that second look. <code>begin_test_attempt</code> refuses a retry once test\nrollouts exist on disk, before anything is spent.</li>\n</ul>\n<p>The seal refuses that mistake by default; it is not unbypassable.\n<code>CAPEVOLVE_ALLOW_TEST_RESCORE=1</code> is a deliberate opt-in override\n(<code>rundir.py:166</code>). Its own message promises the use \"is recorded in the run\" —\nnothing records it (issue #341), so a run that took a second look currently looks\nidentical to an honest one. If you set it, disclose it in the write-up yourself.</p>\n<p>Corollary: <strong>all selection happens before finalize.</strong> Choose the single best\ncandidate on val, <em>then</em> finalize it. Finalists that genuinely need comparing get\ncompared on val — never on test.</p>\n<h2>If finalize refuses</h2>\n<p>A <code>TestSealError</code> is three situations with three different right moves. Tell them\napart from <code>&lt;run&gt;/rollouts/test/</code> and <code>test_used</code> in <code>splits.json</code>:</p>\n<table>\n<thead>\n<tr>\n<th>State</th>\n<th>What happened</th>\n<th>Do this</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>no test rollouts, seal unused</td>\n<td>crashed before scoring</td>\n<td>Re-run finalize — the case seal-on-success exists for.</td>\n</tr>\n<tr>\n<td>test rollouts exist, seal unused</td>\n<td>crashed after scoring, before commit</td>\n<td>Do <strong>not</strong> re-score. Read the rollouts under <code>&lt;run&gt;/rollouts/test/</code> and report what that attempt already computed.</td>\n</tr>\n<tr>\n<td><code>test_used: true</code></td>\n<td>the run is finalized</td>\n<td>Read <code>final.json</code> and regenerate the human artifact with <code>report</code> alone; <code>cap-evolve run --resume</code> skips finalize for you.</td>\n</tr>\n</tbody>\n</table>\n<p>Never delete test rollouts or edit <code>splits.json</code> to get past the error — that\nmanufactures a clean-looking number from a split that has already been seen.</p>\n<h2>Dual-mode</h2>\n<p>This phase runs two ways from the <strong>same</strong> SKILL.md: standalone as the slash command <code>/cap-evolve:finalize</code> (the <code>argument-hint</code> shows its run.py args), and orchestrator-callable — <code>cap-evolve run</code> / the <code>orchestrate</code> skill invokes the same <code>scripts/run.py</code> headlessly and threads the run dir between phases.</p>\n<h2>How to run</h2>\n<pre><code>python scripts/run.py --run-dir .capevolve/run_XXXX --project .capevolve/project --n-trials 3\n</code></pre>\n<p>Multiple trials give the headline an honest <code>stderr</code> and a pass^k reliability\nfigure instead of one noisy point. Under <code>cap-evolve run</code> the count comes from\n<code>num_trials</code> in <code>capevolve.yaml</code> and <strong>defaults to 1</strong> — set it to ≥3, or the\norchestrated headline ships with <code>stderr</code> 0: the exact single point this warns\nagainst. If the split was configured with no holdout (test == train/val) the\nnumber is a <em>fit</em> metric, not a held-out result; the dashboard flags it, so say so\nin the summary too.</p>\n<p>Then read the result instead of just filing it: test ≈ val means the val gain\ngeneralized, test ≪ val means search overfit val — a real finding, not a reason to\nre-score.</p>\n<h2>References</h2>\n<ul>\n<li><code>references/concepts.md</code> — why each peek biases the estimate, the\ntrain-fits / val-selects / test-estimates rationale, no-holdout runs, and how\nthis maps to public benchmark protocol, with sources.</li>\n</ul>\n","files":[{"path":"meta.yaml","sizeBytes":299,"isText":true},{"path":"references/concepts.md","sizeBytes":4502,"isText":true},{"path":"scripts/abstract.py","sizeBytes":165,"isText":true},{"path":"scripts/_bootstrap.py","sizeBytes":3694,"isText":true},{"path":"scripts/check.py","sizeBytes":1653,"isText":true},{"path":"scripts/run.py","sizeBytes":1963,"isText":true},{"path":"SKILL.md","sizeBytes":5822,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-01T18:59:49.782096Z","sha256":"BC96143F9B8ACF7C8F157A1B868984DF4FBF9169461FB58A73CB19BDDF2978E9","sizeBytes":9267},"review":null,"source":{"repositoryUrl":"https://github.com/skillberry-ai/cap-evolve","path":"skills/phases/finalize","license":"Apache-2.0","commit":"da4781c8c51a48f2c2bd7fb0cfc779d3940bf05b","subtreeSha":"DC441A3DCCDF92D70317EC7A95058B8BE683B83517CE930FB25F42EAC52FBDBB","lastSyncedAt":"2026-09-26T23:11:58.154287Z"},"reviewedAt":"2026-09-01T19:01:23.396074Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/skillberry-ai/cap-evolve/tree/main/skills/phases/finalize"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install skillberry-ai-cap-evolve@llmmart"},{"target":"git","command":"git clone https://github.com/skillberry-ai/cap-evolve.git"}]}