{"slug":"baseline","title":"baseline","summary":"Establish the starting point. Use after implement-and-check and before any algorithm. Creates the run directory, freezes the seeded train/val/test split (written once), scores the unmodified seed capability on val, and records it as the candidate every algorithm must beat. Report","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-01T18:59:10.032243Z","repo":{"url":"https://github.com/skillberry-ai/cap-evolve","stars":61,"forks":16,"license":"Apache-2.0","updatedAt":"2026-09-26T21:16:54Z"},"bodyHtml":"<hr>\n<h2>name: baseline\ndescription: Establish the starting point. Use after implement-and-check and before any algorithm. Creates the run directory, freezes the seeded train/val/test split (written once), scores the unmodified seed capability on val, and records it as the candidate every algorithm must beat. Reports the remaining headroom so a saturated seed stops the run before it spends budget.\ncomponent: phase\nargument-hint: \"--base .capevolve --project DIR --capability DIR [--seed N] [--ratios a,b,c] [--n-trials N] [--split-ids FILE] [--resume] [--reuse-baseline DIR]\"\nallowed-tools: Read, Write, Bash\nprovides: [splits, baseline, candidate, scores, traces]\nneeds: [project, tasks]\nsources: []</h2>\n<h1>baseline — freeze splits, score the seed</h1>\n<p>baseline is the first phase that touches data, so it owns the run's one\nirreversible decision: <strong>the split</strong>. It writes <code>splits.json</code> once (seeded), scores\nthe <em>unmodified</em> seed capability on val, and records that score as the bar every\nalgorithm must beat.</p>\n<p>Run <code>implement-and-check</code> first. baseline re-runs that check itself and exits\nnon-zero before creating a run dir if it is red — a split frozen against a broken\nadapter poisons every number measured afterwards.</p>\n<h2>Why it matters</h2>\n<ul>\n<li><strong>Fair comparison point.</strong> Every algorithm hill-climbs <em>against</em> the baseline\nval score; a candidate that does not beat it is not progress.</li>\n<li><strong>Headroom.</strong> The printed JSON carries <code>headroom</code> (<code>1 - val</code>) and\n<code>headroom_verdict</code>: <code>saturated</code> means the seed is already at the ceiling and\nfurther iterations buy noise — stop; <code>floor</code> (val at 0) usually means a\nmis-wired adapter rather than a hard task — re-check before spending budget;\n<code>ok</code> means proceed. The same verdict is logged as a <code>headroom</code> event so the\norchestrator can stop on it with no human reading the number.</li>\n</ul>\n<h2>Splitting choices</h2>\n<ul>\n<li><strong>Seeded ratio split</strong> (default <code>0.5 / 0.25 / 0.25</code>): deterministic given\n<code>--seed</code>. Reproducible runs partition identically.</li>\n<li><strong>Pinned split</strong> (<code>--split-ids</code>): a JSON <code>{train,val,test}</code> of ids — use a\nbenchmark's official split, or set all three equal to fit the whole set with\n<strong>no holdout</strong> (the test number is then a <em>fit</em> metric, not a held-out result;\nthe run dir records a <code>splits_warning</code> saying so).</li>\n<li>A ratio split that leaves val or test empty is <strong>refused</strong> — the gate would\nhave nothing to decide on and the sealed test number would cover no tasks.\nBelow 5 val tasks baseline warns: the gate's bar is optimistic at that <code>n</code>, and\na candidate that improves exactly one val task cannot reliably clear it at all\n(issue #351), so size val with the decisions it has to make in mind.</li>\n</ul>\n<h2>Reusing a prior baseline (<code>--reuse-baseline PRIOR_RUN_DIR</code>)</h2>\n<p>Re-scoring the seed is wasteful when the split + seed are unchanged.\n<code>--reuse-baseline &lt;prior run_* dir&gt;</code> (spec key <code>reuse_baseline</code>) copies that run's\n<code>splits.json</code>, <code>baseline.json</code>, seed snapshot and seed val rollouts into the fresh\nrun dir and skips the baseline eval; the copied <code>test_used</code> flag is reset so this\nrun can still finalize on test exactly once. <code>--resume</code> is the same-run variant:\nreopen an existing run dir, skip the eval when <code>baseline.json</code> is already there.\nBudget flags (<code>--max-iterations</code>, <code>--stall</code>, <code>--max-usd</code>, …) are accepted here\nbecause the run dir owns the budget and later phases read it from there.</p>\n<p>Runs standalone (<code>/cap-evolve:baseline</code>) or headlessly via <code>cap-evolve run</code>; same\n<code>scripts/run.py</code> either way.</p>\n<h2>How to run</h2>\n<pre><code>python scripts/run.py --base .capevolve --project .capevolve/project \\\n    --capability seed_capability --seed 0 --ratios 0.5,0.25,0.25 \\\n    --max-iterations 10 --stall 2\n</code></pre>\n<p>Prints the run-dir path (used by the algorithm + finalize), the split sizes, the\nbaseline val and the headroom verdict. Use <code>--n-trials ≥ 3</code> for stochastic\ntargets so the baseline carries a real <code>stderr</code> rather than 0.</p>\n<p>The one failure mode nothing later can repair is re-splitting after this phase: a\ntask migrating out of test leaks the held-out set, and every later number —\nincluding the sealed test score — becomes unfalsifiable.</p>\n<h2>References</h2>\n<ul>\n<li><code>references/concepts.md</code> — why the split is frozen once and seeded, how to read\nthe headroom verdict, and why no-holdout runs are fit metrics, with sources.</li>\n</ul>\n","files":[{"path":"meta.yaml","sizeBytes":352,"isText":true},{"path":"references/concepts.md","sizeBytes":3513,"isText":true},{"path":"scripts/abstract.py","sizeBytes":165,"isText":true},{"path":"scripts/_bootstrap.py","sizeBytes":3694,"isText":true},{"path":"scripts/check.py","sizeBytes":5431,"isText":true},{"path":"scripts/run.py","sizeBytes":12940,"isText":true},{"path":"SKILL.md","sizeBytes":4288,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-26T23:12:18.995838Z","sha256":"EDF2AEB5064829049CF83E9E78E9D8B4711D31586785A82C858FA06D114CBBED","sizeBytes":12790},"review":null,"source":{"repositoryUrl":"https://github.com/skillberry-ai/cap-evolve","path":"skills/phases/baseline","license":"Apache-2.0","commit":"da4781c8c51a48f2c2bd7fb0cfc779d3940bf05b","subtreeSha":"8AF3B7F7282B3568F925D3D2EE0EBCEDE996333FD853ED92F84D786B6FA80A83","lastSyncedAt":"2026-09-26T23:11:58.154287Z"},"reviewedAt":"2026-09-26T23:12:48.006076Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/skillberry-ai/cap-evolve/tree/main/skills/phases/baseline"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install skillberry-ai-cap-evolve@llmmart"},{"target":"git","command":"git clone https://github.com/skillberry-ai/cap-evolve.git"}]}