autoresearch
Imported from paulrberg/agent-skills/skills/autoresearch.
Install
npx skills add https://github.com/PaulRBerg/agent-skills/tree/main/skills/autoresearch
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install paulrberg-agent-skills@llmmart
git clone https://github.com/PaulRBerg/agent-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole paulrberg/agent-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Autoresearch
Use the parent for research decisions and Codex handoff workers for experiment execution. Measure consistently, retain only verified improvements, and stop on explicit resource or convergence limits.
Orchestration Contract
Start a new autoresearch session in Plan mode. Also require Plan mode before materially changing an approved session's objective, primary metric or direction, benchmark or correctness commands, write scope, or hard resource, cost, regression, or convergence limits. If Plan mode is required but inactive, ask the user to switch and stop. An unchanged approved contract may resume and receive new hypothesis batches outside Plan mode.
Always invoke $codex-handoff for implementation. Never fall back to direct parent implementation or another handoff
mechanism. Follow its host selection, plan manifest, team sizing, validation ownership, reconciliation, failure, and
completion contracts.
The parent owns the session contract, evidence synthesis, hypothesis selection and ordering, and stop decisions. It may delegate read-only repository investigation when useful, but research workers return evidence rather than hypotheses or plans. Keep the parent's execution work to orchestration, integrity checks, compact result review, and selection of the next batch.
Implementation workers execute parent-supplied ordered hypothesis batches. Default to one worker for a sequential search and pack multiple related hypotheses into one brief up to codex-handoff's sizing limit; never map one worker to each idea by default. Split only when hypotheses are genuinely independent, dependency waves require it, or one brief would be oversized. A worker may make local adjustments needed to execute an assigned hypothesis, but it must report rather than execute a materially different research direction.
Plan and Session Contract
Resolve the objective, primary metric and direction, benchmark and correctness commands, allowed/off-limits paths, run/runtime/command/cost/regression limits, convergence window, and reporting cadence before approval. Infer safe facts from the request and repository; ask only when a missing choice changes the experiment.
Defaults: 20 runs, two hours wall time, 10 minutes per benchmark, five minutes per correctness check, no new paid API
spend, and convergence after five consecutive valid runs without a new retained best. Explicit --max-runs and
--max-runtime values are hard limits.
The Plan-mode response must include the resolved contract, an evidence-backed ordered initial hypothesis batch with its completion or early-stop criteria, and codex-handoff's required plan section and manifest. Adding, removing, or reordering hypotheses inside the approved contract is follow-on planning, not a material contract change.
Delegated Execution
Read references/worker-loop.md before constructing an implementation brief. Inline its applicable instructions with
the approved contract, ordered batch, exact paths and commands, current session and best-result state or first-batch
status, and batch stopping criteria. Codex-handoff owns the remaining prompt and result fields.
The first implementation worker creates the isolation and session artifacts and records the unchanged baseline. Each worker leaves detailed measurements and logs in those artifacts and returns only the compact batch receipt required by the worker reference. The parent reconciles that receipt with the session module's JSON status, reads raw benchmark output only when a decision or integrity check requires it, then selects another batch or stops. Further batches under the unchanged contract remain follow-on work within the approved outcome.
Progress and Completion
Use codex-handoff's host-native progress surface. Send parent-authored updates only from settled evidence at the baseline, completed batch, material best change, blocker, or final stop; do not relay per-run narration. Render the session module's exact bar, counts, metrics, budgets, and convergence facts, and name the next parent-selected batch without recording it as settled work.
Finish with ### 🏁 Autoresearch complete — <stop reason>, baseline/best/delta/confidence, status counts, kept-file
tree, exact checks, worktree/branch, and remaining cleanup or integration. Keep METRIC lines, JSONL, commands, and
diagnostics undecorated. A resource limit is not convergence.
Files (agent-skills)
-
agents
-
openai.yaml 42 B
policy: allow_implicit_invocation: true
-
-
references
-
loop-rules.md 2 KB
# Loop Rules Reference Read only when a keep/discard decision is ambiguous, a benchmark is noisy, or the search repeats itself. ## Agent Judgment - The declared primary metric supplies direction; secondary metrics enforce explicit budgets and explain tradeoffs. - Keep only candidates that pass correctness checks and hard constraints. - Re-run gains within the measured noise floor. Module confidence is evidence, not a substitute for measurement. - Prefer less code for equivalent results. Do not retain complexity for an unconfirmed marginal gain. - Treat crashes, timeouts, missing metrics, and failed checks as failed experiments. Fix one trivial experiment mistake; otherwise record the outcome and revert it. The session module calculates retained bests, confidence, budget state, and the no-improvement window from validated records. Do not reproduce those calculations manually or override a reported segment/budget fact. ## Thrash and Noise Treat three variants of the same mechanism or failure as thrashing: record the lesson and choose a structurally different hypothesis. When ideas run out, inspect profiles, source, dependencies, or relevant papers before trying random parameters. Establish noise from repeated unchanged or best-known runs. Prefer medians for short noisy workloads and keep the sampling method stable. Do not move goalposts after seeing a result; changing the primary metric requires a new module segment and baseline. ## Ideas Backlog Add an item to `autoresearch.ideas.md` when promising work requires profiling, a coupled refactor, or a prerequisite. Remove tried or invalidated items on resume. The backlog never expands session limits. ## Safe Revert Before each experiment, record exact tracked and untracked in-scope paths and their state. Restore only those tracked paths and delete only newly created in-scope files. Never clean, checkout, or reset the repository broadly. Preserve `autoresearch.md`, `autoresearch.sh`, `autoresearch.checks.sh`, `autoresearch.jsonl`, and `autoresearch.ideas.md` through experiment reverts. -
worker-loop.md 5.1 KB
# Autoresearch Worker Loop The parent reads this reference when constructing an implementation batch through codex-handoff and inlines the applicable instructions in the worker brief. Workers execute the approved hypotheses; they do not choose the research direction. ## Brief Contract Require the brief to provide the approved session contract, ordered hypotheses, exact write scope and commands, current session and retained-best state or first-batch status, and batch stopping criteria. Return `blocked` when any of these inputs is missing or contradictory rather than inventing a contract. Execute the supplied hypotheses in order against the current retained best. Make only local implementation adjustments needed to test an assigned mechanism. If a result invalidates a remaining hypothesis, stop the batch and report that fact. Record promising new directions in `autoresearch.ideas.md` or the result, but do not execute them until the parent assigns them. ## Isolation For the first batch, prefer a dedicated branch in a separate Git worktree. Record its starting commit, path, initial status, allowed paths, and session files in `autoresearch.md`. If isolation is unavailable, require a clean worktree or explicit authorization to share it. Never run repository-wide clean, checkout, stash, or reset commands. Revert only paths changed by the current experiment from its recorded pre-run state, and remove only newly created in-scope paths. Preserve unrelated files and all session evidence. ## Session Module Create `autoresearch.md`, deterministic `autoresearch.sh`, optional `autoresearch.checks.sh`, append-only `autoresearch.jsonl`, and optional `autoresearch.ideas.md`. Resolve the module from the owning `SKILL.md` and initialize the JSONL before the baseline: ```sh uv run "<skill-dir>/scripts/autoresearch-session.py" init \ --file autoresearch.jsonl --metric <name> --direction <higher|lower> \ --max-runs <n> --max-runtime-seconds <seconds> \ --max-cost <amount> --convergence-runs <n> ``` The first config record declares `direction`. Record each completed attempt only after assigning its status: ```sh uv run "<skill-dir>/scripts/autoresearch-session.py" record \ --file autoresearch.jsonl --metric <number> \ --status <keep|discard|crash|checks_failed> \ [--commit <id>] [--description <text>] \ [--elapsed-seconds <n>] [--estimated-cost <amount>] ``` Zero and negative metrics are valid. The worker owns mechanical `keep` versus `discard` judgment under the approved contract; the module validates records and uses the declared direction. When the primary metric changes under a newly approved contract, pass `--metric-name <new>` and `--direction <higher|lower>` on the first new record. The module appends a new segment config. Use `status --format json` for best/delta/MAD/confidence, counts, convergence, budgets, and exact progress rendering: ```sh uv run "<skill-dir>/scripts/autoresearch-session.py" status --file autoresearch.jsonl ``` `scripts/confidence.sh [jsonl]` and `scripts/summary.sh [jsonl]` remain compatibility adapters. Malformed records or violated invariants fail; noisy, equivalent, or worker-discarded results are reported facts, not helper failures. ## Batch Loop 1. Inspect all in-scope source plus relevant tests or profiles. For the first batch, create isolation and session files, then record an unchanged baseline. 2. Before each hypothesis, snapshot allowed paths, implement the assigned mechanism, and run the benchmark within its timeout. 3. Parse the declared metric. Missing metrics, crashes, timeouts, and failed correctness checks cannot be improvements. 4. Run correctness checks for every candidate that might be retained. 5. Use the session status plus repeated measurements to judge noise or equivalence. Keep only a verified improvement within all hard constraints; prefer simpler code when results are equivalent. Otherwise perform the scoped revert. 6. Append the assigned record and update `autoresearch.md` when evidence changes the retained best or rules out an approach. 7. Stop at the first batch criterion, hard session limit, user interruption, satisfied target, helper-reported convergence, or result that invalidates the remaining batch. Do not substitute an unassigned hypothesis. 8. Read `<skill-dir>/references/loop-rules.md` only for ambiguous keep/discard judgment, noise handling, backlog maintenance, or thrash recovery. ## Batch Result Run the session module's JSON status after the last settled attempt. Preserve full commands, measurements, diagnostics, and lessons in the session artifacts. Return codex-handoff's required result fields plus a compact autoresearch receipt: - each attempted hypothesis in order with its `keep`, `discard`, `crash`, or `checks_failed` status and metric when available; - baseline, retained best, delta, confidence, status counts, budget state, and convergence state; - the exact batch stop reason and any remaining hypotheses invalidated or not attempted; and - suggested next directions, without implementing them. Do not return raw benchmark logs unless they are necessary evidence for a blocker or integrity failure.
-
-
scripts
-
autoresearch-session.py 14.1 KB
#!/usr/bin/env python3 """Maintain and summarize an append-only autoresearch session.""" from __future__ import annotations import argparse import datetime as dt import json import math import os import statistics import sys from pathlib import Path from typing import Any STATUSES = {"keep", "discard", "crash", "checks_failed"} DIRECTIONS = {"higher", "lower"} class SessionError(ValueError): pass def utc_now() -> str: return dt.datetime.now(dt.timezone.utc).isoformat().replace("+00:00", "Z") def read_records(path: Path) -> list[dict[str, Any]]: try: lines = path.read_text(encoding="utf-8").splitlines() except OSError as exc: raise SessionError(str(exc)) from exc records: list[dict[str, Any]] = [] for line_number, line in enumerate(lines, start=1): if not line.strip(): continue try: record = json.loads(line) except json.JSONDecodeError as exc: raise SessionError(f"line {line_number}: malformed JSON: {exc.msg}") from exc if not isinstance(record, dict): raise SessionError(f"line {line_number}: record must be an object") records.append(record) validate_records(records) return records def validate_records(records: list[dict[str, Any]]) -> None: if not records or records[0].get("type") != "config": raise SessionError("first record must be a config record") configs: dict[int, dict[str, Any]] = {} last_run = 0 for index, record in enumerate(records, start=1): if record.get("type") == "config": segment = record.get("segment") if not isinstance(segment, int) or segment < 0 or segment in configs: raise SessionError(f"record {index}: invalid or duplicate config segment") if record.get("direction") not in DIRECTIONS: raise SessionError(f"record {index}: direction must be higher or lower") if not isinstance(record.get("metricName"), str) or not record["metricName"]: raise SessionError(f"record {index}: metricName must be non-empty") configs[segment] = record continue status = record.get("status") if status not in STATUSES: raise SessionError(f"record {index}: invalid status") if not isinstance(record.get("run"), int) or record["run"] != last_run + 1: raise SessionError(f"record {index}: run numbers must be consecutive from 1") last_run = record["run"] segment = record.get("segment") if segment not in configs: raise SessionError(f"record {index}: run references an unknown segment") metric = record.get("metric") if isinstance(metric, bool) or not isinstance(metric, (int, float)) or not math.isfinite(metric): raise SessionError(f"record {index}: metric must be a finite number") for field in ("elapsedSeconds", "estimatedCost"): value = record.get(field, 0) if isinstance(value, bool) or not isinstance(value, (int, float)) or value < 0 or not math.isfinite(value): raise SessionError(f"record {index}: {field} must be a non-negative number") def append_record(path: Path, record: dict[str, Any]) -> None: encoded = (json.dumps(record, separators=(",", ":"), ensure_ascii=False) + "\n").encode() descriptor = os.open(path, os.O_WRONLY | os.O_APPEND) try: os.write(descriptor, encoded) os.fsync(descriptor) finally: os.close(descriptor) def current_config(records: list[dict[str, Any]]) -> dict[str, Any]: return max((record for record in records if record.get("type") == "config"), key=lambda item: item["segment"]) def summarize(records: list[dict[str, Any]]) -> dict[str, Any]: config = current_config(records) segment = config["segment"] runs = [record for record in records if record.get("status") in STATUSES and record["segment"] == segment] all_runs = [record for record in records if record.get("status") in STATUSES] valid = [record for record in runs if record["status"] in {"keep", "discard"}] kept = [record for record in valid if record["status"] == "keep"] baseline = valid[0] if valid else None candidates = kept or ([baseline] if baseline else []) direction = config["direction"] best = (max if direction == "higher" else min)(candidates, key=lambda item: item["metric"]) if candidates else None metrics = [record["metric"] for record in valid] mad = None if len(metrics) >= 3: median = statistics.median(metrics) mad = statistics.median(abs(value - median) for value in metrics) delta = None if not baseline or not best else best["metric"] - baseline["metric"] improvement = None if delta is None else (delta if direction == "higher" else -delta) confidence = None if improvement is None or mad in (None, 0) else improvement / mad confidence_level = None if confidence is not None: confidence_level = "likely_real" if confidence >= 2 else "marginal" if confidence >= 1 else "within_noise" consecutive = 0 best_so_far: float | None = None for record in valid: if record["status"] == "keep" and ( best_so_far is None or (direction == "higher" and record["metric"] > best_so_far) or (direction == "lower" and record["metric"] < best_so_far) ): best_so_far = record["metric"] consecutive = 0 else: consecutive += 1 total_runs = len(all_runs) max_runs = config.get("maxRuns") elapsed = sum(float(record.get("elapsedSeconds", 0)) for record in all_runs) cost = sum(float(record.get("estimatedCost", 0)) for record in all_runs) convergence_runs = config.get("convergenceRuns") filled = None if not max_runs else min(10, math.floor(10 * total_runs / max_runs + 0.5)) counts = {status: sum(record["status"] == status for record in runs) for status in sorted(STATUSES)} return { "schemaVersion": 1, "segment": segment, "metricName": config["metricName"], "direction": direction, "runsCompleted": total_runs, "segmentRuns": len(runs), "counts": counts, "baseline": baseline["metric"] if baseline else None, "best": best["metric"] if best else None, "bestRun": best["run"] if best else None, "delta": delta, "mad": mad, "confidence": confidence, "confidenceLevel": confidence_level, "consecutiveValidWithoutBest": consecutive, "converged": bool(convergence_runs is not None and consecutive >= convergence_runs), "budgets": { "runs": {"used": total_runs, "limit": max_runs, "exhausted": bool(max_runs is not None and total_runs >= max_runs)}, "runtimeSeconds": { "used": elapsed, "limit": config.get("maxRuntimeSeconds"), "exhausted": bool(config.get("maxRuntimeSeconds") is not None and elapsed >= config["maxRuntimeSeconds"]), }, "cost": { "used": cost, "limit": config.get("maxCost"), "exhausted": bool(config.get("maxCost") is not None and cost > config["maxCost"]), }, }, "progress": None if filled is None else {"filled": filled, "empty": 10 - filled, "bar": "█" * filled + "░" * (10 - filled)}, "runs": runs, } def render_summary(summary: dict[str, Any]) -> str: counts = summary["counts"] progress = f" [{summary['progress']['bar']}]" if summary["progress"] else "" lines = [ f"AUTORESEARCH SUMMARY{progress}", f"Metric: {summary['metricName']} ({summary['direction']} is better), segment {summary['segment']}", ( f"Runs: {summary['segmentRuns']} segment | {counts['keep']} kept | {counts['discard']} discarded | " f"{counts['crash']} crashed | {counts['checks_failed']} checks_failed" ), f"Baseline: {summary['baseline']}", f"Best: {summary['best']}", ] if summary["confidence"] is None: reason = "zero noise" if summary["mad"] == 0 else "insufficient valid runs" lines.append(f"Confidence: N/A ({reason})") else: lines.append(f"Confidence: {summary['confidence']:.2f}x ({summary['confidenceLevel'].upper()})") return "\n".join(lines) + "\n" def render_confidence(summary: dict[str, Any]) -> tuple[str, int]: valid_count = sum(summary["counts"][status] for status in ("keep", "discard")) if valid_count < 3: return f"Insufficient valid data in current segment: {valid_count} runs (need >= 3)\n", 1 if summary["mad"] == 0: return "MAD is zero — no measurable noise in the data\n", 0 if summary["confidence"] is None: return "No kept result available for confidence calculation\n", 0 lines = [ f"Confidence: {summary['confidence']:.2f}x ({summary['confidenceLevel'].upper()})", f"Baseline: {summary['baseline']}", f"Best: {summary['best']} ({summary['direction']} is better)", f"Delta: {summary['delta']}", f"MAD: {summary['mad']}", ] return "\n".join(lines) + "\n", 0 def init_command(args: argparse.Namespace) -> int: if args.file.exists(): raise SessionError(f"session file already exists: {args.file}") args.file.parent.mkdir(parents=True, exist_ok=True) config = { "type": "config", "schemaVersion": 1, "segment": 0, "metricName": args.metric, "direction": args.direction, "maxRuns": args.max_runs, "maxRuntimeSeconds": args.max_runtime_seconds, "maxCost": args.max_cost, "convergenceRuns": args.convergence_runs, "createdAt": utc_now(), } descriptor = os.open(args.file, os.O_WRONLY | os.O_CREAT | os.O_EXCL, 0o600) try: os.write(descriptor, (json.dumps(config, separators=(",", ":")) + "\n").encode()) os.fsync(descriptor) finally: os.close(descriptor) print(json.dumps(config, indent=2)) return 0 def record_command(args: argparse.Namespace) -> int: records = read_records(args.file) config = current_config(records) metric_name = args.metric_name or config["metricName"] if metric_name != config["metricName"]: if args.direction is None: raise SessionError("--direction is required when --metric-name starts a new segment") config = { **{key: config.get(key) for key in ("maxRuns", "maxRuntimeSeconds", "maxCost", "convergenceRuns")}, "type": "config", "schemaVersion": 1, "segment": config["segment"] + 1, "metricName": metric_name, "direction": args.direction, "createdAt": utc_now(), } append_record(args.file, config) records.append(config) elif args.direction is not None and args.direction != config["direction"]: raise SessionError("direction cannot change without a new metric segment") run_number = 1 + sum(record.get("status") in STATUSES for record in records) record = { "type": "run", "run": run_number, "segment": config["segment"], "metric": args.metric, "status": args.status, "commit": args.commit, "description": args.description, "elapsedSeconds": args.elapsed_seconds, "estimatedCost": args.estimated_cost, "recordedAt": utc_now(), } append_record(args.file, record) print(json.dumps(record, indent=2)) return 0 def status_command(args: argparse.Namespace) -> int: summary = summarize(read_records(args.file)) if args.format == "json": print(json.dumps(summary, indent=2, ensure_ascii=False)) return 0 if not summary["runs"]: print("No experiment results found.") return 1 if args.format == "summary": print(render_summary(summary), end="") return 0 output, returncode = render_confidence(summary) print(output, end="") return returncode def build_parser() -> argparse.ArgumentParser: parser = argparse.ArgumentParser(description=__doc__) subparsers = parser.add_subparsers(dest="command", required=True) init = subparsers.add_parser("init") init.add_argument("--file", type=Path, default=Path("autoresearch.jsonl")) init.add_argument("--metric", required=True) init.add_argument("--direction", choices=sorted(DIRECTIONS), required=True) init.add_argument("--max-runs", type=int, default=20) init.add_argument("--max-runtime-seconds", type=float, default=7200) init.add_argument("--max-cost", type=float, default=0) init.add_argument("--convergence-runs", type=int, default=5) init.set_defaults(handler=init_command) record = subparsers.add_parser("record") record.add_argument("--file", type=Path, default=Path("autoresearch.jsonl")) record.add_argument("--metric", type=float, required=True) record.add_argument("--metric-name") record.add_argument("--direction", choices=sorted(DIRECTIONS)) record.add_argument("--status", choices=sorted(STATUSES), required=True) record.add_argument("--commit") record.add_argument("--description", default="") record.add_argument("--elapsed-seconds", type=float, default=0) record.add_argument("--estimated-cost", type=float, default=0) record.set_defaults(handler=record_command) status = subparsers.add_parser("status") status.add_argument("--file", type=Path, default=Path("autoresearch.jsonl")) status.add_argument("--format", choices=("json", "summary", "confidence"), default="json") status.set_defaults(handler=status_command) return parser def main() -> int: parser = build_parser() args = parser.parse_args() try: if getattr(args, "max_runs", 1) is not None and getattr(args, "max_runs", 1) <= 0: raise SessionError("--max-runs must be positive") if getattr(args, "convergence_runs", 1) is not None and getattr(args, "convergence_runs", 1) <= 0: raise SessionError("--convergence-runs must be positive") return args.handler(args) except SessionError as exc: print(f"ERROR: {exc}", file=sys.stderr) return 64 if __name__ == "__main__": raise SystemExit(main()) -
confidence.sh 250 B
#!/bin/bash # Compatibility wrapper for MAD confidence output. set -euo pipefail script_dir="$(cd "$(dirname "$0")" && pwd -P)" exec python3 "$script_dir/autoresearch-session.py" status \ --file "${1:-autoresearch.jsonl}" \ --format confidence -
summary.sh 247 B
#!/bin/bash # Compatibility wrapper for the session dashboard. set -euo pipefail script_dir="$(cd "$(dirname "$0")" && pwd -P)" exec python3 "$script_dir/autoresearch-session.py" status \ --file "${1:-autoresearch.jsonl}" \ --format summary
-
-
SKILL.md 4.8 KB
--- argument-hint: <goal> [--max-runs N] [--max-runtime DURATION] compatibility: Requires Plan mode to establish or materially change a session contract and a codex-handoff-compatible host for implementation. name: autoresearch skill-dependencies: - codex-handoff description: Use for autoresearch or "optimize X overnight/in a loop"; plans bounded measurable experiment batches, then delegates execution through codex-handoff. --- # Autoresearch Use the parent for research decisions and Codex handoff workers for experiment execution. Measure consistently, retain only verified improvements, and stop on explicit resource or convergence limits. ## Orchestration Contract Start a new autoresearch session in Plan mode. Also require Plan mode before materially changing an approved session's objective, primary metric or direction, benchmark or correctness commands, write scope, or hard resource, cost, regression, or convergence limits. If Plan mode is required but inactive, ask the user to switch and stop. An unchanged approved contract may resume and receive new hypothesis batches outside Plan mode. Always invoke `$codex-handoff` for implementation. Never fall back to direct parent implementation or another handoff mechanism. Follow its host selection, plan manifest, team sizing, validation ownership, reconciliation, failure, and completion contracts. The parent owns the session contract, evidence synthesis, hypothesis selection and ordering, and stop decisions. It may delegate read-only repository investigation when useful, but research workers return evidence rather than hypotheses or plans. Keep the parent's execution work to orchestration, integrity checks, compact result review, and selection of the next batch. Implementation workers execute parent-supplied ordered hypothesis batches. Default to one worker for a sequential search and pack multiple related hypotheses into one brief up to codex-handoff's sizing limit; never map one worker to each idea by default. Split only when hypotheses are genuinely independent, dependency waves require it, or one brief would be oversized. A worker may make local adjustments needed to execute an assigned hypothesis, but it must report rather than execute a materially different research direction. ## Plan and Session Contract Resolve the objective, primary metric and direction, benchmark and correctness commands, allowed/off-limits paths, run/runtime/command/cost/regression limits, convergence window, and reporting cadence before approval. Infer safe facts from the request and repository; ask only when a missing choice changes the experiment. Defaults: 20 runs, two hours wall time, 10 minutes per benchmark, five minutes per correctness check, no new paid API spend, and convergence after five consecutive valid runs without a new retained best. Explicit `--max-runs` and `--max-runtime` values are hard limits. The Plan-mode response must include the resolved contract, an evidence-backed ordered initial hypothesis batch with its completion or early-stop criteria, and codex-handoff's required plan section and manifest. Adding, removing, or reordering hypotheses inside the approved contract is follow-on planning, not a material contract change. ## Delegated Execution Read `references/worker-loop.md` before constructing an implementation brief. Inline its applicable instructions with the approved contract, ordered batch, exact paths and commands, current session and best-result state or first-batch status, and batch stopping criteria. Codex-handoff owns the remaining prompt and result fields. The first implementation worker creates the isolation and session artifacts and records the unchanged baseline. Each worker leaves detailed measurements and logs in those artifacts and returns only the compact batch receipt required by the worker reference. The parent reconciles that receipt with the session module's JSON status, reads raw benchmark output only when a decision or integrity check requires it, then selects another batch or stops. Further batches under the unchanged contract remain follow-on work within the approved outcome. ## Progress and Completion Use codex-handoff's host-native progress surface. Send parent-authored updates only from settled evidence at the baseline, completed batch, material best change, blocker, or final stop; do not relay per-run narration. Render the session module's exact bar, counts, metrics, budgets, and convergence facts, and name the next parent-selected batch without recording it as settled work. Finish with `### 🏁 Autoresearch complete — <stop reason>`, baseline/best/delta/confidence, status counts, kept-file tree, exact checks, worktree/branch, and remaining cleanup or integration. Keep `METRIC` lines, JSONL, commands, and diagnostics undecorated. A resource limit is not convergence.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.