Claude Skill

autoresearch

Imported from paulrberg/agent-skills/skills/autoresearch.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download paulrberg-agent-skills-skills_autoresearch-913232a.zip · 10 KB
Part of paulrberg/agent-skills — 42 skills

Install

skills CLI npx skills add https://github.com/PaulRBerg/agent-skills/tree/main/skills/autoresearch
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install paulrberg-agent-skills@llmmart
Git git clone https://github.com/PaulRBerg/agent-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole paulrberg/agent-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Autoresearch

Use the parent for research decisions and Codex handoff workers for experiment execution. Measure consistently, retain only verified improvements, and stop on explicit resource or convergence limits.

Orchestration Contract

Start a new autoresearch session in Plan mode. Also require Plan mode before materially changing an approved session's objective, primary metric or direction, benchmark or correctness commands, write scope, or hard resource, cost, regression, or convergence limits. If Plan mode is required but inactive, ask the user to switch and stop. An unchanged approved contract may resume and receive new hypothesis batches outside Plan mode.

Always invoke $codex-handoff for implementation. Never fall back to direct parent implementation or another handoff mechanism. Follow its host selection, plan manifest, team sizing, validation ownership, reconciliation, failure, and completion contracts.

The parent owns the session contract, evidence synthesis, hypothesis selection and ordering, and stop decisions. It may delegate read-only repository investigation when useful, but research workers return evidence rather than hypotheses or plans. Keep the parent's execution work to orchestration, integrity checks, compact result review, and selection of the next batch.

Implementation workers execute parent-supplied ordered hypothesis batches. Default to one worker for a sequential search and pack multiple related hypotheses into one brief up to codex-handoff's sizing limit; never map one worker to each idea by default. Split only when hypotheses are genuinely independent, dependency waves require it, or one brief would be oversized. A worker may make local adjustments needed to execute an assigned hypothesis, but it must report rather than execute a materially different research direction.

Plan and Session Contract

Resolve the objective, primary metric and direction, benchmark and correctness commands, allowed/off-limits paths, run/runtime/command/cost/regression limits, convergence window, and reporting cadence before approval. Infer safe facts from the request and repository; ask only when a missing choice changes the experiment.

Defaults: 20 runs, two hours wall time, 10 minutes per benchmark, five minutes per correctness check, no new paid API spend, and convergence after five consecutive valid runs without a new retained best. Explicit --max-runs and --max-runtime values are hard limits.

The Plan-mode response must include the resolved contract, an evidence-backed ordered initial hypothesis batch with its completion or early-stop criteria, and codex-handoff's required plan section and manifest. Adding, removing, or reordering hypotheses inside the approved contract is follow-on planning, not a material contract change.

Delegated Execution

Read references/worker-loop.md before constructing an implementation brief. Inline its applicable instructions with the approved contract, ordered batch, exact paths and commands, current session and best-result state or first-batch status, and batch stopping criteria. Codex-handoff owns the remaining prompt and result fields.

The first implementation worker creates the isolation and session artifacts and records the unchanged baseline. Each worker leaves detailed measurements and logs in those artifacts and returns only the compact batch receipt required by the worker reference. The parent reconciles that receipt with the session module's JSON status, reads raw benchmark output only when a decision or integrity check requires it, then selects another batch or stops. Further batches under the unchanged contract remain follow-on work within the approved outcome.

Progress and Completion

Use codex-handoff's host-native progress surface. Send parent-authored updates only from settled evidence at the baseline, completed batch, material best change, blocker, or final stop; do not relay per-run narration. Render the session module's exact bar, counts, metrics, budgets, and convergence facts, and name the next parent-selected batch without recording it as settled work.

Finish with ### 🏁 Autoresearch complete — <stop reason>, baseline/best/delta/confidence, status counts, kept-file tree, exact checks, worktree/branch, and remaining cleanup or integration. Keep METRIC lines, JSONL, commands, and diagnostics undecorated. A resource limit is not convergence.

Files (agent-skills)
  • agents
    • openai.yaml 42 B
      policy:
        allow_implicit_invocation: true
      
  • references
    • loop-rules.md 2 KB
      # Loop Rules Reference
      
      Read only when a keep/discard decision is ambiguous, a benchmark is noisy, or the search repeats itself.
      
      ## Agent Judgment
      
      - The declared primary metric supplies direction; secondary metrics enforce explicit budgets and explain tradeoffs.
      - Keep only candidates that pass correctness checks and hard constraints.
      - Re-run gains within the measured noise floor. Module confidence is evidence, not a substitute for measurement.
      - Prefer less code for equivalent results. Do not retain complexity for an unconfirmed marginal gain.
      - Treat crashes, timeouts, missing metrics, and failed checks as failed experiments. Fix one trivial experiment mistake;
        otherwise record the outcome and revert it.
      
      The session module calculates retained bests, confidence, budget state, and the no-improvement window from validated
      records. Do not reproduce those calculations manually or override a reported segment/budget fact.
      
      ## Thrash and Noise
      
      Treat three variants of the same mechanism or failure as thrashing: record the lesson and choose a structurally
      different hypothesis. When ideas run out, inspect profiles, source, dependencies, or relevant papers before trying
      random parameters.
      
      Establish noise from repeated unchanged or best-known runs. Prefer medians for short noisy workloads and keep the
      sampling method stable. Do not move goalposts after seeing a result; changing the primary metric requires a new module
      segment and baseline.
      
      ## Ideas Backlog
      
      Add an item to `autoresearch.ideas.md` when promising work requires profiling, a coupled refactor, or a prerequisite.
      Remove tried or invalidated items on resume. The backlog never expands session limits.
      
      ## Safe Revert
      
      Before each experiment, record exact tracked and untracked in-scope paths and their state. Restore only those tracked
      paths and delete only newly created in-scope files. Never clean, checkout, or reset the repository broadly.
      
      Preserve `autoresearch.md`, `autoresearch.sh`, `autoresearch.checks.sh`, `autoresearch.jsonl`, and
      `autoresearch.ideas.md` through experiment reverts.
      
    • worker-loop.md 5.1 KB
      # Autoresearch Worker Loop
      
      The parent reads this reference when constructing an implementation batch through codex-handoff and inlines the
      applicable instructions in the worker brief. Workers execute the approved hypotheses; they do not choose the research
      direction.
      
      ## Brief Contract
      
      Require the brief to provide the approved session contract, ordered hypotheses, exact write scope and commands, current
      session and retained-best state or first-batch status, and batch stopping criteria. Return `blocked` when any of these
      inputs is missing or contradictory rather than inventing a contract.
      
      Execute the supplied hypotheses in order against the current retained best. Make only local implementation adjustments
      needed to test an assigned mechanism. If a result invalidates a remaining hypothesis, stop the batch and report that
      fact. Record promising new directions in `autoresearch.ideas.md` or the result, but do not execute them until the parent
      assigns them.
      
      ## Isolation
      
      For the first batch, prefer a dedicated branch in a separate Git worktree. Record its starting commit, path, initial
      status, allowed paths, and session files in `autoresearch.md`. If isolation is unavailable, require a clean worktree or
      explicit authorization to share it.
      
      Never run repository-wide clean, checkout, stash, or reset commands. Revert only paths changed by the current experiment
      from its recorded pre-run state, and remove only newly created in-scope paths. Preserve unrelated files and all session
      evidence.
      
      ## Session Module
      
      Create `autoresearch.md`, deterministic `autoresearch.sh`, optional `autoresearch.checks.sh`, append-only
      `autoresearch.jsonl`, and optional `autoresearch.ideas.md`. Resolve the module from the owning `SKILL.md` and initialize
      the JSONL before the baseline:
      
      ```sh
      uv run "<skill-dir>/scripts/autoresearch-session.py" init \
        --file autoresearch.jsonl --metric <name> --direction <higher|lower> \
        --max-runs <n> --max-runtime-seconds <seconds> \
        --max-cost <amount> --convergence-runs <n>
      ```
      
      The first config record declares `direction`. Record each completed attempt only after assigning its status:
      
      ```sh
      uv run "<skill-dir>/scripts/autoresearch-session.py" record \
        --file autoresearch.jsonl --metric <number> \
        --status <keep|discard|crash|checks_failed> \
        [--commit <id>] [--description <text>] \
        [--elapsed-seconds <n>] [--estimated-cost <amount>]
      ```
      
      Zero and negative metrics are valid. The worker owns mechanical `keep` versus `discard` judgment under the approved
      contract; the module validates records and uses the declared direction. When the primary metric changes under a newly
      approved contract, pass `--metric-name <new>` and `--direction <higher|lower>` on the first new record. The module
      appends a new segment config.
      
      Use `status --format json` for best/delta/MAD/confidence, counts, convergence, budgets, and exact progress rendering:
      
      ```sh
      uv run "<skill-dir>/scripts/autoresearch-session.py" status --file autoresearch.jsonl
      ```
      
      `scripts/confidence.sh [jsonl]` and `scripts/summary.sh [jsonl]` remain compatibility adapters. Malformed records or
      violated invariants fail; noisy, equivalent, or worker-discarded results are reported facts, not helper failures.
      
      ## Batch Loop
      
      1. Inspect all in-scope source plus relevant tests or profiles. For the first batch, create isolation and session files,
         then record an unchanged baseline.
      2. Before each hypothesis, snapshot allowed paths, implement the assigned mechanism, and run the benchmark within its
         timeout.
      3. Parse the declared metric. Missing metrics, crashes, timeouts, and failed correctness checks cannot be improvements.
      4. Run correctness checks for every candidate that might be retained.
      5. Use the session status plus repeated measurements to judge noise or equivalence. Keep only a verified improvement
         within all hard constraints; prefer simpler code when results are equivalent. Otherwise perform the scoped revert.
      6. Append the assigned record and update `autoresearch.md` when evidence changes the retained best or rules out an
         approach.
      7. Stop at the first batch criterion, hard session limit, user interruption, satisfied target, helper-reported
         convergence, or result that invalidates the remaining batch. Do not substitute an unassigned hypothesis.
      8. Read `<skill-dir>/references/loop-rules.md` only for ambiguous keep/discard judgment, noise handling, backlog
         maintenance, or thrash recovery.
      
      ## Batch Result
      
      Run the session module's JSON status after the last settled attempt. Preserve full commands, measurements, diagnostics,
      and lessons in the session artifacts. Return codex-handoff's required result fields plus a compact autoresearch receipt:
      
      - each attempted hypothesis in order with its `keep`, `discard`, `crash`, or `checks_failed` status and metric when
        available;
      - baseline, retained best, delta, confidence, status counts, budget state, and convergence state;
      - the exact batch stop reason and any remaining hypotheses invalidated or not attempted; and
      - suggested next directions, without implementing them.
      
      Do not return raw benchmark logs unless they are necessary evidence for a blocker or integrity failure.
      
  • scripts
    • autoresearch-session.py 14.1 KB
      #!/usr/bin/env python3
      """Maintain and summarize an append-only autoresearch session."""
      
      from __future__ import annotations
      
      import argparse
      import datetime as dt
      import json
      import math
      import os
      import statistics
      import sys
      from pathlib import Path
      from typing import Any
      
      
      STATUSES = {"keep", "discard", "crash", "checks_failed"}
      DIRECTIONS = {"higher", "lower"}
      
      
      class SessionError(ValueError):
          pass
      
      
      def utc_now() -> str:
          return dt.datetime.now(dt.timezone.utc).isoformat().replace("+00:00", "Z")
      
      
      def read_records(path: Path) -> list[dict[str, Any]]:
          try:
              lines = path.read_text(encoding="utf-8").splitlines()
          except OSError as exc:
              raise SessionError(str(exc)) from exc
          records: list[dict[str, Any]] = []
          for line_number, line in enumerate(lines, start=1):
              if not line.strip():
                  continue
              try:
                  record = json.loads(line)
              except json.JSONDecodeError as exc:
                  raise SessionError(f"line {line_number}: malformed JSON: {exc.msg}") from exc
              if not isinstance(record, dict):
                  raise SessionError(f"line {line_number}: record must be an object")
              records.append(record)
          validate_records(records)
          return records
      
      
      def validate_records(records: list[dict[str, Any]]) -> None:
          if not records or records[0].get("type") != "config":
              raise SessionError("first record must be a config record")
          configs: dict[int, dict[str, Any]] = {}
          last_run = 0
          for index, record in enumerate(records, start=1):
              if record.get("type") == "config":
                  segment = record.get("segment")
                  if not isinstance(segment, int) or segment < 0 or segment in configs:
                      raise SessionError(f"record {index}: invalid or duplicate config segment")
                  if record.get("direction") not in DIRECTIONS:
                      raise SessionError(f"record {index}: direction must be higher or lower")
                  if not isinstance(record.get("metricName"), str) or not record["metricName"]:
                      raise SessionError(f"record {index}: metricName must be non-empty")
                  configs[segment] = record
                  continue
              status = record.get("status")
              if status not in STATUSES:
                  raise SessionError(f"record {index}: invalid status")
              if not isinstance(record.get("run"), int) or record["run"] != last_run + 1:
                  raise SessionError(f"record {index}: run numbers must be consecutive from 1")
              last_run = record["run"]
              segment = record.get("segment")
              if segment not in configs:
                  raise SessionError(f"record {index}: run references an unknown segment")
              metric = record.get("metric")
              if isinstance(metric, bool) or not isinstance(metric, (int, float)) or not math.isfinite(metric):
                  raise SessionError(f"record {index}: metric must be a finite number")
              for field in ("elapsedSeconds", "estimatedCost"):
                  value = record.get(field, 0)
                  if isinstance(value, bool) or not isinstance(value, (int, float)) or value < 0 or not math.isfinite(value):
                      raise SessionError(f"record {index}: {field} must be a non-negative number")
      
      
      def append_record(path: Path, record: dict[str, Any]) -> None:
          encoded = (json.dumps(record, separators=(",", ":"), ensure_ascii=False) + "\n").encode()
          descriptor = os.open(path, os.O_WRONLY | os.O_APPEND)
          try:
              os.write(descriptor, encoded)
              os.fsync(descriptor)
          finally:
              os.close(descriptor)
      
      
      def current_config(records: list[dict[str, Any]]) -> dict[str, Any]:
          return max((record for record in records if record.get("type") == "config"), key=lambda item: item["segment"])
      
      
      def summarize(records: list[dict[str, Any]]) -> dict[str, Any]:
          config = current_config(records)
          segment = config["segment"]
          runs = [record for record in records if record.get("status") in STATUSES and record["segment"] == segment]
          all_runs = [record for record in records if record.get("status") in STATUSES]
          valid = [record for record in runs if record["status"] in {"keep", "discard"}]
          kept = [record for record in valid if record["status"] == "keep"]
          baseline = valid[0] if valid else None
          candidates = kept or ([baseline] if baseline else [])
          direction = config["direction"]
          best = (max if direction == "higher" else min)(candidates, key=lambda item: item["metric"]) if candidates else None
          metrics = [record["metric"] for record in valid]
          mad = None
          if len(metrics) >= 3:
              median = statistics.median(metrics)
              mad = statistics.median(abs(value - median) for value in metrics)
          delta = None if not baseline or not best else best["metric"] - baseline["metric"]
          improvement = None if delta is None else (delta if direction == "higher" else -delta)
          confidence = None if improvement is None or mad in (None, 0) else improvement / mad
          confidence_level = None
          if confidence is not None:
              confidence_level = "likely_real" if confidence >= 2 else "marginal" if confidence >= 1 else "within_noise"
      
          consecutive = 0
          best_so_far: float | None = None
          for record in valid:
              if record["status"] == "keep" and (
                  best_so_far is None
                  or (direction == "higher" and record["metric"] > best_so_far)
                  or (direction == "lower" and record["metric"] < best_so_far)
              ):
                  best_so_far = record["metric"]
                  consecutive = 0
              else:
                  consecutive += 1
      
          total_runs = len(all_runs)
          max_runs = config.get("maxRuns")
          elapsed = sum(float(record.get("elapsedSeconds", 0)) for record in all_runs)
          cost = sum(float(record.get("estimatedCost", 0)) for record in all_runs)
          convergence_runs = config.get("convergenceRuns")
          filled = None if not max_runs else min(10, math.floor(10 * total_runs / max_runs + 0.5))
          counts = {status: sum(record["status"] == status for record in runs) for status in sorted(STATUSES)}
          return {
              "schemaVersion": 1,
              "segment": segment,
              "metricName": config["metricName"],
              "direction": direction,
              "runsCompleted": total_runs,
              "segmentRuns": len(runs),
              "counts": counts,
              "baseline": baseline["metric"] if baseline else None,
              "best": best["metric"] if best else None,
              "bestRun": best["run"] if best else None,
              "delta": delta,
              "mad": mad,
              "confidence": confidence,
              "confidenceLevel": confidence_level,
              "consecutiveValidWithoutBest": consecutive,
              "converged": bool(convergence_runs is not None and consecutive >= convergence_runs),
              "budgets": {
                  "runs": {"used": total_runs, "limit": max_runs, "exhausted": bool(max_runs is not None and total_runs >= max_runs)},
                  "runtimeSeconds": {
                      "used": elapsed,
                      "limit": config.get("maxRuntimeSeconds"),
                      "exhausted": bool(config.get("maxRuntimeSeconds") is not None and elapsed >= config["maxRuntimeSeconds"]),
                  },
                  "cost": {
                      "used": cost,
                      "limit": config.get("maxCost"),
                      "exhausted": bool(config.get("maxCost") is not None and cost > config["maxCost"]),
                  },
              },
              "progress": None if filled is None else {"filled": filled, "empty": 10 - filled, "bar": "█" * filled + "░" * (10 - filled)},
              "runs": runs,
          }
      
      
      def render_summary(summary: dict[str, Any]) -> str:
          counts = summary["counts"]
          progress = f" [{summary['progress']['bar']}]" if summary["progress"] else ""
          lines = [
              f"AUTORESEARCH SUMMARY{progress}",
              f"Metric: {summary['metricName']} ({summary['direction']} is better), segment {summary['segment']}",
              (
                  f"Runs: {summary['segmentRuns']} segment | {counts['keep']} kept | {counts['discard']} discarded | "
                  f"{counts['crash']} crashed | {counts['checks_failed']} checks_failed"
              ),
              f"Baseline: {summary['baseline']}",
              f"Best: {summary['best']}",
          ]
          if summary["confidence"] is None:
              reason = "zero noise" if summary["mad"] == 0 else "insufficient valid runs"
              lines.append(f"Confidence: N/A ({reason})")
          else:
              lines.append(f"Confidence: {summary['confidence']:.2f}x ({summary['confidenceLevel'].upper()})")
          return "\n".join(lines) + "\n"
      
      
      def render_confidence(summary: dict[str, Any]) -> tuple[str, int]:
          valid_count = sum(summary["counts"][status] for status in ("keep", "discard"))
          if valid_count < 3:
              return f"Insufficient valid data in current segment: {valid_count} runs (need >= 3)\n", 1
          if summary["mad"] == 0:
              return "MAD is zero — no measurable noise in the data\n", 0
          if summary["confidence"] is None:
              return "No kept result available for confidence calculation\n", 0
          lines = [
              f"Confidence: {summary['confidence']:.2f}x ({summary['confidenceLevel'].upper()})",
              f"Baseline:   {summary['baseline']}",
              f"Best:       {summary['best']} ({summary['direction']} is better)",
              f"Delta:      {summary['delta']}",
              f"MAD:        {summary['mad']}",
          ]
          return "\n".join(lines) + "\n", 0
      
      
      def init_command(args: argparse.Namespace) -> int:
          if args.file.exists():
              raise SessionError(f"session file already exists: {args.file}")
          args.file.parent.mkdir(parents=True, exist_ok=True)
          config = {
              "type": "config",
              "schemaVersion": 1,
              "segment": 0,
              "metricName": args.metric,
              "direction": args.direction,
              "maxRuns": args.max_runs,
              "maxRuntimeSeconds": args.max_runtime_seconds,
              "maxCost": args.max_cost,
              "convergenceRuns": args.convergence_runs,
              "createdAt": utc_now(),
          }
          descriptor = os.open(args.file, os.O_WRONLY | os.O_CREAT | os.O_EXCL, 0o600)
          try:
              os.write(descriptor, (json.dumps(config, separators=(",", ":")) + "\n").encode())
              os.fsync(descriptor)
          finally:
              os.close(descriptor)
          print(json.dumps(config, indent=2))
          return 0
      
      
      def record_command(args: argparse.Namespace) -> int:
          records = read_records(args.file)
          config = current_config(records)
          metric_name = args.metric_name or config["metricName"]
          if metric_name != config["metricName"]:
              if args.direction is None:
                  raise SessionError("--direction is required when --metric-name starts a new segment")
              config = {
                  **{key: config.get(key) for key in ("maxRuns", "maxRuntimeSeconds", "maxCost", "convergenceRuns")},
                  "type": "config",
                  "schemaVersion": 1,
                  "segment": config["segment"] + 1,
                  "metricName": metric_name,
                  "direction": args.direction,
                  "createdAt": utc_now(),
              }
              append_record(args.file, config)
              records.append(config)
          elif args.direction is not None and args.direction != config["direction"]:
              raise SessionError("direction cannot change without a new metric segment")
          run_number = 1 + sum(record.get("status") in STATUSES for record in records)
          record = {
              "type": "run",
              "run": run_number,
              "segment": config["segment"],
              "metric": args.metric,
              "status": args.status,
              "commit": args.commit,
              "description": args.description,
              "elapsedSeconds": args.elapsed_seconds,
              "estimatedCost": args.estimated_cost,
              "recordedAt": utc_now(),
          }
          append_record(args.file, record)
          print(json.dumps(record, indent=2))
          return 0
      
      
      def status_command(args: argparse.Namespace) -> int:
          summary = summarize(read_records(args.file))
          if args.format == "json":
              print(json.dumps(summary, indent=2, ensure_ascii=False))
              return 0
          if not summary["runs"]:
              print("No experiment results found.")
              return 1
          if args.format == "summary":
              print(render_summary(summary), end="")
              return 0
          output, returncode = render_confidence(summary)
          print(output, end="")
          return returncode
      
      
      def build_parser() -> argparse.ArgumentParser:
          parser = argparse.ArgumentParser(description=__doc__)
          subparsers = parser.add_subparsers(dest="command", required=True)
          init = subparsers.add_parser("init")
          init.add_argument("--file", type=Path, default=Path("autoresearch.jsonl"))
          init.add_argument("--metric", required=True)
          init.add_argument("--direction", choices=sorted(DIRECTIONS), required=True)
          init.add_argument("--max-runs", type=int, default=20)
          init.add_argument("--max-runtime-seconds", type=float, default=7200)
          init.add_argument("--max-cost", type=float, default=0)
          init.add_argument("--convergence-runs", type=int, default=5)
          init.set_defaults(handler=init_command)
      
          record = subparsers.add_parser("record")
          record.add_argument("--file", type=Path, default=Path("autoresearch.jsonl"))
          record.add_argument("--metric", type=float, required=True)
          record.add_argument("--metric-name")
          record.add_argument("--direction", choices=sorted(DIRECTIONS))
          record.add_argument("--status", choices=sorted(STATUSES), required=True)
          record.add_argument("--commit")
          record.add_argument("--description", default="")
          record.add_argument("--elapsed-seconds", type=float, default=0)
          record.add_argument("--estimated-cost", type=float, default=0)
          record.set_defaults(handler=record_command)
      
          status = subparsers.add_parser("status")
          status.add_argument("--file", type=Path, default=Path("autoresearch.jsonl"))
          status.add_argument("--format", choices=("json", "summary", "confidence"), default="json")
          status.set_defaults(handler=status_command)
          return parser
      
      
      def main() -> int:
          parser = build_parser()
          args = parser.parse_args()
          try:
              if getattr(args, "max_runs", 1) is not None and getattr(args, "max_runs", 1) <= 0:
                  raise SessionError("--max-runs must be positive")
              if getattr(args, "convergence_runs", 1) is not None and getattr(args, "convergence_runs", 1) <= 0:
                  raise SessionError("--convergence-runs must be positive")
              return args.handler(args)
          except SessionError as exc:
              print(f"ERROR: {exc}", file=sys.stderr)
              return 64
      
      
      if __name__ == "__main__":
          raise SystemExit(main())
      
    • confidence.sh 250 B
      #!/bin/bash
      # Compatibility wrapper for MAD confidence output.
      
      set -euo pipefail
      
      script_dir="$(cd "$(dirname "$0")" && pwd -P)"
      exec python3 "$script_dir/autoresearch-session.py" status \
        --file "${1:-autoresearch.jsonl}" \
        --format confidence
      
    • summary.sh 247 B
      #!/bin/bash
      # Compatibility wrapper for the session dashboard.
      
      set -euo pipefail
      
      script_dir="$(cd "$(dirname "$0")" && pwd -P)"
      exec python3 "$script_dir/autoresearch-session.py" status \
        --file "${1:-autoresearch.jsonl}" \
        --format summary
      
  • SKILL.md 4.8 KB
    ---
    argument-hint: <goal> [--max-runs N] [--max-runtime DURATION]
    compatibility:
      Requires Plan mode to establish or materially change a session contract and a codex-handoff-compatible host for
      implementation.
    name: autoresearch
    skill-dependencies:
      - codex-handoff
    description:
      Use for autoresearch or "optimize X overnight/in a loop"; plans bounded measurable experiment batches, then delegates
      execution through codex-handoff.
    ---
    
    # Autoresearch
    
    Use the parent for research decisions and Codex handoff workers for experiment execution. Measure consistently, retain
    only verified improvements, and stop on explicit resource or convergence limits.
    
    ## Orchestration Contract
    
    Start a new autoresearch session in Plan mode. Also require Plan mode before materially changing an approved session's
    objective, primary metric or direction, benchmark or correctness commands, write scope, or hard resource, cost,
    regression, or convergence limits. If Plan mode is required but inactive, ask the user to switch and stop. An unchanged
    approved contract may resume and receive new hypothesis batches outside Plan mode.
    
    Always invoke `$codex-handoff` for implementation. Never fall back to direct parent implementation or another handoff
    mechanism. Follow its host selection, plan manifest, team sizing, validation ownership, reconciliation, failure, and
    completion contracts.
    
    The parent owns the session contract, evidence synthesis, hypothesis selection and ordering, and stop decisions. It may
    delegate read-only repository investigation when useful, but research workers return evidence rather than hypotheses or
    plans. Keep the parent's execution work to orchestration, integrity checks, compact result review, and selection of the
    next batch.
    
    Implementation workers execute parent-supplied ordered hypothesis batches. Default to one worker for a sequential search
    and pack multiple related hypotheses into one brief up to codex-handoff's sizing limit; never map one worker to each
    idea by default. Split only when hypotheses are genuinely independent, dependency waves require it, or one brief would
    be oversized. A worker may make local adjustments needed to execute an assigned hypothesis, but it must report rather
    than execute a materially different research direction.
    
    ## Plan and Session Contract
    
    Resolve the objective, primary metric and direction, benchmark and correctness commands, allowed/off-limits paths,
    run/runtime/command/cost/regression limits, convergence window, and reporting cadence before approval. Infer safe facts
    from the request and repository; ask only when a missing choice changes the experiment.
    
    Defaults: 20 runs, two hours wall time, 10 minutes per benchmark, five minutes per correctness check, no new paid API
    spend, and convergence after five consecutive valid runs without a new retained best. Explicit `--max-runs` and
    `--max-runtime` values are hard limits.
    
    The Plan-mode response must include the resolved contract, an evidence-backed ordered initial hypothesis batch with its
    completion or early-stop criteria, and codex-handoff's required plan section and manifest. Adding, removing, or
    reordering hypotheses inside the approved contract is follow-on planning, not a material contract change.
    
    ## Delegated Execution
    
    Read `references/worker-loop.md` before constructing an implementation brief. Inline its applicable instructions with
    the approved contract, ordered batch, exact paths and commands, current session and best-result state or first-batch
    status, and batch stopping criteria. Codex-handoff owns the remaining prompt and result fields.
    
    The first implementation worker creates the isolation and session artifacts and records the unchanged baseline. Each
    worker leaves detailed measurements and logs in those artifacts and returns only the compact batch receipt required by
    the worker reference. The parent reconciles that receipt with the session module's JSON status, reads raw benchmark
    output only when a decision or integrity check requires it, then selects another batch or stops. Further batches under
    the unchanged contract remain follow-on work within the approved outcome.
    
    ## Progress and Completion
    
    Use codex-handoff's host-native progress surface. Send parent-authored updates only from settled evidence at the
    baseline, completed batch, material best change, blocker, or final stop; do not relay per-run narration. Render the
    session module's exact bar, counts, metrics, budgets, and convergence facts, and name the next parent-selected batch
    without recording it as settled work.
    
    Finish with `### 🏁 Autoresearch complete — <stop reason>`, baseline/best/delta/confidence, status counts, kept-file
    tree, exact checks, worktree/branch, and remaining cleanup or integration. Keep `METRIC` lines, JSONL, commands, and
    diagnostics undecorated. A resource limit is not convergence.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related