Claude Skill

lab:autoresearch

Self-improving loop for plugin skills. Reads program.md, proposes one mutation per iteration, evaluates against deterministic scorer, keeps improvements via git, reverts failures. Targets weakest skill+dimension. Use with /loop for overnight runs.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oliver-kriska-claude-elixir-phoenix-lab_autoresearch-9767a82.zip · 21 KB
Part of oliver-kriska/claude-elixir-phoenix — 93 skills

Install

skills CLI npx skills add https://github.com/oliver-kriska/claude-elixir-phoenix/tree/main/lab/autoresearch
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oliver-kriska-claude-elixir-phoenix@llmmart
Git git clone https://github.com/oliver-kriska/claude-elixir-phoenix.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oliver-kriska/claude-elixir-phoenix collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Autoresearch — Plugin Skill Self-Improvement

Iteratively improve plugin skills via the autoresearch pattern: propose one mutation -> eval -> keep/revert -> repeat.

Usage

/lab:autoresearch                           # Targeted: attack weakest skill+dimension
/lab:autoresearch --skill review            # Focus on one skill
/lab:autoresearch --strategy sweep          # Process all skills alphabetically
/lab:autoresearch --dry-run                 # Show what would change, don't commit

For overnight runs:

/loop 5m /lab:autoresearch --strategy sweep --max-iterations 200

Iron Laws

  1. ONE mutation per iteration — if description needs "and", split into two
  2. NEVER mutate read-only files — check program.md before every write
  3. EVAL is deterministic — always use the wrapper script, never LLM-judge
  4. REVERT on regression OR checks failure — no exceptions
  5. LOG every iteration — use keep or revert command (never skip)
  6. CHECK ideas.md before proposing — don't rediscover known optimizations

Wrapper Script Commands

All eval/git/journal operations go through ONE script. Do NOT run these manually.

# Find the weakest skill+dimension
python3 lab/autoresearch/scripts/run-iteration.py target --strategy targeted

# Score a skill (before mutation, to get baseline)
python3 lab/autoresearch/scripts/run-iteration.py score <skill-name>

# After mutation: score + checks + compare → verdict (KEEP or REVERT)
python3 lab/autoresearch/scripts/run-iteration.py eval <skill-name>

# Act on verdict:
python3 lab/autoresearch/scripts/run-iteration.py keep <skill> <dim> <old> <new> \
  --desc "what changed" --asi '{"hypothesis": "why", "mechanism": "how"}'

python3 lab/autoresearch/scripts/run-iteration.py revert <skill> <dim> <old> <new> \
  --desc "what was attempted" --asi '{"hypothesis": "why", "regression": "what broke", "avoid": "do not retry this"}'

# Check overall progress
python3 lab/autoresearch/scripts/run-iteration.py status

Core Loop (ONE iteration)

Step 1: Read State

  1. Read lab/autoresearch/program.md (goals, mutable surface, rules)
  2. Read lab/autoresearch/ideas.md if it exists (deferred optimizations)
  3. Run: python3 lab/autoresearch/scripts/run-iteration.py status

Step 2: Select Target

Run: python3 lab/autoresearch/scripts/run-iteration.py target --strategy targeted

Parse the JSON: skill, dimension, failing_checks. If all_perfect → STOP.

Step 3: Read + Propose

  1. Read target SKILL.md and its references/ listing
  2. Read eval definition from lab/eval/evals/{skill}.json
  3. Check ideas.md for deferred ideas about this skill
  4. Check recent journal entries for prior failures on this skill (avoid repeats)
  5. Consult ${CLAUDE_SKILL_DIR}/references/mutation-strategies.md
  6. Propose exactly ONE change targeting the failing checks

Step 4: Apply + Evaluate

  1. Apply the mutation via Edit tool
  2. Run: python3 lab/autoresearch/scripts/run-iteration.py eval <skill-name>
  3. Parse JSON → check verdict field

Step 5: Keep or Revert

If verdict is KEEP:

python3 lab/autoresearch/scripts/run-iteration.py keep <skill> <dim> <old> <new> \
  --desc "..." --asi '{"hypothesis": "...", "mechanism": "..."}'

If verdict is REVERT:

python3 lab/autoresearch/scripts/run-iteration.py revert <skill> <dim> <old> <new> \
  --desc "..." --asi '{"hypothesis": "...", "regression": "...", "avoid": "..."}'

Step 6: Ideas Backlog

If during analysis you discovered a promising optimization you can't act on now:

  • Append it to lab/autoresearch/ideas.md as a bullet
  • On next resume: prune stale/tried ideas, experiment with the rest

Step 7: Continue or Stop

  • All targets >= 0.95? Print "AUTORESEARCH_COMPLETE"
  • Max iterations reached? Print "AUTORESEARCH_COMPLETE"
  • 50 consecutive discards? Print "AUTORESEARCH_STUCK"
  • Otherwise: immediately start Step 1 again

References

  • ${CLAUDE_SKILL_DIR}/references/mutation-strategies.md — mutation type catalog
  • ${CLAUDE_SKILL_DIR}/references/state-management.md — git protocol, journaling
  • lab/autoresearch/program.md — research agenda (read every iteration)
Files (claude-elixir-phoenix)
  • references
    • mutation-strategies.md 2.4 KB
      # Mutation Strategies
      
      ## Mutation Types
      
      ### compress_section
      
      Remove verbosity while preserving meaning. Target sections exceeding 40 lines.
      **When**: conciseness dimension fails on `max_section_lines`
      
      ### add_iron_law
      
      Add a missing Iron Law from patterns observed in other skills.
      **When**: safety or completeness fails on `has_iron_laws` with low count
      
      ### rewrite_description
      
      Improve frontmatter description for better auto-triggering.
      **When**: triggering fails on `description_keywords` or `description_length`
      **Keywords**: elixir, phoenix, liveview, ecto, oban, genserver, test, deploy, security, etc.
      
      ### move_to_references
      
      Extract detailed content from SKILL.md to references/*.md.
      **When**: conciseness fails on `line_count`
      **Rule**: Iron Laws MUST stay in SKILL.md (agents may not read references/)
      
      ### fix_stale_reference
      
      Update a cross-reference that points to a renamed, moved, or removed file/skill/agent.
      **When**: accuracy fails on `valid_skill_refs`, `valid_agent_refs`, or `valid_file_refs`
      
      ### add_missing_section
      
      Add a required section that doesn't exist.
      **When**: completeness fails on `section_exists`
      
      ### add_table
      
      Convert prose to a decision table for quick scanning.
      **When**: conciseness or completeness — section is verbose but information-dense
      
      ### add_prohibition
      
      Add explicit NEVER/MUST NOT/DO NOT rules.
      **When**: safety fails on content_present for prohibition patterns
      
      ### add_description_structure
      
      Add "Use when..." component to description.
      **When**: specificity fails on `description_structure`
      
      ### add_examples
      
      Add code blocks with concrete patterns.
      **When**: specificity fails on `has_examples`
      
      ### increase_action_density
      
      Convert theory/context into imperative instructions.
      **When**: clarity fails on `action_density`
      
      ## ReflexiCoder Pattern (Post-Discard Analysis)
      
      After a discard, BEFORE proposing the next mutation on the same skill:
      
      1. Read the discarded mutation from results.tsv
      2. Re-run the scorer to see exactly which checks got WORSE
      3. Understand the causal chain: what did the mutation break?
      4. Propose a mutation that achieves the same goal without the regression
      
      ## Strategy Selection
      
      | Strategy | Best For | Convergence |
      |----------|----------|-------------|
      | `targeted` | Known weak spots, early optimization | Fast on low-hanging fruit |
      | `sweep` | Broad improvement, initial pass | Steady, predictable |
      | `random` | Escaping local optima, exploration | Slow but diverse |
      
    • state-management.md 1.1 KB
      # State Management
      
      ## Git Protocol
      
      ### Branch
      
      Create a dedicated branch before starting: `git checkout -b autoresearch/sweep-{date}`
      Branch HEAD = current best. Only improvements committed. Failed experiments reverted.
      
      ### Per-Iteration Protocol
      
      1. Apply mutation to working tree
      2. Run eval
      3. If KEEP: `git add + git commit -m "autoresearch: {skill} {dim} {old}->{new}"`
      4. If DISCARD: `git checkout -- plugins/elixir-phoenix/skills/{skill}/`
      
      ### Merge Protocol (Human)
      
      ```bash
      git log --oneline autoresearch/sweep-{date}
      git diff main..autoresearch/sweep-{date} --stat
      git checkout main && git merge autoresearch/sweep-{date}
      ```
      
      ## Results Journal
      
      **Path**: `lab/autoresearch/results.tsv` (gitignored)
      
      ### Header
      
      ```
      iteration skill dimension mutation_type old_composite new_composite kept timestamp description
      ```
      
      ## Session State
      
      **Path**: `lab/autoresearch/autoresearch.md` (gitignored)
      
      Written after every iteration. Read at start of every iteration.
      Bridges across context resets (/loop restarts).
      
      Contains: iteration number, per-skill best scores, consecutive discards, stuck skills, last mutation.
      
  • scripts
    • checks.sh 3.2 KB
      #!/bin/bash
      # Structural checks that the scorer doesn't cover.
      # Run AFTER eval passes. If checks fail, revert regardless of score.
      # Usage: checks.sh <skill-name>
      #
      # Inspired by pi-autoresearch's autoresearch.checks.sh pattern:
      # catches things the metric can't measure.
      
      set -euo pipefail
      
      SKILL_NAME="$1"
      SKILL_DIR="plugins/elixir-phoenix/skills/${SKILL_NAME}"
      SKILL_FILE="${SKILL_DIR}/SKILL.md"
      PROJECT_ROOT="$(cd "$(dirname "$0")/../../.." && pwd)"
      
      cd "$PROJECT_ROOT"
      
      if [ ! -f "$SKILL_FILE" ]; then
          echo "FAIL: Skill file not found: $SKILL_FILE"
          exit 1
      fi
      
      ERRORS=0
      
      # 1. Markdown lint (if available)
      if command -v npx &>/dev/null && [ -f "node_modules/.bin/markdownlint" ]; then
          if ! npx markdownlint "$SKILL_FILE" --quiet 2>/dev/null; then
              echo "FAIL: markdownlint errors in $SKILL_FILE"
              ERRORS=$((ERRORS + 1))
          fi
      fi
      
      # 2. YAML frontmatter must parse
      python3 -c "
      import yaml
      with open('$SKILL_FILE') as f:
          content = f.read()
      if not content.startswith('---'):
          raise ValueError('No frontmatter')
      end = content.find('---', 3)
      fm = yaml.safe_load(content[3:end])
      assert fm.get('name'), 'Missing name'
      assert fm.get('description'), 'Missing description'
      " 2>/dev/null || {
          echo "FAIL: YAML frontmatter broken in $SKILL_FILE"
          ERRORS=$((ERRORS + 1))
      }
      
      # 3. File size under hard limit (command skills: 185, reference: 150, orchestrators: 535)
      LINES=$(wc -l < "$SKILL_FILE")
      if [ "$LINES" -gt 535 ]; then
          echo "FAIL: $SKILL_FILE is $LINES lines (hard limit: 535)"
          ERRORS=$((ERRORS + 1))
      fi
      
      # 4. All reference files mentioned in SKILL.md exist
      python3 -c "
      import re, os
      with open('$SKILL_FILE') as f:
          content = f.read()
      refs = re.findall(r'CLAUDE_SKILL_DIR\}?/references/([\w-]+\.md)', content)
      missing = [r for r in refs if not os.path.isfile('$SKILL_DIR/references/' + r)]
      if missing:
          print(f'FAIL: Missing reference files: {missing}')
          exit(1)
      " 2>/dev/null || {
          ERRORS=$((ERRORS + 1))
      }
      
      # 5. No conflict markers
      if grep -qE '<<<<<<|======|>>>>>>' "$SKILL_FILE" 2>/dev/null; then
          echo "FAIL: Conflict markers in $SKILL_FILE"
          ERRORS=$((ERRORS + 1))
      fi
      
      # 6. No empty sections (heading followed immediately by another heading)
      python3 -c "
      import re
      with open('$SKILL_FILE') as f:
          lines = f.readlines()
      prev_was_heading = False
      for i, line in enumerate(lines):
          is_heading = line.startswith('## ') or line.startswith('### ')
          if is_heading and prev_was_heading:
              print(f'FAIL: Empty section at line {i}: {lines[i-1].strip()}')
              exit(1)
          prev_was_heading = is_heading and not lines[i+1].strip() if i+1 < len(lines) else is_heading
      " 2>/dev/null || {
          ERRORS=$((ERRORS + 1))
      }
      
      # 7. Protected-section invariant: Iron Laws are append-only (SkillOpt fast/slow split).
      #    A mutation may ADD an Iron Law but never delete or reword an existing one.
      #    Compares the working tree against git HEAD (last kept state). New skills not
      #    yet in HEAD are skipped — there is nothing to protect.
      PROTECT_OUT=$(python3 "$(dirname "$0")/protected_sections.py" "$SKILL_FILE" 2>/dev/null) || {
          echo "${PROTECT_OUT:-FAIL: protected Iron Law eroded}"
          ERRORS=$((ERRORS + 1))
      }
      
      if [ "$ERRORS" -gt 0 ]; then
          echo "CHECKS FAILED: $ERRORS errors"
          exit 1
      fi
      
      echo "CHECKS PASSED"
      exit 0
      
    • protected_sections.py 3.4 KB
      #!/usr/bin/env python3
      """Protected-section invariant for the autoresearch loop.
      
      SkillOpt (arXiv 2605.23904) found that a *structural guarantee* preventing fast
      edits from overwriting slow lessons is worth ~22 points on SpreadsheetBench.
      Their "fast vs slow state" split maps cleanly onto our skills:
      
        - slow state  = Iron Laws  (hard-won prohibitions; must never erode)
        - fast state  = everything else (patterns, examples, prose the loop tunes)
      
      Without this guard, an optimizer chasing the conciseness dimension can legally
      delete or reword an Iron Law, because safety is only a *soft* score penalty and
      our anti-thrash rule keeps a mutation when it helps one dimension more than it
      hurts another. This module makes Iron Laws **append-only**: the loop may add a
      law, but deleting or rewording an existing one forces a REVERT.
      
      The pure function lives here so it is unit-testable; ``checks.sh`` shells out to
      ``__main__`` and turns a non-empty "removed" set into a hard checks failure.
      """
      
      from __future__ import annotations
      
      import re
      import subprocess
      import sys
      
      # Matches the Iron Laws H2 heading line (any suffix, e.g. "— Never Violate These")
      # and captures the section body up to the next H2 heading or end of file.
      _SECTION_RE = re.compile(r"^##\s+Iron Laws[^\n]*\n(.*?)(?=^##\s|\Z)", re.M | re.S)
      
      # A numbered list item: "1. ...", "12. ...". Captures the law text after the prefix.
      _ITEM_RE = re.compile(r"^\s*\d+\.\s+(.*\S)\s*$")
      
      
      def iron_laws(text: str) -> set[str]:
          """Return the set of Iron Law item texts in ``text``.
      
          The leading ``N. `` number prefix is stripped and whitespace normalized, so
          renumbering (inserting a law in the middle) is NOT treated as erosion — only
          deleting or rewording an existing law changes the set.
          """
          match = _SECTION_RE.search(text)
          if not match:
              return set()
          laws: set[str] = set()
          for line in match.group(1).splitlines():
              item = _ITEM_RE.match(line)
              if item:
                  laws.add(" ".join(item.group(1).split()))
          return laws
      
      
      def removed_iron_laws(head_text: str, new_text: str) -> set[str]:
          """Iron Laws present in ``head_text`` but missing from ``new_text``.
      
          A non-empty result means the mutation eroded the protected section.
          """
          return iron_laws(head_text) - iron_laws(new_text)
      
      
      def _git_head_version(path: str) -> str | None:
          """Return the HEAD (last kept) contents of ``path``, or None if untracked."""
          try:
              return subprocess.run(
                  ["git", "show", f"HEAD:{path}"],
                  capture_output=True,
                  text=True,
                  check=True,
              ).stdout
          except subprocess.CalledProcessError:
              return None  # New file not yet in HEAD — nothing to protect.
      
      
      def check_file(path: str) -> set[str]:
          """Compare working-tree ``path`` against HEAD; return eroded Iron Laws."""
          head_text = _git_head_version(path)
          if head_text is None:
              return set()
          with open(path) as f:
              new_text = f.read()
          return removed_iron_laws(head_text, new_text)
      
      
      def main(argv: list[str]) -> int:
          if len(argv) != 2:
              print("usage: protected_sections.py <skill-file>", file=sys.stderr)
              return 2
          removed = check_file(argv[1])
          if removed:
              sample = sorted(removed)[:2]
              print(f"FAIL: protected Iron Law(s) deleted or reworded: {sample}")
              return 1
          return 0
      
      
      if __name__ == "__main__":
          raise SystemExit(main(sys.argv))
      
    • run-iteration.py 25 KB
      #!/usr/bin/env python3
      """Autoresearch iteration wrapper — single command replaces 5+ manual steps.
      
      Inspired by pi-autoresearch's typed MCP tools (init/run/log), adapted for
      our deterministic scorer + git state machine.
      
      Usage:
          # Score a skill and compare against baseline
          python3 lab/autoresearch/scripts/run-iteration.py score <skill-name>
      
          # Evaluate mutation: score + checks + compare + decide keep/revert
          python3 lab/autoresearch/scripts/run-iteration.py eval <skill-name> [--hypothesis "..."]
      
          # Keep current changes (commit + journal + state update)
          python3 lab/autoresearch/scripts/run-iteration.py keep <skill-name> <dimension> <old> <new> --desc "..." [--asi '{"key": "val"}']
      
          # Revert current changes (git checkout + journal)
          python3 lab/autoresearch/scripts/run-iteration.py revert <skill-name> <dimension> <old> <new> --desc "..." [--asi '{"key": "val"}']
      
          # Find the weakest skill+dimension
          python3 lab/autoresearch/scripts/run-iteration.py target [--strategy targeted|sweep|random]
      
          # Show current state summary
          python3 lab/autoresearch/scripts/run-iteration.py status
      """
      
      import argparse
      import json
      import os
      import subprocess
      import sys
      from datetime import datetime, timezone
      
      # Add project root to path
      PROJECT_ROOT = os.path.dirname(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))))
      sys.path.insert(0, PROJECT_ROOT)
      
      from lab.eval.scorer import score_skill, find_eval, find_all_skills, PLUGIN_ROOT
      from lab.eval.schemas import EvalDefinition
      from lab.eval.triggers.deviation_classifier import RESULTS_DIR as TRIGGER_RESULTS_DIR
      from lab.eval.triggers.deviation_types import DeviationType, TriggerDeviation
      
      RESULTS_FILE = os.path.join(PROJECT_ROOT, "lab", "autoresearch", "results.jsonl")
      STATE_FILE = os.path.join(PROJECT_ROOT, "lab", "autoresearch", "autoresearch.md")
      CHECKS_SCRIPT = os.path.join(PROJECT_ROOT, "lab", "autoresearch", "scripts", "checks.sh")
      
      
      def load_deviations(skill_name: str) -> list[TriggerDeviation]:
          """Read cached trigger deviations for one skill. Returns [] if no cache."""
          path = os.path.join(TRIGGER_RESULTS_DIR, f"{skill_name}.json")
          if not os.path.isfile(path):
              return []
          with open(path) as f:
              data = json.load(f)
          return [TriggerDeviation.from_dict(d) for d in data.get("deviations", [])]
      
      
      def pick_dominant_deviation(deviations: list[TriggerDeviation]) -> TriggerDeviation | None:
          """Pick the deviation that should drive the next mutation.
          Priority: high severity > medium > low, then most common type."""
          if not deviations:
              return None
          high = [d for d in deviations if d.severity.value == "high"]
          pool = high or deviations
          # Group by type, pick most common
          from collections import Counter
          type_counts = Counter(d.deviation_type for d in pool)
          dominant_type, _ = type_counts.most_common(1)[0]
          # Return first deviation of that type (preserves competing_skill / matched_keywords)
          for d in pool:
              if d.deviation_type == dominant_type:
                  return d
          return pool[0]
      
      
      def score_one(skill_name: str) -> dict:
          """Score a single skill, return full result dict."""
          skill_path = os.path.join(PLUGIN_ROOT, "skills", skill_name, "SKILL.md")
          eval_path = find_eval(skill_name)
          eval_def = EvalDefinition.from_file(eval_path) if eval_path else None
          result = score_skill(skill_path, eval_def)
          return result.to_dict()
      
      
      def score_all() -> dict[str, dict]:
          """Score all skills, return {name: result_dict}."""
          results = {}
          for skill_path in find_all_skills():
              name = os.path.basename(os.path.dirname(skill_path))
              eval_path = find_eval(name)
              eval_def = EvalDefinition.from_file(eval_path) if eval_path else None
              result = score_skill(skill_path, eval_def)
              results[name] = result.to_dict()
          return results
      
      
      def run_checks(skill_name: str) -> tuple[bool, str]:
          """Run structural checks. Returns (passed, output)."""
          if not os.path.isfile(CHECKS_SCRIPT):
              return True, "No checks.sh found — skipping"
          try:
              result = subprocess.run(
                  ["bash", CHECKS_SCRIPT, skill_name],
                  capture_output=True, text=True, timeout=30, cwd=PROJECT_ROOT
              )
              output = (result.stdout + result.stderr).strip()
              return result.returncode == 0, output
          except subprocess.TimeoutExpired:
              return False, "CHECKS TIMEOUT after 30s"
          except Exception as e:
              return False, f"CHECKS ERROR: {e}"
      
      
      def append_journal(entry: dict):
          """Append one JSONL entry to results file."""
          with open(RESULTS_FILE, "a") as f:
              f.write(json.dumps(entry) + "\n")
      
      
      def read_journal_tail(n: int = 5) -> list[dict]:
          """Read last N journal entries."""
          if not os.path.isfile(RESULTS_FILE):
              return []
          entries = []
          with open(RESULTS_FILE) as f:
              for line in f:
                  line = line.strip()
                  if line:
                      try:
                          entries.append(json.loads(line))
                      except json.JSONDecodeError:
                          pass
          return entries[-n:]
      
      
      def get_iteration_count() -> int:
          """Count total iterations from journal."""
          if not os.path.isfile(RESULTS_FILE):
              return 0
          count = 0
          with open(RESULTS_FILE) as f:
              for line in f:
                  if line.strip():
                      count += 1
          return count
      
      
      def git_commit(skill_name: str, message: str) -> bool:
          """Git add + commit for a skill directory."""
          skill_dir = os.path.join("plugins", "elixir-phoenix", "skills", skill_name)
          try:
              subprocess.run(["git", "add", skill_dir], check=True, capture_output=True, cwd=PROJECT_ROOT)
              subprocess.run(["git", "commit", "-m", message], check=True, capture_output=True, cwd=PROJECT_ROOT)
              return True
          except subprocess.CalledProcessError as e:
              print(f"GIT ERROR: {e.stderr.decode()}", file=sys.stderr)
              return False
      
      
      def git_revert(skill_name: str) -> bool:
          """Git checkout to revert skill changes."""
          skill_dir = os.path.join("plugins", "elixir-phoenix", "skills", skill_name)
          try:
              subprocess.run(["git", "checkout", "--", skill_dir], check=True, capture_output=True, cwd=PROJECT_ROOT)
              return True
          except subprocess.CalledProcessError as e:
              print(f"GIT REVERT ERROR: {e.stderr.decode()}", file=sys.stderr)
              return False
      
      
      def find_weakest(strategy: str = "targeted") -> dict | None:
          """Find weakest skill+dimension. Returns {skill, dimension, score, composite, mode}.
      
          Mode is "structural" for dimension scores < 1.0, or "tournament" when all
          structural dimensions are 1.0 but trigger accuracy is below threshold.
          Returns None when everything is perfect.
          """
          all_scores = score_all()
      
          if strategy == "targeted":
              # First check for structural weaknesses
              weakest = None
              for name, data in all_scores.items():
                  for dim_name, dim_data in data["dimensions"].items():
                      if dim_data["score"] < 1.0:
                          if weakest is None or dim_data["score"] < weakest["dim_score"]:
                              weakest = {
                                  "skill": name,
                                  "dimension": dim_name,
                                  "dim_score": dim_data["score"],
                                  "composite": data["composite"],
                                  "mode": "structural",
                                  "failing_checks": [
                                      a["desc"] for a in dim_data["assertions"] if not a["passed"]
                                  ],
                              }
              if weakest:
                  return weakest
      
              # All structural perfect — check trigger accuracy
              try:
                  from lab.tournament.description_tournament import find_weak_skills
                  weak = find_weak_skills(threshold=0.75)
                  if weak:
                      skill_name, accuracy = weak[0]  # Worst first
                      target = {
                          "skill": skill_name,
                          "dimension": "trigger_accuracy",
                          "dim_score": accuracy,
                          "composite": 1.0,
                          "mode": "tournament",
                          "failing_checks": [f"trigger accuracy {accuracy:.0%} < 75%"],
                      }
                      # Augment with deviation taxonomy — drives mutation strategy
                      deviations = load_deviations(skill_name)
                      dominant = pick_dominant_deviation(deviations)
                      if dominant:
                          target["deviation_type"] = dominant.deviation_type.value
                          target["severity"] = dominant.severity.value
                          target["fix_hint"] = dominant.fix_hint
                          target["competing_skill"] = dominant.competing_skill
                          target["matched_keywords"] = list(dominant.matched_keywords)
                          target["strategy"] = _strategy_for(dominant.deviation_type)
                          target["deviation_count"] = len(deviations)
                      return target
              except ImportError:
                  pass
      
              return None  # All perfect
      
          elif strategy == "sweep":
              for name in sorted(all_scores.keys()):
                  data = all_scores[name]
                  if data["composite"] < 1.0:
                      # Find weakest dimension for this skill
                      worst_dim = min(
                          data["dimensions"].items(),
                          key=lambda x: x[1]["score"]
                      )
                      return {
                          "skill": name,
                          "dimension": worst_dim[0],
                          "dim_score": worst_dim[1]["score"],
                          "composite": data["composite"],
                          "mode": "structural",
                          "failing_checks": [
                              a["desc"] for a in worst_dim[1]["assertions"] if not a["passed"]
                          ],
                      }
              return None  # All perfect
      
          elif strategy == "random":
              import random
              below = [(n, d) for n, d in all_scores.items() if d["composite"] < 1.0]
              if not below:
                  return None
              name, data = random.choice(below)
              dims_below = [(dn, dd) for dn, dd in data["dimensions"].items() if dd["score"] < 1.0]
              if not dims_below:
                  return None
              dim_name, dim_data = random.choice(dims_below)
              return {
                  "skill": name,
                  "dimension": dim_name,
                  "dim_score": dim_data["score"],
                  "composite": data["composite"],
                  "mode": "structural",
                  "failing_checks": [a["desc"] for a in dim_data["assertions"] if not a["passed"]],
              }
      
          return None
      
      
      def cmd_score(args):
          """Score a skill and print result."""
          result = score_one(args.skill)
          print(json.dumps({
              "skill": args.skill,
              "composite": result["composite"],
              "dimensions": {k: v["score"] for k, v in result["dimensions"].items()},
          }))
      
      
      def cmd_eval(args):
          """Evaluate mutation: score + checks + compare against previous."""
          # Score
          result = score_one(args.skill)
          new_composite = result["composite"]
      
          # Run checks
          checks_passed, checks_output = run_checks(args.skill)
      
          # Read previous best from journal
          journal = read_journal_tail(50)
          prev_best = 0.0
          for entry in journal:
              if entry.get("skill") == args.skill and entry.get("kept"):
                  prev_best = max(prev_best, entry.get("new_composite", 0))
          if prev_best == 0.0:
              # No prior journal entry — use current score as baseline comparison
              # (caller should have scored before mutation)
              prev_best = new_composite
      
          improved = new_composite >= prev_best
          verdict = "KEEP" if improved and checks_passed else "REVERT"
          reason = ""
          if not checks_passed:
              verdict = "REVERT"
              reason = f"checks failed: {checks_output}"
          elif not improved:
              reason = f"regression: {prev_best:.3f} -> {new_composite:.3f}"
      
          output = {
              "skill": args.skill,
              "composite": new_composite,
              "previous_best": prev_best,
              "delta": round(new_composite - prev_best, 4),
              "checks_passed": checks_passed,
              "checks_output": checks_output if not checks_passed else "PASSED",
              "verdict": verdict,
              "reason": reason,
              "dimensions": {k: round(v["score"], 4) for k, v in result["dimensions"].items()},
              "failing": [
                  {"dim": k, "check": a["desc"], "evidence": a["evidence"][:80]}
                  for k, v in result["dimensions"].items()
                  for a in v["assertions"] if not a["passed"]
              ],
          }
          print(json.dumps(output))
      
      
      def cmd_keep(args):
          """Keep mutation: commit + journal + state."""
          iteration = get_iteration_count() + 1
          msg = f"autoresearch: {args.skill} {args.dimension} {args.old}->{args.new}"
      
          if not git_commit(args.skill, msg):
              print(json.dumps({"error": "git commit failed"}))
              sys.exit(1)
      
          asi = {}
          if args.asi:
              try:
                  asi = json.loads(args.asi)
              except json.JSONDecodeError:
                  asi = {"raw": args.asi}
      
          entry = {
              "iteration": iteration,
              "skill": args.skill,
              "dimension": args.dimension,
              "old_composite": float(args.old),
              "new_composite": float(args.new),
              "kept": True,
              "timestamp": datetime.now(timezone.utc).isoformat(),
              "description": args.desc,
              "asi": asi,
          }
          if args.deviation_type:
              entry["deviation_type"] = args.deviation_type
          if args.strategy_applied:
              entry["strategy_applied"] = args.strategy_applied
          append_journal(entry)
          print(json.dumps({"status": "kept", "iteration": iteration, "commit_msg": msg}))
      
      
      def cmd_revert(args):
          """Revert mutation: git checkout + journal."""
          iteration = get_iteration_count() + 1
      
          git_revert(args.skill)
      
          asi = {}
          if args.asi:
              try:
                  asi = json.loads(args.asi)
              except json.JSONDecodeError:
                  asi = {"raw": args.asi}
      
          entry = {
              "iteration": iteration,
              "skill": args.skill,
              "dimension": args.dimension,
              "old_composite": float(args.old),
              "new_composite": float(args.new),
              "kept": False,
              "timestamp": datetime.now(timezone.utc).isoformat(),
              "description": args.desc,
              "asi": asi,
          }
          if args.deviation_type:
              entry["deviation_type"] = args.deviation_type
          if args.strategy_applied:
              entry["strategy_applied"] = args.strategy_applied
          append_journal(entry)
          print(json.dumps({"status": "reverted", "iteration": iteration}))
      
      
      def cmd_target(args):
          """Find weakest skill+dimension."""
          if getattr(args, "check_retention", False):
              from lab.autoresearch.retention import is_converged
              if is_converged(streak=args.retention_streak):
                  print(json.dumps({
                      "status": "retention_converged",
                      "message": "AUTORESEARCH_COMPLETE_VIA_RETENTION",
                  }))
                  return
          target = find_weakest(args.strategy)
          if target is None:
              print(json.dumps({"status": "all_perfect", "message": "AUTORESEARCH_COMPLETE"}))
          else:
              print(json.dumps(target))
      
      
      def cmd_retention(args):
          """Compute Retention@K, append to ledger, print status."""
          from lab.autoresearch.retention import check_retention
          result = check_retention(
              k=args.k,
              threshold=args.threshold,
              streak=args.streak,
          )
          print(json.dumps(result, indent=2 if args.pretty else None))
      
      
      # Deviation-type → mutation strategy. Phase 4b dispatch table.
      _STRATEGIES: dict[DeviationType, str] = {
          DeviationType.MISSING_KEYWORD: "inject_keywords",
          DeviationType.SCOPE_TOO_NARROW: "synonym_expand",
          DeviationType.DESCRIPTION_OVERLAP: "disambiguate",
          DeviationType.USE_CASE_GAP: "add_use_when",
          DeviationType.SCOPE_TOO_BROAD: "tighten_scope",
          DeviationType.UNKNOWN: "random_rewrite",
      }
      
      
      def _strategy_for(deviation_type: DeviationType) -> str:
          return _STRATEGIES.get(deviation_type, "random_rewrite")
      
      
      def cmd_deviations(args):
          """Inspect classified trigger deviations across cached results."""
          from collections import Counter
      
          if args.skill:
              devs = load_deviations(args.skill)
              if not devs:
                  print(json.dumps({"skill": args.skill, "deviations": []}))
                  return
              dominant = pick_dominant_deviation(devs)
              out = {
                  "skill": args.skill,
                  "count": len(devs),
                  "by_type": dict(Counter(d.deviation_type.value for d in devs)),
                  "by_severity": dict(Counter(d.severity.value for d in devs)),
                  "deviations": [d.to_dict() for d in devs],
                  "dominant": dominant.to_dict() if dominant else None,
                  "strategy": _strategy_for(dominant.deviation_type) if dominant else None,
              }
              print(json.dumps(out, indent=2 if args.pretty else None))
              return
      
          # Aggregate across all skills
          all_devs: dict[str, list[TriggerDeviation]] = {}
          if not os.path.isdir(TRIGGER_RESULTS_DIR):
              print(json.dumps({"error": "no_trigger_cache"}))
              sys.exit(1)
          for fname in sorted(os.listdir(TRIGGER_RESULTS_DIR)):
              if not fname.endswith(".json") or fname.startswith("_"):
                  continue
              skill = fname[:-5]
              devs = load_deviations(skill)
              if devs:
                  all_devs[skill] = devs
      
          type_counter: Counter = Counter()
          severity_counter: Counter = Counter()
          strategy_counter: Counter = Counter()
          for devs in all_devs.values():
              for d in devs:
                  type_counter[d.deviation_type.value] += 1
                  severity_counter[d.severity.value] += 1
                  strategy_counter[_strategy_for(d.deviation_type)] += 1
      
          output = {
              "skills_with_deviations": len(all_devs),
              "total_deviations": sum(len(d) for d in all_devs.values()),
              "by_type": dict(type_counter),
              "by_severity": dict(severity_counter),
              "by_strategy": dict(strategy_counter),
          }
          if args.histogram:
              total = output["total_deviations"]
              print(f"Deviation distribution across {output['skills_with_deviations']} skills "
                    f"({total} total):")
              for dtype in DeviationType:
                  n = type_counter.get(dtype.value, 0)
                  pct = (n / total * 100) if total else 0
                  strat = _STRATEGIES.get(dtype, "?")
                  print(f"  {dtype.value:22s} {n:4d}  ({pct:5.1f}%)  → {strat}")
          else:
              print(json.dumps(output, indent=2 if args.pretty else None))
      
      
      def cmd_tournament(args):
          """Run description tournament for a skill."""
          from lab.tournament.description_tournament import (
              load_all_descriptions,
              load_trigger_prompts,
              run_tournament,
          )
          from lab.tournament.config import load_config
      
          skill_name = args.skill
          config = load_config()
      
          # Gate: structural score must be 1.000
          result = score_one(skill_name)
          if result["composite"] < 0.999:
              print(json.dumps({
                  "error": "structural_gate_failed",
                  "skill": skill_name,
                  "composite": result["composite"],
                  "message": "Skill must pass structural eval (1.000) before tournament",
              }))
              sys.exit(1)
      
          # Load tournament inputs
          all_descriptions = load_all_descriptions()
          trigger_prompts = load_trigger_prompts(skill_name, split="train")
          if not trigger_prompts:
              print(json.dumps({"error": "no_trigger_file", "skill": skill_name}))
              sys.exit(1)
      
          # Run tournament
          tournament_result = run_tournament(skill_name, all_descriptions, trigger_prompts, config)
      
          # Journal the result
          iteration = get_iteration_count() + 1
          entry = {
              "iteration": iteration,
              "skill": skill_name,
              "dimension": "trigger_accuracy",
              "old_composite": result["composite"],
              "new_composite": result["composite"],  # structural unchanged
              "kept": tournament_result["changed"],
              "timestamp": datetime.now(timezone.utc).isoformat(),
              "description": f"tournament: {tournament_result['passes']} passes, "
                             f"converged={tournament_result['converged']}",
              "asi": {
                  "type": "tournament",
                  "passes": tournament_result["passes"],
                  "converged": tournament_result["converged"],
                  "description_before": tournament_result["description_before"],
                  "description_after": tournament_result["description_after"],
                  "history": tournament_result["history_summary"],
              },
          }
          append_journal(entry)
      
          print(json.dumps(tournament_result, indent=2))
      
      
      def cmd_status(args):
          """Show current state summary."""
          all_scores = score_all()
          perfect = sum(1 for v in all_scores.values() if v["composite"] >= 0.999)
          avg = sum(v["composite"] for v in all_scores.values()) / len(all_scores) if all_scores else 0
          iterations = get_iteration_count()
          journal = read_journal_tail(5)
      
          below = {k: v["composite"] for k, v in all_scores.items() if v["composite"] < 0.95}
      
          output = {
              "total_skills": len(all_scores),
              "perfect": perfect,
              "average": round(avg, 4),
              "iterations": iterations,
              "below_threshold": below,
              "recent": [
                  {"skill": e["skill"], "kept": e["kept"], "desc": e.get("description", "")[:60]}
                  for e in journal
              ],
          }
          print(json.dumps(output, indent=2))
      
      
      def main():
          parser = argparse.ArgumentParser(description="Autoresearch iteration wrapper")
          sub = parser.add_subparsers(dest="command", required=True)
      
          # score
          p_score = sub.add_parser("score", help="Score a skill")
          p_score.add_argument("skill")
          p_score.set_defaults(func=cmd_score)
      
          # eval
          p_eval = sub.add_parser("eval", help="Evaluate mutation (score + checks + compare)")
          p_eval.add_argument("skill")
          p_eval.add_argument("--hypothesis", default="")
          p_eval.set_defaults(func=cmd_eval)
      
          # keep
          p_keep = sub.add_parser("keep", help="Keep mutation (commit + journal)")
          p_keep.add_argument("skill")
          p_keep.add_argument("dimension")
          p_keep.add_argument("old")
          p_keep.add_argument("new")
          p_keep.add_argument("--desc", required=True)
          p_keep.add_argument("--asi", default="{}")
          p_keep.add_argument("--deviation-type", default="",
                              help="Deviation type that drove this mutation (Phase 4b)")
          p_keep.add_argument("--strategy-applied", default="",
                              help="Mutation strategy applied (e.g., inject_keywords)")
          p_keep.set_defaults(func=cmd_keep)
      
          # revert
          p_revert = sub.add_parser("revert", help="Revert mutation (git checkout + journal)")
          p_revert.add_argument("skill")
          p_revert.add_argument("dimension")
          p_revert.add_argument("old")
          p_revert.add_argument("new")
          p_revert.add_argument("--desc", required=True)
          p_revert.add_argument("--asi", default="{}")
          p_revert.add_argument("--deviation-type", default="",
                                help="Deviation type that drove this mutation (Phase 4b)")
          p_revert.add_argument("--strategy-applied", default="",
                                help="Mutation strategy applied (e.g., inject_keywords)")
          p_revert.set_defaults(func=cmd_revert)
      
          # tournament
          p_tournament = sub.add_parser("tournament", help="Run description tournament for a skill")
          p_tournament.add_argument("skill")
          p_tournament.set_defaults(func=cmd_tournament)
      
          # target
          p_target = sub.add_parser("target", help="Find weakest skill+dimension")
          p_target.add_argument("--strategy", default="targeted", choices=["targeted", "sweep", "random"])
          p_target.add_argument("--check-retention", action="store_true",
                                help="Short-circuit to retention_converged when Retention@K stable")
          p_target.add_argument("--retention-streak", type=int, default=2,
                                help="Consecutive iterations above threshold required (default: 2)")
          p_target.set_defaults(func=cmd_target)
      
          # retention
          p_retention = sub.add_parser("retention",
                                        help="Compute Retention@K convergence over top-K trigger accuracy")
          p_retention.add_argument("--k", type=int, default=30, help="Top-K size (default: 30)")
          p_retention.add_argument("--threshold", type=float, default=0.9,
                                    help="Retention threshold (default: 0.9)")
          p_retention.add_argument("--streak", type=int, default=2,
                                    help="Consecutive iterations required for convergence (default: 2)")
          p_retention.add_argument("--pretty", action="store_true", help="Pretty-print JSON")
          p_retention.set_defaults(func=cmd_retention)
      
          # deviations
          p_dev = sub.add_parser("deviations", help="Inspect classified trigger deviations")
          p_dev.add_argument("--skill", help="Show deviations for one skill")
          p_dev.add_argument("--histogram", action="store_true", help="Print type histogram with strategies")
          p_dev.add_argument("--pretty", action="store_true", help="Pretty-print JSON")
          p_dev.set_defaults(func=cmd_deviations)
      
          # status
          p_status = sub.add_parser("status", help="Show current state")
          p_status.set_defaults(func=cmd_status)
      
          args = parser.parse_args()
          args.func(args)
      
      
      if __name__ == "__main__":
          main()
      
    • score-skill.py 1.2 KB
      #!/usr/bin/env python3
      """Thin wrapper around lab/eval/scorer.py for autoresearch use.
      
      Usage:
          python3 lab/autoresearch/scripts/score-skill.py plugins/elixir-phoenix/skills/review/SKILL.md
          python3 lab/autoresearch/scripts/score-skill.py plugins/elixir-phoenix/skills/review/SKILL.md lab/eval/evals/review.json
      
      Outputs single JSON line to stdout. Errors to stderr.
      """
      
      import os
      import sys
      
      # Add project root to path (scripts/ -> autoresearch/ -> lab/ -> project root)
      project_root = os.path.dirname(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))))
      sys.path.insert(0, project_root)
      
      from lab.eval.scorer import score_skill, find_eval
      from lab.eval.schemas import EvalDefinition
      
      
      def main():
          if len(sys.argv) < 2:
              print("Usage: score-skill.py <skill_path> [eval_path]", file=sys.stderr)
              sys.exit(1)
      
          skill_path = sys.argv[1]
          eval_path = sys.argv[2] if len(sys.argv) > 2 else None
      
          if eval_path is None:
              skill_name = os.path.basename(os.path.dirname(skill_path))
              eval_path = find_eval(skill_name)
      
          eval_def = EvalDefinition.from_file(eval_path) if eval_path else None
          score = score_skill(skill_path, eval_def)
          print(score.to_json())
      
      
      if __name__ == "__main__":
          main()
      
  • tests
    • test_deviation_dispatch.py 3.1 KB
      """Tests for Phase 4b — deviation_type → mutation strategy dispatch."""
      
      import importlib.util
      import os
      import sys
      
      PROJECT_ROOT = os.path.dirname(os.path.dirname(os.path.dirname(os.path.dirname(
          os.path.abspath(__file__)))))
      sys.path.insert(0, PROJECT_ROOT)
      
      # run-iteration.py uses a hyphen, so import via spec.
      SCRIPT_PATH = os.path.join(
          PROJECT_ROOT, "lab", "autoresearch", "scripts", "run-iteration.py"
      )
      spec = importlib.util.spec_from_file_location("run_iteration", SCRIPT_PATH)
      run_iteration = importlib.util.module_from_spec(spec)
      spec.loader.exec_module(run_iteration)
      
      from lab.eval.triggers.deviation_types import DeviationType, Severity, TriggerDeviation
      
      
      class TestStrategyDispatch:
          """Each deviation type maps to a distinct mutation strategy."""
      
          def test_missing_keyword_maps_to_inject(self):
              assert run_iteration._strategy_for(DeviationType.MISSING_KEYWORD) == "inject_keywords"
      
          def test_scope_too_narrow_maps_to_synonym(self):
              assert run_iteration._strategy_for(DeviationType.SCOPE_TOO_NARROW) == "synonym_expand"
      
          def test_description_overlap_maps_to_disambiguate(self):
              assert run_iteration._strategy_for(DeviationType.DESCRIPTION_OVERLAP) == "disambiguate"
      
          def test_use_case_gap_maps_to_use_when(self):
              assert run_iteration._strategy_for(DeviationType.USE_CASE_GAP) == "add_use_when"
      
          def test_scope_too_broad_maps_to_tighten(self):
              assert run_iteration._strategy_for(DeviationType.SCOPE_TOO_BROAD) == "tighten_scope"
      
          def test_unknown_falls_back_to_random(self):
              assert run_iteration._strategy_for(DeviationType.UNKNOWN) == "random_rewrite"
      
          def test_every_deviation_type_has_strategy(self):
              for dtype in DeviationType:
                  assert run_iteration._strategy_for(dtype), f"{dtype} has no strategy"
      
      
      class TestPickDominantDeviation:
          def test_returns_none_for_empty(self):
              assert run_iteration.pick_dominant_deviation([]) is None
      
          def test_prefers_high_severity(self):
              devs = [
                  TriggerDeviation("a", [], "low prompt",
                                   DeviationType.SCOPE_TOO_BROAD, Severity.MEDIUM, "h"),
                  TriggerDeviation("a", [], "high prompt",
                                   DeviationType.MISSING_KEYWORD, Severity.HIGH, "h"),
              ]
              dom = run_iteration.pick_dominant_deviation(devs)
              assert dom.severity == Severity.HIGH
      
          def test_prefers_most_common_type(self):
              devs = [
                  TriggerDeviation("a", [], "p1", DeviationType.MISSING_KEYWORD, Severity.HIGH, "h"),
                  TriggerDeviation("a", [], "p2", DeviationType.MISSING_KEYWORD, Severity.HIGH, "h"),
                  TriggerDeviation("a", [], "p3", DeviationType.SCOPE_TOO_NARROW, Severity.HIGH, "h"),
              ]
              dom = run_iteration.pick_dominant_deviation(devs)
              assert dom.deviation_type == DeviationType.MISSING_KEYWORD
      
      
      class TestLoadDeviations:
          def test_returns_empty_list_when_no_cache(self, tmp_path, monkeypatch):
              monkeypatch.setattr(run_iteration, "TRIGGER_RESULTS_DIR", str(tmp_path))
              assert run_iteration.load_deviations("nonexistent") == []
      
    • test_protected_sections.py 3.3 KB
      """Tests for the protected-section (Iron Laws append-only) invariant."""
      
      import os
      import sys
      
      # protected_sections.py lives in scripts/ and is normally run as __main__,
      # so make it importable without turning scripts/ into a package.
      SCRIPTS_DIR = os.path.join(os.path.dirname(os.path.dirname(__file__)), "scripts")
      sys.path.insert(0, SCRIPTS_DIR)
      
      import protected_sections as ps  # noqa: E402
      
      
      HEAD = """\
      ---
      name: security
      ---
      
      # Security
      
      ## Iron Laws — Never Violate These
      
      1. **VALIDATE AT BOUNDARIES** — Never trust client input
      2. **NEVER INTERPOLATE USER INPUT** — Use Ecto's `^` operator
      3. **NO String.to_atom WITH USER INPUT** — Atom exhaustion DoS
      
      ## Quick Patterns
      
      Some prose here.
      """
      
      
      def test_iron_laws_extracts_items():
          laws = ps.iron_laws(HEAD)
          assert len(laws) == 3
          assert "**VALIDATE AT BOUNDARIES** — Never trust client input" in laws
      
      
      def test_iron_laws_empty_when_no_section():
          assert ps.iron_laws("# Skill\n\nNo laws here.\n") == set()
      
      
      def test_iron_laws_stops_at_next_h2():
          # "Some prose here." lives under Quick Patterns, not Iron Laws.
          assert all("prose" not in law for law in ps.iron_laws(HEAD))
      
      
      def test_append_is_allowed():
          new = HEAD.replace(
              "3. **NO String.to_atom WITH USER INPUT** — Atom exhaustion DoS\n",
              "3. **NO String.to_atom WITH USER INPUT** — Atom exhaustion DoS\n"
              "4. **ESCAPE BY DEFAULT** — Never use `raw/1` with untrusted content\n",
          )
          assert ps.removed_iron_laws(HEAD, new) == set()
      
      
      def test_renumber_is_allowed():
          # Insert a law in the middle; existing law texts survive, only prefixes shift.
          new = HEAD.replace(
              "2. **NEVER INTERPOLATE USER INPUT** — Use Ecto's `^` operator\n",
              "2. **NEW MIDDLE LAW** — inserted here\n"
              "3. **NEVER INTERPOLATE USER INPUT** — Use Ecto's `^` operator\n",
          )
          assert ps.removed_iron_laws(HEAD, new) == set()
      
      
      def test_deletion_is_rejected():
          new = HEAD.replace(
              "2. **NEVER INTERPOLATE USER INPUT** — Use Ecto's `^` operator\n", ""
          )
          removed = ps.removed_iron_laws(HEAD, new)
          assert removed == {"**NEVER INTERPOLATE USER INPUT** — Use Ecto's `^` operator"}
      
      
      def test_rewording_is_rejected():
          new = HEAD.replace(
              "2. **NEVER INTERPOLATE USER INPUT** — Use Ecto's `^` operator",
              "2. **INTERPOLATION OK SOMETIMES** — relax the rule",
          )
          assert len(ps.removed_iron_laws(HEAD, new)) == 1
      
      
      def test_check_file_skips_untracked(monkeypatch, tmp_path):
          # File not in HEAD → nothing to protect → no erosion reported.
          monkeypatch.setattr(ps, "_git_head_version", lambda path: None)
          f = tmp_path / "SKILL.md"
          f.write_text("# new skill\n")
          assert ps.check_file(str(f)) == set()
      
      
      def test_main_returns_nonzero_on_erosion(monkeypatch, tmp_path, capsys):
          monkeypatch.setattr(ps, "_git_head_version", lambda path: HEAD)
          f = tmp_path / "SKILL.md"
          f.write_text(HEAD.replace("1. **VALIDATE AT BOUNDARIES** — Never trust client input\n", ""))
          rc = ps.main(["protected_sections.py", str(f)])
          assert rc == 1
          assert "FAIL: protected Iron Law" in capsys.readouterr().out
      
      
      def test_main_returns_zero_when_intact(monkeypatch, tmp_path):
          monkeypatch.setattr(ps, "_git_head_version", lambda path: HEAD)
          f = tmp_path / "SKILL.md"
          f.write_text(HEAD)
          assert ps.main(["protected_sections.py", str(f)]) == 0
      
    • __init__.py 0 B
  • .gitignore 130 B · in bundle
  • program.md 3.9 KB
    # Autoresearch Program — Elixir/Phoenix Plugin Skills
    
    ## Goals (ordered by priority)
    
    1. Fix accuracy issues: stale cross-references, missing agents/skills
    2. Improve conciseness: compress bloated sections, move detail to references/
       (NEVER by trimming a protected section — see "Protected Sections" below)
    3. Strengthen Iron Laws: add missing prohibitions, ensure min coverage
    4. Improve triggering: add domain keywords to generic descriptions
    5. Fill completeness gaps: missing sections, undocumented flags
    6. Improve clarity: raise action density, remove cross-section duplication
    7. Improve specificity: add code examples, concrete patterns over vague guidance
    
    ## Mutable Surface (ONLY these files)
    
    - `plugins/elixir-phoenix/skills/*/SKILL.md`
    - `plugins/elixir-phoenix/skills/*/references/*.md`
    
    ## Read-Only (NEVER mutate)
    
    - `lab/**` (eval infrastructure, this file, scripts)
    - `plugins/elixir-phoenix/agents/**`
    - `plugins/elixir-phoenix/hooks/**`
    - `plugins/elixir-phoenix/.claude-plugin/**`
    - `CLAUDE.md`
    - `CHANGELOG.md`
    - `README.md`
    
    ## Protected Sections (Frozen — append-only)
    
    The `## Iron Laws` section of every SKILL.md is **slow state**: hard-won
    prohibitions that must never erode. The loop MAY append a new Iron Law but MUST
    NEVER delete or reword an existing one. This is a hard invariant, not a scored
    tradeoff — `checks.sh` (check #7, via `scripts/protected_sections.py`) compares
    the mutation against git HEAD and forces REVERT if any existing law disappears.
    
    Rationale: SkillOpt (arXiv 2605.23904) measured that removing this fast/slow
    guarantee cost 22 points on SpreadsheetBench. Conciseness gains must come from
    the fast state (patterns, examples, prose), never from the protected section.
    
    ## Scoring
    
    - 7 dimensions: completeness, accuracy, conciseness, triggering, safety, clarity, specificity
    - Composite = weighted average (0.20, 0.15, 0.15, 0.10, 0.10, 0.15, 0.15)
    - Eval definitions: `lab/eval/evals/{skill}.json` (skill-specific) or default
    - Scorer: `python3 -m lab.eval.scorer {skill_path}`
    
    ## Keep Threshold
    
    Keep if `new_composite >= previous_best_composite`.
    On exact tie: keep (prefer newer — likely simpler or more accurate).
    
    ## Stop Conditions
    
    ### Structural mode (default)
    
    - All target skills at composite >= 0.95
    - 10 consecutive discards on same skill -> skip that skill
    - 50 total consecutive discards -> stop entirely
    - Human interrupts (Ctrl+C)
    
    ### Tournament mode (post-saturation)
    
    - Activated when all structural composites >= 1.000 but trigger accuracy < 0.75
    - Per-skill: incumbent A wins k=2 consecutive rounds -> converged, stop
    - Per-skill: max 20 passes hard ceiling
    - Global: 5 consecutive "all_perfect" target checks -> stop entirely
    
    ## Anti-Thrashing Rules
    
    - Same skill mutated 5+ times without improvement: skip for 10 iterations
    - If composite hasn't improved in 20 iterations: switch strategy
    - NEVER revert a mutation that improved one dimension unless another regressed by MORE
    - After a discard: analyze WHY before next attempt on same skill (ReflexiCoder)
    - NEVER retry the exact same mutation type on the same section twice
    
    ## Meta-Improvement Awareness (from Hyperagents paper)
    
    The eval framework + autoresearch loop IS a meta-improvement.
    It transfers across use cases (skill improvement → user code improvement).
    Do NOT accidentally simplify or remove infrastructure that enables self-improvement:
    
    - lab/eval/ scoring (24 matchers, 8 dimensions) — the evaluation IS the value
    - lab/autoresearch/scripts/ (run-iteration.py, checks.sh) — the loop IS the value
    - ASI metadata in JSONL — failure context IS the value
    - ideas.md backlog — deferred knowledge IS the value
    
    When improving the autoresearch system itself, treat it as a meta-improvement:
    changes to the loop/eval/scorer are higher-value than changes to individual skills.
    
    ## Simplicity Criterion
    
    A 0.01 improvement that adds 10 lines of content? Probably not worth it.
    A 0.01 improvement from removing redundancy? Definitely keep.
    All else equal, shorter is better.
    
  • retention.py 5.7 KB
    """Retention@K convergence metric for autoresearch.
    
    Adapted from TACO (arXiv 2604.19572) — instead of running fixed iterations,
    stop when the top-K skill ranking by trigger accuracy stabilizes:
    
        Retention_K^(i) = |TopK(R^(i-1)) ∩ TopK(R^(i))| / K
    
    When Retention@K ≥ threshold for `convergence_streak` consecutive iterations,
    autoresearch has converged. Defaults match the paper: K=30, threshold=0.9,
    streak=2.
    
    Why one level above per-skill tournaments:
    - `description_tournament.py` already has k=2 consecutive-A-wins convergence
      for a single skill description.
    - Retention@K answers the orthogonal question: "Has the SET of top skills
      stopped reshuffling?" Useful for deciding when to stop running the whole
      autoresearch loop, not when one skill is done.
    """
    
    import json
    import os
    from datetime import datetime, timezone
    
    PROJECT_ROOT = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
    TRIGGER_RESULTS_DIR = os.path.join(PROJECT_ROOT, "lab", "eval", "triggers", "results")
    RETENTION_LEDGER = os.path.join(PROJECT_ROOT, "lab", "autoresearch", "retention.jsonl")
    
    
    def compute_topk_by_trigger(k: int = 30) -> list[tuple[str, float]]:
        """Read cached trigger results, return top-K skills by accuracy.
    
        Returns list of (skill_name, accuracy) sorted descending by accuracy,
        truncated to k entries. Skills without cached results are excluded.
        Ties are broken by skill name (alphabetical) for determinism.
        """
        if not os.path.isdir(TRIGGER_RESULTS_DIR):
            return []
    
        scores: list[tuple[str, float]] = []
        for fname in sorted(os.listdir(TRIGGER_RESULTS_DIR)):
            if not fname.endswith(".json") or fname.startswith("_"):
                continue
            skill = fname[:-5]
            path = os.path.join(TRIGGER_RESULTS_DIR, fname)
            try:
                with open(path) as f:
                    data = json.load(f)
            except (json.JSONDecodeError, OSError):
                continue
            accuracy = data.get("accuracy")
            if accuracy is None:
                continue
            scores.append((skill, float(accuracy)))
    
        # Sort by accuracy desc, then skill name asc (deterministic tiebreak)
        scores.sort(key=lambda x: (-x[1], x[0]))
        return scores[:k]
    
    
    def retention_at_k(prev_topk: list[str], curr_topk: list[str], k: int) -> float:
        """Compute Retention@K — overlap fraction between two top-K lists.
    
        Args:
            prev_topk: Skill names from previous iteration's top-K
            curr_topk: Skill names from current iteration's top-K
            k: Denominator (use the configured K, not actual list lengths,
                so partial pools don't inflate retention)
    
        Returns:
            |prev ∩ curr| / k. Always between 0.0 and 1.0.
        """
        if k <= 0:
            return 0.0
        overlap = len(set(prev_topk) & set(curr_topk))
        return overlap / k
    
    
    def load_last_entry() -> dict | None:
        """Read the most recent retention.jsonl entry. None if ledger empty/missing."""
        if not os.path.isfile(RETENTION_LEDGER):
            return None
        last = None
        with open(RETENTION_LEDGER) as f:
            for line in f:
                line = line.strip()
                if not line:
                    continue
                try:
                    last = json.loads(line)
                except json.JSONDecodeError:
                    continue
        return last
    
    
    def append_entry(entry: dict) -> None:
        """Append one entry to retention.jsonl (creating dir if needed)."""
        os.makedirs(os.path.dirname(RETENTION_LEDGER), exist_ok=True)
        with open(RETENTION_LEDGER, "a") as f:
            f.write(json.dumps(entry) + "\n")
    
    
    def check_retention(k: int = 30, threshold: float = 0.9, streak: int = 2) -> dict:
        """Compute current retention vs previous, append to ledger, report status.
    
        Returns:
            {
                "k": int,
                "threshold": float,
                "top_k": list[str],         # current top-K skill names
                "retention": float | None,  # None on first iteration (no prev)
                "above_threshold": bool,
                "consecutive_above": int,   # streak count after this iteration
                "converged": bool,          # True iff streak >= configured streak
                "iteration": int,           # 1-indexed count of ledger entries
            }
        """
        curr_pairs = compute_topk_by_trigger(k=k)
        curr_topk = [s for s, _ in curr_pairs]
    
        prev = load_last_entry()
        prev_topk = prev["top_k"] if prev else []
        prev_consecutive = prev.get("consecutive_above", 0) if prev else 0
        prev_iteration = prev.get("iteration", 0) if prev else 0
    
        if prev is None:
            retention = None
            above = False
            consecutive = 0
        else:
            retention = retention_at_k(prev_topk, curr_topk, k=k)
            above = retention >= threshold
            consecutive = prev_consecutive + 1 if above else 0
    
        converged = consecutive >= streak
    
        entry = {
            "iteration": prev_iteration + 1,
            "timestamp": datetime.now(timezone.utc).isoformat(),
            "k": k,
            "threshold": threshold,
            "top_k": curr_topk,
            "top_k_with_accuracy": [{"skill": s, "accuracy": a} for s, a in curr_pairs],
            "retention": retention,
            "above_threshold": above,
            "consecutive_above": consecutive,
            "converged": converged,
        }
        append_entry(entry)
    
        return {
            "k": k,
            "threshold": threshold,
            "top_k": curr_topk,
            "retention": retention,
            "above_threshold": above,
            "consecutive_above": consecutive,
            "converged": converged,
            "iteration": entry["iteration"],
        }
    
    
    def is_converged(streak: int = 2) -> bool:
        """Quick check from ledger — has retention converged?"""
        last = load_last_entry()
        if last is None:
            return False
        return last.get("consecutive_above", 0) >= streak
    
  • SKILL.md 4.6 KB
    ---
    name: lab:autoresearch
    description: >
      Self-improving loop for plugin skills. Reads program.md, proposes one
      mutation per iteration, evaluates against deterministic scorer, keeps
      improvements via git, reverts failures. Targets weakest skill+dimension.
      Use with /loop for overnight runs.
    effort: high
    argument-hint: "[--skill NAME] [--strategy targeted|sweep|random] [--dry-run] [--max-iterations N]"
    disable-model-invocation: true
    ---
    
    # Autoresearch — Plugin Skill Self-Improvement
    
    Iteratively improve plugin skills via the autoresearch pattern:
    propose one mutation -> eval -> keep/revert -> repeat.
    
    ## Usage
    
    ```
    /lab:autoresearch                           # Targeted: attack weakest skill+dimension
    /lab:autoresearch --skill review            # Focus on one skill
    /lab:autoresearch --strategy sweep          # Process all skills alphabetically
    /lab:autoresearch --dry-run                 # Show what would change, don't commit
    ```
    
    For overnight runs:
    
    ```
    /loop 5m /lab:autoresearch --strategy sweep --max-iterations 200
    ```
    
    ## Iron Laws
    
    1. **ONE mutation per iteration** — if description needs "and", split into two
    2. **NEVER mutate read-only files** — check program.md before every write
    3. **EVAL is deterministic** — always use the wrapper script, never LLM-judge
    4. **REVERT on regression OR checks failure** — no exceptions
    5. **LOG every iteration** — use `keep` or `revert` command (never skip)
    6. **CHECK ideas.md before proposing** — don't rediscover known optimizations
    
    ## Wrapper Script Commands
    
    All eval/git/journal operations go through ONE script. Do NOT run these manually.
    
    ```bash
    # Find the weakest skill+dimension
    python3 lab/autoresearch/scripts/run-iteration.py target --strategy targeted
    
    # Score a skill (before mutation, to get baseline)
    python3 lab/autoresearch/scripts/run-iteration.py score <skill-name>
    
    # After mutation: score + checks + compare → verdict (KEEP or REVERT)
    python3 lab/autoresearch/scripts/run-iteration.py eval <skill-name>
    
    # Act on verdict:
    python3 lab/autoresearch/scripts/run-iteration.py keep <skill> <dim> <old> <new> \
      --desc "what changed" --asi '{"hypothesis": "why", "mechanism": "how"}'
    
    python3 lab/autoresearch/scripts/run-iteration.py revert <skill> <dim> <old> <new> \
      --desc "what was attempted" --asi '{"hypothesis": "why", "regression": "what broke", "avoid": "do not retry this"}'
    
    # Check overall progress
    python3 lab/autoresearch/scripts/run-iteration.py status
    ```
    
    ## Core Loop (ONE iteration)
    
    ### Step 1: Read State
    
    1. Read `lab/autoresearch/program.md` (goals, mutable surface, rules)
    2. Read `lab/autoresearch/ideas.md` if it exists (deferred optimizations)
    3. Run: `python3 lab/autoresearch/scripts/run-iteration.py status`
    
    ### Step 2: Select Target
    
    Run: `python3 lab/autoresearch/scripts/run-iteration.py target --strategy targeted`
    
    Parse the JSON: `skill`, `dimension`, `failing_checks`. If `all_perfect` → STOP.
    
    ### Step 3: Read + Propose
    
    1. Read target SKILL.md and its references/ listing
    2. Read eval definition from `lab/eval/evals/{skill}.json`
    3. Check `ideas.md` for deferred ideas about this skill
    4. Check recent journal entries for prior failures on this skill (avoid repeats)
    5. Consult `${CLAUDE_SKILL_DIR}/references/mutation-strategies.md`
    6. Propose exactly ONE change targeting the failing checks
    
    ### Step 4: Apply + Evaluate
    
    1. Apply the mutation via Edit tool
    2. Run: `python3 lab/autoresearch/scripts/run-iteration.py eval <skill-name>`
    3. Parse JSON → check `verdict` field
    
    ### Step 5: Keep or Revert
    
    **If verdict is KEEP**:
    
    ```bash
    python3 lab/autoresearch/scripts/run-iteration.py keep <skill> <dim> <old> <new> \
      --desc "..." --asi '{"hypothesis": "...", "mechanism": "..."}'
    ```
    
    **If verdict is REVERT**:
    
    ```bash
    python3 lab/autoresearch/scripts/run-iteration.py revert <skill> <dim> <old> <new> \
      --desc "..." --asi '{"hypothesis": "...", "regression": "...", "avoid": "..."}'
    ```
    
    ### Step 6: Ideas Backlog
    
    If during analysis you discovered a promising optimization you can't act on now:
    
    - Append it to `lab/autoresearch/ideas.md` as a bullet
    - On next resume: prune stale/tried ideas, experiment with the rest
    
    ### Step 7: Continue or Stop
    
    - All targets >= 0.95? Print "AUTORESEARCH_COMPLETE"
    - Max iterations reached? Print "AUTORESEARCH_COMPLETE"
    - 50 consecutive discards? Print "AUTORESEARCH_STUCK"
    - Otherwise: immediately start Step 1 again
    
    ## References
    
    - `${CLAUDE_SKILL_DIR}/references/mutation-strategies.md` — mutation type catalog
    - `${CLAUDE_SKILL_DIR}/references/state-management.md` — git protocol, journaling
    - `lab/autoresearch/program.md` — research agenda (read every iteration)
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related