Claude Skill

ai-research-explore

Rigor Explore compatible skill slug for meaningful and potentially novel deep learning research candidates. Use when the researcher has chosen the task family, dataset, benchmark, evaluation method, provided SOTA references, and wants candidate-only exploration on top of `current

LLM Mart · 0 points · 3 views 11 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download lllllllama-rigorpilot-skills-skills_ai-research-explore-bd91195.zip · 81 KB
Part of lllllllama/rigorpilot-skills — 11 skills

Install

skills CLI npx skills add https://github.com/lllllllama/RigorPilot-Skills/tree/main/skills/ai-research-explore
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install lllllllama-rigorpilot-skills@llmmart
Git git clone https://github.com/lllllllama/RigorPilot-Skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole lllllllama/rigorpilot-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

ai-research-explore

Purpose

Use this as the Rigor Explore compatible skill slug after the researcher explicitly authorizes candidate-only work on top of a durable current_research anchor. The installed slug remains ai-research-explore for compatibility. Rigor Explore is for meaningful and potentially novel deep learning research candidates while preserving scientific rigor, comparability, reproducibility, and auditable collaboration. Novelty and significance remain hypotheses before literature contrast, ablation evidence, and fair comparison. The skill does not promise autonomous discovery, global benchmark completeness, novelty proof, or trusted reproduction success.

Start from the shared operating principles in ../ai-research-reproduction/references/agent-operating-principles.md, then load ../ai-research-reproduction/references/research-rigor-principles.md for research claims and ../ai-research-reproduction/references/deep-learning-experiment-principles.md when experiment details affect comparability or reproducibility.

Fit

Use this skill only when the request has both:

  • Explicit exploration authorization such as candidate-only work, isolated branch or worktree, sweep, several variants, or exploratory ranking.
  • A durable current_research context such as a branch, commit, checkpoint, run record, or already-trained local model state.

Keep narrow code-only requests on explore-code. Keep narrow run-only requests on explore-run. Keep passive repository analysis on analyze-project. Keep README-first reproduction on ai-research-reproduction.

Research Rhythm

Use a two-loop rhythm:

  • Outer loop: understand the repository, freeze task/dataset/evaluation/budget, preserve user ideas, map sources, gate ideas, and decide whether the next experiment is worth running.
  • Inner loop: make one bounded candidate change or run, smoke-check it, collect evidence, rank it against the current anchor, and either stop or return to the outer loop with the new evidence.

This rhythm is a guide, not a rigid autonomous loop. Stop at explicit blockers, unclear scientific meaning, exhausted budget, missing anchor/evaluation, or a human checkpoint.

Workflow

  1. Confirm current_research and explicit explore-lane authorization.
  2. Accept either legacy variant_spec or higher-level research_campaign.
  3. In campaign mode, freeze the task, dataset, benchmark, evaluation source, SOTA reference, and budget before candidate work.
  4. Build only the repo-understanding artifacts needed for the current campaign, usually through analyze-project.
  5. Run bounded, cache-first source lookup when source support matters; prefer local curated literature such as Zotero if available, then seed sources, repo-local locators, public locators, or optional web lookup. Treat lookup as source resolution, not an open-ended literature search.
  6. Preserve researcher-provided ideas, optionally add a small bounded set of single-variable seed ideas, and rank ideas with explicit gates and score breakdowns.
  7. Prefer one clear candidate at a time. Use explore-code for bounded code adaptation and explore-run for short-cycle trials or sweeps.
  8. Use minimal-run-and-audit or run-train only when the exploratory plan requires real execution evidence.
  9. Write candidate-only outputs to analysis_outputs/, sources/, and explore_outputs/ as appropriate; never present exploratory gains as trusted reproduction success. Include SCIENTIFIC_CHANGELOG.md and COMPARABILITY_REPORT.md for candidate scientific meaning and comparison boundaries.

Ranking and Evidence

  • Before execution, prioritize candidates by expected gain, cost, success likelihood, patch surface, dependency drag, evaluation risk, and rollback ease.
  • After execution, rank by real evidence first: command status, observed metrics, artifacts, changed paths, smoke results, and reproducibility notes.
  • Keep researcher-provided evaluation_source and sota_reference frozen for the campaign; do not claim they are globally complete.
  • If the top ideas are too close or the implementation cannot be decomposed into auditable units, stop for a checkpoint instead of silently choosing.

Campaign Inputs

research_campaign is preferred for Rigor Explore campaigns, but it should stay minimal. The durable core is:

  • current_research
  • task_family
  • dataset
  • benchmark
  • evaluation_source
  • sota_reference
  • compute_budget

Use candidate_ideas, variant_spec, research_lookup, idea_policy, idea_generation, source_constraints, feasibility_policy, baseline_gate, and execution_policy as optional guidance, not as fields the agent must fill for every campaign. See references/research-campaign-spec.md for the advanced schema and artifact expectations.

Reference Loading

  • Load references/ai-research-explore-policy.md for lane safety and candidate semantics.
  • Load references/research-campaign-spec.md only when a campaign file is present or the user asks for Rigor Explore campaign governance.
  • Load ../ai-research-reproduction/references/explore-variant-spec.md for run-level variant matrix details.
  • Load ../ai-research-reproduction/references/research-thinking-loop.md before proposing or ranking candidate changes; it is the required greedy observe-ground-design-compare cycle.
  • Load ../ai-research-reproduction/references/research-rigor-principles.md before making novelty, contribution, SOTA, or comparability statements.
  • Consult ~/.rigorpilot/PERSONAL_RIGOR.md if present, under ../ai-research-reproduction/references/continuous-learning-policy.md (advisory only; core wins).
  • Load ../ai-research-reproduction/references/deep-learning-experiment-principles.md when training, evaluation, baseline, ablation, metric, checkpoint, or dataset details matter.
  • Use scripts/orchestrate_explore.py and scripts/write_outputs.py for the existing deterministic artifact workflow.
Files (rigorpilot-skills)
  • agents
    • openai.yaml 423 B
      display_name: Rigor Explore
      short_description: Rigor Explore compatible slug for candidate-only current_research exploration into auditable explore_outputs.
      default_prompt: Use current_research as the explicit exploratory context, coordinate isolated code and run exploration conservatively, treat novelty as a hypothesis until evidence supports it, and write CHANGESET.md TOP_RUNS.md and status.json into explore_outputs.
      
  • references
    • ai-research-explore-policy.md 1.8 KB
      # Research Explore Policy
      
      ## Purpose
      
      Use this skill only when exploratory work has been explicitly authorized on top of `current_research`. In RigorPilot terms, the goal is meaningful and potentially novel candidate work, not verified novelty.
      
      ## Requirements
      
      - keep work on an isolated branch or worktree
      - record `current_research` in a durable form
      - treat all outputs as candidate-only exploratory records
      - coordinate code and run exploration conservatively instead of freeform rewriting
      - keep the trusted lane and exploratory lane clearly separated
      - keep improvement mining bounded to the frozen task family, dataset, benchmark, evaluation source, and provided SOTA references
      - require source-backed idea cards before transplant-style implementation planning
      - keep patch plans minimal, reversible, and auditable
      - keep research lookup free-first and provider-optional; missing external keys must not block the flow
      - prefer local curated literature, including Zotero when available, before broader lookup, without requiring a provider
      - treat `seed_only` lookup records as weak evidence only
      - distinguish `external_provider`, `parsed_locator`, `repo_local_extracted`, and `seed_only` evidence in downstream ranking and support summaries
      
      ## Avoid
      
      - implicit experimentation
      - claiming exploratory gains as trusted reproduction success
      - claiming novelty, contribution, or SOTA superiority before literature contrast, ablation evidence, and fair comparison
      - requiring non-bundled skills to complete the workflow
      - using this skill for narrow code-only or run-only asks
      - open-ended scientific brainstorming without a frozen campaign anchor
      - broad multi-module rewrites or metric-surface edits by default
      - presenting cache-first locator parsing as complete current-literature retrieval
      
    • idea-evaluation-framework.md 1.1 KB
      # Idea Evaluation Framework
      
      `ai-research-explore` uses a bounded, candidate-only evaluation scheme for Rigor Explore idea ranking.
      
      ## Hard Gates
      
      - baseline gate must not be `abandon`
      - `single_variable_fit >= 0.6`
      - `interface_fit >= 0.5`
      - `patch_surface <= 0.7`
      - `dependency_drag <= 0.7`
      - `eval_risk <= 0.6`
      - `short_run_feasibility != blocked`
      
      ## Soft Ranking
      
      Positive contributions:
      
      - `expected_upside`
      - `single_variable_fit`
      - `interface_fit`
      - `rollback_ease`
      - `innovation_story_strength`
      - `source_support_strength`
      - `execution_feasibility`
      
      Negative contributions:
      
      - `implementation_risk`
      - `eval_risk`
      - `patch_surface`
      - `dependency_drag`
      - `execution_cost`
      - `baseline_distance`
      
      ## Provenance
      
      Each ranked card should record where each field came from:
      
      - campaign input
      - read-only repo analysis
      - source lookup cache
      - source mapping and patch planning
      - execution feasibility or smoke evidence
      
      ## Guardrails
      
      - ranking is for candidate prioritization only
      - ranking does not prove novelty
      - ranking does not prove benchmark completeness
      - ranking does not prove verified SOTA superiority
      
      
    • research-campaign-spec.md 9.4 KB
      # Research Campaign Spec
      
      ## Purpose
      
      Use `research_campaign.json` or `research_campaign.yaml` when `ai-research-explore` is operating as Rigor Explore:
      
      - the task family is already chosen
      - the dataset is already chosen
      - the evaluation method is already chosen
      - the provided SOTA table is already frozen by the researcher
      - the remaining work is campaign governance, implementation, and candidate filtering
      
      This file is an advanced reference. The public entrypoint only requires the
      campaign core to be frozen; the detailed blocks below are guidance for richer
      campaigns, not fields the agent must invent on every run.
      
      `variant_spec` still exists, but it is now an optional run-level part of a
      larger campaign.
      
      Rigor Explore treats novelty as a candidate hypothesis. Novelty and
      significance remain hypotheses until supported by literature contrast, ablation
      evidence, and fair comparison.
      
      ## Minimal Shape
      
      ```json
      {
        "current_research": "seg-branch@abc1234",
        "task_family": "segmentation",
        "dataset": "DemoSeg",
        "benchmark": {
          "name": "DemoBench",
          "primary_metric": "miou",
          "metric_goal": "maximize"
        },
        "evaluation_source": {
          "command": "python eval.py --config configs/demo.yaml",
          "path": "eval.py",
          "primary_metric": "miou",
          "metric_goal": "maximize"
        },
        "sota_reference": [
          {
            "name": "Provided SOTA",
            "metric": "miou",
            "value": 80.0,
            "source": "paper-or-table-url"
          }
        ],
        "compute_budget": {
          "max_runtime_hours": 8
        }
      }
      ```
      
      ## Durable Core Fields
      
      - `current_research`
        Durable anchor for the current research state.
      - `task_family`
        The already-chosen task family, such as `segmentation`, `classification`, or `depth`.
      - `dataset`
        The dataset name used for this campaign.
      - `benchmark`
        The benchmark name or descriptor. A dictionary may also carry `primary_metric` and `metric_goal`.
      - `evaluation_source`
        The frozen evaluation contract input. Prefer a command plus an optional path.
      - `sota_reference`
        The user-provided comparison table. `ai-research-explore` treats this as authoritative input and does not prove completeness.
      - `compute_budget`
        The bounded resource envelope for candidate-only work.
      
      ## Optional Guidance Fields
      
      - `candidate_ideas`
        Preferred but optional candidate directions that the researcher already wants to consider. If omitted, or if `idea_generation.allow_synthesized_seed_ideas` stays enabled, the orchestrator may add a small number of conservative single-variable seed ideas.
      - `variant_spec`
        Optional run-level candidate matrix used by `explore-run`.
      
      Optional top-level fields:
      
      - `baseline_gate`
      - `execution_policy`
      - `research_lookup`
      - `idea_policy`
      - `idea_generation`
      - `source_constraints`
      - `feasibility_policy`
      
      ## `evaluation_source`
      
      Supported fields:
      
      - `command`
      - `path`
      - `primary_metric`
      - `metric_goal`
      - `execution_kind`
      - `artifacts`
      - `notes`
      - `split`
      
      This block feeds both:
      
      - `analysis_outputs/EVAL_CONTRACT.md`
      - the baseline gate
      
      ## `sota_reference`
      
      Each item should preferably contain:
      
      - `name`
      - `metric`
      - `value`
      
      Optional fields:
      
      - `source`
      - `notes`
      - `metric_goal`
      
      This is a frozen comparison set for the campaign. It is not a guarantee that the real global SOTA has been fully covered.
      
      ## `candidate_ideas`
      
      Each item should contain:
      
      - `id`
      - `summary`
      - `change_scope`
      - `target_component`
      - `expected_upside`
      - `implementation_risk`
      - `eval_risk`
      - `rollback_ease`
      - `estimated_runtime_cost`
      - `single_variable_fit`
      
      Optional fields:
      
      - `hypothesis`
      - `supporting_changes`
      
      The orchestrator uses these to run the idea gate. It does not treat them as novelty claims. When a researcher idea passes hard gates, final selection stays inside the researcher pool even if synthesized or hybrid ideas are also present for auditability.
      
      ## Optional Policy Blocks
      
      ### `research_lookup`
      
      Use this block to seed auditable lookup records without turning the orchestrator into an open-ended research agent.
      
      Supported fields:
      
      - `source_preference`
      - `local_literature`
      - `queries`
      - `seed_sources`
      - `enable_repo_local_extraction`
      - `optional_providers`
      
      Rigor Explore may prefer local literature context first, including Zotero,
      if available. Local literature should be treated as curated prior knowledge. If
      local literature is unavailable or too sparse, bounded web/source lookup may be
      used. Zotero-first is a source lookup strategy, not a separate main skill. It
      should support meaningful and potentially novel idea generation, not become a
      generic literature search tool.
      
      Example future-compatible lookup hint:
      
      ```yaml
      research_lookup:
        source_preference:
          - local_literature
          - seed_sources
          - repo_local
          - public_locators
          - optional_web
        local_literature:
          enabled: auto
          provider: zotero
          fallback_to_web_when:
            - unavailable
            - too_sparse
            - insufficient_candidate_coverage
      ```
      
      All lookup artifacts are cached into `sources/` with stable names, `sources/records/`, and an `index.json`. Missing optional provider keys, including Zotero, must not block this pass.
      
      ### `idea_policy`
      
      Optional governance hints for idea selection. Current implementations keep hard gates fixed and treat policy hints as future-compatible metadata.
      
      Suggested fields:
      
      - `max_patch_surface`
      - `max_dependency_drag`
      - `require_source_backing`
      
      ### `idea_generation`
      
      Optional hints for bounded idea-space expansion. This block is additive; it should not break the minimal campaign shape.
      
      Supported fields:
      
      - `allow_synthesized_seed_ideas`
      - `max_generated_ideas`
      - `require_diverse_targets`
      
      Default behavior keeps generation conservative:
      
      - prefer single-variable ideas
      - do not modify the frozen eval contract
      - do not jump directly to broad architecture rewrites
      - keep synthesized ideas bounded to repo-local components, existing variant axes, or lookup-backed source hints
      - bind each generated seed to `current_research`, `task_family`, `dataset`, and `evaluation_source` in `IDEA_SEEDS.json`
      
      ### `source_constraints`
      
      Optional hints for transplant scope.
      
      Suggested fields:
      
      - `preferred_repos`
      - `forbidden_paths`
      - `protected_zones`
      
      ### `feasibility_policy`
      
      Optional hints for bounded execution.
      
      Suggested fields:
      
      - `prefer_short_run_only`
      - `require_gpu`
      - `max_short_run_hours`
      
      ## Gates
      
      ### Baseline gate
      
      Default rules:
      
      - `maximize`: abandon if baseline trails provided SOTA by more than `2.0` absolute points
      - `minimize`: abandon if baseline is worse than provided SOTA by more than `5%`
      
      The gate can return:
      
      - `proceed`
      - `borderline`
      - `abandon`
      - `not-applicable`
      
      ### Idea gate
      
      Hard gates:
      
      - `baseline_gate != abandon`
      - `single_variable_fit >= 0.6`
      - `interface_fit >= 0.5`
      - `patch_surface <= 0.7`
      - `dependency_drag <= 0.7`
      - `eval_risk <= 0.6`
      - `short_run_feasibility != blocked`
      
      Soft ranking combines:
      
      - `expected_upside`
      - `single_variable_fit`
      - `groundedness`
      - `novelty_estimate`
      - `interface_fit`
      - `rollback_ease`
      - `source_support_strength`
      - `ablation_clarity`
      - `implementation_story_clarity`
      - `implementation_risk`
      - `eval_risk`
      - `estimated_runtime_cost`
      - `patch_surface`
      - `dependency_drag`
      - `baseline_distance`
      
      `IDEA_SCORES.json` records both raw inputs and explicit score breakdowns. If the active top-two ideas are too close, `ai-research-explore` records a human checkpoint instead of silently training.
      
      If the selected idea cannot be decomposed into implementable atomic units, `ai-research-explore` records an explicit blocker/checkpoint such as `atomic-decomposition-blocked` and stops before broader implementation or execution.
      
      ## Output Expectations
      
      The following artifacts are the full advanced campaign surface. A minimal
      campaign should produce only the files justified by the active work; do not
      inflate the run with empty artifacts just to satisfy this list.
      
      Campaign mode writes:
      
      - `analysis_outputs/RESEARCH_MAP.md`
      - `analysis_outputs/CHANGE_MAP.md`
      - `analysis_outputs/EVAL_CONTRACT.md`
      - `analysis_outputs/SOURCE_INVENTORY.md`
      - `analysis_outputs/SOURCE_SUPPORT.json`
      - `analysis_outputs/IMPROVEMENT_BANK.md`
      - `analysis_outputs/IDEA_CARDS.json`
      - `analysis_outputs/IDEA_SEEDS.json`
      - `analysis_outputs/IDEA_EVALUATION.md`
      - `analysis_outputs/IDEA_SCORES.json`
      - `analysis_outputs/MODULE_CANDIDATES.md`
      - `analysis_outputs/INTERFACE_DIFF.md`
      - `analysis_outputs/ATOMIC_IDEA_MAP.md`
      - `analysis_outputs/ATOMIC_IDEA_MAP.json`
      - `analysis_outputs/IMPLEMENTATION_FIDELITY.md`
      - `analysis_outputs/IMPLEMENTATION_FIDELITY.json`
      - `analysis_outputs/RESOURCE_PLAN.md`
      - `analysis_outputs/status.json`
      - `sources/index.json`
      - `sources/SUMMARY.md`
      - `sources/records/`
      - `explore_outputs/CHANGESET.md`
      - `explore_outputs/IDEA_GATE.md`
      - `explore_outputs/EXPERIMENT_PLAN.md`
      - `explore_outputs/EXPERIMENT_MANIFEST.md`
      - `explore_outputs/EXPERIMENT_LEDGER.md`
      - `explore_outputs/TRANSPLANT_SMOKE_REPORT.md`
      - `explore_outputs/TOP_RUNS.md`
      - `explore_outputs/status.json`
      
      ## Notes
      
      - Keep idea generation bounded and auditable rather than open-ended.
      - Keep evaluation and SOTA inputs human-frozen.
      - `IDEA_SEEDS.json` should expose per-seed bindings such as `context_anchor`, `task_family_binding`, `dataset_binding`, `evaluation_binding`, and `constraint_notes`.
      - `IMPLEMENTATION_FIDELITY.json` should separate `planned_implementation_sites`, `heuristic_implementation_sites`, and `observed_implementation_sites`, and should record `verification_level` as one of `not_checked`, `planned_only`, `heuristic_only`, `executor_observed`, or `diff_verified`.
      - Let `ai-research-explore` focus on understanding, gating, implementation planning, controlled execution, and comparison.
      
      
    • smoke-validation-policy.md 787 B
      # Smoke Validation Policy
      
      `ai-research-explore` should prefer cheap, auditable smoke checks before broader exploratory execution.
      
      ## Required Checks
      
      - syntax parse for candidate Python files
      - import-resolution style sanity for touched modules
      - config path resolution for frozen commands
      - constructor surface availability
      - forward surface availability
      - short-run command verification or explicit planned state
      
      ## Reporting
      
      - write a standardized `TRANSPLANT_SMOKE_REPORT.md`
      - record per-check status and blockers
      - distinguish `planned` from `passed` and `failed`
      
      ## Guardrails
      
      - smoke success does not imply trusted correctness
      - smoke success does not imply global benchmark validity
      - smoke failure should block broader candidate execution when the blocker is structural
      
      
    • source-mapping-policy.md 615 B
      # Source Mapping Policy
      
      Source mapping in `ai-research-explore` is for bounded, auditable adaptation only.
      
      ## Required Outputs
      
      - source repo or source reference id
      - source file and source symbol when known
      - target file and target symbol
      - supporting files
      - interface diff summary
      - minimal reversible patch plan
      
      ## Patch Classes
      
      - `config-only`
      - `import-glue`
      - `module-transplant-shim`
      
      These classes are preferred in that order.
      
      ## Forbidden Defaults
      
      - broad rewrites
      - train-loop redesign
      - metric or leaderboard mutation without explicit campaign permission
      - unscoped multi-module behavior changes
      
      
    • sources-naming-policy.md 895 B
      # Sources Naming Policy
      
      All internal research lookup results for `ai-research-explore` are saved under `sources/`.
      
      ## Naming
      
      - summary/index stay at `sources/`
      - canonical source records live under `sources/records/`
      - record filename format: `kind__slug__sha12.(json|md)`
      - `kind` should be stable across reruns such as `paper`, `repo`, `benchmark`, `module`, `query`
      - `slug` should be lowercase and human-readable
      - `sha12` should be derived from the normalized lookup payload
      
      ## Cache Rules
      
      - check `sources/index.json` before creating a new record
      - reuse an existing record if the normalized payload hash matches
      - preserve source URLs, titles, and lookup queries for auditability
      
      ## Scope
      
      - `sources/` is an exploratory audit trail, not a trusted literature database
      - source cache entries must not imply benchmark completeness
      - source cache entries must not imply novelty proof
      
      
  • scripts
    • lookup
      • providers
        • arxiv_provider.py 2.8 KB
          """Free arXiv metadata provider."""
          
          from __future__ import annotations
          
          import urllib.parse
          import xml.etree.ElementTree as ET
          from typing import Any, Dict
          
          from .base import http_get
          
          
          def resolve_arxiv_record(locator_info: Dict[str, Any]) -> Dict[str, Any]:
              arxiv_id = str(locator_info.get("arxiv_id") or locator_info.get("identifier") or "").strip()
              record = {
                  "provider_type": "arxiv",
                  "source_type": "paper",
                  "locator_type": locator_info.get("locator_type", "arxiv_id"),
                  "raw_locator": locator_info.get("raw_locator", ""),
                  "normalized_id": locator_info.get("normalized_id", f"arxiv:{arxiv_id.lower()}"),
                  "title": f"arXiv:{arxiv_id}" if arxiv_id else "",
                  "url": locator_info.get("url", ""),
                  "authors": [],
                  "year": None,
                  "venue": "arXiv",
                  "doi": "",
                  "arxiv_id": arxiv_id,
                  "parse_status": "parsed-only",
                  "fetch_status": "parsed-only",
                  "evidence_class": "parsed_locator",
                  "provider_metadata": {"resolved_via": "arxiv"},
              }
              if not arxiv_id:
                  return record
              try:
                  payload = http_get(
                      f"https://export.arxiv.org/api/query?id_list={urllib.parse.quote(arxiv_id)}",
                      accept="application/atom+xml, text/xml;q=0.9",
                  )
                  root = ET.fromstring(payload.decode("utf-8", errors="ignore"))
                  namespace = {"atom": "http://www.w3.org/2005/Atom"}
                  entry = root.find("atom:entry", namespace)
                  if entry is None:
                      return {**record, "parse_status": "fetch-failed", "fetch_status": "fetch-failed"}
                  authors = [
                      (node.findtext("atom:name", default="", namespaces=namespace) or "").strip()
                      for node in entry.findall("atom:author", namespace)
                      if (node.findtext("atom:name", default="", namespaces=namespace) or "").strip()
                  ]
                  published = (entry.findtext("atom:published", default="", namespaces=namespace) or "").strip()
                  year = int(published[:4]) if published[:4].isdigit() else None
                  url = record["url"]
                  for item in entry.findall("atom:link", namespace):
                      href = item.attrib.get("href", "")
                      if href:
                          url = href
                          break
                  return {
                      **record,
                      "title": (entry.findtext("atom:title", default="", namespaces=namespace) or "").strip() or record["title"],
                      "summary": (entry.findtext("atom:summary", default="", namespaces=namespace) or "").strip(),
                      "authors": authors,
                      "year": year,
                      "url": url,
                      "parse_status": "resolved",
                      "fetch_status": "network-fetched",
                      "evidence_class": "external_provider",
                  }
              except Exception:
                  return {**record, "parse_status": "fetch-failed", "fetch_status": "fetch-failed"}
          
        • base.py 3 KB
          """Shared transport and HTML helpers for lookup providers."""
          
          from __future__ import annotations
          
          import json
          import urllib.request
          from html.parser import HTMLParser
          from typing import Any, Dict, List, Optional, Tuple
          
          
          REQUEST_TIMEOUT_SECONDS = 6
          USER_AGENT = "ai-research-explore-lookup/2.0"
          
          
          class MetadataHTMLParser(HTMLParser):
              def __init__(self) -> None:
                  super().__init__()
                  self.in_title = False
                  self.title_parts: List[str] = []
                  self.meta: Dict[str, str] = {}
                  self.links: Dict[str, str] = {}
          
              def handle_starttag(self, tag: str, attrs: List[Tuple[str, Optional[str]]]) -> None:
                  lowered = tag.lower()
                  attr_map = {key.lower(): (value or "") for key, value in attrs}
                  if lowered == "title":
                      self.in_title = True
                  if lowered == "meta":
                      name = attr_map.get("name") or attr_map.get("property")
                      content = attr_map.get("content", "").strip()
                      if name and content:
                          self.meta[name.lower()] = content
                  if lowered == "link":
                      rel = attr_map.get("rel", "").lower()
                      href = attr_map.get("href", "").strip()
                      if rel and href:
                          self.links[rel] = href
          
              def handle_endtag(self, tag: str) -> None:
                  if tag.lower() == "title":
                      self.in_title = False
          
              def handle_data(self, data: str) -> None:
                  if self.in_title and data.strip():
                      self.title_parts.append(data.strip())
          
              def title_text(self) -> str:
                  return " ".join(self.title_parts).strip()
          
              def description_text(self) -> str:
                  for key in ("og:description", "description", "twitter:description"):
                      if self.meta.get(key):
                          return self.meta[key]
                  return ""
          
              def canonical_url(self) -> str:
                  return self.links.get("canonical") or self.meta.get("og:url", "")
          
          
          def http_get(url: str, *, accept: str = "application/json, text/plain;q=0.9, text/html;q=0.8") -> bytes:
              request = urllib.request.Request(
                  url,
                  headers={
                      "User-Agent": USER_AGENT,
                      "Accept": accept,
                  },
              )
              with urllib.request.urlopen(request, timeout=REQUEST_TIMEOUT_SECONDS) as response:
                  return response.read()
          
          
          def http_get_json(url: str, *, accept: str = "application/json") -> Dict[str, Any]:
              payload = http_get(url, accept=accept)
              loaded = json.loads(payload.decode("utf-8", errors="ignore"))
              return loaded if isinstance(loaded, dict) else {}
          
          
          def coerce_author_list(values: Any) -> List[str]:
              authors: List[str] = []
              if isinstance(values, list):
                  for item in values:
                      if isinstance(item, dict):
                          given = str(item.get("given") or "").strip()
                          family = str(item.get("family") or "").strip()
                          full = " ".join(part for part in [given, family] if part).strip()
                          if full:
                              authors.append(full)
                      elif str(item).strip():
                          authors.append(str(item).strip())
              return authors
          
          
        • doi_provider.py 2.8 KB
          """Free DOI metadata provider."""
          
          from __future__ import annotations
          
          import json
          import urllib.parse
          from typing import Any, Dict
          
          from .base import coerce_author_list, http_get
          
          
          def resolve_doi_record(locator_info: Dict[str, Any]) -> Dict[str, Any]:
              doi = str(locator_info.get("doi") or locator_info.get("identifier") or "").strip().lower()
              record = {
                  "provider_type": "doi",
                  "source_type": "paper",
                  "locator_type": locator_info.get("locator_type", "doi"),
                  "raw_locator": locator_info.get("raw_locator", ""),
                  "normalized_id": locator_info.get("normalized_id", f"doi:{doi}"),
                  "title": f"DOI:{doi}" if doi else "",
                  "url": locator_info.get("url", f"https://doi.org/{doi}" if doi else ""),
                  "authors": [],
                  "year": None,
                  "venue": "",
                  "doi": doi,
                  "arxiv_id": "",
                  "parse_status": "parsed-only",
                  "fetch_status": "parsed-only",
                  "evidence_class": "parsed_locator",
                  "provider_metadata": {"resolved_via": "doi"},
              }
              if not doi:
                  return record
              try:
                  payload = http_get(
                      f"https://doi.org/{urllib.parse.quote(doi, safe='/')}",
                      accept="application/vnd.citationstyles.csl+json, application/json;q=0.9",
                  )
                  loaded = json.loads(payload.decode("utf-8", errors="ignore"))
                  title = loaded.get("title")
                  if isinstance(title, list):
                      title = title[0] if title else ""
                  venue = loaded.get("container-title")
                  if isinstance(venue, list):
                      venue = venue[0] if venue else ""
                  year = None
                  issued = loaded.get("issued") or {}
                  date_parts = issued.get("date-parts") if isinstance(issued, dict) else None
                  if isinstance(date_parts, list) and date_parts and isinstance(date_parts[0], list) and date_parts[0]:
                      value = date_parts[0][0]
                      year = int(value) if isinstance(value, int) or (isinstance(value, str) and str(value).isdigit()) else None
                  return {
                      **record,
                      "title": str(title or record["title"]),
                      "summary": str(loaded.get("abstract") or ""),
                      "authors": coerce_author_list(loaded.get("author")),
                      "year": year,
                      "venue": str(venue or loaded.get("publisher") or ""),
                      "url": str(loaded.get("URL") or record["url"]),
                      "parse_status": "resolved",
                      "fetch_status": "network-fetched",
                      "evidence_class": "external_provider",
                      "provider_metadata": {
                          "resolved_via": "doi",
                          "publisher": loaded.get("publisher"),
                          "type": loaded.get("type"),
                      },
                  }
              except Exception:
                  return {**record, "parse_status": "fetch-failed", "fetch_status": "fetch-failed"}
          
        • github_provider.py 3.3 KB
          """Free GitHub repository metadata provider."""
          
          from __future__ import annotations
          
          import base64
          from typing import Any, Dict, List
          
          from lookup.normalizers import extract_urls
          
          from .base import http_get_json
          
          
          def _paper_links(links: List[str]) -> List[str]:
              return [
                  link
                  for link in links
                  if "arxiv.org" in link.lower() or "doi.org" in link.lower() or "openreview.net" in link.lower()
              ]
          
          
          def _fetch_readme(owner: str, repo: str) -> Dict[str, Any]:
              try:
                  payload = http_get_json(f"https://api.github.com/repos/{owner}/{repo}/readme")
              except Exception:
                  return {"readme_links": [], "paper_links_in_readme": []}
              content = payload.get("content")
              if not content:
                  return {"readme_links": [], "paper_links_in_readme": []}
              try:
                  decoded = base64.b64decode(str(content).encode("utf-8"), validate=False).decode("utf-8", errors="ignore")
              except Exception:
                  return {"readme_links": [], "paper_links_in_readme": []}
              links = extract_urls(decoded)
              return {
                  "readme_links": links,
                  "paper_links_in_readme": _paper_links(links),
              }
          
          
          def resolve_github_record(locator_info: Dict[str, Any]) -> Dict[str, Any]:
              repo_full_name = str(locator_info.get("repo_full_name") or locator_info.get("identifier") or "").strip()
              owner = str(locator_info.get("owner") or "").strip()
              repo = str(locator_info.get("repo") or "").strip()
              record = {
                  "provider_type": "github",
                  "source_type": "repo",
                  "locator_type": locator_info.get("locator_type", "github_repo_url"),
                  "raw_locator": locator_info.get("raw_locator", ""),
                  "normalized_id": locator_info.get("normalized_id", f"github:{repo_full_name.lower()}"),
                  "title": repo_full_name,
                  "url": locator_info.get("url", ""),
                  "authors": [],
                  "year": None,
                  "venue": "GitHub",
                  "repo_full_name": repo_full_name,
                  "doi": "",
                  "arxiv_id": "",
                  "source_file": locator_info.get("source_file", ""),
                  "parse_status": "parsed-only",
                  "fetch_status": "parsed-only",
                  "evidence_class": "parsed_locator",
                  "provider_metadata": {"resolved_via": "github"},
              }
              if not owner or not repo:
                  return record
              try:
                  payload = http_get_json(f"https://api.github.com/repos/{owner}/{repo}")
                  readme_meta = _fetch_readme(owner, repo)
                  return {
                      **record,
                      "title": str(payload.get("full_name") or record["title"]),
                      "summary": str(payload.get("description") or ""),
                      "url": str(payload.get("html_url") or record["url"]),
                      "repo_full_name": str(payload.get("full_name") or repo_full_name),
                      "parse_status": "resolved",
                      "fetch_status": "network-fetched",
                      "evidence_class": "external_provider",
                      "provider_metadata": {
                          "resolved_via": "github",
                          "default_branch": payload.get("default_branch"),
                          "homepage": payload.get("homepage"),
                          "license": (payload.get("license") or {}).get("spdx_id"),
                          "stargazers_count": payload.get("stargazers_count"),
                          **readme_meta,
                      },
                  }
              except Exception:
                  return {**record, "parse_status": "fetch-failed", "fetch_status": "fetch-failed"}
          
        • optional_provider.py 844 B
          """Optional paid-provider adapter placeholder."""
          
          from __future__ import annotations
          
          import os
          from typing import Any, Dict, Optional
          
          
          OPTIONAL_PROVIDER_ENV_VARS = {
              "openrouter": "RESEARCH_LOOKUP_OPENROUTER_API_KEY",
              "perplexity": "RESEARCH_LOOKUP_PERPLEXITY_API_KEY",
              "parallel": "RESEARCH_LOOKUP_PARALLEL_API_KEY",
          }
          
          
          def resolve_optional_record(locator_info: Dict[str, Any], lookup_config: Dict[str, Any]) -> Optional[Dict[str, Any]]:
              configured = lookup_config.get("optional_providers") if isinstance(lookup_config, dict) else None
              providers = [str(item).strip().lower() for item in (configured or []) if str(item).strip()]
              for provider_name in providers:
                  env_name = OPTIONAL_PROVIDER_ENV_VARS.get(provider_name)
                  if env_name and os.environ.get(env_name):
                      return None
              return None
          
        • url_provider.py 2 KB
          """Generic URL metadata provider."""
          
          from __future__ import annotations
          
          import urllib.parse
          from typing import Any, Dict
          
          from lookup.normalizers import canonicalize_url
          
          from .base import MetadataHTMLParser, http_get
          
          
          def resolve_url_record(locator_info: Dict[str, Any]) -> Dict[str, Any]:
              url = canonicalize_url(locator_info.get("url") or locator_info.get("raw_locator") or "")
              parsed = urllib.parse.urlsplit(url) if url else None
              record = {
                  "provider_type": "url",
                  "source_type": "web",
                  "locator_type": locator_info.get("locator_type", "url"),
                  "raw_locator": locator_info.get("raw_locator", ""),
                  "normalized_id": locator_info.get("normalized_id", f"url:{url}" if url else ""),
                  "title": url,
                  "url": url,
                  "authors": [],
                  "year": None,
                  "venue": parsed.netloc if parsed else "",
                  "repo_full_name": "",
                  "doi": "",
                  "arxiv_id": "",
                  "parse_status": "parsed-only",
                  "fetch_status": "parsed-only",
                  "evidence_class": "parsed_locator",
                  "provider_metadata": {"resolved_via": "url", "host": parsed.netloc.lower() if parsed else ""},
              }
              if not url:
                  return record
              try:
                  payload = http_get(url, accept="text/html, application/xhtml+xml;q=0.9")
                  parser = MetadataHTMLParser()
                  parser.feed(payload.decode("utf-8", errors="ignore"))
                  canonical = parser.canonical_url() or url
                  return {
                      **record,
                      "title": parser.meta.get("og:title") or parser.title_text() or url,
                      "summary": parser.description_text(),
                      "url": canonicalize_url(canonical),
                      "parse_status": "resolved",
                      "fetch_status": "network-fetched",
                      "evidence_class": "external_provider",
                      "provider_metadata": {
                          **record["provider_metadata"],
                          "canonical_url": canonicalize_url(canonical),
                      },
                  }
              except Exception:
                  return {**record, "parse_status": "fetch-failed", "fetch_status": "fetch-failed"}
          
        • __init__.py 457 B
          """Provider adapters for free-first research lookup."""
          
          from .arxiv_provider import resolve_arxiv_record
          from .doi_provider import resolve_doi_record
          from .github_provider import resolve_github_record
          from .optional_provider import resolve_optional_record
          from .url_provider import resolve_url_record
          
          __all__ = [
              "resolve_arxiv_record",
              "resolve_doi_record",
              "resolve_github_record",
              "resolve_optional_record",
              "resolve_url_record",
          ]
          
      • cache_store.py 8.4 KB
        """Cache/index helpers for source records."""
        
        from __future__ import annotations
        
        import json
        from datetime import datetime, timezone
        from pathlib import Path
        from typing import Any, Dict, Iterable, List
        
        from .normalizers import slugify, stable_digest, stable_filename
        from .record_schema import normalize_record, record_priority
        
        
        def _normalized_id_from_index_item(item: Dict[str, Any]) -> str:
            if item.get("normalized_id"):
                return str(item.get("normalized_id"))
            provider_type = str(item.get("provider_type") or "seed")
            identifier = str(item.get("provider_identifier") or item.get("source_url") or item.get("query") or "")
            return f"{provider_type}:{identifier}".strip(":")
        
        
        def load_cache_index(sources_dir: Path) -> Dict[str, Any]:
            index_path = sources_dir / "index.json"
            if not index_path.exists():
                return {
                    "schema_version": "2.0",
                    "mode": "free-first-cache-first",
                    "records_dir": "sources/records",
                    "records": [],
                    "record_lookup": {},
                }
            payload = json.loads(index_path.read_text(encoding="utf-8"))
            records = payload.get("records", [])
            payload["record_lookup"] = {
                _normalized_id_from_index_item(item): dict(item)
                for item in records
                if _normalized_id_from_index_item(item)
            }
            return payload
        
        
        def _merge_lists(left: Iterable[Any], right: Iterable[Any]) -> List[Any]:
            merged: List[Any] = []
            for item in list(left) + list(right):
                if item not in merged and item not in ("", None, []):
                    merged.append(item)
            return merged
        
        
        def _prefer_value(existing: Any, incoming: Any) -> Any:
            if incoming not in ("", None, [], {}):
                return incoming
            return existing
        
        
        def merge_records(existing: Dict[str, Any], incoming: Dict[str, Any]) -> Dict[str, Any]:
            normalized_existing = normalize_record(existing)
            normalized_incoming = normalize_record(incoming)
            if record_priority(normalized_incoming) >= record_priority(normalized_existing):
                primary, secondary = normalized_incoming, normalized_existing
            else:
                primary, secondary = normalized_existing, normalized_incoming
            merged = dict(primary)
            for key in (
                "title",
                "summary",
                "url",
                "venue",
                "repo_full_name",
                "doi",
                "arxiv_id",
                "source_repo",
                "source_file",
                "source_symbol",
            ):
                merged[key] = _prefer_value(secondary.get(key), primary.get(key))
            merged["authors"] = _merge_lists(secondary.get("authors", []), primary.get("authors", []))
            merged["origins"] = _merge_lists(secondary.get("origins", []), primary.get("origins", []))
            merged["extracted_from_repo_paths"] = _merge_lists(
                secondary.get("extracted_from_repo_paths", []),
                primary.get("extracted_from_repo_paths", []),
            )
            merged["selection_hints"] = _merge_lists(secondary.get("selection_hints", []), primary.get("selection_hints", []))
            merged["provider_metadata"] = {**secondary.get("provider_metadata", {}), **primary.get("provider_metadata", {})}
            return normalize_record(merged)
        
        
        def store_records(sources_dir: Path, records: Iterable[Dict[str, Any]]) -> Dict[str, Any]:
            sources_dir.mkdir(parents=True, exist_ok=True)
            records_dir = sources_dir / "records"
            records_dir.mkdir(parents=True, exist_ok=True)
        
            existing_index = load_cache_index(sources_dir)
            existing_lookup = existing_index.get("record_lookup", {})
            stored_by_id: Dict[str, Dict[str, Any]] = {}
            cache_hits = 0
            cache_misses = 0
            merge_upgrades = 0
        
            for raw_record in records:
                record = normalize_record(raw_record)
                normalized_id = str(record.get("normalized_id") or "")
                if not normalized_id:
                    identity = {
                        "source_type": record.get("source_type"),
                        "provider_type": record.get("provider_type"),
                        "locator_type": record.get("locator_type"),
                        "raw_locator": record.get("raw_locator"),
                        "url": record.get("url"),
                        "title": record.get("title"),
                    }
                    normalized_id = f"seed:{stable_digest(identity)[:16]}"
                    record["normalized_id"] = normalized_id
                if normalized_id in stored_by_id:
                    stored_by_id[normalized_id] = merge_records(stored_by_id[normalized_id], record)
                    continue
                existing = existing_lookup.get(normalized_id)
                if existing:
                    cache_hits += 1
                    existing_rel = str(existing.get("artifact_path") or "")
                    existing_path = None
                    if existing_rel.startswith("sources/"):
                        existing_path = sources_dir / Path(existing_rel).relative_to("sources")
                    existing_payload = dict(existing)
                    if existing_path and existing_path.exists():
                        existing_payload = json.loads(existing_path.read_text(encoding="utf-8"))
                    merged = merge_records(existing_payload, record)
                    if record_priority(merged) > record_priority(normalize_record(existing_payload)):
                        merge_upgrades += 1
                    merged["cache_hit"] = True
                    stored_by_id[normalized_id] = merged
                else:
                    cache_misses += 1
                    record["cache_hit"] = False
                    stored_by_id[normalized_id] = record
        
            timestamp = datetime.now(timezone.utc).isoformat()
            stored_records: List[Dict[str, Any]] = []
            index_records: List[Dict[str, Any]] = []
            for normalized_id in sorted(stored_by_id):
                record = normalize_record(stored_by_id[normalized_id])
                if not record.get("resolved_at"):
                    record["resolved_at"] = timestamp
                digest = stable_digest(
                    {
                        "normalized_id": normalized_id,
                        "provider_type": record.get("provider_type"),
                        "source_type": record.get("source_type"),
                    }
                )
                source_id = record.get("source_id") or f"{record.get('source_type', 'source')}:{digest[:8]}"
                record["source_id"] = source_id
                slug = slugify(record.get("title") or normalized_id)
                filename = stable_filename(str(record.get("source_type") or "source"), slug, digest)
                artifact_path = records_dir / filename
                record["artifact_path"] = f"sources/records/{filename}"
                record["artifact_abspath"] = str(artifact_path)
                record["digest"] = digest
                artifact_path.write_text(
                    json.dumps({"schema_version": "2.0", **record}, indent=2, ensure_ascii=False),
                    encoding="utf-8",
                )
                stored_records.append(record)
                index_records.append(
                    {
                        "source_id": source_id,
                        "source_type": record.get("source_type"),
                        "provider_type": record.get("provider_type"),
                        "locator_type": record.get("locator_type"),
                        "raw_locator": record.get("raw_locator"),
                        "normalized_id": normalized_id,
                        "title": record.get("title"),
                        "url": record.get("url"),
                        "repo_full_name": record.get("repo_full_name"),
                        "doi": record.get("doi"),
                        "arxiv_id": record.get("arxiv_id"),
                        "evidence_class": record.get("evidence_class"),
                        "evidence_weight": record.get("evidence_weight"),
                        "parse_status": record.get("parse_status"),
                        "cache_hit": record.get("cache_hit"),
                        "artifact_path": record.get("artifact_path"),
                        "source_repo": record.get("source_repo"),
                        "source_file": record.get("source_file"),
                        "source_symbol": record.get("source_symbol"),
                        "resolved_at": record.get("resolved_at"),
                    }
                )
        
            index_payload = {
                "schema_version": "2.0",
                "mode": "free-first-cache-first",
                "records_dir": "sources/records",
                "records": index_records,
                "stats": {
                    "record_count": len(index_records),
                    "cache_hits": cache_hits,
                    "cache_misses": cache_misses,
                    "merge_upgrades": merge_upgrades,
                },
            }
            index_path = sources_dir / "index.json"
            index_path.write_text(json.dumps(index_payload, indent=2, ensure_ascii=False), encoding="utf-8")
            return {
                "index_path": str(index_path),
                "records_dir": str(records_dir),
                "records": stored_records,
                "cache_hits": cache_hits,
                "cache_misses": cache_misses,
                "merge_upgrades": merge_upgrades,
            }
        
      • inventory_writer.py 3.1 KB
        """Human-readable lookup inventory writers."""
        
        from __future__ import annotations
        
        from pathlib import Path
        from typing import Any, Dict, Sequence
        
        
        def write_sources_summary(sources_dir: Path, records: Sequence[Dict[str, Any]]) -> Path:
            lines = [
                "# Sources Summary",
                "",
                "Research lookup for `ai-research-explore` is free-first, cache-first, and auditable.",
                "",
                "## Cached Records",
                "",
            ]
            if not records:
                lines.append("- None.")
            else:
                for item in records:
                    triple = " / ".join(
                        part
                        for part in [
                            str(item.get("source_repo") or ""),
                            str(item.get("source_file") or ""),
                            str(item.get("source_symbol") or ""),
                        ]
                        if part
                    ) or "no-source-triple"
                    lines.append(
                        f"- `{item.get('source_id')}` `{item.get('provider_type')}` `{item.get('title')}` -> `{item.get('artifact_path')}` evidence={item.get('evidence_class', 'unknown')} triple={triple}"
                    )
            summary_path = sources_dir / "SUMMARY.md"
            summary_path.write_text("\n".join(lines), encoding="utf-8")
            return summary_path
        
        
        def write_source_inventory(
            analysis_output_dir: Path,
            *,
            records: Sequence[Dict[str, Any]],
            repo_local_extractions: Sequence[Dict[str, Any]],
            cache_stats: Dict[str, Any],
        ) -> Path:
            by_class: Dict[str, int] = {}
            for item in records:
                key = str(item.get("evidence_class") or "unknown")
                by_class[key] = by_class.get(key, 0) + 1
            lines = [
                "# Source Inventory",
                "",
                "Human-readable inventory for free-first, provider-optional research lookup.",
                "",
                "## Evidence Breakdown",
                "",
            ]
            if by_class:
                for key in sorted(by_class):
                    lines.append(f"- `{key}`: {by_class[key]}")
            else:
                lines.append("- None.")
            lines.extend(
                [
                    "",
                    "## Cache Stats",
                    "",
                    f"- Cache hits: {cache_stats.get('cache_hits', 0)}",
                    f"- Cache misses: {cache_stats.get('cache_misses', 0)}",
                    f"- Merge upgrades: {cache_stats.get('merge_upgrades', 0)}",
                    "",
                    "## Repo-local Extractions",
                    "",
                ]
            )
            if repo_local_extractions:
                for item in repo_local_extractions[:20]:
                    paths = ", ".join(item.get("extracted_from_repo_paths", [])) or "unknown-path"
                    lines.append(f"- `{item.get('raw_locator') or item.get('query')}` from `{paths}`")
            else:
                lines.append("- None.")
            lines.extend(["", "## Resolved Records", ""])
            if records:
                for item in records:
                    lines.append(
                        f"- `{item.get('source_id')}` `{item.get('source_type')}` `{item.get('provider_type')}` `{item.get('title')}` evidence=`{item.get('evidence_class')}`"
                    )
            else:
                lines.append("- None.")
            path = analysis_output_dir / "SOURCE_INVENTORY.md"
            path.write_text("\n".join(lines), encoding="utf-8")
            return path
        
        
      • normalizers.py 4.8 KB
        """Locator parsing and normalization helpers for research lookup."""
        
        from __future__ import annotations
        
        import hashlib
        import json
        import re
        import urllib.parse
        from typing import Any, Dict, Optional
        
        
        ARXIV_ID_RE = re.compile(r"(?:arxiv:|arxiv\.org/(?:abs|pdf)/)?(?P<id>\d{4}\.\d{4,5}(?:v\d+)?)", re.IGNORECASE)
        DOI_RE = re.compile(r"(?:https?://(?:dx\.)?doi\.org/)?(?P<doi>10\.\d{4,9}/[-._;()/:A-Z0-9]+)", re.IGNORECASE)
        GITHUB_URL_RE = re.compile(
            r"^https?://github\.com/(?P<owner>[^/\s]+)/(?P<repo>[^/\s#]+)(?:/(?P<rest>.*))?$",
            re.IGNORECASE,
        )
        HTTP_URL_RE = re.compile(r"^https?://", re.IGNORECASE)
        URL_RE = re.compile(r"https?://[^\s<>\]\"')]+", re.IGNORECASE)
        
        
        def slugify(value: str) -> str:
            slug = re.sub(r"[^a-z0-9]+", "-", str(value).lower()).strip("-")
            return slug[:48] or "source"
        
        
        def stable_digest(payload: Dict[str, Any]) -> str:
            text = json.dumps(payload, sort_keys=True, ensure_ascii=False)
            return hashlib.sha1(text.encode("utf-8")).hexdigest()
        
        
        def stable_filename(kind: str, slug: str, digest: str, suffix: str = "json") -> str:
            return f"{kind}__{slug}__{digest[:12]}.{suffix}"
        
        
        def ensure_http_url(value: str) -> str:
            text = str(value or "").strip()
            if not text:
                return ""
            if HTTP_URL_RE.match(text):
                return text
            if text.lower().startswith("doi:"):
                return f"https://doi.org/{text[4:].strip()}"
            return text
        
        
        def canonicalize_url(value: str) -> str:
            text = ensure_http_url(value)
            if not text:
                return ""
            parsed = urllib.parse.urlsplit(text)
            path = parsed.path or "/"
            return urllib.parse.urlunsplit(
                (
                    parsed.scheme.lower(),
                    parsed.netloc.lower(),
                    path.rstrip("/") or "/",
                    parsed.query,
                    "",
                )
            )
        
        
        def extract_urls(text: str) -> list[str]:
            found: list[str] = []
            for match in URL_RE.finditer(str(text or "")):
                url = match.group(0).rstrip(".,);]")
                if url not in found:
                    found.append(url)
            return found
        
        
        def parse_arxiv_locator(locator: str) -> Optional[Dict[str, Any]]:
            text = str(locator or "").strip()
            match = ARXIV_ID_RE.search(text)
            if not match:
                return None
            arxiv_id = match.group("id")
            locator_type = "arxiv_url" if "arxiv.org" in text.lower() else "arxiv_id"
            return {
                "provider_type": "arxiv",
                "source_type": "paper",
                "locator_type": locator_type,
                "raw_locator": text,
                "normalized_id": f"arxiv:{arxiv_id.lower()}",
                "identifier": arxiv_id,
                "arxiv_id": arxiv_id,
                "url": f"https://arxiv.org/abs/{arxiv_id}",
            }
        
        
        def parse_doi_locator(locator: str) -> Optional[Dict[str, Any]]:
            text = str(locator or "").strip()
            match = DOI_RE.search(text)
            if not match:
                return None
            doi = match.group("doi").lower()
            locator_type = "doi_url" if "doi.org" in text.lower() else "doi"
            return {
                "provider_type": "doi",
                "source_type": "paper",
                "locator_type": locator_type,
                "raw_locator": text,
                "normalized_id": f"doi:{doi}",
                "identifier": doi,
                "doi": doi,
                "url": f"https://doi.org/{doi}",
            }
        
        
        def parse_github_repo_locator(locator: str) -> Optional[Dict[str, Any]]:
            text = canonicalize_url(locator)
            match = GITHUB_URL_RE.match(text)
            if not match:
                return None
            owner = match.group("owner")
            repo = (match.group("repo") or "").removesuffix(".git")
            rest = match.group("rest") or ""
            source_file = ""
            if rest.startswith("blob/"):
                parts = rest.split("/", 3)
                if len(parts) == 4:
                    source_file = parts[3]
            return {
                "provider_type": "github",
                "source_type": "repo",
                "locator_type": "github_repo_url",
                "raw_locator": str(locator or "").strip(),
                "normalized_id": f"github:{owner.lower()}/{repo.lower()}",
                "identifier": f"{owner}/{repo}",
                "repo_full_name": f"{owner}/{repo}",
                "owner": owner,
                "repo": repo,
                "source_file": source_file,
                "url": f"https://github.com/{owner}/{repo}",
            }
        
        
        def parse_generic_url(locator: str) -> Optional[Dict[str, Any]]:
            text = canonicalize_url(locator)
            if not HTTP_URL_RE.match(text):
                return None
            parsed = urllib.parse.urlsplit(text)
            return {
                "provider_type": "url",
                "source_type": "web",
                "locator_type": "url",
                "raw_locator": str(locator or "").strip(),
                "normalized_id": f"url:{text}",
                "identifier": text,
                "host": parsed.netloc.lower(),
                "url": text,
            }
        
        
        def detect_locator(locator: str) -> Optional[Dict[str, Any]]:
            for parser in (parse_github_repo_locator, parse_arxiv_locator, parse_doi_locator, parse_generic_url):
                parsed = parser(locator)
                if parsed:
                    return parsed
            return None
        
      • record_schema.py 3.4 KB
        """Lookup record schema helpers."""
        
        from __future__ import annotations
        
        from typing import Any, Dict
        
        
        EVIDENCE_CLASS_ALIASES = {
            "external-provider": "external_provider",
            "parsed-locator": "parsed_locator",
            "repo-local-extracted": "repo_local_extracted",
            "seed-only": "seed_only",
        }
        
        EVIDENCE_CLASS_PRIORITY = {
            "seed_only": 0,
            "repo_local_extracted": 1,
            "parsed_locator": 2,
            "external_provider": 3,
        }
        
        DEFAULT_RECORD_FIELDS = {
            "source_type": "web",
            "provider_type": "seed",
            "locator_type": "seed",
            "raw_locator": "",
            "normalized_id": "",
            "title": "",
            "summary": "",
            "url": "",
            "authors": [],
            "year": None,
            "venue": "",
            "repo_full_name": "",
            "doi": "",
            "arxiv_id": "",
            "evidence_class": "seed_only",
            "evidence_weight": 0.2,
            "resolved_at": "",
            "cache_hit": False,
            "parse_status": "seed-only",
            "fetch_status": "seed-only",
            "provider_metadata": {},
            "source_repo": "",
            "source_file": "",
            "source_symbol": "",
            "origins": [],
            "extracted_from_repo_paths": [],
            "selection_hints": [],
        }
        
        
        def normalize_evidence_class(value: Any, default: str = "seed_only") -> str:
            text = str(value or "").strip().lower().replace("-", "_")
            text = EVIDENCE_CLASS_ALIASES.get(text, text)
            if text in EVIDENCE_CLASS_PRIORITY:
                return text
            return default
        
        
        def evidence_weight_for_class(evidence_class: Any, parse_status: Any = "") -> float:
            normalized = normalize_evidence_class(evidence_class)
            status = str(parse_status or "").strip().lower()
            if normalized == "external_provider":
                return 1.0 if status in {"resolved", "network-fetched"} else 0.9
            if normalized == "parsed_locator":
                return 0.65
            if normalized == "repo_local_extracted":
                return 0.45
            return 0.2
        
        
        def metadata_completeness(record: Dict[str, Any]) -> int:
            score = 0
            for key in ("title", "summary", "url", "repo_full_name", "doi", "arxiv_id", "venue"):
                if record.get(key):
                    score += 1
            authors = record.get("authors")
            if isinstance(authors, list) and authors:
                score += 1
            if record.get("year"):
                score += 1
            return score
        
        
        def normalize_record(record: Dict[str, Any]) -> Dict[str, Any]:
            normalized = {**DEFAULT_RECORD_FIELDS, **record}
            normalized["evidence_class"] = normalize_evidence_class(normalized.get("evidence_class"))
            normalized["evidence_weight"] = round(
                float(
                    normalized.get("evidence_weight")
                    or evidence_weight_for_class(normalized["evidence_class"], normalized.get("parse_status"))
                ),
                4,
            )
            normalized["authors"] = [str(item) for item in normalized.get("authors", []) if str(item).strip()]
            normalized["origins"] = [str(item) for item in normalized.get("origins", []) if str(item).strip()]
            normalized["extracted_from_repo_paths"] = [
                str(item).replace("\\", "/")
                for item in normalized.get("extracted_from_repo_paths", [])
                if str(item).strip()
            ]
            normalized["selection_hints"] = [str(item) for item in normalized.get("selection_hints", []) if str(item).strip()]
            return normalized
        
        
        def record_priority(record: Dict[str, Any]) -> tuple[int, int]:
            normalized = normalize_record(record)
            return (
                EVIDENCE_CLASS_PRIORITY.get(normalized["evidence_class"], 0),
                metadata_completeness(normalized),
            )
        
      • repo_extractors.py 3 KB
        """Repo-local source extraction for free-first research lookup."""
        
        from __future__ import annotations
        
        from pathlib import Path
        from typing import Any, Dict, List
        
        from .normalizers import ARXIV_ID_RE, DOI_RE, detect_locator, extract_urls
        
        
        IGNORED_DIRS = {
            ".git",
            ".venv",
            "__pycache__",
            "node_modules",
            "dist",
            "build",
            ".mypy_cache",
            ".pytest_cache",
        }
        
        
        def _candidate_paths(repo_path: Path) -> List[Path]:
            patterns = [
                "README*",
                "*.md",
                "*.rst",
                "*.yaml",
                "*.yml",
                "*.toml",
                "*.ini",
                "*.py",
            ]
            results: List[Path] = []
            seen: set[Path] = set()
            for pattern in patterns:
                for path in repo_path.rglob(pattern):
                    if not path.is_file():
                        continue
                    if any(part in IGNORED_DIRS for part in path.relative_to(repo_path).parts):
                        continue
                    if path.stat().st_size > 256_000:
                        continue
                    if path in seen:
                        continue
                    seen.add(path)
                    results.append(path)
            return sorted(results)[:120]
        
        
        def _extract_locators(text: str) -> List[str]:
            found: List[str] = []
            for url in extract_urls(text):
                if url not in found:
                    found.append(url)
            for pattern in (ARXIV_ID_RE, DOI_RE):
                for match in pattern.finditer(text):
                    raw = match.group(0).strip()
                    if raw and raw not in found:
                        found.append(raw)
            return found
        
        
        def _classify_kind(locator: str) -> str:
            parsed = detect_locator(locator)
            if not parsed:
                return "web"
            source_type = parsed.get("source_type")
            if source_type in {"paper", "repo", "web"}:
                return str(source_type)
            return "web"
        
        
        def extract_repo_local_seeds(repo_path: Path) -> List[Dict[str, Any]]:
            repo_root = Path(repo_path).resolve()
            seeds: List[Dict[str, Any]] = []
            seen_locators: set[str] = set()
            for path in _candidate_paths(repo_root):
                try:
                    text = path.read_text(encoding="utf-8", errors="ignore")
                except OSError:
                    continue
                locators = _extract_locators(text)
                relative_path = path.relative_to(repo_root).as_posix()
                for locator in locators:
                    if locator in seen_locators:
                        continue
                    seen_locators.add(locator)
                    seeds.append(
                        {
                            "kind": _classify_kind(locator),
                            "title": locator,
                            "summary": f"Repo-local extracted source from `{relative_path}`.",
                            "query": locator,
                            "source_url": locator if locator.lower().startswith("http") else "",
                            "source_repo": "",
                            "source_file": "",
                            "source_symbol": "",
                            "origin": "repo_local_extracted",
                            "raw_locator": locator,
                            "extracted_from_repo_paths": [relative_path],
                        }
                    )
            return seeds
        
      • source_support.py 4.4 KB
        """Machine-readable source support artifacts for downstream passes."""
        
        from __future__ import annotations
        
        import json
        import re
        from pathlib import Path
        from typing import Any, Dict, List, Sequence
        
        from .record_schema import normalize_evidence_class
        
        
        def _tokenize(value: Any) -> List[str]:
            return [token for token in re.split(r"[^a-z0-9]+", str(value or "").lower()) if len(token) > 2]
        
        
        def _record_haystack(record: Dict[str, Any]) -> str:
            return " ".join(
                [
                    str(record.get("title") or ""),
                    str(record.get("summary") or ""),
                    str(record.get("url") or ""),
                    str(record.get("repo_full_name") or ""),
                    str(record.get("doi") or ""),
                    str(record.get("arxiv_id") or ""),
                    str(record.get("source_repo") or ""),
                    str(record.get("source_file") or ""),
                    str(record.get("source_symbol") or ""),
                ]
            ).lower()
        
        
        def _match_records(tokens: Sequence[str], records: Sequence[Dict[str, Any]]) -> List[Dict[str, Any]]:
            matches: List[tuple[int, float, Dict[str, Any]]] = []
            for record in records:
                haystack = _record_haystack(record)
                score = sum(1 for token in tokens if token in haystack)
                if score > 0:
                    matches.append((score, float(record.get("evidence_weight") or 0.0), record))
            matches.sort(key=lambda item: (-item[0], -item[1], item[2].get("source_id", "")))
            return [record for _score, _weight, record in matches[:6]]
        
        
        def build_source_support(
            campaign: Dict[str, Any],
            records: Sequence[Dict[str, Any]],
            repo_local_extractions: Sequence[Dict[str, Any]],
            cache_stats: Dict[str, Any],
        ) -> Dict[str, Any]:
            by_evidence_class: Dict[str, List[str]] = {}
            by_type: Dict[str, List[str]] = {}
            for item in records:
                evidence = normalize_evidence_class(item.get("evidence_class"))
                by_evidence_class.setdefault(evidence, []).append(str(item.get("source_id")))
                source_type = str(item.get("source_type") or "unknown")
                by_type.setdefault(source_type, []).append(str(item.get("source_id")))
        
            support_index_by_candidate_idea: Dict[str, Dict[str, Any]] = {}
            support_index_by_target_component: Dict[str, Dict[str, Any]] = {}
            for idea in campaign.get("candidate_ideas", []):
                idea_id = str(idea.get("id") or "idea")
                tokens = _tokenize(idea.get("summary")) + _tokenize(idea.get("target_component")) + _tokenize(idea.get("change_scope"))
                matched = _match_records(tokens, records)
                support_index_by_candidate_idea[idea_id] = {
                    "matched_source_ids": [item.get("source_id") for item in matched],
                    "matched_external_source_ids": [
                        item.get("source_id")
                        for item in matched
                        if normalize_evidence_class(item.get("evidence_class")) == "external_provider"
                    ],
                    "matched_repo_local_source_ids": [
                        item.get("source_id")
                        for item in matched
                        if normalize_evidence_class(item.get("evidence_class")) == "repo_local_extracted"
                    ],
                    "matched_parsed_locator_ids": [
                        item.get("source_id")
                        for item in matched
                        if normalize_evidence_class(item.get("evidence_class")) == "parsed_locator"
                    ],
                }
                component = str(idea.get("target_component") or "unspecified")
                support_index_by_target_component.setdefault(component, {"matched_source_ids": []})
                for source_id in support_index_by_candidate_idea[idea_id]["matched_source_ids"]:
                    if source_id not in support_index_by_target_component[component]["matched_source_ids"]:
                        support_index_by_target_component[component]["matched_source_ids"].append(source_id)
        
            return {
                "schema_version": "1.0",
                "records": list(records),
                "records_by_evidence_class": by_evidence_class,
                "records_by_type": by_type,
                "support_index_by_candidate_idea": support_index_by_candidate_idea,
                "support_index_by_target_component": support_index_by_target_component,
                "repo_local_extractions": list(repo_local_extractions),
                "cache_stats": cache_stats,
            }
        
        
        def write_source_support(analysis_output_dir: Path, support_bundle: Dict[str, Any]) -> Path:
            path = analysis_output_dir / "SOURCE_SUPPORT.json"
            path.write_text(json.dumps(support_bundle, indent=2, ensure_ascii=False), encoding="utf-8")
            return path
        
      • __init__.py 624 B
        """Internal lookup helpers for ai-research-explore."""
        
        from .cache_store import load_cache_index, store_records
        from .inventory_writer import write_source_inventory, write_sources_summary
        from .normalizers import detect_locator, ensure_http_url
        from .repo_extractors import extract_repo_local_seeds
        from .source_support import build_source_support, write_source_support
        
        __all__ = [
            "build_source_support",
            "detect_locator",
            "ensure_http_url",
            "extract_repo_local_seeds",
            "load_cache_index",
            "store_records",
            "write_source_inventory",
            "write_source_support",
            "write_sources_summary",
        ]
        
        
    • passes
      • atomic_idea_decomposition.py 12.5 KB
        """Atomic academic concept decomposition for ai-research-explore."""
        
        from __future__ import annotations
        
        import json
        import re
        from pathlib import Path
        from typing import Any, Dict, List, Sequence
        
        
        BLOCKED_SCOPES = {"unspecified", "broad_rewrite", "rewrite-everything"}
        
        
        def tokenize(value: Any) -> List[str]:
            return [token for token in re.split(r"[^a-z0-9]+", str(value or "").lower()) if len(token) > 2]
        
        
        def unique_preserving(values: Sequence[str], *, limit: int | None = None) -> List[str]:
            ordered: List[str] = []
            for value in values:
                if not value or value in ordered:
                    continue
                ordered.append(value)
                if limit is not None and len(ordered) >= limit:
                    break
            return ordered
        
        
        def humanize_slug(value: str) -> str:
            text = str(value or "").replace("_", " ").replace("-", " ").strip()
            return " ".join(part.capitalize() for part in text.split()) or "Atomic Concept"
        
        
        def classify_surface(path: str) -> str:
            lowered = str(path).lower()
            if any(token in lowered for token in ("config", ".yaml", ".yml", ".json", ".toml", ".ini")):
                return "config"
            if any(token in lowered for token in ("data", "dataset", "loader", "transform")):
                return "data interface"
            if any(token in lowered for token in ("eval", "metric", "benchmark", "validation", "test")):
                return "evaluation adapter"
            if any(token in lowered for token in ("train", "trainer", "loss", "optim", "schedule")):
                return "training"
            return "model"
        
        
        def formula_support(selected_idea: Dict[str, Any], lookup_bundle: Dict[str, Any]) -> List[Dict[str, Any]]:
            source_ids = list(selected_idea.get("source_reference", []) or [])
            records = {item.get("source_id"): item for item in lookup_bundle.get("records", []) if item.get("source_id")}
            support: List[Dict[str, Any]] = []
            for source_id in source_ids[:3]:
                record = records.get(source_id) or {}
                support.append(
                    {
                        "source_id": source_id,
                        "title": record.get("title") or "Unresolved source reference",
                        "evidence_class": record.get("evidence_class") or "unresolved",
                        "note": "Use this source only as bounded academic support; it is not a novelty proof.",
                    }
                )
            if not support:
                support.append(
                    {
                        "source_id": "none",
                        "title": "No directly matched formula-level source",
                        "evidence_class": "none",
                        "note": "Concept remains grounded in repo-local constraints rather than directly verified source material.",
                    }
                )
            return support
        
        
        def code_support(selected_idea: Dict[str, Any], source_mapping: Dict[str, Any]) -> List[Dict[str, Any]]:
            selected_source = source_mapping.get("selected_source_record", {}) or {}
            target_location_map = source_mapping.get("target_location_map", []) or []
            support: List[Dict[str, Any]] = []
            if selected_source:
                support.append(
                    {
                        "source_id": selected_source.get("source_id") or "selected-source-record",
                        "source_repo": selected_source.get("source_repo") or "",
                        "source_file": selected_source.get("source_file") or "",
                        "source_symbol": selected_source.get("source_symbol") or "",
                        "note": "Candidate source triple for a bounded transplant or adaptation path.",
                    }
                )
            if target_location_map:
                support.append(
                    {
                        "source_id": "repo-local-target-map",
                        "source_repo": "current-research",
                        "source_file": target_location_map[0].get("file") or "",
                        "source_symbol": target_location_map[0].get("target_symbol") or "",
                        "note": "Repo-local implementation target inferred from source mapping.",
                    }
                )
            if not support:
                support.append(
                    {
                        "source_id": "none",
                        "source_repo": "current-research",
                        "source_file": "",
                        "source_symbol": "",
                        "note": "No concrete code support could be resolved.",
                    }
                )
            return support
        
        
        def unit_validation_strategy(surface: str, selected_idea: Dict[str, Any], variant_spec: Dict[str, Any]) -> str:
            if surface == "config":
                return f"Verify that `{selected_idea.get('change_scope', 'candidate change')}` can be isolated through config or CLI overrides without touching the frozen eval contract."
            if surface == "evaluation adapter":
                return "Keep evaluation surfaces protected; validate only compatibility and do not change metric semantics."
            if variant_spec.get("base_command"):
                return f"Smoke the implementation through `{variant_spec['base_command']}` with a short-run gate before any broader candidate run."
            return "Use a bounded static and short-run validation path before claiming the idea is implementable."
        
        
        def implementation_risk(selected_idea: Dict[str, Any], surface: str) -> float:
            base = float(selected_idea.get("implementation_risk") or 0.4)
            if surface in {"training", "evaluation adapter"}:
                base += 0.10
            return max(0.0, min(1.0, round(base, 4)))
        
        
        def scientific_meaning_risk(selected_idea: Dict[str, Any], surface: str) -> float:
            base = float(selected_idea.get("eval_risk") or 0.3)
            if surface == "evaluation adapter":
                base += 0.20
            if surface == "training":
                base += 0.10
            return max(0.0, min(1.0, round(base, 4)))
        
        
        def build_atomic_units(
            *,
            selected_idea: Dict[str, Any],
            analysis_data: Dict[str, Any],
            source_mapping: Dict[str, Any],
            lookup_bundle: Dict[str, Any],
            variant_spec: Dict[str, Any],
        ) -> tuple[List[Dict[str, Any]], List[str]]:
            blockers: List[str] = []
            change_scope = str(selected_idea.get("change_scope") or "unspecified")
            target_component = str(selected_idea.get("target_component") or "unspecified")
            if change_scope in BLOCKED_SCOPES:
                blockers.append("selected-idea-change-scope-too-broad-for-atomic-decomposition")
            target_location_map = source_mapping.get("target_location_map", []) or []
            module_candidates = source_mapping.get("module_candidates", []) or []
            if not target_location_map and not module_candidates:
                blockers.append("no-target-surface-for-atomic-decomposition")
        
            code_files = unique_preserving(
                [
                    str(item.get("file") or "")
                    for item in target_location_map
                    if str(item.get("role") or "") == "code"
                ]
                + [str(item.get("target_file") or "") for item in module_candidates],
                limit=5,
            )
            config_files = unique_preserving(
                [
                    str(item.get("file") or "")
                    for item in target_location_map
                    if str(item.get("role") or "") == "config"
                ]
                + list(analysis_data.get("config_binding_hints", [])),
                limit=4,
            )
            code_symbols = unique_preserving(
                [str(item.get("target_symbol") or "") for item in target_location_map if item.get("target_symbol")]
                + list(analysis_data.get("constructor_candidates", []))
                + list(analysis_data.get("forward_candidates", [])),
                limit=6,
            )
        
            units: List[Dict[str, Any]] = []
            if code_files:
                surface = classify_surface(code_files[0])
                units.append(
                    {
                        "atomic_id": f"{selected_idea.get('id', 'idea')}-atomic-01",
                        "concept_name": humanize_slug(change_scope or target_component),
                        "concept_summary": str(selected_idea.get("summary") or "Bounded implementation concept"),
                        "why_needed": f"Translate the selected idea into repo-local `{surface}` logic without broad architectural rewrites.",
                        "formula_support": formula_support(selected_idea, lookup_bundle),
                        "code_support": code_support(selected_idea, source_mapping),
                        "expected_code_surface": surface,
                        "target_file_candidates": code_files,
                        "target_symbol_candidates": code_symbols or unique_preserving([target_component], limit=3),
                        "validation_strategy": unit_validation_strategy(surface, selected_idea, variant_spec),
                        "implementation_risk": implementation_risk(selected_idea, surface),
                        "scientific_meaning_risk": scientific_meaning_risk(selected_idea, surface),
                    }
                )
            if config_files:
                units.append(
                    {
                        "atomic_id": f"{selected_idea.get('id', 'idea')}-atomic-02",
                        "concept_name": f"{humanize_slug(change_scope)} Control Surface",
                        "concept_summary": f"Expose `{change_scope}` as a single-variable ablation surface rather than an entangled rewrite.",
                        "why_needed": "Keep attribution clear, rollback easy, and short-run feasibility auditable.",
                        "formula_support": formula_support(selected_idea, lookup_bundle),
                        "code_support": code_support(selected_idea, source_mapping),
                        "expected_code_surface": "config",
                        "target_file_candidates": config_files,
                        "target_symbol_candidates": unique_preserving(
                            list(analysis_data.get("config_binding_hints", [])) + [change_scope, target_component],
                            limit=6,
                        ),
                        "validation_strategy": unit_validation_strategy("config", selected_idea, variant_spec),
                        "implementation_risk": implementation_risk(selected_idea, "config"),
                        "scientific_meaning_risk": scientific_meaning_risk(selected_idea, "config"),
                    }
                )
        
            if not units:
                blockers.append("selected-idea-could-not-be-split-into-implementable-atomic-units")
            return units, unique_preserving(blockers)
        
        
        def write_atomic_markdown(output_dir: Path, payload: Dict[str, Any]) -> Path:
            lines = [
                "# Atomic Idea Map",
                "",
                f"- Status: `{payload.get('status', 'blocked')}`",
                f"- Selected idea: `{payload.get('selected_idea_id', 'none')}`",
                "",
                "## Blockers",
                "",
            ]
            blockers = payload.get("blockers", [])
            if blockers:
                lines.extend(f"- {item}" for item in blockers)
            else:
                lines.append("- None.")
            lines.extend(["", "## Atomic Units", ""])
            atomic_units = payload.get("atomic_units", [])
            if not atomic_units:
                lines.append("- None.")
            else:
                for unit in atomic_units:
                    lines.extend(
                        [
                            f"### {unit['atomic_id']} - {unit['concept_name']}",
                            "",
                            f"- Summary: {unit['concept_summary']}",
                            f"- Why needed: {unit['why_needed']}",
                            f"- Expected code surface: `{unit['expected_code_surface']}`",
                            f"- Target file candidates: {', '.join(unit.get('target_file_candidates', [])) or 'none'}",
                            f"- Target symbol candidates: {', '.join(unit.get('target_symbol_candidates', [])) or 'none'}",
                            f"- Validation strategy: {unit['validation_strategy']}",
                            f"- Implementation risk: `{unit['implementation_risk']}`",
                            f"- Scientific meaning risk: `{unit['scientific_meaning_risk']}`",
                            "",
                        ]
                    )
            path = output_dir / "ATOMIC_IDEA_MAP.md"
            path.write_text("\n".join(lines), encoding="utf-8")
            return path
        
        
        def run_atomic_idea_decomposition_pass(
            *,
            analysis_output_dir: Path,
            selected_idea: Dict[str, Any],
            analysis_data: Dict[str, Any],
            source_mapping: Dict[str, Any],
            lookup_bundle: Dict[str, Any],
            current_research: str,
            variant_spec: Dict[str, Any],
        ) -> Dict[str, Any]:
            del current_research
            atomic_units, blockers = build_atomic_units(
                selected_idea=selected_idea,
                analysis_data=analysis_data,
                source_mapping=source_mapping,
                lookup_bundle=lookup_bundle,
                variant_spec=variant_spec,
            )
            payload = {
                "schema_version": "1.0",
                "status": "blocked" if blockers else "ready",
                "selected_idea_id": str(selected_idea.get("id") or ""),
                "atomic_units": atomic_units,
                "atomic_unit_count": len(atomic_units),
                "blockers": blockers,
            }
            json_path = analysis_output_dir / "ATOMIC_IDEA_MAP.json"
            json_path.write_text(json.dumps(payload, indent=2, ensure_ascii=False), encoding="utf-8")
            markdown_path = write_atomic_markdown(analysis_output_dir, payload)
            return {
                **payload,
                "artifact_paths": [str(markdown_path), str(json_path)],
                "artifact_path": str(json_path),
            }
        
        
      • candidate_idea_generation.py 21.6 KB
        """Candidate idea generation pass for ai-research-explore."""
        
        from __future__ import annotations
        
        import json
        import re
        from pathlib import Path
        from typing import Any, Dict, List, Sequence, Tuple
        
        
        DEFAULT_POLICY = {
            "allow_synthesized_seed_ideas": True,
            "max_generated_ideas": 3,
            "require_diverse_targets": True,
        }
        REWRITE_RISK_TOKENS = {"rewrite", "architecture", "backbone", "full-model", "all-modules", "trainer-core"}
        SKIP_COMPONENT_TOKENS = {"eval", "metric", "benchmark", "leaderboard"}
        
        
        def clamp(value: Any, default: float = 0.5) -> float:
            try:
                numeric = float(value)
            except (TypeError, ValueError):
                numeric = default
            return max(0.0, min(1.0, numeric))
        
        
        def tokenize(value: Any) -> List[str]:
            return [token for token in re.split(r"[^a-z0-9]+", str(value or "").lower()) if len(token) > 2]
        
        
        def unique_preserving(values: Sequence[str], *, limit: int | None = None) -> List[str]:
            ordered: List[str] = []
            for value in values:
                if not value or value in ordered:
                    continue
                ordered.append(value)
                if limit is not None and len(ordered) >= limit:
                    break
            return ordered
        
        
        def normalize_policy(raw: Any) -> Dict[str, Any]:
            policy = dict(DEFAULT_POLICY)
            if isinstance(raw, dict):
                policy.update(raw)
            policy["allow_synthesized_seed_ideas"] = bool(policy.get("allow_synthesized_seed_ideas", True))
            try:
                policy["max_generated_ideas"] = max(0, int(policy.get("max_generated_ideas", 3)))
            except (TypeError, ValueError):
                policy["max_generated_ideas"] = 3
            policy["require_diverse_targets"] = bool(policy.get("require_diverse_targets", True))
            return policy
        
        
        def stringify_binding(value: Any) -> str:
            if isinstance(value, dict):
                for key in ("name", "id", "path", "label"):
                    if value.get(key):
                        return str(value[key])
                items = [f"{key}={value[key]}" for key in sorted(value) if value.get(key) not in {None, ""}]
                return ", ".join(items) or "unspecified"
            if isinstance(value, list):
                return ", ".join(str(item) for item in value if str(item).strip()) or "unspecified"
            text = str(value or "").strip()
            return text or "unspecified"
        
        
        def evaluation_binding_text(evaluation_source: Dict[str, Any]) -> str:
            command = str(evaluation_source.get("command") or "").strip()
            path = str(evaluation_source.get("path") or "").strip()
            metric = str(evaluation_source.get("primary_metric") or "").strip()
            parts = []
            if path:
                parts.append(f"path={path}")
            if command:
                parts.append(f"command={command}")
            if metric:
                parts.append(f"metric={metric}")
            return " | ".join(parts) or "unspecified"
        
        
        def normalize_context_bindings(
            *,
            current_research: str,
            task_family: str,
            dataset: Any,
            evaluation_source: Dict[str, Any],
        ) -> Dict[str, Any]:
            dataset_binding = stringify_binding(dataset)
            evaluation_binding = evaluation_binding_text(evaluation_source)
            task_binding = str(task_family or "").strip() or "unspecified"
            evaluation_tokens = unique_preserving(
                tokenize(evaluation_source.get("command"))
                + tokenize(evaluation_source.get("path"))
                + tokenize(evaluation_source.get("primary_metric"))
                + tokenize(evaluation_source.get("split")),
                limit=10,
            )
            task_tokens = unique_preserving(tokenize(task_binding) + tokenize(dataset_binding), limit=10)
            return {
                "context_anchor": str(current_research or "").strip() or "unspecified",
                "task_family_binding": task_binding,
                "dataset_binding": dataset_binding,
                "evaluation_binding": evaluation_binding,
                "evaluation_tokens": evaluation_tokens,
                "task_tokens": task_tokens,
            }
        
        
        def context_constraint_notes(context: Dict[str, Any], *, axis: str, target_component: str) -> List[str]:
            notes = [
                f"Anchor all generated work to current_research `{context['context_anchor']}`.",
                f"Keep the candidate inside task family `{context['task_family_binding']}` and dataset `{context['dataset_binding']}`.",
                f"Preserve the frozen evaluation binding `{context['evaluation_binding']}`.",
                f"Prefer the single-variable axis `{axis}` around `{target_component}` and keep rollback easy.",
            ]
            return unique_preserving(notes, limit=4)
        
        
        def module_component_candidates(analysis_data: Dict[str, Any]) -> List[str]:
            all_candidates: List[str] = []
            preferred_candidates: List[str] = []
            for path in analysis_data.get("module_files", []):
                stem = Path(str(path)).stem.replace("_", "-")
                if not stem:
                    continue
                all_candidates.append(stem)
                if stem not in {"model", "train", "eval"}:
                    preferred_candidates.append(stem)
            for item in analysis_data.get("constructor_candidates", []):
                token = str(item).split(":", 1)[-1].split(".", 1)[0].replace("_", "-")
                if token:
                    preferred_candidates.append(token)
                    all_candidates.append(token)
            for item in analysis_data.get("forward_candidates", []):
                token = str(item).split(":", 1)[-1].split(".", 1)[0].replace("_", "-")
                if token:
                    preferred_candidates.append(token)
                    all_candidates.append(token)
            candidates = preferred_candidates or all_candidates
            return unique_preserving(candidates, limit=6)
        
        
        def component_pool(
            researcher_candidate_ideas: Sequence[Dict[str, Any]],
            analysis_data: Dict[str, Any],
            improvement_bank: Sequence[Dict[str, Any]],
        ) -> List[str]:
            candidates: List[str] = []
            for idea in researcher_candidate_ideas:
                candidates.append(str(idea.get("target_component") or ""))
            for item in improvement_bank:
                candidates.append(str(item.get("target_component") or ""))
            candidates.extend(module_component_candidates(analysis_data))
            return unique_preserving([item for item in candidates if item and item != "unspecified"], limit=10)
        
        
        def contextual_component_pool(components: Sequence[str], context: Dict[str, Any]) -> List[str]:
            task_tokens = set(context.get("task_tokens", []))
            evaluation_tokens = set(context.get("evaluation_tokens", []))
        
            def score(component: str) -> Tuple[int, int, str]:
                tokens = set(tokenize(component))
                task_overlap = len(tokens & task_tokens)
                evaluation_overlap = len(tokens & evaluation_tokens)
                return (task_overlap, -evaluation_overlap, component)
        
            ordered = sorted(unique_preserving(components), key=score, reverse=True)
            if ordered:
                return ordered
            fallback = str(context.get("task_family_binding") or "training-config").replace(" ", "-").lower()
            if not fallback or fallback == "unspecified":
                fallback = "training-config"
            return [fallback]
        
        
        def source_hint_for_component(component: str, improvement_bank: Sequence[Dict[str, Any]], analysis_data: Dict[str, Any]) -> str:
            lowered = component.lower()
            for item in improvement_bank:
                if lowered and lowered in str(item.get("target_component") or "").lower():
                    refs = item.get("external_source_reference") or item.get("source_reference") or []
                    if refs:
                        return f"Bound to source support from {', '.join(refs[:2])}."
            for path in analysis_data.get("module_files", []):
                if lowered and lowered in str(path).lower():
                    return f"Anchored to repo-local component `{path}`."
            return "Anchored to repo-local structure and frozen evaluation constraints."
        
        
        def feasibility_hint_for_scope(change_scope: str, variant_spec: Dict[str, Any], analysis_data: Dict[str, Any]) -> str:
            if change_scope in (variant_spec.get("variant_axes") or {}):
                return f"Variant axis `{change_scope}` already exists in variant_spec, so short-run feasibility can stay command-level."
            if analysis_data.get("config_binding_hints"):
                return f"Likely feasible through existing config bindings such as `{analysis_data['config_binding_hints'][0]}`."
            return "Feasibility remains heuristic; keep the patch single-variable and reversible."
        
        
        def make_seed_id(prefix: str, index: int) -> str:
            return f"{prefix}-{index:03d}"
        
        
        def broad_rewrite_risk(change_scope: str, target_component: str) -> bool:
            tokens = set(tokenize(change_scope) + tokenize(target_component))
            return bool(tokens & REWRITE_RISK_TOKENS)
        
        
        def eval_contract_risk(target_component: str, evaluation_tokens: Sequence[str]) -> bool:
            tokens = set(tokenize(target_component))
            return bool(tokens & (SKIP_COMPONENT_TOKENS | set(evaluation_tokens)))
        
        
        def axis_seed(
            *,
            axis: str,
            axis_values: Sequence[Any],
            target_component: str,
            source_support_hint: str,
            feasibility_hint: str,
            index: int,
            seed_origin: str,
            context: Dict[str, Any],
            campaign_idea_id: str = "",
        ) -> Dict[str, Any]:
            value_summary = ", ".join(str(value) for value in list(axis_values)[:3]) or "bounded values"
            return {
                "id": make_seed_id("idea-seed", index),
                "summary": (
                    f"Probe `{axis}` as a single-variable change around `{target_component}` while keeping "
                    f"`{context['evaluation_binding']}` unchanged."
                ),
                "change_scope": axis,
                "target_component": target_component,
                "expected_upside": clamp(0.55 if seed_origin == "synthesized" else 0.60, default=0.55),
                "implementation_risk": clamp(0.28 if seed_origin == "hybrid" else 0.22, default=0.25),
                "eval_risk": clamp(0.18, default=0.18),
                "rollback_ease": clamp(0.88, default=0.88),
                "estimated_runtime_cost": clamp(0.30 if len(list(axis_values)) <= 2 else 0.38, default=0.35),
                "single_variable_fit": clamp(0.94, default=0.94),
                "seed_origin": seed_origin,
                "campaign_idea_id": campaign_idea_id or None,
                "source_support_hint": source_support_hint,
                "feasibility_hint": f"{feasibility_hint} Candidate values: {value_summary}.",
                "context_anchor": context["context_anchor"],
                "task_family_binding": context["task_family_binding"],
                "dataset_binding": context["dataset_binding"],
                "evaluation_binding": context["evaluation_binding"],
                "constraint_notes": context_constraint_notes(context, axis=axis, target_component=target_component),
            }
        
        
        def fallback_seed(
            *,
            target_component: str,
            source_support_hint: str,
            feasibility_hint: str,
            index: int,
            context: Dict[str, Any],
        ) -> Dict[str, Any]:
            axis = f"single-variable-{target_component.replace(' ', '-').lower()}"
            return {
                "id": make_seed_id("idea-seed", index),
                "summary": (
                    f"Introduce one bounded follow-up around `{target_component}` while preserving "
                    f"`{context['evaluation_binding']}` and avoiding architecture rewrites."
                ),
                "change_scope": axis,
                "target_component": target_component,
                "expected_upside": 0.50,
                "implementation_risk": 0.24,
                "eval_risk": 0.18,
                "rollback_ease": 0.90,
                "estimated_runtime_cost": 0.32,
                "single_variable_fit": 0.90,
                "seed_origin": "synthesized",
                "campaign_idea_id": None,
                "source_support_hint": source_support_hint,
                "feasibility_hint": feasibility_hint,
                "context_anchor": context["context_anchor"],
                "task_family_binding": context["task_family_binding"],
                "dataset_binding": context["dataset_binding"],
                "evaluation_binding": context["evaluation_binding"],
                "constraint_notes": context_constraint_notes(context, axis=axis, target_component=target_component),
            }
        
        
        def reject_seed(seed: Dict[str, Any], reason: str) -> Dict[str, Any]:
            rejected = dict(seed)
            rejected["rejection_reason"] = reason
            return rejected
        
        
        def existing_signatures(researcher_candidate_ideas: Sequence[Dict[str, Any]]) -> set[Tuple[str, str]]:
            return {
                (
                    str(item.get("change_scope") or "").lower(),
                    str(item.get("target_component") or "").lower(),
                )
                for item in researcher_candidate_ideas
            }
        
        
        def build_generated_ideas(
            *,
            current_research: str,
            task_family: str,
            dataset: Any,
            evaluation_source: Dict[str, Any],
            variant_spec: Dict[str, Any],
            researcher_candidate_ideas: Sequence[Dict[str, Any]],
            improvement_bank: Sequence[Dict[str, Any]],
            analysis_data: Dict[str, Any],
            policy: Dict[str, Any],
        ) -> Tuple[List[Dict[str, Any]], List[Dict[str, Any]], Dict[str, Any]]:
            context = normalize_context_bindings(
                current_research=current_research,
                task_family=task_family,
                dataset=dataset,
                evaluation_source=evaluation_source,
            )
            if not policy["allow_synthesized_seed_ideas"] or policy["max_generated_ideas"] <= 0:
                return [], [], context
        
            generated: List[Dict[str, Any]] = []
            rejected: List[Dict[str, Any]] = []
            existing = existing_signatures(researcher_candidate_ideas)
            used_targets: set[str] = set()
            axis_map = {
                key: list(value) if isinstance(value, (list, tuple)) else [value]
                for key, value in (variant_spec.get("variant_axes") or {}).items()
            }
            components = contextual_component_pool(
                component_pool(researcher_candidate_ideas, analysis_data, improvement_bank),
                context,
            )
            if context.get("task_family_binding") == "unspecified" and analysis_data.get("config_binding_hints"):
                components = unique_preserving(["training-config", *components], limit=10)
            researcher_ideas = list(researcher_candidate_ideas)
            evaluation_tokens = context.get("evaluation_tokens", [])
        
            if researcher_ideas:
                for idea in researcher_ideas:
                    for axis, axis_values in sorted(axis_map.items()):
                        if len(generated) >= policy["max_generated_ideas"]:
                            break
                        target_component = str(idea.get("target_component") or components[0] or "training-config")
                        signature = (str(axis).lower(), str(target_component).lower())
                        if signature in existing:
                            continue
                        seed = axis_seed(
                            axis=axis,
                            axis_values=list(axis_values),
                            target_component=target_component,
                            source_support_hint=source_hint_for_component(target_component, improvement_bank, analysis_data),
                            feasibility_hint=feasibility_hint_for_scope(axis, variant_spec, analysis_data),
                            index=len(generated) + 1,
                            seed_origin="hybrid",
                            context=context,
                            campaign_idea_id=str(idea.get("id") or ""),
                        )
                        if broad_rewrite_risk(seed["change_scope"], seed["target_component"]):
                            rejected.append(reject_seed(seed, "broad-architecture-rewrite-risk"))
                            continue
                        if eval_contract_risk(seed["target_component"], evaluation_tokens):
                            rejected.append(reject_seed(seed, "frozen-eval-contract-risk"))
                            continue
                        if policy["require_diverse_targets"] and seed["target_component"] in used_targets:
                            rejected.append(reject_seed(seed, "diverse-targets-required"))
                            continue
                        generated.append(seed)
                        used_targets.add(seed["target_component"])
                    if len(generated) >= policy["max_generated_ideas"]:
                        break
        
            if len(generated) < policy["max_generated_ideas"]:
                for axis, axis_values in sorted(axis_map.items()):
                    if len(generated) >= policy["max_generated_ideas"]:
                        break
                    target_component = next(
                        (
                            component
                            for component in components
                            if not (policy["require_diverse_targets"] and component in used_targets)
                            and not eval_contract_risk(component, evaluation_tokens)
                        ),
                        components[0],
                    )
                    signature = (str(axis).lower(), str(target_component).lower())
                    if signature in existing:
                        continue
                    seed = axis_seed(
                        axis=axis,
                        axis_values=list(axis_values),
                        target_component=target_component,
                        source_support_hint=source_hint_for_component(target_component, improvement_bank, analysis_data),
                        feasibility_hint=feasibility_hint_for_scope(axis, variant_spec, analysis_data),
                        index=len(generated) + 1,
                        seed_origin="synthesized",
                        context=context,
                    )
                    if broad_rewrite_risk(seed["change_scope"], seed["target_component"]):
                        rejected.append(reject_seed(seed, "broad-architecture-rewrite-risk"))
                        continue
                    if eval_contract_risk(seed["target_component"], evaluation_tokens):
                        rejected.append(reject_seed(seed, "frozen-eval-contract-risk"))
                        continue
                    if policy["require_diverse_targets"] and seed["target_component"] in used_targets:
                        rejected.append(reject_seed(seed, "diverse-targets-required"))
                        continue
                    generated.append(seed)
                    used_targets.add(seed["target_component"])
        
            if len(generated) < policy["max_generated_ideas"] and not axis_map:
                for component in components:
                    if len(generated) >= policy["max_generated_ideas"]:
                        break
                    if eval_contract_risk(component, evaluation_tokens):
                        continue
                    seed = fallback_seed(
                        target_component=component,
                        source_support_hint=source_hint_for_component(component, improvement_bank, analysis_data),
                        feasibility_hint="No explicit variant axis was provided, so this stays a repo-local bounded follow-up.",
                        index=len(generated) + 1,
                        context=context,
                    )
                    if policy["require_diverse_targets"] and seed["target_component"] in used_targets:
                        rejected.append(reject_seed(seed, "diverse-targets-required"))
                        continue
                    generated.append(seed)
                    used_targets.add(seed["target_component"])
        
            return generated, rejected, context
        
        
        def diversity_summary(researcher_ideas: Sequence[Dict[str, Any]], generated: Sequence[Dict[str, Any]]) -> Dict[str, Any]:
            all_ideas = list(researcher_ideas) + list(generated)
            targets = unique_preserving([str(item.get("target_component") or "") for item in all_ideas if item.get("target_component")])
            scopes = unique_preserving([str(item.get("change_scope") or "") for item in all_ideas if item.get("change_scope")])
            by_origin: Dict[str, int] = {}
            for item in all_ideas:
                origin = str(item.get("seed_origin") or "researcher")
                by_origin[origin] = by_origin.get(origin, 0) + 1
            return {
                "unique_target_components": targets,
                "unique_change_scopes": scopes,
                "by_seed_origin": by_origin,
                "researcher_idea_count": len(researcher_ideas),
                "generated_idea_count": len(generated),
                "synthesized_idea_count": sum(1 for item in generated if item.get("seed_origin") == "synthesized"),
                "hybrid_idea_count": sum(1 for item in generated if item.get("seed_origin") == "hybrid"),
            }
        
        
        def write_seed_artifact(
            output_dir: Path,
            *,
            policy: Dict[str, Any],
            researcher_ideas: Sequence[Dict[str, Any]],
            generated: Sequence[Dict[str, Any]],
            rejected: Sequence[Dict[str, Any]],
            diversity: Dict[str, Any],
        ) -> Path:
            payload = {
                "schema_version": "1.0",
                "generation_policy": policy,
                "researcher_ideas": list(researcher_ideas),
                "generated_ideas": list(generated),
                "all_seed_ideas": [*researcher_ideas, *generated],
                "diversity_summary": diversity,
                "rejected_seed_ideas": list(rejected),
            }
            path = output_dir / "IDEA_SEEDS.json"
            path.write_text(json.dumps(payload, indent=2, ensure_ascii=False), encoding="utf-8")
            return path
        
        
        def run_candidate_idea_generation_pass(
            *,
            analysis_output_dir: Path,
            current_research: str,
            task_family: str,
            dataset: Any,
            evaluation_source: Dict[str, Any],
            variant_spec: Dict[str, Any],
            analysis_data: Dict[str, Any],
            improvement_bank: Dict[str, Any],
            researcher_candidate_ideas: Sequence[Dict[str, Any]],
            idea_generation: Any,
        ) -> Dict[str, Any]:
            policy = normalize_policy(idea_generation)
            generated, rejected, context = build_generated_ideas(
                current_research=current_research,
                task_family=task_family,
                dataset=dataset,
                evaluation_source=evaluation_source,
                variant_spec=variant_spec,
                researcher_candidate_ideas=researcher_candidate_ideas,
                improvement_bank=improvement_bank.get("items", []),
                analysis_data=analysis_data,
                policy=policy,
            )
            diversity = diversity_summary(researcher_candidate_ideas, generated)
            path = write_seed_artifact(
                analysis_output_dir,
                policy=policy,
                researcher_ideas=researcher_candidate_ideas,
                generated=generated,
                rejected=rejected,
                diversity=diversity,
            )
            return {
                "schema_version": "1.0",
                "artifact_path": str(path),
                "generation_policy": policy,
                "researcher_ideas": list(researcher_candidate_ideas),
                "generated_ideas": generated,
                "all_seed_ideas": [*researcher_candidate_ideas, *generated],
                "diversity_summary": diversity,
                "rejected_seed_ideas": rejected,
                "context_bindings": {
                    "context_anchor": context["context_anchor"],
                    "task_family_binding": context["task_family_binding"],
                    "dataset_binding": context["dataset_binding"],
                    "evaluation_binding": context["evaluation_binding"],
                },
            }
        
      • execution_feasibility.py 19.5 KB
        """Execution feasibility and smoke-validation pass for ai-research-explore."""
        
        from __future__ import annotations
        
        import ast
        import contextlib
        import hashlib
        import importlib.util
        import io
        import os
        import platform
        import re
        import shutil
        import subprocess
        import sys
        from pathlib import Path
        from typing import Any, Dict, List, Sequence
        
        
        UNSAFE_RUNTIME_IMPORT_FILES = {
            "train.py",
            "eval.py",
            "main.py",
            "__main__.py",
        }
        
        
        def exec_module_silenced(spec: importlib.machinery.ModuleSpec, module: Any) -> None:
            # Probed repo modules may print at import time; swallow that output so the
            # orchestrator's stdout stays a clean JSON payload.
            sink = io.StringIO()
            with contextlib.redirect_stdout(sink), contextlib.redirect_stderr(sink):
                spec.loader.exec_module(module)
        
        
        def safe_float(value: Any) -> float:
            try:
                return float(value)
            except (TypeError, ValueError):
                return 0.0
        
        
        def memory_info() -> Dict[str, Any]:
            try:
                import psutil  # type: ignore
        
                vm = psutil.virtual_memory()
                return {
                    "total_gb": round(vm.total / (1024 ** 3), 2),
                    "available_gb": round(vm.available / (1024 ** 3), 2),
                    "percent_used": round(vm.percent, 2),
                }
            except Exception:
                return {
                    "total_gb": None,
                    "available_gb": None,
                    "percent_used": None,
                }
        
        
        def disk_info(root: Path) -> Dict[str, Any]:
            usage = shutil.disk_usage(root)
            return {
                "total_gb": round(usage.total / (1024 ** 3), 2),
                "available_gb": round(usage.free / (1024 ** 3), 2),
                "percent_used": round((usage.used / usage.total) * 100.0, 2) if usage.total else 0.0,
            }
        
        
        def detect_nvidia() -> List[Dict[str, Any]]:
            try:
                result = subprocess.run(
                    [
                        "nvidia-smi",
                        "--query-gpu=name,memory.total",
                        "--format=csv,noheader,nounits",
                    ],
                    check=True,
                    capture_output=True,
                    text=True,
                )
            except (FileNotFoundError, subprocess.CalledProcessError):
                return []
            gpus: List[Dict[str, Any]] = []
            for line in result.stdout.splitlines():
                if not line.strip():
                    continue
                parts = [part.strip() for part in line.split(",")]
                if not parts:
                    continue
                gpus.append(
                    {
                        "name": parts[0],
                        "memory_gb": round(safe_float(parts[1]) / 1024.0, 2) if len(parts) > 1 else None,
                        "backend": "CUDA",
                    }
                )
            return gpus
        
        
        def detect_rocm() -> List[Dict[str, Any]]:
            try:
                result = subprocess.run(
                    ["rocm-smi", "--showproductname"],
                    check=True,
                    capture_output=True,
                    text=True,
                )
            except (FileNotFoundError, subprocess.CalledProcessError):
                return []
            gpus: List[Dict[str, Any]] = []
            for line in result.stdout.splitlines():
                lowered = line.lower()
                if "card series" in lowered:
                    gpus.append({"name": line.split(":", 1)[-1].strip(), "memory_gb": None, "backend": "ROCm"})
            return gpus
        
        
        def detect_resources(root: Path) -> Dict[str, Any]:
            nvidia_gpus = detect_nvidia()
            rocm_gpus = detect_rocm()
            available_backends = sorted({gpu["backend"] for gpu in nvidia_gpus + rocm_gpus})
            return {
                "os": {
                    "system": platform.system(),
                    "release": platform.release(),
                    "machine": platform.machine(),
                    "python_version": platform.python_version(),
                },
                "cpu": {
                    "logical_cores": os.cpu_count(),
                },
                "memory": memory_info(),
                "disk": disk_info(root),
                "gpu": {
                    "nvidia_gpus": nvidia_gpus,
                    "amd_gpus": rocm_gpus,
                    "available_backends": available_backends,
                    "total_gpus": len(nvidia_gpus) + len(rocm_gpus),
                },
            }
        
        
        def parse_command_paths(command: str) -> List[str]:
            paths: List[str] = []
            for token in re.findall(r"[\w./\\-]+\.(?:py|ya?ml|json|toml|ini)", command):
                cleaned = token.strip().strip("\"'").replace("\\", "/")
                if cleaned and cleaned not in paths:
                    paths.append(cleaned)
            return paths
        
        
        def syntax_check(repo_path: Path, smoke_plan: Sequence[Dict[str, Any]]) -> Dict[str, Any]:
            python_files: List[str] = []
            for check in smoke_plan:
                if check.get("name") == "syntax-parse":
                    python_files.extend(check.get("scope", []))
            unique_files: List[str] = []
            for item in python_files:
                if item not in unique_files:
                    unique_files.append(item)
            blockers: List[str] = []
            passed: List[str] = []
            for rel in unique_files:
                path = repo_path / rel
                if not path.exists():
                    blockers.append(f"missing:{rel}")
                    continue
                try:
                    ast.parse(path.read_text(encoding="utf-8", errors="ignore"))
                    passed.append(rel)
                except SyntaxError as exc:
                    blockers.append(f"syntax:{rel}:{exc.lineno}")
            return {
                "name": "syntax-parse",
                "status": "passed" if not blockers else "failed",
                "passed": passed,
                "blockers": blockers,
            }
        
        
        def config_check(repo_path: Path, base_command: str) -> Dict[str, Any]:
            blockers: List[str] = []
            passed: List[str] = []
            for rel in parse_command_paths(base_command):
                path = repo_path / rel
                if path.exists():
                    passed.append(rel)
                else:
                    blockers.append(rel)
            return {
                "name": "config-path",
                "status": "passed" if not blockers else "failed",
                "passed": passed,
                "blockers": blockers,
            }
        
        
        def surface_check(name: str, values: Sequence[str], *, optional: bool = False) -> Dict[str, Any]:
            if values:
                return {
                    "name": name,
                    "status": "passed",
                    "passed": list(values),
                    "blockers": [],
                }
            if optional:
                return {
                    "name": name,
                    "status": "passed",
                    "passed": [],
                    "blockers": [],
                    "notes": [f"missing-{name}"],
                }
            return {
                "name": name,
                "status": "planned",
                "passed": [],
                "blockers": [f"missing-{name}"],
            }
        
        
        def import_resolution_check(target_location_map: Sequence[Dict[str, Any]]) -> Dict[str, Any]:
            scopes = [item["file"] for item in target_location_map if str(item["file"]).endswith(".py")]
            return {
                "name": "import-resolution",
                "status": "passed" if scopes else "planned",
                "passed": scopes,
                "blockers": [] if scopes else ["no-python-targets"],
            }
        
        
        def safe_runtime_targets(target_location_map: Sequence[Dict[str, Any]]) -> List[Dict[str, Any]]:
            targets: List[Dict[str, Any]] = []
            for item in target_location_map:
                file_name = Path(str(item.get("file") or "")).name.lower()
                if str(item.get("role") or "") != "code":
                    continue
                if not str(item.get("file") or "").endswith(".py"):
                    continue
                if file_name in UNSAFE_RUNTIME_IMPORT_FILES:
                    continue
                targets.append(item)
            return targets
        
        
        def import_probe_check(repo_path: Path, target_location_map: Sequence[Dict[str, Any]]) -> Dict[str, Any]:
            targets = safe_runtime_targets(target_location_map)
            if not targets:
                return {
                    "name": "import-probe",
                    "status": "passed",
                    "passed": [],
                    "blockers": [],
                    "notes": ["no-safe-import-targets"],
                }
            passed: List[str] = []
            blockers: List[str] = []
            sys_path_added = False
            repo_root = str(repo_path)
            if repo_root not in sys.path:
                sys.path.insert(0, repo_root)
                sys_path_added = True
            try:
                for item in targets:
                    rel = str(item.get("file") or "")
                    module_path = repo_path / rel
                    if not module_path.exists():
                        blockers.append(f"missing:{rel}")
                        continue
                    module_name = f"_research_explore_smoke_{hashlib.sha1(rel.encode('utf-8')).hexdigest()[:12]}"
                    try:
                        spec = importlib.util.spec_from_file_location(module_name, module_path)
                        if spec is None or spec.loader is None:
                            blockers.append(f"import-spec:{rel}")
                            continue
                        module = importlib.util.module_from_spec(spec)
                        exec_module_silenced(spec, module)
                        passed.append(rel)
                    except ModuleNotFoundError as exc:
                        blockers.append(f"missing-dependency:{rel}:{exc.name or 'unknown'}")
                    except Exception as exc:  # pragma: no cover - defensive, exercised via repo fixtures
                        blockers.append(f"import-error:{rel}:{exc.__class__.__name__}")
                    finally:
                        sys.modules.pop(module_name, None)
            finally:
                if sys_path_added:
                    try:
                        sys.path.remove(repo_root)
                    except ValueError:
                        pass
            hard_blockers = [item for item in blockers if not item.startswith("missing-dependency:")]
            soft_blockers = [item for item in blockers if item.startswith("missing-dependency:")]
            return {
                "name": "import-probe",
                "status": "failed" if hard_blockers else "planned" if soft_blockers else "passed",
                "passed": passed,
                "blockers": hard_blockers,
                "notes": soft_blockers,
            }
        
        
        def constructor_probe_check(repo_path: Path, target_location_map: Sequence[Dict[str, Any]]) -> Dict[str, Any]:
            targets = safe_runtime_targets(target_location_map)
            if not targets:
                return {
                    "name": "constructor-probe",
                    "status": "passed",
                    "passed": [],
                    "blockers": [],
                    "notes": ["constructor-probe-not-applicable"],
                }
            passed: List[str] = []
            blockers: List[str] = []
            soft_notes: List[str] = []
            sys_path_added = False
            repo_root = str(repo_path)
            if repo_root not in sys.path:
                sys.path.insert(0, repo_root)
                sys_path_added = True
            try:
                for item in targets:
                    rel = str(item.get("file") or "")
                    target_symbol = str(item.get("target_symbol") or "")
                    symbol_root = target_symbol
                    if ":" in symbol_root:
                        symbol_root = symbol_root.split(":", 1)[1]
                    symbol_root = symbol_root.split(".", 1)[0].strip()
                    if not symbol_root or symbol_root == "unspecified-symbol":
                        soft_notes.append(f"unresolved-target-symbol:{rel}")
                        continue
                    module_path = repo_path / rel
                    module_name = f"_research_explore_ctor_{hashlib.sha1(rel.encode('utf-8')).hexdigest()[:12]}"
                    try:
                        spec = importlib.util.spec_from_file_location(module_name, module_path)
                        if spec is None or spec.loader is None:
                            blockers.append(f"constructor-spec:{rel}")
                            continue
                        module = importlib.util.module_from_spec(spec)
                        exec_module_silenced(spec, module)
                        if hasattr(module, symbol_root):
                            passed.append(f"{rel}:{symbol_root}")
                        else:
                            blockers.append(f"missing-symbol:{rel}:{symbol_root}")
                    except ModuleNotFoundError as exc:
                        soft_notes.append(f"missing-dependency:{rel}:{exc.name or 'unknown'}")
                    except Exception as exc:  # pragma: no cover - defensive, exercised via repo fixtures
                        blockers.append(f"constructor-error:{rel}:{exc.__class__.__name__}")
                    finally:
                        sys.modules.pop(module_name, None)
            finally:
                if sys_path_added:
                    try:
                        sys.path.remove(repo_root)
                    except ValueError:
                        pass
            return {
                "name": "constructor-probe",
                "status": "failed" if blockers else "passed",
                "passed": passed,
                "blockers": blockers,
                "notes": soft_notes,
            }
        
        
        def short_run_check(executed_runs: Sequence[Dict[str, Any]], variant_matrix: Dict[str, Any]) -> Dict[str, Any]:
            if executed_runs:
                statuses = [item.get("status", "unknown") for item in executed_runs]
                return {
                    "name": "short-run-command",
                    "status": "passed" if any(status in {"success", "partial"} for status in statuses) else "failed",
                    "passed": [item.get("id", "unknown") for item in executed_runs],
                    "blockers": [] if any(status in {"success", "partial"} for status in statuses) else statuses,
                }
            if variant_matrix.get("base_command"):
                return {
                    "name": "short-run-command",
                    "status": "planned",
                    "passed": [],
                    "blockers": ["not-executed-yet"],
                }
            return {
                "name": "short-run-command",
                "status": "failed",
                "passed": [],
                "blockers": ["missing-base-command"],
            }
        
        
        def recommend_strategy(resources: Dict[str, Any]) -> Dict[str, Any]:
            logical_cores = resources["cpu"].get("logical_cores") or 1
            available_memory = resources["memory"].get("available_gb") or 0.0
            backends = resources["gpu"].get("available_backends", [])
            if logical_cores >= 8:
                parallel_strategy = "high-parallelism"
                suggested_workers = max(1, logical_cores - 2)
            elif logical_cores >= 4:
                parallel_strategy = "moderate-parallelism"
                suggested_workers = max(1, logical_cores - 1)
            else:
                parallel_strategy = "low-parallelism"
                suggested_workers = 1
            memory_strategy = "memory-abundant" if available_memory >= 16 else "moderate-memory" if available_memory >= 4 else "memory-constrained"
            acceleration = (
                f"Use {', '.join(backends)} acceleration for short-run probes."
                if backends
                else "No GPU backend detected; keep early exploratory runs small and CPU-safe."
            )
            return {
                "parallel_strategy": parallel_strategy,
                "suggested_workers": suggested_workers,
                "memory_strategy": memory_strategy,
                "acceleration_suggestion": acceleration,
            }
        
        
        def feasibility_decision(
            *,
            campaign: Dict[str, Any],
            variant_matrix: Dict[str, Any],
            resources: Dict[str, Any],
        ) -> Dict[str, Any]:
            budget_hours = safe_float((campaign.get("compute_budget") or {}).get("max_runtime_hours")) or 0.0
            executed_budget = safe_float((campaign.get("execution_policy") or {}).get("max_executed_variants")) * safe_float(
                (campaign.get("execution_policy") or {}).get("variant_timeout")
            )
            estimated_hours = executed_budget / 3600.0 if executed_budget else 0.0
            short_run_status = "proceed"
            full_run_status = "proceed"
            blockers: List[str] = []
            if not variant_matrix.get("base_command"):
                short_run_status = "blocked"
                full_run_status = "blocked"
                blockers.append("missing-base-command")
            if budget_hours and estimated_hours > budget_hours:
                full_run_status = "borderline"
            if resources["gpu"].get("total_gpus", 0) == 0 and variant_matrix.get("variant_count", 0) > 2:
                full_run_status = "borderline"
            return {
                "short_run_feasibility": short_run_status,
                "full_run_feasibility": full_run_status,
                "estimated_short_run_hours": round(estimated_hours, 4),
                "budget_hours": budget_hours,
                "blockers": blockers,
            }
        
        
        def write_resource_plan(output_dir: Path, resources: Dict[str, Any], recommendations: Dict[str, Any], feasibility: Dict[str, Any]) -> Path:
            lines = [
                "# Resource Plan",
                "",
                f"- OS: `{resources['os']['system']} {resources['os']['release']}`",
                f"- CPU logical cores: `{resources['cpu'].get('logical_cores')}`",
                f"- Memory available (GB): `{resources['memory'].get('available_gb')}`",
                f"- Disk available (GB): `{resources['disk'].get('available_gb')}`",
                f"- GPU backends: `{', '.join(resources['gpu'].get('available_backends', [])) or 'none'}`",
                f"- Short-run feasibility: `{feasibility['short_run_feasibility']}`",
                f"- Full-run feasibility: `{feasibility['full_run_feasibility']}`",
                "",
                "## Recommendations",
                "",
                f"- Parallel strategy: `{recommendations['parallel_strategy']}` with `{recommendations['suggested_workers']}` workers",
                f"- Memory strategy: `{recommendations['memory_strategy']}`",
                f"- Acceleration: {recommendations['acceleration_suggestion']}",
                "",
            ]
            if feasibility["blockers"]:
                lines.extend(["## Blockers", "", *[f"- {item}" for item in feasibility["blockers"]], ""])
            path = output_dir / "RESOURCE_PLAN.md"
            path.write_text("\n".join(lines), encoding="utf-8")
            return path
        
        
        def summarize_smoke(checks: Sequence[Dict[str, Any]], ignored_blockers: Sequence[str]) -> Dict[str, Any]:
            blockers = [
                blocker
                for item in checks
                for blocker in item.get("blockers", [])
                if blocker not in ignored_blockers
            ]
            statuses = {item["status"] for item in checks}
            if statuses <= {"passed"}:
                status = "passed"
            elif statuses <= {"passed", "planned"}:
                status = "planned"
            else:
                status = "failed"
            return {
                "checks": list(checks),
                "status": status,
                "blockers": blockers,
            }
        
        
        def run_execution_feasibility_pass(
            *,
            analysis_output_dir: Path,
            repo_path: Path,
            campaign: Dict[str, Any],
            analysis_data: Dict[str, Any],
            variant_matrix: Dict[str, Any],
            source_mapping: Dict[str, Any],
            executed_runs: Sequence[Dict[str, Any]],
        ) -> Dict[str, Any]:
            resources = detect_resources(analysis_output_dir.parent)
            recommendations = recommend_strategy(resources)
            feasibility = feasibility_decision(campaign=campaign, variant_matrix=variant_matrix, resources=resources)
            static_checks = [
                syntax_check(repo_path, source_mapping.get("smoke_plan", [])),
                import_resolution_check(source_mapping.get("target_location_map", [])),
                config_check(repo_path, str(variant_matrix.get("base_command") or "")),
                surface_check("constructor-surface", analysis_data.get("constructor_candidates", [])[:4], optional=True),
                surface_check("forward-surface", analysis_data.get("forward_candidates", [])[:4], optional=True),
            ]
            runtime_checks = [
                import_probe_check(repo_path, source_mapping.get("target_location_map", [])),
                constructor_probe_check(repo_path, source_mapping.get("target_location_map", [])),
                short_run_check(executed_runs, variant_matrix),
            ]
            static_smoke = summarize_smoke(
                static_checks,
                ignored_blockers=("no-python-targets", "missing-constructor-surface", "missing-forward-surface"),
            )
            runtime_smoke = summarize_smoke(runtime_checks, ignored_blockers=("not-executed-yet",))
            overall_status = "failed"
            if static_smoke["status"] == "passed" and runtime_smoke["status"] == "passed":
                overall_status = "passed"
            elif static_smoke["status"] in {"passed", "planned"} and runtime_smoke["status"] in {"passed", "planned"}:
                overall_status = "planned"
            smoke_report = {
                "static_smoke": static_smoke,
                "runtime_smoke": runtime_smoke,
                "status": overall_status,
                "blockers": [*static_smoke["blockers"], *runtime_smoke["blockers"]],
            }
            resource_plan_path = write_resource_plan(analysis_output_dir, resources, recommendations, feasibility)
            return {
                "schema_version": "1.0",
                "artifact_path": str(resource_plan_path),
                "resources": resources,
                "recommendations": recommendations,
                "feasibility": feasibility,
                "static_smoke": static_smoke,
                "runtime_smoke": runtime_smoke,
                "smoke_report": smoke_report,
            }
        
        
      • idea_cards.py 1.7 KB
        """Hypothesis-card pass for ai-research-explore."""
        
        from __future__ import annotations
        
        import json
        from pathlib import Path
        from typing import Any, Dict, List, Sequence
        
        
        REQUIRED_FIELDS = [
            "id",
            "summary",
            "rationale",
            "target_component",
            "source_reference",
            "expected_upside",
            "single_variable_fit",
            "implementation_risk",
            "eval_risk",
            "rollback_ease",
            "patch_surface",
            "dependency_drag",
            "validation_path",
            "innovation_note",
        ]
        
        ALLOWED_PATCH_CLASSES = {
            "config-only",
            "import-glue",
            "module-transplant-shim",
        }
        
        
        def build_cards(improvement_items: Sequence[Dict[str, Any]]) -> List[Dict[str, Any]]:
            cards: List[Dict[str, Any]] = []
            for item in improvement_items:
                card = dict(item)
                missing = [field for field in REQUIRED_FIELDS if field not in card]
                if missing:
                    raise ValueError(f"Improvement item `{item.get('id', 'unknown')}` is missing required card fields: {missing}")
                patch_class = str(card.get("patch_class") or "").strip().lower()
                card["patch_class"] = patch_class if patch_class in ALLOWED_PATCH_CLASSES else "config-only"
                card["patch_class_source"] = "campaign" if patch_class in ALLOWED_PATCH_CLASSES else "default"
                card.setdefault("short_run_feasibility", "plausible")
                cards.append(card)
            return cards
        
        
        def run_idea_card_pass(*, analysis_output_dir: Path, improvement_items: Sequence[Dict[str, Any]]) -> Dict[str, Any]:
            cards = build_cards(improvement_items)
            path = analysis_output_dir / "IDEA_CARDS.json"
            path.write_text(json.dumps(cards, indent=2, ensure_ascii=False), encoding="utf-8")
            return {
                "schema_version": "1.0",
                "artifact_path": str(path),
                "cards": cards,
            }
        
        
      • idea_ranking.py 8.4 KB
        """Idea evaluation and ranking pass for ai-research-explore."""
        
        from __future__ import annotations
        
        import json
        from pathlib import Path
        from typing import Any, Dict, List, Sequence, Tuple
        
        
        POSITIVE_WEIGHTS = {
            "expected_upside": 14.0,
            "single_variable_fit": 10.0,
            "groundedness": 10.0,
            "source_support_strength": 9.0,
            "interface_fit": 9.0,
            "rollback_ease": 6.0,
            "novelty_estimate": 5.0,
            "ablation_clarity": 8.0,
            "implementation_story_clarity": 8.0,
            "execution_feasibility": 7.0,
        }
        NEGATIVE_WEIGHTS = {
            "implementation_risk": 8.0,
            "eval_risk": 7.0,
            "patch_surface": 5.0,
            "dependency_drag": 5.0,
            "execution_cost": 4.0,
            "baseline_distance": 3.0,
        }
        
        
        def clamp(value: Any, default: float = 0.5) -> float:
            try:
                numeric = float(value)
            except (TypeError, ValueError):
                numeric = default
            return max(0.0, min(1.0, numeric))
        
        
        def hard_gate_failures(card: Dict[str, Any], baseline_gate: Dict[str, Any]) -> List[str]:
            failures: List[str] = []
            if baseline_gate.get("decision") == "abandon":
                failures.append("baseline-gate-abandon")
            if clamp(card.get("single_variable_fit"), default=0.8) < 0.6:
                failures.append("single-variable-fit")
            if clamp(card.get("interface_fit"), default=0.5) < 0.5:
                failures.append("interface-fit")
            if clamp(card.get("patch_surface"), default=0.4) > 0.7:
                failures.append("patch-surface")
            if clamp(card.get("dependency_drag"), default=0.2) > 0.7:
                failures.append("dependency-drag")
            if clamp(card.get("eval_risk"), default=0.5) > 0.6:
                failures.append("eval-risk")
            if str(card.get("short_run_feasibility") or "plausible") == "blocked":
                failures.append("short-run-feasibility")
            return failures
        
        
        def normalized_score(score_points: float) -> float:
            max_positive = sum(POSITIVE_WEIGHTS.values())
            max_negative = sum(NEGATIVE_WEIGHTS.values())
            return round((score_points + max_negative) / (max_positive + max_negative), 4)
        
        
        def score_payload(value: float, weight: float, direction: str) -> Dict[str, Any]:
            signed = round(weight * value, 4)
            contribution = signed if direction == "positive" else -signed
            return {
                "value": round(value, 4),
                "weight": weight,
                "direction": direction,
                "contribution": round(contribution, 4),
            }
        
        
        def evaluate_card(card: Dict[str, Any], baseline_gate: Dict[str, Any]) -> Dict[str, Any]:
            raw_scores: Dict[str, float] = {}
            breakdown: Dict[str, Dict[str, Any]] = {}
            score_points = 0.0
            execution_feasibility = card.get("execution_feasibility_score", 1.0 - clamp(card.get("execution_cost"), default=0.5))
            for key, weight in POSITIVE_WEIGHTS.items():
                raw_value = execution_feasibility if key == "execution_feasibility" else clamp(card.get(key), default=0.5)
                raw_scores[key] = round(raw_value, 4)
                item = score_payload(raw_value, weight, "positive")
                breakdown[key] = item
                score_points += item["contribution"]
            for key, weight in NEGATIVE_WEIGHTS.items():
                raw_value = clamp(card.get(key), default=0.5)
                raw_scores[key] = round(raw_value, 4)
                item = score_payload(raw_value, weight, "negative")
                breakdown[key] = item
                score_points += item["contribution"]
            failures = hard_gate_failures(card, baseline_gate)
            evaluated = dict(card)
            evaluated["hard_gate_failures"] = failures
            evaluated["hard_gate_passed"] = not failures
            evaluated["score_inputs"] = raw_scores
            evaluated["score_breakdown"] = breakdown
            evaluated["weighted_total"] = round(score_points, 4)
            evaluated["idea_score"] = normalized_score(score_points)
            evaluated["seed_origin"] = str(card.get("seed_origin") or "researcher")
            return evaluated
        
        
        def selection_pool(eligible: Sequence[Dict[str, Any]]) -> Tuple[str, List[Dict[str, Any]], str]:
            researcher_eligible = [
                item
                for item in eligible
                if str(item.get("seed_origin") or "researcher") == "researcher"
            ]
            if researcher_eligible:
                return (
                    "researcher",
                    researcher_eligible,
                    "researcher hard precedence kept final selection inside the researcher-provided pool.",
                )
            return (
                "all-eligible",
                list(eligible),
                "No researcher idea passed hard gates, so the full eligible pool remained available.",
            )
        
        
        def ranking_sort_key(item: Dict[str, Any]) -> Tuple[int, float, float, float, float, str]:
            return (
                1 if item["hard_gate_passed"] else 0,
                item["idea_score"],
                item.get("expected_upside", 0.0),
                item.get("groundedness", 0.0),
                1.0 - item.get("implementation_risk", 1.0),
                item.get("id", ""),
            )
        
        
        def pool_priority(item: Dict[str, Any], active_pool: str) -> int:
            if active_pool == "researcher":
                return 1 if str(item.get("seed_origin") or "researcher") == "researcher" else 0
            return 1 if item.get("hard_gate_passed") else 0
        
        
        def write_evaluation_markdown(
            output_dir: Path,
            ranked_cards: Sequence[Dict[str, Any]],
            baseline_gate: Dict[str, Any],
            *,
            selected_idea: Dict[str, Any] | None,
            active_selection_pool: str,
            selection_reason: str,
        ) -> Path:
            lines = [
                "# Idea Evaluation",
                "",
                f"- Baseline gate: `{baseline_gate.get('decision', 'not-applicable')}`",
                "- Hard gates: baseline_gate != abandon, single_variable_fit >= 0.6, interface_fit >= 0.5, patch_surface <= 0.7, dependency_drag <= 0.7, eval_risk <= 0.6, short_run_feasibility != blocked.",
                "- Soft scoring uses explicit breakdown fields rather than a black-box total.",
                f"- Active selection pool: `{active_selection_pool}`",
                f"- Selection reason: {selection_reason}",
                "",
                "## Ranked Cards",
                "",
            ]
            if not ranked_cards:
                lines.append("- None.")
            else:
                for item in ranked_cards:
                    lines.append(
                        f"- `{item['id']}` origin=`{item.get('seed_origin', 'researcher')}` score=`{item['idea_score']}` hard_gate=`{item['hard_gate_passed']}` failures={','.join(item['hard_gate_failures']) or 'none'} summary={item['summary']}"
                    )
            lines.extend(["", "## Selected Idea", ""])
            if selected_idea is None:
                lines.append("- None.")
            else:
                lines.append(
                    f"- `{selected_idea['id']}` origin=`{selected_idea.get('seed_origin', 'researcher')}` score=`{selected_idea['idea_score']}`"
                )
            path = output_dir / "IDEA_EVALUATION.md"
            path.write_text("\n".join(lines), encoding="utf-8")
            return path
        
        
        def run_idea_ranking_pass(
            *,
            analysis_output_dir: Path,
            cards: Sequence[Dict[str, Any]],
            baseline_gate: Dict[str, Any],
        ) -> Dict[str, Any]:
            ranked = [evaluate_card(card, baseline_gate) for card in cards]
            eligible = [item for item in ranked if item["hard_gate_passed"]]
            active_selection_pool, active_candidates, selection_reason = selection_pool(eligible)
            active_candidates.sort(key=ranking_sort_key, reverse=True)
            selected = active_candidates[0] if active_candidates else None
            top_diff = None
            if len(active_candidates) >= 2:
                top_diff = round(active_candidates[0]["idea_score"] - active_candidates[1]["idea_score"], 4)
            ranked.sort(
                key=lambda item: (
                    pool_priority(item, active_selection_pool),
                    *ranking_sort_key(item),
                ),
                reverse=True,
            )
            if selected is not None:
                selected = dict(selected)
                selected["selection_pool"] = active_selection_pool
                selected["selection_reason"] = selection_reason
                selected["selected_via_hard_precedence"] = active_selection_pool == "researcher"
            scores_path = analysis_output_dir / "IDEA_SCORES.json"
            scores_path.write_text(json.dumps(ranked, indent=2, ensure_ascii=False), encoding="utf-8")
            markdown_path = write_evaluation_markdown(
                analysis_output_dir,
                ranked,
                baseline_gate,
                selected_idea=selected,
                active_selection_pool=active_selection_pool,
                selection_reason=selection_reason,
            )
            return {
                "schema_version": "1.0",
                "artifact_paths": [str(markdown_path), str(scores_path)],
                "ranked_ideas": ranked,
                "selected_idea": selected,
                "decision": "selected" if selected else "not-configured",
                "top_idea_score_diff": top_diff,
                "active_selection_pool": active_selection_pool,
                "selection_reason": selection_reason,
                "selected_idea_breakdown": selected.get("score_breakdown") if selected else {},
            }
        
        
      • implementation_fidelity.py 16.2 KB
        """Implementation fidelity checks for ai-research-explore."""
        
        from __future__ import annotations
        
        import json
        from pathlib import Path
        from typing import Any, Dict, List, Sequence, Tuple
        
        
        VERIFICATION_MODE_BY_LEVEL = {
            "not_checked": "not_checked",
            "planned_only": "not_checked",
            "heuristic_only": "heuristic",
            "executor_observed": "observed",
            "diff_verified": "observed",
        }
        
        
        def unique_preserving(values: Sequence[str]) -> List[str]:
            ordered: List[str] = []
            for value in values:
                text = str(value or "").strip()
                if not text or text in ordered:
                    continue
                ordered.append(text)
            return ordered
        
        
        def prefixed_site_entries(values: Sequence[str], *, label: str, source: str) -> List[Dict[str, str]]:
            entries: List[Dict[str, str]] = []
            for value in values:
                text = str(value or "").strip()
                if not text:
                    continue
                entries.append(
                    {
                        "site": f"{label}:{text}",
                        "path": text,
                        "source": source,
                    }
                )
            return entries
        
        
        def expected_files_for_unit(unit: Dict[str, Any]) -> List[str]:
            return unique_preserving(str(item) for item in unit.get("target_file_candidates", []))
        
        
        def path_matches_expected(path: str, expected_files: Sequence[str]) -> bool:
            normalized_path = str(path or "").strip().lower()
            expected_lookup = {str(item or "").strip().lower() for item in expected_files if str(item or "").strip()}
            return bool(normalized_path and normalized_path in expected_lookup)
        
        
        def filter_path_entries(entries: Sequence[Dict[str, str]], expected_files: Sequence[str]) -> List[Dict[str, str]]:
            return [entry for entry in entries if path_matches_expected(entry.get("path", ""), expected_files)]
        
        
        def collect_run_path_entries(executed_runs: Sequence[Dict[str, Any]]) -> Tuple[List[Dict[str, str]], List[Dict[str, str]]]:
            diff_entries: List[Dict[str, str]] = []
            touched_entries: List[Dict[str, str]] = []
            for item in executed_runs:
                run_id = str(item.get("id") or "executed-run")
                diff_entries.extend(
                    prefixed_site_entries(
                        item.get("changed_files", []) or [],
                        label="executor-changed-file",
                        source=f"executed_runs[{run_id}].changed_files",
                    )
                )
                diff_entries.extend(
                    prefixed_site_entries(
                        item.get("new_files", []) or [],
                        label="executor-new-file",
                        source=f"executed_runs[{run_id}].new_files",
                    )
                )
                diff_entries.extend(
                    prefixed_site_entries(
                        item.get("deleted_files", []) or [],
                        label="executor-deleted-file",
                        source=f"executed_runs[{run_id}].deleted_files",
                    )
                )
                touched_entries.extend(
                    prefixed_site_entries(
                        item.get("touched_paths", []) or [],
                        label="executor-touched-path",
                        source=f"executed_runs[{run_id}].touched_paths",
                    )
                )
            return diff_entries, touched_entries
        
        
        def planned_site_entries(
            unit: Dict[str, Any],
            *,
            source_mapping: Dict[str, Any],
            code_plan: Dict[str, Any],
        ) -> List[Dict[str, str]]:
            expected_files = expected_files_for_unit(unit)
            entries = prefixed_site_entries(
                expected_files,
                label="expected-target-file",
                source="atomic_bundle.target_file_candidates",
            )
            entries.extend(
                filter_path_entries(
                    prefixed_site_entries(
                        code_plan.get("candidate_edit_targets", []) or [],
                        label="code-plan-target",
                        source="code_plan.candidate_edit_targets",
                    ),
                    expected_files,
                )
            )
            entries.extend(
                filter_path_entries(
                    prefixed_site_entries(
                        [
                            item.get("file")
                            for item in source_mapping.get("target_location_map", []) or []
                            if str(item.get("file") or "").strip()
                        ],
                        label="source-mapping-target",
                        source="source_mapping.target_location_map",
                    ),
                    expected_files,
                )
            )
            entries.extend(
                filter_path_entries(
                    prefixed_site_entries(
                        [
                            target
                            for item in source_mapping.get("minimal_patch_plan", []) or []
                            for target in (item.get("target_files", []) or [])
                            if str(target or "").strip()
                        ],
                        label="minimal-patch-target",
                        source="source_mapping.minimal_patch_plan",
                    ),
                    expected_files,
                )
            )
            deduped: List[Dict[str, str]] = []
            seen: set[tuple[str, str]] = set()
            for entry in entries:
                marker = (entry["site"], entry["source"])
                if marker in seen:
                    continue
                deduped.append(entry)
                seen.add(marker)
            return deduped
        
        
        def heuristic_site_entries(
            unit: Dict[str, Any],
            *,
            selected_idea: Dict[str, Any],
            experiment_manifest: Dict[str, Any],
            executed_runs: Sequence[Dict[str, Any]],
        ) -> List[Dict[str, str]]:
            change_scope = str(selected_idea.get("change_scope") or "").strip()
            surface = str(unit.get("expected_code_surface") or "")
            entries: List[Dict[str, str]] = []
            if surface == "config" and change_scope:
                config_overrides = experiment_manifest.get("config_overrides", {}) or {}
                if change_scope in config_overrides:
                    entries.append(
                        {
                            "site": f"config-override:{change_scope}",
                            "path": "",
                            "source": "experiment_manifest.config_overrides",
                        }
                    )
            if surface == "config" and change_scope:
                for item in executed_runs:
                    run_id = str(item.get("id") or "executed-run")
                    axes = item.get("axes", {}) or {}
                    if change_scope in axes:
                        entries.append(
                            {
                                "site": f"executed-axis:{change_scope}",
                                "path": "",
                                "source": f"executed_runs[{run_id}].axes",
                            }
                        )
            deduped: List[Dict[str, str]] = []
            seen: set[tuple[str, str]] = set()
            for entry in entries:
                marker = (entry["site"], entry["source"])
                if marker in seen:
                    continue
                deduped.append(entry)
                seen.add(marker)
            return deduped
        
        
        def observed_site_entries(unit: Dict[str, Any], executed_runs: Sequence[Dict[str, Any]]) -> Tuple[List[Dict[str, str]], str]:
            expected_files = expected_files_for_unit(unit)
            diff_entries, touched_entries = collect_run_path_entries(executed_runs)
            matched_diff_entries = filter_path_entries(diff_entries, expected_files)
            if matched_diff_entries:
                return matched_diff_entries, "diff_verified"
            matched_touched_entries = filter_path_entries(touched_entries, expected_files)
            if matched_touched_entries:
                return matched_touched_entries, "executor_observed"
            return [], "not_checked"
        
        
        def common_failure_mode(unit: Dict[str, Any]) -> str:
            surface = str(unit.get("expected_code_surface") or "model")
            if surface == "config":
                return "runtime override exists but no repo-local config/code diff was verified"
            if surface == "evaluation adapter":
                return "metric surface drifted instead of preserving the frozen eval contract"
            if surface == "training":
                return "training hook changed without a bounded single-variable story"
            return "target module was planned, but no repo-local implementation diff was verified"
        
        
        def unit_state(
            *,
            phase: str,
            surface: str,
            planned_entries: Sequence[Dict[str, str]],
            heuristic_entries: Sequence[Dict[str, str]],
            observed_entries: Sequence[Dict[str, str]],
            observed_level: str,
            executed_runs: Sequence[Dict[str, Any]],
        ) -> Tuple[str, str, str]:
            if observed_entries:
                if observed_level == "diff_verified":
                    return (
                        "likely-implemented",
                        "diff_verified",
                        "Observed matching repo-local diff evidence for this atomic unit.",
                    )
                return (
                    "partial",
                    "executor_observed",
                    "Observed executor-emitted touched-path evidence for this atomic unit, but the change type stayed coarser than a diff-verified file class.",
                )
        
            if heuristic_entries:
                if surface == "config":
                    return (
                        "partial",
                        "heuristic_only",
                        "Runtime overrides or executed axes lined up with the planned config surface, but no repo-local diff was verified.",
                    )
                return (
                    "unclear",
                    "heuristic_only",
                    "Heuristic execution signals existed, but no repo-local implementation diff was verified.",
                )
        
            if phase == "pre-execution":
                if planned_entries:
                    return (
                        "not-started",
                        "planned_only",
                        "Only implementation expectations are available at pre-execution time.",
                    )
                return (
                    "not-started",
                    "not_checked",
                    "No implementation evidence has been checked yet.",
                )
        
            if executed_runs:
                if planned_entries:
                    return (
                        "unclear",
                        "planned_only",
                        "Execution completed, but no observed implementation evidence matched the planned sites.",
                    )
                return (
                    "unclear",
                    "not_checked",
                    "Execution completed, but no implementation evidence matched this atomic unit.",
                )
        
            if planned_entries:
                return (
                    "not-started",
                    "planned_only",
                    "Only implementation expectations are available; no executor evidence exists yet.",
                )
            return (
                "not-started",
                "not_checked",
                "No implementation evidence was available.",
            )
        
        
        def summarize_fidelity(units: Sequence[Dict[str, Any]]) -> Dict[str, Any]:
            states: Dict[str, int] = {}
            verification_levels: Dict[str, int] = {}
            verification_modes: Dict[str, int] = {}
            for unit in units:
                state = str(unit.get("fidelity_state") or "unknown")
                states[state] = states.get(state, 0) + 1
                level = str(unit.get("verification_level") or "unknown")
                verification_levels[level] = verification_levels.get(level, 0) + 1
                mode = str(unit.get("verification_mode") or "unknown")
                verification_modes[mode] = verification_modes.get(mode, 0) + 1
            return {
                "unit_count": len(units),
                "states": states,
                "verification_levels": verification_levels,
                "verification_modes": verification_modes,
            }
        
        
        def site_values(entries: Sequence[Dict[str, str]]) -> List[str]:
            return unique_preserving(entry.get("site", "") for entry in entries)
        
        
        def write_fidelity_markdown(output_dir: Path, payload: Dict[str, Any]) -> Path:
            lines = [
                "# Implementation Fidelity",
                "",
                f"- Status: `{payload.get('status', 'ready')}`",
                f"- Phase: `{payload.get('phase', 'pre-execution')}`",
                f"- Selected idea: `{payload.get('selected_idea_id', 'none')}`",
                "",
                "## Summary",
                "",
                f"- Atomic unit count: `{payload.get('fidelity_summary', {}).get('unit_count', 0)}`",
                f"- States: `{payload.get('fidelity_summary', {}).get('states', {})}`",
                f"- Verification levels: `{payload.get('fidelity_summary', {}).get('verification_levels', {})}`",
                "",
                "## Units",
                "",
            ]
            fidelity_units = payload.get("fidelity_units", [])
            if not fidelity_units:
                lines.append("- None.")
            else:
                for unit in fidelity_units:
                    lines.extend(
                        [
                            f"### {unit['atomic_id']}",
                            "",
                            f"- Expected implementation site: surface=`{unit['expected_implementation_site']['surface']}` files={', '.join(unit['expected_implementation_site'].get('files', [])) or 'none'} symbols={', '.join(unit['expected_implementation_site'].get('symbols', [])) or 'none'}",
                            f"- Planned implementation sites: {', '.join(unit.get('planned_implementation_sites', [])) or 'none'}",
                            f"- Heuristic implementation sites: {', '.join(unit.get('heuristic_implementation_sites', [])) or 'none'}",
                            f"- Observed implementation sites: {', '.join(unit.get('observed_implementation_sites', [])) or 'none'}",
                            f"- Fidelity state: `{unit['fidelity_state']}`",
                            f"- Common failure mode: {unit['common_failure_mode']}",
                            f"- Verification level: `{unit['verification_level']}`",
                            f"- Verification note: {unit['verification_note']}",
                            "",
                        ]
                    )
            path = output_dir / "IMPLEMENTATION_FIDELITY.md"
            path.write_text("\n".join(lines), encoding="utf-8")
            return path
        
        
        def run_implementation_fidelity_pass(
            *,
            analysis_output_dir: Path,
            selected_idea: Dict[str, Any],
            atomic_bundle: Dict[str, Any],
            source_mapping: Dict[str, Any],
            code_plan: Dict[str, Any],
            experiment_manifest: Dict[str, Any],
            executed_runs: Sequence[Dict[str, Any]],
            phase: str,
        ) -> Dict[str, Any]:
            fidelity_units: List[Dict[str, Any]] = []
            for unit in atomic_bundle.get("atomic_units", []):
                planned_entries = planned_site_entries(unit, source_mapping=source_mapping, code_plan=code_plan)
                heuristic_entries = heuristic_site_entries(
                    unit,
                    selected_idea=selected_idea,
                    experiment_manifest=experiment_manifest,
                    executed_runs=executed_runs,
                )
                observed_entries, observed_level = observed_site_entries(unit, executed_runs)
                fidelity_state, verification_level, verification_note = unit_state(
                    phase=phase,
                    surface=str(unit.get("expected_code_surface") or "model"),
                    planned_entries=planned_entries,
                    heuristic_entries=heuristic_entries,
                    observed_entries=observed_entries,
                    observed_level=observed_level,
                    executed_runs=executed_runs,
                )
                fidelity_units.append(
                    {
                        "atomic_id": unit.get("atomic_id"),
                        "concept_name": unit.get("concept_name"),
                        "expected_implementation_site": {
                            "surface": unit.get("expected_code_surface"),
                            "files": list(unit.get("target_file_candidates", [])),
                            "symbols": list(unit.get("target_symbol_candidates", [])),
                        },
                        "planned_implementation_sites": site_values(planned_entries),
                        "heuristic_implementation_sites": site_values(heuristic_entries),
                        "observed_implementation_sites": site_values(observed_entries),
                        "actual_observed_implementation_site": site_values(observed_entries),
                        "evidence_provenance": {
                            "planned": planned_entries,
                            "heuristic": heuristic_entries,
                            "observed": observed_entries,
                        },
                        "fidelity_state": fidelity_state,
                        "common_failure_mode": common_failure_mode(unit),
                        "verification_note": verification_note,
                        "verification_level": verification_level,
                        "verification_mode": VERIFICATION_MODE_BY_LEVEL.get(verification_level, "unknown"),
                    }
                )
            payload = {
                "schema_version": "1.0",
                "status": "blocked" if atomic_bundle.get("status") == "blocked" else "ready",
                "phase": phase,
                "selected_idea_id": str(selected_idea.get("id") or ""),
                "fidelity_units": fidelity_units,
                "fidelity_summary": summarize_fidelity(fidelity_units),
                "blockers": list(atomic_bundle.get("blockers", [])),
            }
            json_path = analysis_output_dir / "IMPLEMENTATION_FIDELITY.json"
            json_path.write_text(json.dumps(payload, indent=2, ensure_ascii=False), encoding="utf-8")
            markdown_path = write_fidelity_markdown(analysis_output_dir, payload)
            return {
                **payload,
                "artifact_paths": [str(markdown_path), str(json_path)],
                "artifact_path": str(json_path),
            }
        
      • improvement_bank.py 18.8 KB
        """Improvement mining pass for ai-research-explore."""
        
        from __future__ import annotations
        
        import re
        from pathlib import Path
        from typing import Any, Dict, List, Optional, Sequence
        
        from lookup.record_schema import normalize_evidence_class
        
        
        def tokenize(value: str) -> List[str]:
            return [token for token in re.split(r"[^a-z0-9]+", str(value).lower()) if len(token) > 2]
        
        
        def clamp(value: Any, default: float = 0.5) -> float:
            try:
                numeric = float(value)
            except (TypeError, ValueError):
                numeric = default
            return max(0.0, min(1.0, numeric))
        
        
        def record_evidence_class(record: Dict[str, Any]) -> str:
            if record.get("evidence_class"):
                return normalize_evidence_class(record.get("evidence_class"))
            return "seed_only" if str(record.get("provider_type") or "seed") == "seed" else "external_provider"
        
        
        def record_evidence_weight(record: Dict[str, Any]) -> float:
            if record.get("evidence_weight") is not None:
                return clamp(record.get("evidence_weight"), default=0.2)
            evidence = record_evidence_class(record)
            if evidence == "external_provider":
                return 1.0
            if evidence == "parsed_locator":
                return 0.65
            if evidence == "repo_local_extracted":
                return 0.45
            return 0.2
        
        
        def baseline_distance_score(baseline_gate: Dict[str, Any]) -> float:
            gap = baseline_gate.get("gap_to_sota")
            relative = baseline_gate.get("relative_gap_to_sota")
            if gap is not None:
                return clamp(float(gap) / 5.0, default=0.0)
            if relative is not None:
                return clamp(float(relative) / 0.10, default=0.0)
            return 0.0
        
        
        def match_sources(idea: Dict[str, Any], source_records: Sequence[Dict[str, Any]]) -> List[Dict[str, Any]]:
            tokens = set(
                tokenize(idea.get("summary"))
                + tokenize(idea.get("target_component"))
                + tokenize(idea.get("change_scope"))
            )
            matches: List[tuple[int, float, Dict[str, Any]]] = []
            for record in source_records:
                haystack = " ".join(
                    [
                        str(record.get("title") or ""),
                        str(record.get("summary") or ""),
                        str(record.get("query") or ""),
                    ]
                ).lower()
                score = sum(1 for token in tokens if token in haystack)
                if score > 0:
                    matches.append((score, record_evidence_weight(record), record))
            matches.sort(key=lambda item: (-item[0], -item[1], item[2].get("source_id", "")))
            return [record for _, _weight, record in matches[:4]]
        
        
        def matched_sources_for_idea(idea: Dict[str, Any], lookup_bundle: Dict[str, Any]) -> List[Dict[str, Any]]:
            source_records = lookup_bundle.get("records", [])
            support_bundle = lookup_bundle.get("support_bundle", {})
            support_index = support_bundle.get("support_index_by_candidate_idea", {})
            idea_support = support_index.get(str(idea.get("id") or ""), {})
            matched_ids = idea_support.get("matched_source_ids", [])
            lookup = {item.get("source_id"): item for item in source_records if item.get("source_id")}
            matched = [lookup[source_id] for source_id in matched_ids if source_id in lookup]
            return matched or match_sources(idea, source_records)
        
        
        def interface_fit_seed(
            idea: Dict[str, Any],
            analysis_data: Dict[str, Any],
            code_plan: Dict[str, Any],
        ) -> float:
            tokens = set(tokenize(idea.get("target_component")) + tokenize(idea.get("summary")))
            score = 0.50
            candidate_targets = code_plan.get("candidate_edit_targets", [])
            if any(any(token in path.lower() for token in tokens) for path in candidate_targets):
                score += 0.25
            symbol_hints = analysis_data.get("symbol_hints", [])
            if any(any(token in str(item).lower() for token in tokens) for item in symbol_hints):
                score += 0.20
            module_files = analysis_data.get("module_files", [])
            if any(any(token in path.lower() for token in tokens) for path in module_files):
                score += 0.15
            if analysis_data.get("constructor_candidates"):
                score += 0.05
            if analysis_data.get("forward_candidates"):
                score += 0.05
            return clamp(score, default=0.5)
        
        
        def patch_surface_score(code_plan: Dict[str, Any], idea: Dict[str, Any]) -> float:
            target_count = len(code_plan.get("candidate_edit_targets", []))
            support_count = len(idea.get("supporting_changes", []) or [])
            return clamp(0.15 + 0.08 * target_count + 0.05 * support_count, default=0.4)
        
        
        def dependency_drag_score(idea: Dict[str, Any], matched_sources: Sequence[Dict[str, Any]]) -> float:
            return clamp(
                0.10
                + 0.08 * len(idea.get("supporting_changes", []) or [])
                + 0.04 * max(0, len(matched_sources) - 1),
                default=0.2,
            )
        
        
        def source_support_strength(matched_sources: Sequence[Dict[str, Any]]) -> float:
            external_weight = sum(
                record_evidence_weight(item)
                for item in matched_sources
                if record_evidence_class(item) == "external_provider"
            )
            parsed_weight = sum(
                record_evidence_weight(item)
                for item in matched_sources
                if record_evidence_class(item) == "parsed_locator"
            )
            repo_local_weight = sum(
                record_evidence_weight(item)
                for item in matched_sources
                if record_evidence_class(item) == "repo_local_extracted"
            )
            seed_weight = sum(
                record_evidence_weight(item)
                for item in matched_sources
                if record_evidence_class(item) == "seed_only"
            )
            return clamp(
                0.08
                + 0.28 * min(external_weight, 2.0)
                + 0.12 * min(parsed_weight, 2.0)
                + 0.08 * min(repo_local_weight, 2.0)
                + 0.04 * min(seed_weight, 2.0),
                default=0.25,
            )
        
        
        def groundedness_score(
            *,
            source_support: float,
            interface_fit: float,
            single_variable_fit: float,
            eval_risk: float,
        ) -> float:
            return clamp(
                0.15
                + 0.35 * source_support
                + 0.20 * interface_fit
                + 0.20 * single_variable_fit
                + 0.10 * (1.0 - eval_risk),
                default=0.45,
            )
        
        
        def novelty_estimate(idea: Dict[str, Any], matched_sources: Sequence[Dict[str, Any]]) -> float:
            summary_tokens = set(tokenize(idea.get("summary")))
            novelty_terms = {"novel", "cross", "adapter", "transplant", "hybrid", "augment", "replace", "rank"}
            origin = str(idea.get("seed_origin") or "researcher")
            score = 0.20
            if novelty_terms & summary_tokens:
                score += 0.20
            if idea.get("target_component") and idea.get("change_scope") not in {"", "unspecified"}:
                score += 0.15
            if origin in {"synthesized", "hybrid"}:
                score += 0.10
            if matched_sources:
                score += 0.05
            return clamp(score, default=0.4)
        
        
        def ablation_clarity(
            *,
            single_variable_fit: float,
            rollback_ease: float,
            change_scope: str,
        ) -> float:
            score = 0.20 + 0.45 * single_variable_fit + 0.20 * rollback_ease
            if change_scope and change_scope != "unspecified":
                score += 0.10
            return clamp(score, default=0.5)
        
        
        def implementation_story_clarity(
            *,
            interface_fit: float,
            patch_surface: float,
            target_component: str,
            matched_sources: Sequence[Dict[str, Any]],
        ) -> float:
            score = 0.20 + 0.35 * interface_fit + 0.20 * (1.0 - patch_surface)
            if target_component and target_component != "unspecified":
                score += 0.10
            if matched_sources:
                score += 0.10
            return clamp(score, default=0.45)
        
        
        def innovation_story_strength(
            idea: Dict[str, Any],
            matched_sources: Sequence[Dict[str, Any]],
            novelty: float,
            story_clarity: float,
        ) -> float:
            summary_tokens = set(tokenize(idea.get("summary")))
            novelty_terms = {"novel", "cross", "adapter", "transplant", "hybrid", "augment", "improve", "replace"}
            base = 0.30 + 0.10 * len(matched_sources) + 0.30 * novelty + 0.20 * story_clarity
            if novelty_terms & summary_tokens:
                base += 0.10
            return clamp(base, default=0.5)
        
        
        def build_rationale(idea: Dict[str, Any], baseline_gate: Dict[str, Any], matched_sources: Sequence[Dict[str, Any]]) -> str:
            source_ids = ", ".join(item["source_id"] for item in matched_sources[:3]) or "no-source-id"
            external_ids = ", ".join(
                item["source_id"]
                for item in matched_sources
                if record_evidence_class(item) == "external_provider"
            ) or "no-external-source-id"
            if baseline_gate.get("decision") == "proceed":
                return (
                    f"Baseline gate permits follow-up work; candidate stays within the frozen evaluation contract and is "
                    f"supported by external source references {external_ids} and bounded lookup references {source_ids}."
                )
            if baseline_gate.get("decision") == "borderline":
                return (
                    f"Baseline is borderline; keep the patch surface tight and rely on source references {source_ids} "
                    f"before widening execution."
                )
            return f"Use source references {source_ids} to keep the candidate auditable and bounded."
        
        
        def build_validation_path(campaign: Dict[str, Any], idea: Dict[str, Any]) -> str:
            evaluation_source = campaign.get("evaluation_source", {})
            command = str(evaluation_source.get("command") or "frozen-eval-command")
            return f"Preserve `{command}` and verify `{idea.get('change_scope') or 'candidate change'}` via short-run gate before any wider run."
        
        
        def candidate_list(campaign: Dict[str, Any], candidate_ideas: Optional[Sequence[Dict[str, Any]]]) -> List[Dict[str, Any]]:
            if candidate_ideas is not None:
                return [dict(item) for item in candidate_ideas]
            return [dict(item) for item in campaign.get("candidate_ideas", [])]
        
        
        def build_improvement_bank(
            campaign: Dict[str, Any],
            analysis_data: Dict[str, Any],
            code_plan: Dict[str, Any],
            lookup_bundle: Dict[str, Any],
            baseline_gate: Dict[str, Any],
            candidate_ideas: Optional[Sequence[Dict[str, Any]]] = None,
        ) -> List[Dict[str, Any]]:
            bank: List[Dict[str, Any]] = []
            baseline_distance = baseline_distance_score(baseline_gate)
            for idea in candidate_list(campaign, candidate_ideas):
                matched_sources = matched_sources_for_idea(idea, lookup_bundle)
                external_source_reference = [
                    item["source_id"]
                    for item in matched_sources
                    if record_evidence_class(item) == "external_provider"
                ]
                parsed_locator_reference = [
                    item["source_id"]
                    for item in matched_sources
                    if record_evidence_class(item) == "parsed_locator"
                ]
                repo_local_source_reference = [
                    item["source_id"]
                    for item in matched_sources
                    if record_evidence_class(item) == "repo_local_extracted"
                ]
                seed_only_reference = [
                    item["source_id"]
                    for item in matched_sources
                    if record_evidence_class(item) == "seed_only"
                ]
                source_evidence_summary = {
                    "external_provider_records": len(external_source_reference),
                    "parsed_locator_records": len(parsed_locator_reference),
                    "repo_local_extracted_records": len(repo_local_source_reference),
                    "seed_only_records": len(seed_only_reference),
                    "weighted_external_support": round(
                        sum(record_evidence_weight(item) for item in matched_sources if record_evidence_class(item) == "external_provider"),
                        4,
                    ),
                    "weighted_parsed_locator_support": round(
                        sum(record_evidence_weight(item) for item in matched_sources if record_evidence_class(item) == "parsed_locator"),
                        4,
                    ),
                    "weighted_repo_local_support": round(
                        sum(record_evidence_weight(item) for item in matched_sources if record_evidence_class(item) == "repo_local_extracted"),
                        4,
                    ),
                    "weighted_seed_support": round(
                        sum(record_evidence_weight(item) for item in matched_sources if record_evidence_class(item) == "seed_only"),
                        4,
                    ),
                }
                expected_upside = clamp(idea.get("expected_upside"), default=0.5)
                single_variable_fit = clamp(idea.get("single_variable_fit"), default=0.8)
                implementation_risk = clamp(idea.get("implementation_risk"), default=0.5)
                eval_risk = clamp(idea.get("eval_risk"), default=0.5)
                rollback_ease = clamp(idea.get("rollback_ease"), default=0.5)
                execution_cost = clamp(idea.get("estimated_runtime_cost"), default=0.5)
                patch_surface = patch_surface_score(code_plan, idea)
                dependency_drag = dependency_drag_score(idea, matched_sources)
                interface_fit = interface_fit_seed(idea, analysis_data, code_plan)
                source_support = source_support_strength(matched_sources)
                groundedness = groundedness_score(
                    source_support=source_support,
                    interface_fit=interface_fit,
                    single_variable_fit=single_variable_fit,
                    eval_risk=eval_risk,
                )
                novelty = novelty_estimate(idea, matched_sources)
                story_clarity = implementation_story_clarity(
                    interface_fit=interface_fit,
                    patch_surface=patch_surface,
                    target_component=str(idea.get("target_component") or ""),
                    matched_sources=matched_sources,
                )
                ablation = ablation_clarity(
                    single_variable_fit=single_variable_fit,
                    rollback_ease=rollback_ease,
                    change_scope=str(idea.get("change_scope") or ""),
                )
                innovation_story = innovation_story_strength(idea, matched_sources, novelty, story_clarity)
                record = {
                    "id": str(idea.get("id") or "idea"),
                    "summary": str(idea.get("summary") or "Candidate improvement"),
                    "rationale": build_rationale(idea, baseline_gate, matched_sources),
                    "target_component": str(idea.get("target_component") or "unspecified"),
                    "change_scope": str(idea.get("change_scope") or "unspecified"),
                    "seed_origin": str(idea.get("seed_origin") or "researcher"),
                    "source_reference": [item["source_id"] for item in matched_sources],
                    "external_source_reference": external_source_reference,
                    "parsed_locator_reference": parsed_locator_reference,
                    "repo_local_source_reference": repo_local_source_reference,
                    "seed_only_source_reference": seed_only_reference,
                    "source_evidence_summary": source_evidence_summary,
                    "expected_upside": expected_upside,
                    "single_variable_fit": single_variable_fit,
                    "implementation_risk": implementation_risk,
                    "eval_risk": eval_risk,
                    "rollback_ease": rollback_ease,
                    "patch_surface": patch_surface,
                    "dependency_drag": dependency_drag,
                    "interface_fit": interface_fit,
                    "execution_cost": execution_cost,
                    "innovation_story_strength": innovation_story,
                    "source_support_strength": source_support,
                    "novelty_estimate": novelty,
                    "groundedness": groundedness,
                    "ablation_clarity": ablation,
                    "implementation_story_clarity": story_clarity,
                    "baseline_distance": baseline_distance,
                    "validation_path": build_validation_path(campaign, idea),
                    "innovation_note": (
                        f"Candidate-only story for `{idea.get('change_scope') or 'change'}`; do not present as verified novelty."
                    ),
                    "source_support_hint": str(idea.get("source_support_hint") or ""),
                    "feasibility_hint": str(idea.get("feasibility_hint") or ""),
                    "provenance": {
                        "campaign_idea_id": str(idea.get("campaign_idea_id") or idea.get("id") or "idea"),
                        "matched_source_ids": [item["source_id"] for item in matched_sources],
                        "matched_external_source_ids": external_source_reference,
                        "analysis_files": analysis_data.get("module_files", [])[:4],
                        "seed_origin": str(idea.get("seed_origin") or "researcher"),
                        "selection_origin": str(idea.get("selection_origin") or "campaign"),
                    },
                }
                bank.append(record)
            return bank
        
        
        def write_improvement_bank(output_dir: Path, bank: Sequence[Dict[str, Any]]) -> Path:
            lines = [
                "# Improvement Bank",
                "",
                "Structured candidate improvements for the third research scenario.",
                "",
            ]
            if not bank:
                lines.append("- None.")
            else:
                for item in bank:
                    lines.extend(
                        [
                            f"## {item['id']}",
                            "",
                            f"- Summary: {item['summary']}",
                            f"- Seed origin: `{item.get('seed_origin', 'researcher')}`",
                            f"- Target component: `{item['target_component']}`",
                            f"- Source references: {', '.join(item['source_reference']) or 'none'}",
                            f"- External source references: {', '.join(item.get('external_source_reference', [])) or 'none'}",
                            f"- Parsed locator references: {', '.join(item.get('parsed_locator_reference', [])) or 'none'}",
                            f"- Repo-local source references: {', '.join(item.get('repo_local_source_reference', [])) or 'none'}",
                            f"- Evidence summary: external={item.get('source_evidence_summary', {}).get('external_provider_records', 0)} parsed={item.get('source_evidence_summary', {}).get('parsed_locator_records', 0)} repo_local={item.get('source_evidence_summary', {}).get('repo_local_extracted_records', 0)} seed={item.get('source_evidence_summary', {}).get('seed_only_records', 0)}",
                            f"- Single-variable fit: `{item['single_variable_fit']}`",
                            f"- Interface fit: `{item['interface_fit']}`",
                            f"- Groundedness: `{item['groundedness']}`",
                            f"- Novelty estimate: `{item['novelty_estimate']}`",
                            f"- Ablation clarity: `{item['ablation_clarity']}`",
                            f"- Implementation story clarity: `{item['implementation_story_clarity']}`",
                            f"- Patch surface: `{item['patch_surface']}`",
                            f"- Dependency drag: `{item['dependency_drag']}`",
                            f"- Validation path: {item['validation_path']}",
                            f"- Rationale: {item['rationale']}",
                            "",
                        ]
                    )
            path = output_dir / "IMPROVEMENT_BANK.md"
            path.write_text("\n".join(lines), encoding="utf-8")
            return path
        
        
        def run_improvement_bank_pass(
            *,
            analysis_output_dir: Path,
            campaign: Dict[str, Any],
            analysis_data: Dict[str, Any],
            code_plan: Dict[str, Any],
            lookup_bundle: Dict[str, Any],
            baseline_gate: Dict[str, Any],
            candidate_ideas: Optional[Sequence[Dict[str, Any]]] = None,
        ) -> Dict[str, Any]:
            bank = build_improvement_bank(
                campaign,
                analysis_data,
                code_plan,
                lookup_bundle,
                baseline_gate,
                candidate_ideas=candidate_ideas,
            )
            path = write_improvement_bank(analysis_output_dir, bank)
            return {
                "schema_version": "1.0",
                "artifact_path": str(path),
                "items": bank,
            }
        
        
      • lookup_sources.py 16.2 KB
        """Free-first, cache-first research lookup pass for ai-research-explore."""
        
        from __future__ import annotations
        
        import json
        import re
        from pathlib import Path
        from typing import Any, Dict, List, Optional, Sequence
        
        from lookup import (
            build_source_support,
            detect_locator,
            ensure_http_url,
            extract_repo_local_seeds,
            store_records,
            write_source_inventory,
            write_source_support,
            write_sources_summary,
        )
        from lookup.normalizers import stable_digest
        from lookup.providers import (
            resolve_arxiv_record,
            resolve_doi_record,
            resolve_github_record,
            resolve_optional_record,
            resolve_url_record,
        )
        from lookup.record_schema import normalize_record
        
        
        def dedupe_preserving_order(items: Sequence[Dict[str, Any]]) -> List[Dict[str, Any]]:
            seen: set[str] = set()
            ordered: List[Dict[str, Any]] = []
            for item in items:
                key = stable_digest(item)
                if key in seen:
                    continue
                seen.add(key)
                ordered.append(dict(item))
            return ordered
        
        
        def command_paths(command: str) -> List[str]:
            paths: List[str] = []
            for token in re.findall(r"[\w./\\-]+\.(?:py|ya?ml|json|toml|ini|csv|md)", command):
                cleaned = token.strip().strip("\"'").replace("\\", "/")
                if cleaned and cleaned not in paths:
                    paths.append(cleaned)
            return paths
        
        
        def collect_seed_records(
            campaign: Dict[str, Any],
            analysis_data: Dict[str, Any],
            code_plan: Dict[str, Any],
        ) -> List[Dict[str, Any]]:
            evaluation_source = campaign.get("evaluation_source", {})
            benchmark = campaign.get("benchmark")
            benchmark_name = benchmark.get("name") if isinstance(benchmark, dict) else benchmark
            records: List[Dict[str, Any]] = [
                {
                    "kind": "benchmark",
                    "title": str(benchmark_name or "frozen-benchmark"),
                    "summary": f"Frozen benchmark for {campaign.get('task_family') or 'research campaign'}.",
                    "query": str(benchmark_name or campaign.get("dataset") or campaign.get("task_family") or "benchmark"),
                    "source_url": "",
                    "source_repo": "",
                    "source_file": "",
                    "source_symbol": "",
                    "origin": "seed_only",
                },
                {
                    "kind": "evaluation",
                    "title": str(evaluation_source.get("path") or "evaluation-source"),
                    "summary": str(evaluation_source.get("command") or "Frozen evaluation source."),
                    "query": str(evaluation_source.get("command") or evaluation_source.get("path") or "evaluation"),
                    "source_url": "",
                    "source_repo": "",
                    "source_file": str(evaluation_source.get("path") or ""),
                    "source_symbol": "",
                    "origin": "seed_only",
                },
            ]
            for path in command_paths(str(evaluation_source.get("command") or "")):
                records.append(
                    {
                        "kind": "module",
                        "title": path,
                        "summary": "Path referenced by the frozen evaluation contract.",
                        "query": path,
                        "source_url": "",
                        "source_repo": "",
                        "source_file": path,
                        "source_symbol": "",
                        "origin": "seed_only",
                    }
                )
            for item in campaign.get("sota_reference", []):
                records.append(
                    {
                        "kind": "paper" if item.get("source") else "benchmark",
                        "title": str(item.get("name") or "provided-sota"),
                        "summary": f"Frozen comparison entry for {item.get('metric') or 'metric'}.",
                        "query": str(item.get("name") or item.get("source") or item.get("metric") or "sota"),
                        "source_url": ensure_http_url(str(item.get("source") or "")) if item.get("source") else "",
                        "source_repo": "",
                        "source_file": "",
                        "source_symbol": "",
                        "origin": "seed_only",
                    }
                )
            for item in campaign.get("candidate_ideas", []):
                records.append(
                    {
                        "kind": "query",
                        "title": str(item.get("id") or item.get("summary") or "candidate-idea"),
                        "summary": str(item.get("summary") or ""),
                        "query": " ".join(
                            token
                            for token in [
                                str(campaign.get("task_family") or ""),
                                str(campaign.get("dataset") or ""),
                                str(item.get("target_component") or ""),
                                str(item.get("change_scope") or ""),
                                str(item.get("summary") or ""),
                            ]
                            if token
                        ),
                        "source_url": ensure_http_url(str(item.get("source") or "")) if item.get("source") else "",
                        "source_repo": str(item.get("source_repo") or ""),
                        "source_file": str(item.get("source_file") or ""),
                        "source_symbol": str(item.get("source_symbol") or ""),
                        "origin": "seed_only",
                    }
                )
            for item in code_plan.get("source_repo_refs", []):
                records.append(
                    {
                        "kind": "repo",
                        "title": str(item.get("repo") or "source-repo"),
                        "summary": str(item.get("note") or "Source repository reference for exploratory adaptation."),
                        "query": str(item.get("repo") or item.get("ref") or "repo"),
                        "source_url": ensure_http_url(str(item.get("url") or "")) if item.get("url") else "",
                        "source_repo": str(item.get("repo") or ""),
                        "source_file": "",
                        "source_symbol": "",
                        "origin": "seed_only",
                    }
                )
            for path in analysis_data.get("module_files", [])[:6]:
                records.append(
                    {
                        "kind": "module",
                        "title": str(path),
                        "summary": "Task-relevant module candidate from read-only repo analysis.",
                        "query": str(path),
                        "source_url": "",
                        "source_repo": "",
                        "source_file": str(path),
                        "source_symbol": "",
                        "origin": "seed_only",
                    }
                )
            for path in analysis_data.get("metric_files", [])[:4]:
                records.append(
                    {
                        "kind": "module",
                        "title": str(path),
                        "summary": "Metric or evaluation-related file from read-only repo analysis.",
                        "query": str(path),
                        "source_url": "",
                        "source_repo": "",
                        "source_file": str(path),
                        "source_symbol": "",
                        "origin": "seed_only",
                    }
                )
            lookup_config = campaign.get("research_lookup", {})
            if isinstance(lookup_config, dict):
                for item in lookup_config.get("seed_sources", []) or []:
                    if not isinstance(item, dict):
                        continue
                    records.append(
                        {
                            "kind": str(item.get("kind") or "paper"),
                            "title": str(item.get("title") or item.get("name") or "seed-source"),
                            "summary": str(item.get("summary") or item.get("notes") or ""),
                            "query": str(item.get("query") or item.get("title") or ""),
                            "source_url": ensure_http_url(str(item.get("url") or item.get("source") or "")) if item.get("url") or item.get("source") else "",
                            "source_repo": str(item.get("source_repo") or item.get("repo") or ""),
                            "source_file": str(item.get("source_file") or item.get("file") or ""),
                            "source_symbol": str(item.get("source_symbol") or item.get("symbol") or ""),
                            "origin": "seed_only",
                        }
                    )
                for query in lookup_config.get("queries", []) or []:
                    if not query:
                        continue
                    records.append(
                        {
                            "kind": "query",
                            "title": str(query),
                            "summary": "Explicit research lookup query provided by the campaign.",
                            "query": str(query),
                            "source_url": ensure_http_url(str(query)),
                            "source_repo": "",
                            "source_file": "",
                            "source_symbol": "",
                            "origin": "seed_only",
                        }
                    )
            return dedupe_preserving_order(records)
        
        
        def candidate_locators(raw: Dict[str, Any]) -> List[str]:
            locators: List[str] = []
            for value in [
                raw.get("raw_locator"),
                raw.get("source_url"),
                raw.get("query"),
                raw.get("title"),
            ]:
                text = str(value or "").strip()
                if text and text not in locators:
                    locators.append(text)
            return locators
        
        
        def resolve_provider_record(raw: Dict[str, Any], lookup_config: Dict[str, Any]) -> Dict[str, Any]:
            locator_info: Optional[Dict[str, Any]] = None
            for locator in candidate_locators(raw):
                locator_info = detect_locator(locator)
                if locator_info:
                    break
            if locator_info:
                optional_record = resolve_optional_record(locator_info, lookup_config)
                resolved = optional_record or {}
                if not resolved:
                    provider_type = locator_info.get("provider_type")
                    if provider_type == "github":
                        resolved = resolve_github_record(locator_info)
                    elif provider_type == "arxiv":
                        resolved = resolve_arxiv_record(locator_info)
                    elif provider_type == "doi":
                        resolved = resolve_doi_record(locator_info)
                    elif provider_type == "url":
                        resolved = resolve_url_record(locator_info)
                    else:
                        resolved = {}
                origin = str(raw.get("origin") or "seed_only")
                if resolved.get("parse_status") == "resolved":
                    evidence_class = "external_provider"
                elif origin == "repo_local_extracted":
                    evidence_class = "repo_local_extracted"
                else:
                    evidence_class = "parsed_locator"
                record = {
                    "source_type": resolved.get("source_type") or raw.get("kind") or "web",
                    "provider_type": resolved.get("provider_type") or "seed",
                    "provider_identifier": locator_info.get("identifier") or resolved.get("normalized_id") or "",
                    "provider_locator": locator_info.get("raw_locator") or "",
                    "locator_type": resolved.get("locator_type") or locator_info.get("locator_type") or "seed",
                    "raw_locator": locator_info.get("raw_locator") or "",
                    "normalized_id": resolved.get("normalized_id") or locator_info.get("normalized_id") or "",
                    "title": resolved.get("title") or raw.get("title") or "",
                    "summary": resolved.get("summary") or raw.get("summary") or "",
                    "query": str(raw.get("query") or raw.get("title") or ""),
                    "url": resolved.get("url") or locator_info.get("url") or raw.get("source_url") or "",
                    "authors": resolved.get("authors") or [],
                    "year": resolved.get("year"),
                    "venue": resolved.get("venue") or "",
                    "repo_full_name": resolved.get("repo_full_name") or raw.get("source_repo") or "",
                    "doi": resolved.get("doi") or "",
                    "arxiv_id": resolved.get("arxiv_id") or "",
                    "evidence_class": evidence_class,
                    "parse_status": resolved.get("parse_status") or "parsed-only",
                    "fetch_status": resolved.get("fetch_status") or "parsed-only",
                    "provider_metadata": resolved.get("provider_metadata") or {},
                    "source_repo": raw.get("source_repo") or resolved.get("repo_full_name") or "",
                    "source_file": raw.get("source_file") or resolved.get("source_file") or locator_info.get("source_file") or "",
                    "source_symbol": raw.get("source_symbol") or "",
                    "origins": [origin],
                    "extracted_from_repo_paths": list(raw.get("extracted_from_repo_paths") or []),
                    "selection_hints": [str(raw.get("query") or ""), str(raw.get("title") or "")],
                }
                return normalize_record(record)
        
            return normalize_record(
                {
                    "source_type": raw.get("kind") or "query",
                    "provider_type": "seed",
                    "provider_identifier": str(raw.get("query") or raw.get("title") or ""),
                    "provider_locator": str(raw.get("source_url") or ""),
                    "locator_type": "seed",
                    "raw_locator": str(raw.get("source_url") or raw.get("query") or raw.get("title") or ""),
                    "normalized_id": "",
                    "title": str(raw.get("title") or raw.get("query") or "seed-source"),
                    "summary": str(raw.get("summary") or ""),
                    "query": str(raw.get("query") or raw.get("title") or ""),
                    "url": str(raw.get("source_url") or ""),
                    "repo_full_name": str(raw.get("source_repo") or ""),
                    "evidence_class": "repo_local_extracted" if str(raw.get("origin") or "") == "repo_local_extracted" else "seed_only",
                    "parse_status": "seed-only",
                    "fetch_status": "seed-only",
                    "source_repo": str(raw.get("source_repo") or ""),
                    "source_file": str(raw.get("source_file") or ""),
                    "source_symbol": str(raw.get("source_symbol") or ""),
                    "origins": [str(raw.get("origin") or "seed_only")],
                    "extracted_from_repo_paths": list(raw.get("extracted_from_repo_paths") or []),
                    "selection_hints": [str(raw.get("query") or ""), str(raw.get("title") or "")],
                }
            )
        
        
        def run_lookup_pass(
            *,
            sources_dir: Path,
            repo_path: Path,
            analysis_output_dir: Optional[Path],
            campaign: Dict[str, Any],
            analysis_data: Dict[str, Any],
            code_plan: Dict[str, Any],
        ) -> Dict[str, Any]:
            sources_dir.mkdir(parents=True, exist_ok=True)
            output_dir = analysis_output_dir or (sources_dir.parent / "analysis_outputs")
            output_dir.mkdir(parents=True, exist_ok=True)
            lookup_config = campaign.get("research_lookup", {}) if isinstance(campaign.get("research_lookup"), dict) else {}
            seed_records = collect_seed_records(campaign, analysis_data, code_plan)
            repo_local_seeds = extract_repo_local_seeds(repo_path) if lookup_config.get("enable_repo_local_extraction", True) else []
            raw_records = dedupe_preserving_order([*seed_records, *repo_local_seeds])
            resolved_records = [resolve_provider_record(raw, lookup_config) for raw in raw_records]
            stored_bundle = store_records(sources_dir, resolved_records)
            records = stored_bundle["records"]
            cache_stats = {
                "cache_hits": stored_bundle.get("cache_hits", 0),
                "cache_misses": stored_bundle.get("cache_misses", 0),
                "merge_upgrades": stored_bundle.get("merge_upgrades", 0),
            }
            summary_path = write_sources_summary(sources_dir, records)
            inventory_path = write_source_inventory(
                output_dir,
                records=records,
                repo_local_extractions=repo_local_seeds,
                cache_stats=cache_stats,
            )
            support_bundle = build_source_support(campaign, records, repo_local_seeds, cache_stats)
            support_path = write_source_support(output_dir, support_bundle)
            records_by_evidence_class = sorted({str(item.get("evidence_class") or "") for item in records if item.get("evidence_class")})
            return {
                "schema_version": "2.0",
                "mode": "free-first-cache-first",
                "sources_dir": str(sources_dir),
                "records_dir": stored_bundle.get("records_dir"),
                "index_path": stored_bundle.get("index_path"),
                "summary_path": str(summary_path),
                "inventory_path": str(inventory_path),
                "support_path": str(support_path),
                "support_bundle": support_bundle,
                "records": records,
                "records_by_kind": sorted({str(item.get("source_type") or "") for item in records if item.get("source_type")}),
                "records_by_provider": sorted({str(item.get("provider_type") or "") for item in records if item.get("provider_type")}),
                "records_by_evidence_class": records_by_evidence_class,
                "cache_hits": stored_bundle.get("cache_hits", 0),
                "cache_misses": stored_bundle.get("cache_misses", 0),
                "repo_extracted_locators": [item.get("raw_locator") for item in repo_local_seeds],
                "repo_local_extractions": repo_local_seeds,
                "queries": [item.get("query") for item in raw_records if item.get("kind") == "query"],
                "optional_provider_used": False,
            }
        
        
      • source_mapping.py 19.1 KB
        """Source module lookup and interface diff pass for ai-research-explore."""
        
        from __future__ import annotations
        
        import re
        from pathlib import Path
        from typing import Any, Dict, List, Sequence
        
        from lookup.record_schema import normalize_evidence_class
        
        ALLOWED_PATCH_CLASSES = {
            "config-only",
            "import-glue",
            "module-transplant-shim",
        }
        
        
        def first_items(values: Sequence[Any], limit: int) -> List[Any]:
            ordered: List[Any] = []
            for item in values:
                if item not in ordered:
                    ordered.append(item)
                if len(ordered) >= limit:
                    break
            return ordered
        
        
        def select_source_record(selected_idea: Dict[str, Any], lookup_bundle: Dict[str, Any]) -> Dict[str, Any]:
            source_lookup = {item["source_id"]: item for item in lookup_bundle.get("records", [])}
            references = list(selected_idea.get("source_reference", []) or [])
            matched = [source_lookup[source_id] for source_id in references if source_id in source_lookup]
            for item in matched:
                if item.get("source_repo") and item.get("source_file") and item.get("source_symbol"):
                    return item
            tokens: List[str] = []
            for raw in [
                str(selected_idea.get("summary") or ""),
                str(selected_idea.get("target_component") or ""),
                str(selected_idea.get("change_scope") or ""),
            ]:
                for token in re.split(r"[^a-z0-9]+", raw.lower()):
                    if len(token) > 2 and token not in tokens:
                        tokens.append(token)
            scored_candidates: List[tuple[int, int, float, Dict[str, Any]]] = []
            for item in lookup_bundle.get("records", []):
                if not (item.get("source_repo") and item.get("source_file") and item.get("source_symbol")):
                    continue
                haystack = " ".join(
                    [
                        str(item.get("title") or ""),
                        str(item.get("summary") or ""),
                        str(item.get("query") or ""),
                        str(item.get("source_repo") or ""),
                        str(item.get("source_file") or ""),
                        str(item.get("source_symbol") or ""),
                    ]
                ).lower()
                score = sum(1 for token in tokens if token in haystack)
                if score > 0:
                    evidence = normalize_evidence_class(item.get("evidence_class"))
                    evidence_priority = {
                        "external_provider": 3,
                        "parsed_locator": 2,
                        "repo_local_extracted": 1,
                        "seed_only": 0,
                    }.get(evidence, 0)
                    scored_candidates.append((score, evidence_priority, float(item.get("evidence_weight") or 0.0), item))
            scored_candidates.sort(key=lambda pair: (-pair[0], -pair[1], -pair[2], pair[3].get("source_id", "")))
            if scored_candidates:
                return scored_candidates[0][3]
            return matched[0] if matched else {}
        
        
        def source_blockers(source_record: Dict[str, Any]) -> List[str]:
            blockers: List[str] = []
            if not source_record.get("source_repo"):
                blockers.append("missing-source-repo")
            if not source_record.get("source_file"):
                blockers.append("missing-source-file")
            if not source_record.get("source_symbol"):
                blockers.append("missing-source-symbol")
            return blockers
        
        
        def normalize_patch_class(value: Any) -> str:
            text = str(value or "").strip().lower()
            return text if text in ALLOWED_PATCH_CLASSES else ""
        
        
        def best_target_symbol(selected_idea: Dict[str, Any], analysis_data: Dict[str, Any]) -> str:
            component = str(selected_idea.get("target_component") or "").lower()
            for item in analysis_data.get("constructor_candidates", []):
                if component and component in str(item).lower():
                    return str(item)
            for item in analysis_data.get("forward_candidates", []):
                if component and component in str(item).lower():
                    return str(item)
            for item in first_items(analysis_data.get("constructor_candidates", []), 1):
                return str(item)
            for item in first_items(analysis_data.get("forward_candidates", []), 1):
                return str(item)
            return "unspecified-symbol"
        
        
        def build_target_location_map(
            selected_idea: Dict[str, Any],
            analysis_data: Dict[str, Any],
            code_plan: Dict[str, Any],
        ) -> List[Dict[str, Any]]:
            target_symbol = best_target_symbol(selected_idea, analysis_data)
            config_hints = list(analysis_data.get("config_binding_hints", []))
            target_component = str(selected_idea.get("target_component") or "").lower()
            prefer_config_only = "config" in target_component
            results: List[Dict[str, Any]] = []
            candidate_paths = first_items(code_plan.get("candidate_edit_targets", []), 4)
            if not config_hints:
                config_hints = [
                    path
                    for path in candidate_paths
                    if any(token in str(path).lower() for token in ("config", ".yaml", ".yml", ".json", ".toml", ".ini"))
                ]
            if prefer_config_only and config_hints:
                candidate_paths = first_items(
                    [path for path in candidate_paths if path in config_hints] or list(config_hints),
                    4,
                )
            for path in candidate_paths:
                results.append(
                    {
                        "file": path,
                        "target_symbol": target_symbol,
                        "role": "config" if path in config_hints else "code",
                        "reason": f"Selected for `{selected_idea.get('change_scope', 'candidate change')}` with target component `{selected_idea.get('target_component', 'unspecified')}`.",
                    }
                )
            return results
        
        
        def build_module_candidates(
            selected_idea: Dict[str, Any],
            lookup_bundle: Dict[str, Any],
            analysis_data: Dict[str, Any],
            code_plan: Dict[str, Any],
        ) -> List[Dict[str, Any]]:
            target_location_map = build_target_location_map(selected_idea, analysis_data, code_plan)
            target_symbol = target_location_map[0]["target_symbol"] if target_location_map else "unspecified-symbol"
            module_files = analysis_data.get("module_files", []) or code_plan.get("candidate_edit_targets", [])
            source_record = select_source_record(selected_idea, lookup_bundle)
            source_reference = [source_record.get("source_id")] if source_record.get("source_id") else []
            transplant_ready = not source_blockers(source_record)
            candidates: List[Dict[str, Any]] = []
            for path in first_items(module_files, 3):
                candidates.append(
                    {
                        "idea_id": selected_idea.get("id"),
                        "source_repo": (source_record or {}).get("source_repo") or code_plan.get("source_repo_refs", [{}])[0].get("repo", "current-research"),
                        "source_reference": source_reference,
                        "source_file": (source_record or {}).get("source_file") or "",
                        "source_symbol": (source_record or {}).get("source_symbol") or "",
                        "target_file": path,
                        "target_symbol": target_symbol,
                        "supporting_files": first_items(analysis_data.get("config_binding_hints", []), 3),
                        "source_triple_ready": transplant_ready,
                        "why_fit": (
                            f"`{path}` is in the allowed change zone and matches `{selected_idea.get('target_component', 'component')}`."
                            if transplant_ready
                            else "The target location is plausible, but the source transplant triple is incomplete."
                        ),
                        "why_not_fit": (
                            "Keep the patch reversible and avoid changing evaluation or leaderboard parsing files."
                            if transplant_ready
                            else f"Transplant path blocked: {', '.join(source_blockers(source_record))}."
                        ),
                    }
                )
            return candidates
        
        
        def build_interface_diff(
            selected_idea: Dict[str, Any],
            analysis_data: Dict[str, Any],
            module_candidates: Sequence[Dict[str, Any]],
        ) -> Dict[str, Any]:
            transplant_ready = bool(module_candidates[0].get("source_triple_ready")) if module_candidates else False
            return {
                "selected_idea": selected_idea.get("id"),
                "transplant_ready": transplant_ready,
                "constructor_surface": first_items(analysis_data.get("constructor_candidates", []), 6),
                "forward_surface": first_items(analysis_data.get("forward_candidates", []), 6),
                "config_surface": first_items(analysis_data.get("config_binding_hints", []), 6),
                "metric_surface": first_items(analysis_data.get("metric_files", []), 4),
                "required_shims": [
                    "Keep constructor wiring mechanical and reversible.",
                    "Prefer import/registry glue over behavioral rewrites.",
                    "Preserve frozen evaluation and metric parsing surfaces.",
                    *(
                        []
                        if transplant_ready
                        else ["Do not enter the transplant path until source_repo + source_file + source_symbol are all present."]
                    ),
                ],
                "candidate_targets": [item["target_file"] for item in module_candidates],
            }
        
        
        def build_minimal_patch_plan(
            selected_idea: Dict[str, Any],
            target_location_map: Sequence[Dict[str, Any]],
            interface_diff: Dict[str, Any],
            selected_source_record: Dict[str, Any],
        ) -> List[Dict[str, Any]]:
            if not target_location_map:
                return []
            plan: List[Dict[str, Any]] = []
            config_targets = [item for item in target_location_map if item["role"] == "config"]
            code_targets = [item for item in target_location_map if item["role"] == "code"]
            triple_blockers = source_blockers(selected_source_record)
            if config_targets:
                plan.append(
                    {
                        "change_type": "config-only",
                        "target_files": [item["file"] for item in config_targets],
                        "rationale": f"Expose `{selected_idea.get('change_scope', 'candidate change')}` through existing config bindings.",
                        "rollback": "Revert config keys to the baseline values.",
                        "smoke_checks": ["config-path", "short-run-command"],
                    }
                )
            if code_targets and not triple_blockers:
                plan.append(
                    {
                        "change_type": "import-glue",
                        "target_files": [code_targets[0]["file"]],
                        "rationale": "Wire the candidate module through the smallest registry or import boundary.",
                        "rollback": "Remove the candidate import and restore the baseline registry entry.",
                        "smoke_checks": ["syntax-parse", "import-resolution", "constructor-surface"],
                    }
                )
                plan.append(
                    {
                        "change_type": "module-transplant-shim",
                        "target_files": [code_targets[0]["file"]],
                        "rationale": "Add a thin compatibility shim only if constructor/forward surfaces do not match.",
                        "rollback": "Delete the shim and point the call-site back to the baseline module.",
                        "smoke_checks": ["forward-surface", "short-run-command"],
                    }
                )
            elif code_targets and triple_blockers:
                plan.append(
                    {
                        "change_type": "transplant-blocked",
                        "target_files": [code_targets[0]["file"]],
                        "rationale": "Do not enter the transplant path until source_repo + source_file + source_symbol are all available.",
                        "rollback": "No-op; keep the baseline module path unchanged.",
                        "smoke_checks": ["static-smoke-only"],
                        "blockers": triple_blockers,
                    }
                )
            if interface_diff.get("metric_surface"):
                plan.append(
                    {
                        "change_type": "protected-zone-no-touch",
                        "target_files": interface_diff["metric_surface"],
                        "rationale": "Metric and leaderboard surfaces are protected unless the campaign explicitly allows them.",
                        "rollback": "No-op; these files should remain unchanged.",
                        "smoke_checks": ["metric-surface-protected"],
                    }
                )
            return plan
        
        
        def build_smoke_plan(
            selected_idea: Dict[str, Any],
            target_location_map: Sequence[Dict[str, Any]],
            variant_matrix: Dict[str, Any],
        ) -> List[Dict[str, Any]]:
            return [
                {
                    "name": "syntax-parse",
                    "scope": [item["file"] for item in target_location_map if item["file"].endswith(".py")],
                    "reason": "Candidate patch must keep Python files parseable.",
                },
                {
                    "name": "import-resolution",
                    "scope": [item["file"] for item in target_location_map if item["file"].endswith(".py")],
                    "reason": "Candidate patch must not break module loading paths.",
                },
                {
                    "name": "config-path",
                    "scope": [item["file"] for item in target_location_map if item["role"] == "config"],
                    "reason": "Frozen command/config references must still resolve.",
                },
                {
                    "name": "constructor-surface",
                    "scope": [str(selected_idea.get("target_component") or "candidate-component")],
                    "reason": "Constructor wiring should remain mechanical and reversible.",
                },
                {
                    "name": "forward-surface",
                    "scope": [str(selected_idea.get("target_component") or "candidate-component")],
                    "reason": "Forward path should stay single-variable and attribution-friendly.",
                },
                {
                    "name": "short-run-command",
                    "scope": [str(variant_matrix.get("base_command") or "no-base-command")],
                    "reason": "Use short-run smoke before any broader candidate run.",
                },
            ]
        
        
        def resolve_patch_class(
            selected_idea: Dict[str, Any],
            minimal_patch_plan: Sequence[Dict[str, Any]],
        ) -> Dict[str, Any]:
            change_types = {str(item.get("change_type") or "") for item in minimal_patch_plan}
            requested_patch_class = normalize_patch_class(selected_idea.get("patch_class"))
            # A researcher-requested class the plan can satisfy must not be escalated
            # to the transplant path (which demands a full source triple); per
            # source-mapping-policy, config-only stays the least invasive choice.
            if requested_patch_class and requested_patch_class in change_types:
                return {
                    "requested_patch_class": requested_patch_class,
                    "resolved_patch_class": requested_patch_class,
                    "patch_class_source": "campaign",
                    "requires_source_triple": requested_patch_class == "module-transplant-shim",
                }
            if "module-transplant-shim" in change_types or "transplant-blocked" in change_types:
                return {
                    "requested_patch_class": requested_patch_class,
                    "resolved_patch_class": "module-transplant-shim",
                    "patch_class_source": "source-mapping",
                    "requires_source_triple": True,
                }
            if "import-glue" in change_types:
                return {
                    "requested_patch_class": requested_patch_class,
                    "resolved_patch_class": "import-glue",
                    "patch_class_source": "source-mapping",
                    "requires_source_triple": False,
                }
            return {
                "requested_patch_class": requested_patch_class,
                "resolved_patch_class": requested_patch_class or "config-only",
                "patch_class_source": "campaign" if requested_patch_class else "source-mapping",
                "requires_source_triple": False,
            }
        
        
        def write_module_candidates(output_dir: Path, module_candidates: Sequence[Dict[str, Any]]) -> Path:
            lines = [
                "# Module Candidates",
                "",
            ]
            if not module_candidates:
                lines.append("- None.")
            else:
                for item in module_candidates:
                    lines.extend(
                        [
                            f"## {item['idea_id']} -> {item['target_file']}",
                            "",
                            f"- Source repo: `{item['source_repo']}`",
                            f"- Source file: `{item['source_file']}`",
                            f"- Source symbol: `{item['source_symbol']}`",
                            f"- Target symbol: `{item['target_symbol']}`",
                            f"- Supporting files: {', '.join(item['supporting_files']) or 'none'}",
                            f"- Why fit: {item['why_fit']}",
                            f"- Why not fit: {item['why_not_fit']}",
                            "",
                        ]
                    )
            path = output_dir / "MODULE_CANDIDATES.md"
            path.write_text("\n".join(lines), encoding="utf-8")
            return path
        
        
        def write_interface_diff(output_dir: Path, interface_diff: Dict[str, Any]) -> Path:
            constructor_surface = [f"- {item}" for item in interface_diff.get("constructor_surface", [])] or ["- none"]
            forward_surface = [f"- {item}" for item in interface_diff.get("forward_surface", [])] or ["- none"]
            config_surface = [f"- {item}" for item in interface_diff.get("config_surface", [])] or ["- none"]
            required_shims = [f"- {item}" for item in interface_diff.get("required_shims", [])] or ["- none"]
            lines = [
                "# Interface Diff",
                "",
                f"- Selected idea: `{interface_diff.get('selected_idea', 'none')}`",
                "",
                "## Constructor Surface",
                "",
                *constructor_surface,
                "",
                "## Forward Surface",
                "",
                *forward_surface,
                "",
                "## Config Surface",
                "",
                *config_surface,
                "",
                "## Required Shims",
                "",
                *required_shims,
                "",
            ]
            path = output_dir / "INTERFACE_DIFF.md"
            path.write_text("\n".join(lines), encoding="utf-8")
            return path
        
        
        def run_source_mapping_pass(
            *,
            analysis_output_dir: Path,
            selected_idea: Dict[str, Any],
            analysis_data: Dict[str, Any],
            code_plan: Dict[str, Any],
            lookup_bundle: Dict[str, Any],
            variant_matrix: Dict[str, Any],
        ) -> Dict[str, Any]:
            target_location_map = build_target_location_map(selected_idea, analysis_data, code_plan)
            selected_source_record = select_source_record(selected_idea, lookup_bundle)
            module_candidates = build_module_candidates(selected_idea, lookup_bundle, analysis_data, code_plan)
            interface_diff = build_interface_diff(selected_idea, analysis_data, module_candidates)
            minimal_patch_plan = build_minimal_patch_plan(selected_idea, target_location_map, interface_diff, selected_source_record)
            smoke_plan = build_smoke_plan(selected_idea, target_location_map, variant_matrix)
            patch_class = resolve_patch_class(selected_idea, minimal_patch_plan)
            module_candidates_path = write_module_candidates(analysis_output_dir, module_candidates)
            interface_diff_path = write_interface_diff(analysis_output_dir, interface_diff)
            return {
                "schema_version": "1.0",
                "artifact_paths": [str(module_candidates_path), str(interface_diff_path)],
                "selected_source_record": selected_source_record or {},
                "transplant_ready": not source_blockers(selected_source_record),
                "source_blockers": source_blockers(selected_source_record),
                "target_location_map": target_location_map,
                "supporting_changes": code_plan.get("supporting_changes", []),
                "patch_surface_summary": code_plan.get("patch_surface_summary", {}),
                "module_candidates": module_candidates,
                "interface_diff": interface_diff,
                "minimal_patch_plan": minimal_patch_plan,
                "smoke_plan": smoke_plan,
                "requested_patch_class": patch_class["requested_patch_class"],
                "resolved_patch_class": patch_class["resolved_patch_class"],
                "patch_class_source": patch_class["patch_class_source"],
                "requires_source_triple": patch_class["requires_source_triple"],
            }
        
        
      • __init__.py 909 B
        """Internal passes for ai-research-explore orchestration."""
        
        from .atomic_idea_decomposition import run_atomic_idea_decomposition_pass
        from .candidate_idea_generation import run_candidate_idea_generation_pass
        from .execution_feasibility import run_execution_feasibility_pass
        from .idea_cards import run_idea_card_pass
        from .idea_ranking import run_idea_ranking_pass
        from .implementation_fidelity import run_implementation_fidelity_pass
        from .improvement_bank import run_improvement_bank_pass
        from .lookup_sources import run_lookup_pass
        from .source_mapping import run_source_mapping_pass
        
        __all__ = [
            "run_atomic_idea_decomposition_pass",
            "run_candidate_idea_generation_pass",
            "run_execution_feasibility_pass",
            "run_idea_card_pass",
            "run_idea_ranking_pass",
            "run_implementation_fidelity_pass",
            "run_improvement_bank_pass",
            "run_lookup_pass",
            "run_source_mapping_pass",
        ]
        
        
    • orchestrate_explore.py 121.9 KB
      #!/usr/bin/env python3
      """Plan or execute explicit exploratory research work on top of current_research."""
      
      from __future__ import annotations
      
      import argparse
      import hashlib
      import json
      import re
      import subprocess
      import sys
      import tempfile
      import time
      from pathlib import Path
      from typing import Any, Dict, List, Optional, Sequence, Tuple
      
      from passes import (
          run_atomic_idea_decomposition_pass,
          run_candidate_idea_generation_pass,
          run_execution_feasibility_pass,
          run_idea_card_pass,
          run_idea_ranking_pass,
          run_implementation_fidelity_pass,
          run_improvement_bank_pass,
          run_lookup_pass,
          run_source_mapping_pass,
      )
      
      SHARED_SCRIPTS = Path(__file__).resolve().parents[3] / "shared" / "scripts"
      if not all((SHARED_SCRIPTS / name).is_file() for name in (
          "runtime_runner.py", "model_adapter.py", "command_utils.py", "resource_monitor.py"
      )):
          SHARED_SCRIPTS = (Path(__file__).resolve().parents[2] / "ai-research-reproduction"
                            / "_bundled" / "shared" / "scripts")
      if not (SHARED_SCRIPTS / "model_adapter.py").is_file():
          raise RuntimeError("Shared runtime missing: install all RigorPilot skills, including ai-research-reproduction.")
      if str(SHARED_SCRIPTS) not in sys.path:
          sys.path.insert(0, str(SHARED_SCRIPTS))
      
      from model_adapter import ModelAdapterError, load_model_profile, missing_capabilities
      
      
      DURABLE_ANCHOR_HASH_RE = re.compile(r"^[0-9a-fA-F]{7,40}$")
      EXTERNAL_REFERENCE_PREFIXES = ("run:", "checkpoint:", "branch:", "commit:", "model:", "state:")
      DEFAULT_BASELINE_GATE = {
          "maximize": {"borderline_gap": 1.0, "abandon_gap": 2.0},
          "minimize": {"borderline_relative_gap": 0.02, "abandon_relative_gap": 0.05},
      }
      DEFAULT_EXECUTION_POLICY = {
          "run_selected_variants": False,
          "max_executed_variants": 1,
          "variant_timeout": 60,
          "run_full_after_short_run": False,
      }
      DEFAULT_IDEA_GENERATION_POLICY = {
          "allow_synthesized_seed_ideas": True,
          "max_generated_ideas": 3,
          "require_diverse_targets": True,
      }
      
      
      def run_json(script: Path, args: List[str]) -> Dict[str, Any]:
          result = subprocess.run([sys.executable, str(script), *args], check=True, capture_output=True, text=True)
          return json.loads(result.stdout)
      
      
      def add_model_profile_args(args: List[str], profile_json: str, required_capabilities: List[str]) -> List[str]:
          rendered = list(args)
          if profile_json:
              rendered.extend(["--model-profile-json", profile_json])
          for capability in required_capabilities:
              rendered.extend(["--require-model-capability", capability])
          return rendered
      
      
      def run_text(command: List[str], cwd: Optional[Path] = None) -> str:
          result = subprocess.run(command, check=True, capture_output=True, text=True, cwd=str(cwd) if cwd else None)
          return result.stdout.strip()
      
      
      def write_bundle(script: Path, output_dir: Path, context: Dict[str, Any]) -> None:
          output_dir.mkdir(parents=True, exist_ok=True)
          with tempfile.NamedTemporaryFile("w", encoding="utf-8", suffix=".json", delete=False) as handle:
              context_path = Path(handle.name)
              handle.write(json.dumps(context, indent=2, ensure_ascii=False))
      
          try:
              subprocess.run(
                  [
                      sys.executable,
                      str(script),
                      "--context-json",
                      str(context_path),
                      "--output-dir",
                      str(output_dir),
                  ],
                  check=True,
              )
          finally:
              if context_path.exists():
                  context_path.unlink()
      
      
      def slugify(value: str) -> str:
          slug = re.sub(r"[^a-z0-9]+", "-", value.lower()).strip("-")
          return slug[:40] or "current-research"
      
      
      def choose_experiment_branch(current_research: str, explicit_branch: str) -> str:
          if explicit_branch:
              return explicit_branch
          return f"exp/ai-research-explore-{slugify(current_research)}"
      
      
      def maybe_git_root(repo_path: Path) -> Optional[Path]:
          try:
              return Path(run_text(["git", "rev-parse", "--show-toplevel"], cwd=repo_path)).resolve()
          except (subprocess.CalledProcessError, FileNotFoundError):
              return None
      
      
      def build_context_id(current_research: str, experiment_branch: str) -> str:
          digest = hashlib.sha1(f"{current_research}::{experiment_branch}".encode("utf-8")).hexdigest()[:12]
          return f"ai-research-explore-{digest}"
      
      
      def experiment_worktree_root(git_root: Path, experiment_branch: str) -> Path:
          base_dir = git_root.parent / f".{git_root.name}-explore-worktrees" / slugify(experiment_branch)
          return base_dir / git_root.name
      
      
      def validate_current_research(repo_path: Path, current_research: str) -> Dict[str, Any]:
          value = current_research.strip()
          if not value:
              raise ValueError("`current_research` is required.")
      
          literal_path = Path(value)
          if literal_path.is_absolute() and literal_path.exists():
              return {"kind": "path", "value": value, "resolved_path": str(literal_path.resolve())}
      
          repo_relative = (repo_path / value).resolve()
          if repo_relative.exists():
              return {"kind": "repo-path", "value": value, "resolved_path": str(repo_relative)}
      
          git_root = maybe_git_root(repo_path)
          if git_root:
              try:
                  resolved_ref = run_text(["git", "rev-parse", "--verify", f"{value}^{{commit}}"], cwd=git_root)
                  return {
                      "kind": "git-ref",
                      "value": value,
                      "resolved_ref": resolved_ref,
                      "git_root": str(git_root),
                  }
              except subprocess.CalledProcessError:
                  pass
      
          if "@" in value:
              left, _, right = value.partition("@")
              if left and right and (
                  DURABLE_ANCHOR_HASH_RE.fullmatch(right.strip())
                  or any(right.strip().startswith(prefix) for prefix in EXTERNAL_REFERENCE_PREFIXES)
              ):
                  return {"kind": "named-anchor", "value": value}
      
          raise ValueError(
              "`current_research` should map to a durable branch, commit, checkpoint, run record, or trained model state."
          )
      
      
      def validate_existing_worktree(worktree_root: Path, expected_branch: str) -> Dict[str, Any]:
          actual_root = Path(run_text(["git", "rev-parse", "--show-toplevel"], cwd=worktree_root)).resolve()
          actual_branch = run_text(["git", "symbolic-ref", "--quiet", "--short", "HEAD"], cwd=worktree_root)
          if actual_root != worktree_root.resolve():
              raise ValueError(f"Existing experiment workspace `{worktree_root}` is not a valid git worktree root.")
          if actual_branch != expected_branch:
              raise ValueError(
                  f"Existing experiment workspace `{worktree_root}` is on branch `{actual_branch}`, expected `{expected_branch}`."
              )
          return {
              "workspace_root": str(actual_root),
              "worktree_root": str(actual_root),
              "mode": "worktree",
          }
      
      
      def ensure_experiment_workspace(repo_path: Path, experiment_branch: str) -> Dict[str, Any]:
          git_root = maybe_git_root(repo_path)
          if git_root is None:
              raise ValueError("Explore orchestration requires a git repository so the isolated experiment branch can be created.")
      
          head_sha = run_text(["git", "rev-parse", "HEAD"], cwd=git_root)
          try:
              current_branch = run_text(["git", "symbolic-ref", "--quiet", "--short", "HEAD"], cwd=git_root)
          except subprocess.CalledProcessError:
              current_branch = "DETACHED"
      
          branch_ref = f"refs/heads/{experiment_branch}"
          created_branch = False
          try:
              branch_sha = run_text(["git", "rev-parse", "--verify", branch_ref], cwd=git_root)
              branch_exists = True
          except subprocess.CalledProcessError:
              branch_sha = head_sha
              branch_exists = False
      
          if current_branch == experiment_branch:
              isolated_workspace = experiment_branch.startswith(("exp/", "explore/"))
              return {
                  "mode": "branch",
                  "workspace_root": str(git_root),
                  "worktree_root": None,
                  "branch": experiment_branch,
                  "branch_ref": branch_ref,
                  "branch_sha": branch_sha,
                  "head_sha": head_sha,
                  "current_branch": current_branch,
                  "created_branch": created_branch,
                  "isolated_workspace": isolated_workspace,
              }
      
          worktree_root = experiment_worktree_root(git_root, experiment_branch)
          if worktree_root.exists():
              worktree_info = validate_existing_worktree(worktree_root, experiment_branch)
          else:
              worktree_root.parent.mkdir(parents=True, exist_ok=True)
              if branch_exists:
                  run_text(["git", "worktree", "add", str(worktree_root), experiment_branch], cwd=git_root)
              else:
                  run_text(["git", "worktree", "add", "-b", experiment_branch, str(worktree_root), head_sha], cwd=git_root)
                  created_branch = True
              branch_sha = run_text(["git", "rev-parse", "--verify", branch_ref], cwd=git_root)
              worktree_info = validate_existing_worktree(worktree_root, experiment_branch)
      
          return {
              "mode": worktree_info["mode"],
              "workspace_root": worktree_info["workspace_root"],
              "worktree_root": worktree_info["worktree_root"],
              "branch": experiment_branch,
              "branch_ref": branch_ref,
              "branch_sha": branch_sha,
              "head_sha": head_sha,
              "current_branch": current_branch,
              "created_branch": created_branch,
              "isolated_workspace": True,
          }
      
      
      def normalize_task_family(value: Any) -> Optional[str]:
          text = str(value or "").strip().lower()
          return text or None
      
      
      def safe_float(value: Any) -> Optional[float]:
          if value is None:
              return None
          if isinstance(value, (int, float)):
              return float(value)
          try:
              return float(str(value))
          except ValueError:
              return None
      
      
      def clamp_score(value: Optional[float], default: float = 0.5) -> float:
          if value is None:
              return default
          return max(0.0, min(1.0, float(value)))
      
      
      def normalize_metric_goal(value: Any) -> str:
          text = str(value or "maximize").strip().lower()
          if text in {"min", "minimize", "lower", "lower_is_better"}:
              return "minimize"
          return "maximize"
      
      
      def load_structured_file(path: Path) -> Dict[str, Any]:
          text = path.read_text(encoding="utf-8-sig")
          if path.suffix.lower() in {".yaml", ".yml"}:
              try:
                  import yaml  # type: ignore
              except ImportError as exc:
                  raise RuntimeError("YAML input requires PyYAML to be installed.") from exc
              payload = yaml.safe_load(text) or {}
          else:
              payload = json.loads(text)
          if not isinstance(payload, dict):
              raise ValueError(f"Structured input `{path}` must contain a top-level object.")
          return payload
      
      
      def normalize_variant_spec(spec: Dict[str, Any], current_research: str) -> Dict[str, Any]:
          normalized = dict(spec)
          explicit_value = normalized.get("current_research") or normalized.get("baseline_ref")
          if explicit_value and explicit_value != current_research:
              raise ValueError(
                  f"Variant spec current research `{explicit_value}` does not match current_research `{current_research}`."
              )
          normalized["current_research"] = current_research
          normalized.setdefault("baseline_ref", current_research)
          # Explicit nulls, empty lists, and scalar axis values are all valid-looking
          # campaign inputs; coerce them so downstream passes never see a non-list.
          raw_axes = normalized.get("variant_axes") or {}
          normalized["variant_axes"] = {
              key: list(value) if isinstance(value, (list, tuple)) else [value]
              for key, value in raw_axes.items()
          }
          normalized["subset_sizes"] = list(normalized.get("subset_sizes") or [None])
          normalized["short_run_steps"] = list(normalized.get("short_run_steps") or [None])
          return normalized
      
      
      def load_variant_spec(path: Path, current_research: str) -> Dict[str, Any]:
          return normalize_variant_spec(load_structured_file(path), current_research)
      
      
      def normalize_evaluation_source(raw: Any, variant_spec: Dict[str, Any]) -> Dict[str, Any]:
          if isinstance(raw, str):
              source = {"command": raw}
          elif isinstance(raw, dict):
              source = dict(raw)
          else:
              source = {}
      
          primary_metric = source.get("primary_metric") or variant_spec.get("primary_metric")
          metric_goal = normalize_metric_goal(source.get("metric_goal") or variant_spec.get("metric_goal"))
          execution_kind = str(source.get("execution_kind") or "").strip().lower()
          return {
              "command": str(source.get("command") or ""),
              "path": str(source.get("path") or ""),
              "primary_metric": primary_metric,
              "metric_goal": metric_goal,
              "execution_kind": execution_kind or None,
              "artifacts": list(source.get("artifacts", []) or []),
              "notes": list(source.get("notes", []) or []),
              "split": str(source.get("split") or ""),
          }
      
      
      def bind_evaluation_command_to_variant_spec(
          variant_spec: Dict[str, Any],
          evaluation_source: Dict[str, Any],
      ) -> Dict[str, Any]:
          if variant_spec.get("base_command") or not evaluation_source.get("command"):
              return variant_spec
      
          normalized = dict(variant_spec)
          normalized["base_command"] = str(evaluation_source["command"]).strip()
          normalized["base_command_source"] = "evaluation_source"
          if evaluation_source.get("primary_metric") and not normalized.get("primary_metric"):
              normalized["primary_metric"] = evaluation_source["primary_metric"]
          if evaluation_source.get("metric_goal") and not normalized.get("metric_goal"):
              normalized["metric_goal"] = evaluation_source["metric_goal"]
          if evaluation_source.get("execution_kind") and not normalized.get("execution_kind"):
              normalized["execution_kind"] = evaluation_source["execution_kind"]
          return normalized
      
      
      def normalize_sota_reference(items: Any, primary_metric: Optional[str], metric_goal: str) -> List[Dict[str, Any]]:
          if not isinstance(items, list):
              return []
          normalized: List[Dict[str, Any]] = []
          for index, item in enumerate(items, start=1):
              if isinstance(item, (int, float)):
                  normalized.append(
                      {
                          "id": f"sota-{index:03d}",
                          "name": f"SOTA reference {index}",
                          "metric": primary_metric,
                          "metric_goal": metric_goal,
                          "value": float(item),
                          "source": "",
                          "notes": "",
                      }
                  )
                  continue
              if not isinstance(item, dict):
                  continue
              value = safe_float(item.get("value"))
              if value is None:
                  continue
              normalized.append(
                  {
                      "id": str(item.get("id") or f"sota-{index:03d}"),
                      "name": str(item.get("name") or item.get("paper") or f"SOTA reference {index}"),
                      "metric": str(item.get("metric") or primary_metric or ""),
                      "metric_goal": normalize_metric_goal(item.get("metric_goal") or metric_goal),
                      "value": value,
                      "source": str(item.get("source") or item.get("url") or ""),
                      "notes": str(item.get("notes") or ""),
                  }
              )
          return normalized
      
      
      def normalize_compute_budget(raw: Any) -> Dict[str, Any]:
          if not isinstance(raw, dict):
              return {}
          budget = dict(raw)
          if "max_runtime_hours" in budget:
              runtime = safe_float(budget.get("max_runtime_hours"))
              if runtime is not None:
                  budget["max_runtime_hours"] = runtime
          return budget
      
      
      def normalize_baseline_gate(raw: Any, metric_goal: str) -> Dict[str, Any]:
          gate = dict(raw) if isinstance(raw, dict) else {}
          defaults = DEFAULT_BASELINE_GATE[metric_goal]
          normalized = {
              "metric_goal": metric_goal,
              "borderline_gap": safe_float(gate.get("borderline_gap")),
              "abandon_gap": safe_float(gate.get("abandon_gap")),
              "borderline_relative_gap": safe_float(gate.get("borderline_relative_gap")),
              "abandon_relative_gap": safe_float(gate.get("abandon_relative_gap")),
              "timeout": int(gate.get("timeout") or 60),
              "max_steps": int(gate.get("max_steps") or 0),
          }
          if metric_goal == "maximize":
              normalized["borderline_gap"] = normalized["borderline_gap"] if normalized["borderline_gap"] is not None else defaults["borderline_gap"]
              normalized["abandon_gap"] = normalized["abandon_gap"] if normalized["abandon_gap"] is not None else defaults["abandon_gap"]
          else:
              normalized["borderline_relative_gap"] = normalized["borderline_relative_gap"] if normalized["borderline_relative_gap"] is not None else defaults["borderline_relative_gap"]
              normalized["abandon_relative_gap"] = normalized["abandon_relative_gap"] if normalized["abandon_relative_gap"] is not None else defaults["abandon_relative_gap"]
          return normalized
      
      
      def normalize_execution_policy(raw: Any, args: argparse.Namespace) -> Dict[str, Any]:
          policy = dict(DEFAULT_EXECUTION_POLICY)
          if isinstance(raw, dict):
              policy.update(raw)
          if args.run_selected_variants:
              policy["run_selected_variants"] = True
          if args.max_executed_variants is not None:
              policy["max_executed_variants"] = int(args.max_executed_variants)
          if args.variant_timeout is not None:
              policy["variant_timeout"] = int(args.variant_timeout)
          max_executed_variants = policy.get("max_executed_variants")
          variant_timeout = policy.get("variant_timeout")
          full_run_timeout = policy.get("full_run_timeout")
          return {
              "run_selected_variants": bool(policy.get("run_selected_variants", False)),
              "max_executed_variants": int(max_executed_variants) if max_executed_variants is not None else 1,
              "variant_timeout": int(variant_timeout) if variant_timeout is not None else 60,
              "run_full_after_short_run": bool(policy.get("run_full_after_short_run", False)),
              "full_run_timeout": (
                  int(full_run_timeout)
                  if full_run_timeout is not None
                  else int(variant_timeout)
                  if variant_timeout is not None
                  else 60
              ),
          }
      
      
      def stringify_campaign_binding(value: Any) -> str:
          if isinstance(value, dict):
              for key in ("name", "id", "path", "label"):
                  if value.get(key):
                      return str(value[key])
              items = [f"{key}={value[key]}" for key in sorted(value) if value.get(key) not in {None, ""}]
              return ", ".join(items) or "unspecified"
          if isinstance(value, list):
              return ", ".join(str(item) for item in value if str(item).strip()) or "unspecified"
          text = str(value or "").strip()
          return text or "unspecified"
      
      
      def evaluation_binding_text(evaluation_source: Dict[str, Any]) -> str:
          command = str(evaluation_source.get("command") or "").strip()
          path = str(evaluation_source.get("path") or "").strip()
          metric = str(evaluation_source.get("primary_metric") or "").strip()
          parts = []
          if path:
              parts.append(f"path={path}")
          if command:
              parts.append(f"command={command}")
          if metric:
              parts.append(f"metric={metric}")
          return " | ".join(parts) or "unspecified"
      
      
      def normalize_candidate_ideas(
          raw: Any,
          variant_spec: Dict[str, Any],
          *,
          current_research: str,
          task_family: str,
          dataset: Any,
          evaluation_source: Dict[str, Any],
      ) -> List[Dict[str, Any]]:
          if not isinstance(raw, list) or not raw:
              return []
      
          dataset_binding = stringify_campaign_binding(dataset)
          evaluation_binding = evaluation_binding_text(evaluation_source)
          task_binding = str(task_family or "").strip() or "unspecified"
          normalized: List[Dict[str, Any]] = []
          for index, item in enumerate(raw, start=1):
              if not isinstance(item, dict):
                  continue
              change_scope = str(item.get("change_scope") or "unspecified")
              target_component = str(item.get("target_component") or "unspecified")
              normalized.append(
                  {
                      "id": str(item.get("id") or f"idea-{index:03d}"),
                      "summary": str(item.get("summary") or item.get("description") or f"Candidate idea {index}"),
                      "change_scope": change_scope,
                      "target_component": target_component,
                      "expected_upside": clamp_score(safe_float(item.get("expected_upside")), default=0.5),
                      "implementation_risk": clamp_score(safe_float(item.get("implementation_risk")), default=0.5),
                      "eval_risk": clamp_score(safe_float(item.get("eval_risk")), default=0.5),
                      "rollback_ease": clamp_score(safe_float(item.get("rollback_ease")), default=0.5),
                      "estimated_runtime_cost": clamp_score(safe_float(item.get("estimated_runtime_cost")), default=0.5),
                      "single_variable_fit": clamp_score(safe_float(item.get("single_variable_fit")), default=0.8),
                      "hypothesis": str(item.get("hypothesis") or item.get("summary") or ""),
                      "supporting_changes": list(item.get("supporting_changes", []) or []),
                      "seed_origin": "researcher",
                      "campaign_idea_id": str(item.get("id") or f"idea-{index:03d}"),
                      "source_support_hint": str(item.get("source_support_hint") or ""),
                      "feasibility_hint": str(item.get("feasibility_hint") or ""),
                      "source": str(item.get("source") or ""),
                      "source_repo": str(item.get("source_repo") or ""),
                      "source_file": str(item.get("source_file") or ""),
                      "source_symbol": str(item.get("source_symbol") or ""),
                      "selection_origin": "campaign",
                      "context_anchor": str(item.get("context_anchor") or current_research),
                      "task_family_binding": str(item.get("task_family_binding") or task_binding),
                      "dataset_binding": str(item.get("dataset_binding") or dataset_binding),
                      "evaluation_binding": str(item.get("evaluation_binding") or evaluation_binding),
                      "constraint_notes": list(item.get("constraint_notes", []) or [
                          f"Anchor this candidate to current_research `{current_research}`.",
                          f"Keep the candidate inside task family `{task_binding}` and dataset `{dataset_binding}`.",
                          f"Preserve the frozen evaluation binding `{evaluation_binding}`.",
                          f"Keep `{change_scope}` around `{target_component}` single-variable and reversible.",
                      ]),
                  }
              )
          return normalized
      
      
      def normalize_idea_generation(raw: Any) -> Dict[str, Any]:
          policy = dict(DEFAULT_IDEA_GENERATION_POLICY)
          if isinstance(raw, dict):
              policy.update(raw)
          try:
              policy["max_generated_ideas"] = max(0, int(policy.get("max_generated_ideas", 3)))
          except (TypeError, ValueError):
              policy["max_generated_ideas"] = 3
          policy["allow_synthesized_seed_ideas"] = bool(policy.get("allow_synthesized_seed_ideas", True))
          policy["require_diverse_targets"] = bool(policy.get("require_diverse_targets", True))
          return policy
      
      
      def normalize_campaign(args: argparse.Namespace) -> Tuple[Dict[str, Any], bool]:
          if args.research_campaign_json:
              raw_campaign = load_structured_file(Path(args.research_campaign_json).resolve())
              compatibility_mode = False
          else:
              raw_campaign = {}
              compatibility_mode = True
      
          current_research = str(raw_campaign.get("current_research") or args.current_research or "").strip()
          if not current_research:
              raise ValueError("Either --current-research or --research-campaign-json with current_research is required.")
      
          if args.variant_spec_json:
              variant_spec = load_variant_spec(Path(args.variant_spec_json).resolve(), current_research)
          else:
              variant_spec = normalize_variant_spec(
                  raw_campaign.get("variant_spec", {}) if isinstance(raw_campaign.get("variant_spec"), dict) else {},
                  current_research,
              )
      
          evaluation_source = normalize_evaluation_source(raw_campaign.get("evaluation_source", {}), variant_spec)
          variant_spec = bind_evaluation_command_to_variant_spec(variant_spec, evaluation_source)
          metric_goal = normalize_metric_goal(evaluation_source.get("metric_goal") or variant_spec.get("metric_goal"))
          candidate_ideas = normalize_candidate_ideas(
              raw_campaign.get("candidate_ideas", []),
              variant_spec,
              current_research=current_research,
              task_family=str(raw_campaign.get("task_family") or ""),
              dataset=raw_campaign.get("dataset"),
              evaluation_source=evaluation_source,
          )
          execution_policy = normalize_execution_policy(raw_campaign.get("execution_policy", {}), args)
          sota_reference = normalize_sota_reference(raw_campaign.get("sota_reference", []), evaluation_source.get("primary_metric"), metric_goal)
          idea_generation = normalize_idea_generation(raw_campaign.get("idea_generation", {}))
      
          campaign = {
              "schema_version": "1.0",
              "mode": "legacy" if compatibility_mode else "campaign",
              "current_research": current_research,
              "task_family": normalize_task_family(raw_campaign.get("task_family")),
              "dataset": raw_campaign.get("dataset"),
              "benchmark": raw_campaign.get("benchmark"),
              "evaluation_source": evaluation_source,
              "sota_reference": sota_reference,
              "candidate_ideas": candidate_ideas,
              "researcher_candidate_ideas": candidate_ideas,
              "compute_budget": normalize_compute_budget(raw_campaign.get("compute_budget", {})),
              "variant_spec": variant_spec,
              "baseline_gate": normalize_baseline_gate(raw_campaign.get("baseline_gate", {}), metric_goal),
              "execution_policy": execution_policy,
              "research_lookup": dict(raw_campaign.get("research_lookup", {})) if isinstance(raw_campaign.get("research_lookup"), dict) else {},
              "idea_policy": dict(raw_campaign.get("idea_policy", {})) if isinstance(raw_campaign.get("idea_policy"), dict) else {},
              "idea_generation": idea_generation,
              "source_constraints": dict(raw_campaign.get("source_constraints", {})) if isinstance(raw_campaign.get("source_constraints"), dict) else {},
              "feasibility_policy": dict(raw_campaign.get("feasibility_policy", {})) if isinstance(raw_campaign.get("feasibility_policy"), dict) else {},
          }
          return campaign, compatibility_mode
      
      
      def build_stage_trace_entry(stage: str, tool: str, summary: str, status: str = "completed") -> Dict[str, Any]:
          return {"stage": stage, "tool": tool, "status": status, "summary": summary}
      
      
      def normalize_flag_name(key: str) -> str:
          return "--" + re.sub(r"[^a-z0-9]+", "-", key.lower()).strip("-")
      
      
      def quote_cli_value(value: Any) -> str:
          text = str(value)
          if any(char.isspace() for char in text):
              return f"\"{text}\""
          return text
      
      
      def maybe_append_cli_arg(command: str, flag: Any, value: Any) -> str:
          if flag in {None, False, ""} or value is None:
              return command
          return f"{command} {flag} {quote_cli_value(value)}"
      
      
      def compose_variant_command(base_command: str, variant: Dict[str, Any], spec: Dict[str, Any]) -> str:
          command = base_command.strip()
          axis_flag_map = spec.get("axis_flag_map") or {}
          for key, value in sorted(variant.get("axes", {}).items()):
              flag = axis_flag_map.get(key) or normalize_flag_name(key)
              command = maybe_append_cli_arg(command, flag, value)
      
          command = maybe_append_cli_arg(command, spec.get("subset_size_flag", "--subset-size"), variant.get("subset_size"))
          command = maybe_append_cli_arg(command, spec.get("short_run_steps_flag", "--max-steps"), variant.get("short_run_steps"))
          return command
      
      
      def summarize_variant_result(result: Dict[str, Any]) -> str:
          metric = result.get("best_metric")
          if metric:
              return f"status={result.get('status', 'unknown')}, stop={result.get('stop_reason', 'unknown')}, metric={metric['name']}={metric['value']}"
          return f"status={result.get('status', 'unknown')}, stop={result.get('stop_reason', 'unknown')}"
      
      
      def infer_execution_kind(base_command: Optional[str], spec_or_source: Dict[str, Any]) -> str:
          explicit = str(spec_or_source.get("execution_kind") or "").strip().lower()
          if explicit in {"train", "training"}:
              return "training"
          if explicit in {"run", "verify", "eval", "inference", "non_training", "non-training"}:
              return "non_training"
      
          lowered = str(base_command or "").lower()
          if any(token in lowered for token in [" train", "trainer", "fit", "fine-tune", "finetune"]):
              return "training"
          return "non_training"
      
      
      def extract_metric_policy(variant_matrix: Dict[str, Any], variant_spec: Dict[str, Any], campaign: Dict[str, Any]) -> Dict[str, Any]:
          matrix_policy = dict(variant_matrix.get("metric_policy", {}))
          evaluation_source = campaign.get("evaluation_source", {})
          primary_metric = matrix_policy.get("primary_metric") or evaluation_source.get("primary_metric") or variant_spec.get("primary_metric")
          metric_goal = normalize_metric_goal(
              matrix_policy.get("metric_goal") or evaluation_source.get("metric_goal") or variant_spec.get("metric_goal")
          )
          return {"primary_metric": primary_metric, "metric_goal": metric_goal}
      
      
      def extract_comparison_metric_policy(campaign: Dict[str, Any], metric_policy: Dict[str, Any]) -> Dict[str, Any]:
          evaluation_source = campaign.get("evaluation_source", {})
          return {
              "primary_metric": evaluation_source.get("primary_metric") or metric_policy.get("primary_metric"),
              "metric_goal": normalize_metric_goal(evaluation_source.get("metric_goal") or metric_policy.get("metric_goal")),
          }
      
      
      def default_metric_payload(item: Dict[str, Any]) -> Tuple[Optional[float], Optional[str]]:
          metric = item.get("best_metric")
          if isinstance(metric, dict):
              return safe_float(metric.get("value")), metric.get("name")
          return None, None
      
      
      def metric_payload_for_policy(item: Dict[str, Any], primary_metric: Optional[str]) -> Tuple[Optional[float], Optional[str], bool]:
          observed_metrics = item.get("observed_metrics", {})
          if primary_metric and isinstance(observed_metrics, dict) and primary_metric in observed_metrics:
              return safe_float(observed_metrics[primary_metric]), primary_metric, True
      
          best_metric = item.get("best_metric")
          if primary_metric and isinstance(best_metric, dict) and best_metric.get("name") == primary_metric:
              return safe_float(best_metric.get("value")), primary_metric, True
      
          fallback_value, fallback_name = default_metric_payload(item)
          return fallback_value, fallback_name, False
      
      
      def decorate_run_with_metric_policy(item: Dict[str, Any], metric_policy: Dict[str, Any]) -> Dict[str, Any]:
          primary_metric = metric_policy.get("primary_metric")
          metric_goal = normalize_metric_goal(metric_policy.get("metric_goal"))
          ranking_value, ranking_name, matched_primary_metric = metric_payload_for_policy(item, primary_metric)
      
          decorated = dict(item)
          decorated["ranking_metric"] = {
              "name": ranking_name,
              "value": ranking_value,
              "goal": metric_goal,
          } if ranking_name and ranking_value is not None else None
          decorated["ranking_metric_name"] = ranking_name
          decorated["ranking_metric_goal"] = metric_goal
          decorated["matched_primary_metric"] = matched_primary_metric if primary_metric else ranking_value is not None
          decorated["metric_policy_applied"] = bool(primary_metric)
          return decorated
      
      
      def rank_executed_runs(executed_runs: List[Dict[str, Any]], metric_policy: Dict[str, Any]) -> List[Dict[str, Any]]:
          status_rank = {"success": 3, "partial": 2, "blocked": 1, "not_run": 0}
          metric_goal = normalize_metric_goal(metric_policy.get("metric_goal"))
      
          def adjust_for_goal(value: Optional[float]) -> float:
              numeric_value = safe_float(value)
              if numeric_value is None:
                  return float("-inf")
              return numeric_value if metric_goal == "maximize" else -numeric_value
      
          decorated = [decorate_run_with_metric_policy(item, metric_policy) for item in executed_runs]
      
          def sort_key(item: Dict[str, Any]) -> Tuple[int, int, float, float]:
              ranking_metric = item.get("ranking_metric")
              ranking_value = ranking_metric.get("value") if isinstance(ranking_metric, dict) else None
              fallback_value, _fallback_name = default_metric_payload(item)
              return (
                  status_rank.get(item.get("status", "not_run"), 0),
                  1 if item.get("matched_primary_metric") else 0,
                  adjust_for_goal(ranking_value),
                  adjust_for_goal(fallback_value),
              )
      
          return sorted(decorated, key=sort_key, reverse=True)
      
      
      def build_variant_matrix(planner_script: Path, variant_spec: Dict[str, Any]) -> Dict[str, Any]:
          if not variant_spec.get("base_command"):
              current_research = variant_spec["current_research"]
              return {
                  "schema_version": "1.0",
                  "current_research": current_research,
                  "baseline_ref": variant_spec.get("baseline_ref", current_research),
                  "base_command": None,
                  "raw_variant_count": 0,
                  "variant_count": 0,
                  "pruned_variant_count": 0,
                  "variant_budget": {
                      "max_variants": int(variant_spec.get("max_variants") or 0),
                      "max_short_cycle_runs": int(variant_spec.get("max_short_cycle_runs") or 0),
                  },
                  "selection_policy": {
                      "factors": ["cost", "success_rate", "expected_gain"],
                      "weights": variant_spec.get("selection_weights", {}),
                  },
                  "metric_policy": {
                      "primary_metric": variant_spec.get("primary_metric"),
                      "metric_goal": normalize_metric_goal(variant_spec.get("metric_goal")),
                  },
                  "variants": [],
              }
      
          with tempfile.NamedTemporaryFile("w", encoding="utf-8", suffix=".json", delete=False) as handle:
              temp_spec_path = Path(handle.name)
              handle.write(json.dumps(variant_spec, indent=2, ensure_ascii=False))
      
          try:
              matrix = run_json(planner_script, ["--spec-json", str(temp_spec_path), "--json"])
          finally:
              if temp_spec_path.exists():
                  temp_spec_path.unlink()
          return matrix
      
      
      def execute_variant_candidates(
          *,
          train_execute_script: Path,
          run_execute_script: Path,
          repo_path: Path,
          variant_matrix: Dict[str, Any],
          variant_spec: Dict[str, Any],
          current_research: str,
          timeout: int,
          max_executed_variants: int,
          runtime_root: Path,
          model_profile_json: str,
          required_model_capabilities: List[str],
          gpu_monitor_enabled: bool,
          campaign: Optional[Dict[str, Any]] = None,
      ) -> Tuple[List[Dict[str, Any]], List[Dict[str, Any]]]:
          base_command = variant_matrix.get("base_command")
          variants = variant_matrix.get("variants", [])
          if not base_command or not variants or max_executed_variants <= 0:
              return [], []
      
          execution_kind = infer_execution_kind(base_command, variant_spec)
          metric_policy = extract_metric_policy(variant_matrix, variant_spec, campaign or {"evaluation_source": {}})
          executed_runs: List[Dict[str, Any]] = []
          stage_trace: List[Dict[str, Any]] = []
          for variant in variants[:max_executed_variants]:
              command = compose_variant_command(base_command, variant, variant_spec)
              if execution_kind == "training":
                  run_mode = "short_run_verification" if variant.get("short_run_steps") is not None else "startup_verification"
                  run_args = [
                          "--repo",
                          str(repo_path),
                          "--command",
                          command,
                          "--timeout",
                          str(timeout),
                          "--lane",
                          "explore",
                          "--run-mode",
                          run_mode,
                          "--dataset",
                          "current_research",
                          "--checkpoint-source",
                          current_research,
                          "--max-steps",
                          str(variant.get("short_run_steps") or 0),
                          "--runtime-root",
                          str(runtime_root),
                      ]
                  if not gpu_monitor_enabled:
                      run_args.append("--no-gpu-monitor")
                  payload = run_json(train_execute_script, add_model_profile_args(run_args, model_profile_json, required_model_capabilities))
                  tool_name = "run-train/scripts/run_training.py"
              else:
                  run_mode = "candidate_verify"
                  run_args = [
                          "--repo",
                          str(repo_path),
                          "--command",
                          command,
                          "--timeout",
                          str(timeout),
                          "--runtime-root",
                          str(runtime_root),
                      ]
                  payload = run_json(run_execute_script, add_model_profile_args(run_args, model_profile_json, required_model_capabilities))
                  payload.setdefault("stop_reason", "command_completed" if payload.get("status") == "success" else "command_checked")
                  tool_name = "minimal-run-and-audit/scripts/run_command.py"
              summary = summarize_variant_result(payload)
              executed_runs.append(
                  {
                      "id": variant.get("id", "unknown"),
                      "metric": payload.get("best_metric", {}).get("value") if payload.get("best_metric") else payload.get("status", "unknown"),
                      "metric_name": payload.get("best_metric", {}).get("name") if payload.get("best_metric") else None,
                      "summary": summary,
                      "status": payload.get("status", "unknown"),
                      "stop_reason": payload.get("stop_reason", "unknown"),
                      "runtime": {
                          "run_id": payload.get("runtime_run_id"),
                          "status": payload.get("runtime_status"),
                          "run_dir": payload.get("runtime_dir"),
                          "state_path": payload.get("runtime_state_path"),
                          "events_path": payload.get("runtime_events_path"),
                          "stdout_log_path": payload.get("stdout_log_path"),
                          "stderr_log_path": payload.get("stderr_log_path"),
                          "resources_log_path": payload.get("resources_log_path"),
                          "resource_summary": payload.get("resource_summary", {}),
                          "model_adapter": payload.get("model_adapter"),
                      },
                      "command": command,
                      "axes": variant.get("axes", {}),
                      "subset_size": variant.get("subset_size"),
                      "short_run_steps": variant.get("short_run_steps"),
                      "best_metric": payload.get("best_metric"),
                      "observed_metrics": payload.get("observed_metrics", {}),
                      "best_checkpoint": payload.get("best_checkpoint"),
                      "changed_files": payload.get("changed_files", []),
                      "new_files": payload.get("new_files", []),
                      "deleted_files": payload.get("deleted_files", []),
                      "touched_paths": payload.get("touched_paths", []),
                      "touched_symbols": payload.get("touched_symbols", []),
                      "evidence_capture": payload.get("evidence_capture", {}),
                  }
              )
              stage_trace.append(
                  build_stage_trace_entry(
                      "variant-execution",
                      tool_name,
                      f"Executed `{variant.get('id', 'unknown')}` with mode `{run_mode}` and observed {summary}.",
                  )
              )
      
          return rank_executed_runs(executed_runs, metric_policy), stage_trace
      
      
      def build_analysis_context(campaign: Dict[str, Any], metric_policy: Dict[str, Any], current_research: str) -> Dict[str, Any]:
          evaluation_source = dict(campaign.get("evaluation_source", {}))
          if metric_policy.get("primary_metric") and not evaluation_source.get("primary_metric"):
              evaluation_source["primary_metric"] = metric_policy["primary_metric"]
          if metric_policy.get("metric_goal") and not evaluation_source.get("metric_goal"):
              evaluation_source["metric_goal"] = metric_policy["metric_goal"]
          return {
              "current_research": current_research,
              "task_family": campaign.get("task_family"),
              "dataset": campaign.get("dataset"),
              "benchmark": campaign.get("benchmark"),
              "evaluation_source": evaluation_source,
          }
      
      
      def run_analysis_pass(
          analysis_script: Path,
          workspace_repo_path: Path,
          analysis_output_dir: Path,
          analysis_context: Dict[str, Any],
      ) -> Dict[str, Any]:
          with tempfile.NamedTemporaryFile("w", encoding="utf-8", suffix=".json", delete=False) as handle:
              context_path = Path(handle.name)
              handle.write(json.dumps(analysis_context, indent=2, ensure_ascii=False))
      
          try:
              return run_json(
                  analysis_script,
                  [
                      "--repo",
                      str(workspace_repo_path),
                      "--output-dir",
                      str(analysis_output_dir),
                      "--analysis-context-json",
                      str(context_path),
                  ],
              )
          finally:
              if context_path.exists():
                  context_path.unlink()
      
      
      def run_code_plan_pass(
          *,
          code_planner_script: Path,
          workspace_repo_path: Path,
          current_research: str,
          experiment_branch: str,
          task_family: str,
          variant_spec: Dict[str, Any],
          selected_idea: Optional[Dict[str, Any]] = None,
          analysis_data: Optional[Dict[str, Any]] = None,
      ) -> Dict[str, Any]:
          temp_paths: List[Path] = []
          args = [
              "--repo",
              str(workspace_repo_path),
              "--current-research",
              current_research,
              "--experiment-branch",
              experiment_branch,
              "--task-family",
              task_family,
              "--json",
          ]
      
          with tempfile.NamedTemporaryFile("w", encoding="utf-8", suffix=".json", delete=False) as handle:
              variant_spec_path = Path(handle.name)
              handle.write(json.dumps(variant_spec, indent=2, ensure_ascii=False))
          temp_paths.append(variant_spec_path)
          args.extend(["--variant-spec-json", str(variant_spec_path)])
      
          if selected_idea:
              with tempfile.NamedTemporaryFile("w", encoding="utf-8", suffix=".json", delete=False) as handle:
                  idea_card_path = Path(handle.name)
                  handle.write(json.dumps(selected_idea, indent=2, ensure_ascii=False))
              temp_paths.append(idea_card_path)
              args.extend(["--idea-card-json", str(idea_card_path)])
      
          if analysis_data:
              with tempfile.NamedTemporaryFile("w", encoding="utf-8", suffix=".json", delete=False) as handle:
                  analysis_path = Path(handle.name)
                  handle.write(json.dumps(analysis_data, indent=2, ensure_ascii=False))
              temp_paths.append(analysis_path)
              args.extend(["--analysis-json", str(analysis_path)])
      
          try:
              return run_json(code_planner_script, args)
          finally:
              for path in temp_paths:
                  if path.exists():
                      path.unlink()
      
      
      def best_sota_reference(sota_reference: Sequence[Dict[str, Any]], metric_policy: Dict[str, Any]) -> Optional[Dict[str, Any]]:
          primary_metric = metric_policy.get("primary_metric")
          metric_goal = normalize_metric_goal(metric_policy.get("metric_goal"))
          candidates = [
              item
              for item in sota_reference
              if safe_float(item.get("value")) is not None and (not primary_metric or item.get("metric") in {primary_metric, "", None})
          ]
          if not candidates:
              return None
          reverse = metric_goal == "maximize"
          return sorted(candidates, key=lambda item: safe_float(item.get("value")) or 0.0, reverse=reverse)[0]
      
      
      def run_baseline_evaluation(
          *,
          train_execute_script: Path,
          run_execute_script: Path,
          repo_path: Path,
          current_research: str,
          evaluation_source: Dict[str, Any],
          baseline_gate_cfg: Dict[str, Any],
          runtime_root: Path,
          model_profile_json: str,
          required_model_capabilities: List[str],
          gpu_monitor_enabled: bool,
      ) -> Tuple[Dict[str, Any], Dict[str, Any], float]:
          command = str(evaluation_source.get("command") or "").strip()
          if not command:
              return (
                  {
                      "decision": "not-applicable",
                      "reason": "No evaluation command was provided in evaluation_source.",
                      "metric_name": evaluation_source.get("primary_metric"),
                      "metric_value": None,
                      "runtime_seconds": 0.0,
                  },
                  {},
                  0.0,
              )
      
          execution_kind = infer_execution_kind(command, evaluation_source)
          start = time.perf_counter()
          if execution_kind == "training":
              max_steps = int(baseline_gate_cfg.get("max_steps") or 0)
              run_mode = "short_run_verification" if max_steps > 0 else "startup_verification"
              run_args = [
                      "--repo",
                      str(repo_path),
                      "--command",
                      command,
                      "--timeout",
                      str(int(baseline_gate_cfg.get("timeout") or 60)),
                      "--lane",
                      "explore",
                      "--run-mode",
                      run_mode,
                      "--dataset",
                      str(evaluation_source.get("split") or "baseline"),
                      "--checkpoint-source",
                      current_research,
                      "--max-steps",
                      str(max_steps),
                      "--runtime-root",
                      str(runtime_root),
                  ]
              if not gpu_monitor_enabled:
                  run_args.append("--no-gpu-monitor")
              payload = run_json(train_execute_script, add_model_profile_args(run_args, model_profile_json, required_model_capabilities))
          else:
              run_args = [
                      "--repo",
                      str(repo_path),
                      "--command",
                      command,
                      "--timeout",
                      str(int(baseline_gate_cfg.get("timeout") or 60)),
                      "--runtime-root",
                      str(runtime_root),
                  ]
              payload = run_json(run_execute_script, add_model_profile_args(run_args, model_profile_json, required_model_capabilities))
              payload.setdefault("stop_reason", "command_completed" if payload.get("status") == "success" else "command_checked")
          runtime_seconds = round(time.perf_counter() - start, 3)
      
          primary_metric = evaluation_source.get("primary_metric")
          metric_value, metric_name, matched_primary = metric_payload_for_policy(payload, primary_metric)
          baseline_metric_name = metric_name or primary_metric
          baseline_gate = {
              "decision": "not-applicable",
              "reason": "Evaluation ran, but no comparable SOTA reference was available.",
              "metric_name": baseline_metric_name,
              "metric_value": metric_value,
              "matched_primary_metric": matched_primary,
              "status": payload.get("status", "unknown"),
              "stop_reason": payload.get("stop_reason", "unknown"),
              "runtime_seconds": runtime_seconds,
              "execution_kind": execution_kind,
          }
          return baseline_gate, payload, runtime_seconds
      
      
      def compare_baseline_to_sota(
          baseline_gate: Dict[str, Any],
          baseline_payload: Dict[str, Any],
          metric_policy: Dict[str, Any],
          sota_reference: Sequence[Dict[str, Any]],
          baseline_gate_cfg: Dict[str, Any],
      ) -> Dict[str, Any]:
          baseline_value = safe_float(baseline_gate.get("metric_value"))
          metric_name = baseline_gate.get("metric_name") or metric_policy.get("primary_metric")
          if baseline_value is None or not metric_name:
              baseline_gate["decision"] = "not-applicable"
              baseline_gate["reason"] = "Baseline evaluation did not produce the primary metric."
              return baseline_gate
      
          reference = best_sota_reference(sota_reference, metric_policy)
          if not reference:
              baseline_gate["decision"] = "not-applicable"
              baseline_gate["reason"] = "No comparable SOTA reference was provided."
              return baseline_gate
      
          metric_goal = normalize_metric_goal(metric_policy.get("metric_goal"))
          sota_value = float(reference["value"])
          baseline_gate["reference"] = reference
          if metric_goal == "maximize":
              gap = round(sota_value - baseline_value, 4)
              baseline_gate["gap_to_sota"] = gap
              if gap > float(baseline_gate_cfg["abandon_gap"]):
                  baseline_gate["decision"] = "abandon"
                  baseline_gate["reason"] = f"Baseline `{metric_name}={baseline_value}` trails provided SOTA `{sota_value}` by `{gap}` absolute points."
              elif gap > float(baseline_gate_cfg["borderline_gap"]):
                  baseline_gate["decision"] = "borderline"
                  baseline_gate["reason"] = f"Baseline `{metric_name}={baseline_value}` is within a plausible improvement range but still `{gap}` points off the provided SOTA."
              else:
                  baseline_gate["decision"] = "proceed"
                  baseline_gate["reason"] = f"Baseline `{metric_name}={baseline_value}` is close enough to the provided SOTA `{sota_value}` to justify follow-up work."
          else:
              relative_gap = 0.0 if sota_value == 0 else round(max(0.0, (baseline_value - sota_value) / abs(sota_value)), 4)
              baseline_gate["relative_gap_to_sota"] = relative_gap
              if relative_gap > float(baseline_gate_cfg["abandon_relative_gap"]):
                  baseline_gate["decision"] = "abandon"
                  baseline_gate["reason"] = f"Baseline `{metric_name}={baseline_value}` is worse than the provided SOTA `{sota_value}` by `{relative_gap:.2%}`."
              elif relative_gap > float(baseline_gate_cfg["borderline_relative_gap"]):
                  baseline_gate["decision"] = "borderline"
                  baseline_gate["reason"] = f"Baseline `{metric_name}={baseline_value}` is close enough to the provided SOTA `{sota_value}` to review manually before scaling."
              else:
                  baseline_gate["decision"] = "proceed"
                  baseline_gate["reason"] = f"Baseline `{metric_name}={baseline_value}` is close enough to the provided SOTA `{sota_value}` to justify follow-up work."
          baseline_gate["observed_metrics"] = baseline_payload.get("observed_metrics", {})
          baseline_gate["best_metric"] = baseline_payload.get("best_metric")
          baseline_gate["best_checkpoint"] = baseline_payload.get("best_checkpoint")
          return baseline_gate
      
      
      def score_candidate_idea(idea: Dict[str, Any]) -> float:
          score = (
              0.40 * clamp_score(safe_float(idea.get("expected_upside")), default=0.5)
              + 0.20 * clamp_score(safe_float(idea.get("single_variable_fit")), default=0.8)
              + 0.15 * clamp_score(safe_float(idea.get("rollback_ease")), default=0.5)
              - 0.10 * clamp_score(safe_float(idea.get("implementation_risk")), default=0.5)
              - 0.10 * clamp_score(safe_float(idea.get("eval_risk")), default=0.5)
              - 0.05 * clamp_score(safe_float(idea.get("estimated_runtime_cost")), default=0.5)
          )
          return round(score, 4)
      
      
      def build_idea_gate(candidate_ideas: Sequence[Dict[str, Any]]) -> Dict[str, Any]:
          ranked = [dict(item, idea_score=score_candidate_idea(item)) for item in candidate_ideas]
          ranked.sort(
              key=lambda item: (
                  -item["idea_score"],
                  -item.get("expected_upside", 0.0),
                  item.get("implementation_risk", 1.0),
                  item.get("estimated_runtime_cost", 1.0),
                  item.get("id", ""),
              )
          )
          top_diff = None
          if len(ranked) >= 2:
              top_diff = round(ranked[0]["idea_score"] - ranked[1]["idea_score"], 4)
          return {
              "decision": "selected" if ranked else "not-configured",
              "ranked_ideas": ranked,
              "selected_idea": ranked[0] if ranked else None,
              "top_idea_score_diff": top_diff,
          }
      
      
      def human_checkpoint_state(
          *,
          compatibility_mode: bool,
          eval_contract_complete: bool,
          baseline_gate: Dict[str, Any],
          idea_gate: Dict[str, Any],
      ) -> Tuple[str, List[str]]:
          if compatibility_mode:
              return "not-required", []
      
          reasons: List[str] = []
          if not eval_contract_complete:
              reasons.append("eval-contract-incomplete")
          if baseline_gate.get("decision") == "borderline":
              reasons.append("baseline-borderline")
          top_diff = safe_float(idea_gate.get("top_idea_score_diff"))
          if top_diff is not None and top_diff < 0.05:
              reasons.append("idea-selection-confirmation-required")
          if not reasons:
              return "not-required", []
          if len(reasons) == 1:
              return reasons[0], reasons
          return "multiple-reasons", reasons
      
      
      def build_config_diff_summary(selected_idea: Optional[Dict[str, Any]], variant_matrix: Dict[str, Any]) -> List[str]:
          lines: List[str] = []
          if selected_idea:
              lines.append(f"Primary change scope: `{selected_idea.get('change_scope', 'unspecified')}`.")
          if variant_matrix.get("variants"):
              variant = variant_matrix["variants"][0]
              for key, value in sorted(variant.get("axes", {}).items()):
                  lines.append(f"Set `{key}` to `{value}` for the leading short-run candidate.")
              if variant.get("subset_size") is not None:
                  lines.append(f"Use subset size `{variant['subset_size']}` during the short-run gate.")
              if variant.get("short_run_steps") is not None:
                  lines.append(f"Cap short-run execution at `{variant['short_run_steps']}` steps.")
          if not lines:
              lines.append("No config overrides were derived from the current campaign.")
          return lines
      
      
      def feasibility_score(short_run_feasibility: str) -> float:
          if short_run_feasibility == "proceed":
              return 1.0
          if short_run_feasibility == "borderline":
              return 0.5
          return 0.0
      
      
      def enrich_cards_with_feasibility(
          cards: Sequence[Dict[str, Any]],
          feasibility_bundle: Dict[str, Any],
      ) -> List[Dict[str, Any]]:
          short_run_feasibility = str(feasibility_bundle.get("feasibility", {}).get("short_run_feasibility") or "plausible")
          score = feasibility_score(short_run_feasibility)
          enriched: List[Dict[str, Any]] = []
          for item in cards:
              card = dict(item)
              card["short_run_feasibility"] = short_run_feasibility
              card["execution_feasibility_score"] = score
              enriched.append(card)
          return enriched
      
      
      def merge_selected_idea_with_source_mapping(
          selected_idea: Optional[Dict[str, Any]],
          source_mapping: Dict[str, Any],
      ) -> Optional[Dict[str, Any]]:
          if selected_idea is None:
              return None
          merged = dict(selected_idea)
          merged["requested_patch_class"] = source_mapping.get("requested_patch_class") or str(merged.get("patch_class") or "")
          merged["patch_class"] = source_mapping.get("resolved_patch_class") or str(merged.get("patch_class") or "config-only")
          merged["patch_class_source"] = source_mapping.get("patch_class_source") or ("campaign" if merged.get("patch_class") else "default")
          merged["requires_source_triple"] = bool(source_mapping.get("requires_source_triple"))
          return merged
      
      
      def observed_changed_files_from_fidelity(implementation_fidelity: Dict[str, Any]) -> List[str]:
          observed: List[str] = []
          for unit in implementation_fidelity.get("fidelity_units", []) or []:
              for site in unit.get("observed_implementation_sites", []) or unit.get("actual_observed_implementation_site", []):
                  text = str(site or "").strip()
                  if not text:
                      continue
                  _label, _sep, path = text.partition(":")
                  candidate = path or text
                  candidate = candidate.strip()
                  if candidate and candidate not in observed:
                      observed.append(candidate)
          return observed
      
      
      def build_experiment_manifest(
          *,
          current_research: str,
          selected_idea: Optional[Dict[str, Any]],
          code_plan: Dict[str, Any],
          campaign: Dict[str, Any],
          metric_policy: Dict[str, Any],
          analysis_output_dir: Path,
          variant_matrix: Dict[str, Any],
          source_mapping: Optional[Dict[str, Any]] = None,
          feasibility_bundle: Optional[Dict[str, Any]] = None,
          atomic_bundle: Optional[Dict[str, Any]] = None,
          implementation_fidelity: Optional[Dict[str, Any]] = None,
      ) -> Dict[str, Any]:
          mapping = source_mapping or {}
          atomic = atomic_bundle or {}
          fidelity = implementation_fidelity or {}
          planned_changed_files = [item.get("file") for item in mapping.get("target_location_map", [])[:3] if item.get("file")]
          observed_changed_files = observed_changed_files_from_fidelity(fidelity)
          if selected_idea is None:
              return {
                  "status": "blocked",
                  "parent_baseline": current_research,
                  "idea_id": None,
                  "hypothesis": "",
                  "changed_files": [],
                  "planned_changed_files": [],
                  "observed_changed_files": [],
                  "config_overrides": {},
                  "dataset": campaign.get("dataset"),
                  "eval_contract_ref": str((analysis_output_dir / "EVAL_CONTRACT.md").as_posix()),
                  "improvement_bank_ref": str((analysis_output_dir / "IMPROVEMENT_BANK.md").as_posix()),
                  "idea_cards_ref": str((analysis_output_dir / "IDEA_CARDS.json").as_posix()),
                  "idea_scores_ref": str((analysis_output_dir / "IDEA_SCORES.json").as_posix()),
                  "idea_seeds_ref": str((analysis_output_dir / "IDEA_SEEDS.json").as_posix()),
                  "module_candidates_ref": str((analysis_output_dir / "MODULE_CANDIDATES.md").as_posix()),
                  "interface_diff_ref": str((analysis_output_dir / "INTERFACE_DIFF.md").as_posix()),
                  "resource_plan_ref": str((analysis_output_dir / "RESOURCE_PLAN.md").as_posix()),
                  "atomic_idea_map_ref": str((analysis_output_dir / "ATOMIC_IDEA_MAP.json").as_posix()),
                  "implementation_fidelity_ref": str((analysis_output_dir / "IMPLEMENTATION_FIDELITY.json").as_posix()),
                  "primary_metric": metric_policy.get("primary_metric"),
                  "seed_policy": "inherit-baseline-seeds",
                  "budget": campaign.get("compute_budget", {}),
                  "promotion_rule": "No promotion; experiment manifest is blocked until one idea passes the idea gate.",
                  "supporting_changes": mapping.get("supporting_changes", []),
                  "selected_source_reference": [],
                  "selected_source_record": mapping.get("selected_source_record", {}),
                  "target_location_map": mapping.get("target_location_map", []),
                  "minimal_patch_plan": mapping.get("minimal_patch_plan", []),
                  "smoke_validation_plan": mapping.get("smoke_plan", []),
                  "feasibility_summary": (feasibility_bundle or {}).get("feasibility", {}),
                  "atomic_idea_summary": {
                      "status": atomic.get("status", "blocked"),
                      "atomic_unit_count": atomic.get("atomic_unit_count", 0),
                  },
                  "implementation_fidelity_summary": fidelity.get("fidelity_summary", {}),
                  "blockers": ["no-selected-idea"],
              }
          idea = selected_idea
          manifest_blockers = list(mapping.get("source_blockers", [])) if mapping.get("requires_source_triple") else []
          manifest_blockers.extend(list(atomic.get("blockers", [])))
          manifest_blockers = [item for item in manifest_blockers if item]
          return {
              "status": "blocked" if manifest_blockers else "ready",
              "parent_baseline": current_research,
              "idea_id": idea.get("id"),
              "hypothesis": idea.get("hypothesis") or idea.get("summary"),
              "changed_files": observed_changed_files,
              "planned_changed_files": planned_changed_files,
              "observed_changed_files": observed_changed_files,
              "config_overrides": variant_matrix.get("variants", [{}])[0].get("axes", {}) if variant_matrix.get("variants") else {},
              "dataset": campaign.get("dataset"),
              "eval_contract_ref": str((analysis_output_dir / "EVAL_CONTRACT.
    • write_outputs.py 1.2 KB
      #!/usr/bin/env python3
      """Compatibility wrapper for research exploration output bundles."""
      
      from __future__ import annotations
      
      import importlib.util
      from pathlib import Path
      
      
      def load_shared_module():
          module_path = Path(__file__).resolve().parents[3] / "shared" / "scripts" / "write_explore_bundle.py"
          if not module_path.is_file():
              module_path = (Path(__file__).resolve().parents[2] / "ai-research-reproduction"
                             / "_bundled" / "shared" / "scripts" / "write_explore_bundle.py")
          if not module_path.is_file():
              raise RuntimeError("Shared writer missing: install all RigorPilot skills, including ai-research-reproduction.")
          spec = importlib.util.spec_from_file_location("write_explore_bundle", module_path)
          if spec is None or spec.loader is None:
              raise RuntimeError(f"Unable to load shared writer module from {module_path}")
          module = importlib.util.module_from_spec(spec)
          spec.loader.exec_module(module)
          return module
      
      
      def main() -> int:
          module = load_shared_module()
          return module.main(default_mode="research", default_output_dir="explore_outputs")
      
      
      if __name__ == "__main__":
          raise SystemExit(main())
      
  • SKILL.md 6.5 KB
    ---
    name: ai-research-explore
    description: Rigor Explore compatible skill slug for meaningful and potentially novel deep learning research candidates. Use when the researcher has chosen the task family, dataset, benchmark, evaluation method, provided SOTA references, and wants candidate-only exploration on top of `current_research` with auditable repo understanding, idea gating, fair comparison, and governed experiments written to `explore_outputs/`. Do not use for README-first trusted reproduction, open-ended direction finding, narrow code-only or run-only exploration, passive repo analysis, verified novelty claims, or implicit experimentation.
    ---
    
    # ai-research-explore
    
    ## Purpose
    
    Use this as the Rigor Explore compatible skill slug after the researcher
    explicitly authorizes candidate-only work on top of a durable
    `current_research` anchor. The installed slug remains `ai-research-explore` for
    compatibility. Rigor Explore is for meaningful and potentially novel deep
    learning research candidates while preserving scientific rigor, comparability,
    reproducibility, and auditable collaboration. Novelty and significance remain
    hypotheses before literature contrast, ablation evidence, and fair comparison.
    The skill does not promise autonomous discovery, global benchmark completeness,
    novelty proof, or trusted reproduction success.
    
    Start from the shared operating principles in
    `../ai-research-reproduction/references/agent-operating-principles.md`, then load
    `../ai-research-reproduction/references/research-rigor-principles.md` for research claims and
    `../ai-research-reproduction/references/deep-learning-experiment-principles.md` when experiment
    details affect comparability or reproducibility.
    
    ## Fit
    
    Use this skill only when the request has both:
    
    - Explicit exploration authorization such as candidate-only work, isolated
      branch or worktree, sweep, several variants, or exploratory ranking.
    - A durable `current_research` context such as a branch, commit, checkpoint,
      run record, or already-trained local model state.
    
    Keep narrow code-only requests on `explore-code`. Keep narrow run-only requests
    on `explore-run`. Keep passive repository analysis on `analyze-project`. Keep
    README-first reproduction on `ai-research-reproduction`.
    
    ## Research Rhythm
    
    Use a two-loop rhythm:
    
    - Outer loop: understand the repository, freeze task/dataset/evaluation/budget,
      preserve user ideas, map sources, gate ideas, and decide whether the next
      experiment is worth running.
    - Inner loop: make one bounded candidate change or run, smoke-check it, collect
      evidence, rank it against the current anchor, and either stop or return to the
      outer loop with the new evidence.
    
    This rhythm is a guide, not a rigid autonomous loop. Stop at explicit blockers,
    unclear scientific meaning, exhausted budget, missing anchor/evaluation, or a
    human checkpoint.
    
    ## Workflow
    
    1. Confirm `current_research` and explicit explore-lane authorization.
    2. Accept either legacy `variant_spec` or higher-level `research_campaign`.
    3. In campaign mode, freeze the task, dataset, benchmark, evaluation source,
       SOTA reference, and budget before candidate work.
    4. Build only the repo-understanding artifacts needed for the current campaign,
       usually through `analyze-project`.
    5. Run bounded, cache-first source lookup when source support matters; prefer
       local curated literature such as Zotero if available, then seed sources,
       repo-local locators, public locators, or optional web lookup. Treat lookup as
       source resolution, not an open-ended literature search.
    6. Preserve researcher-provided ideas, optionally add a small bounded set of
       single-variable seed ideas, and rank ideas with explicit gates and score
       breakdowns.
    7. Prefer one clear candidate at a time. Use `explore-code` for bounded code
       adaptation and `explore-run` for short-cycle trials or sweeps.
    8. Use `minimal-run-and-audit` or `run-train` only when the exploratory plan
       requires real execution evidence.
    9. Write candidate-only outputs to `analysis_outputs/`, `sources/`, and
       `explore_outputs/` as appropriate; never present exploratory gains as trusted
       reproduction success. Include `SCIENTIFIC_CHANGELOG.md` and
       `COMPARABILITY_REPORT.md` for candidate scientific meaning and comparison
       boundaries.
    
    ## Ranking and Evidence
    
    - Before execution, prioritize candidates by expected gain, cost, success
      likelihood, patch surface, dependency drag, evaluation risk, and rollback
      ease.
    - After execution, rank by real evidence first: command status, observed
      metrics, artifacts, changed paths, smoke results, and reproducibility notes.
    - Keep researcher-provided `evaluation_source` and `sota_reference` frozen for
      the campaign; do not claim they are globally complete.
    - If the top ideas are too close or the implementation cannot be decomposed into
      auditable units, stop for a checkpoint instead of silently choosing.
    
    ## Campaign Inputs
    
    `research_campaign` is preferred for Rigor Explore campaigns, but it should
    stay minimal. The durable core is:
    
    - `current_research`
    - `task_family`
    - `dataset`
    - `benchmark`
    - `evaluation_source`
    - `sota_reference`
    - `compute_budget`
    
    Use `candidate_ideas`, `variant_spec`, `research_lookup`, `idea_policy`,
    `idea_generation`, `source_constraints`, `feasibility_policy`, `baseline_gate`,
    and `execution_policy` as optional guidance, not as fields the agent must fill
    for every campaign. See `references/research-campaign-spec.md` for the advanced
    schema and artifact expectations.
    
    ## Reference Loading
    
    - Load `references/ai-research-explore-policy.md` for lane safety and candidate
      semantics.
    - Load `references/research-campaign-spec.md` only when a campaign file is
      present or the user asks for Rigor Explore campaign governance.
    - Load `../ai-research-reproduction/references/explore-variant-spec.md` for run-level variant matrix
      details.
    - Load `../ai-research-reproduction/references/research-thinking-loop.md` before proposing or ranking candidate changes; it is the required greedy observe-ground-design-compare cycle.
    - Load `../ai-research-reproduction/references/research-rigor-principles.md` before making novelty, contribution, SOTA, or comparability statements.
    - Consult `~/.rigorpilot/PERSONAL_RIGOR.md` if present, under `../ai-research-reproduction/references/continuous-learning-policy.md` (advisory only; core wins).
    - Load `../ai-research-reproduction/references/deep-learning-experiment-principles.md` when training,
      evaluation, baseline, ablation, metric, checkpoint, or dataset details matter.
    - Use `scripts/orchestrate_explore.py` and `scripts/write_outputs.py` for the
      existing deterministic artifact workflow.
    
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related