minimal-run-and-audit
Rigor Run skill for README-first deep learning repo reproduction. Use when the task is specifically to capture or normalize evidence from the selected smoke test or documented inference or evaluation command and write standardized `repro_outputs/` files, including patch notes whe
Install
npx skills add https://github.com/lllllllama/RigorPilot-Skills/tree/main/skills/minimal-run-and-audit
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install lllllllama-rigorpilot-skills@llmmart
git clone https://github.com/lllllllama/RigorPilot-Skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole lllllllama/rigorpilot-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
minimal-run-and-audit
Use this as the Rigor Run skill. The installed slug remains
minimal-run-and-audit for compatibility.
Use the shared operating principles in
../ai-research-reproduction/references/agent-operating-principles.md; this skill should make run
evidence auditable without turning every command into a rigid protocol.
When to apply
- After a reproduction target and setup plan exist.
- When the main skill needs execution evidence and normalized outputs.
- When a smoke test, documented inference run, documented evaluation run, or other short non-training verification is appropriate.
- When the user already knows what command should be attempted and wants execution plus reporting only.
When not to apply
- During initial repo scanning.
- When environment or assets are still undefined enough to make execution meaningless.
- When the task is a literature lookup rather than repository execution.
- When the user is still deciding which reproduction target should count as the main run.
Clear boundaries
- This skill owns normalized reporting for an attempted command.
- It may receive execution evidence from the main skill or a thin helper.
- It does not choose the overall target on its own.
- It does not perform broad paper analysis.
- It does not own training startup, resume, or long-running training state.
- It should not normalize risky code edits into acceptable practice.
- It must not hide changes that alter evaluation, preprocessing, checkpoints, metrics, or other scientific meaning.
Input expectations
- selected reproduction goal
- runnable commands or smoke commands
- environment and asset assumptions
- optional patch metadata
Output expectations
- execution result summary
- standardized
repro_outputs/files SCIENTIFIC_CHANGELOG.mdfor changed scientific meaning and evidence statusCOMPARABILITY_REPORT.mdfor README/paper/baseline comparability- clear distinction between verified, partial, and blocked states
PATCHES.mdwhen repo files changed
Notes
Use references/reporting-policy.md, ../ai-research-reproduction/references/research-rigor-principles.md, scripts/run_command.py, and scripts/write_outputs.py.
Files (rigorpilot-skills)
-
agents
-
openai.yaml 313 B
display_name: Rigor Run short_description: Rigor Run mode for selected inference, evaluation, smoke, or sanity execution evidence. default_prompt: Run the selected smoke, inference, evaluation, or sanity command conservatively, capture execution evidence, and write SUMMARY.md COMMANDS.md LOG.md and status.json.
-
-
references
-
reporting-policy.md 747 B
# Reporting Policy ## Tone Keep reports short, factual, and easy to audit. ## Requirements - separate facts from inferences - mention the documented command explicitly - mention whether the non-training run was full, partial, smoke-only, sanity-only, or blocked - explain the main blocker without burying it - when patches were applied, mention patch state briefly in `SUMMARY.md` and keep the full audit in `PATCHES.md` ## Output priorities 1. clear overall result 2. copyable commands 3. concise process trace 4. stable machine-readable state 5. patch evidence when relevant ## Avoid - long narrative journals - vague "it should work" language - hiding unsupported assumptions - treating training startup or resume as part of this skill
-
-
scripts
-
run_command.py 13.3 KB
#!/usr/bin/env python3 """Execute a short non-training command and normalize the evidence.""" from __future__ import annotations import argparse import json import re import subprocess import sys from pathlib import Path from typing import Any, Dict, Iterable, List, Optional, Tuple SHARED_SCRIPTS = Path(__file__).resolve().parents[3] / "shared" / "scripts" if not all((SHARED_SCRIPTS / name).is_file() for name in ( "runtime_runner.py", "model_adapter.py", "command_utils.py", "resource_monitor.py" )): SHARED_SCRIPTS = (Path(__file__).resolve().parents[2] / "ai-research-reproduction" / "_bundled" / "shared" / "scripts") if not (SHARED_SCRIPTS / "model_adapter.py").is_file(): raise RuntimeError("Shared runtime missing: install all RigorPilot skills, including ai-research-reproduction.") if str(SHARED_SCRIPTS) not in sys.path: sys.path.insert(0, str(SHARED_SCRIPTS)) from runtime_runner import run_persistent_command from model_adapter import ModelAdapterError, load_model_profile, missing_capabilities METRIC_RE = re.compile( r"\b([A-Za-z][A-Za-z0-9_.-]{1,31})\s*[:=]\s*(-?\d+(?:\.\d+)?(?:[eE][+-]?\d+)?)" ) def combine_logs(parts: Iterable[str]) -> str: return "\n".join(part for part in parts if part).strip() def decode_stream(value: Any) -> str: # On POSIX, subprocess.TimeoutExpired carries captured output as bytes # even when the run was started with text=True. if isinstance(value, bytes): return value.decode("utf-8", errors="replace") return value or "" def parse_metrics(text: str) -> Dict[str, Any]: observed_metrics: Dict[str, float] = {} best_metric: Optional[Dict[str, Any]] = None for match in METRIC_RE.finditer(text): name = match.group(1) value = float(match.group(2)) observed_metrics[name] = value priority_names = [ name for name in observed_metrics if not any(token in name.lower() for token in {"loss", "lr", "time", "mem", "epoch", "step", "iter", "iteration"}) ] if priority_names: chosen = priority_names[-1] best_metric = {"name": chosen, "value": observed_metrics[chosen]} elif observed_metrics: chosen = list(observed_metrics)[-1] best_metric = {"name": chosen, "value": observed_metrics[chosen]} return { "observed_metrics": observed_metrics, "best_metric": best_metric, } def run_git(repo: Path, args: List[str]) -> subprocess.CompletedProcess[str]: try: return subprocess.run( ["git", *args], cwd=repo, capture_output=True, text=True, timeout=15, check=False, ) except (FileNotFoundError, subprocess.TimeoutExpired) as exc: # A missing or hanging git binary must degrade to the documented # "git-unavailable" evidence path, not crash the runner. return subprocess.CompletedProcess(["git", *args], returncode=127, stdout="", stderr=str(exc)) def git_status_snapshot(repo: Path) -> Tuple[Optional[Dict[str, str]], Dict[str, Any]]: probe = run_git(repo, ["rev-parse", "--is-inside-work-tree"]) if probe.returncode != 0 or probe.stdout.strip() != "true": return None, { "collection_method": "git-status-diff", "available": False, "reason": "git-unavailable-or-not-a-worktree", } result = run_git(repo, ["status", "--porcelain=v1", "--untracked-files=all"]) if result.returncode != 0: return None, { "collection_method": "git-status-diff", "available": False, "reason": "git-status-failed", "stderr": result.stderr.strip(), } snapshot: Dict[str, str] = {} for raw_line in result.stdout.splitlines(): line = raw_line.rstrip() if len(line) < 4: continue status = line[:2] path = line[3:] if " -> " in path: _old, _arrow, path = path.partition(" -> ") normalized = path.replace("\\", "/").strip() if normalized: snapshot[normalized] = status return snapshot, { "collection_method": "git-status-diff", "available": True, "status_entries": len(snapshot), } def diff_status_snapshots( before: Optional[Dict[str, str]], after: Optional[Dict[str, str]], ) -> Dict[str, List[str]]: if before is None or after is None: return { "changed_files": [], "new_files": [], "deleted_files": [], "touched_paths": [], "touched_symbols": [], } changed_files: List[str] = [] new_files: List[str] = [] deleted_files: List[str] = [] for path, status in after.items(): previous_status = before.get(path) if previous_status == status: continue normalized_status = status.replace(" ", "") if "D" in normalized_status: deleted_files.append(path) continue if "?" in normalized_status or "A" in normalized_status: new_files.append(path) continue changed_files.append(path) touched_paths = [] for path in [*changed_files, *new_files, *deleted_files]: if path not in touched_paths: touched_paths.append(path) return { "changed_files": changed_files, "new_files": new_files, "deleted_files": deleted_files, "touched_paths": touched_paths, "touched_symbols": [], } def exclude_runtime_snapshot( repo: Path, runtime_dir: Path, snapshot: Optional[Dict[str, str]], ) -> Optional[Dict[str, str]]: if snapshot is None: return None try: prefix = runtime_dir.resolve().relative_to(repo.resolve()).as_posix().rstrip("/") + "/" except ValueError: return snapshot return {path: status for path, status in snapshot.items() if not path.startswith(prefix)} def execute_command( repo: Path, command: str, timeout: int, shell_mode: str = "direct", runtime_root: Optional[Path] = None, model_adapter: Optional[Dict[str, Any]] = None, monitor_gpu: bool = False, ) -> Dict[str, Any]: before_status, before_capture = git_status_snapshot(repo) selected_runtime_root = (runtime_root or (repo / "repro_outputs" / "_runtime")).resolve() execution = run_persistent_command( repo=repo, command=command, timeout=timeout, runtime_root=selected_runtime_root, shell_mode=shell_mode, model_adapter=model_adapter, monitor_gpu=monitor_gpu, ) after_status, after_capture = git_status_snapshot(repo) after_status = exclude_runtime_snapshot(repo, Path(execution["runtime_dir"]), after_status) if after_status is not None: after_capture["status_entries"] = len(after_status) after_capture["runtime_artifacts_excluded"] = True execution.update(diff_status_snapshots(before_status, after_status)) execution["evidence_capture"] = { **after_capture, "before_status_entries": before_capture.get("status_entries"), } return execution def decide_outcome(command: str, timeout: int, execution: Dict[str, Any], metric_data: Dict[str, Any]) -> Dict[str, Any]: combined_text = combine_logs( [ f"STDOUT:\n{execution['stdout'].strip()}" if execution.get("stdout", "").strip() else "", f"STDERR:\n{execution['stderr'].strip()}" if execution.get("stderr", "").strip() else "", ] ) if execution.get("launch_error"): return { "status": "blocked", "documented_command_status": "blocked", "main_blocker": f"Executable not found for command: {execution['launch_error']}", "execution_log": [f"Command failed before launch: {execution['launch_error']}"], "monitoring_scope": "no_run", } if execution.get("cancelled"): return { "status": "partial", "documented_command_status": "partial", "main_blocker": "The selected command was cancelled through the runtime control file.", "execution_log": [combined_text] if combined_text else ["Command cancelled."], "monitoring_scope": "runtime_cancel", } if execution.get("timed_out"): return { "status": "partial", "documented_command_status": "partial", "main_blocker": f"Selected command did not finish within {timeout} seconds.", "execution_log": [combined_text or f"Command timed out after {timeout} seconds."], "monitoring_scope": f"timeout:{timeout}s", } if execution.get("returncode") == 0: return { "status": "success", "documented_command_status": "success", "main_blocker": "None.", "execution_log": [combined_text] if combined_text else [], "monitoring_scope": "process_completion", } return { "status": "partial", "documented_command_status": "partial", "main_blocker": f"Selected command exited with code {execution.get('returncode')}.", "execution_log": [combined_text] if combined_text else [f"Command `{command}` exited non-zero."], "monitoring_scope": "process_completion", } def main() -> int: parser = argparse.ArgumentParser(description="Run a short non-training command and summarize the evidence.") parser.add_argument("--repo", required=True, help="Path to the target repository.") parser.add_argument("--command", required=True, help="Command to execute.") parser.add_argument("--timeout", type=int, default=60, help="Execution timeout in seconds.") parser.add_argument( "--shell-mode", choices=["direct", "native"], default="direct", help="Use direct argv execution by default; native shell execution requires explicit opt-in.", ) parser.add_argument( "--runtime-root", default="", help="Directory for persistent runtime state and streamed logs (default: <repo>/repro_outputs/_runtime).", ) parser.add_argument("--model-profile-json", default="", help="Optional provider-neutral model identity/capability profile.") parser.add_argument( "--require-model-capability", action="append", default=[], help="Required model capability; repeat as needed.", ) parser.add_argument("--monitor-gpu", action="store_true", help="Sample NVIDIA device-level telemetry when available.") args = parser.parse_args() if args.timeout <= 0: parser.error("--timeout must be greater than zero") repo = Path(args.repo).resolve() runtime_root = Path(args.runtime_root).resolve() if args.runtime_root else None try: model_adapter = load_model_profile(Path(args.model_profile_json) if args.model_profile_json else None) missing = missing_capabilities(model_adapter, args.require_model_capability) except ModelAdapterError as exc: parser.error(str(exc)) if missing: parser.error(f"model profile is missing required capabilities: {', '.join(missing)}") execution = execute_command( repo, args.command, args.timeout, args.shell_mode, runtime_root, model_adapter, args.monitor_gpu, ) metric_data = parse_metrics(combine_logs([execution.get("stdout", ""), execution.get("stderr", "")])) outcome = decide_outcome(args.command, args.timeout, execution, metric_data) payload = { "status": outcome["status"], "documented_command_status": outcome["documented_command_status"], "main_blocker": outcome["main_blocker"], "execution_log": outcome["execution_log"], "monitoring_scope": outcome["monitoring_scope"], "execution_mode": execution.get("execution_mode", args.shell_mode), "runtime_run_id": execution.get("runtime_run_id"), "runtime_dir": execution.get("runtime_dir"), "runtime_status": execution.get("runtime_status"), "runtime_state_path": execution.get("runtime_state_path"), "runtime_events_path": execution.get("runtime_events_path"), "stdout_log_path": execution.get("stdout_log_path"), "stderr_log_path": execution.get("stderr_log_path"), "stdout_truncated": execution.get("stdout_truncated", False), "stderr_truncated": execution.get("stderr_truncated", False), "cancelled": execution.get("cancelled", False), "duration_seconds": execution.get("duration_seconds"), "runtime_attempt": execution.get("runtime_attempt", 1), "runtime_retry_of": execution.get("runtime_retry_of"), "resources_log_path": execution.get("resources_log_path"), "resource_summary": execution.get("resource_summary", {}), "model_adapter": execution.get("model_adapter"), "best_metric": metric_data["best_metric"], "observed_metrics": metric_data["observed_metrics"], "changed_files": execution.get("changed_files", []), "new_files": execution.get("new_files", []), "deleted_files": execution.get("deleted_files", []), "touched_paths": execution.get("touched_paths", []), "touched_symbols": execution.get("touched_symbols", []), "evidence_capture": execution.get("evidence_capture", {}), } print(json.dumps(payload, indent=2, ensure_ascii=False)) return 0 if __name__ == "__main__": raise SystemExit(main()) -
write_outputs.py 1.1 KB
#!/usr/bin/env python3 """Compatibility wrapper for trusted verify output bundles.""" from __future__ import annotations import importlib.util from pathlib import Path def load_shared_module(): module_path = Path(__file__).resolve().parents[3] / "shared" / "scripts" / "write_run_bundle.py" if not module_path.is_file(): module_path = (Path(__file__).resolve().parents[2] / "ai-research-reproduction" / "_bundled" / "shared" / "scripts" / "write_run_bundle.py") if not module_path.is_file(): raise RuntimeError("Shared writer missing: install all RigorPilot skills, including ai-research-reproduction.") spec = importlib.util.spec_from_file_location("write_run_bundle", module_path) if spec is None or spec.loader is None: raise RuntimeError(f"Unable to load shared writer module from {module_path}") module = importlib.util.module_from_spec(spec) spec.loader.exec_module(module) return module def main() -> int: module = load_shared_module() return module.main(default_mode="repro", default_output_dir="repro_outputs") if __name__ == "__main__": raise SystemExit(main())
-
-
SKILL.md 2.7 KB
--- name: minimal-run-and-audit description: Rigor Run skill for README-first deep learning repo reproduction. Use when the task is specifically to capture or normalize evidence from the selected smoke test or documented inference or evaluation command and write standardized `repro_outputs/` files, including patch notes when repository files changed. Do not use for training execution, initial repo intake, generic environment setup, paper lookup, target selection, hidden scientific-meaning changes, or end-to-end orchestration by itself. --- # minimal-run-and-audit Use this as the Rigor Run skill. The installed slug remains `minimal-run-and-audit` for compatibility. Use the shared operating principles in `../ai-research-reproduction/references/agent-operating-principles.md`; this skill should make run evidence auditable without turning every command into a rigid protocol. ## When to apply - After a reproduction target and setup plan exist. - When the main skill needs execution evidence and normalized outputs. - When a smoke test, documented inference run, documented evaluation run, or other short non-training verification is appropriate. - When the user already knows what command should be attempted and wants execution plus reporting only. ## When not to apply - During initial repo scanning. - When environment or assets are still undefined enough to make execution meaningless. - When the task is a literature lookup rather than repository execution. - When the user is still deciding which reproduction target should count as the main run. ## Clear boundaries - This skill owns normalized reporting for an attempted command. - It may receive execution evidence from the main skill or a thin helper. - It does not choose the overall target on its own. - It does not perform broad paper analysis. - It does not own training startup, resume, or long-running training state. - It should not normalize risky code edits into acceptable practice. - It must not hide changes that alter evaluation, preprocessing, checkpoints, metrics, or other scientific meaning. ## Input expectations - selected reproduction goal - runnable commands or smoke commands - environment and asset assumptions - optional patch metadata ## Output expectations - execution result summary - standardized `repro_outputs/` files - `SCIENTIFIC_CHANGELOG.md` for changed scientific meaning and evidence status - `COMPARABILITY_REPORT.md` for README/paper/baseline comparability - clear distinction between verified, partial, and blocked states - `PATCHES.md` when repo files changed ## Notes Use `references/reporting-policy.md`, `../ai-research-reproduction/references/research-rigor-principles.md`, `scripts/run_command.py`, and `scripts/write_outputs.py`.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.