Claude opencode Skill

rpi

Coordinate one RPI traversal: one bounded Plan and Implement experiment, then fresh Validate and a bounded repair phase to convergence. Triggers: "run rpi", "run one traversal", "execute this plan", orchestration or worker delegation that implements changes.

LLM Mart · 0 points · 9 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download boshu2-agentops-skills_rpi-9ac484e.zip · 24 KB
boshu2/agentops 445 41 forks Apache-2.0 Updated 1d ago
Part of boshu2/agentops — 73 skills

Install

skills CLI npx skills add https://github.com/boshu2/agentops/tree/main/skills/rpi
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install boshu2-agentops@llmmart
Git git clone https://github.com/boshu2/agentops.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole boshu2/agentops collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

RPI

Own the authorized outcome through finish. Use the native coding agent and shell. BD or the caller's tracker owns work and handoffs; Git owns content and delivery. AgentOps supplies a small charter and fresh judgment, not a scheduler.

Operating charter

  1. Use the existing accepted outcome, scope and real bounds. A clear change needs no Plan, Recall or Learn worksheet. Resolve uncertainty only when it could change the implementation or acceptance decision.
  2. Take the smallest acceptance-advancing action. Plan shapes missing intent or revises a disproved approach. Once an implementer can act and a validator can judge, implement; do not keep improving the plan. Approach revisions preserve acceptance and authorized scope. Acceptance changes need caller authority.
  3. Implement and repair ordinary known defects directly. A known test failure needs a fix and a discriminating check, not another planning phase, council or helper.
  4. Use focused checks during edits and complete required integration checks before final judgment. Reuse valid exact-input receipts; rerun affected checks after changes. Reserve capacity for integration, review and repair. Keep the final subject unchanged while it is being judged.
  5. Obtain Validate from one fresh author-distinct context in the author's model family unless the caller selects required additional legs. There is no fixed ten-minute cap; explicitly required reviewers remain required. Risk deepens evidence, not reviewer multiplication. Repair actionable findings within authority and remaining bounds, then revalidate the changed exact subject.
  6. Stop at completed acceptance, cancellation, refusal, a spent real bound or an unresolved causal stall after the help below. Adjacent improvements are not permission to expand the goal. Report them briefly only when useful; do not turn them into another work batch.

Context and handoffs

Load required contracts once per context, then read only what the next decision needs. A reference link is available context, not a reading list. Search before opening large files; expand only for consequential uncertainty. Keep successful output compact at the tool boundary; retain full logs for inspection. Reuse the worker's component-check list and current receipts instead of rediscovering them. Use native completion watches or bounded waits for ongoing checks and helpers. At completion, verify the expected subject and required results; a quiet or partial status is not success. Inspect further for a failure, suspected stall or decision need. Keep required user updates concise rather than narrating each poll.

When delegation is authorized and useful, select the runtime's task-only dispatch option for independent work; a short prompt in a full-history fork still carries full history. Supply accepted intent/scope, exact subject, relevant evidence, remaining bounds, result consumer and check ownership. Resume an author for direct repair when useful. Validators always receive fresh context without the author's desired verdict. Observe actual dispatch settings; prompt wording proves neither isolation nor smaller inherited context.

Return concise findings, check facts and evidence references in the existing handoff; disclose missing or truncated evidence. Derive the combined subject's manifest and applicable orphan scan at the integration/judgment boundary. Unjudged worker increments supply content identity and check facts, not duplicate final evidence bundles. A separately judged subject still needs complete proof. Machine evidence such as verdict.v2 is optional unless requested or required by a declared consumer. When no machine artifact is requested or required, return the result without creating one.

Causal stall and bounds

Unknown cause, recurrence, no progress or a wrong objective admits at most one bounded fresh helper for that incident within authority and bounds. Give it the failed assumption, evidence and one discriminating question. Resume only with a different testable approach; an unhelpful answer ends the attempt. Do not chain helpers or rename the incident. Known failures get direct repair. Cancellation, refusal and spent hard time/cost/quota skip help.

Respect actual caller/native limits, including explicit repair-round bounds. Retries, compaction, helpers and new subjects never renew them; retry count alone is not a spent budget. If interruption threatens evidence, preserve accepted intent, exact subject, useful receipts, unresolved cause, bounds and helper use in the native handoff. Prompt text proves no native enforcement. Outer-goal guidance remains optional.

Evidence and boundaries

Bind accepted intent, complete changed paths, exact subject and factual receipts for the fresh validator; disclose affected orphaned acceptance evidence. Use existing provenance helpers rather than a new evidence format. Requested proof uses caller-selected protected external non-Git storage; preserve legacy .agents/ evidence. Missing identity, freshness or proof means NOT_PROVEN; proven failed acceptance or scope violation means FAIL. PASS needs every criterion verified and empty not_checked. Authors cannot issue binding PASS.

Memory, specialists and runtime adapters are on demand; no-match and no-change are valid. Read boundaries when authority, scope, evidence or delivery is at issue. The optional fixed-dispatch adapter is not the native execution engine. Do not invent a runtime, hidden machine artifact or workflow to finish an ordinary change.

Report the result, strongest checks and material limits. Plans, activity, reviews and saved pages earn no capability credit; NOT_PLANNED and NOT_BUILT are progress descriptions, not semantic verdicts.

Files (agentops)
  • agents
    • openai.yaml 43 B
      policy:
        allow_implicit_invocation: false
      
  • references
    • boundaries.md 4.7 KB
      # Ownership boundaries for the lean RPI core
      
      RPI owns the authorized outcome through implementation, checks, direct repairs
      and fresh final judgment. Plan shapes missing intent and may revise an approach
      falsified by evidence within unchanged accepted outcome/scope. Implement edits
      and collects facts. Validate independently judges the exact subject and alone
      authors semantic `verdict.v2` when persistence is selected. Memory is optional;
      its operation references own recall, mining and curation.
      
      ## Native authority
      
      BD or the caller's tracker owns work/status/dependencies/handoffs. Git and
      repository policy own content/history and delivery. Native runtimes and callers
      own aggregate budgets, work selection, queues, claims, stops and subsequent
      outcomes. A skill grants no extra Git, tracker, publishing or credential
      permission. Existing caller authorization remains usable; do not invent another
      approval step merely because a phase changed. Keep one authoritative work
      account, not a parallel AgentOps ledger.
      
      The runtime derives exact intent/subject identity, complete changed paths,
      receipts and observed context identities. Never invent a model/context identity
      or transcribe a fictional runtime packet. New requested proof uses protected
      external non-Git storage; preserve legacy `.agents/` proof under owner policy.
      Plans, dashboards and reviews count as subject completion only when requested.
      
      ## Direct repair and help
      
      Known failures get direct repair. Evidence that disproves an assumption permits
      approach revision under unchanged acceptance and scope. Acceptance or authority
      expansion needs caller approval; useful source/generated changes already covered
      by a scope class do not. Cheap discriminating checks precede expensive judgment.
      Reserve finishing capacity and use valid exact-input receipts when applicable.
      
      Unknown cause, recurrence, no progress or wrong objective warrants causal
      examination. A genuine stall admits at most one bounded fresh helper per incident
      inside existing authority and bounds. Do not build a helper chain for known
      failures or rename an unresolved incident. An unhelpful helper ends the attempt.
      Cancellation, refusal or spent real limits skip help; retry counts alone are not
      spent time/quota. Compact native recovery state preserves evidence, not new budget.
      
      ## Optional specialists and adapters
      
      Anti-ceremony, premortem, council, research, factories and runtime adapters are
      optional. Risk deepens evidence inspection without mandatory specialist dispatch.
      No Recall or Learn toll applies to trivial work. A selected factory remains
      behind its own coordinator, doctor and supervisor doors; its reconciler creates
      and repairs sessions. Concurrent writers require authorized disjoint source and
      regeneration scope and isolation. Pass bounded task evidence, not the author's
      desired verdict. Do not start another runtime merely because it exists.
      
      The optional `run_once.py` developer adapter retains its explicitly selected
      fixed-dispatch and finite-round contract in [bounded-adapter.md](bounded-adapter.md).
      It does not restrict native approach revision or implement direct repair for you.
      
      ## Fresh judgment
      
      The author cannot issue binding PASS. Judge legs read; implementers fix. Default
      to a fresh author-distinct same-family reviewer. Cross-model review is opt-in;
      an explicitly required unavailable leg leaves NOT_PROVEN. No fixed ten-minute
      cap applies, and no invocation renews caller/native limits.
      
      PASS needs exact subject continuity, complete changed-path coverage, unchanged
      acceptance, distinct context IDs and attested freshness, nonempty checked scope,
      evidence for every criterion and empty `not_checked`. Incomplete proof remains
      NOT_PROVEN; failed acceptance or proven out-of-scope changes remain FAIL.
      Necessary findings never become optional to get green. Judge disagreement stays
      visible and never becomes PASS by preference or majority vote.
      
      Validate returns judgment, not a repair or delivery instruction. RPI completes
      existing authorized work before reporting, within real bounds. Report the
      subject, strongest evidence and any remaining acceptance gaps; persist a machine
      artifact only for a declared consumer or caller request. A new subject requires
      new final judgment. Mutating checks run on a disposable copy or committed subject
      so they cannot overwrite the judged working tree.
      
      ## Observed guardrails and limits
      
      The July 2026 unlisted-regeneration incident supports scope as a class; it does
      not authorize unrelated files. The July mutating-check incident supports the
      quarantine; it does not require rerunning every expensive check. The planning
      spiral supports smallest useful action; it does not forbid revising a falsified
      approach. These rules protect actual work and may be revised by later evidence.
      
    • bounded-adapter.md 3.9 KB
      # Optional fixed-dispatch reference adapter
      
      The grandfathered `scripts/run_once.py` is a pure developer reference for callers
      that explicitly select fixed dispatch and a finite list of supplied review rounds.
      It invokes an explicit anti-ceremony function, Plan and Implement at most once;
      its repair evaluator consumes supplied evidence and cannot execute agents, infer
      causes, fix subjects or enforce aggregate budgets. Installed native RPI follows
      its operating charter, not this adapter. The old phase lock and default two-round
      limit apply only to this selected adapter, never as a restriction on native
      implementation's direct repairs or evidence-driven approach revision.
      
      Its existing deterministic tests guard exact evidence and finite consumption;
      they do not prove native agent behavior or practical benefit. It stops when
      converged, stopped by the law, or out of `repair_rounds` and never extends the
      caller's bound. The adapter preserves these narrower admission semantics:
      
      ## The convergence law
      
      A repair round is admitted only while all hold:
      
      1. `rounds_used < repair_rounds` (caller-declared, default 2).
      2. New digest-bound evidence proves closure of a named acceptance finding or,
         for `NOT_PROVEN`, resolves a named proof gap. A changed digest or a smaller
         finding count alone is not useful progress. Generated-only changes qualify
         only when the evidence proves that they repair required behavior or parity.
         An unchanged subject previously judged FAIL cannot be repaired by a new label
         or verdict flip; changed bytes still require acceptance proof.
      3. No finding id closed in an earlier round reopens. No closed finding class
         recurs, and no introduced regression or new finding of unknown cause is
         admitted. Before/after reproduction or equivalent causal evidence under the
         same acceptance must distinguish a pre-existing discovery from a regression;
         neither counts, timestamps, nor a new id establish that distinction.
      
      Keep the union of every required judge's findings, keyed by stable
      `findings[].id`; do not hide a necessary finding as optional. Newly exposed
      pre-existing defects may increase the open count while another acceptance gap
      is demonstrably closed. Their evidence must prove prior existence;
      unknown cause stops repair for causal examination even if another gap closed.
      Validators reuse a short stable `class` for each kind of defect. A reopened id
      or returning class warrants causal HOLD in a selected outer goal. Recurrence
      alone does not prove that the design is wrong and never auto-reopens Plan.
      
      Reuse existing check receipts, findings summaries, and evidence references for
      this reasoning. In the pure reference, decoded receipt bindings use `ref`,
      `subject_digest`, and `resolves` for ids actually closed. `preexisting` ids must
      bind reproduction to the prior subject digest; `introduced` ids bind causal
      comparison to the current digest and stop repair. These are supplied receipt
      facts, not new persisted verdict fields or a lifecycle schema. The reference
      cannot prove a receipt's truth or infer cause from wording.
      
      Converged: the fresh validator returns PASS and every required cross-family
      validator does too, over the exact subject and all acceptance with empty
      `not_checked`. On any violation RPI stops and reports the current status.
      `checked` carries one line per round (`repair round N: k open findings`); open
      findings ride in the result and the report. A reworded finding with the same id
      is the same finding. Acceptance and its digest stay fixed. The orchestrating
      context fixes; judge legs only read. RPI convenes no further judge of its own,
      does not escalate, and does not auto-replan.
      
      
      An unknown cause or recurrence returns evidence to the native caller; the
      adapter dispatches no helper. The native charter decides whether a genuine
      causal stall merits its single bounded consultation. This does not revive a
      spent caller bound or change a completed verdict.
      
    • outer-goal.md 1.8 KB
      # Optional outer goal
      
      Use this reference only when the caller explicitly selected a sustained goal or
      several outcomes. The native caller/controller keeps work selection, aggregate
      budgets, stops and delivery authority. No AO scheduler, new command or goal
      ledger is required. The RPI charter already owns a single authorized outcome
      through finish; an outer goal is not permission needed for ordinary repair.
      
      Carry accepted terminal outcome and scope, measured remaining allowance and the
      current causal incident in the native work/handoff source. Choose the smallest
      acceptance-advancing action or consequential uncertainty. Reserve capacity for
      integration, final fresh judgment, required repairs and a useful handoff before
      spending the whole allowance on discovery or reviews.
      
      Apply the charter's at-most-one bounded helper to a genuine causal stall. Known
      failures get direct repair. A repeated wakeup, new context or renamed finding is
      not a new incident. A helper with no useful new approach ends that attempt;
      report the unresolved cause and required caller decision. Cancellation, refusal
      and spent real bounds skip help. Native blocked-status thresholds are bookkeeping,
      not permission to renew time, cost, quota or helper use.
      
      Report observed native enforcement and unmeasured limits honestly. No objective
      text, saved plan or simulated stop proves aggregate runtime enforcement. The
      caller may authorize a new outcome or scope; agents may revise an approach when
      evidence disproves an assumption within unchanged authority.
      
      Existing host/user policies may impose stricter helper or stopping requirements.
      This repository charter does not update those installed host instructions.
      Inspect and report the effective upstream rule instead of claiming the lean
      contract overrides it or that an optional guide enforces native controls.
      
    • rpi.feature 2.9 KB · in bundle
  • scripts
    • run_once.py 24.2 KB
      #!/usr/bin/env python3
      """Pure reference behavior for one RPI invocation and its bounded repair phase.
      
      The caller supplies one anti-ceremony guard and the three core phase functions.
      This module invokes the guard once before Plan, dispatches Plan and Implement at
      most once, and never chooses a retry, a budget, or a next action.
      
      Under ADR-0017 (loop as control flow, not knowledge) the traversal no longer
      ends at the first validation result. `run_repair_phase` models the bounded
      repair phase as pure data: it consumes validate rounds that already happened and
      decides, under the convergence law, whether another repair round is admitted.
      It performs no I/O, dispatches nothing, and owns no budget of its own — the
      caller declares `repair_rounds`.
      """
      
      from __future__ import annotations
      
      from collections.abc import Callable, Mapping, Sequence
      import re
      from typing import Any
      
      
      # The exact-identity property is BYTE-addressed: Validate snapshots the resolved
      # intent bytes under `sha256(bytes)` and stores them as `<digest>.intent`
      # (validate.py snapshot_intent), then re-derives that same digest from the same
      # bytes when it binds runtime facts into the verdict. RPI is a dispatcher, not a
      # second digest authority — it carries the digest Plan declares over the bytes it
      # snapshotted, and cross-checks Validate's independently re-derived value against
      # it.
      #
      # This module previously computed its own `sha256(canonical-JSON(mapping))` here
      # and hard-compared that against Validate's `sha256(raw bytes)`. The two can
      # never agree unless the source is byte-identical canonical JSON, so the composed
      # contract was broken; both unit suites stayed green only because the RPI test
      # mocked Validate with THIS module's digest function. A canonical-JSON digest is
      # also the wrong identity in principle: two different source files that parse to
      # the same mapping share it, which is precisely the collision exact identity
      # exists to forbid.
      DIGEST_PATTERN = re.compile(r"^[0-9a-f]{64}$")
      
      
      def valid_digest(value: Any) -> bool:
          """True for a lowercase hex SHA-256, the only shape an identity may take."""
          return isinstance(value, str) and bool(DIGEST_PATTERN.match(value))
      
      
      def valid_string_list(value: Any) -> bool:
          """True for the guard contract's JSON-shaped string lists."""
          return isinstance(value, list) and all(
              isinstance(item, str) and bool(item.strip()) for item in value
          )
      
      
      def guard_result(value: Any) -> dict[str, Any]:
          """Return one valid artifact-free anti-ceremony decision."""
          if not isinstance(value, Mapping):
              raise ValueError("anti-ceremony guard must return a mapping")
          result = dict(value)
          expected = {
              "decision",
              "reason",
              "frozen_outcome",
              "parked_process_work",
              "remaining_proof",
              "stop_condition",
          }
          if set(result) != expected:
              raise ValueError("anti-ceremony guard returned the wrong fields")
          if result["decision"] not in {"CONTINUE", "STOP"}:
              raise ValueError("anti-ceremony decision must be CONTINUE or STOP")
          reason = result["reason"]
          if (
              not isinstance(reason, str)
              or not reason.strip()
              or "\n" in reason
              or reason[-1] not in ".!?"
              or sum(reason.count(mark) for mark in ".!?") != 1
          ):
              raise ValueError("anti-ceremony reason must be exactly one sentence")
          if not isinstance(result["frozen_outcome"], str) or not result["frozen_outcome"].strip():
              raise ValueError("anti-ceremony frozen_outcome must be a nonempty string")
          if not valid_string_list(result["parked_process_work"]):
              raise ValueError("anti-ceremony parked_process_work must be a string list")
          if not valid_string_list(result["remaining_proof"]):
              raise ValueError("anti-ceremony remaining_proof must be a string list")
          if not isinstance(result["stop_condition"], str) or not result["stop_condition"].strip():
              raise ValueError("anti-ceremony stop_condition must be a nonempty string")
          return result
      
      
      def report(
          status: str,
          *,
          intent_ref: str | None = None,
          acceptance_digest: str | None = None,
          subject_digest: str | None = None,
          verdict_ref: str | None = None,
          verdict_digest: str | None = None,
          checked: list[str] | None = None,
          not_checked: list[str] | None = None,
      ) -> dict[str, Any]:
          return {
              "schema_version": "rpi-report.v1",
              "status": status,
              "intent_ref": intent_ref,
              "acceptance_digest": acceptance_digest,
              "subject_manifest_digest": subject_digest,
              "verdict_ref": verdict_ref,
              "verdict_digest": verdict_digest,
              "checked": checked or [],
              "not_checked": not_checked or [],
          }
      
      
      def invoke_once(
          intent: Any,
          anti_ceremony_guard: Callable[[Any], Mapping[str, Any]],
          plan_phase: Callable[[Any], Mapping[str, Any] | None],
          implement_phase: Callable[[Mapping[str, Any]], Mapping[str, Any] | None],
          validate_phase: Callable[[Mapping[str, Any], Mapping[str, Any]], Mapping[str, Any]],
      ) -> dict[str, Any]:
          """Invoke the guard once, then dispatch each core phase at most once."""
          admission = guard_result(anti_ceremony_guard(intent))
          if admission["decision"] == "STOP":
              return report(
                  "NOT_PLANNED",
                  checked=[f"anti-ceremony guard: STOP — {admission['reason']}"],
                  not_checked=["plan", "implement", "validate"],
              )
          resolved_intent = plan_phase(intent)
          if resolved_intent is None:
              return report("NOT_PLANNED", not_checked=["implement", "validate"])
          resolved_intent = dict(resolved_intent)
          intent_ref = resolved_intent.get("intent_ref")
          if not isinstance(intent_ref, str) or not intent_ref:
              intent_ref = "caller"
          acceptance_digest = resolved_intent.get("acceptance_digest")
          if not valid_digest(acceptance_digest):
              raise ValueError(
                  "Plan must declare acceptance_digest as the SHA-256 of the exact resolved "
                  "intent bytes it snapshotted (validate.py snapshot-intent emits it)"
              )
      
          subject = implement_phase(resolved_intent)
          if subject is None:
              return report(
                  "NOT_BUILT",
                  intent_ref=intent_ref,
                  acceptance_digest=acceptance_digest,
                  checked=["plan"],
                  not_checked=["validate"],
              )
          subject = dict(subject)
      
          validation = dict(validate_phase(resolved_intent, subject))
          status = validation.get("verdict")
          if status not in {"PASS", "FAIL", "NOT_PROVEN"}:
              raise ValueError("Validate must return PASS, FAIL, or NOT_PROVEN")
          # Validate re-derives this from the snapshot bytes independently; equality
          # here is the composed exact-identity check, not a self-comparison.
          if validation.get("acceptance_digest") != acceptance_digest:
              raise ValueError("Validate verdict does not match the resolved intent digest")
          subject_digest = validation.get("subject_manifest_digest")
          if not valid_digest(subject_digest):
              raise ValueError("Validate must return the exact subject manifest digest")
          candidate_digest = subject.get("subject_manifest_digest")
          if candidate_digest is not None and subject_digest != candidate_digest:
              raise ValueError("Validate result does not match the implemented subject digest")
          author_context_id = validation.get("author_context_id")
          validator_context_id = validation.get("validator_context_id")
          freshness = validation.get("freshness_attestation")
          if (
              not isinstance(author_context_id, str)
              or not author_context_id
              or not isinstance(validator_context_id, str)
              or not validator_context_id
              or author_context_id == validator_context_id
              or not isinstance(freshness, Mapping)
              or freshness.get("source") not in {"runtime", "caller"}
              or not isinstance(freshness.get("attester_identity"), str)
              or not freshness.get("attester_identity")
          ):
              raise ValueError("Validate must return distinct context identities and explicit freshness")
          verdict_digest = validation.get("verdict_digest")
          verdict_ref = validation.get("verdict_ref")
          if (verdict_digest is None) != (verdict_ref is None):
              raise ValueError("Validate must return both verdict_ref and verdict_digest when persistence is requested")
          if verdict_ref is not None and (
              not isinstance(verdict_ref, str)
              or not verdict_ref
              or not valid_digest(verdict_digest)
          ):
              raise ValueError("Persisted verdict identity is invalid")
          return report(
              status,
              intent_ref=intent_ref,
              acceptance_digest=acceptance_digest,
              subject_digest=subject_digest,
              verdict_ref=verdict_ref,
              verdict_digest=verdict_digest,
              checked=list(validation.get("checked") or []),
              not_checked=list(validation.get("not_checked") or []),
          )
      
      
      # ---------------------------------------------------------------------------
      # The bounded repair phase (ADR-0017)
      # ---------------------------------------------------------------------------
      #
      # The 2026-07-14 cathedral cut removed the iterate loop together with the
      # unproven compounding claim, although ADR-0011 demoted only the latter. What
      # comes back is control flow, not knowledge: a repair round is admitted only
      # while every condition of the convergence law holds. Byte movement and finding
      # counts are identity/accounting facts, not evidence of acceptance progress.
      #
      # Recurrence is checked before progress. It requires causal examination by the
      # caller, not an automatic claim that the design is wrong. This pure reference
      # neither diagnoses causes nor dispatches a HOLD helper.
      
      REPAIR_ROUNDS_DEFAULT = 2
      
      #: Terminal reasons `run_repair_phase` may report. `converged` is the only
      #: success; the rest are law stops the caller owns the response to.
      STOP_REASONS = (
          "converged",
          "diversity_unsatisfied",
          "repair_budget_exhausted",
          "reopened_finding",
          "recurring_finding_class",
          "introduced_regression",
          "new_finding_requires_causal_review",
          "no_acceptance_progress",
          "not_converged",
      )
      
      _STATUS_RANK = {"PASS": 0, "NOT_PROVEN": 1, "FAIL": 2}
      
      
      def _leg_status(leg: Mapping[str, Any]) -> str:
          """Read a validate leg's semantic verdict under either spelling."""
          status = leg.get("status", leg.get("verdict"))
          if status not in _STATUS_RANK:
              raise ValueError("each validate result must report PASS, FAIL, or NOT_PROVEN")
          return str(status)
      
      
      def normalize_round(value: Any) -> dict[str, Any]:
          """Fold one validation round's legs into the facts the law reasons over.
      
          A round is one or more validate results (the fresh validator, plus the
          cross-family validator when the caller selects one). Open findings
          are the UNION of the legs' stable `findings[].id`; the round's status is the
          worst leg's; the digest is the subject every leg judged.
          """
          legs: list[Mapping[str, Any]]
          if isinstance(value, Mapping):
              legs = [value]
          elif isinstance(value, Sequence) and not isinstance(value, (str, bytes)):
              legs = list(value)
          else:
              raise ValueError("a validation round must be a validate result or a list of them")
          if not legs:
              raise ValueError("a validation round must contain at least one validate result")
      
          open_findings: dict[str, dict[str, Any]] = {}
          families: list[str] = []
          evidence_refs: list[dict[str, Any]] = []
          checked: list[str] = []
          not_checked: list[str] = []
          digest: Any = None
          status = "PASS"
          for leg in legs:
              if not isinstance(leg, Mapping):
                  raise ValueError("each validate result must be a mapping")
              leg_status = _leg_status(leg)
              if _STATUS_RANK[leg_status] > _STATUS_RANK[status]:
                  status = leg_status
              if "findings" not in leg:
                  raise ValueError("each validate leg must carry a findings list (empty on PASS)")
              raw_findings = leg["findings"]
              if not isinstance(raw_findings, (list, tuple)):
                  raise ValueError("findings must be a list")
              leg_ids: set[str] = set()
              for finding in raw_findings:
                  if not isinstance(finding, Mapping):
                      raise ValueError("each finding must be a mapping")
                  finding_id = finding.get("id")
                  if not isinstance(finding_id, str) or not finding_id.strip():
                      raise ValueError("each finding must carry a stable nonempty id")
                  if finding_id in leg_ids:
                      raise ValueError(f"finding id {finding_id!r} appears twice in one validate leg")
                  leg_ids.add(finding_id)
                  if "class" in finding and (
                      not isinstance(finding["class"], str) or not finding["class"].strip()
                  ):
                      raise ValueError("finding class must be a nonempty string when present")
                  # Wording can differ, but another leg cannot erase a class used to
                  # detect recurrence or silently replace it with a conflicting one.
                  existing = open_findings.get(finding_id, {})
                  if existing.get("class") and finding.get("class") not in {None, existing["class"]}:
                      raise ValueError(f"finding id {finding_id!r} has conflicting classes")
                  open_findings[finding_id] = {**existing, **finding}
              if leg_status == "PASS" and leg_ids:
                  raise ValueError("a PASS leg cannot carry open findings")
              if leg_status == "FAIL" and not leg_ids:
                  raise ValueError("a FAIL leg must name at least one finding")
              family = leg.get("validator_family")
              if isinstance(family, str) and family and family not in families:
                  families.append(family)
              if "evidence_refs" not in leg:
                  raise ValueError("each validate leg must carry an evidence_refs list (empty if none)")
              raw_evidence = leg["evidence_refs"]
              if not isinstance(raw_evidence, (list, tuple)):
                  raise ValueError("evidence_refs must be a list")
              for ref in raw_evidence:
                  # Evidence is either a bare label (unbound; it can never admit an
                  # unchanged digest) or a binding {ref, subject_digest, resolves}.
                  if isinstance(ref, str):
                      entry: dict[str, Any] = {"ref": ref}
                  elif isinstance(ref, Mapping):
                      if not isinstance(ref.get("ref"), str) or not ref["ref"].strip():
                          raise ValueError("each evidence binding must carry a nonempty ref")
                      entry = dict(ref)
                      # These are decoded facts from existing check receipts, not
                      # additional verdict.v2 fields or a persisted receipt schema.
                      for key in ("resolves", "preexisting", "introduced"):
                          ids = entry.get(key)
                          if ids is not None and not valid_string_list(ids):
                              raise ValueError(f"evidence.{key} must be a list of finding ids")
                      if "subject_digest" in entry and not valid_digest(entry["subject_digest"]):
                          raise ValueError("evidence.subject_digest must be a valid digest")
                  else:
                      raise ValueError("each evidence ref must be a string or a binding mapping")
                  existing_evidence = next((e for e in evidence_refs if e["ref"] == entry["ref"]), None)
                  if existing_evidence is None:
                      evidence_refs.append(entry)
                  elif existing_evidence != entry:
                      raise ValueError(f"evidence ref {entry['ref']!r} has conflicting bindings")
              leg_digest = leg.get("subject_digest", leg.get("subject_manifest_digest"))
              if not valid_digest(leg_digest):
                  raise ValueError("each validate leg must carry a valid subject digest")
              if digest is not None and leg_digest != digest:
                  raise ValueError("validate legs disagree about the subject digest")
              digest = leg_digest
              for key, sink in (("checked", checked), ("not_checked", not_checked)):
                  items = leg.get(key, [])
                  if not valid_string_list(items):
                      raise ValueError(f"{key} must be a list of strings")
                  sink.extend(items)
              # Each required leg must carry its own visible proof surface; a peer's
              # receipts cannot repair a deficient PASS. Exact identities and all
              # criterion proofs remain Validate's upstream contract, not a claim
              # that this pure reference attested or re-executed them.
              if leg_status == "PASS" and (
                  leg.get("not_checked") or not leg.get("checked") or not raw_evidence
              ) and status == "PASS":
                  status = "NOT_PROVEN"
      
          return {
              "status": status,
              "open_findings": list(open_findings.values()),
              "open_ids": set(open_findings),
              "subject_digest": digest,
              "evidence_refs": evidence_refs,
              "families": families,
              "checked": checked,
              "not_checked": not_checked,
          }
      
      
      def law_violation(
          previous: Mapping[str, Any],
          current: Mapping[str, Any],
          closed_ids: set[str],
          closed_classes: set[str] | None = None,
      ) -> str | None:
          """Return the violated convergence-law condition, or None when all hold.
      
          Condition 1 (the caller's `repair_rounds`) is a precondition on admission
          and is checked by `run_repair_phase` before a round is consumed; conditions
          2 and 3 are properties of the round that was produced. The existing receipt
          binding names a gap that a fresh judge actually closed; the function does
          not prove that closure or infer a new finding's cause from prose or counts.
          """
          reopened = current["open_ids"] & closed_ids
          if reopened:
              return "reopened_finding"
          current_classes = {f.get("class") for f in current["open_findings"] if f.get("class")}
          if current_classes & (closed_classes or set()):
              return "recurring_finding_class"
          new_ids = current["open_ids"] - previous["open_ids"]
          introduced = {
              fid
              for evidence in current["evidence_refs"]
              if evidence.get("subject_digest") == current["subject_digest"]
              for fid in evidence.get("introduced", [])
          }
          if introduced & current["open_ids"]:
              return "introduced_regression"
          preexisting = {
              fid
              for evidence in current["evidence_refs"]
              if evidence.get("subject_digest") == previous["subject_digest"]
              for fid in evidence.get("preexisting", [])
          }
          if new_ids - preexisting:
              return "new_finding_requires_causal_review"
          previous_refs = {e["ref"] for e in previous["evidence_refs"]}
          resolved = previous["open_ids"] - current["open_ids"]
          # New digest-bound evidence is required even when the bytes or count moved.
          # It names a gap actually closed this round, not a renamed or still-open
          # finding. New findings remain visible; their count is not a regression
          # diagnosis. Fresh judgment must establish acceptance relevance and cause.
          binding_evidence = [
              e
              for e in current["evidence_refs"]
              if e["ref"] not in previous_refs
              and e.get("subject_digest") == current["subject_digest"]
              and resolved & set(e.get("resolves") or [])
          ]
          if binding_evidence and (
              current["subject_digest"] != previous["subject_digest"]
              or (previous["status"] == "NOT_PROVEN" and current["status"] != "FAIL")
          ):
              return None
          return "no_acceptance_progress"
      
      
      def run_repair_phase(
          validations: Sequence[Any],
          *,
          repair_rounds: int = REPAIR_ROUNDS_DEFAULT,
          risky_surface: bool = False,
          cross_model: bool = False,
          intent_ref: str | None = None,
          acceptance_digest: str | None = None,
          verdict_ref: str | None = None,
          verdict_digest: str | None = None,
      ) -> dict[str, Any]:
          """Walk already-produced validation rounds under the convergence law.
      
          `validations[0]` is the traversal's first fresh validation; every later
          element is a repair round the orchestrator produced after fixing findings.
          ``cross_model`` requires a second family; risk alone does not select it.
          ``risky_surface`` remains an accepted compatibility hint with no effect on
          family selection. This pure reference consumes declared family facts; it
          does not dispatch models or attest fresh context identities.
      
          Returns a mapping with:
      
          - ``report``: the exact nine-key `rpi-report.v1` object. `checked` opens
            with one `repair round N: k open findings` line per round; open findings
            never enter `not_checked`, which keeps its meaning (unverified in-scope
            acceptance).
          - ``open_findings``: the findings still open at the stop, deduplicated by id.
          - ``rounds_used``: repair rounds actually spent (the first validation is
            round 0 and spends none).
          - ``stop_reason``: one of :data:`STOP_REASONS`.
          """
          if not validations:
              raise ValueError("the repair phase needs at least one validation round")
          if not isinstance(repair_rounds, int) or isinstance(repair_rounds, bool) or repair_rounds < 0:
              raise ValueError("repair_rounds must be a non-negative integer")
      
          checked: list[str] = []
          closed_ids: set[str] = set()
          closed_classes: set[str] = set()
          rounds_used = 0
          current = normalize_round(validations[0])
          previous = current
          stop_reason = "not_converged"
          law_stopped = False
      
          for index, raw_candidate in enumerate(validations):
              if index > 0:
                  # Condition 1: the caller's bound, checked before the round is even
                  # normalized, so a round past the bound is never consumed.
                  if rounds_used >= repair_rounds:
                      stop_reason = "repair_budget_exhausted"
                      break
                  candidate = normalize_round(raw_candidate)
                  rounds_used += 1
                  current = candidate
                  checked.append(
                      f"repair round {rounds_used}: {len(current['open_ids'])} open findings"
                  )
                  violation = law_violation(previous, current, closed_ids, closed_classes)
                  if violation is not None:
                      stop_reason = violation
                      law_stopped = True
                      break
                  closed_ids |= previous["open_ids"] - current["open_ids"]
                  # A class closes only when none of its findings remain open. A
                  # later return, even under a new id, warrants causal examination.
                  closed_classes |= (
                      {f.get("class") for f in previous["open_findings"] if f.get("class")}
                      - {f.get("class") for f in current["open_findings"] if f.get("class")}
                  )
              else:
                  checked.append(f"repair round 0: {len(current['open_ids'])} open findings")
      
              converged, reason = _converged(current, cross_model)
              previous = current
              if converged:
                  stop_reason = "converged"
                  break
              if reason is not None:
                  stop_reason = reason
                  break
          else:
              stop_reason = "not_converged"
      
          if stop_reason == "not_converged" and rounds_used >= repair_rounds and current["open_ids"]:
              # Findings remain and the caller's bound is spent: name it as such.
              stop_reason = "repair_budget_exhausted"
      
          status = current["status"]
          if stop_reason == "diversity_unsatisfied" or (law_stopped and status == "PASS"):
              # A PASS produced by a law-violating round cannot certify anything: a
              # PASS over unchanged bytes after a FAIL is a flip, not a proof. A FAIL
              # that also broke the law stays a FAIL; the subject is still wrong.
              status = "NOT_PROVEN"
      
          return {
              "report": report(
                  status,
                  intent_ref=intent_ref,
                  acceptance_digest=acceptance_digest,
                  subject_digest=current["subject_digest"],
                  verdict_ref=verdict_ref,
                  verdict_digest=verdict_digest,
                  checked=checked + current["checked"],
                  not_checked=list(current["not_checked"]),
              ),
              "open_findings": list(current["open_findings"]),
              "rounds_used": rounds_used,
              "stop_reason": stop_reason,
          }
      
      
      def _converged(current: Mapping[str, Any], cross_model: bool) -> tuple[bool, str | None]:
          """Converged ⇔ fresh PASS, plus a cross-family PASS when selected."""
          if current["status"] != "PASS":
              return False, None
          if cross_model and len(current["families"]) < 2:
              # Fresh same-family judgment is valid by default, but cannot satisfy
              # an explicitly selected second family.
              return False, "diversity_unsatisfied"
          return True, None
      
    • validate.sh 1.2 KB
      #!/usr/bin/env bash
      set -euo pipefail
      skill_dir="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
      # ADR-0017's lean amendment removes the native phase lock while preserving
      # exact fresh judgment and real bounds. The pure adapter is checked separately.
      grep -q '^name: rpi$' "$skill_dir/SKILL.md"
      grep -Fq 'dependencies: [plan, implement, validate]' "$skill_dir/SKILL.md"
      grep -Fq 'Own the authorized outcome through finish.' "$skill_dir/SKILL.md"
      grep -Fq 'ordinary known' "$skill_dir/SKILL.md"
      grep -Fq 'Acceptance changes need caller authority.' "$skill_dir/SKILL.md"
      grep -Fq 'at most one bounded' "$skill_dir/SKILL.md"
      grep -Fq 'reviewers remain required.' "$skill_dir/SKILL.md"
      grep -Fq 'author-distinct' "$skill_dir/SKILL.md"
      grep -Fq 'no fixed ten-minute cap' "$skill_dir/SKILL.md"
      grep -Fq 'empty' "$skill_dir/SKILL.md"
      grep -Fq 'When no machine' "$skill_dir/SKILL.md"
      if grep -Eq 'Plan is closed for that intent|dependencies:.*anti-ceremony|plan_packet_digest' "$skill_dir/SKILL.md"; then
        echo 'rpi retains a retired phase lock, mandatory specialist or planning packet' >&2
        exit 1
      fi
      for ref in boundaries bounded-adapter outer-goal; do
        test -s "$skill_dir/references/$ref.md"
      done
      echo 'rpi lean skill contract: PASS'
      
  • tests
    • test_run_once.py 36.8 KB
      from __future__ import annotations
      
      import hashlib
      import importlib.util
      from pathlib import Path
      import tempfile
      import unittest
      
      
      MODULE_PATH = Path(__file__).parents[1] / "scripts" / "run_once.py"
      SPEC = importlib.util.spec_from_file_location("rpi_run_once", MODULE_PATH)
      assert SPEC and SPEC.loader
      MODULE = importlib.util.module_from_spec(SPEC)
      SPEC.loader.exec_module(MODULE)
      
      # The unchanged Validate reference oracle, not a stand-in. The composed-contract test
      # below drives RPI against this module's actual identity functions; that is the
      # only shape that can catch a disagreement between the two skills, which is
      # exactly the defect that hid here (both suites green over a broken contract
      # because the fake Validate borrowed RPI's own digest function).
      VALIDATE_PATH = Path(__file__).parents[2] / "validate" / "tests" / "validate.py"
      VALIDATE_SPEC = importlib.util.spec_from_file_location("ao_validate", VALIDATE_PATH)
      assert VALIDATE_SPEC and VALIDATE_SPEC.loader
      VALIDATE = importlib.util.module_from_spec(VALIDATE_SPEC)
      VALIDATE_SPEC.loader.exec_module(VALIDATE)
      
      # A literal, independently written digest. Fakes must never derive an expected
      # identity by calling the code under test — that is how the original defect
      # stayed invisible.
      INTENT_DIGEST = "c" * 64
      
      
      def validation_round(
          status,
          finding_ids,
          *,
          digest="a" * 64,
          evidence=("acceptance-receipt",),
          family="fresh",
          summaries=None,
          checked=("acceptance",),
          not_checked=(),
      ):
          """One validate leg's result, as the repair phase consumes it (pure data)."""
          summaries = summaries or {}
          return {
              "status": status,
              "findings": [
                  {"id": fid, "summary": summaries.get(fid, f"finding {fid}")}
                  for fid in finding_ids
              ],
              "subject_digest": digest,
              "evidence_refs": list(evidence),
              "validator_family": family,
              "checked": list(checked),
              "not_checked": list(not_checked),
          }
      
      
      def continue_guard(intent):
          return {
              "decision": "CONTINUE",
              "reason": "The frozen outcome still requires implementation proof.",
              "frozen_outcome": str(intent),
              "parked_process_work": [],
              "remaining_proof": ["implementation", "fresh validation"],
              "stop_condition": "Stop after one fresh validation result.",
          }
      
      
      class RunOnceTests(unittest.TestCase):
          def phases(self, verdict: str = "PASS"):
              calls: list[str] = []
      
              def plan(intent):
                  calls.append("plan")
                  return {
                      "intent_ref": "bead:agentops-test",
                      "intent": intent,
                      "acceptance": ["works"],
                      "acceptance_digest": INTENT_DIGEST,
                  }
      
              def implement(_plan):
                  calls.append("implement")
                  return {"subject_manifest_digest": "a" * 64, "checks": ["focused"]}
      
              def validate(_plan, _candidate):
                  calls.append("validate")
                  return {
                      "verdict": verdict,
                      "acceptance_digest": INTENT_DIGEST,
                      "subject_manifest_digest": "a" * 64,
                      "author_context_id": "author-ctx",
                      "validator_context_id": "validator-ctx",
                      "freshness_attestation": {
                          "source": "runtime",
                          "attester_identity": "runtime:rpi-test",
                      },
                      "verdict_digest": "b" * 64,
                      "verdict_ref": "/tmp/verdict.json",
                      "checked": ["acceptance"],
                      "not_checked": [],
                  }
      
              return calls, plan, implement, validate
      
          def test_anti_ceremony_guard_runs_once_before_plan(self):
              calls, plan, implement, validate = self.phases()
      
              def anti_ceremony(intent):
                  calls.append("anti-ceremony")
                  return {
                      "decision": "CONTINUE",
                      "reason": "The frozen outcome still requires implementation proof.",
                      "frozen_outcome": intent,
                      "parked_process_work": [],
                      "remaining_proof": ["implementation", "fresh validation"],
                      "stop_condition": "Stop after one fresh validation result.",
                  }
      
              result = MODULE.invoke_once(
                  "intent",
                  anti_ceremony,
                  plan,
                  implement,
                  validate,
              )
      
              self.assertEqual(
                  calls,
                  ["anti-ceremony", "plan", "implement", "validate"],
              )
              self.assertEqual(result["status"], "PASS")
      
          def test_anti_ceremony_stop_dispatches_no_core_phase(self):
              calls: list[str] = []
              reason = "The proposed traversal would create only process artifacts."
      
              def anti_ceremony(_intent):
                  calls.append("anti-ceremony")
                  return {
                      "decision": "STOP",
                      "reason": reason,
                      "frozen_outcome": "Ship the already-proved caller outcome",
                      "parked_process_work": ["another plan", "another audit"],
                      "remaining_proof": [],
                      "stop_condition": "Stop before Plan.",
                  }
      
              result = MODULE.invoke_once(
                  "intent",
                  anti_ceremony,
                  lambda _intent: calls.append("plan"),
                  lambda _plan: calls.append("implement"),
                  lambda _plan, _candidate: calls.append("validate"),
              )
      
              self.assertEqual(calls, ["anti-ceremony"])
              self.assertEqual(result["status"], "NOT_PLANNED")
              self.assertEqual(
                  result["checked"],
                  [f"anti-ceremony guard: STOP — {reason}"],
              )
              self.assertEqual(result["not_checked"], ["plan", "implement", "validate"])
      
          def test_each_phase_runs_once_and_pass_reports(self):
              calls, plan, implement, validate = self.phases()
              result = MODULE.invoke_once("intent", continue_guard, plan, implement, validate)
              self.assertEqual(calls, ["plan", "implement", "validate"])
              self.assertEqual(result["status"], "PASS")
              self.assertEqual(result["intent_ref"], "bead:agentops-test")
              self.assertEqual(result["acceptance_digest"], INTENT_DIGEST)
              self.assertNotIn("next_action", result)
      
          def test_fail_from_one_experiment_feeds_the_repair_phase(self):
              """Replaces the old stop-on-FAIL test (ADR-0017).
      
              One experiment still dispatches Plan and Implement exactly once, and the
              FAIL it produces is no longer terminal by itself: it is the first round
              handed to the bounded repair phase, which owns the stop decision.
              """
              calls, plan, implement, validate = self.phases("FAIL")
              result = MODULE.invoke_once("intent", continue_guard, plan, implement, validate)
              self.assertEqual(calls, ["plan", "implement", "validate"])
              self.assertEqual(result["status"], "FAIL")
      
              outcome = MODULE.run_repair_phase(
                  [
                      validation_round("FAIL", ["f1"], digest="a" * 64),
                      validation_round("PASS", [], digest="d" * 64, evidence=(
                          {"ref": "fixed-f1", "subject_digest": "d" * 64, "resolves": ["f1"]},
                      )),
                  ],
                  repair_rounds=2,
                  intent_ref=result["intent_ref"],
                  acceptance_digest=result["acceptance_digest"],
              )
              self.assertEqual(outcome["stop_reason"], "converged")
              self.assertEqual(outcome["report"]["status"], "PASS")
              self.assertEqual(outcome["rounds_used"], 1)
      
          def test_fresh_validation_does_not_require_persisted_verdict(self):
              calls, plan, implement, validate = self.phases()
      
              def inline_result(resolved, subject):
                  result = validate(resolved, subject)
                  result.pop("verdict_digest")
                  result.pop("verdict_ref")
                  return result
      
              result = MODULE.invoke_once("intent", continue_guard, plan, implement, inline_result)
      
              self.assertEqual(calls, ["plan", "implement", "validate"])
              self.assertEqual(result["status"], "PASS")
              self.assertEqual(result["subject_manifest_digest"], "a" * 64)
              self.assertIsNone(result["verdict_ref"])
              self.assertIsNone(result["verdict_digest"])
      
          def test_fresh_validation_requires_distinct_contexts_and_attestation(self):
              _calls, plan, implement, validate = self.phases()
      
              for field, value in (
                  ("validator_context_id", "author-ctx"),
                  ("freshness_attestation", None),
              ):
                  with self.subTest(field=field):
                      def invalid(resolved, subject, field=field, value=value):
                          result = validate(resolved, subject)
                          result[field] = value
                          return result
      
                      with self.assertRaisesRegex(ValueError, "distinct context identities"):
                          MODULE.invoke_once("intent", continue_guard, plan, implement, invalid)
      
          def test_missing_plan_stops_before_implement(self):
              calls: list[str] = []
              result = MODULE.invoke_once(
                  "intent",
                  continue_guard,
                  lambda _intent: None,
                  lambda _plan: calls.append("implement"),
                  lambda _plan, _candidate: calls.append("validate"),
              )
              self.assertEqual(calls, [])
              self.assertEqual(result["status"], "NOT_PLANNED")
      
          def test_missing_candidate_stops_before_validate(self):
              calls: list[str] = []
              result = MODULE.invoke_once(
                  "intent",
                  continue_guard,
                  lambda _intent: {
                      "intent_ref": "caller",
                      "acceptance": ["works"],
                      "acceptance_digest": INTENT_DIGEST,
                  },
                  lambda _plan: None,
                  lambda _plan, _candidate: calls.append("validate"),
              )
              self.assertEqual(calls, [])
              self.assertEqual(result["status"], "NOT_BUILT")
      
          def test_validate_cannot_report_a_different_intent(self):
              calls, plan, implement, validate = self.phases()
      
              def mismatched(resolved, subject):
                  result = validate(resolved, subject)
                  result["acceptance_digest"] = "f" * 64
                  return result
      
              with self.assertRaisesRegex(ValueError, "resolved intent digest"):
                  MODULE.invoke_once("intent", continue_guard, plan, implement, mismatched)
      
          def test_plan_without_a_declared_digest_is_a_contract_error(self):
              _calls, _plan, implement, validate = self.phases()
      
              def undeclared(_intent):
                  return {"intent_ref": "caller", "acceptance": ["works"]}
      
              with self.assertRaisesRegex(ValueError, "acceptance_digest"):
                  MODULE.invoke_once("intent", continue_guard, undeclared, implement, validate)
      
          def test_plan_digest_must_be_a_sha256(self):
              _calls, _plan, implement, validate = self.phases()
      
              for bogus in ("", "not-a-digest", "C" * 64, "a" * 63, 12345):
                  with self.subTest(digest=bogus):
                      def undeclared(_intent, value=bogus):
                          return {
                              "intent_ref": "caller",
                              "acceptance": ["works"],
                              "acceptance_digest": value,
                          }
      
                      with self.assertRaisesRegex(ValueError, "acceptance_digest"):
                          MODULE.invoke_once("intent", continue_guard, undeclared, implement, validate)
      
      
      class RepairPhaseTests(unittest.TestCase):
          """The bounded repair phase and its convergence law (ADR-0017).
      
          RPI is no longer single-pass: a `FAIL` or `NOT_PROVEN` with findings may be
          repaired and re-validated while acceptance progress and the bound hold. The loop is
          modelled here as pure data — a sequence of already-produced validate rounds
          — so the stop semantics are executable without Git, `ao`, or a tracker.
          """
      
          def repair(self, rounds, **kwargs):
              kwargs.setdefault("intent_ref", "bead:agentops-test")
              kwargs.setdefault("acceptance_digest", INTENT_DIGEST)
              return MODULE.run_repair_phase(rounds, **kwargs)
      
          def test_repair_rounds_zero_with_findings_is_budget_exhausted(self):
              outcome = self.repair([validation_round("FAIL", ["f1"])], repair_rounds=0)
              self.assertEqual(outcome["stop_reason"], "repair_budget_exhausted")
              self.assertEqual(outcome["rounds_used"], 0)
              self.assertEqual(outcome["report"]["status"], "FAIL")
      
          def test_a_pass_over_unchanged_bytes_after_a_fail_is_a_flip_not_a_proof(self):
              outcome = self.repair(
                  [
                      validation_round("FAIL", ["f1"], digest="a" * 64),
                      validation_round("PASS", [], digest="a" * 64),
                  ]
              )
              self.assertEqual(outcome["stop_reason"], "no_acceptance_progress")
              self.assertEqual(outcome["report"]["status"], "NOT_PROVEN")
      
          def test_new_evidence_must_resolve_a_prior_finding_to_admit_an_unchanged_digest(self):
              outcome = self.repair(
                  [
                      validation_round("NOT_PROVEN", ["gap"], digest="a" * 64),
                      validation_round(
                          "NOT_PROVEN", ["gap"], digest="a" * 64,
                          evidence=({"ref": "receipt-2", "subject_digest": "a" * 64, "resolves": ["gap"]},),
                      ),
                  ]
              )
              # "resolves" claims gap, but gap is still open: nothing was resolved.
              self.assertEqual(outcome["stop_reason"], "no_acceptance_progress")
      
          def test_a_bare_new_evidence_label_does_not_admit_an_unchanged_digest(self):
              outcome = self.repair(
                  [
                      validation_round("NOT_PROVEN", ["gap", "other"], digest="a" * 64),
                      validation_round("NOT_PROVEN", ["other"], digest="a" * 64, evidence=("receipt-2",)),
                  ]
              )
              self.assertEqual(outcome["stop_reason"], "no_acceptance_progress")
      
          def test_evidence_bound_to_another_digest_does_not_admit(self):
              outcome = self.repair(
                  [
                      validation_round("NOT_PROVEN", ["gap", "other"], digest="a" * 64),
                      validation_round(
                          "NOT_PROVEN", ["other"], digest="a" * 64,
                          evidence=({"ref": "receipt-2", "subject_digest": "b" * 64, "resolves": ["gap"]},),
                      ),
                  ]
              )
              self.assertEqual(outcome["stop_reason"], "no_acceptance_progress")
      
          def test_missing_findings_or_evidence_keys_are_rejected(self):
              for key in ("findings", "evidence_refs"):
                  with self.subTest(missing=key):
                      bad = validation_round("PASS", [])
                      del bad[key]
                      with self.assertRaisesRegex(ValueError, key):
                          self.repair([bad])
              bad = validation_round("PASS", [])
              bad["checked"] = "acceptance"
              with self.assertRaisesRegex(ValueError, "checked must be a list"):
                  self.repair([bad])
      
          def test_scalar_evidence_refs_are_rejected(self):
              bad = validation_round("FAIL", ["f1"])
              bad["evidence_refs"] = "receipt-1"
              with self.assertRaisesRegex(ValueError, "evidence_refs must be a list"):
                  self.repair([bad])
      
          def test_new_evidence_does_not_admit_a_current_fail_over_unchanged_bytes(self):
              outcome = self.repair(
                  [
                      validation_round("NOT_PROVEN", ["gap", "bug"], digest="a" * 64),
                      validation_round("FAIL", ["bug"], digest="a" * 64, evidence=("receipt-2",)),
                  ]
              )
              self.assertEqual(outcome["stop_reason"], "no_acceptance_progress")
              self.assertEqual(outcome["report"]["status"], "FAIL")
      
          def test_malformed_rounds_are_rejected_not_swallowed(self):
              cases = {
                  "pass with findings": validation_round("PASS", ["f1"]),
                  "fail without findings": validation_round("FAIL", []),
                  "missing digest": validation_round("FAIL", ["f1"], digest=None),
              }
              for name, bad in cases.items():
                  with self.subTest(case=name):
                      with self.assertRaises(ValueError):
                          self.repair([bad])
              duplicate = validation_round("FAIL", ["f1"])
              duplicate["findings"].append({"id": "f1", "summary": "again"})
              with self.assertRaisesRegex(ValueError, "twice"):
                  self.repair([duplicate])
      
          def test_rounds_past_the_bound_are_never_normalized(self):
              outcome = self.repair(
                  [
                      validation_round("FAIL", ["f1", "f2"], digest="a" * 64),
                      validation_round("FAIL", ["f1"], digest="b" * 64, evidence=(
                          {"ref": "fixed-f2", "subject_digest": "b" * 64, "resolves": ["f2"]},
                      )),
                      {"status": "garbage-that-would-raise"},
                  ],
                  repair_rounds=1,
              )
              self.assertEqual(outcome["stop_reason"], "repair_budget_exhausted")
      
          def test_a_first_round_pass_converges_without_spending_a_repair_round(self):
              outcome = self.repair([validation_round("PASS", [])])
              self.assertEqual(outcome["stop_reason"], "converged")
              self.assertEqual(outcome["report"]["status"], "PASS")
              self.assertEqual(outcome["rounds_used"], 0)
              self.assertEqual(outcome["open_findings"], [])
              self.assertEqual(
                  outcome["report"]["checked"][0], "repair round 0: 0 open findings"
              )
      
          def test_repair_stops_at_the_declared_repair_rounds_budget(self):
              outcome = self.repair(
                  [
                      validation_round("FAIL", ["f1", "f2"], digest="a" * 64),
                      validation_round("FAIL", ["f1"], digest="b" * 64, evidence=(
                          {"ref": "fixed-f2", "subject_digest": "b" * 64, "resolves": ["f2"]},
                      )),
                      validation_round("FAIL", ["f1"], digest="c" * 64),
                  ],
                  repair_rounds=1,
              )
              self.assertEqual(outcome["stop_reason"], "repair_budget_exhausted")
              self.assertEqual(outcome["rounds_used"], 1)
              self.assertEqual(outcome["report"]["status"], "FAIL")
              self.assertEqual(
                  outcome["report"]["checked"][:2],
                  ["repair round 0: 2 open findings", "repair round 1: 1 open findings"],
              )
      
          def test_new_finding_with_unknown_cause_requires_causal_review(self):
              outcome = self.repair(
                  [
                      validation_round("FAIL", ["f1"], digest="a" * 64),
                      validation_round("FAIL", ["f1", "f2"], digest="b" * 64),
                  ]
              )
              self.assertEqual(outcome["stop_reason"], "new_finding_requires_causal_review")
              self.assertEqual(outcome["report"]["status"], "FAIL")
              self.assertEqual(outcome["rounds_used"], 1)
              self.assertEqual(
                  sorted(f["id"] for f in outcome["open_findings"]), ["f1", "f2"]
              )
      
          def test_repair_stops_when_a_closed_finding_reopens(self):
              outcome = self.repair(
                  [
                      validation_round("FAIL", ["f1", "f2"], digest="a" * 64),
                      validation_round("FAIL", ["f1"], digest="b" * 64, evidence=(
                          {"ref": "fixed-f2", "subject_digest": "b" * 64, "resolves": ["f2"]},
                      )),
                      validation_round("FAIL", ["f2"], digest="c" * 64),
                  ]
              )
              self.assertEqual(outcome["stop_reason"], "reopened_finding")
              self.assertEqual(outcome["rounds_used"], 2)
              self.assertEqual(outcome["report"]["status"], "FAIL")
      
          def test_digest_and_count_movement_without_bound_proof_are_not_progress(self):
              for remaining in (["f1", "f2"], ["f1"], []):
                  with self.subTest(remaining=remaining):
                      outcome = self.repair([
                          validation_round("FAIL", ["f1", "f2"], digest="a" * 64),
                          validation_round("FAIL" if remaining else "PASS", remaining,
                                           digest="b" * 64, evidence=("another-check-label",)),
                      ])
                      self.assertEqual(outcome["stop_reason"], "no_acceptance_progress")
                      self.assertNotEqual(outcome["report"]["status"], "PASS")
      
          def test_discovered_preexisting_defects_may_grow_count_with_real_progress(self):
              outcome = self.repair([
                  validation_round("FAIL", ["fixed"], digest="a" * 64),
                  validation_round("FAIL", ["discovered-1", "discovered-2"], digest="b" * 64,
                                   evidence=(
                      {"ref": "fixed-check", "subject_digest": "b" * 64, "resolves": ["fixed"]},
                      {"ref": "reproduced-on-prior", "subject_digest": "a" * 64,
                       "preexisting": ["discovered-1", "discovered-2"]},
                  )),
              ])
              self.assertEqual(outcome["stop_reason"], "not_converged")
              self.assertEqual(outcome["report"]["status"], "FAIL")
              self.assertEqual(len(outcome["open_findings"]), 2)
              self.assertEqual(outcome["rounds_used"], 1)
      
          def test_new_finding_cannot_hide_behind_another_resolved_gap(self):
              for proof in ((), ({"ref": "wrong-baseline", "subject_digest": "c" * 64,
                                 "preexisting": ["new"]},)):
                  with self.subTest(proof=proof):
                      outcome = self.repair([
                          validation_round("FAIL", ["fixed"], digest="a" * 64),
                          validation_round("FAIL", ["new"], digest="b" * 64, evidence=(
                              {"ref": "fixed-check", "subject_digest": "b" * 64, "resolves": ["fixed"]},
                          ) + proof),
                      ])
                      self.assertEqual(outcome["stop_reason"], "new_finding_requires_causal_review")
      
          def test_introduced_regression_stops_even_if_another_gap_closed(self):
              outcome = self.repair([
                  validation_round("FAIL", ["fixed"], digest="a" * 64),
                  validation_round("FAIL", ["regression"], digest="b" * 64, evidence=(
                      {"ref": "fixed-check", "subject_digest": "b" * 64, "resolves": ["fixed"]},
                      {"ref": "before-after-check", "subject_digest": "b" * 64,
                       "introduced": ["regression"]},
                  )),
              ])
              self.assertEqual(outcome["stop_reason"], "introduced_regression")
              self.assertEqual(outcome["report"]["status"], "FAIL")
      
          def test_recurring_class_requires_causal_review_without_design_diagnosis(self):
              first = validation_round("FAIL", ["f1", "f2"], digest="a" * 64)
              first["findings"][0]["class"] = "deadline-bypass"
              recurrence = validation_round("FAIL", ["new-id"], digest="c" * 64)
              recurrence["findings"][0]["class"] = "deadline-bypass"
              outcome = self.repair([
                  first,
                  validation_round("FAIL", ["f2"], digest="b" * 64, evidence=(
                      {"ref": "fixed-f1", "subject_digest": "b" * 64, "resolves": ["f1"]},
                  )),
                  recurrence,
              ])
              self.assertEqual(outcome["stop_reason"], "recurring_finding_class")
              self.assertEqual(outcome["rounds_used"], 2)
              self.assertNotIn("design", str(outcome))
      
          def test_old_receipt_does_not_prove_new_acceptance_progress(self):
              receipt = {"ref": "old-check", "subject_digest": "b" * 64, "resolves": ["f2"]}
              outcome = self.repair([
                  validation_round("FAIL", ["f1", "f2"], digest="a" * 64, evidence=(receipt,)),
                  validation_round("FAIL", ["f1"], digest="b" * 64, evidence=(receipt,)),
              ])
              self.assertEqual(outcome["stop_reason"], "no_acceptance_progress")
      
          def test_peer_leg_cannot_silently_override_a_receipts_causal_binding(self):
              with self.assertRaisesRegex(ValueError, "conflicting bindings"):
                  self.repair([[
                      validation_round("FAIL", ["f1"], evidence=(
                          {"ref": "comparison", "subject_digest": "a" * 64, "preexisting": ["f1"]},
                      )),
                      validation_round("FAIL", ["f1"], family="other", evidence=(
                          {"ref": "comparison", "subject_digest": "a" * 64, "introduced": ["f1"]},
                      )),
                  ]])
      
          def test_peer_leg_cannot_silently_replace_a_recurrence_class(self):
              first = validation_round("FAIL", ["f1"])
              first["findings"][0]["class"] = "deadline-bypass"
              peer = validation_round("FAIL", ["f1"], family="other")
              self.assertEqual(MODULE.normalize_round([first, peer])["open_findings"][0]["class"],
                               "deadline-bypass")
              peer["findings"][0]["class"] = "cosmetic"
              with self.assertRaisesRegex(ValueError, "conflicting classes"):
                  self.repair([[first, peer]])
      
          def test_pass_cannot_converge_with_unverified_acceptance_or_missing_proof(self):
              cases = (
                  validation_round("PASS", [], not_checked=("required-cancellation-case",)),
                  validation_round("PASS", [], checked=()),
                  validation_round("PASS", [], evidence=()),
              )
              for raw in cases:
                  with self.subTest(raw=raw):
                      for round_value in (raw, [raw, validation_round("PASS", [], family="other")]):
                          outcome = self.repair([round_value])
                          self.assertEqual(outcome["report"]["status"], "NOT_PROVEN")
                          self.assertNotEqual(outcome["stop_reason"], "converged")
                          self.assertEqual(outcome["report"]["not_checked"], raw["not_checked"])
      
          def test_repair_stops_when_the_digest_is_unchanged_and_no_new_evidence(self):
              outcome = self.repair(
                  [
                      validation_round("FAIL", ["f1"], digest="a" * 64, evidence=["r1"]),
                      validation_round("FAIL", ["f1"], digest="a" * 64, evidence=["r1"]),
                  ]
              )
              self.assertEqual(outcome["stop_reason"], "no_acceptance_progress")
              self.assertEqual(outcome["rounds_used"], 1)
      
          def test_not_proven_is_resolved_by_new_evidence_with_an_unchanged_digest(self):
              outcome = self.repair(
                  [
                      validation_round(
                          "NOT_PROVEN", ["gap1"], digest="a" * 64, evidence=["r1"]
                      ),
                      validation_round(
                          "PASS", [], digest="a" * 64,
                          evidence=["r1", {"ref": "r2", "subject_digest": "a" * 64, "resolves": ["gap1"]}],
                      ),
                  ]
              )
              self.assertEqual(outcome["stop_reason"], "converged")
              self.assertEqual(outcome["report"]["status"], "PASS")
              self.assertEqual(outcome["rounds_used"], 1)
              self.assertEqual(outcome["report"]["subject_manifest_digest"], "a" * 64)
      
          def test_new_evidence_does_not_rescue_a_fail_round(self):
              """Evidence alone cannot repair an unchanged subject already judged FAIL.
      
              A FAIL means the subject is wrong; both a changed subject and proven
              acceptance progress are needed, not an extra unbound evidence label.
              """
              outcome = self.repair(
                  [
                      validation_round("FAIL", ["f1"], digest="a" * 64, evidence=["r1"]),
                      validation_round("FAIL", ["f1"], digest="a" * 64, evidence=["r1", "r2"]),
                  ]
              )
              self.assertEqual(outcome["stop_reason"], "no_acceptance_progress")
      
          def test_a_reworded_summary_with_the_same_id_is_the_same_finding(self):
              outcome = self.repair(
                  [
                      validation_round(
                          "FAIL", ["f1"], digest="a" * 64, summaries={"f1": "gate fails"}
                      ),
                      validation_round(
                          "FAIL",
                          ["f1"],
                          digest="b" * 64,
                          summaries={"f1": "the deterministic gate still rejects the tree"},
                      ),
                  ]
              )
              self.assertEqual(outcome["rounds_used"], 1)
              self.assertEqual(outcome["stop_reason"], "no_acceptance_progress")
              self.assertEqual([f["id"] for f in outcome["open_findings"]], ["f1"])
              self.assertEqual(
                  outcome["open_findings"][0]["summary"],
                  "the deterministic gate still rejects the tree",
              )
      
          def test_open_findings_are_the_union_of_fresh_and_cross_family_ids(self):
              outcome = self.repair(
                  [
                      validation_round("FAIL", ["f1", "f2", "f3"], digest="a" * 64),
                      [
                          validation_round("FAIL", ["f1"], digest="b" * 64, family="fresh", evidence=(
                              {"ref": "fixed-f3", "subject_digest": "b" * 64, "resolves": ["f3"]},
                          )),
                          validation_round("FAIL", ["f2"], digest="b" * 64, family="codex"),
                      ],
                  ]
              )
              self.assertEqual(outcome["rounds_used"], 1)
              self.assertEqual(sorted(f["id"] for f in outcome["open_findings"]), ["f1", "f2"])
              self.assertEqual(
                  outcome["report"]["checked"][1], "repair round 1: 2 open findings"
              )
      
          def test_a_generated_only_change_needs_proven_acceptance_progress(self):
              outcome = self.repair(
                  [
                      validation_round("FAIL", ["f1"], digest="a" * 64),
                      validation_round("PASS", [], digest="b" * 64, evidence=(
                          {"ref": "projection-parity", "subject_digest": "b" * 64, "resolves": ["f1"]},
                      )),
                  ]
              )
              self.assertEqual(outcome["stop_reason"], "converged")
              self.assertEqual(outcome["rounds_used"], 1)
              self.assertEqual(
                  outcome["report"]["checked"][:2],
                  ["repair round 0: 1 open findings", "repair round 1: 0 open findings"],
              )
      
          def test_open_findings_never_land_in_not_checked(self):
              outcome = self.repair(
                  [
                      validation_round(
                          "FAIL",
                          ["f1"],
                          digest="a" * 64,
                          not_checked=["edge case acceptance"],
                      )
                  ],
                  repair_rounds=0,
              )
              self.assertEqual(outcome["report"]["not_checked"], ["edge case acceptance"])
              self.assertEqual([f["id"] for f in outcome["open_findings"]], ["f1"])
              self.assertNotIn("finding f1", outcome["report"]["not_checked"])
      
          def test_a_risky_surface_uses_one_fresh_family_by_default(self):
              for family in ("codex", "claude"):
                  with self.subTest(family=family):
                      single = self.repair(
                          [validation_round("PASS", [], family=family)], risky_surface=True
                      )
                      self.assertEqual(single["stop_reason"], "converged")
                      self.assertEqual(single["report"]["status"], "PASS")
      
          def test_selected_cross_model_review_needs_a_second_family(self):
              single = self.repair(
                  [validation_round("PASS", [], family="claude")], cross_model=True
              )
              self.assertEqual(single["stop_reason"], "diversity_unsatisfied")
              self.assertEqual(single["report"]["status"], "NOT_PROVEN")
      
              same_family = self.repair(
                  [[validation_round("PASS", [], family="claude"),
                    validation_round("PASS", [], family="claude")]],
                  cross_model=True,
              )
              self.assertEqual(same_family["stop_reason"], "diversity_unsatisfied")
              self.assertEqual(same_family["report"]["status"], "NOT_PROVEN")
      
              crossed = self.repair(
                  [
                      [
                          validation_round("PASS", [], family="fresh"),
                          validation_round("PASS", [], family="codex"),
                      ]
                  ],
                  cross_model=True,
              )
              self.assertEqual(crossed["stop_reason"], "converged")
              self.assertEqual(crossed["report"]["status"], "PASS")
      
          def test_selected_cross_model_disagreement_keeps_the_failure(self):
              split = self.repair(
                  [[validation_round("PASS", [], family="claude"),
                    validation_round("FAIL", ["f1"], family="codex")]],
                  cross_model=True,
              )
              self.assertEqual(split["report"]["status"], "FAIL")
              self.assertEqual([f["id"] for f in split["open_findings"]], ["f1"])
      
          def test_the_report_keeps_the_nine_key_rpi_report_shape(self):
              outcome = self.repair([validation_round("PASS", [])])
              self.assertEqual(
                  sorted(outcome["report"]),
                  sorted(
                      [
                          "schema_version",
                          "status",
                          "intent_ref",
                          "acceptance_digest",
                          "subject_manifest_digest",
                          "verdict_ref",
                          "verdict_digest",
                          "checked",
                          "not_checked",
                      ]
                  ),
              )
              self.assertEqual(outcome["report"]["schema_version"], "rpi-report.v1")
      
          def test_the_repair_phase_needs_at_least_one_validation_round(self):
              with self.assertRaisesRegex(ValueError, "at least one validation round"):
                  self.repair([])
      
      
      class ComposedIdentityContractTests(unittest.TestCase):
          """RPI against the REAL Validate identity functions.
      
          This is the test the defect needed. RPI used to digest a canonical-JSON
          re-serialization of the parsed intent mapping while Validate digested the raw
          intent bytes, and RPI hard-compared the two. Nothing caught it because RPI's
          own suite mocked Validate with RPI's digest function, so the mock agreed with
          the code under test by construction. Here the digest crosses the skill
          boundary in both directions with no shared helper.
          """
      
          # Deliberately NOT canonical JSON: real intent sources have indentation,
          # trailing newlines, and key order. A canonical-JSON digest of the parsed
          # mapping differs from sha256(these bytes), so this payload discriminates
          # between the two implementations instead of accidentally agreeing.
          INTENT_BYTES = b'{\n  "acceptance": ["works"],\n  "intent_ref": "bead:agentops-test"\n}\n'
      
          def test_rpi_carries_the_digest_validate_derives_from_the_snapshot_bytes(self):
              with tempfile.TemporaryDirectory() as tmp:
                  intent_dir = Path(tmp) / "intents"
      
                  def plan(_intent):
                      # Plan resolves the intent and snapshots the EXACT bytes through
                      # Validate's own store, which is what defines the identity.
                      path, _existed = VALIDATE.snapshot_intent(self.INTENT_BYTES, intent_dir)
                      return {
                          "intent_ref": str(path),
                          "acceptance": ["works"],
                          "acceptance_digest": hashlib.sha256(self.INTENT_BYTES).hexdigest(),
                      }
      
                  def implement(_plan):
                      return {"subject_manifest_digest": "a" * 64}
      
                  def validate(resolved, _candidate):
                      # Validate re-reads the snapshot from disk and re-derives the
                      # digest through its own runtime-fact binder — no value is passed
                      # through from Plan, so agreement is earned, not assumed.
                      replayed = Path(resolved["intent_ref"]).read_bytes()
                      bound = VALIDATE.bind_runtime_facts(
                          {"verdict": "PASS"},
                          replayed,
                          None,
                          None,
                          None,
                          None,
                          None,
                          None,
                      )
                      return {
                          "verdict": "PASS",
                          "acceptance_digest": bound["acceptance_digest"],
                          "subject_manifest_digest": "a" * 64,
                          "author_context_id": "author-ctx",
                          "validator_context_id": "validator-ctx",
                          "freshness_attestation": {
                              "source": "runtime",
                              "attester_identity": "runtime:composed-test",
                          },
                          "verdict_digest": "b" * 64,
                          "verdict_ref": str(Path(tmp) / "verdict.json"),
                          "checked": ["acceptance"],
                          "not_checked": [],
                      }
      
                  result = MODULE.invoke_once(
                      self.INTENT_BYTES,
                      continue_guard,
                      plan,
                      implement,
                      validate,
                  )
      
              self.assertEqual(result["status"], "PASS")
              self.assertEqual(
                  result["acceptance_digest"],
                  hashlib.sha256(self.INTENT_BYTES).hexdigest(),
              )
      
          def test_the_canonical_json_digest_is_not_the_intent_identity(self):
              """Pins the two digests apart so the defect cannot silently return.
      
              If someone reintroduces a canonical-JSON digest of the parsed mapping as
              the acceptance identity, this fails: the byte digest and the value digest
              are different numbers for the same intent.
              """
              import json
      
              byte_digest = hashlib.sha256(self.INTENT_BYTES).hexdigest()
              value_digest = VALIDATE.digest_value(json.loads(self.INTENT_BYTES))
              self.assertNotEqual(byte_digest, value_digest)
      
              # And the collision the byte digest forbids: two distinct sources that
              # parse to the same mapping must NOT share an acceptance identity.
              reordered = b'{"intent_ref": "bead:agentops-test", "acceptance": ["works"]}'
              self.assertEqual(json.loads(reordered), json.loads(self.INTENT_BYTES))
              self.assertNotEqual(byte_digest, hashlib.sha256(reordered).hexdigest())
              self.assertEqual(value_digest, VALIDATE.digest_value(json.loads(reordered)))
      
      
      if __name__ == "__main__":
          unittest.main()
      
  • SKILL.md 6.7 KB
    ---
    name: rpi
    description: 'Apply the outcome-to-judgment charter. Use when: the caller explicitly selects RPI; ordinary coding, delegation and native goals do not require this workflow.'
    practices:
    - bdd-gherkin
    - tdd
    - design-by-contract
    hexagonal_role: domain
    consumes:
    - plan
    - implement
    - validate
    produces:
    - rpi-report.v1
    context_rel:
    - kind: customer-of
      with: plan
    - kind: customer-of
      with: implement
    - kind: customer-of
      with: validate
    skill_api_version: 1
    user-invocable: true
    disable-model-invocation: true
    metadata:
      graph_root: true
      tier: meta
      dependencies: [plan, implement, validate]
      capabilities: [own_authorized_outcome, report]
      effects: [dispatch_core_phases]
      canonical_status: canonical
      disposition: keep_strategy
    output_contract: 'concise human-readable result; optional rpi-report.v1 when a caller or declared consumer requests machine-readable evidence'
    ---
    
    # RPI
    
    Own the authorized outcome through finish. Use the native coding agent and
    shell. BD or the caller's tracker owns work and handoffs; Git owns content and
    delivery. AgentOps supplies a small charter and fresh judgment, not a scheduler.
    
    ## Operating charter
    
    1. Use the existing accepted outcome, scope and real bounds. A clear change
       needs no Plan, Recall or Learn worksheet. Resolve uncertainty only when it
       could change the implementation or acceptance decision.
    2. Take the smallest acceptance-advancing action. [Plan](../plan/SKILL.md)
       shapes missing intent or revises a disproved approach. Once an implementer
       can act and a validator can judge, implement; do not keep improving the plan.
       Approach revisions preserve acceptance and authorized scope.
       Acceptance changes need caller authority.
    3. [Implement](../implement/SKILL.md) and repair ordinary known defects directly.
       A known test failure needs a fix and a discriminating check, not another
       planning phase, council or helper.
    4. Use focused checks during edits and complete required integration checks
       before final judgment. Reuse valid exact-input receipts; rerun affected
       checks after changes. Reserve capacity for integration, review and repair.
       Keep the final subject unchanged while it is being judged.
    5. Obtain [Validate](../validate/SKILL.md) from one fresh author-distinct context
       in the author's model family unless the caller selects required additional
       legs. There is no fixed ten-minute cap; explicitly required
       reviewers remain required. Risk deepens evidence, not reviewer multiplication. Repair
       actionable findings within authority and remaining bounds, then revalidate
       the changed exact subject.
    6. Stop at completed acceptance, cancellation, refusal, a spent real bound or
       an unresolved causal stall after the help below. Adjacent improvements are
       not permission to expand the goal. Report them briefly only when useful;
       do not turn them into another work batch.
    
    ## Context and handoffs
    
    Load required contracts once per context, then read only what the next decision
    needs. A reference link is available context, not a reading list. Search before
    opening large files; expand only for consequential uncertainty. Keep successful
    output compact at the tool boundary; retain full logs for inspection. Reuse the
    worker's component-check list and current receipts instead of rediscovering them.
    Use native completion watches or bounded waits for ongoing checks and helpers.
    At completion, verify the expected subject and required results; a quiet or
    partial status is not success. Inspect further for a failure, suspected stall or
    decision need. Keep required user updates concise rather than narrating each poll.
    
    When delegation is authorized and useful, select the runtime's task-only
    dispatch option for independent work; a short prompt in a full-history fork
    still carries full history. Supply accepted intent/scope, exact subject,
    relevant evidence, remaining bounds, result consumer and check ownership.
    Resume an author for direct repair when useful. Validators always receive fresh
    context without the author's desired verdict. Observe actual dispatch settings;
    prompt wording proves neither isolation nor smaller inherited context.
    
    Return concise findings, check facts and evidence references in the existing
    handoff; disclose missing or truncated evidence. Derive the combined subject's
    manifest and applicable orphan scan at the integration/judgment boundary.
    Unjudged worker increments supply content identity and check facts, not duplicate
    final evidence bundles. A separately judged subject still needs complete proof.
    Machine evidence such as `verdict.v2` is optional
    unless requested or required by a declared consumer. When no machine
    artifact is requested or required, return the result without creating one.
    
    ## Causal stall and bounds
    
    Unknown cause, recurrence, no progress or a wrong objective admits
    at most one bounded fresh helper for that incident within authority and bounds.
    Give it the failed assumption, evidence and one discriminating question. Resume
    only with a different testable approach; an unhelpful answer ends the attempt.
    Do not chain helpers or rename the incident. Known failures get direct repair.
    Cancellation, refusal and spent hard time/cost/quota skip help.
    
    Respect actual caller/native limits, including explicit repair-round bounds.
    Retries, compaction, helpers and new subjects never renew them; retry count
    alone is not a spent budget. If interruption threatens evidence, preserve
    accepted intent, exact subject, useful receipts, unresolved cause, bounds and
    helper use in the native handoff. Prompt text proves no native enforcement.
    [Outer-goal guidance](references/outer-goal.md) remains optional.
    
    ## Evidence and boundaries
    
    Bind accepted intent, complete changed paths, exact subject and factual receipts
    for the fresh validator; disclose affected orphaned acceptance evidence. Use
    existing provenance helpers rather than a new evidence format. Requested proof
    uses caller-selected protected external non-Git storage; preserve legacy
    `.agents/` evidence. Missing identity, freshness or proof means NOT_PROVEN;
    proven failed acceptance or scope violation means FAIL. PASS needs every
    criterion verified and empty `not_checked`. Authors cannot issue binding PASS.
    
    [Memory](../memory/SKILL.md), specialists and runtime adapters are on demand;
    no-match and no-change are valid. Read [boundaries](references/boundaries.md)
    when authority, scope, evidence or delivery is at issue. The optional
    [fixed-dispatch adapter](references/bounded-adapter.md) is not the native
    execution engine. Do not invent a runtime, hidden machine artifact or workflow
    to finish an ordinary change.
    
    Report the result, strongest checks and material limits. Plans, activity,
    reviews and saved pages earn no capability credit; NOT_PLANNED and NOT_BUILT
    are progress descriptions, not semantic verdicts.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related