rpi
Coordinate one RPI traversal: one bounded Plan and Implement experiment, then fresh Validate and a bounded repair phase to convergence. Triggers: "run rpi", "run one traversal", "execute this plan", orchestration or worker delegation that implements changes.
Install
npx skills add https://github.com/boshu2/agentops/tree/main/images/gemini/skills/rpi
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install boshu2-agentops@llmmart
git clone https://github.com/boshu2/agentops.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole boshu2/agentops collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
RPI
Own the authorized outcome through finish. Use the native coding agent and shell. BD or the caller's tracker owns work and handoffs; Git owns content and delivery. AgentOps supplies a small charter and fresh judgment, not a scheduler.
Operating charter
- Use the existing accepted outcome, scope and real bounds. A clear change needs no Plan, Recall or Learn worksheet. Resolve uncertainty only when it could change the implementation or acceptance decision.
- Take the smallest acceptance-advancing action. Plan shapes missing intent or revises a disproved approach. Once an implementer can act and a validator can judge, implement; do not keep improving the plan. Approach revisions preserve acceptance and authorized scope. Acceptance changes need caller authority.
- Implement and repair ordinary known defects directly. A known test failure needs a fix and a discriminating check, not another planning phase, council or helper.
- Use focused checks during edits and complete required integration checks before final judgment. Reuse valid exact-input receipts; rerun affected checks after changes. Reserve capacity for integration, review and repair. Keep the final subject unchanged while it is being judged.
- Obtain Validate from one fresh author-distinct context in the author's model family unless the caller selects required additional legs. There is no fixed ten-minute cap; explicitly required reviewers remain required. Risk deepens evidence, not reviewer multiplication. Repair actionable findings within authority and remaining bounds, then revalidate the changed exact subject.
- Stop at completed acceptance, cancellation, refusal, a spent real bound or an unresolved causal stall after the help below. Adjacent improvements are not permission to expand the goal. Report them briefly only when useful; do not turn them into another work batch.
Context and handoffs
Load required contracts once per context, then read only what the next decision needs. A reference link is available context, not a reading list. Search before opening large files; expand only for consequential uncertainty. Keep successful output compact at the tool boundary; retain full logs for inspection. Reuse the worker's component-check list and current receipts instead of rediscovering them. Use native completion watches or bounded waits for ongoing checks and helpers. At completion, verify the expected subject and required results; a quiet or partial status is not success. Inspect further for a failure, suspected stall or decision need. Keep required user updates concise rather than narrating each poll.
When delegation is authorized and useful, select the runtime's task-only dispatch option for independent work; a short prompt in a full-history fork still carries full history. Supply accepted intent/scope, exact subject, relevant evidence, remaining bounds, result consumer and check ownership. Resume an author for direct repair when useful. Validators always receive fresh context without the author's desired verdict. Observe actual dispatch settings; prompt wording proves neither isolation nor smaller inherited context.
Return concise findings, check facts and evidence references in the existing
handoff; disclose missing or truncated evidence. Derive the combined subject's
manifest and applicable orphan scan at the integration/judgment boundary.
Unjudged worker increments supply content identity and check facts, not duplicate
final evidence bundles. A separately judged subject still needs complete proof.
Machine evidence such as verdict.v2 is optional
unless requested or required by a declared consumer. When no machine
artifact is requested or required, return the result without creating one.
Causal stall and bounds
Unknown cause, recurrence, no progress or a wrong objective admits at most one bounded fresh helper for that incident within authority and bounds. Give it the failed assumption, evidence and one discriminating question. Resume only with a different testable approach; an unhelpful answer ends the attempt. Do not chain helpers or rename the incident. Known failures get direct repair. Cancellation, refusal and spent hard time/cost/quota skip help.
Respect actual caller/native limits, including explicit repair-round bounds. Retries, compaction, helpers and new subjects never renew them; retry count alone is not a spent budget. If interruption threatens evidence, preserve accepted intent, exact subject, useful receipts, unresolved cause, bounds and helper use in the native handoff. Prompt text proves no native enforcement. Outer-goal guidance remains optional.
Evidence and boundaries
Bind accepted intent, complete changed paths, exact subject and factual receipts
for the fresh validator; disclose affected orphaned acceptance evidence. Use
existing provenance helpers rather than a new evidence format. Requested proof
uses caller-selected protected external non-Git storage; preserve legacy
.agents/ evidence. Missing identity, freshness or proof means NOT_PROVEN;
proven failed acceptance or scope violation means FAIL. PASS needs every
criterion verified and empty not_checked. Authors cannot issue binding PASS.
Memory, specialists and runtime adapters are on demand; no-match and no-change are valid. Read boundaries when authority, scope, evidence or delivery is at issue. The optional fixed-dispatch adapter is not the native execution engine. Do not invent a runtime, hidden machine artifact or workflow to finish an ordinary change.
Report the result, strongest checks and material limits. Plans, activity, reviews and saved pages earn no capability credit; NOT_PLANNED and NOT_BUILT are progress descriptions, not semantic verdicts.
Files (agentops)
-
references
-
boundaries.md 4.7 KB
# Ownership boundaries for the lean RPI core RPI owns the authorized outcome through implementation, checks, direct repairs and fresh final judgment. Plan shapes missing intent and may revise an approach falsified by evidence within unchanged accepted outcome/scope. Implement edits and collects facts. Validate independently judges the exact subject and alone authors semantic `verdict.v2` when persistence is selected. Memory is optional; its operation references own recall, mining and curation. ## Native authority BD or the caller's tracker owns work/status/dependencies/handoffs. Git and repository policy own content/history and delivery. Native runtimes and callers own aggregate budgets, work selection, queues, claims, stops and subsequent outcomes. A skill grants no extra Git, tracker, publishing or credential permission. Existing caller authorization remains usable; do not invent another approval step merely because a phase changed. Keep one authoritative work account, not a parallel AgentOps ledger. The runtime derives exact intent/subject identity, complete changed paths, receipts and observed context identities. Never invent a model/context identity or transcribe a fictional runtime packet. New requested proof uses protected external non-Git storage; preserve legacy `.agents/` proof under owner policy. Plans, dashboards and reviews count as subject completion only when requested. ## Direct repair and help Known failures get direct repair. Evidence that disproves an assumption permits approach revision under unchanged acceptance and scope. Acceptance or authority expansion needs caller approval; useful source/generated changes already covered by a scope class do not. Cheap discriminating checks precede expensive judgment. Reserve finishing capacity and use valid exact-input receipts when applicable. Unknown cause, recurrence, no progress or wrong objective warrants causal examination. A genuine stall admits at most one bounded fresh helper per incident inside existing authority and bounds. Do not build a helper chain for known failures or rename an unresolved incident. An unhelpful helper ends the attempt. Cancellation, refusal or spent real limits skip help; retry counts alone are not spent time/quota. Compact native recovery state preserves evidence, not new budget. ## Optional specialists and adapters Anti-ceremony, premortem, council, research, factories and runtime adapters are optional. Risk deepens evidence inspection without mandatory specialist dispatch. No Recall or Learn toll applies to trivial work. A selected factory remains behind its own coordinator, doctor and supervisor doors; its reconciler creates and repairs sessions. Concurrent writers require authorized disjoint source and regeneration scope and isolation. Pass bounded task evidence, not the author's desired verdict. Do not start another runtime merely because it exists. The optional `run_once.py` developer adapter retains its explicitly selected fixed-dispatch and finite-round contract in [bounded-adapter.md](bounded-adapter.md). It does not restrict native approach revision or implement direct repair for you. ## Fresh judgment The author cannot issue binding PASS. Judge legs read; implementers fix. Default to a fresh author-distinct same-family reviewer. Cross-model review is opt-in; an explicitly required unavailable leg leaves NOT_PROVEN. No fixed ten-minute cap applies, and no invocation renews caller/native limits. PASS needs exact subject continuity, complete changed-path coverage, unchanged acceptance, distinct context IDs and attested freshness, nonempty checked scope, evidence for every criterion and empty `not_checked`. Incomplete proof remains NOT_PROVEN; failed acceptance or proven out-of-scope changes remain FAIL. Necessary findings never become optional to get green. Judge disagreement stays visible and never becomes PASS by preference or majority vote. Validate returns judgment, not a repair or delivery instruction. RPI completes existing authorized work before reporting, within real bounds. Report the subject, strongest evidence and any remaining acceptance gaps; persist a machine artifact only for a declared consumer or caller request. A new subject requires new final judgment. Mutating checks run on a disposable copy or committed subject so they cannot overwrite the judged working tree. ## Observed guardrails and limits The July 2026 unlisted-regeneration incident supports scope as a class; it does not authorize unrelated files. The July mutating-check incident supports the quarantine; it does not require rerunning every expensive check. The planning spiral supports smallest useful action; it does not forbid revising a falsified approach. These rules protect actual work and may be revised by later evidence. -
bounded-adapter.md 3.9 KB
# Optional fixed-dispatch reference adapter The grandfathered `scripts/run_once.py` is a pure developer reference for callers that explicitly select fixed dispatch and a finite list of supplied review rounds. It invokes an explicit anti-ceremony function, Plan and Implement at most once; its repair evaluator consumes supplied evidence and cannot execute agents, infer causes, fix subjects or enforce aggregate budgets. Installed native RPI follows its operating charter, not this adapter. The old phase lock and default two-round limit apply only to this selected adapter, never as a restriction on native implementation's direct repairs or evidence-driven approach revision. Its existing deterministic tests guard exact evidence and finite consumption; they do not prove native agent behavior or practical benefit. It stops when converged, stopped by the law, or out of `repair_rounds` and never extends the caller's bound. The adapter preserves these narrower admission semantics: ## The convergence law A repair round is admitted only while all hold: 1. `rounds_used < repair_rounds` (caller-declared, default 2). 2. New digest-bound evidence proves closure of a named acceptance finding or, for `NOT_PROVEN`, resolves a named proof gap. A changed digest or a smaller finding count alone is not useful progress. Generated-only changes qualify only when the evidence proves that they repair required behavior or parity. An unchanged subject previously judged FAIL cannot be repaired by a new label or verdict flip; changed bytes still require acceptance proof. 3. No finding id closed in an earlier round reopens. No closed finding class recurs, and no introduced regression or new finding of unknown cause is admitted. Before/after reproduction or equivalent causal evidence under the same acceptance must distinguish a pre-existing discovery from a regression; neither counts, timestamps, nor a new id establish that distinction. Keep the union of every required judge's findings, keyed by stable `findings[].id`; do not hide a necessary finding as optional. Newly exposed pre-existing defects may increase the open count while another acceptance gap is demonstrably closed. Their evidence must prove prior existence; unknown cause stops repair for causal examination even if another gap closed. Validators reuse a short stable `class` for each kind of defect. A reopened id or returning class warrants causal HOLD in a selected outer goal. Recurrence alone does not prove that the design is wrong and never auto-reopens Plan. Reuse existing check receipts, findings summaries, and evidence references for this reasoning. In the pure reference, decoded receipt bindings use `ref`, `subject_digest`, and `resolves` for ids actually closed. `preexisting` ids must bind reproduction to the prior subject digest; `introduced` ids bind causal comparison to the current digest and stop repair. These are supplied receipt facts, not new persisted verdict fields or a lifecycle schema. The reference cannot prove a receipt's truth or infer cause from wording. Converged: the fresh validator returns PASS and every required cross-family validator does too, over the exact subject and all acceptance with empty `not_checked`. On any violation RPI stops and reports the current status. `checked` carries one line per round (`repair round N: k open findings`); open findings ride in the result and the report. A reworded finding with the same id is the same finding. Acceptance and its digest stay fixed. The orchestrating context fixes; judge legs only read. RPI convenes no further judge of its own, does not escalate, and does not auto-replan. An unknown cause or recurrence returns evidence to the native caller; the adapter dispatches no helper. The native charter decides whether a genuine causal stall merits its single bounded consultation. This does not revive a spent caller bound or change a completed verdict. -
outer-goal.md 1.8 KB
# Optional outer goal Use this reference only when the caller explicitly selected a sustained goal or several outcomes. The native caller/controller keeps work selection, aggregate budgets, stops and delivery authority. No AO scheduler, new command or goal ledger is required. The RPI charter already owns a single authorized outcome through finish; an outer goal is not permission needed for ordinary repair. Carry accepted terminal outcome and scope, measured remaining allowance and the current causal incident in the native work/handoff source. Choose the smallest acceptance-advancing action or consequential uncertainty. Reserve capacity for integration, final fresh judgment, required repairs and a useful handoff before spending the whole allowance on discovery or reviews. Apply the charter's at-most-one bounded helper to a genuine causal stall. Known failures get direct repair. A repeated wakeup, new context or renamed finding is not a new incident. A helper with no useful new approach ends that attempt; report the unresolved cause and required caller decision. Cancellation, refusal and spent real bounds skip help. Native blocked-status thresholds are bookkeeping, not permission to renew time, cost, quota or helper use. Report observed native enforcement and unmeasured limits honestly. No objective text, saved plan or simulated stop proves aggregate runtime enforcement. The caller may authorize a new outcome or scope; agents may revise an approach when evidence disproves an assumption within unchanged authority. Existing host/user policies may impose stricter helper or stopping requirements. This repository charter does not update those installed host instructions. Inspect and report the effective upstream rule instead of claiming the lean contract overrides it or that an optional guide enforces native controls. -
rpi.feature 2.9 KB · in bundle
-
-
SKILL.md 6.7 KB
--- name: rpi description: 'Apply the outcome-to-judgment charter. Use when: the caller explicitly selects RPI; ordinary coding, delegation and native goals do not require this workflow.' practices: - bdd-gherkin - tdd - design-by-contract hexagonal_role: domain consumes: - plan - implement - validate produces: - rpi-report.v1 context_rel: - kind: customer-of with: plan - kind: customer-of with: implement - kind: customer-of with: validate skill_api_version: 1 user-invocable: true disable-model-invocation: true metadata: graph_root: true tier: meta dependencies: [plan, implement, validate] capabilities: [own_authorized_outcome, report] effects: [dispatch_core_phases] canonical_status: canonical disposition: keep_strategy output_contract: 'concise human-readable result; optional rpi-report.v1 when a caller or declared consumer requests machine-readable evidence' --- # RPI Own the authorized outcome through finish. Use the native coding agent and shell. BD or the caller's tracker owns work and handoffs; Git owns content and delivery. AgentOps supplies a small charter and fresh judgment, not a scheduler. ## Operating charter 1. Use the existing accepted outcome, scope and real bounds. A clear change needs no Plan, Recall or Learn worksheet. Resolve uncertainty only when it could change the implementation or acceptance decision. 2. Take the smallest acceptance-advancing action. [Plan](../plan/SKILL.md) shapes missing intent or revises a disproved approach. Once an implementer can act and a validator can judge, implement; do not keep improving the plan. Approach revisions preserve acceptance and authorized scope. Acceptance changes need caller authority. 3. [Implement](../implement/SKILL.md) and repair ordinary known defects directly. A known test failure needs a fix and a discriminating check, not another planning phase, council or helper. 4. Use focused checks during edits and complete required integration checks before final judgment. Reuse valid exact-input receipts; rerun affected checks after changes. Reserve capacity for integration, review and repair. Keep the final subject unchanged while it is being judged. 5. Obtain [Validate](../validate/SKILL.md) from one fresh author-distinct context in the author's model family unless the caller selects required additional legs. There is no fixed ten-minute cap; explicitly required reviewers remain required. Risk deepens evidence, not reviewer multiplication. Repair actionable findings within authority and remaining bounds, then revalidate the changed exact subject. 6. Stop at completed acceptance, cancellation, refusal, a spent real bound or an unresolved causal stall after the help below. Adjacent improvements are not permission to expand the goal. Report them briefly only when useful; do not turn them into another work batch. ## Context and handoffs Load required contracts once per context, then read only what the next decision needs. A reference link is available context, not a reading list. Search before opening large files; expand only for consequential uncertainty. Keep successful output compact at the tool boundary; retain full logs for inspection. Reuse the worker's component-check list and current receipts instead of rediscovering them. Use native completion watches or bounded waits for ongoing checks and helpers. At completion, verify the expected subject and required results; a quiet or partial status is not success. Inspect further for a failure, suspected stall or decision need. Keep required user updates concise rather than narrating each poll. When delegation is authorized and useful, select the runtime's task-only dispatch option for independent work; a short prompt in a full-history fork still carries full history. Supply accepted intent/scope, exact subject, relevant evidence, remaining bounds, result consumer and check ownership. Resume an author for direct repair when useful. Validators always receive fresh context without the author's desired verdict. Observe actual dispatch settings; prompt wording proves neither isolation nor smaller inherited context. Return concise findings, check facts and evidence references in the existing handoff; disclose missing or truncated evidence. Derive the combined subject's manifest and applicable orphan scan at the integration/judgment boundary. Unjudged worker increments supply content identity and check facts, not duplicate final evidence bundles. A separately judged subject still needs complete proof. Machine evidence such as `verdict.v2` is optional unless requested or required by a declared consumer. When no machine artifact is requested or required, return the result without creating one. ## Causal stall and bounds Unknown cause, recurrence, no progress or a wrong objective admits at most one bounded fresh helper for that incident within authority and bounds. Give it the failed assumption, evidence and one discriminating question. Resume only with a different testable approach; an unhelpful answer ends the attempt. Do not chain helpers or rename the incident. Known failures get direct repair. Cancellation, refusal and spent hard time/cost/quota skip help. Respect actual caller/native limits, including explicit repair-round bounds. Retries, compaction, helpers and new subjects never renew them; retry count alone is not a spent budget. If interruption threatens evidence, preserve accepted intent, exact subject, useful receipts, unresolved cause, bounds and helper use in the native handoff. Prompt text proves no native enforcement. [Outer-goal guidance](references/outer-goal.md) remains optional. ## Evidence and boundaries Bind accepted intent, complete changed paths, exact subject and factual receipts for the fresh validator; disclose affected orphaned acceptance evidence. Use existing provenance helpers rather than a new evidence format. Requested proof uses caller-selected protected external non-Git storage; preserve legacy `.agents/` evidence. Missing identity, freshness or proof means NOT_PROVEN; proven failed acceptance or scope violation means FAIL. PASS needs every criterion verified and empty `not_checked`. Authors cannot issue binding PASS. [Memory](../memory/SKILL.md), specialists and runtime adapters are on demand; no-match and no-change are valid. Read [boundaries](references/boundaries.md) when authority, scope, evidence or delivery is at issue. The optional [fixed-dispatch adapter](references/bounded-adapter.md) is not the native execution engine. Do not invent a runtime, hidden machine artifact or workflow to finish an ordinary change. Report the result, strongest checks and material limits. Plans, activity, reviews and saved pages earn no capability credit; NOT_PLANNED and NOT_BUILT are progress descriptions, not semantic verdicts.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.