codex-handoff
Imported from paulrberg/agent-skills/skills/codex-handoff.
Install
npx skills add https://github.com/PaulRBerg/agent-skills/tree/main/skills/codex-handoff
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install paulrberg-agent-skills@llmmart
git clone https://github.com/PaulRBerg/agent-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole paulrberg/agent-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Codex Handoff
Codex-handoff orchestrates read-only investigation or implementation within the current session after explicit plan approval. Task-handoff instead writes a decision-complete file for a fresh, separate session; use it when work continues later or elsewhere, and use an in-session handoff skill to implement an approved plan now.
If these instructions are already present in the conversation from a slash or dollar invocation, follow them directly; do not invoke this skill again through a skill tool.
Follow the shared contract below, select exactly one host adapter, and use it for every host-specific action.
Host Selection
Inspect the callable orchestration tools, not environment variables, process ancestry, or a user-supplied host name:
spawn_agent,wait_agent,send_message,followup_taskavailable: readreferences/codex-cli-host.mdcompletely.- Otherwise, Claude Code's Agent and Bash tools available: read
references/claude-code-host.mdcompletely. - Neither: stop with a compatibility error. Native Codex multi-agent support is mandatory on the Codex host; never fall back to a nested Codex CLI process.
Use exactly one adapter for every host-specific action; never load both or combine their launch, progress, retry, permission, or result-transport mechanics. The adapter may specialize host mechanics and manifest configuration but cannot weaken this shared contract.
Contract
- Run only after explicit invocation. Research-only = requested outcome is findings/evidence/assessment only, no repo changes or plan requested. All handoffs may run in any host mode; implementation handoffs must pass through the Plan Phase and receive explicit user approval before launch.
- Reuse an already approved plan when the outcome and material constraints are unchanged. Explicit user instructions take precedence over skill defaults; ask again only for an unresolved decision or action outside that authorization.
- The parent owns decisions, the final plan, and orchestration. Delegate investigation to read-only research agents only when task and repository evidence make it useful before planning.
- Research agents gather evidence and report findings only — never edit files, make design decisions, or return plans.
- Implementation agents implement their assigned part of the approved plan: inspect, edit, validate, never redesign or return another plan.
- Use the smallest effective implementation team (one agent is valid); add agents only when decomposition materially improves latency, correctness, or verification. Never exceed eight implementation agents total.
- Size every brief before finalizing the team: estimate its wall-clock time and split any brief likely to exceed roughly 25-30 minutes into parallel disjoint scopes or dependency waves — add an integration agent if needed — instead of one monolithic agent.
- Use at most three research agents, stable IDs
R1-R3, counted separately from the eight implementation agents. - Keep the parent's own implementation work to orchestration, integrity checks, failure handling, and conditional polish passes.
- Treat an explicit user model preference (e.g. GPT-6 Luna) as an orchestration constraint on every research and implementation agent unless scoped narrower; don't substitute the adapter's usual Luna/Sol/Astra selection. If the host can't launch that model, report the incompatibility and ask before falling back.
- Treat the approved outcome — not the initial manifest or its write scopes — as the authorization boundary: when implementation reveals a related in-repository fix or evidence change required for that outcome, the parent may extend the handoff and launch follow-on agents for the newly discovered scope without asking again. The worker that discovered the need still stops at its assigned scope and returns evidence; the parent owns scope expansion, repository coordination, and delegation.
- Size verification to the requested outcome. Never add validation machinery (gates, manifests, checkpoints, hash pins, journals, receipts) unless the approved plan explicitly calls for it. An explicit user request to hurry or wrap up overrides optional repeat checks and required polish passes: commit the validated work and report what was skipped or left unverified.
Use $ARGUMENTS as the task when present; otherwise use the active user request. A task naming another skill follows
Companion Skills.
Companion Skills
A task that names another skill alongside this handoff — for example,
run $fresh-eyes-sweep and delegate the implementation — is a composition: the companion skill defines the work; this
handoff owns delegation mechanics.
- Load the companion skill and run its discovery, judgment, and planning phases in the parent. Route read-only discovery
through this skill's research agents when materially faster, then fold the companion's method, findings, and
constraints into the handoff plan. A companion missing from the host's skill list may be installed but hidden by
disable-model-invocation: true: read itsSKILL.mddirectly from the host skill root (~/.claude/skills/<name>/SKILL.mdin Claude Code,~/.agents/skills/<name>/SKILL.mdin Codex) before concluding it is absent. - When the companion's discovery is itself the bulk of the work — an audit or sweep over a whole repository or large file set — the parent maps the scope and slices it instead of reading it inline. After plan approval, each implementation agent audits and fixes its own slice under the companion's rules, inlined in its brief.
- This contract overrides the companion's overlapping plan approval, agent limits and stable IDs, single validation owner, result fields, failure classification, commit ownership, and completion reporting, even when the companion prescribes its own subagent, validation, or commit mechanics. Its user-decision gates still bind; this skill's plan approval satisfies coincident gates, and a plan-only companion implements only when the combined invocation requested implementation and the user approved the plan.
- Agents cannot load skills. Never brief one to "use skill X" by name; inline the specific companion instructions, conventions, or excerpts it needs.
- A companion requirement to run
$code-polishor$agents-brain polishmarks that pass required in the Plan Phase; it still runs once, per Completion. Companion commit instructions never add commits. - Satisfy both contracts at Completion: produce the companion's required report artifacts (ledgers, tables, verdicts) alongside the selected adapter's completion report.
Research Phase
For a research-only task, launch one to three research agents and stop after returning the consolidated investigation; never enter the Plan Phase or launch implementation agents. For an implementation handoff, trigger research when scope is uncertain, the task crosses multiple or unfamiliar subsystems, or gathering the needed evidence serially would be materially slower for the parent. Zero research agents is the default for implementation handoffs. The parent alone decides the research count from task and repository evidence; never ask the user to opt in or name agents. Either launch the research wave immediately or proceed straight to planning.
When triggered, assign up to three agents stable IDs R1-R3 and launch them immediately through the selected
adapter's read-only mechanism. Give each agent a self-contained prompt containing:
- the open questions and exact investigation scope;
- relevant repository constraints and known concurrent-work boundaries;
- a strict read-only authority boundary;
- the stopping rule that it must return evidence rather than a plan or design;
- its time budget, with the instruction to use it: an early
blockedreturn citing only time is not a valid stop; and - exact result fields:
status,findings,open_questions,evidence, andblockers.
A research agent that returns blocked citing only its time budget while most of that budget is unused and no concrete
obstacle is named has not settled its scope, and neither has one that returns completed while its evidence still lists
uninspected scope paths: continue the same agent once through the adapter's same-agent mechanism with the uncovered
files and the remaining budget. This continuation is not a new research agent.
When every required research agent settles, fold its findings and evidence into the implementation plan or the research-only response. Surface open questions or blockers through the host's user-question mechanism only when they change scope or approach. Do not reconcile the working tree — research agents change nothing; any reported edit is a contract violation.
For a research-only task, synthesize the evidence and finish with ### 🔎 Research handoff — <completed|blocked>, the
agent count, findings, evidence, open questions, and blockers. This replaces the Plan Phase and the selected adapter's
implementation completion report. If the investigation shows that changes are needed, report them as findings and stop —
do not produce an implementation plan or begin edits.
Plan Phase
Enter this phase only for an implementation handoff.
When research — delegated or the parent's own — contradicts a fact the user stated explicitly (quantities, which items, which accounts), ask through the host's user-question mechanism before writing the plan; never widen the plan's default scope to fit the research.
Produce a decision-complete plan with this section and the selected adapter's exact manifest table:
## Codex Handoff
- Research: `<none | R1..Rn — key findings used>`
- Companion skills: `<none | $x — phases the parent runs / what is folded into briefs>`
- Strategy: `<sequential|parallel|hybrid>`
- Agents: `<1-8>` — `<why this is the smallest effective count>`
- Validation owner: `<agent-id|parent>` — `<aggregate checks it runs once>`
<host-adapter manifest table>
- Code polish: `<required|not required>` — `<reason>`
- Agent-context polish: `<required|not required>` — `<reason>`
While finalizing the plan, record the union of every manifest write scope with
ai-coord draft --name <plan-slug> '<label>' '<path>'..., using --recursive only for directory scopes. Derive
<plan-slug> from the plan's short identity (for example debarrel-lib): lowercase it, replace characters outside
[A-Za-z0-9._-] with -, and truncate to 40 characters. For two or more Git roots, use
ai-coord bundle draft --name <plan-slug> '<label>' '<absolute-path>'.... In the plan's "Wait out conflicting agents"
section, write the exact promote command ai-coord start --draft <plan-slug> (or
ai-coord bundle start --draft <plan-slug>) and retain the explicit ai-coord start '<label>' '<path>'... fallback (or
ai-coord bundle start '<label>' '<absolute-path>'...) over the same union. A fresh implementation session must receive
these commands in the plan itself, without reconstructing scopes from prose. Use the fallback only when promotion
reports no draft named .... Named drafts grant no authority and expire after seven days.
Choose the execution shape from repository evidence and the approved work:
- Sequential: one agent depends on another, write scopes overlap, or a later agent owns integration or aggregate validation.
- Parallel: independent work only, with explicitly disjoint write scopes. Agents may inspect shared context but must not write outside their assigned scope.
- Hybrid: dependency-ordered waves — run independent agents within a wave in parallel, reconcile the entire wave, then start its dependents.
A wave finishes with its slowest agent. Keep the highest-tier agent's scope minimal and move deferrable validation to the validation owner. If parallel work does not collectively prove the overall plan, reserve a later sequential agent for integration and aggregate validation. Use stable agent IDs and explicit dependencies across the whole handoff.
Assign aggregate validation to exactly one owner: package- or repository-wide checks run once, by the integration agent when one exists, otherwise by the parent during post-wave reconciliation. Every other agent runs only the narrowest checks that prove its own edits, such as file-scoped formatting, lint, or typecheck plus targeted tests.
Require $code-polish for nonlocal invariants, concurrency or state machines, migrations or parsing, auth or security,
retry or error semantics, and public API or data-contract changes. File count alone is not a trigger.
Require $agents-brain polish when approved work changes a target its polish workflow supports: README.md, AGENTS.md or
CLAUDE.md, a durable context doc, an existing project-installed skill under .agents/skills, or an existing git-tracked
source-catalog skill under skills/ where polish is prose-only. Installed copies under managed agent-config roots
remain excluded. Mark both passes required when both trigger rules apply; mark neither when neither applies.
Do not launch implementation agents until the user approves the plan. Read-only research is the only pre-approval exception.
Implementation Prompt Contract
Build a self-contained, outcome-first prompt for every implementation agent. Include:
- The approved overall outcome plus the agent's implementation brief, dependencies, and completion evidence.
- Its exact write scope, relevant repository constraints, known dirty-work boundaries, and prerequisite agent results.
- Its validation assignment per the Plan Phase's single validation owner: scoped checks it must run and, unless it owns validation, that it must not run aggregate checks. Never brief new validation machinery the approved plan does not call for.
- A pacing estimate matching its manifest sizing and any user- or adapter-imposed hard runtime limit. A soft estimate alone is not a stop condition; report partial evidence and the concrete blocker or exhausted hard limit when blocked.
- This authority boundary: inspect, edit only within the assigned scope, and validate locally; never commit, push, deploy, make external writes, or broaden scope, even when repository or host instructions favor committing finished work promptly. Committing stays with the parent after reconciliation.
- The selected adapter's delegation and coordination context, including why the parent session and disjoint siblings
are not conflicting work and what unrelated exact-scope claim would justify returning
blocked. Delegates must not run coordination lifecycle commands: these are rejected with exit 64. Permit onlyai-coord status,ai-coord touched,ai-coord inbox,ai-coord msg, andai-coord finding. - This stopping rule: implement the approved plan exactly; if infeasible or requiring redesign, return
blockedwith evidence instead of proposing a replacement plan. Continue after progress updates while authorized work remains; a milestone, offer to continue, or list of nonblocking decisions is not a completed result. - A requirement to return every result field:
status(completedorblocked),summary,changed_fileslisting only files actually touched,verificationlisting every command and outcome,residual_risks, andblockers.
Keep each agent prompt as compact as completeness allows: the shared outcome summary plus that agent's own brief, scope, and constraints — never restate the full plan text per agent.
Add the selected adapter's command, permission, transport, and host-tool constraints without restating this contract.
Execution and Reconciliation
Before implementation wave 1, the parent promotes the plan's named draft to acquire the full manifest write-scope union.
Use the plan's recorded explicit start fallback only when promotion reports no draft named ...; require READY before
launching agents. When promotion or start queues or blocks instead, never end the turn to pause: run ai-coord wait
through the adapter's wait mechanics. It also returns on non-readiness wake events (message, unknown coverage, work
release, 300-second default timeout); on each wake, handle MESSAGE events through ai-coord inbox, re-submit the
recorded promote or start command, and diagnose stale blockers.
Launch agents through the selected adapter in the approved strategy and dependency waves. Do not add agents or change models, efforts, scopes, or validation ownership merely because a worker is slow or quiet. Do revise the manifest and launch a narrowly scoped follow-on agent when completed work discovers an unplanned prerequisite covered by the approved outcome. Preserve stable IDs, dependency order, the eight-agent limit, and one aggregate-validation owner; include follow-on agents in the final counts and report.
For each completed agent: require every shared result field and treat changed_files as its authoritative post-pass
scope; confirm reported files exist or were intentionally deleted, stay within scope, and carry verification evidence
matching the assignment; and pass relevant completed results to dependent agents.
After every implementation wave, reconcile all results with the current manifest and visible working tree without folding in unrelated concurrent changes. When the parent owns validation, run the assigned aggregate checks once during this reconciliation. Attribute aggregate-check failures before blocking: first rule out effects of the handoff's changes, formatters, hooks, and generators, including failures in downstream files outside its write scopes. Continue past a failure only when evidence establishes that it is unrelated and the handoff's own checks still pass. Unexpected out-of-scope edits, same-wave overlap, or a failure attributable to the handoff are blockers; do not start dependents or polish, and do not silently take over implementation.
Failure Classification
- A
status: blockedresult identifying a related in-repository fix or evidence change outside the worker's scope that is necessary for the approved outcome is follow-on work under the Contract's scope-expansion authority, not a request for fresh authorization: let already-started independent agents finish, gate dependents, extend the manifest with the smallest sufficient scope, satisfy repository coordination for that scope, and launch a new or reused implementation agent. Repeat until the outcome is complete or a genuine authorization boundary is reached. - Ask the user only when continuation would change the approved outcome, require a material redesign or unrelated work, or cross an existing confirmation boundary such as destructive action, purchase, deployment, or external write. Never silently take over implementation or relaunch solely on a larger model.
- Treat a tool or infrastructure failure as retryable only when adapter-specific evidence supports that classification. Inspect partial edits first, then use the adapter's same-agent mechanism for exactly one verify-and-continue attempt; this is not a new agent against the eight-agent limit. A second infrastructure failure blocks that agent and its dependents.
- Never classify an ordinary timeout, a returned blocker, silence, or task-level validation failure as infrastructure failure. Continue only work proven independent.
Skill Evolution Review
Keep verified repairs to skills used during the handoff separate from the optional review below. When user or repository instructions already authorize repairs, the parent owns their completion; subagents report evidence without expanding their write scopes. One verified occurrence is enough, and a blocked main task does not prevent independent repairs. Complete the handoff's required work or establish its blocker, then finish independent repairs before the final report under the applicable maintenance policy. Plan Mode still prohibits edits.
After every required agent succeeds and the task is verified — never for a blocked, failed, or partial handoff — the
parent alone judges skill-evolution opportunities; agents never make the recommendation. Recommend only a stable,
reusable workflow credibly likely to recur; reject one-offs, rare contingencies, incidental cleanup, and speculative
value. For a new skill, state repo-local vs. global ~/projects/agent-skills placement; for a revision, name the exact
skills and why. When a proposal clears this bar, append at most one compact suggestion (≤2 short sentences) to the
adapter's existing completion report without changing its format, then offer $task-handoff as the next action. Never
auto-invoke $task-handoff or create/revise anything during this review; when nothing qualifies, stay silent — no
placeholder.
Completion
- After every required agent completes, deduplicate the union of reported
changed_filesand confirm the combined verification evidence proves the approved plan. - Before the completion report, fix remaining same-pattern sites the approved outcome covers through follow-on agents per Failure Classification; never list them as optional or out-of-scope items.
- If any required agent failed, or the user explicitly asked to hurry or wrap up, skip every planned polish pass and
report the skip. Otherwise, invoke each required pass once with only its applicable paths from that union:
$code-polishfirst in its default simplify-then-review mode, then$agents-brain polishwith its eligible context targets. Invoke only one when only one is required. Do not seed either pass with paths outside the union or let it broaden beyond its declared workflow authority. - Reconcile in-scope files actually changed by each polish pass into the final changed-files set and verification. A required polish pass that blocks, fails, or writes outside its supported scope blocks later polish and cross-repository commits.
- If approved work changes repositories on this machine other than the one where the handoff began, invoke
$commitfrom each additional repository after its work, validation, and required polish complete, scoped to files changed there; do not commit incomplete, blocked, unexpected, or out-of-scope changes. Push when the request or standing user instructions authorize it. - When the handoff pushed commits and the repository defines CI workflows, such as
.github/workflows, watch the pushed head's runs before the completion report (gh run list --commit <sha>, thengh run watch <run-id>, in the background when the host supports it). Fix failures attributable to the handoff as follow-on work and report the CI outcome. When changed code behaves differently by platform and local checks covered only one, name the unverified platforms as a risk. - Finish with the selected adapter's completion report, including strategy, wave and agent counts, each agent's
requested configuration, status, and summary, plus combined changed files, verification, polish when run (listing each
pass and outcome), automatic cross-repository commit hashes when any, and
Issues and caveatswhen present. Writenonefor other applicable empty values and never expose machine result payloads. - Group issues and caveats as
Resolved(verified fixes with evidence) andOpen(remaining problems, limitations, or unverified assumptions, with impact and next step). Omit empty groups and the whole section when empty. Report each item once; put neutral context and agreed decisions under changes or scope. Reserveblockerfor something preventing required work andriskfor a specific potential adverse outcome. A workaround leaves an item open when the underlying issue still affects the result.
Files (agent-skills)
-
agents
-
openai.yaml 42 B
policy: allow_implicit_invocation: true
-
-
references
-
claude-code-host.md 15.8 KB
# Claude Code Host Adapter Load this adapter only when `SKILL.md` selects the Claude Code orchestration surface. Do not read or apply the Codex CLI adapter in the same handoff. This adapter requires Git, `/bin/bash`, Python 3, and an authenticated Codex CLI with dangerous bypass support. Claude Code 2.1.98+ is recommended for live progress through the Monitor tool. Stop with a compatibility error when a required runner prerequisite is unavailable. ## Research Mechanics For each research agent selected by the shared contract, resolve `../scripts/run-codex-handoff.sh` relative to this file and use the implementation launch template with `--read-only`. Give every agent separate `<agent-id>.progress.jsonl`, `<agent-id>.result.json`, and `<agent-id>.stderr.log` artifacts. Start all selected agents as background Bash tasks (`run_in_background: true`) in the same turn, then watch the wave through the implementation watcher and Monitor flow below. The runner-enforced read-only sandbox permits this launch in any host mode. Give each research agent a self-contained prompt containing the open questions, exact investigation scope, read-only boundary, relevant repository constraints, and stopping rule from the shared prompt contract. Require every field in `research-result.schema.json`; prohibit plans, design decisions, and edits. When the user has not explicitly included research agents in a model preference, select research configuration from these tiers: | Investigation | Model | Effort | Baseline timeout | | -------------------------------------- | ------------- | -------- | ---------------- | | Bounded, routine survey | `gpt-6-luna` | `high` | 10 minutes | | Involved survey across unfamiliar code | `gpt-6.1-sol` | `medium` | 15 minutes | Under this default selection, use Luna for bounded surveys and Sol for involved ones; Astra is implementation-only — research gathers evidence, the parent synthesizes. Never select `low`, `ultra`, or `max`. Research should normally use shorter budgets than implementation; keep the baseline between 10 and 15 minutes unless repository evidence says otherwise. When the research wave settles, parse each result against `research-result.schema.json`, read its stderr artifact for failure forensics, and return the findings to the shared Research Phase for the plan or research-only response. Do not reconcile the working tree. ## Plan Manifest and Configuration Use this exact host-specific table inside the shared `## Codex Handoff` plan section: ```markdown | Agent | Wave | Depends on | Scope | Model | Effort | Timeout | Implementation brief | Completion evidence | | ----- | ---- | ---------- | ------------------ | ---------------------------------------- | ----------------------- | ------------------- | ------------------------------------------------------ | ----------------------------------- | | `A1` | `1` | `none` | `<files/behavior>` | `<gpt-6-luna\|gpt-6.1-sol\|gpt-6-astra>` | `<medium\|high\|xhigh>` | `<minutes> minutes` | `<outcome, edits, constraints, and stopping criteria>` | `<commands and observable results>` | ``` When the user has not specified a model preference, select implementation configuration from these tiers: | Work | Model | Effort | Baseline timeout | | --------------------------------------------------------------------------------- | ------------- | ------------------ | ---------------- | | Bounded, routine implementation | `gpt-6-luna` | `high` | 10 minutes | | Everyday or involved implementation | `gpt-6.1-sol` | `medium` or `high` | 20 minutes | | Semantic or cross-cutting implementation | `gpt-6.1-sol` | `xhigh` | 40 minutes | | Hardest implementation: interacting invariants or difficult algorithmic reasoning | `gpt-6-astra` | `xhigh` | 40 minutes | An explicit user model preference replaces this task-complexity model selection, but effort and timeout still follow the applicable work tier. Never select `low`, `ultra`, or `max`. Adjust a timeout when repository evidence shows that required validation needs materially more or less time. The timeout is a kill-switch, not pacing: Codex never sees it and an early finish costs nothing, so size it only to bound how long a hung agent can block its wave. Keep the highest-tier agent's scope minimal and move deferrable validation to the validation owner. ## Execution Mechanics ### Launch Resolve `../scripts/run-codex-handoff.sh` to an absolute path relative to this file; never search for it in the target repository. Each invocation is one Codex agent. Without `--read-only`, the runner deliberately disables Codex approvals and sandboxing. Use that mode only after the user approves the plan and accepts that agents can read, modify, or delete any files accessible to the host account. The runner pins every Codex process to the `default` service tier, overriding inherited fast or priority selection without changing persisted Codex configuration. Before implementation wave 1, the Claude parent promotes the named draft recorded over the full manifest write-scope union during the shared Plan Phase: `ai-coord start --draft <plan-slug>` (or `ai-coord bundle start --draft <plan-slug>` for two or more Git roots). Only when promotion reports `no draft named ...`, use the plan's explicit `ai-coord start '<label>' '<path>'...` fallback (or `ai-coord bundle start '<label>' '<absolute-path>'...`) over that union. Name exact files individually and use `--recursive` only for true subtrees; require `READY` before launch. When the claim queues or blocks, run `ai-coord wait` as a background Bash task (`run_in_background: true`) so its return wakes the session, and apply the shared wake handling; never end the turn to pause. Hold that claim through reconciliation, required polish, and commit; the parent claim authorizes each delegate's assigned writes and is not a conflict. One work item per session requires the full union up front. When follow-on work expands the scope, do so only at a wave boundary: run `ai-coord done`, then start a fresh item over the enlarged union before launching the next wave. For every agent, create separate per-agent artifact paths ending in `<agent-id>.progress.jsonl`, `<agent-id>.result.json`, and `<agent-id>.stderr.log` under `${TMPDIR:-/tmp}`. Convert its approved whole-minute timeout to seconds only at the wrapper boundary, then start the runner from anywhere inside the target Git worktree as a background Bash task (`run_in_background: true`) with a description like `Codex A1/3: <scope> (<model>, <effort>, ≤<minutes>m)`: ```bash bash <skill-dir>/scripts/run-codex-handoff.sh \ --model <agent-model> \ --effort <agent-effort> \ --timeout-seconds <agent-minutes-times-60> \ --coord-identity claude/<parent-session-id> \ --progress-file <agent-progress-file> \ --result-file <agent-result-file> \ 2> <agent-stderr-file> <<'CODEX_PROMPT' <agent implementation prompt> CODEX_PROMPT ``` `--result-file` keeps structured JSON out of stdout, and redirecting stderr keeps wrapper diagnostics out of the background task display. Do not set a Bash-tool timeout; the wrapper's `--timeout-seconds` is the sole timeout authority and always terminates itself. Start sequential agents only after reconciling their dependencies. Start every agent in a parallel wave in the same turn. Pass the same `--coord-identity` on every fresh or resumed implementation launch. Research launches omit it because read-only agents make no writes and remain separately visible. Delegate prompts forbid ai-coord lifecycle commands. Delegates launched with `--coord-identity` share the orchestrating session's coordination identity. The guard now rejects a delegate's `ai-coord draft`, `ai-coord start`, `ai-coord bundle draft`, `ai-coord bundle start`, `ai-coord wait`, or `ai-coord done` with exit 64 and `lifecycle commands are not allowed from a delegate of <client>/<session>; the parent's claim covers this work`. Require every implementation prompt to say this explicitly, permitting only `ai-coord status`, `ai-coord touched`, `ai-coord inbox`, `ai-coord msg`, and `ai-coord finding`; also state that the parent's claim authorizes the assigned writes rather than conflicting with them. Delegate claims and work no longer appear as separate `ai-coord status` work rows; transient Codex thread inventory may remain visible while delegates run, and per-delegate progress lives in handoff artifacts. Add these host constraints to the shared implementation prompt: - Honor `~/.codex/rules/*.rules`, which the CLI enforces even under the bypass flag. Non-interactive runs reject `prompt`-gated commands outright. Skim existing rules and include relevant restrictions in the prompt. - Baseline command conventions: use `rg`, not `grep` variants; use `uv run python` and `uv add` or `uv run --with`, never bare Python or pip; keep Bash-only constructs inside an explicit `bash <<'EOF'` block; avoid recursive removal, worktree-destroying or history-rewriting Git, secret-reading commands, and package deploy or release scripts. - Require every field in `result.schema.json`. The wrapper passes that schema to Codex and writes the structured result to the selected artifact. ### Watch Research and implementation waves share this watcher. Read `progress-events.md` for the progress-event, sentinel, settlement, and quiet/failure contracts. Resolve `../scripts/watch-codex-wave.sh` relative to this file. Arm one Monitor per wave around one watcher invocation, passing each agent's stable ID, budget in seconds, and progress path as a repeated triple: ```sh bash <skill-dir>/scripts/watch-codex-wave.sh \ --agent A1 <budget-seconds> <A1.progress.jsonl> \ --agent A2 <budget-seconds> <A2.progress.jsonl> ``` The watcher tolerates delayed file creation and emits stable JSONL `watcher.digest`, `watcher.sentinel`, and `watcher.settlement` records. It owns elapsed time, event counts, last relevant activity, settled percentage, and the ten-cell bar. Set the Monitor `timeout_ms` above the wave's largest budget plus the 120-second no-sentinel grace. On each digest or settlement, post one short wave-status block using those exact facts. If Monitor is unavailable, run the same watcher in a foreground command; do not recreate its loop or arithmetic. Once Monitor is armed, wait for Monitor events. Do not launch Bash sleeps, tail artifacts, poll result or progress files, or add any second wait loop. Inspect artifacts only after settlement. The watcher settles an agent as failed with reason `no-sentinel` once elapsed exceeds its budget plus 120 seconds of grace. Silence is never evidence of safety buffering or model rerouting. Keep watching until the wrapper sentinel or approved timeout; never cancel, retry, extend, or downgrade because of silence. Report `no recent activity` during quiet periods. ### Collect and Reconcile When a sentinel arrives, read the result artifact and the stderr artifact for the `codex-handoff: elapsed=<seconds>s` line or failure forensics. Do not read or print background-task output; artifact-mode stdout is intentionally empty. Parse implementation results against `result.schema.json` before applying the shared reconciliation rules. At each wave boundary, if `ai-coord status` shows that a delegate narrowed the parent's claim, re-run the parent's full `ai-coord start` before continuing. This recovery applies only if a delegate bypassed the guard, for example with an unrelated identity; normally the lifecycle attempt is rejected with exit 64 in the delegate's stderr artifact. Treat timeouts, nonzero runner exits, and watcher `no-sentinel` settlements as failed settlements, not returned plan blockers. For a `handoff.failed` sentinel with reason `error`, inspect stderr first. When it evidences a transport, stream, or API death and no Codex-reported task failure, inspect partial edits with `git status` and `git diff`. Extract the session ID from the progress file's `thread.started` event and perform the shared one allowed same-agent continuation through `--resume <session-id>` with a fresh budget and a short verify-and-continue prompt naming the partially edited files. Fall back to one fresh relaunch only when no session ID is recoverable. Returned `blocked` results and timeouts are never infrastructure failures. ### Commit Delegated Work Delegated writes are attributed to the parent session, so they cannot create a delegate residual or stale-dirt baseline for those paths. Reconcile, perform required polish, and commit under the held parent claim; release it only afterward. Legacy troubleshooting: validate an unexpected residual session ID against `thread.started` and its dirt blob hashes against the reconciled diff. Ask before clearing coordination state. If matching auto-baselines exclude verified delegate edits, re-prepare with `--no-auto-baseline` and disclose it. ## Status Reporting These dashboards and the shared completion report are mandatory. Host-rendered background-task and Monitor banners are transport notifications, not status reports. Do not expose task IDs, raw JSON, sentinels, or monitor payloads. Use this legend consistently: 🔎 research · 🚀 kickoff · ⏳ running · ✅ completed · ⛔ blocked · ⏱️ timed out · 💥 runner error · 🧹 polish · 🏁 final report. Keep each update to one compact rendered block. Prefix every wave-scoped kickoff, digest, and completion update with the watcher's exact ten-cell bar, percentage, and settled counts. Progress means sentinel settlement, including failed sentinels; never infer it from elapsed time, event count, or activity. Kickoff, once per wave: ```markdown ### 🚀 Wave 1/2 [░░░░░░░░░░] 0% (0/3 settled) — 3 agents launched | Agent | Scope | Model · effort | Budget | State | | ----- | ------------------- | --------------------- | ------ | ----------- | | A1 | `internal/pricing` | `gpt-6.1-sol` · high | ≤20m | 🚀 launched | | A2 | `internal/backfill` | `gpt-6.1-sol` · high | ≤20m | 🚀 launched | | A3 | `internal/evidence` | `gpt-6.1-sol` · xhigh | ≤40m | 🚀 launched | ``` Research waves use 🔎 in their heading and investigation scopes in their rows. Wave status, on each digest or completion: ```markdown ### ⏳ Wave 1/2 [███░░░░░░░] 33% (1/3 settled) — 15m elapsed | Agent · model/effort | Status | Activity | | ---------------------- | ---------- | -------------------------- | | A1 · gpt-6.1-sol/high | ⏳ 15m/20m | ran `cargo test` | | A2 · gpt-6.1-sol/high | ✅ 8m | done — 3 files, tests pass | | A3 · gpt-6.1-sol/xhigh | ⏳ 15m/40m | no recent activity | ``` At full settlement, use the final watcher settlement record. A wave with failures still reaches 100%; its heading and rows must expose those failures. ## Completion Report Render `### 🏁 Codex handoff [██████████] 100% (<settled>/<total> settled) — <completed|blocked>`. Include strategy, agent count, and wave count, then one row per agent with result, requested model and effort, timeout budget versus actual elapsed, output tokens when available, and summary. For a resumed retry, report its sentinel's output-token total minus the prior run's total as that attempt's usage. Follow the table with `### 📦 Changed`, `### 🧪 Verification`, `### 🧹 Polish` when applicable, automatic cross-repository commit hashes when any, and `### Issues and caveats` with the shared contract's `Resolved` and `Open` groups. Omit empty issue groups and the whole section when empty; write `none` for other applicable empty values. Never expose result JSON. -
codex-cli-host.md 8.5 KB
# Codex CLI Host Adapter Load this adapter only when `SKILL.md` selects Codex's native orchestration tools. Do not read or apply the Claude Code adapter in the same handoff. Native multi-agent support is required. Stop with a compatibility blocker if the orchestration tools become unavailable. Never invoke `codex exec` or fall back to any nested CLI process. ## Native Agent Configuration When the user has not specified a model preference, use these tiers for research and implementation: | Work | Model | Effort | | --------------------------------------------------------------------------------- | ------------- | ------------------ | | Bounded research or routine implementation | `gpt-6-luna` | `high` | | Involved research or implementation | `gpt-6.1-sol` | `medium` or `high` | | Semantic or cross-cutting implementation | `gpt-6.1-sol` | `xhigh` | | Hardest implementation: interacting invariants or difficult algorithmic reasoning | `gpt-6-astra` | `xhigh` | Under this default selection, Astra at `xhigh` is the ceiling and Astra is implementation-only. Research agents use Luna or Sol — research gathers evidence, the parent synthesizes. Never select `low`, `ultra`, or `max`. Keep the highest-tier agent's scope minimal and move deferrable validation to the validation owner. Spawn every research or implementation worker with a self-contained prompt and `fork_turns: "none"`. This avoids copying the parent conversation and permits explicit `model` and `reasoning_effort` selection. Use a stable lowercase task name derived from its manifest ID and scope, and preserve the visible `R1` or `A1` ID in the prompt and report. Never exceed the active-agent concurrency limit reported by the harness. Reserve one slot for the parent and account for other active workers when reported. Split a wider manifest into dependency-preserving waves; the eight-agent shared limit is total implementation agents, not concurrent width. With no reported concurrency cap, launch one worker at a time. Codex subagents inherit the parent sandbox and approval policy: research stays under the parent's read-only controls, and implementation cannot bypass the permissions selected for the approved parent turn. For every research-only handoff, the shared prompt's strict no-edit boundary is mandatory; treat any reported edit as a contract violation. ## Research Mechanics For each selected research agent, call `spawn_agent` with `fork_turns: "none"`, the selected model and effort, and the shared self-contained research prompt. Start all agents that fit the current concurrency allowance without waiting between launches; place any remainder in a later research wave. Wait for native results with `wait_agent`. Fold the returned findings into the parent plan or research-only response per the shared Research Phase. Do not create progress, result, stderr, sentinel, or watcher artifacts. Treat any reported edit as a contract violation. ## Plan Manifest Use this exact host-specific table inside the shared `## Codex Handoff` plan section: ```markdown | Agent | Wave | Depends on | Scope | Model | Effort | Implementation brief | Completion evidence | | ----- | ---- | ---------- | ------------------ | ---------------------------------------- | ----------------------- | ------------------------------------------------------ | ----------------------------------- | | `A1` | `1` | `none` | `<files/behavior>` | `<gpt-6-luna\|gpt-6.1-sol\|gpt-6-astra>` | `<medium\|high\|xhigh>` | `<outcome, edits, constraints, and stopping criteria>` | `<commands and observable results>` | ``` Use the native configuration table above for every manifest row unless the user's explicit preference overrides its model selection. Do not add artificial timeout budgets: native agent lifetime and waiting are owned by the harness. ## Execution Mechanics The ai-coord session that performs writes owns the claim. Native Codex subagents inherit the parent session identity, so the parent owns the coordination claim for every delegated write scope. The guard now rejects a delegate's `ai-coord draft`, `ai-coord start`, `ai-coord bundle draft`, `ai-coord bundle start`, `ai-coord wait`, or `ai-coord done` with exit 64 and `lifecycle commands are not allowed from a delegate of <client>/<session>; the parent's claim covers this work`. Every worker prompt must forbid those lifecycle commands and permit only `ai-coord status`, `ai-coord touched`, `ai-coord inbox`, `ai-coord msg`, and `ai-coord finding`. Include this fact in every worker prompt so the parent's claim is treated as authorization rather than a conflict; unrelated claims on the exact assigned scope can still block work. Before implementation wave 1, the parent promotes the named draft recorded over the full manifest write-scope union during the shared Plan Phase: `ai-coord start --draft <plan-slug>` (or `ai-coord bundle start --draft <plan-slug>` for two or more Git roots). Only when promotion reports `no draft named ...`, use the plan's explicit `ai-coord start '<label>' '<path>'...` fallback (or `ai-coord bundle start '<label>' '<absolute-path>'...`) over that union; require `READY` before launch. When the claim queues or blocks, run `ai-coord wait` as a foreground command with a command timeout above its `-t` value (300 seconds by default), apply the shared wake handling, and repeat until `READY`; never end the turn between waits. After plan approval, call `spawn_agent` for each implementation worker with: - `fork_turns: "none"`; - the model and `reasoning_effort` from its approved manifest row; - a stable task name and a self-contained prompt satisfying the shared implementation prompt contract. Start all independent workers that fit the concurrency allowance without waiting between calls. Reconcile the entire wave before launching dependents. Never spawn more workers merely because a thread is quiet. Use `wait_agent` with `timeout_ms: 900000` while any agent is running; it returns early for mailbox updates, completed results, or user steering. Codex's native thread UI is the progress surface: do not reproduce it with custom dashboards, polling loops, wrapper artifacts, or synthetic percentages. Ground any concise user update in an actual agent result or harness state. Practice wait economy: when `wait_agent` returns without a settled result, an actionable mailbox message, or user steering, immediately call it again after the permitted fifteen-minute status update when one is due — no analysis, extra narration, or `list_agents` round-trips. Reserve reasoning and user-visible status for settlements, steering-worthy evidence, or that one compact update per roughly fifteen minutes of elapsed wave time; every idle wakeup otherwise costs a full model turn. Use `send_message` only to steer a currently running agent when new evidence shows it is off track or missing material context. Do not use it for routine check-ins, completed agents, or retries. ## Collection and Failure Handling Read each completed agent's final message and require every field in the shared result contract. Apply the shared scope, validation, dependency-gating, and working-tree reconciliation rules before starting the next wave. A returned `status: blocked` is a plan blocker, not an infrastructure failure. An agent-tool error or a final result missing required fields is an infrastructure failure only when the harness evidence supports that classification. After inspecting partial edits, use exactly one `followup_task` on that same agent with a short verify-and-continue prompt naming the partial files and missing evidence. Do not spawn a replacement agent. A second infrastructure failure blocks that agent and its dependents. ## Completion Report Rely on native thread rendering while work runs. At settlement, render `### 🏁 Codex handoff — <completed|blocked>` with the strategy, total agent count, and wave count. Include a compact per-agent table with model, effort, result, and summary, then `### 📦 Changed`, `### 🧪 Verification`, `### 🧹 Polish` when applicable, automatic cross-repository commit hashes when any, and `### Issues and caveats` with the shared contract's `Resolved` and `Open` groups. Omit empty issue groups and the whole section when empty; write `none` for other applicable empty values. -
progress-events.md 6.2 KB
# Progress Stream Reference When `run-codex-handoff.sh` is invoked with `--progress-file PATH`, the file is a live JSONL stream: every line Codex emits under `codex exec --json`, followed by one wrapper-authored sentinel. Pre-launch validation failures exit nonzero before the stream exists and write no sentinel; after the run starts, the wrapper writes exactly one. Tail it for real-time watching and post-mortems. Pass `--result-file PATH` separately to keep the final structured result in an artifact and leave background-task stdout empty. Sessions persist: record the `thread.started` session ID, then pass `--resume SESSION_ID` to continue that session with the same wrapper controls and a new stdin prompt. ## Codex events One JSON object per line, each with a top-level `type` ([non-interactive mode docs](https://learn.chatgpt.com/docs/non-interactive-mode)): | Event | Meaning | | -------------------------------------------------- | ------------------------------------------------------------- | | `thread.started`, `turn.started` | Session/turn lifecycle | | `turn.completed` | Turn finished; carries `usage` with `output_tokens` etc. | | `turn.failed` | Turn failed; carries error details | | `item.started` / `item.updated` / `item.completed` | Work items; `item.type` identifies the activity | | `error` | Unrecoverable stream error; the wrapper still owns settlement | Item types in Codex CLI 0.156.1: `agent_message` (assistant text), `reasoning`, `command_execution` (has `command` and `status`), `file_change`, `mcp_tool_call`, `collab_tool_call`, `web_search`, `todo_list` (plan updates), and `error` (non-fatal item error) ([0.156.1 event definitions](https://github.com/openai/codex/blob/rust-v0.156.1/codex-rs/exec/src/exec_events.rs)). Example: ```json { "type": "item.completed", "item": { "id": "item_3", "type": "agent_message", "text": "Repo contains docs and sdk." } } ``` ### Intentional visibility gap The app-server protocol documents separate `model/safetyBuffering/updated` and `model/rerouted` notifications ([turn events](https://learn.chatgpt.com/docs/app-server#turn-events)), but they are not part of the documented `codex exec --json` event set. Verified against Codex CLI 0.156.1; later versions may differ, so treat the forwarded event set as version-dependent, not guaranteed. Do not invent equivalent JSONL events or infer a safety check from silence. A quiet period may be ordinary work or transient buffering, and an independent server-side policy reroute may leave the responding model unknowable. In status digests, say `no recent activity` and keep watching until the wrapper sentinel or approved timeout. Do not cancel, retry, extend, downgrade to a suggested faster model, or relaunch because the stream is quiet; preserve normal timeout and failure handling. ## Wrapper sentinel The wrapper appends exactly one terminal line per run; its presence — not process state — is the completion signal: | Sentinel | Emitted when | | ----------------------------------------------------------------------- | -------------------------------------------------------------- | | `{"type":"handoff.completed","elapsed_seconds":N,"output_tokens":M}` | Success; last `turn.completed` value (thread-cumulative total) | | `{"type":"handoff.failed","reason":"timeout","elapsed_seconds":N}` | Wrapper timeout hit | | `{"type":"handoff.failed","reason":"error","rc":R,"elapsed_seconds":N}` | Codex nonzero exit or missing result | | `{"type":"handoff.failed","reason":"cancelled","elapsed_seconds":N}` | Wrapper received INT/TERM | The result JSON itself is in the path passed to `--result-file`, not in this progress file. Without `--result-file`, the wrapper writes the result to stdout for backward compatibility. Token accounting is best-effort; `output_tokens` is omitted when no value parses. `turn.completed` usage is the thread-cumulative total, so a `--resume` run's count includes every prior run of that thread — that run's own usage is its sentinel total minus the prior run's sentinel total. ## Wave watcher Use one bundled watcher per wave. Pass repeated agent ID, budget-seconds, and progress-file triples: ```sh bash scripts/watch-codex-wave.sh \ --agent A1 1200 /tmp/A1.progress.jsonl \ --agent A2 2400 /tmp/A2.progress.jsonl ``` Its stdout is machine-readable JSONL. `watcher.digest` reports elapsed/budget, event count, last relevant activity, and delayed-file state. `watcher.sentinel` preserves the wrapper sentinel and reason. `watcher.settlement` supplies exact settled counts, percentage, and ten-cell bar. A completed wave exits `0`; any failed agent sentinel settles normally and makes the watcher exit `1` after all agents settle. If an unsettled agent exceeds its budget plus a 120-second grace, the watcher synthesizes `{"type":"handoff.failed","reason":"no-sentinel"}` and settles it as failed; a wrapper sentinel arriving after that settlement is ignored. This only backstops a dead wrapper; the wrapper remains the timeout authority. Malformed or otherwise invalid progress emits `watcher.failed` and exits as an invariant failure, not an agent result. `watcher.digest.lastActivity` is deliberately privacy-minimal. It carries only `type` for messages, reasoning, web searches, todo lists, and item or stream errors; `type` plus `status` for file changes; `type`, `command`, and `status` for command execution; `type`, `server`, `tool`, and `status` for MCP calls; and `type`, `tool`, and `status` for collaboration calls. Missing fields are omitted. Never include message or reasoning text, search queries, arguments, results, prompts, thread IDs, agent states, or error messages. Neither a stream `error` nor an item `error` settles an agent; only the wrapper sentinel does. -
research-result.schema.json 747 B
{ "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "additionalProperties": false, "required": ["status", "findings", "open_questions", "evidence", "blockers"], "properties": { "status": { "type": "string", "enum": ["completed", "blocked"] }, "findings": { "type": "string", "minLength": 1 }, "open_questions": { "type": "array", "items": { "type": "string", "minLength": 1 } }, "evidence": { "type": "array", "items": { "type": "string", "minLength": 1 } }, "blockers": { "type": "array", "items": { "type": "string", "minLength": 1 } } } } -
result.schema.json 1.2 KB
{ "$schema": "https://json-schema.org/draft/2020-12/schema", "type": "object", "additionalProperties": false, "required": ["status", "summary", "changed_files", "verification", "residual_risks", "blockers"], "properties": { "status": { "type": "string", "enum": ["completed", "blocked"] }, "summary": { "type": "string", "minLength": 1 }, "changed_files": { "type": "array", "items": { "type": "string", "minLength": 1 } }, "verification": { "type": "array", "items": { "type": "object", "additionalProperties": false, "required": ["command", "outcome", "details"], "properties": { "command": { "type": "string", "minLength": 1 }, "outcome": { "type": "string", "enum": ["passed", "failed", "not_run"] }, "details": { "type": "string" } } } }, "residual_risks": { "type": "array", "items": { "type": "string", "minLength": 1 } }, "blockers": { "type": "array", "items": { "type": "string", "minLength": 1 } } } }
-
-
scripts
-
run-codex-handoff.sh 12.1 KB
#!/usr/bin/env bash set -euo pipefail usage() { cat <<'EOF' Usage: run-codex-handoff.sh --model MODEL --effort EFFORT --timeout-seconds SECONDS [--resume SESSION_ID] [--coord-identity CLIENT/SESSION_ID] [--read-only] [--progress-file PATH] [--result-file PATH] Read an approved implementation prompt from stdin and run one Codex implementation turn in the current Git worktree. Sessions persist; use --resume SESSION_ID to continue an existing session. With --coord-identity, the spawned Codex process runs under the specified claude or codex ai-coord session identity. With --read-only, Codex runs a read-only research session instead. With --progress-file, Codex runs with --json and streams JSONL events to PATH; the wrapper appends one terminal {"type":"handoff.completed"|"handoff.failed"} sentinel line and leaves the file in place for inspection. With --result-file, the structured result is written to PATH and stdout stays empty so background-task interfaces do not display raw JSON. Allowed models: gpt-6-astra, gpt-6.1-sol, gpt-6-luna Allowed efforts: medium, high, xhigh, max EOF } model="" effort="" timeout_seconds="" resume_session="" coord_identity="" coord_client="" coord_session="" progress_file="" result_output_file="" read_only=0 while [[ $# -gt 0 ]]; do case "$1" in --model) [[ $# -ge 2 ]] || { echo "ERROR: --model requires a value" >&2; exit 64; } model="$2" shift 2 ;; --model=*) model="${1#*=}" shift ;; --effort) [[ $# -ge 2 ]] || { echo "ERROR: --effort requires a value" >&2; exit 64; } effort="$2" shift 2 ;; --effort=*) effort="${1#*=}" shift ;; --timeout-seconds) [[ $# -ge 2 ]] || { echo "ERROR: --timeout-seconds requires a value" >&2; exit 64; } timeout_seconds="$2" shift 2 ;; --timeout-seconds=*) timeout_seconds="${1#*=}" shift ;; --resume) [[ $# -ge 2 ]] || { echo "ERROR: --resume requires a value" >&2; exit 64; } [[ -n "$2" ]] || { echo "ERROR: --resume session ID must be non-empty" >&2; exit 64; } resume_session="$2" shift 2 ;; --resume=*) resume_session="${1#*=}" [[ -n "$resume_session" ]] || { echo "ERROR: --resume session ID must be non-empty" >&2; exit 64; } shift ;; --coord-identity) [[ $# -ge 2 ]] || { echo "ERROR: --coord-identity requires a value" >&2; exit 64; } [[ -n "$2" ]] || { echo "ERROR: --coord-identity must be CLIENT/SESSION_ID" >&2; exit 64; } coord_identity="$2" shift 2 ;; --coord-identity=*) coord_identity="${1#*=}" [[ -n "$coord_identity" ]] || { echo "ERROR: --coord-identity must be CLIENT/SESSION_ID" >&2; exit 64; } shift ;; --progress-file) [[ $# -ge 2 ]] || { echo "ERROR: --progress-file requires a value" >&2; exit 64; } progress_file="$2" shift 2 ;; --progress-file=*) progress_file="${1#*=}" shift ;; --result-file) [[ $# -ge 2 ]] || { echo "ERROR: --result-file requires a value" >&2; exit 64; } result_output_file="$2" shift 2 ;; --result-file=*) result_output_file="${1#*=}" shift ;; --read-only) read_only=1 shift ;; -h | --help) usage exit 0 ;; *) echo "ERROR: unknown argument: $1" >&2 usage >&2 exit 64 ;; esac done case "$model" in gpt-6-astra | gpt-6.1-sol | gpt-6-luna) ;; *) echo "ERROR: --model must be gpt-6-astra, gpt-6.1-sol, or gpt-6-luna" >&2 exit 64 ;; esac case "$effort" in medium | high | xhigh | max) ;; *) echo "ERROR: --effort must be medium, high, xhigh, or max" >&2 exit 64 ;; esac if [[ ! "$timeout_seconds" =~ ^[1-9][0-9]*$ ]]; then echo "ERROR: --timeout-seconds must be a positive integer" >&2 exit 64 fi if [[ -n "$coord_identity" ]]; then if [[ "$coord_identity" != */* ]]; then echo "ERROR: --coord-identity must be CLIENT/SESSION_ID" >&2 exit 64 fi coord_client="${coord_identity%%/*}" coord_session="${coord_identity#*/}" case "$coord_client" in claude | codex) ;; *) echo "ERROR: --coord-identity client must be claude or codex" >&2 exit 64 ;; esac if [[ -z "$coord_session" ]]; then echo "ERROR: --coord-identity session ID must be non-empty" >&2 exit 64 fi fi script_dir="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd -P)" schema_file="$script_dir/../references/result.schema.json" if [[ $read_only -eq 1 ]]; then schema_file="$script_dir/../references/research-result.schema.json" fi [[ -f "$schema_file" ]] || { echo "ERROR: missing result schema: $schema_file" >&2; exit 66; } repo_root="$(git rev-parse --show-toplevel 2>/dev/null || true)" [[ -n "$repo_root" ]] || { echo "ERROR: run from inside a Git worktree" >&2; exit 66; } codex_bin="$(command -v codex || true)" [[ -n "$codex_bin" ]] || { echo "ERROR: codex not found in PATH" >&2; exit 69; } if ! "$codex_bin" login status >/dev/null 2>&1; then echo "ERROR: Codex CLI is not authenticated; run 'codex login'" >&2 exit 69 fi top_help="$("$codex_bin" --help 2>/dev/null || true)" exec_help="$("$codex_bin" exec --help 2>/dev/null || true)" if [[ $read_only -eq 1 ]]; then required_top_flag="--sandbox" else required_top_flag="--dangerously-bypass-approvals-and-sandbox" fi [[ "$top_help" == *"$required_top_flag"* ]] || { echo "ERROR: codex lacks $required_top_flag" >&2 exit 69 } required_flags="--color --cd --model --output-schema --output-last-message" if [[ -n "$progress_file" ]]; then required_flags="$required_flags --json" fi for required_flag in $required_flags; do if [[ "$exec_help" != *"$required_flag"* ]]; then echo "ERROR: codex exec lacks required flag: $required_flag" >&2 exit 69 fi done if [[ -n "$progress_file" ]]; then if ! : >"$progress_file" 2>/dev/null; then echo "ERROR: cannot create progress file: $progress_file" >&2 exit 66 fi fi if [[ -n "$result_output_file" ]]; then if [[ "$result_output_file" == "$progress_file" ]]; then echo "ERROR: --result-file and --progress-file must be different paths" >&2 exit 64 fi if ! : >"$result_output_file" 2>/dev/null; then echo "ERROR: cannot create result file: $result_output_file" >&2 exit 66 fi fi prompt_file="$(mktemp "${TMPDIR:-/tmp}/codex-handoff.prompt.XXXXXX")" stdout_file="$(mktemp "${TMPDIR:-/tmp}/codex-handoff.stdout.XXXXXX")" stderr_file="$(mktemp "${TMPDIR:-/tmp}/codex-handoff.stderr.XXXXXX")" result_file="$(mktemp "${TMPDIR:-/tmp}/codex-handoff.result.XXXXXX")" timeout_marker="$(mktemp "${TMPDIR:-/tmp}/codex-handoff.timeout.XXXXXX")" rm -f "$timeout_marker" codex_pid="" started_at="" sentinel_done=0 # Append exactly one terminal sentinel line to the progress file so watchers # never have to rely on process state to detect completion. emit_sentinel() { if [[ -z "$progress_file" || $sentinel_done -eq 1 ]]; then return 0 fi sentinel_done=1 printf '%s\n' "$1" >>"$progress_file" 2>/dev/null || true } elapsed_now() { if [[ -n "$started_at" ]]; then printf '%s' $((SECONDS - started_at)) else printf '0' fi } report_elapsed() { echo "codex-handoff: elapsed=$(elapsed_now)s" >&2 } # Surface what Codex was doing when it died: final messages and failures from # the JSONL stream are the only forensics available after a kill. report_last_activity() { if [[ -n "$progress_file" && -s "$progress_file" ]]; then echo "--- Codex last activity (from progress file) ---" >&2 grep -E '"type":"(turn\.failed|error|agent_message)"' "$progress_file" 2>/dev/null | tail -n 5 >&2 || true fi } cleanup() { if [[ -n "$codex_pid" ]] && kill -0 "$codex_pid" >/dev/null 2>&1; then kill -TERM "$codex_pid" >/dev/null 2>&1 || true cleanup_started_at=$SECONDS while kill -0 "$codex_pid" >/dev/null 2>&1 && ((SECONDS - cleanup_started_at < 5)); do sleep 0.2 done kill -KILL "$codex_pid" >/dev/null 2>&1 || true wait "$codex_pid" >/dev/null 2>&1 || true fi rm -f "$prompt_file" "$stdout_file" "$stderr_file" "$result_file" "$timeout_marker" } forward_signal() { signal="$1" if [[ -n "$codex_pid" ]]; then kill "-$signal" "$codex_pid" >/dev/null 2>&1 || true fi if [[ -n "$started_at" ]]; then emit_sentinel "{\"type\":\"handoff.failed\",\"reason\":\"cancelled\",\"elapsed_seconds\":$(elapsed_now)}" report_elapsed fi case "$signal" in INT) exit 130 ;; TERM) exit 143 ;; esac } trap cleanup EXIT trap 'forward_signal INT' INT trap 'forward_signal TERM' TERM cat >"$prompt_file" if [[ ! -s "$prompt_file" ]]; then echo "ERROR: empty prompt; provide the approved implementation brief on stdin" >&2 exit 64 fi if [[ $read_only -eq 1 ]]; then codex_args=(--sandbox read-only exec) else codex_args=(--dangerously-bypass-approvals-and-sandbox exec) fi codex_args+=( --color never -C "$repo_root") if [[ -n "$resume_session" ]]; then codex_args+=(resume "$resume_session") fi codex_args+=( -m "$model" -c "model_reasoning_effort=\"$effort\"" -c 'service_tier="default"' -c 'notify=[]' --output-schema "$schema_file" --output-last-message "$result_file") codex_stdout="$stdout_file" if [[ -n "$progress_file" ]]; then codex_args+=(--json) codex_stdout="$progress_file" fi codex_args+=(-) if [[ -n "$coord_identity" ]]; then AI_COORD_CLIENT="$coord_client" AI_COORD_SESSION_ID="$coord_session" \ "$codex_bin" "${codex_args[@]}" \ <"$prompt_file" >"$codex_stdout" 2>"$stderr_file" & else "$codex_bin" "${codex_args[@]}" \ <"$prompt_file" >"$codex_stdout" 2>"$stderr_file" & fi codex_pid=$! started_at=$SECONDS timed_out=0 while kill -0 "$codex_pid" >/dev/null 2>&1; do if ((SECONDS - started_at >= timeout_seconds)); then timed_out=1 : >"$timeout_marker" kill -TERM "$codex_pid" >/dev/null 2>&1 || true break fi sleep 0.2 done if [[ $timed_out -eq 1 ]]; then grace_started_at=$SECONDS while kill -0 "$codex_pid" >/dev/null 2>&1 && ((SECONDS - grace_started_at < 5)); do sleep 0.2 done kill -KILL "$codex_pid" >/dev/null 2>&1 || true fi set +e wait "$codex_pid" codex_rc=$? set -e codex_pid="" elapsed=$(elapsed_now) if [[ -f "$timeout_marker" ]]; then emit_sentinel "{\"type\":\"handoff.failed\",\"reason\":\"timeout\",\"elapsed_seconds\":$elapsed}" report_elapsed echo "ERROR: Codex timed out after ${timeout_seconds}s" >&2 report_last_activity [[ ! -s "$stderr_file" ]] || tail -n 200 "$stderr_file" >&2 exit 124 fi if [[ $codex_rc -ne 0 ]]; then emit_sentinel "{\"type\":\"handoff.failed\",\"reason\":\"error\",\"rc\":$codex_rc,\"elapsed_seconds\":$elapsed}" report_elapsed echo "ERROR: Codex exited with status $codex_rc" >&2 report_last_activity if [[ -z "$progress_file" && -s "$stdout_file" ]]; then echo "--- Codex stdout (last 200 lines) ---" >&2 tail -n 200 "$stdout_file" >&2 fi if [[ -s "$stderr_file" ]]; then echo "--- Codex stderr (last 200 lines) ---" >&2 tail -n 200 "$stderr_file" >&2 fi exit "$codex_rc" fi if [[ ! -s "$result_file" ]]; then emit_sentinel "{\"type\":\"handoff.failed\",\"reason\":\"error\",\"rc\":70,\"elapsed_seconds\":$elapsed}" report_elapsed echo "ERROR: Codex completed without a structured result" >&2 report_last_activity [[ ! -s "$stderr_file" ]] || tail -n 200 "$stderr_file" >&2 exit 70 fi if [[ -n "$result_output_file" ]]; then if ! cp "$result_file" "$result_output_file"; then emit_sentinel "{\"type\":\"handoff.failed\",\"reason\":\"error\",\"rc\":74,\"elapsed_seconds\":$elapsed}" report_elapsed echo "ERROR: cannot write result file: $result_output_file" >&2 exit 74 fi else cat "$result_file" # Codex writes the result without a trailing newline; add one so the JSON # stays on its own line when stdout and stderr share a terminal. [[ -z "$(tail -c 1 "$result_file")" ]] || echo fi # turn.completed usage is the thread-cumulative total, so the last value is the # run's final count; for resumed sessions it also includes prior runs. output_tokens="" if [[ -n "$progress_file" ]]; then output_tokens="$(grep -F '"type":"turn.completed"' "$progress_file" 2>/dev/null | grep -o '"output_tokens":[0-9]*' | tail -n 1 | cut -d: -f2 || true)" fi if [[ -n "$output_tokens" ]]; then emit_sentinel "{\"type\":\"handoff.completed\",\"elapsed_seconds\":$elapsed,\"output_tokens\":$output_tokens}" else emit_sentinel "{\"type\":\"handoff.completed\",\"elapsed_seconds\":$elapsed}" fi report_elapsed -
watch-codex-wave.sh 8.4 KB
#!/bin/bash # Watch one Codex handoff wave and emit machine-readable JSONL records. set -euo pipefail exec python3 - "$@" <<'PY' from __future__ import annotations import json import math import signal import sys import time from pathlib import Path ACTIVITY_FIELDS = { "agent_message": (), "reasoning": (), "command_execution": ("command", "status"), "file_change": ("status",), "mcp_tool_call": ("server", "tool", "status"), "collab_tool_call": ("tool", "status"), "web_search": (), "todo_list": (), "error": (), } def fail(message: str, code: int = 64) -> None: print(f"ERROR: {message}", file=sys.stderr) raise SystemExit(code) def emit(record: dict) -> None: print(json.dumps(record, separators=(",", ":"), sort_keys=True), flush=True) agents: list[dict] = [] digest_seconds = 300.0 poll_seconds = 1.0 no_sentinel_grace_seconds = 120.0 arguments = sys.argv[1:] index = 0 while index < len(arguments): argument = arguments[index] if argument == "--agent": if index + 3 >= len(arguments): fail("--agent requires ID BUDGET_SECONDS PROGRESS_FILE") agent_id, raw_budget, raw_path = arguments[index + 1 : index + 4] try: budget = float(raw_budget) except ValueError: fail(f"invalid budget for {agent_id}: {raw_budget}") if not agent_id or budget <= 0: fail("agent ID must be non-empty and budget must be positive") agents.append( { "id": agent_id, "budget": budget, "path": Path(raw_path), "offset": 0, "partial": "", "events": 0, "lastActivity": None, "settled": False, "sentinel": None, } ) index += 4 elif argument in {"--digest-seconds", "--poll-seconds", "--no-sentinel-grace-seconds"}: if index + 1 >= len(arguments): fail(f"{argument} requires a value") try: value = float(arguments[index + 1]) except ValueError: fail(f"invalid value for {argument}: {arguments[index + 1]}") if value < 0 or (value == 0 and argument != "--no-sentinel-grace-seconds"): fail(f"{argument} must be {'non-negative' if argument == '--no-sentinel-grace-seconds' else 'positive'}") if argument == "--digest-seconds": digest_seconds = value elif argument == "--poll-seconds": poll_seconds = value else: no_sentinel_grace_seconds = value index += 2 else: fail(f"unknown argument: {argument}") if not agents: fail("at least one --agent triple is required") ids = [agent["id"] for agent in agents] if len(ids) != len(set(ids)): fail("agent IDs must be unique") paths = [str(agent["path"].resolve()) for agent in agents] if len(paths) != len(set(paths)): fail("progress files must be unique") started = time.monotonic() def elapsed() -> int: return math.floor(time.monotonic() - started) def settlement() -> None: settled = sum(agent["settled"] for agent in agents) total = len(agents) filled = min(10, math.floor(10 * settled / total + 0.5)) percentage = math.floor(100 * settled / total + 0.5) emit( { "type": "watcher.settlement", "elapsedSeconds": elapsed(), "settled": settled, "total": total, "settledPercentage": percentage, "bar": "█" * filled + "░" * (10 - filled), } ) def cancel(_signum, _frame) -> None: emit({"type": "watcher.cancelled", "elapsedSeconds": elapsed(), "settled": sum(agent["settled"] for agent in agents), "total": len(agents)}) raise SystemExit(143) signal.signal(signal.SIGTERM, cancel) signal.signal(signal.SIGINT, cancel) def process_line(agent: dict, line: str) -> None: if not line.strip(): return try: event = json.loads(line) except json.JSONDecodeError as exc: emit({"type": "watcher.failed", "agentId": agent["id"], "reason": "malformed-progress", "line": line, "elapsedSeconds": elapsed()}) fail(f"{agent['id']} progress contains malformed JSON: {exc}", 65) if not isinstance(event, dict): emit({"type": "watcher.failed", "agentId": agent["id"], "reason": "non-object-progress", "elapsedSeconds": elapsed()}) fail(f"{agent['id']} progress event must be an object", 65) event_type = event.get("type") if event_type in {"handoff.completed", "handoff.failed"}: if agent["settled"]: # A real sentinel landing after the no-sentinel backstop keeps the # backstop verdict; only a second wrapper sentinel is an invariant failure. if agent["sentinel"].get("reason") == "no-sentinel": return emit({"type": "watcher.failed", "agentId": agent["id"], "reason": "duplicate-sentinel", "elapsedSeconds": elapsed()}) fail(f"{agent['id']} progress contains multiple sentinels", 65) agent["settled"] = True agent["sentinel"] = event emit( { "type": "watcher.sentinel", "agentId": agent["id"], "status": "completed" if event_type == "handoff.completed" else "failed", "reason": event.get("reason"), "elapsedSeconds": elapsed(), "budgetSeconds": agent["budget"], "eventCount": agent["events"], "sentinel": event, } ) settlement() return agent["events"] += 1 if event_type == "error": agent["lastActivity"] = {"type": "error"} return item = event.get("item") if isinstance(event.get("item"), dict) else {} item_type = item.get("type") fields = ACTIVITY_FIELDS.get(item_type) if fields is not None: activity = {"type": item_type} for field in fields: if item.get(field) is not None: activity[field] = item[field] agent["lastActivity"] = activity def read_new(agent: dict) -> None: path = agent["path"] if not path.exists(): return try: with path.open("r", encoding="utf-8") as handle: handle.seek(agent["offset"]) chunk = handle.read() agent["offset"] = handle.tell() except OSError as exc: emit({"type": "watcher.failed", "agentId": agent["id"], "reason": "unreadable-progress", "elapsedSeconds": elapsed()}) fail(f"cannot read {path}: {exc}", 66) text = agent["partial"] + chunk lines = text.split("\n") agent["partial"] = lines.pop() for line in lines: process_line(agent, line) next_digest = started + digest_seconds while not all(agent["settled"] for agent in agents): for agent in agents: read_new(agent) now = time.monotonic() for agent in agents: if agent["settled"] or now - started <= agent["budget"] + no_sentinel_grace_seconds: continue sentinel = {"type": "handoff.failed", "reason": "no-sentinel"} agent["settled"] = True agent["sentinel"] = sentinel emit( { "type": "watcher.sentinel", "agentId": agent["id"], "status": "failed", "reason": "no-sentinel", "elapsedSeconds": elapsed(), "budgetSeconds": agent["budget"], "eventCount": agent["events"], "sentinel": sentinel, } ) settlement() if now >= next_digest and not all(agent["settled"] for agent in agents): for agent in agents: if agent["settled"]: continue emit( { "type": "watcher.digest", "agentId": agent["id"], "elapsedSeconds": elapsed(), "budgetSeconds": agent["budget"], "eventCount": agent["events"], "lastActivity": agent["lastActivity"], "noRecentActivity": agent["lastActivity"] is None, "progressFileExists": agent["path"].exists(), } ) while next_digest <= now: next_digest += digest_seconds if not all(agent["settled"] for agent in agents): time.sleep(poll_seconds) raise SystemExit(1 if any(agent["sentinel"].get("type") == "handoff.failed" for agent in agents) else 0) PY
-
-
SKILL.md 23.8 KB
--- argument-hint: "[task]" compatibility: The Claude Code host requires Git, /bin/bash, Python 3, and an authenticated Codex CLI with dangerous bypass support; the Codex CLI host requires native subagents. metadata: install-targets: claude-code codex name: codex-handoff skill-dependencies: - agents-brain - code-polish - commit description: Orchestrate read-only Codex research in any mode, or one to eight Codex agents to implement approved plans from Claude Code or Codex CLI. --- # Codex Handoff Codex-handoff orchestrates read-only investigation or implementation within the current session after explicit plan approval. Task-handoff instead writes a decision-complete file for a fresh, separate session; use it when work continues later or elsewhere, and use an in-session handoff skill to implement an approved plan now. If these instructions are already present in the conversation from a slash or dollar invocation, follow them directly; do not invoke this skill again through a skill tool. Follow the shared contract below, select exactly one host adapter, and use it for every host-specific action. ## Host Selection Inspect the callable orchestration tools, not environment variables, process ancestry, or a user-supplied host name: - `spawn_agent`, `wait_agent`, `send_message`, `followup_task` available: read `references/codex-cli-host.md` completely. - Otherwise, Claude Code's Agent and Bash tools available: read `references/claude-code-host.md` completely. - Neither: stop with a compatibility error. Native Codex multi-agent support is mandatory on the Codex host; never fall back to a nested Codex CLI process. Use exactly one adapter for every host-specific action; never load both or combine their launch, progress, retry, permission, or result-transport mechanics. The adapter may specialize host mechanics and manifest configuration but cannot weaken this shared contract. ## Contract - Run only after explicit invocation. Research-only = requested outcome is findings/evidence/assessment only, no repo changes or plan requested. All handoffs may run in any host mode; implementation handoffs must pass through the Plan Phase and receive explicit user approval before launch. - Reuse an already approved plan when the outcome and material constraints are unchanged. Explicit user instructions take precedence over skill defaults; ask again only for an unresolved decision or action outside that authorization. - The parent owns decisions, the final plan, and orchestration. Delegate investigation to read-only research agents only when task and repository evidence make it useful before planning. - Research agents gather evidence and report findings only — never edit files, make design decisions, or return plans. - Implementation agents implement their assigned part of the approved plan: inspect, edit, validate, never redesign or return another plan. - Use the smallest effective implementation team (one agent is valid); add agents only when decomposition materially improves latency, correctness, or verification. Never exceed eight implementation agents total. - Size every brief before finalizing the team: estimate its wall-clock time and split any brief likely to exceed roughly 25-30 minutes into parallel disjoint scopes or dependency waves — add an integration agent if needed — instead of one monolithic agent. - Use at most three research agents, stable IDs `R1`-`R3`, counted separately from the eight implementation agents. - Keep the parent's own implementation work to orchestration, integrity checks, failure handling, and conditional polish passes. - Treat an explicit user model preference (e.g. GPT-6 Luna) as an orchestration constraint on every research and implementation agent unless scoped narrower; don't substitute the adapter's usual Luna/Sol/Astra selection. If the host can't launch that model, report the incompatibility and ask before falling back. - Treat the approved outcome — not the initial manifest or its write scopes — as the authorization boundary: when implementation reveals a related in-repository fix or evidence change required for that outcome, the parent may extend the handoff and launch follow-on agents for the newly discovered scope without asking again. The worker that discovered the need still stops at its assigned scope and returns evidence; the parent owns scope expansion, repository coordination, and delegation. - Size verification to the requested outcome. Never add validation machinery (gates, manifests, checkpoints, hash pins, journals, receipts) unless the approved plan explicitly calls for it. An explicit user request to hurry or wrap up overrides optional repeat checks and required polish passes: commit the validated work and report what was skipped or left unverified. Use `$ARGUMENTS` as the task when present; otherwise use the active user request. A task naming another skill follows Companion Skills. ## Companion Skills A task that names another skill alongside this handoff — for example, `run $fresh-eyes-sweep and delegate the implementation` — is a composition: the companion skill defines the work; this handoff owns delegation mechanics. - Load the companion skill and run its discovery, judgment, and planning phases in the parent. Route read-only discovery through this skill's research agents when materially faster, then fold the companion's method, findings, and constraints into the handoff plan. A companion missing from the host's skill list may be installed but hidden by `disable-model-invocation: true`: read its `SKILL.md` directly from the host skill root (`~/.claude/skills/<name>/SKILL.md` in Claude Code, `~/.agents/skills/<name>/SKILL.md` in Codex) before concluding it is absent. - When the companion's discovery is itself the bulk of the work — an audit or sweep over a whole repository or large file set — the parent maps the scope and slices it instead of reading it inline. After plan approval, each implementation agent audits and fixes its own slice under the companion's rules, inlined in its brief. - This contract overrides the companion's overlapping plan approval, agent limits and stable IDs, single validation owner, result fields, failure classification, commit ownership, and completion reporting, even when the companion prescribes its own subagent, validation, or commit mechanics. Its user-decision gates still bind; this skill's plan approval satisfies coincident gates, and a plan-only companion implements only when the combined invocation requested implementation and the user approved the plan. - Agents cannot load skills. Never brief one to "use skill X" by name; inline the specific companion instructions, conventions, or excerpts it needs. - A companion requirement to run `$code-polish` or `$agents-brain polish` marks that pass required in the Plan Phase; it still runs once, per Completion. Companion commit instructions never add commits. - Satisfy both contracts at Completion: produce the companion's required report artifacts (ledgers, tables, verdicts) alongside the selected adapter's completion report. ## Research Phase For a research-only task, launch one to three research agents and stop after returning the consolidated investigation; never enter the Plan Phase or launch implementation agents. For an implementation handoff, trigger research when scope is uncertain, the task crosses multiple or unfamiliar subsystems, or gathering the needed evidence serially would be materially slower for the parent. Zero research agents is the default for implementation handoffs. The parent alone decides the research count from task and repository evidence; never ask the user to opt in or name agents. Either launch the research wave immediately or proceed straight to planning. When triggered, assign up to three agents stable IDs `R1`-`R3` and launch them immediately through the selected adapter's read-only mechanism. Give each agent a self-contained prompt containing: - the open questions and exact investigation scope; - relevant repository constraints and known concurrent-work boundaries; - a strict read-only authority boundary; - the stopping rule that it must return evidence rather than a plan or design; - its time budget, with the instruction to use it: an early `blocked` return citing only time is not a valid stop; and - exact result fields: `status`, `findings`, `open_questions`, `evidence`, and `blockers`. A research agent that returns `blocked` citing only its time budget while most of that budget is unused and no concrete obstacle is named has not settled its scope, and neither has one that returns `completed` while its evidence still lists uninspected scope paths: continue the same agent once through the adapter's same-agent mechanism with the uncovered files and the remaining budget. This continuation is not a new research agent. When every required research agent settles, fold its findings and evidence into the implementation plan or the research-only response. Surface open questions or blockers through the host's user-question mechanism only when they change scope or approach. Do not reconcile the working tree — research agents change nothing; any reported edit is a contract violation. For a research-only task, synthesize the evidence and finish with `### 🔎 Research handoff — <completed|blocked>`, the agent count, findings, evidence, open questions, and blockers. This replaces the Plan Phase and the selected adapter's implementation completion report. If the investigation shows that changes are needed, report them as findings and stop — do not produce an implementation plan or begin edits. ## Plan Phase Enter this phase only for an implementation handoff. When research — delegated or the parent's own — contradicts a fact the user stated explicitly (quantities, which items, which accounts), ask through the host's user-question mechanism before writing the plan; never widen the plan's default scope to fit the research. Produce a decision-complete plan with this section and the selected adapter's exact manifest table: ```markdown ## Codex Handoff - Research: `<none | R1..Rn — key findings used>` - Companion skills: `<none | $x — phases the parent runs / what is folded into briefs>` - Strategy: `<sequential|parallel|hybrid>` - Agents: `<1-8>` — `<why this is the smallest effective count>` - Validation owner: `<agent-id|parent>` — `<aggregate checks it runs once>` <host-adapter manifest table> - Code polish: `<required|not required>` — `<reason>` - Agent-context polish: `<required|not required>` — `<reason>` ``` While finalizing the plan, record the union of every manifest write scope with `ai-coord draft --name <plan-slug> '<label>' '<path>'...`, using `--recursive` only for directory scopes. Derive `<plan-slug>` from the plan's short identity (for example `debarrel-lib`): lowercase it, replace characters outside `[A-Za-z0-9._-]` with `-`, and truncate to 40 characters. For two or more Git roots, use `ai-coord bundle draft --name <plan-slug> '<label>' '<absolute-path>'...`. In the plan's "Wait out conflicting agents" section, write the exact promote command `ai-coord start --draft <plan-slug>` (or `ai-coord bundle start --draft <plan-slug>`) and retain the explicit `ai-coord start '<label>' '<path>'...` fallback (or `ai-coord bundle start '<label>' '<absolute-path>'...`) over the same union. A fresh implementation session must receive these commands in the plan itself, without reconstructing scopes from prose. Use the fallback only when promotion reports `no draft named ...`. Named drafts grant no authority and expire after seven days. Choose the execution shape from repository evidence and the approved work: - Sequential: one agent depends on another, write scopes overlap, or a later agent owns integration or aggregate validation. - Parallel: independent work only, with explicitly disjoint write scopes. Agents may inspect shared context but must not write outside their assigned scope. - Hybrid: dependency-ordered waves — run independent agents within a wave in parallel, reconcile the entire wave, then start its dependents. A wave finishes with its slowest agent. Keep the highest-tier agent's scope minimal and move deferrable validation to the validation owner. If parallel work does not collectively prove the overall plan, reserve a later sequential agent for integration and aggregate validation. Use stable agent IDs and explicit dependencies across the whole handoff. Assign aggregate validation to exactly one owner: package- or repository-wide checks run once, by the integration agent when one exists, otherwise by the parent during post-wave reconciliation. Every other agent runs only the narrowest checks that prove its own edits, such as file-scoped formatting, lint, or typecheck plus targeted tests. Require `$code-polish` for nonlocal invariants, concurrency or state machines, migrations or parsing, auth or security, retry or error semantics, and public API or data-contract changes. File count alone is not a trigger. Require `$agents-brain polish` when approved work changes a target its polish workflow supports: README.md, AGENTS.md or CLAUDE.md, a durable context doc, an existing project-installed skill under `.agents/skills`, or an existing git-tracked source-catalog skill under `skills/` where polish is prose-only. Installed copies under managed agent-config roots remain excluded. Mark both passes required when both trigger rules apply; mark neither when neither applies. Do not launch implementation agents until the user approves the plan. Read-only research is the only pre-approval exception. ## Implementation Prompt Contract Build a self-contained, outcome-first prompt for every implementation agent. Include: 1. The approved overall outcome plus the agent's implementation brief, dependencies, and completion evidence. 2. Its exact write scope, relevant repository constraints, known dirty-work boundaries, and prerequisite agent results. 3. Its validation assignment per the Plan Phase's single validation owner: scoped checks it must run and, unless it owns validation, that it must not run aggregate checks. Never brief new validation machinery the approved plan does not call for. 4. A pacing estimate matching its manifest sizing and any user- or adapter-imposed hard runtime limit. A soft estimate alone is not a stop condition; report partial evidence and the concrete blocker or exhausted hard limit when blocked. 5. This authority boundary: inspect, edit only within the assigned scope, and validate locally; never commit, push, deploy, make external writes, or broaden scope, even when repository or host instructions favor committing finished work promptly. Committing stays with the parent after reconciliation. 6. The selected adapter's delegation and coordination context, including why the parent session and disjoint siblings are not conflicting work and what unrelated exact-scope claim would justify returning `blocked`. Delegates must not run coordination lifecycle commands: these are rejected with exit 64. Permit only `ai-coord status`, `ai-coord touched`, `ai-coord inbox`, `ai-coord msg`, and `ai-coord finding`. 7. This stopping rule: implement the approved plan exactly; if infeasible or requiring redesign, return `blocked` with evidence instead of proposing a replacement plan. Continue after progress updates while authorized work remains; a milestone, offer to continue, or list of nonblocking decisions is not a completed result. 8. A requirement to return every result field: `status` (`completed` or `blocked`), `summary`, `changed_files` listing only files actually touched, `verification` listing every command and outcome, `residual_risks`, and `blockers`. Keep each agent prompt as compact as completeness allows: the shared outcome summary plus that agent's own brief, scope, and constraints — never restate the full plan text per agent. Add the selected adapter's command, permission, transport, and host-tool constraints without restating this contract. ## Execution and Reconciliation Before implementation wave 1, the parent promotes the plan's named draft to acquire the full manifest write-scope union. Use the plan's recorded explicit start fallback only when promotion reports `no draft named ...`; require `READY` before launching agents. When promotion or start queues or blocks instead, never end the turn to pause: run `ai-coord wait` through the adapter's wait mechanics. It also returns on non-readiness wake events (message, unknown coverage, work release, 300-second default timeout); on each wake, handle `MESSAGE` events through `ai-coord inbox`, re-submit the recorded promote or start command, and diagnose stale blockers. Launch agents through the selected adapter in the approved strategy and dependency waves. Do not add agents or change models, efforts, scopes, or validation ownership merely because a worker is slow or quiet. Do revise the manifest and launch a narrowly scoped follow-on agent when completed work discovers an unplanned prerequisite covered by the approved outcome. Preserve stable IDs, dependency order, the eight-agent limit, and one aggregate-validation owner; include follow-on agents in the final counts and report. For each completed agent: require every shared result field and treat `changed_files` as its authoritative post-pass scope; confirm reported files exist or were intentionally deleted, stay within scope, and carry verification evidence matching the assignment; and pass relevant completed results to dependent agents. After every implementation wave, reconcile all results with the current manifest and visible working tree without folding in unrelated concurrent changes. When the parent owns validation, run the assigned aggregate checks once during this reconciliation. Attribute aggregate-check failures before blocking: first rule out effects of the handoff's changes, formatters, hooks, and generators, including failures in downstream files outside its write scopes. Continue past a failure only when evidence establishes that it is unrelated and the handoff's own checks still pass. Unexpected out-of-scope edits, same-wave overlap, or a failure attributable to the handoff are blockers; do not start dependents or polish, and do not silently take over implementation. ## Failure Classification - A `status: blocked` result identifying a related in-repository fix or evidence change outside the worker's scope that is necessary for the approved outcome is follow-on work under the Contract's scope-expansion authority, not a request for fresh authorization: let already-started independent agents finish, gate dependents, extend the manifest with the smallest sufficient scope, satisfy repository coordination for that scope, and launch a new or reused implementation agent. Repeat until the outcome is complete or a genuine authorization boundary is reached. - Ask the user only when continuation would change the approved outcome, require a material redesign or unrelated work, or cross an existing confirmation boundary such as destructive action, purchase, deployment, or external write. Never silently take over implementation or relaunch solely on a larger model. - Treat a tool or infrastructure failure as retryable only when adapter-specific evidence supports that classification. Inspect partial edits first, then use the adapter's same-agent mechanism for exactly one verify-and-continue attempt; this is not a new agent against the eight-agent limit. A second infrastructure failure blocks that agent and its dependents. - Never classify an ordinary timeout, a returned blocker, silence, or task-level validation failure as infrastructure failure. Continue only work proven independent. ## Skill Evolution Review Keep verified repairs to skills used during the handoff separate from the optional review below. When user or repository instructions already authorize repairs, the parent owns their completion; subagents report evidence without expanding their write scopes. One verified occurrence is enough, and a blocked main task does not prevent independent repairs. Complete the handoff's required work or establish its blocker, then finish independent repairs before the final report under the applicable maintenance policy. Plan Mode still prohibits edits. After every required agent succeeds and the task is verified — never for a blocked, failed, or partial handoff — the parent alone judges skill-evolution opportunities; agents never make the recommendation. Recommend only a stable, reusable workflow credibly likely to recur; reject one-offs, rare contingencies, incidental cleanup, and speculative value. For a new skill, state repo-local vs. global `~/projects/agent-skills` placement; for a revision, name the exact skills and why. When a proposal clears this bar, append at most one compact suggestion (≤2 short sentences) to the adapter's existing completion report without changing its format, then offer `$task-handoff` as the next action. Never auto-invoke `$task-handoff` or create/revise anything during this review; when nothing qualifies, stay silent — no placeholder. ## Completion - After every required agent completes, deduplicate the union of reported `changed_files` and confirm the combined verification evidence proves the approved plan. - Before the completion report, fix remaining same-pattern sites the approved outcome covers through follow-on agents per Failure Classification; never list them as optional or out-of-scope items. - If any required agent failed, or the user explicitly asked to hurry or wrap up, skip every planned polish pass and report the skip. Otherwise, invoke each required pass once with only its applicable paths from that union: `$code-polish` first in its default simplify-then-review mode, then `$agents-brain polish` with its eligible context targets. Invoke only one when only one is required. Do not seed either pass with paths outside the union or let it broaden beyond its declared workflow authority. - Reconcile in-scope files actually changed by each polish pass into the final changed-files set and verification. A required polish pass that blocks, fails, or writes outside its supported scope blocks later polish and cross-repository commits. - If approved work changes repositories on this machine other than the one where the handoff began, invoke `$commit` from each additional repository after its work, validation, and required polish complete, scoped to files changed there; do not commit incomplete, blocked, unexpected, or out-of-scope changes. Push when the request or standing user instructions authorize it. - When the handoff pushed commits and the repository defines CI workflows, such as `.github/workflows`, watch the pushed head's runs before the completion report (`gh run list --commit <sha>`, then `gh run watch <run-id>`, in the background when the host supports it). Fix failures attributable to the handoff as follow-on work and report the CI outcome. When changed code behaves differently by platform and local checks covered only one, name the unverified platforms as a risk. - Finish with the selected adapter's completion report, including strategy, wave and agent counts, each agent's requested configuration, status, and summary, plus combined changed files, verification, polish when run (listing each pass and outcome), automatic cross-repository commit hashes when any, and `Issues and caveats` when present. Write `none` for other applicable empty values and never expose machine result payloads. - Group issues and caveats as `Resolved` (verified fixes with evidence) and `Open` (remaining problems, limitations, or unverified assumptions, with impact and next step). Omit empty groups and the whole section when empty. Report each item once; put neutral context and agreed decisions under changes or scope. Reserve `blocker` for something preventing required work and `risk` for a specific potential adverse outcome. A workaround leaves an item open when the underlying issue still affects the result.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.