codex-delegate
Delegate a coding task to the OpenAI Codex CLI as a background implementer, then review its diff and land it yourself. Use this whenever the user wants to hand implementation work to Codex — phrasings like "have Codex do X", "delegate this to Codex", "run it through Codex", or "u
Install
npx skills add https://github.com/amElnagdy/delegate-skills/tree/master/skills/codex-delegate
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install amelnagdy-delegate-skills@llmmart
git clone https://github.com/amElnagdy/delegate-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole amelnagdy/delegate-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Codex Delegate
You are the orchestrator. This skill lets you hand a bounded coding task to a separate implementer — the OpenAI Codex CLI — then review what it produced and land it yourself. You write the brief and own the judgment; Codex does the typing in its own sandbox; you verify and commit.
Nothing here is specific to one orchestrating agent. The loop needs only the ability to run a shell command and read a file, so it works the same whether you are Claude Code, OpenCode with a selected model, or any comparable agent. (It is designed for and run on Claude Code; treat other orchestrators as designed-for, not yet proven.)
When NOT to use this
- The task is small enough to just do inline — delegation overhead is not worth it.
- The
codexCLI is not installed or not authenticated (runcodex login). - You want to write the code yourself, or you only need a review (use Codex's own
reviewcommand).
Prerequisites (check once)
codex --versionsucceeds. If not, install (npm i -g @openai/codex) andcodex login.- Confirm which
codexis on PATH. Multiple installs are common (e.g. a current npm/nvm copy and a stale Homebrew one).command -v codexshows the active one andcodex --versionits version — an old binary predates flags this skill relies on (codex exec --json,-o,exec resume). The relay also records the version it ran intoresult.json, so a stale binary is visible after the fact. - You are in (or will point
--cdat) the target git repository.
The loop
Run these five steps per task. Steps 1, 4, and 5 are your judgment; 2 and 3 are mechanical.
1. Write the brief
Codex sees only the text you send — no repo memory, no chat history, no shared context. Everything the task needs goes in the brief: the goal, the current state, what to change, what to leave untouched, the project's actual gate commands (discover them from the repo's CLAUDE.md/AGENTS.md/Makefile — do not assume), and a report contract. Tell Codex it will not commit (you will). Keep one task per brief. Full guidance and a template: references/writing-the-brief.md.
2. Dispatch
Send the brief to Codex with the bundled helper. It wraps codex exec, captures the run, and writes a
structured result.json — so your only job is "run a command, read a file." (<skill-dir> below is
this skill's installed directory — the folder containing this SKILL.md, i.e. the directory you loaded
the skill from. Claude Code prints it as "Base directory for this skill" when the skill loads; on other
orchestrators use that same directory — if unsure where it landed, run
find ~ -name relay.mjs -path '*codex-delegate*' and substitute the directory above it.)
node "<skill-dir>/scripts/relay.mjs" --brief brief.txt --cd /path/to/repo
# read-only (review/diagnosis, no edits): add --read-only
# isolated review (skip ambient MCP/user config): add --ignore-user-config
# continue the exact Codex session: add --session <threadId> (from result.json; send only the delta brief)
# fallback when no thread id is available: add --resume-last
# hard time limit (watchdog): add --timeout 2h (default: off; implementation runs routinely need 1-2h)
# see all options: node .../relay.mjs --help
The helper defaults to a write-capable (workspace-write) sandbox and writes its artifacts to a temp
dir, so the repo under review stays clean. It never commits — see step 5. Mechanics, flags, and the
result.json shape: references/dispatch-and-poll.md.
3. Wait for completion
The helper blocks until Codex finishes, so back it with whatever your orchestrator offers and resume when it returns:
- Claude Code: run the Bash call with
run_in_background: true; you are notified on completion. - Plain shell / other agents: run it in the foreground for short tasks, or background it and poll
the result file —
… &in bash/zsh (including Git Bash/WSL), or your shell's equivalent (Start-Jobin PowerShell,start /bin cmd). The run is done whenresult.jsonexists with astatus. (A pre-run usage error — bad args or an empty brief — instead exits with code 2 and a stderr message and writes no result file, so check the exit code too. A missingcodexbinary exits 127 but does write aresult.jsonwith statuscodex_unavailable.)
Do not trust progress trackers over reality: a run is finished when result.json is written and the
process has exited. Read the working tree, not a status line. The implementer's full report is
the finalMessage field in result.json (also printed in full on stdout between the report markers).
4. Review — do not trust the self-report
Codex's result.json includes its own summary and gate claims. Re-verify, don't accept:
- Re-run the project's gates yourself (the test/lint/build commands from step 1). Never take "gates passed" on faith.
- Read the diff against the brief: did Codex do what was asked, nothing more (scope creep) and
nothing less?
touchedFilesin the result is your starting point. - Run the relevant guard skills on the diff if you have them installed (clean-code-guard,
test-guard, etc. from
guard-skills) — this skill produces the work; those skills judge it. - For schema/migration changes, round-trip them; for removals, grep for dangling references.
Full checklist: references/review-and-land.md.
5. Land it
Because Codex's sandbox cannot reliably write .git (it varies by version, OS, and path), the
orchestrator commits. Only after the gates pass and the diff holds:
- Commit the verified work yourself, with a clear message.
- If it needs changes, send a delta brief with
--session <threadId>from the priorresult.json(use--resume-lastonly when no thread id is available), and review again.
Read-only second opinions
The relay doubles as a clean way to get an adversarial second opinion with no write risk: dispatch
--read-only with a brief that lists the agreed points, then each contested point with both
positions, and ask Codex to defend or concede each — deliverable in its final message, touching no
files. Any delegation skill whose implementer offers a read-only mode supports the same use, but
check how hard that mode's guarantee is first: Codex's sandbox enforces it, while Grok's is
best-effort and only flagged after the fact (readOnlyViolation) — for those implementers,
verify touchedFiles came back empty instead of assuming no edits.
Authorization model
Delegation is something the human opts into. Once they have ("run this queue", "proceed"), committing verified, gate-passing work is the agreed contract — that is the whole point. Two limits on that mandate: surface, don't absorb (report Codex's design decisions, defensible-but-unasked turns, and non-blocking nitpicks rather than silently keeping them) and stop for scope changes (if correct completion needs going beyond the brief, ask — don't expand the mandate yourself). The full treatment is in references/review-and-land.md.
If you have the openai-codex plugin
The official openai-codex Claude Code plugin is excellent and complementary — codex-delegate
builds on the same codex CLI, it doesn't replace the plugin. They point in different directions:
- The plugin's
codex:codex-rescueagent is a forwarder: it hands one task to Codex and returns the output. It deliberately does not poll, review, or commit. - The plugin's review command and stop-review gate run the inverse direction: Codex reviews your work.
codex-delegateis the orchestration loop in the other direction: you drive Codex to implement across one task or a queue, and you review and land each result. That loop — brief → dispatch → poll → review → commit, with the orchestrator owning the commit — is what the plugin leaves to you, and what this skill encodes.
If you have the plugin installed, its companion CLI is an optional alternative dispatch backend; the
bundled relay.mjs is the default because it adds no install of its own beyond the codex binary
(Node and git, which the relay also needs, are prerequisites for every skill here).
References
- references/writing-the-brief.md — how to write a brief Codex can execute blind: structure, XML blocks, the report contract, embedding the real gate commands.
- references/dispatch-and-poll.md —
relay.mjsflags, theresult.jsoncontract, backgrounding per orchestrator, and recovery when a run misbehaves. - references/review-and-land.md — the review checklist, the commit boundary, and the exact-session rework cycle.
- references/multi-task-queues.md — running a sequential queue: carrying constraints forward, progress tracking, and the end-of-run coherence check.
Files (delegate-skills)
-
references
-
dispatch-and-poll.md 12.3 KB
# Dispatch and poll `scripts/relay.mjs` is the dispatch layer. It wraps `codex exec`, runs the brief in a sandbox, captures everything, and writes a structured `result.json`. Your job collapses to: run one command, then read one file. Everything Codex-specific lives in the helper, which is what keeps the loop portable across orchestrators. ## Before the first run: check the binary Two gotchas, both worth 30 seconds: ```bash command -v codex # the active binary; a stale install (e.g. Homebrew) can shadow a current one codex --version # an old binary predates `exec --json`, `-o`, and `exec resume` codex login status # must be authenticated ``` The Codex CLI moves fast and behavior shifts between versions, so the helper records the version it actually ran into `result.json` — if something behaves oddly, check which binary answered. ## Dispatching ```bash node "<skill-dir>/scripts/relay.mjs" --brief brief.txt --cd /path/to/repo ``` (`<skill-dir>` is wherever this skill is installed — the folder containing its `SKILL.md`. On Claude Code it's the printed "Base directory for this skill"; on other orchestrators substitute that install path. See [`SKILL.md`](../SKILL.md) if you need to locate it.) Options: | Flag | Effect | | --- | --- | | `--brief <file>` | The brief. Omit it to read the brief from stdin (`node relay.mjs … < brief.txt`). | | `--cd <dir>` | Working root for Codex (default: current directory). | | `--lane <name>` | Fleet lane from `delegate-setup` config. Applies that lane's dials; fails if the lane's `implementer` is not this relay. Explicit dial flags win. | | `--model <name>` | Codex model (default: Codex's own configured default). | | `--effort <level>` | Reasoning effort, passed to Codex as `-c model_reasoning_effort=<level>` (default: Codex's own configured default). The relay accepts a bare token; Codex and the model own the supported levels. Applies to fresh and resumed runs. | | `--sandbox <mode>` | `read-only` \| `workspace-write` \| `danger-full-access` (default: `workspace-write`). | | `--read-only` | Shortcut for `--sandbox read-only` — review/diagnosis with no edits. | | `--resume-last` | Continue the most recent Codex session; send only the delta brief (see review-and-land). "Most recent" is global, so an unrelated Codex run can steal it — prefer `--session`. | | `--session <id>` | Continue one specific thread by id (the `threadId` from a prior `result.json`); send only the delta brief. Mutually exclusive with `--resume-last`; an empty id is rejected. | | `--clean-env` | Pass only runtime basics (`PATH`, home, locale, temp, `CODEX_HOME`, and Windows equivalents) to Codex and its version preflight. This changes inherited variables only; it does not protect files or other same-user secrets. | | `--keep-env <name>` | Keep one additional variable under `--clean-env`; repeat for each required environment-backed auth, custom-provider credential, proxy, certificate, or MCP variable. The name must be set and use portable environment-variable syntax. | | `--skip-git-repo-check` | Allow running outside a git repo. | | `--timeout <dur>` | Relay-side watchdog (e.g. `30m`, `2h`); on expiry the child is killed and `result.json` gets `status: "timeout"`. Off by default. | | `--out-dir <dir>` | Where artifacts go (default: a fresh dir under the system temp dir). | Artifacts default to the system temp dir on purpose: the repo under review stays clean, so the touched-files report shows only Codex's edits and nothing of the helper's own. `--clean-env` is not a broader security boundary: Codex can still access files and other same-user secrets available through `HOME`, `CODEX_HOME`, OS facilities, and the selected sandbox. File- or OS-backed auth and normal configuration still load, but direct environment-backed auth (`CODEX_API_KEY` or `CODEX_ACCESS_TOKEN`) needs that variable named with `--keep-env`. `OPENAI_API_KEY` can still matter as a custom-provider credential; provider, proxy, certificate, or MCP settings that reference any stripped variable likewise need it named with `--keep-env`. The same filtered environment is used for preflight and dispatch. On Windows, a sandboxed run (`read-only` or `workspace-write`) receives a `PATH` with every `WindowsApps` entry removed, again for both preflight and dispatch. The Microsoft Store installs apps — PowerShell 7 included — under a folder whose ACLs deny execution to the restricted token Codex's sandbox runs commands with. Codex prefers a `pwsh` found on `PATH`, so a Store-installed `pwsh` makes every sandboxed command fail with `0xC0070005` while the model keeps working blind: it cannot run a single gate and either gives up or reports work it never verified. Without those entries Codex falls back to System32 `powershell.exe`, which the sandbox can launch. `--sandbox danger-full-access` runs without the restricted token and gets `PATH` unchanged. To see whether a machine is affected: `(Get-Command pwsh -All).Source` — any result under a `WindowsApps` folder would hit this. ## The result `<out-dir>/result.json` is the contract. Fields: - `schema` — the result-format version (currently `delegate-relay.result.v1`) - `status` — `completed` | `failed` | `timeout` | `aborted` | `codex_unavailable` - `exitCode` — mirrors Codex's exit code; `128` plus the signal number if the child was killed; `127` if `codex` isn't on PATH; on a `timeout` the relay forces a non-zero code even when the child exited `0` after the watchdog's SIGTERM - `signal` — the signal that killed the child, otherwise `null` - `codexVersion` — the binary that actually ran - `threadId` — feed this to a later `--session <id>` (exact thread; preferred) or `--resume-last` (global "most recent", which another Codex run can steal) - `finalMessage` — Codex's own final report (the `<structured_output_contract>` you asked for) - `touchedFiles` — `git status --porcelain` lines in the working root: your review starting point. `null` (not `[]`) when git can't report — `git` missing, or a non-repo run under `--skip-git-repo-check`; `[]` means git ran and the tree is clean - `briefPath` / `eventsPath` / `finalPath` — the exact brief relay sent, the raw JSONL event stream, and the final-message file - `stderrPath` — the complete stderr stream in `stderr.txt`, including diagnostics omitted from the summary - `workdir`, `sandbox`, `model`, `effort`, `resumeLast`, `session`, `cleanEnv`, `keepEnv`, `startedAt`, `finishedAt` — `sandbox` is the applied mode, or a note that Codex used its active config on an unqualified resume; `session` is the explicit session id, or `null` for fresh and `--resume-last` runs; `keepEnv` records names only, never values - `stderrTail` — up to 20 nonblank lines from the final 64 KiB of stderr; a partial first line is marked as truncated. Present on every run that did not complete (`failed`, `timeout`, `aborted`), absent on `completed`, `codex_unavailable`, and launch failures. Read `stderrPath` for the complete diagnostics. - `error` — present on a launch failure, and on `timeout` and `aborted` runs The helper also prints a summary to stdout and exits with Codex's exit code, so a wrapping script can branch on success/failure directly. ## Waiting for completion The helper blocks until Codex finishes. Back it with whatever your orchestrator offers: - **Claude Code:** run the `Bash` call with `run_in_background: true`; you're notified on completion, then read `result.json`. - **Plain shell / other agents:** foreground for short tasks, or background and poll — `node relay.mjs … &` in bash/zsh (including Git Bash/WSL), or your shell's equivalent (`Start-Job` in PowerShell, `start /b` in cmd). A run is done when `result.json` exists with a `status`. **But** a pre-run usage error (bad args, empty brief) exits with code 2 *before* writing any file — so check the exit code too, don't only watch for the file. (A missing `codex` binary exits 127 but *does* write a `result.json` with status `codex_unavailable`.) Trust the working tree and the process state over any progress display. A run is finished when the process has exited and `result.json` is written — not when a status line says so. ## When a run misbehaves - **`status: codex_unavailable` (exit 127):** `codex` isn't on PATH or isn't found. Install (`npm i -g @openai/codex`) and `codex login`, then re-dispatch. - **an `error` mentioning `version preflight` (`failed`, or `timeout` at exit 124):** the bounded `codex --version` probe exited non-zero or hung past its cap (10s, or `--timeout` when shorter), so codex was never dispatched; only the relay's own artifacts may already exist under `--out-dir`. Check the install by running `codex --version` yourself. - **`status: failed`:** read `result.json`'s `stderrTail`, the full `stderrPath` log, and the tail of `eventsPath` for the cause. Common causes: an auth lapse, an invalid `--model` or unsupported `--effort`, or a sandbox that blocked something the task needed. Fix the cause and re-dispatch; don't paper over it by doing the work yourself unless that's what the user wants. - **`status: timeout`:** the `--timeout` watchdog killed the run. The working tree may hold a half-applied change — inspect it before deciding between a longer `--timeout`, a smaller brief, or a resume. - **`status: aborted`:** the relay itself was killed (its parent's timeout, a stopped task, a closed terminal) and forwarded the kill to codex. The result is written before the relay exits; inspect the working tree before re-dispatching. On native Windows a hard kill of the relay is uncatchable (Node supports no `SIGTERM` handler there), so this status may never get written - a relay process that is gone without a `result.json` is an aborted run; inspect the working tree and `events.jsonl` directly. - **`status: failed` with `signal: "SIGKILL"`:** the host ended the child — commonly the OOM killer or a supervisor timeout, not an implementer error. Free up host memory or split the task into smaller briefs, then re-dispatch. - **Empty `finalMessage`:** Codex exited before producing a final message. Treat as a failed run; the events log usually shows where it stopped. ## Recovering lost work `events.jsonl` in the run directory records every event the implementer streamed. If finished work is lost — the run killed late, or the working tree damaged afterward — read the event log before re-dispatching: it identifies which files and tool commands were involved, which scopes what needs redoing. It cannot rebuild the changes themselves — Codex's JSON stream currently reports a file change as its path and kind only, without the diff contents — so when the tree still holds the work, preserve the tree, and otherwise re-dispatch with the log as the map of what was lost. ## What the helper is doing (and the alternatives) Under the hood the helper runs roughly: ```bash codex exec --json -o <final.txt> -s workspace-write [-m model] [-c model_reasoning_effort=<level>] - < brief.txt # fresh run codex exec [-s mode] resume --last --json -o <final.txt> [-m model] [-c model_reasoning_effort=<level>] - < delta-brief.txt # resume ``` On resume, the helper places an explicit `--sandbox`/`--read-only` or fleet-lane sandbox before the `resume` subcommand so Codex applies it to the resumed turn. Without one, Codex uses its active config. The helper sets the child process's working directory instead of forwarding `-C`. Two alternatives exist if you ever want them, but the helper is the recommended path: - **Raw `codex exec`** — fine for one-offs; you give up the captured `result.json`, touched-files summary, and thread-id extraction the helper does for you. - **The openai-codex Claude Code plugin's companion CLI** (`task`/`status`/`result`) — richer job tracking if you have that plugin installed. It runs Codex as a background job behind a broker process, so you track jobs through `queued`/`running` states; the bundled helper instead spawns `codex` in-process and blocks until completion, so the only state to track is whether `result.json` exists — which is why it's the default here. ## The commit boundary The helper never commits — by design, not omission. Whether Codex's sandbox can write `.git` varies by version, OS, and execution path, so relying on it is a coin flip. The robust contract is: Codex edits the working tree, the orchestrator reviews and commits. See [review-and-land.md](review-and-land.md). -
multi-task-queues.md 3.7 KB
# Multi-task queues The single-task loop scales to a queue, and that's where delegation pays off most — a removal split across layers, a migration touching many files, a refactor sweep. The discipline that makes a queue trustworthy is sequencing and bookkeeping, not parallelism. ## Run sequentially, one commit per task Resist the urge to fan out the whole queue at once. Run tasks **one at a time, in dependency order**, landing each (review + gates + commit) before dispatching the next. Three reasons: - **Later tasks assume earlier ones landed.** Task 3's brief can say "the X added in the previous step exists" only if the previous step actually committed. - **One commit per task** keeps the history reviewable and any single step revertible. - **Each review is honest.** A clean working tree before each dispatch means the next task's `touchedFiles` shows only *its* changes, not a pile-up from earlier tasks. Parallelism is occasionally worth it for genuinely independent tasks on separate files, but it sacrifices the clean-tree-per-task property and makes review harder. Default to sequential. ## Carry decided constraints forward Implementation surfaces facts the original plan didn't have: a helper got named, a fixture lives in a specific place, an interface was chosen. When a later task depends on one of those, **fold it into that task's brief** as an explicit line. Codex has no memory of the earlier run, so a constraint that emerged in task 2 must be restated in task 5's brief or it won't hold. This is the queue equivalent of keeping briefs self-contained. ## Keep a progress file For anything longer than two or three tasks — especially a run the human steps away from — maintain a single progress file alongside the work. It's the durable record that survives your own context limits and lets the human catch up at a glance. A shape that works: - **Status table** — each task: queued / at-implementer / reviewed+committed (with the commit hash). - **Per-task review notes** — what landed, what you verified, the gate outcome. One short paragraph. - **"Needs your eyes"** — design decisions Codex made, non-blocking nitpicks, anything you want the human to overrule or confirm. This is the section they read first. - **End-of-run checklist** — what happens after the last task (push, open/update the PR, manual checks the human should do). Update it as each task lands, not in a batch at the end — if the run is interrupted, the file is still accurate. ## Close with a coherence check Per-task review proves each step in isolation; it doesn't prove the steps cohere. After the last task, verify the whole: - Run the full test/build once more on the final tree — not just the last task's slice. - Do a repo-wide check for the thing the queue was about (e.g. after a removal, grep the entire tree for any surviving reference; after a rename, confirm no stragglers). - For schema work, replay all the new migrations from a clean state and check for drift. - Then push and open or update the PR, with a description that reflects what actually shipped. ## When to stop and ask Proceed without asking on anything that follows from the agreed plan — that's the point of the human opting into the queue. Stop and surface when: - A task can't be completed correctly within its brief's scope (a scope change is the human's call). - A review finds something that calls the *plan* into question, not just the implementation. - The gates reveal a problem that affects tasks already "done." Then report where you are, what's committed, and what the open question is — and wait. A queue that quietly works around a broken assumption produces a lot of commits in the wrong direction. -
review-and-land.md 7.7 KB
# Review and land Codex did the typing; you own the judgment. This is where delegation earns its keep or quietly ships a mistake. The discipline is simple to state and easy to skip under time pressure: **verify against reality, never against the self-report — and read the diff as generated code, which fails in ways a green gate can't see.** ## Check the tests before trusting the gates If the diff touches existing tests, review those edits *first* — before the gate re-run means anything. A weakened assertion, an added skip, or a deleted test makes the gate measure less than it did before the run; green is only meaningful if the yardstick wasn't shortened. - **Unbriefed edits to existing tests are a contract change, not part of the fix.** The brief asked for an implementation; nothing in it authorized moving the goalposts. Flag them, don't absorb them. - **Skipped, disabled, or commented-out tests added in this diff:** treat the underlying test as failing until proven otherwise, whatever the annotation's comment claims. - **Loosened assertions** (exact match relaxed to contains/truthy, error-type checks broadened, tolerance widened): same treatment. ## Re-run the gates yourself `result.json` carries Codex's own claim that the gates passed. Treat that as a claim, not evidence — re-run the project's actual test/lint/build commands in the working tree and read the output. And keep the result in proportion: **passing is necessary, not sufficient.** An implementer can *game* a gate, not just misreport it — that is what the test check above and the sweep below exist to catch. For changes with their own verification shape, go further: - **Migrations / schema:** round-trip them (apply, reverse, re-apply on a scratch target) and check for drift, rather than trusting that "the migration is reversible." - **Removals / renames:** grep the codebase for dangling references to whatever was removed. - **Anything stateful:** exercise the actual behavior, don't just confirm it compiles. ## Read the diff against the brief Open the diff (`touchedFiles` in the result is your starting list) and hold it against what you asked for: - **Scope creep** — did Codex change things the brief said to leave untouched? Unasked refactors, renames, "while I was here" edits. These are the most common quality problem in delegated work. - **Scope shortfall** — did it do the whole task, including the edge cases and cleanup, or stop at the first plausible version? - **Quiet judgment calls** — sometimes Codex makes a defensible decision the brief didn't anticipate. Don't just accept it because it looks reasonable; understand it and decide. ## The implementer sweep Generated code fails in systematic ways that gates are structurally blind to — each of these can sit in a diff whose tests are all green. Walk them against every diff before you commit: - **Hardcoded success or fixture data** on a path the brief says does real work — a canned `{status: "ok"}` or default return passes tests *by design*. If Codex couldn't implement something, the diff should fail loudly, not pretend. - **Catch-all error handling that returns a default** instead of propagating — the suppressed failure is exactly what the gate would have caught. A broad catch is only acceptable with a recovery path the contract documents. - **Unverified imports and API calls** — confirm every new dependency, method, and signature exists in the *installed* version (read the lockfile or the package, don't trust plausibility). - **Dead weight** — unused imports, helpers nothing calls, unreachable branches, "Step 1/Step 2" comment scaffolding, comments that restate the line below them. - **A second way to do what the file already does** — a new HTTP client, error idiom, or logging style introduced beside the existing one instead of reusing it. - **New tests that assert internals** — asserting that an internal helper was called, or mocking the project's own functions to isolate a "unit." Green, brittle, and worthless as regression cover. - **Near-duplicate test bodies** differing by one value — fold into one data-driven test or drop the copies; bloat reads as coverage but isn't. - **Speculative surface** — optional parameters, config flags, or abstractions with no caller in this diff or the repo. Delegated work gets the concrete behavior the brief asked for, nothing extra. - **Guards for impossible cases** — null/type checks for values the code's own contract already excludes. Noise that buries the validation that matters at real trust boundaries. Anything the sweep catches goes back to Codex as a delta brief (below) or gets fixed in the tree before commit — and either way is reported to the user (see "Surface, don't absorb"). If the `guard-skills` package is installed, run the relevant guard on the diff for the full treatment — `clean-code-guard` on production code, `test-guard` on tests, `docs-guard` on documentation. The sweep above is the built-in floor; the guards go deeper. ## The commit boundary When the gates pass and the diff holds, **you commit** — the orchestrator, never Codex. This isn't a workaround for a missing feature; it's the deliberate boundary. Codex's sandbox can't reliably write `.git`, and more importantly, committing should be the act of the party that verified the work. Write a clear message describing what landed. If your project attributes co-authorship, that's the place for it. From dispatch until that commit, the uncommitted working tree is the authoritative copy of the implementer's work — the only one you can commit from, and often the only copy at all. Never run `git checkout`, `reset`, `clean`, or a branch switch in the workspace between those two points — however messy an interrupted run looks, inspect it first: `git status`, `git diff`, `git diff --cached` for anything the implementer staged (plain `git diff` is blind to the index), and open any untracked files (`??` in `git status`) directly — they are the implementer's new files, and no diff shows their contents. The tree is evidence, not clutter. After that inspection the verdict can legitimately be to discard — work built on a premise you have since corrected, for example — and then `git checkout`/`clean` is the right tool. The ban is on reflexive cleanup before anyone has looked. ## Reworking: send the delta, not the whole task If the review turns up problems, don't restate the entire brief. Continue the same Codex session with just the correction: ```bash echo "The fix is right, but the test mocks the DB session - use the real migrated fixture instead, and drop the now-unused import." | node "<skill-dir>/scripts/relay.mjs" --resume-last --cd /path/to/repo ``` (`<skill-dir>` is this skill's install directory — see [dispatch-and-poll.md](dispatch-and-poll.md).) `--session <threadId>` (or `--resume-last`) keeps Codex's context from the first run, so a short delta is enough. Then review again — rework gets the same gate-rerun, test check, diff-read, and sweep as the original, no shortcuts. Repeat until it's right, then commit. ## Surface, don't absorb The human opted into delegation, so committing verified, gate-passing work is the agreed contract. But keep them in the loop on anything that changes the shape of the work: - **Report design decisions** Codex made, and any defensible-but-unrequested turns it took. - **Note non-blocking nitpicks** you chose not to block on, so the human can overrule you. - **Stop and ask** if correct completion requires going beyond the brief — don't expand the mandate on your own. A scope change is the human's call, not yours or Codex's. For a multi-task run, capture these in the progress file rather than letting them scroll past — see [multi-task-queues.md](multi-task-queues.md). -
writing-the-brief.md 5.9 KB
# Writing the brief A brief is the entire task as Codex will see it. Codex runs in a fresh process with **no memory of your conversation, no access to your prior notes, and no shared context** — only the text you send and whatever it can read from the working tree (including the repo's own `AGENTS.md`, which it picks up automatically). If a constraint isn't in the brief or discoverable in the repo, it doesn't exist for Codex. The single most common failure is a brief that assumes context Codex doesn't have. ## The shape that works Codex responds best to compact, block-structured prompts with XML tags rather than long prose. State the task, what "done" looks like, how to behave by default, and the few constraints that actually matter. Add a block only when the task needs it — don't ship empty ceremony. ```xml <task> One or two sentences: the concrete job and where it lives. Then the specifics — current state, what to change, and explicitly what to leave untouched. The "leave untouched" list is what keeps Codex from wandering into unrelated refactors. </task> <verification_loop> Run these before finishing and fix anything they surface, don't just report it: <the project's real test command> <the project's real lint/format command> <the project's real build/typecheck command> Confirm the working tree shows only the intended changes afterward. </verification_loop> <action_safety> Keep changes scoped to the task. No unrelated refactors, renames, or cleanup unless required for correctness. Do NOT run git add or git commit — you cannot reliably write .git, and the orchestrator commits after reviewing. Leave the work uncommitted in the working tree. </action_safety> <structured_output_contract> End with a report in this exact shape: 1. What changed and why 2. Files touched 3. Gate outcomes (paste the test/lint counts) 4. Anything you deviated on, left open, or want a decision on </structured_output_contract> ``` That four-block skeleton covers most implementation tasks. Reach for the extra blocks when the task profile calls for them: - **Debugging / open-ended fixes** — add `<completeness_contract>` (resolve fully, don't stop at the first plausible fix) and `<missing_context_gating>` (don't guess missing repo facts; find them or state what's unknown). - **Review / diagnosis (read-only)** — add `<grounding_rules>` (ground every claim in evidence; label inferences) and run with `--read-only` so Codex can't edit. - **Research / recommendations** — add `<research_mode>` (separate observed facts, inferences, open questions). ## Discover the real gates — don't hardcode `<verification_loop>` is only useful if it names the project's *actual* commands. Read the repo's `CLAUDE.md` / `AGENTS.md` / `Makefile` / `package.json` first and copy the real ones in (`make test`, `npm run lint`, `cargo test`, `pytest -q`, whatever it is). A brief that says "run the tests" without naming them gets you a Codex that guesses — or skips. ## Honor the repo's conventions Codex reads the repo's `AGENTS.md` automatically, so house rules there (style, forbidden patterns, commit conventions) already apply. If the project forbids certain things in code — say, spec/ticket IDs in comments, process language like "MVP"/"for now"/"phase N", or specific test conventions, whatever the repo's own conventions ban — restate the load-bearing ones in the brief too, because Codex's compliance is only as reliable as what's in front of it. ## One task per brief Keep each brief to a single, bounded job. "Review this, fix what you find, update the docs, and suggest a roadmap" produces a muddled run; split it into separate dispatches. One brief → one Codex run → one commit keeps review and rollback clean, and lets a later task assume the earlier one landed. ## Premises freeze at dispatch The implementer starts from the brief's facts and there is no steering channel mid-run. Audit the fact block before sending — ownership, target branch, constraints, anything a judgment call rests on. If a premise turns out wrong while the run is live, stop the run and re-dispatch a corrected brief rather than discounting the output afterward; for a write-capable run, inspect the working tree and reconcile any partial or premise-contaminated edits — keep or revert them — before the re-dispatch. ## Expect environment preamble in the reply Codex's final message may carry environment noise on top of your requested report — a banner injected by the repo's `AGENTS.md`, extra text from an MCP tool or extension you've configured, and similar local additions. That comes from your own Codex setup, not a relay defect. The `<structured_output_contract>` is your defense: ask for a clearly delimited report section so you can find the real output regardless of what wraps it. ## A worked example ```xml <task> In the payments service at services/billing/, the refund path double-charges when a refund is retried after a network timeout (the idempotency key isn't checked before re-submitting). Make the refund submission idempotent: check for an existing refund by idempotency key before creating a new one. Touch only services/billing/refund.py and its tests. Leave the charge path, the API routes, and the data models untouched. </task> <verification_loop> Run and make green before finishing: pytest tests/billing/ -q ruff check services/billing/ Confirm git status shows only refund.py and its test file changed. </verification_loop> <action_safety> Scope strictly to the refund idempotency fix. No unrelated refactors. Do NOT git add or commit; leave changes in the working tree for review. </action_safety> <structured_output_contract> Report: (1) the root cause and your fix, (2) files touched, (3) pytest + ruff outcomes with counts, (4) anything you left open or want decided. </structured_output_contract> ``` Send this with `relay.mjs` (see [dispatch-and-poll.md](dispatch-and-poll.md)); review the result and commit it yourself (see [review-and-land.md](review-and-land.md)).
-
-
scripts
-
relay.mjs 33.6 KB · in bundle
-
-
SKILL.md 10 KB
--- name: codex-delegate description: >- Delegate a coding task to the OpenAI Codex CLI as a background implementer, then review its diff and land it yourself. Use this whenever the user wants to hand implementation work to Codex — phrasings like "have Codex do X", "delegate this to Codex", "run it through Codex", or "use Codex to implement/fix/refactor" — or to run a queue of coding tasks through Codex while staying the reviewer. Prefer it over a one-shot Codex forwarder (such as the codex-rescue agent) when the user will review the diff and commit it themselves. DO NOT USE for tasks small enough to do inline, or when the user wants the code written directly without delegating. license: MIT compatibility: Requires the `codex` CLI (OpenAI Codex) installed and authenticated, Node 18+, and git. The orchestrating agent must be able to run shell commands and read files. Shell examples assume bash/zsh (macOS/Linux, or Git Bash/WSL on Windows). metadata: version: 0.5.0 --- # Codex Delegate You are the **orchestrator**. This skill lets you hand a bounded coding task to a separate **implementer** — the OpenAI Codex CLI — then review what it produced and land it yourself. You write the brief and own the judgment; Codex does the typing in its own sandbox; you verify and commit. Nothing here is specific to one orchestrating agent. The loop needs only the ability to run a shell command and read a file, so it works the same whether you are Claude Code, OpenCode with a selected model, or any comparable agent. (It is designed for and run on Claude Code; treat other orchestrators as designed-for, not yet proven.) ## When NOT to use this - The task is small enough to just do inline — delegation overhead is not worth it. - The `codex` CLI is not installed or not authenticated (run `codex login`). - You want to write the code yourself, or you only need a review (use Codex's own `review` command). ## Prerequisites (check once) 1. `codex --version` succeeds. If not, install (`npm i -g @openai/codex`) and `codex login`. 2. **Confirm which `codex` is on PATH.** Multiple installs are common (e.g. a current npm/nvm copy and a stale Homebrew one). `command -v codex` shows the active one and `codex --version` its version — an old binary predates flags this skill relies on (`codex exec --json`, `-o`, `exec resume`). The relay also records the version it ran into `result.json`, so a stale binary is visible after the fact. 3. You are in (or will point `--cd` at) the target git repository. ## The loop Run these five steps per task. Steps 1, 4, and 5 are your judgment; 2 and 3 are mechanical. ### 1. Write the brief Codex sees **only** the text you send — no repo memory, no chat history, no shared context. Everything the task needs goes in the brief: the goal, the current state, what to change, what to leave untouched, the project's **actual** gate commands (discover them from the repo's CLAUDE.md/AGENTS.md/Makefile — do not assume), and a report contract. Tell Codex it will **not** commit (you will). Keep one task per brief. Full guidance and a template: [references/writing-the-brief.md](references/writing-the-brief.md). ### 2. Dispatch Send the brief to Codex with the bundled helper. It wraps `codex exec`, captures the run, and writes a structured `result.json` — so your only job is "run a command, read a file." (`<skill-dir>` below is this skill's installed directory — the folder containing this `SKILL.md`, i.e. the directory you loaded the skill from. Claude Code prints it as "Base directory for this skill" when the skill loads; on other orchestrators use that same directory — if unsure where it landed, run `find ~ -name relay.mjs -path '*codex-delegate*'` and substitute the directory above it.) ```bash node "<skill-dir>/scripts/relay.mjs" --brief brief.txt --cd /path/to/repo # read-only (review/diagnosis, no edits): add --read-only # isolated review (skip ambient MCP/user config): add --ignore-user-config # continue the exact Codex session: add --session <threadId> (from result.json; send only the delta brief) # fallback when no thread id is available: add --resume-last # hard time limit (watchdog): add --timeout 2h (default: off; implementation runs routinely need 1-2h) # see all options: node .../relay.mjs --help ``` The helper defaults to a write-capable (`workspace-write`) sandbox and writes its artifacts to a temp dir, so the repo under review stays clean. It **never commits** — see step 5. Mechanics, flags, and the `result.json` shape: [references/dispatch-and-poll.md](references/dispatch-and-poll.md). ### 3. Wait for completion The helper blocks until Codex finishes, so back it with whatever your orchestrator offers and resume when it returns: - **Claude Code:** run the Bash call with `run_in_background: true`; you are notified on completion. - **Plain shell / other agents:** run it in the foreground for short tasks, or background it and poll the result file — `… &` in bash/zsh (including Git Bash/WSL), or your shell's equivalent (`Start-Job` in PowerShell, `start /b` in cmd). The run is done when `result.json` exists with a `status`. (A pre-run usage error — bad args or an empty brief — instead exits with code 2 and a stderr message and writes no result file, so check the exit code too. A missing `codex` binary exits 127 but *does* write a `result.json` with status `codex_unavailable`.) Do not trust progress trackers over reality: a run is finished when `result.json` is written and the process has exited. Read the working tree, not a status line. The implementer's full report is the `finalMessage` field in `result.json` (also printed in full on stdout between the report markers). ### 4. Review — do not trust the self-report Codex's `result.json` includes its own summary and gate claims. **Re-verify, don't accept:** - **Re-run the project's gates yourself** (the test/lint/build commands from step 1). Never take "gates passed" on faith. - **Read the diff** against the brief: did Codex do what was asked, nothing more (scope creep) and nothing less? `touchedFiles` in the result is your starting point. - **Run the relevant guard skills** on the diff if you have them installed (clean-code-guard, test-guard, etc. from `guard-skills`) — this skill produces the work; those skills judge it. - For schema/migration changes, round-trip them; for removals, grep for dangling references. Full checklist: [references/review-and-land.md](references/review-and-land.md). ### 5. Land it Because Codex's sandbox cannot reliably write `.git` (it varies by version, OS, and path), **the orchestrator commits.** Only after the gates pass and the diff holds: - Commit the verified work yourself, with a clear message. - If it needs changes, send a delta brief with `--session <threadId>` from the prior `result.json` (use `--resume-last` only when no thread id is available), and review again. ## Read-only second opinions The relay doubles as a clean way to get an adversarial second opinion with no write risk: dispatch `--read-only` with a brief that lists the agreed points, then each contested point with both positions, and ask Codex to defend or concede each — deliverable in its final message, touching no files. Any delegation skill whose implementer offers a read-only mode supports the same use, but check how hard that mode's guarantee is first: Codex's sandbox enforces it, while Grok's is best-effort and only flagged after the fact (`readOnlyViolation`) — for those implementers, verify `touchedFiles` came back empty instead of assuming no edits. ## Authorization model Delegation is something the human opts into. Once they have ("run this queue", "proceed"), committing verified, gate-passing work is the agreed contract — that is the whole point. Two limits on that mandate: **surface, don't absorb** (report Codex's design decisions, defensible-but-unasked turns, and non-blocking nitpicks rather than silently keeping them) and **stop for scope changes** (if correct completion needs going beyond the brief, ask — don't expand the mandate yourself). The full treatment is in [references/review-and-land.md](references/review-and-land.md). ## If you have the openai-codex plugin The official openai-codex Claude Code plugin is excellent and **complementary** — `codex-delegate` builds on the same `codex` CLI, it doesn't replace the plugin. They point in different directions: - The plugin's `codex:codex-rescue` agent is a **forwarder**: it hands one task to Codex and returns the output. It deliberately does not poll, review, or commit. - The plugin's review command and stop-review gate run the **inverse** direction: **Codex reviews your work**. - `codex-delegate` is the **orchestration loop in the other direction**: *you* drive Codex to implement across one task or a queue, and *you* review and land each result. That loop — brief → dispatch → poll → review → commit, with the orchestrator owning the commit — is what the plugin leaves to you, and what this skill encodes. If you have the plugin installed, its companion CLI is an optional alternative dispatch backend; the bundled `relay.mjs` is the default because it adds no install of its own beyond the `codex` binary (Node and `git`, which the relay also needs, are prerequisites for every skill here). ## References - [references/writing-the-brief.md](references/writing-the-brief.md) — how to write a brief Codex can execute blind: structure, XML blocks, the report contract, embedding the real gate commands. - [references/dispatch-and-poll.md](references/dispatch-and-poll.md) — `relay.mjs` flags, the `result.json` contract, backgrounding per orchestrator, and recovery when a run misbehaves. - [references/review-and-land.md](references/review-and-land.md) — the review checklist, the commit boundary, and the exact-session rework cycle. - [references/multi-task-queues.md](references/multi-task-queues.md) — running a sequential queue: carrying constraints forward, progress tracking, and the end-of-run coherence check.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.