codex-ab
Run an A/B codex review experiment — holistic codex review vs 3 focused dimension passes (security, ecto, liveview) on the branch diff, classify findings, report a panel-value verdict. Use when the branch is fresh, before any codex review runs.
Install
npx skills add https://github.com/oliver-kriska/claude-elixir-phoenix/tree/main/.claude/skills/codex-ab
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oliver-kriska-claude-elixir-phoenix@llmmart
git clone https://github.com/oliver-kriska/claude-elixir-phoenix.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oliver-kriska/claude-elixir-phoenix collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Codex Panel A/B (contributor instrument — verdict decided)
Answer one question with evidence: do dimension-focused codex passes find
real issues that one holistic codex exec review misses? Runs both on the
same diff, then classifies every focused finding against the holistic pass.
DECIDED 2026-07-10 after 4 runs (2 fresh): panel KILLED. Fresh-only 1 real miss / 1 false positive plus one zero-value run at 4× cost — real misses did not outnumber FPs. Kept as contributor tooling (NOT distributed) for one possible retest: a UI-heavy diff with a single extra liveview-focused pass (2× cost). Scoreboard:
.claude/research/2026-07-03-codex-review-integration.md§7.
Usage
/codex-ab # A/B against main (~5 min, 4 codex runs)
/codex-ab develop # explicit base branch
Iron Laws
- FRESH DIFF ONLY — ask the user to confirm this branch has NOT been
codex-reviewed yet (cloud or
/phx:codex-loop). A drained diff returns NO FINDINGS everywhere and proves nothing — wasted quota - Verify every REAL MISS in the code before counting it — a focused finding only scores if the issue actually exists at that file:line
- Read ONLY the findings
.mdfiles — streams are diverted to.logfiles; never cat a log into context (10k+ lines each) - Exactly 4 codex runs, never re-run dimensions — bounded quota
- Persist the verdict — an unrecorded experiment is wasted quota
Workflow
Step 1: Preflight
Run command -v codex — missing → STOP with install hint. Then:
git status --shortdirty → warn (codex flags local dirt as findings)- Ask: "Has codex already reviewed this branch (PR review or codex-loop)?" If yes → STOP, explain the fresh-diff requirement (Iron Law 1)
Step 2: Run the A/B (background, ~5 min)
bash ${CLAUDE_SKILL_DIR}/scripts/codex-panel-ab.sh {base} \
.claude/reviews/codex-ab-$(date +%Y-%m-%d-%H%M)
Use run_in_background — it runs 1 holistic codex exec review + 3
focused codex exec workers (security / ecto / liveview) in parallel,
all streams redirected. Do other work or wait; never poll.
Step 3: Classify
Read the 4 findings files (holistic.md, security.md, ecto.md,
liveview.md — small). For EACH focused finding:
| Class | Meaning | Test |
|---|---|---|
| DUPLICATE | Holistic already found it | Same file + same defect |
| REAL MISS | Genuine issue holistic missed | Read the code at file:line — defect confirmed (Iron Law 2) |
| FALSE POSITIVE | Manufactured, pre-existing, or wrong | Code check fails, or issue exists on base branch too |
Step 4: Verdict
Present:
## Codex Panel A/B — {branch} vs {base}
| dimension | findings | duplicate | real miss | false positive |
Holistic-only findings: {n}
Verdict this run: {REAL MISS count} real miss vs {FP count} false positive
Decision rule: build --codex-panel only if real misses outnumber false
positives across 2-3 fresh branches.
Write the verdict table to .claude/reviews/codex-ab-{date}/VERDICT.md.
Suggest repeating on the next 1–2 fresh branches before deciding.
Integration
fresh branch → /codex-ab (YOU ARE HERE) → verdict logged
├─ real misses win across runs → build /phx:review --codex-panel
└─ duplicates/FPs win → keep holistic /phx:codex-loop, drop panel idea
└─ OUTCOME 2026-07-10: this branch won — panel dropped
References
${CLAUDE_SKILL_DIR}/scripts/codex-panel-ab.sh— the 4-run harness- Related:
/phx:codex-loop(holistic fix loop),/phx:review --codex
Files (claude-elixir-phoenix)
-
scripts
-
codex-panel-ab.sh 3.3 KB
#!/usr/bin/env bash # codex-panel-ab.sh — A/B: one holistic `codex exec review` vs a 3-dimension # focused panel (plain `codex exec` workers), on the current branch's diff # against a base branch. Decides whether /phx:review deserves --codex-panel. # # Usage: run from the target project repo, on a FRESH (not yet # codex-reviewed) branch: # bash codex-panel-ab.sh [base-branch] [output-dir] # Defaults: base=main, output=./codex-panel-ab-<timestamp>/ # # Cost: 4 codex runs (~2-5 min each, run in parallel). Read-only sandbox. # Output: 4 findings .md files + stream .log files + a comparison stub. set -uo pipefail BASE="${1:-main}" OUT="${2:-codex-panel-ab-$(date +%Y-%m-%d-%H%M)}" mkdir -p "$OUT" command -v codex >/dev/null || { echo "codex CLI not found"; exit 1; } git rev-parse --verify "origin/$BASE" >/dev/null 2>&1 || { echo "no origin/$BASE"; exit 1; } [[ -n "$(git status --short)" ]] && echo "WARN: dirty tree — codex may flag local dirt" GUARD="Report each finding as '- [P1|P2|P3] title — file:line' plus 2-3 sentences citing the actual code. Scope discipline: ONLY issues introduced by this diff (git diff origin/$BASE...HEAD). If you find no genuine issues in your focus area, output exactly 'NO FINDINGS' — do NOT manufacture findings, do NOT report pre-existing issues outside the diff." focus_prompt() { # $1 = dimension description echo "Review ONLY the changes introduced by this branch relative to origin/$BASE (run: git diff origin/$BASE...HEAD) with a strict focus on $1. $GUARD" } echo "Running 1 holistic + 3 focused codex reviews against origin/$BASE (parallel, ~5 min)..." codex exec review --base "$BASE" --ephemeral \ -o "$OUT/holistic.md" > "$OUT/holistic.log" 2>&1 & codex exec --ephemeral -s read-only -o "$OUT/security.md" \ "$(focus_prompt "SECURITY: authorization gaps, SQL/tsquery injection, unsafe atom creation from user input, XSS, secrets exposure, DoS vectors")" \ > "$OUT/security.log" 2>&1 & codex exec --ephemeral -s read-only -o "$OUT/ecto.md" \ "$(focus_prompt "ECTO AND DATA CORRECTNESS: N+1 queries, missing preloads, row multiplication from has_many joins, implicit cross joins, float money, unpinned query values, migration hazards, constraint gaps")" \ > "$OUT/ecto.log" 2>&1 & codex exec --ephemeral -s read-only -o "$OUT/liveview.md" \ "$(focus_prompt "LIVEVIEW AND CONCURRENCY: unconditional queries in mount, missing streams for large lists, unauthorized handle_event, assign bloat, PubSub double-subscribe, unsupervised processes, race conditions")" \ > "$OUT/liveview.log" 2>&1 & wait echo "Done. Findings:" for f in holistic security ecto liveview; do n=$(grep -c '^- \[P' "$OUT/$f.md" 2>/dev/null || echo 0) echo " $f: $n finding(s) ($OUT/$f.md)" done cat > "$OUT/COMPARE.md" <<'EOF' # Comparison checklist For each focused finding, classify against holistic.md: - DUPLICATE — holistic already found it → panel adds no value here - REAL MISS — genuine issue holistic missed → +1 for --codex-panel - FALSE POSITIVE — manufactured/pre-existing/wrong → -1 for --codex-panel Verdict rule of thumb: build --codex-panel only if REAL MISS > FALSE POSITIVE across 2+ fresh branches. Log the result to lab/findings/interesting.jsonl. EOF echo "Next: fill in $OUT/COMPARE.md (classify each focused finding vs holistic)."
-
-
SKILL.md 3.9 KB
--- name: codex-ab description: Run an A/B codex review experiment — holistic codex review vs 3 focused dimension passes (security, ecto, liveview) on the branch diff, classify findings, report a panel-value verdict. Use when the branch is fresh, before any codex review runs. effort: medium argument-hint: "[base-branch]" --- # Codex Panel A/B (contributor instrument — verdict decided) Answer one question with evidence: **do dimension-focused codex passes find real issues that one holistic `codex exec review` misses?** Runs both on the same diff, then classifies every focused finding against the holistic pass. > **DECIDED 2026-07-10 after 4 runs (2 fresh): panel KILLED.** Fresh-only > 1 real miss / 1 false positive plus one zero-value run at 4× cost — > real misses did not outnumber FPs. Kept as contributor tooling (NOT > distributed) for one possible retest: a UI-heavy diff with a single > extra liveview-focused pass (2× cost). Scoreboard: > `.claude/research/2026-07-03-codex-review-integration.md` §7. ## Usage ``` /codex-ab # A/B against main (~5 min, 4 codex runs) /codex-ab develop # explicit base branch ``` ## Iron Laws 1. **FRESH DIFF ONLY** — ask the user to confirm this branch has NOT been codex-reviewed yet (cloud or `/phx:codex-loop`). A drained diff returns NO FINDINGS everywhere and proves nothing — wasted quota 2. **Verify every REAL MISS in the code before counting it** — a focused finding only scores if the issue actually exists at that file:line 3. **Read ONLY the findings `.md` files** — streams are diverted to `.log` files; never cat a log into context (10k+ lines each) 4. **Exactly 4 codex runs, never re-run dimensions** — bounded quota 5. **Persist the verdict** — an unrecorded experiment is wasted quota ## Workflow ### Step 1: Preflight Run `command -v codex` — missing → STOP with install hint. Then: - `git status --short` dirty → warn (codex flags local dirt as findings) - Ask: "Has codex already reviewed this branch (PR review or codex-loop)?" If yes → STOP, explain the fresh-diff requirement (Iron Law 1) ### Step 2: Run the A/B (background, ~5 min) ```bash bash ${CLAUDE_SKILL_DIR}/scripts/codex-panel-ab.sh {base} \ .claude/reviews/codex-ab-$(date +%Y-%m-%d-%H%M) ``` Use `run_in_background` — it runs 1 holistic `codex exec review` + 3 focused `codex exec` workers (security / ecto / liveview) in parallel, all streams redirected. Do other work or wait; never poll. ### Step 3: Classify Read the 4 findings files (`holistic.md`, `security.md`, `ecto.md`, `liveview.md` — small). For EACH focused finding: | Class | Meaning | Test | |-------|---------|------| | DUPLICATE | Holistic already found it | Same file + same defect | | REAL MISS | Genuine issue holistic missed | Read the code at file:line — defect confirmed (Iron Law 2) | | FALSE POSITIVE | Manufactured, pre-existing, or wrong | Code check fails, or issue exists on base branch too | ### Step 4: Verdict Present: ```markdown ## Codex Panel A/B — {branch} vs {base} | dimension | findings | duplicate | real miss | false positive | Holistic-only findings: {n} Verdict this run: {REAL MISS count} real miss vs {FP count} false positive Decision rule: build --codex-panel only if real misses outnumber false positives across 2-3 fresh branches. ``` Write the verdict table to `.claude/reviews/codex-ab-{date}/VERDICT.md`. Suggest repeating on the next 1–2 fresh branches before deciding. ## Integration ```text fresh branch → /codex-ab (YOU ARE HERE) → verdict logged ├─ real misses win across runs → build /phx:review --codex-panel └─ duplicates/FPs win → keep holistic /phx:codex-loop, drop panel idea └─ OUTCOME 2026-07-10: this branch won — panel dropped ``` ## References - `${CLAUDE_SKILL_DIR}/scripts/codex-panel-ab.sh` — the 4-run harness - Related: `/phx:codex-loop` (holistic fix loop), `/phx:review --codex` -
triggers.json 719 B
{ "skill": "codex-ab", "should_trigger": [ "Run the codex A/B experiment on this branch before I request any review", "Compare a holistic codex review against focused security and ecto passes on my diff", "codex-ab against develop", "Test whether focused codex reviewers find more than the holistic one on this fresh branch", "Run the codex panel experiment and give me the verdict table" ], "should_not_trigger": [ "Run a codex review on my changes and fix whatever it finds", "Review my changes with the specialist agents before committing", "Comment @codex review on PR 42 and watch for its feedback", "A/B test the new search backend against Algolia in production" ] }
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.