Claude Skill

codex-ab

Run an A/B codex review experiment — holistic codex review vs 3 focused dimension passes (security, ecto, liveview) on the branch diff, classify findings, report a panel-value verdict. Use when the branch is fresh, before any codex review runs.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oliver-kriska-claude-elixir-phoenix-.claude_skills_codex-ab-9767a82.zip · 4 KB
Part of oliver-kriska/claude-elixir-phoenix — 93 skills

Install

skills CLI npx skills add https://github.com/oliver-kriska/claude-elixir-phoenix/tree/main/.claude/skills/codex-ab
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oliver-kriska-claude-elixir-phoenix@llmmart
Git git clone https://github.com/oliver-kriska/claude-elixir-phoenix.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oliver-kriska/claude-elixir-phoenix collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Codex Panel A/B (contributor instrument — verdict decided)

Answer one question with evidence: do dimension-focused codex passes find real issues that one holistic codex exec review misses? Runs both on the same diff, then classifies every focused finding against the holistic pass.

DECIDED 2026-07-10 after 4 runs (2 fresh): panel KILLED. Fresh-only 1 real miss / 1 false positive plus one zero-value run at 4× cost — real misses did not outnumber FPs. Kept as contributor tooling (NOT distributed) for one possible retest: a UI-heavy diff with a single extra liveview-focused pass (2× cost). Scoreboard: .claude/research/2026-07-03-codex-review-integration.md §7.

Usage

/codex-ab            # A/B against main (~5 min, 4 codex runs)
/codex-ab develop    # explicit base branch

Iron Laws

  1. FRESH DIFF ONLY — ask the user to confirm this branch has NOT been codex-reviewed yet (cloud or /phx:codex-loop). A drained diff returns NO FINDINGS everywhere and proves nothing — wasted quota
  2. Verify every REAL MISS in the code before counting it — a focused finding only scores if the issue actually exists at that file:line
  3. Read ONLY the findings .md files — streams are diverted to .log files; never cat a log into context (10k+ lines each)
  4. Exactly 4 codex runs, never re-run dimensions — bounded quota
  5. Persist the verdict — an unrecorded experiment is wasted quota

Workflow

Step 1: Preflight

Run command -v codex — missing → STOP with install hint. Then:

  • git status --short dirty → warn (codex flags local dirt as findings)
  • Ask: "Has codex already reviewed this branch (PR review or codex-loop)?" If yes → STOP, explain the fresh-diff requirement (Iron Law 1)

Step 2: Run the A/B (background, ~5 min)

bash ${CLAUDE_SKILL_DIR}/scripts/codex-panel-ab.sh {base} \
  .claude/reviews/codex-ab-$(date +%Y-%m-%d-%H%M)

Use run_in_background — it runs 1 holistic codex exec review + 3 focused codex exec workers (security / ecto / liveview) in parallel, all streams redirected. Do other work or wait; never poll.

Step 3: Classify

Read the 4 findings files (holistic.md, security.md, ecto.md, liveview.md — small). For EACH focused finding:

Class Meaning Test
DUPLICATE Holistic already found it Same file + same defect
REAL MISS Genuine issue holistic missed Read the code at file:line — defect confirmed (Iron Law 2)
FALSE POSITIVE Manufactured, pre-existing, or wrong Code check fails, or issue exists on base branch too

Step 4: Verdict

Present:

## Codex Panel A/B — {branch} vs {base}
| dimension | findings | duplicate | real miss | false positive |
Holistic-only findings: {n}
Verdict this run: {REAL MISS count} real miss vs {FP count} false positive
Decision rule: build --codex-panel only if real misses outnumber false
positives across 2-3 fresh branches.

Write the verdict table to .claude/reviews/codex-ab-{date}/VERDICT.md. Suggest repeating on the next 1–2 fresh branches before deciding.

Integration

fresh branch → /codex-ab (YOU ARE HERE) → verdict logged
   ├─ real misses win across runs → build /phx:review --codex-panel
   └─ duplicates/FPs win → keep holistic /phx:codex-loop, drop panel idea
       └─ OUTCOME 2026-07-10: this branch won — panel dropped

References

  • ${CLAUDE_SKILL_DIR}/scripts/codex-panel-ab.sh — the 4-run harness
  • Related: /phx:codex-loop (holistic fix loop), /phx:review --codex
Files (claude-elixir-phoenix)
  • scripts
    • codex-panel-ab.sh 3.3 KB
      #!/usr/bin/env bash
      # codex-panel-ab.sh — A/B: one holistic `codex exec review` vs a 3-dimension
      # focused panel (plain `codex exec` workers), on the current branch's diff
      # against a base branch. Decides whether /phx:review deserves --codex-panel.
      #
      # Usage: run from the target project repo, on a FRESH (not yet
      # codex-reviewed) branch:
      #   bash codex-panel-ab.sh [base-branch] [output-dir]
      # Defaults: base=main, output=./codex-panel-ab-<timestamp>/
      #
      # Cost: 4 codex runs (~2-5 min each, run in parallel). Read-only sandbox.
      # Output: 4 findings .md files + stream .log files + a comparison stub.
      set -uo pipefail
      
      BASE="${1:-main}"
      OUT="${2:-codex-panel-ab-$(date +%Y-%m-%d-%H%M)}"
      mkdir -p "$OUT"
      
      command -v codex >/dev/null || { echo "codex CLI not found"; exit 1; }
      git rev-parse --verify "origin/$BASE" >/dev/null 2>&1 || { echo "no origin/$BASE"; exit 1; }
      [[ -n "$(git status --short)" ]] && echo "WARN: dirty tree — codex may flag local dirt"
      
      GUARD="Report each finding as '- [P1|P2|P3] title — file:line' plus 2-3 sentences citing the actual code. Scope discipline: ONLY issues introduced by this diff (git diff origin/$BASE...HEAD). If you find no genuine issues in your focus area, output exactly 'NO FINDINGS' — do NOT manufacture findings, do NOT report pre-existing issues outside the diff."
      
      focus_prompt() { # $1 = dimension description
        echo "Review ONLY the changes introduced by this branch relative to origin/$BASE (run: git diff origin/$BASE...HEAD) with a strict focus on $1. $GUARD"
      }
      
      echo "Running 1 holistic + 3 focused codex reviews against origin/$BASE (parallel, ~5 min)..."
      
      codex exec review --base "$BASE" --ephemeral \
        -o "$OUT/holistic.md" > "$OUT/holistic.log" 2>&1 &
      
      codex exec --ephemeral -s read-only -o "$OUT/security.md" \
        "$(focus_prompt "SECURITY: authorization gaps, SQL/tsquery injection, unsafe atom creation from user input, XSS, secrets exposure, DoS vectors")" \
        > "$OUT/security.log" 2>&1 &
      
      codex exec --ephemeral -s read-only -o "$OUT/ecto.md" \
        "$(focus_prompt "ECTO AND DATA CORRECTNESS: N+1 queries, missing preloads, row multiplication from has_many joins, implicit cross joins, float money, unpinned query values, migration hazards, constraint gaps")" \
        > "$OUT/ecto.log" 2>&1 &
      
      codex exec --ephemeral -s read-only -o "$OUT/liveview.md" \
        "$(focus_prompt "LIVEVIEW AND CONCURRENCY: unconditional queries in mount, missing streams for large lists, unauthorized handle_event, assign bloat, PubSub double-subscribe, unsupervised processes, race conditions")" \
        > "$OUT/liveview.log" 2>&1 &
      
      wait
      echo "Done. Findings:"
      for f in holistic security ecto liveview; do
        n=$(grep -c '^- \[P' "$OUT/$f.md" 2>/dev/null || echo 0)
        echo "  $f: $n finding(s)  ($OUT/$f.md)"
      done
      
      cat > "$OUT/COMPARE.md" <<'EOF'
      # Comparison checklist
      
      For each focused finding, classify against holistic.md:
      - DUPLICATE — holistic already found it → panel adds no value here
      - REAL MISS — genuine issue holistic missed → +1 for --codex-panel
      - FALSE POSITIVE — manufactured/pre-existing/wrong → -1 for --codex-panel
      
      Verdict rule of thumb: build --codex-panel only if REAL MISS > FALSE POSITIVE
      across 2+ fresh branches. Log the result to lab/findings/interesting.jsonl.
      EOF
      echo "Next: fill in $OUT/COMPARE.md (classify each focused finding vs holistic)."
      
  • SKILL.md 3.9 KB
    ---
    name: codex-ab
    description: Run an A/B codex review experiment — holistic codex review vs 3 focused dimension passes (security, ecto, liveview) on the branch diff, classify findings, report a panel-value verdict. Use when the branch is fresh, before any codex review runs.
    effort: medium
    argument-hint: "[base-branch]"
    ---
    
    # Codex Panel A/B (contributor instrument — verdict decided)
    
    Answer one question with evidence: **do dimension-focused codex passes find
    real issues that one holistic `codex exec review` misses?** Runs both on the
    same diff, then classifies every focused finding against the holistic pass.
    
    > **DECIDED 2026-07-10 after 4 runs (2 fresh): panel KILLED.** Fresh-only
    > 1 real miss / 1 false positive plus one zero-value run at 4× cost —
    > real misses did not outnumber FPs. Kept as contributor tooling (NOT
    > distributed) for one possible retest: a UI-heavy diff with a single
    > extra liveview-focused pass (2× cost). Scoreboard:
    > `.claude/research/2026-07-03-codex-review-integration.md` §7.
    
    ## Usage
    
    ```
    /codex-ab            # A/B against main (~5 min, 4 codex runs)
    /codex-ab develop    # explicit base branch
    ```
    
    ## Iron Laws
    
    1. **FRESH DIFF ONLY** — ask the user to confirm this branch has NOT been
       codex-reviewed yet (cloud or `/phx:codex-loop`). A drained diff returns
       NO FINDINGS everywhere and proves nothing — wasted quota
    2. **Verify every REAL MISS in the code before counting it** — a focused
       finding only scores if the issue actually exists at that file:line
    3. **Read ONLY the findings `.md` files** — streams are diverted to `.log`
       files; never cat a log into context (10k+ lines each)
    4. **Exactly 4 codex runs, never re-run dimensions** — bounded quota
    5. **Persist the verdict** — an unrecorded experiment is wasted quota
    
    ## Workflow
    
    ### Step 1: Preflight
    
    Run `command -v codex` — missing → STOP with install hint. Then:
    
    - `git status --short` dirty → warn (codex flags local dirt as findings)
    - Ask: "Has codex already reviewed this branch (PR review or codex-loop)?"
      If yes → STOP, explain the fresh-diff requirement (Iron Law 1)
    
    ### Step 2: Run the A/B (background, ~5 min)
    
    ```bash
    bash ${CLAUDE_SKILL_DIR}/scripts/codex-panel-ab.sh {base} \
      .claude/reviews/codex-ab-$(date +%Y-%m-%d-%H%M)
    ```
    
    Use `run_in_background` — it runs 1 holistic `codex exec review` + 3
    focused `codex exec` workers (security / ecto / liveview) in parallel,
    all streams redirected. Do other work or wait; never poll.
    
    ### Step 3: Classify
    
    Read the 4 findings files (`holistic.md`, `security.md`, `ecto.md`,
    `liveview.md` — small). For EACH focused finding:
    
    | Class | Meaning | Test |
    |-------|---------|------|
    | DUPLICATE | Holistic already found it | Same file + same defect |
    | REAL MISS | Genuine issue holistic missed | Read the code at file:line — defect confirmed (Iron Law 2) |
    | FALSE POSITIVE | Manufactured, pre-existing, or wrong | Code check fails, or issue exists on base branch too |
    
    ### Step 4: Verdict
    
    Present:
    
    ```markdown
    ## Codex Panel A/B — {branch} vs {base}
    | dimension | findings | duplicate | real miss | false positive |
    Holistic-only findings: {n}
    Verdict this run: {REAL MISS count} real miss vs {FP count} false positive
    Decision rule: build --codex-panel only if real misses outnumber false
    positives across 2-3 fresh branches.
    ```
    
    Write the verdict table to `.claude/reviews/codex-ab-{date}/VERDICT.md`.
    Suggest repeating on the next 1–2 fresh branches before deciding.
    
    ## Integration
    
    ```text
    fresh branch → /codex-ab (YOU ARE HERE) → verdict logged
       ├─ real misses win across runs → build /phx:review --codex-panel
       └─ duplicates/FPs win → keep holistic /phx:codex-loop, drop panel idea
           └─ OUTCOME 2026-07-10: this branch won — panel dropped
    ```
    
    ## References
    
    - `${CLAUDE_SKILL_DIR}/scripts/codex-panel-ab.sh` — the 4-run harness
    - Related: `/phx:codex-loop` (holistic fix loop), `/phx:review --codex`
    
  • triggers.json 719 B
    {
      "skill": "codex-ab",
      "should_trigger": [
        "Run the codex A/B experiment on this branch before I request any review",
        "Compare a holistic codex review against focused security and ecto passes on my diff",
        "codex-ab against develop",
        "Test whether focused codex reviewers find more than the holistic one on this fresh branch",
        "Run the codex panel experiment and give me the verdict table"
      ],
      "should_not_trigger": [
        "Run a codex review on my changes and fix whatever it finds",
        "Review my changes with the specialist agents before committing",
        "Comment @codex review on PR 42 and watch for its feedback",
        "A/B test the new search backend against Algolia in production"
      ]
    }
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related