Claude Skill

ab-testing

Use when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go significant. NOT recurring metric tracking (that is `analytics`), NOT north-star/KPI trees (that is `kp

LLM Mart · 0 points · 0 views 2 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download ericrisco-rsc-harness-skills_ab-testing-953fef5.zip · 13 KB
Part of ericrisco/rsc-harness — 46 skills

Install

skills CLI npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/ab-testing
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
Git git clone https://github.com/ericrisco/rsc-harness.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

A/B testing — design and read a defensible experiment

An experiment without a pre-committed sample size and a single primary metric is not an experiment. It is a dashboard you stare at until it tells you what you wanted to hear. The discipline lives almost entirely before traffic ships: a falsifiable hypothesis, one primary metric, a sample size derived from the smallest effect worth detecting, and a stop rule you cannot renegotiate at 2pm on day four.

Pre-test checklist — every line true before any traffic

Each one is a place experiments die silently.

  • A falsifiable hypothesis — names the change, the direction, and the metric it moves.
  • Exactly ONE primary metric. More than one primary = multiple comparisons = inflated false positives.
  • Guardrail metrics — what you refuse to harm (latency, refunds, unsubscribes) even for a win.
  • The randomization unit = the analysis unit (usually the user). Mixing them is pseudoreplication.
  • An MDE — the smallest lift that would change a decision. Not "any difference."
  • A computed sample size and the duration it implies at your real daily eligible traffic.
  • A fixed stop rule — a date or an n you commit to before launch. No "we'll see how it looks."

Step 1 — Hypothesis and metrics

State a null you can reject. "The new checkout button changes purchase conversion" with H0: conversion equal across arms, H1: it differs. Vague aspirations ("improve the funnel") have no rejection region.

Pick one primary metric and freeze it. Why: every extra primary metric is another coin flip at α, so three "primary" metrics turn a 5% false-positive rate into roughly 14%. Demote the rest to secondary.

Randomize on the same unit you analyze on. If a user sees the variant on every visit, randomize by user, not by session — analyzing 50k sessions from 8k users treats correlated observations as independent and fabricates significance.

Bad:  "We think the redesign will improve engagement and revenue and retention."  (no null, 3 primaries, no number)
Good: "H0: 30-day purchase conversion is equal between control and the new one-click button.
       H1: it differs. Primary: purchase conversion. Guardrails: refund rate, p95 checkout latency.
       Randomize by user_id. MDE: +1.5pp absolute on a 12% baseline."

Step 2 — Sample size from MDE, baseline, and power

Defaults: power 0.80, α 0.05 (two-sided). The MDE is yours to choose — it is the smallest effect that would actually change what you do.

Rule: required n scales with ~1/MDE². Why: halving the smallest effect you care to detect roughly quadruples the traffic and time. This is the single most expensive decision in the design, so set the MDE to a business threshold, never to "whatever is small."

For a conversion rate (proportion):

from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import proportion_effectsize

p1, p2 = 0.12, 0.135                       # baseline, baseline + MDE (1.5pp)
h = proportion_effectsize(p1, p2)          # Cohen's h (arcsine transform)
n = NormalIndPower().solve_power(effect_size=h, alpha=0.05, power=0.80, ratio=1.0)
print(int(-(-n // 1)))                      # n PER ARM, rounded up

For a continuous metric (revenue per user, time on page) use Welch-style sizing:

from statsmodels.stats.power import TTestIndPower

effect = mde_in_units / pooled_std         # Cohen's d
n = TTestIndPower().solve_power(effect_size=effect, alpha=0.05, power=0.80, ratio=1.0)

Then convert n to a calendar plan: days = ceil((n_per_arm * num_arms) / daily_eligible_users). If that is 9 days, run a clean two full weeks anyway — weekday/weekend mix is part of the population, and a 6-day test oversamples whoever shows up Tuesday. Full worked example (12% baseline, +1.5pp MDE, 80% power) plus runnable sizing, n→duration, CUPED θ and SRM snippets: references/sample-size-and-cuped.md.

Step 3 — Run discipline

Fixed horizon is the default. Commit to the n/date from Step 2 and read the result once, at the end.

Do not peek and stop at first significance. Why: checking repeatedly and stopping the moment p < 0.05 inflates the Type-I error far above 5% — with enough looks, a null test crosses 0.05 most of the time. If you genuinely need to stop early, use a sequential / always-valid method (confidence sequences, e.g. Netflix's anytime-valid CIs) that holds Type-I error under continuous monitoring. Sequential is strong for killing losers early and weak for calling winners early — for a confident win, the fixed-horizon read is tighter.

Gate on SRM before you trust anything. Compute a chi-square test on the observed split versus the intended ratio. If p < 0.001 the assignment or logging is broken — a bot filter dropping one arm, a redirect, a caching bug. Fix the instrumentation and rerun; do not "adjust for it."

The peeking Type-I math, sequential/always-valid options, SRM diagnosis, novelty/primacy effects, Simpson's paradox in segments and HARKing all live in references/pitfalls.md.

Step 4 — Analyze

Pick the test by metric type:

Metric type Test
Binary conversion (proportion) Two-proportion z-test (statsmodels.stats.proportion.proportions_ztest)
Continuous, roughly normal / large n Welch's t-test (scipy.stats.ttest_ind(..., equal_var=False))
Continuous, heavy-tailed / skewed (revenue) Mann-Whitney U, or t-test on a log/winsorized metric

Report lift + confidence interval + p-value together. Never p alone. Why: p < 0.05 with a CI of [+0.1pp, +5pp] is "statistically there, practically a coin toss" — the CI tells you the size, p only tells you it is not exactly zero. Practical significance = compare the CI to your MDE: if the whole interval sits above the MDE, ship; if it straddles the MDE, you detected something too small to matter.

Multiple comparisons. Two regimes:

  • Small set of pre-declared decision metrics → Bonferroni (divide α by the count). Conservative, simple.
  • Large exploratory scan of many metrics/segments → Benjamini-Hochberg (FDR). It keeps far more power than Bonferroni on big scans (in a 20-effect example, ~17 detected vs ~12 under Bonferroni).

Step 5 — CUPED variance reduction

CUPED (Controlled-experiment Using Pre-Experiment Data) subtracts predictable pre-period noise so the same traffic buys more power — or the same power needs less traffic. The adjusted metric:

Y_cuped = Y − θ · (X − E[X])        where  θ = Cov(Y, X) / Var(X)

Estimate θ by regressing the in-experiment metric Y on the pre-experiment covariate X (e.g. each user's spend in the 4 weeks before the test), then analyze Y_cuped with the same test as Step 4.

When it pays: recurring users with a strong pre-period signal. Reported wins — Netflix ~40% variance reduction on engagement, Statsig 50%+ on common metrics → significance in roughly half the time/traffic.

When it does nothing — do not bother: brand-new users (no pre-period data), a covariate uncorrelated with the outcome, or — the cardinal sin — a covariate measured after assignment, which biases the estimate. The covariate MUST be pre-treatment and independent of which arm a user lands in. Runnable θ-via-OLS snippet in references/sample-size-and-cuped.md.

Anti-patterns

Bad Why it is wrong Do instead
Peek daily, stop the day p < 0.05 Repeated looks inflate Type-I error far above α Fix n/date up front; or a sequential method that holds α
No sample size set before launch You will stop on noise and call it a win Compute n from MDE/baseline/power in Step 2
Several "primary" metrics Each is a coin flip at α; 3 metrics ≈ 14% false-positive One frozen primary; the rest are secondary
Ignore the observed split An SRM means assignment/logging is broken; results are garbage Chi-square SRM gate before reading anything
Report only the p-value Hides effect size — p < 0.05 can be practically zero Always lift + CI + p; compare CI to MDE
CUPED on a post-assignment covariate Covariate correlated with the arm biases θ Use only pre-treatment, assignment-independent covariates
Call a winner from an underpowered test "Not significant" then ≠ "no effect"; you lacked power Reach planned n, or report the CI and say "inconclusive, here is the range"
Decide the hypothesis after seeing results (HARKing) Turns the whole analysis into a fishing expedition Pre-register hypothesis + primary metric before launch
Run 6 days because it "looks significant" Oversamples one weekday slice of the population Run full weeks; honor the fixed horizon

Checkable artifact

When this skill emits a Python sizing/analysis script or an experiment-design doc, run scripts/verify.sh from your project root. It confirms the script executes under python3 and prints a numeric sample size, and that any design doc names a primary metric, an MDE, and power/alpha. It is read-only and soft-passes when no artifact is present (a design-only conversation).

Files (rsc-harness)
  • evals
    • cases.yaml 3.5 KB
      skill: ab-testing
      
      should_trigger:
        - prompt: "How many users per arm do I need to detect a 2% lift on a 10% baseline conversion?"
          why: Core sample-size sizing from baseline + MDE + power — the central pre-test calculation this skill owns.
        - prompt: "Is this A/B result statistically significant? Control converted 540/10000, variant 590/10000."
          why: Read-out of a finished proportion test — two-proportion z-test, lift, CI, p-value together.
        - prompt: "Our test has run two weeks and still isn't significant, what do we do now?"
          why: Non-obvious. "Inconclusive" almost always means underpowered, peeked, or an SRM — a design diagnosis, not a re-run.
        - prompt: "Can we use CUPED to get faster results on our checkout experiment?"
          why: Variance reduction via a pre-experiment covariate — when it helps, the formula, and when it does nothing.
        - prompt: "We keep peeking at the dashboard and stopping the moment it hits p<0.05 — is that a problem?"
          why: Non-obvious optional-stopping trap; peeking inflates Type-I error far above alpha. Fixed horizon or always-valid inference.
        - prompt: "El experiment de la landing no surt significatiu, com el dissenyo bé des de zero?"
          why: Catalan trigger — rescue an inconclusive experiment by redesigning the hypothesis/MDE/sample-size up front.
        - prompt: "Cuántos usuarios necesito para detectar un 1.5pp de mejora con 80% de potencia?"
          why: Spanish trigger for sample-size computation from MDE and power.
      
      should_not_trigger:
        - prompt: "Build a dashboard tracking weekly signups and churn over time."
          route_to: analytics
          why: Recurring metric tracking and event instrumentation, not a controlled experiment with a hypothesis.
        - prompt: "Define our north-star metric and the KPI tree underneath it."
          route_to: kpi-framework
          why: Deciding which metrics matter, not designing or analyzing a test of a change.
        - prompt: "Forecast next quarter's revenue from the last 24 months of data."
          route_to: forecasting
          why: Time-series projection into the future, not a randomized comparison of two arms.
        - prompt: "Clean and dedupe this messy events CSV before any analysis."
          route_to: data-cleaning
          why: Data preparation, not experiment design or read-out.
        - prompt: "Turn last month's experiment results into a stakeholder slide deck."
          route_to: reporting
          why: Periodic stakeholder communication of results, not the statistical design/analysis itself.
      
      capability:
        - scenario: >
            Design, size, and plan the analysis for a test of a new one-click checkout button. Baseline
            purchase conversion is 12%; the team wants to detect a +1.5pp absolute lift at 80% power and
            alpha 0.05, with about 1,800 eligible users per day. Include a plan to speed it up with CUPED.
          must_include:
            - A falsifiable hypothesis with H0/H1 and exactly one primary metric (purchase conversion), plus guardrail metrics
            - Randomization unit equals analysis unit (by user), not by session
            - Sample-size calculation naming statsmodels proportion_effectsize and NormalIndPower / solve_power
            - Converting n per arm to a duration in days from daily eligible traffic, rounded up to whole weeks
            - A fixed-horizon stop rule and an explicit no-peeking warning (Type-I inflation)
            - An SRM chi-square check before trusting the result
            - CUPED plan with a pre-treatment covariate and the formula Y_cuped = Y - theta*(X - E[X])
            - Reporting the result as lift + confidence interval + p-value, compared against the MDE for practical significance
      
    • README.md 636 B
      # Evals — ab-testing
      
      These cases are routing and capability checks for the catalog harness, not an automated test runner.
      `should_trigger` and `should_not_trigger` are judged by feeding each prompt to the router and confirming
      it lands on `ab-testing` (or, for the negatives, on the named sibling such as `analytics` or
      `forecasting`). The single `capability` case is graded by hand or with the catalog's eval script: run the
      scenario through the skill and check the produced design/analysis against every line in `must_include` —
      a pass needs all of them present, not just most. There is no `pytest` here; the rubric is the spec.
      
  • references
    • pitfalls.md 4.4 KB
      # Pitfalls — why "inconclusive" tests usually mean a broken design
      
      This is the diagnosis layer. When a test "won't go significant" or "flip-flops," it is almost always one
      of the failure modes below, not bad luck.
      
      ## Peeking and optional stopping
      
      A fixed-horizon test is designed so that, under the null, you cross p < 0.05 about 5% of the time **when
      you look exactly once, at the planned n**. If you look every day and stop the first time you see p < 0.05,
      you take many independent-ish shots at that 5% door. With enough looks the cumulative chance of crossing
      0.05 under a true null climbs well past 5% — a pure-noise test can be "significant" the majority of the
      time if you peek long enough. The reported p-value is then a lie about a single planned look.
      
      Fixes, in order of preference:
      
      1. **Fixed horizon.** Commit to n/date; read once. Tightest read for declaring a winner.
      2. **Sequential / always-valid inference.** Confidence sequences and anytime-valid CIs (e.g. Netflix's
         anytime-valid approach) stay valid under continuous monitoring — you may look every day and the
         guarantee holds. Trade-off: wider intervals, so they are excellent for **stopping a clear loser early**
         and comparatively weak for **declaring a winner early**. Use them when early stopping has real value,
         not as a license to peek on a fixed-horizon design.
      
      Never: a fixed-horizon design read repeatedly and stopped at first significance.
      
      ## Multiple comparisons
      
      Every metric, every variant, every segment is another hypothesis test. Test 20 independent nulls at
      α = 0.05 and you expect ~1 false positive by chance alone. Two corrections:
      
      - **Bonferroni** — divide α by the number of tests. Controls the chance of *any* false positive
        (family-wise error). Simple and conservative; right for a small, pre-declared set of decision metrics.
      - **Benjamini-Hochberg (FDR)** — controls the *expected fraction* of false positives among the discoveries.
        Far more powerful on large exploratory scans. Illustrative example: scanning 20 metrics with true
        effects, an FDR procedure recovered ~17 while Bonferroni's stricter threshold recovered only ~12.
      
      Rule: Bonferroni for the handful of metrics a decision rests on; Benjamini-Hochberg for the big
      exploratory sweep where you can tolerate a known false-discovery rate.
      
      ## Sample ratio mismatch (SRM)
      
      The observed split materially diverges from the intended ratio (chi-square p < 0.001). This is not a
      statistical nuance — it is a signal the experiment is mechanically broken: a bot filter dropping one arm,
      a redirect, differential caching, a bucketing-hash bug, or logging that fires on one variant only.
      Consequence: the arms are no longer comparable, so the treatment effect is confounded. Do not adjust,
      reweight, or "note it as a caveat." Find the cause, fix instrumentation, and rerun. The SRM chi-square
      snippet is in `sample-size-and-cuped.md`.
      
      ## Underpowered ≠ "no effect"
      
      A non-significant result with a small sample means you could not have detected even a meaningful effect —
      absence of evidence, not evidence of absence. Before concluding "no difference," check whether the test
      reached its planned n and report the CI: "the effect is somewhere in [−0.8pp, +1.2pp]" is honest;
      "no effect" is not.
      
      ## Novelty and primacy effects
      
      A shiny new variant can spike because regulars click it out of curiosity (novelty) — or dip because they
      are disoriented by a changed UI (primacy). Both fade. If the effect is concentrated in the first days and
      decays, you are measuring reaction-to-change, not the steady-state effect. Run long enough for the curve
      to flatten, and inspect new-user vs returning-user segments separately.
      
      ## Simpson's paradox across segments
      
      An aggregate lift can reverse inside every segment (or vice versa) when segment mix differs between arms —
      often itself a symptom of SRM or a non-random assignment. If the overall result and the per-segment
      results disagree in direction, distrust the aggregate and find what is unbalanced between the arms before
      believing either number.
      
      ## HARKing — hypothesizing after results are known
      
      Slicing the data until something is significant, then writing the hypothesis to match, is a fishing
      expedition wearing a lab coat. Any "finding" discovered this way must be treated as a hypothesis for a
      *fresh* experiment, never as a confirmed result. Pre-register the hypothesis and primary metric before
      launch so the analysis you run is the analysis you committed to.
      
    • sample-size-and-cuped.md 4.7 KB
      # Sample size, duration, CUPED, and SRM — runnable snippets
      
      Versions assumed: **statsmodels 0.14.6** (stable), **scipy 1.17.1**. Install: `pip install "statsmodels>=0.14" "scipy>=1.15"`.
      
      All snippets are self-contained and print the numbers they compute.
      
      ## 1. Sample size for a conversion rate (proportion)
      
      `proportion_effectsize` applies the arcsine (Cohen's h) transform so the normal-approximation power
      calculation is valid for proportions.
      
      ```python
      import math
      from statsmodels.stats.power import NormalIndPower
      from statsmodels.stats.proportion import proportion_effectsize
      
      baseline = 0.12          # control conversion
      mde_abs  = 0.015         # minimum detectable effect, absolute (1.5 percentage points)
      alpha    = 0.05          # two-sided
      power    = 0.80
      
      variant = baseline + mde_abs
      h = proportion_effectsize(baseline, variant)
      n_per_arm = NormalIndPower().solve_power(effect_size=h, alpha=alpha, power=power, ratio=1.0)
      n_per_arm = math.ceil(n_per_arm)
      print(f"n per arm = {n_per_arm}, total = {2 * n_per_arm}")
      ```
      
      Worked numbers for the example above: h ≈ 0.0449, **n per arm ≈ 7,773**, total ≈ 15,546.
      
      ## 2. Sample size for a continuous metric
      
      Effect size is Cohen's d = (mean difference you care about) / (pooled standard deviation). Estimate the
      std from historical data on the same metric.
      
      ```python
      import math
      from statsmodels.stats.power import TTestIndPower
      
      mde_units  = 0.50        # e.g. detect a $0.50 lift in revenue per user
      pooled_std = 4.0         # historical std of revenue per user
      d = mde_units / pooled_std
      n_per_arm = math.ceil(TTestIndPower().solve_power(effect_size=d, alpha=0.05, power=0.80, ratio=1.0))
      print(f"n per arm = {n_per_arm}")
      ```
      
      ## 3. n → calendar duration
      
      ```python
      import math
      
      n_per_arm = 7773
      num_arms  = 2
      daily_eligible = 1800            # users entering the experiment per day
      
      days = math.ceil((n_per_arm * num_arms) / daily_eligible)
      print(f"raw days = {days}; run full weeks ->", math.ceil(days / 7) * 7)
      ```
      
      Round **up to whole weeks**. A test that mathematically needs 9 days should run 14: weekday and weekend
      users are different populations, and stopping mid-week oversamples one slice.
      
      ## 4. Analyze a finished proportion test
      
      ```python
      import numpy as np
      from statsmodels.stats.proportion import proportions_ztest, proportion_confint
      
      conv  = np.array([540, 590])     # successes: control, variant
      nobs  = np.array([10000, 10000]) # exposures per arm
      
      stat, pval = proportions_ztest(conv, nobs)
      p_c, p_v = conv / nobs
      lift_abs = p_v - p_c
      lift_rel = lift_abs / p_c
      # CI on the difference of two proportions (Wald)
      se = ((p_c*(1-p_c)/nobs[0]) + (p_v*(1-p_v)/nobs[1])) ** 0.5
      lo, hi = lift_abs - 1.96*se, lift_abs + 1.96*se
      print(f"control={p_c:.3%} variant={p_v:.3%} abs lift={lift_abs:+.3%} ({lift_rel:+.1%} rel)")
      print(f"p={pval:.4f}  95% CI on abs lift=[{lo:+.3%}, {hi:+.3%}]")
      ```
      
      Read it as a triple: lift, CI, p. Compare the CI to your MDE — if the interval includes effects below
      the MDE, you have not shown a *practically* meaningful result even if p < 0.05.
      
      ## 5. CUPED — θ via OLS, then test the adjusted metric
      
      ```python
      import numpy as np
      import statsmodels.api as sm
      from scipy import stats
      
      # y  = in-experiment metric per user (e.g. spend during the test)
      # x  = SAME metric measured in the pre-experiment window (must be pre-assignment)
      # arm = 0 control, 1 variant
      
      X = sm.add_constant(x)
      theta = sm.OLS(y, X).fit().params[1]          # slope = Cov(y,x)/Var(x)
      y_cuped = y - theta * (x - x.mean())          # variance-reduced metric
      
      t, p = stats.ttest_ind(y_cuped[arm == 1], y_cuped[arm == 0], equal_var=False)
      var_reduction = 1 - y_cuped.var() / y.var()
      print(f"theta={theta:.3f}  variance reduction={var_reduction:.1%}  t={t:.2f} p={p:.4f}")
      ```
      
      `var_reduction` is the fraction of variance the covariate explained — that is the power you bought for
      free. Zero or negative means the covariate is useless here (new users, or `x` uncorrelated with `y`):
      drop CUPED, it cannot help and a post-assignment covariate would actively bias the estimate.
      
      ## 6. Sample ratio mismatch (SRM) gate
      
      Run this *before* trusting any result. A significant deviation from the intended split means the
      experiment is broken at the assignment/logging layer.
      
      ```python
      from scipy import stats
      
      observed = [10000, 9430]          # users actually seen in each arm
      expected_ratio = [0.5, 0.5]       # intended split
      total = sum(observed)
      expected = [total * r for r in expected_ratio]
      
      chi2, p = stats.chisquare(f_obs=observed, f_exp=expected)
      print(f"chi2={chi2:.1f} p={p:.6f}", "-> SRM, do not trust results" if p < 0.001 else "-> split OK")
      ```
      
      If `p < 0.001`: stop, find why one arm is short (bot filtering, redirect, caching, a broken bucketing
      hash), fix instrumentation, and rerun. Never reweight your way out of an SRM.
      
  • scripts
    • verify.sh 3.8 KB
      #!/usr/bin/env bash
      set -euo pipefail
      
      # verify.sh — ab-testing skill gate. Run from your PROJECT root.
      #
      # What it does (read-only, idempotent, NEVER runs anything destructive):
      #   1. Discovers experiment artifacts produced by this skill:
      #        - Python sizing/analysis scripts whose names hint at experiments
      #          (*sample*size*.py, *ab*test*.py, *experiment*.py, *power*.py, *cuped*.py, *srm*.py)
      #        - experiment-design docs (*experiment*.md, *ab*test*.md, *design*.md)
      #   2. For each Python artifact: checks scipy + statsmodels import, then executes it under
      #      python3 and asserts it prints at least one integer (a sample size / count).  [warn-only]
      #   3. For each design doc: greps for a primary metric, an MDE, and power/alpha.   [warn-only]
      #
      # Exit code: non-zero ONLY when a discovered Python artifact crashes (non-zero exit) under an
      # environment that HAS python3 + scipy + statsmodels. Missing interpreter/libs -> skip, never fail.
      # Missing fields in a doc -> advisory warn. An empty or clean target exits 0 (no false failure).
      # Stock macOS bash 3.2 compatible (no mapfile, no associative arrays).
      
      YELLOW=$'\033[33m'; GREEN=$'\033[32m'; RED=$'\033[31m'; NC=$'\033[0m'
      EXIT=0
      skip() { printf '%s[skip]%s %s\n' "$YELLOW" "$NC" "$*"; }
      note() { printf '%s[warn]%s %s\n' "$YELLOW" "$NC" "$*"; }
      ok()   { printf '%s[ok]%s %s\n'   "$GREEN"  "$NC" "$*"; }
      err()  { printf '%s[fail]%s %s\n' "$RED"    "$NC" "$*"; EXIT=1; }
      
      ROOT="$(pwd)"
      
      # --- discover python sizing/analysis artifacts ---
      PY_FILES=()
      while IFS= read -r -d '' f; do
        PY_FILES+=("$f")
      done < <(
        find "$ROOT" \
          \( -path '*/node_modules/*' -o -path '*/.git/*' -o -path '*/.venv/*' -o -path '*/venv/*' -o -path '*/site-packages/*' \) -prune -o \
          -type f \( -iname '*sample*size*.py' -o -iname '*ab*test*.py' -o -iname '*experiment*.py' \
                     -o -iname '*power*.py' -o -iname '*cuped*.py' -o -iname '*srm*.py' \) -print0 2>/dev/null
      )
      
      # --- discover experiment-design docs ---
      DOC_FILES=()
      while IFS= read -r -d '' f; do
        DOC_FILES+=("$f")
      done < <(
        find "$ROOT" \
          \( -path '*/node_modules/*' -o -path '*/.git/*' -o -path '*/.venv/*' \) -prune -o \
          -type f \( -iname '*experiment*.md' -o -iname '*ab*test*.md' -o -iname '*design*.md' \) -print0 2>/dev/null
      )
      
      if [ "${#PY_FILES[@]}" -eq 0 ] && [ "${#DOC_FILES[@]}" -eq 0 ]; then
        skip "no experiment scripts or design docs found under $ROOT — design-only conversation"
        ok "verify.sh passed (empty target)"
        exit 0
      fi
      
      # --- python artifacts ---
      if [ "${#PY_FILES[@]}" -gt 0 ]; then
        if ! command -v python3 >/dev/null 2>&1; then
          skip "python3 not installed — cannot execute sizing scripts"
        elif ! python3 -c 'import scipy, statsmodels' >/dev/null 2>&1; then
          skip "scipy/statsmodels not importable — install with: pip install scipy statsmodels"
        else
          for f in "${PY_FILES[@]}"; do
            out="$(python3 "$f" 2>&1)" || { err "$f: crashed under python3"; printf '%s\n' "$out" | sed 's/^/    /'; continue; }
            if printf '%s' "$out" | grep -Eq '[0-9]+'; then
              ok "$f: executed and printed a number"
            else
              note "$f: ran but printed no numeric output — a sizing/analysis script should print an n or p-value"
            fi
          done
        fi
      fi
      
      # --- design docs ---
      for f in "${DOC_FILES[@]}"; do
        miss=""
        grep -Eiq 'primary[[:space:]]+metric|metric[[:space:]]*:'                  "$f" || miss="$miss primary-metric"
        grep -Eiq 'mde|minimum[[:space:]]+detectable[[:space:]]+effect'           "$f" || miss="$miss MDE"
        grep -Eiq 'power|alpha|significance[[:space:]]+level'                      "$f" || miss="$miss power/alpha"
        if [ -n "$miss" ]; then
          note "$f: experiment-design doc missing:$miss"
        else
          ok "$f: names a primary metric, an MDE, and power/alpha"
        fi
      done
      
      printf '\n'
      if [ "$EXIT" -eq 0 ]; then ok "verify.sh passed"; else err "verify.sh found failures"; fi
      exit "$EXIT"
      
  • SKILL.md 9.6 KB
    ---
    name: ab-testing
    description: "Use when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go significant. NOT recurring metric tracking (that is `analytics`), NOT north-star/KPI trees (that is `kpi-framework`), NOT projecting metrics forward (that is `forecasting`)."
    tags: [ab-testing, experimentation, statistics, cuped, sample-size, hypothesis-testing]
    recommends: [analytics, kpi-framework, forecasting, data-cleaning, python, reporting]
    origin: risco
    ---
    
    # A/B testing — design and read a defensible experiment
    
    An experiment without a pre-committed sample size and a single primary metric is not an experiment.
    It is a dashboard you stare at until it tells you what you wanted to hear. The discipline lives almost
    entirely *before* traffic ships: a falsifiable hypothesis, one primary metric, a sample size derived
    from the smallest effect worth detecting, and a stop rule you cannot renegotiate at 2pm on day four.
    
    ## Pre-test checklist — every line true before any traffic
    
    Each one is a place experiments die silently.
    
    - [ ] A **falsifiable hypothesis** — names the change, the direction, and the metric it moves.
    - [ ] Exactly **ONE primary metric**. More than one primary = multiple comparisons = inflated false positives.
    - [ ] **Guardrail metrics** — what you refuse to harm (latency, refunds, unsubscribes) even for a win.
    - [ ] The **randomization unit = the analysis unit** (usually the user). Mixing them is pseudoreplication.
    - [ ] An **MDE** — the smallest lift that would change a decision. Not "any difference."
    - [ ] A **computed sample size** and the **duration** it implies at your real daily eligible traffic.
    - [ ] A **fixed stop rule** — a date or an n you commit to before launch. No "we'll see how it looks."
    
    ## Step 1 — Hypothesis and metrics
    
    State a null you can reject. "The new checkout button changes purchase conversion" with H0: conversion
    equal across arms, H1: it differs. Vague aspirations ("improve the funnel") have no rejection region.
    
    Pick one primary metric and freeze it. Why: every extra primary metric is another coin flip at α, so
    three "primary" metrics turn a 5% false-positive rate into roughly 14%. Demote the rest to secondary.
    
    Randomize on the same unit you analyze on. If a user sees the variant on every visit, randomize by user,
    not by session — analyzing 50k sessions from 8k users treats correlated observations as independent and
    fabricates significance.
    
    ```text
    Bad:  "We think the redesign will improve engagement and revenue and retention."  (no null, 3 primaries, no number)
    Good: "H0: 30-day purchase conversion is equal between control and the new one-click button.
           H1: it differs. Primary: purchase conversion. Guardrails: refund rate, p95 checkout latency.
           Randomize by user_id. MDE: +1.5pp absolute on a 12% baseline."
    ```
    
    ## Step 2 — Sample size from MDE, baseline, and power
    
    Defaults: power 0.80, α 0.05 (two-sided). The MDE is yours to choose — it is the smallest effect that
    would actually change what you do.
    
    Rule: required n scales with ~1/MDE². Why: halving the smallest effect you care to detect roughly
    **quadruples** the traffic and time. This is the single most expensive decision in the design, so set the
    MDE to a business threshold, never to "whatever is small."
    
    For a conversion rate (proportion):
    
    ```python
    from statsmodels.stats.power import NormalIndPower
    from statsmodels.stats.proportion import proportion_effectsize
    
    p1, p2 = 0.12, 0.135                       # baseline, baseline + MDE (1.5pp)
    h = proportion_effectsize(p1, p2)          # Cohen's h (arcsine transform)
    n = NormalIndPower().solve_power(effect_size=h, alpha=0.05, power=0.80, ratio=1.0)
    print(int(-(-n // 1)))                      # n PER ARM, rounded up
    ```
    
    For a continuous metric (revenue per user, time on page) use Welch-style sizing:
    
    ```python
    from statsmodels.stats.power import TTestIndPower
    
    effect = mde_in_units / pooled_std         # Cohen's d
    n = TTestIndPower().solve_power(effect_size=effect, alpha=0.05, power=0.80, ratio=1.0)
    ```
    
    Then convert n to a calendar plan: `days = ceil((n_per_arm * num_arms) / daily_eligible_users)`. If that
    is 9 days, run a clean **two full weeks** anyway — weekday/weekend mix is part of the population, and a
    6-day test oversamples whoever shows up Tuesday. Full worked example (12% baseline, +1.5pp MDE, 80%
    power) plus runnable sizing, n→duration, CUPED θ and SRM snippets: `references/sample-size-and-cuped.md`.
    
    ## Step 3 — Run discipline
    
    **Fixed horizon is the default.** Commit to the n/date from Step 2 and read the result once, at the end.
    
    **Do not peek and stop at first significance.** Why: checking repeatedly and stopping the moment p < 0.05
    inflates the Type-I error far above 5% — with enough looks, a null test crosses 0.05 most of the time.
    If you genuinely need to stop early, use a *sequential / always-valid* method (confidence sequences,
    e.g. Netflix's anytime-valid CIs) that holds Type-I error under continuous monitoring. Sequential is
    strong for **killing losers early** and weak for **calling winners early** — for a confident win, the
    fixed-horizon read is tighter.
    
    **Gate on SRM before you trust anything.** Compute a chi-square test on the observed split versus the
    intended ratio. If p < 0.001 the assignment or logging is broken — a bot filter dropping one arm, a
    redirect, a caching bug. Fix the instrumentation and rerun; do not "adjust for it."
    
    The peeking Type-I math, sequential/always-valid options, SRM diagnosis, novelty/primacy effects,
    Simpson's paradox in segments and HARKing all live in `references/pitfalls.md`.
    
    ## Step 4 — Analyze
    
    Pick the test by metric type:
    
    | Metric type | Test |
    |---|---|
    | Binary conversion (proportion) | Two-proportion z-test (`statsmodels.stats.proportion.proportions_ztest`) |
    | Continuous, roughly normal / large n | Welch's t-test (`scipy.stats.ttest_ind(..., equal_var=False)`) |
    | Continuous, heavy-tailed / skewed (revenue) | Mann-Whitney U, or t-test on a log/winsorized metric |
    
    Report **lift + confidence interval + p-value together**. Never p alone. Why: p < 0.05 with a CI of
    [+0.1pp, +5pp] is "statistically there, practically a coin toss" — the CI tells you the size, p only
    tells you it is not exactly zero. **Practical significance** = compare the CI to your MDE: if the whole
    interval sits above the MDE, ship; if it straddles the MDE, you detected *something* too small to matter.
    
    **Multiple comparisons.** Two regimes:
    - Small set of pre-declared **decision** metrics → **Bonferroni** (divide α by the count). Conservative, simple.
    - Large **exploratory** scan of many metrics/segments → **Benjamini-Hochberg (FDR)**. It keeps far more
      power than Bonferroni on big scans (in a 20-effect example, ~17 detected vs ~12 under Bonferroni).
    
    ## Step 5 — CUPED variance reduction
    
    CUPED (Controlled-experiment Using Pre-Experiment Data) subtracts predictable pre-period noise so the
    same traffic buys more power — or the same power needs less traffic. The adjusted metric:
    
    ```text
    Y_cuped = Y − θ · (X − E[X])        where  θ = Cov(Y, X) / Var(X)
    ```
    
    Estimate θ by regressing the in-experiment metric `Y` on the **pre-experiment** covariate `X` (e.g. each
    user's spend in the 4 weeks before the test), then analyze `Y_cuped` with the same test as Step 4.
    
    When it pays: recurring users with a strong pre-period signal. Reported wins — Netflix ~40% variance
    reduction on engagement, Statsig 50%+ on common metrics → significance in roughly half the time/traffic.
    
    When it does **nothing** — do not bother: brand-new users (no pre-period data), a covariate uncorrelated
    with the outcome, or — the cardinal sin — a covariate measured *after* assignment, which biases the
    estimate. The covariate MUST be pre-treatment and independent of which arm a user lands in. Runnable
    θ-via-OLS snippet in `references/sample-size-and-cuped.md`.
    
    ## Anti-patterns
    
    | Bad | Why it is wrong | Do instead |
    |---|---|---|
    | Peek daily, stop the day p < 0.05 | Repeated looks inflate Type-I error far above α | Fix n/date up front; or a sequential method that holds α |
    | No sample size set before launch | You will stop on noise and call it a win | Compute n from MDE/baseline/power in Step 2 |
    | Several "primary" metrics | Each is a coin flip at α; 3 metrics ≈ 14% false-positive | One frozen primary; the rest are secondary |
    | Ignore the observed split | An SRM means assignment/logging is broken; results are garbage | Chi-square SRM gate before reading anything |
    | Report only the p-value | Hides effect size — p < 0.05 can be practically zero | Always lift + CI + p; compare CI to MDE |
    | CUPED on a post-assignment covariate | Covariate correlated with the arm biases θ | Use only pre-treatment, assignment-independent covariates |
    | Call a winner from an underpowered test | "Not significant" then ≠ "no effect"; you lacked power | Reach planned n, or report the CI and say "inconclusive, here is the range" |
    | Decide the hypothesis after seeing results (HARKing) | Turns the whole analysis into a fishing expedition | Pre-register hypothesis + primary metric before launch |
    | Run 6 days because it "looks significant" | Oversamples one weekday slice of the population | Run full weeks; honor the fixed horizon |
    
    ## Checkable artifact
    
    When this skill emits a Python sizing/analysis script or an experiment-design doc, run
    `scripts/verify.sh` from your project root. It confirms the script executes under `python3` and prints a
    numeric sample size, and that any design doc names a primary metric, an MDE, and power/alpha. It is
    read-only and soft-passes when no artifact is present (a design-only conversation).
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related