ab-testing
Use when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go significant. NOT recurring metric tracking (that is `analytics`), NOT north-star/KPI trees (that is `kp
Install
npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/ab-testing
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
git clone https://github.com/ericrisco/rsc-harness.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
A/B testing — design and read a defensible experiment
An experiment without a pre-committed sample size and a single primary metric is not an experiment. It is a dashboard you stare at until it tells you what you wanted to hear. The discipline lives almost entirely before traffic ships: a falsifiable hypothesis, one primary metric, a sample size derived from the smallest effect worth detecting, and a stop rule you cannot renegotiate at 2pm on day four.
Pre-test checklist — every line true before any traffic
Each one is a place experiments die silently.
- A falsifiable hypothesis — names the change, the direction, and the metric it moves.
- Exactly ONE primary metric. More than one primary = multiple comparisons = inflated false positives.
- Guardrail metrics — what you refuse to harm (latency, refunds, unsubscribes) even for a win.
- The randomization unit = the analysis unit (usually the user). Mixing them is pseudoreplication.
- An MDE — the smallest lift that would change a decision. Not "any difference."
- A computed sample size and the duration it implies at your real daily eligible traffic.
- A fixed stop rule — a date or an n you commit to before launch. No "we'll see how it looks."
Step 1 — Hypothesis and metrics
State a null you can reject. "The new checkout button changes purchase conversion" with H0: conversion equal across arms, H1: it differs. Vague aspirations ("improve the funnel") have no rejection region.
Pick one primary metric and freeze it. Why: every extra primary metric is another coin flip at α, so three "primary" metrics turn a 5% false-positive rate into roughly 14%. Demote the rest to secondary.
Randomize on the same unit you analyze on. If a user sees the variant on every visit, randomize by user, not by session — analyzing 50k sessions from 8k users treats correlated observations as independent and fabricates significance.
Bad: "We think the redesign will improve engagement and revenue and retention." (no null, 3 primaries, no number)
Good: "H0: 30-day purchase conversion is equal between control and the new one-click button.
H1: it differs. Primary: purchase conversion. Guardrails: refund rate, p95 checkout latency.
Randomize by user_id. MDE: +1.5pp absolute on a 12% baseline."
Step 2 — Sample size from MDE, baseline, and power
Defaults: power 0.80, α 0.05 (two-sided). The MDE is yours to choose — it is the smallest effect that would actually change what you do.
Rule: required n scales with ~1/MDE². Why: halving the smallest effect you care to detect roughly quadruples the traffic and time. This is the single most expensive decision in the design, so set the MDE to a business threshold, never to "whatever is small."
For a conversion rate (proportion):
from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import proportion_effectsize
p1, p2 = 0.12, 0.135 # baseline, baseline + MDE (1.5pp)
h = proportion_effectsize(p1, p2) # Cohen's h (arcsine transform)
n = NormalIndPower().solve_power(effect_size=h, alpha=0.05, power=0.80, ratio=1.0)
print(int(-(-n // 1))) # n PER ARM, rounded up
For a continuous metric (revenue per user, time on page) use Welch-style sizing:
from statsmodels.stats.power import TTestIndPower
effect = mde_in_units / pooled_std # Cohen's d
n = TTestIndPower().solve_power(effect_size=effect, alpha=0.05, power=0.80, ratio=1.0)
Then convert n to a calendar plan: days = ceil((n_per_arm * num_arms) / daily_eligible_users). If that
is 9 days, run a clean two full weeks anyway — weekday/weekend mix is part of the population, and a
6-day test oversamples whoever shows up Tuesday. Full worked example (12% baseline, +1.5pp MDE, 80%
power) plus runnable sizing, n→duration, CUPED θ and SRM snippets: references/sample-size-and-cuped.md.
Step 3 — Run discipline
Fixed horizon is the default. Commit to the n/date from Step 2 and read the result once, at the end.
Do not peek and stop at first significance. Why: checking repeatedly and stopping the moment p < 0.05 inflates the Type-I error far above 5% — with enough looks, a null test crosses 0.05 most of the time. If you genuinely need to stop early, use a sequential / always-valid method (confidence sequences, e.g. Netflix's anytime-valid CIs) that holds Type-I error under continuous monitoring. Sequential is strong for killing losers early and weak for calling winners early — for a confident win, the fixed-horizon read is tighter.
Gate on SRM before you trust anything. Compute a chi-square test on the observed split versus the intended ratio. If p < 0.001 the assignment or logging is broken — a bot filter dropping one arm, a redirect, a caching bug. Fix the instrumentation and rerun; do not "adjust for it."
The peeking Type-I math, sequential/always-valid options, SRM diagnosis, novelty/primacy effects,
Simpson's paradox in segments and HARKing all live in references/pitfalls.md.
Step 4 — Analyze
Pick the test by metric type:
| Metric type | Test |
|---|---|
| Binary conversion (proportion) | Two-proportion z-test (statsmodels.stats.proportion.proportions_ztest) |
| Continuous, roughly normal / large n | Welch's t-test (scipy.stats.ttest_ind(..., equal_var=False)) |
| Continuous, heavy-tailed / skewed (revenue) | Mann-Whitney U, or t-test on a log/winsorized metric |
Report lift + confidence interval + p-value together. Never p alone. Why: p < 0.05 with a CI of [+0.1pp, +5pp] is "statistically there, practically a coin toss" — the CI tells you the size, p only tells you it is not exactly zero. Practical significance = compare the CI to your MDE: if the whole interval sits above the MDE, ship; if it straddles the MDE, you detected something too small to matter.
Multiple comparisons. Two regimes:
- Small set of pre-declared decision metrics → Bonferroni (divide α by the count). Conservative, simple.
- Large exploratory scan of many metrics/segments → Benjamini-Hochberg (FDR). It keeps far more power than Bonferroni on big scans (in a 20-effect example, ~17 detected vs ~12 under Bonferroni).
Step 5 — CUPED variance reduction
CUPED (Controlled-experiment Using Pre-Experiment Data) subtracts predictable pre-period noise so the same traffic buys more power — or the same power needs less traffic. The adjusted metric:
Y_cuped = Y − θ · (X − E[X]) where θ = Cov(Y, X) / Var(X)
Estimate θ by regressing the in-experiment metric Y on the pre-experiment covariate X (e.g. each
user's spend in the 4 weeks before the test), then analyze Y_cuped with the same test as Step 4.
When it pays: recurring users with a strong pre-period signal. Reported wins — Netflix ~40% variance reduction on engagement, Statsig 50%+ on common metrics → significance in roughly half the time/traffic.
When it does nothing — do not bother: brand-new users (no pre-period data), a covariate uncorrelated
with the outcome, or — the cardinal sin — a covariate measured after assignment, which biases the
estimate. The covariate MUST be pre-treatment and independent of which arm a user lands in. Runnable
θ-via-OLS snippet in references/sample-size-and-cuped.md.
Anti-patterns
| Bad | Why it is wrong | Do instead |
|---|---|---|
| Peek daily, stop the day p < 0.05 | Repeated looks inflate Type-I error far above α | Fix n/date up front; or a sequential method that holds α |
| No sample size set before launch | You will stop on noise and call it a win | Compute n from MDE/baseline/power in Step 2 |
| Several "primary" metrics | Each is a coin flip at α; 3 metrics ≈ 14% false-positive | One frozen primary; the rest are secondary |
| Ignore the observed split | An SRM means assignment/logging is broken; results are garbage | Chi-square SRM gate before reading anything |
| Report only the p-value | Hides effect size — p < 0.05 can be practically zero | Always lift + CI + p; compare CI to MDE |
| CUPED on a post-assignment covariate | Covariate correlated with the arm biases θ | Use only pre-treatment, assignment-independent covariates |
| Call a winner from an underpowered test | "Not significant" then ≠ "no effect"; you lacked power | Reach planned n, or report the CI and say "inconclusive, here is the range" |
| Decide the hypothesis after seeing results (HARKing) | Turns the whole analysis into a fishing expedition | Pre-register hypothesis + primary metric before launch |
| Run 6 days because it "looks significant" | Oversamples one weekday slice of the population | Run full weeks; honor the fixed horizon |
Checkable artifact
When this skill emits a Python sizing/analysis script or an experiment-design doc, run
scripts/verify.sh from your project root. It confirms the script executes under python3 and prints a
numeric sample size, and that any design doc names a primary metric, an MDE, and power/alpha. It is
read-only and soft-passes when no artifact is present (a design-only conversation).
Files (rsc-harness)
-
evals
-
cases.yaml 3.5 KB
skill: ab-testing should_trigger: - prompt: "How many users per arm do I need to detect a 2% lift on a 10% baseline conversion?" why: Core sample-size sizing from baseline + MDE + power — the central pre-test calculation this skill owns. - prompt: "Is this A/B result statistically significant? Control converted 540/10000, variant 590/10000." why: Read-out of a finished proportion test — two-proportion z-test, lift, CI, p-value together. - prompt: "Our test has run two weeks and still isn't significant, what do we do now?" why: Non-obvious. "Inconclusive" almost always means underpowered, peeked, or an SRM — a design diagnosis, not a re-run. - prompt: "Can we use CUPED to get faster results on our checkout experiment?" why: Variance reduction via a pre-experiment covariate — when it helps, the formula, and when it does nothing. - prompt: "We keep peeking at the dashboard and stopping the moment it hits p<0.05 — is that a problem?" why: Non-obvious optional-stopping trap; peeking inflates Type-I error far above alpha. Fixed horizon or always-valid inference. - prompt: "El experiment de la landing no surt significatiu, com el dissenyo bé des de zero?" why: Catalan trigger — rescue an inconclusive experiment by redesigning the hypothesis/MDE/sample-size up front. - prompt: "Cuántos usuarios necesito para detectar un 1.5pp de mejora con 80% de potencia?" why: Spanish trigger for sample-size computation from MDE and power. should_not_trigger: - prompt: "Build a dashboard tracking weekly signups and churn over time." route_to: analytics why: Recurring metric tracking and event instrumentation, not a controlled experiment with a hypothesis. - prompt: "Define our north-star metric and the KPI tree underneath it." route_to: kpi-framework why: Deciding which metrics matter, not designing or analyzing a test of a change. - prompt: "Forecast next quarter's revenue from the last 24 months of data." route_to: forecasting why: Time-series projection into the future, not a randomized comparison of two arms. - prompt: "Clean and dedupe this messy events CSV before any analysis." route_to: data-cleaning why: Data preparation, not experiment design or read-out. - prompt: "Turn last month's experiment results into a stakeholder slide deck." route_to: reporting why: Periodic stakeholder communication of results, not the statistical design/analysis itself. capability: - scenario: > Design, size, and plan the analysis for a test of a new one-click checkout button. Baseline purchase conversion is 12%; the team wants to detect a +1.5pp absolute lift at 80% power and alpha 0.05, with about 1,800 eligible users per day. Include a plan to speed it up with CUPED. must_include: - A falsifiable hypothesis with H0/H1 and exactly one primary metric (purchase conversion), plus guardrail metrics - Randomization unit equals analysis unit (by user), not by session - Sample-size calculation naming statsmodels proportion_effectsize and NormalIndPower / solve_power - Converting n per arm to a duration in days from daily eligible traffic, rounded up to whole weeks - A fixed-horizon stop rule and an explicit no-peeking warning (Type-I inflation) - An SRM chi-square check before trusting the result - CUPED plan with a pre-treatment covariate and the formula Y_cuped = Y - theta*(X - E[X]) - Reporting the result as lift + confidence interval + p-value, compared against the MDE for practical significance -
README.md 636 B
# Evals — ab-testing These cases are routing and capability checks for the catalog harness, not an automated test runner. `should_trigger` and `should_not_trigger` are judged by feeding each prompt to the router and confirming it lands on `ab-testing` (or, for the negatives, on the named sibling such as `analytics` or `forecasting`). The single `capability` case is graded by hand or with the catalog's eval script: run the scenario through the skill and check the produced design/analysis against every line in `must_include` — a pass needs all of them present, not just most. There is no `pytest` here; the rubric is the spec.
-
-
references
-
pitfalls.md 4.4 KB
# Pitfalls — why "inconclusive" tests usually mean a broken design This is the diagnosis layer. When a test "won't go significant" or "flip-flops," it is almost always one of the failure modes below, not bad luck. ## Peeking and optional stopping A fixed-horizon test is designed so that, under the null, you cross p < 0.05 about 5% of the time **when you look exactly once, at the planned n**. If you look every day and stop the first time you see p < 0.05, you take many independent-ish shots at that 5% door. With enough looks the cumulative chance of crossing 0.05 under a true null climbs well past 5% — a pure-noise test can be "significant" the majority of the time if you peek long enough. The reported p-value is then a lie about a single planned look. Fixes, in order of preference: 1. **Fixed horizon.** Commit to n/date; read once. Tightest read for declaring a winner. 2. **Sequential / always-valid inference.** Confidence sequences and anytime-valid CIs (e.g. Netflix's anytime-valid approach) stay valid under continuous monitoring — you may look every day and the guarantee holds. Trade-off: wider intervals, so they are excellent for **stopping a clear loser early** and comparatively weak for **declaring a winner early**. Use them when early stopping has real value, not as a license to peek on a fixed-horizon design. Never: a fixed-horizon design read repeatedly and stopped at first significance. ## Multiple comparisons Every metric, every variant, every segment is another hypothesis test. Test 20 independent nulls at α = 0.05 and you expect ~1 false positive by chance alone. Two corrections: - **Bonferroni** — divide α by the number of tests. Controls the chance of *any* false positive (family-wise error). Simple and conservative; right for a small, pre-declared set of decision metrics. - **Benjamini-Hochberg (FDR)** — controls the *expected fraction* of false positives among the discoveries. Far more powerful on large exploratory scans. Illustrative example: scanning 20 metrics with true effects, an FDR procedure recovered ~17 while Bonferroni's stricter threshold recovered only ~12. Rule: Bonferroni for the handful of metrics a decision rests on; Benjamini-Hochberg for the big exploratory sweep where you can tolerate a known false-discovery rate. ## Sample ratio mismatch (SRM) The observed split materially diverges from the intended ratio (chi-square p < 0.001). This is not a statistical nuance — it is a signal the experiment is mechanically broken: a bot filter dropping one arm, a redirect, differential caching, a bucketing-hash bug, or logging that fires on one variant only. Consequence: the arms are no longer comparable, so the treatment effect is confounded. Do not adjust, reweight, or "note it as a caveat." Find the cause, fix instrumentation, and rerun. The SRM chi-square snippet is in `sample-size-and-cuped.md`. ## Underpowered ≠ "no effect" A non-significant result with a small sample means you could not have detected even a meaningful effect — absence of evidence, not evidence of absence. Before concluding "no difference," check whether the test reached its planned n and report the CI: "the effect is somewhere in [−0.8pp, +1.2pp]" is honest; "no effect" is not. ## Novelty and primacy effects A shiny new variant can spike because regulars click it out of curiosity (novelty) — or dip because they are disoriented by a changed UI (primacy). Both fade. If the effect is concentrated in the first days and decays, you are measuring reaction-to-change, not the steady-state effect. Run long enough for the curve to flatten, and inspect new-user vs returning-user segments separately. ## Simpson's paradox across segments An aggregate lift can reverse inside every segment (or vice versa) when segment mix differs between arms — often itself a symptom of SRM or a non-random assignment. If the overall result and the per-segment results disagree in direction, distrust the aggregate and find what is unbalanced between the arms before believing either number. ## HARKing — hypothesizing after results are known Slicing the data until something is significant, then writing the hypothesis to match, is a fishing expedition wearing a lab coat. Any "finding" discovered this way must be treated as a hypothesis for a *fresh* experiment, never as a confirmed result. Pre-register the hypothesis and primary metric before launch so the analysis you run is the analysis you committed to. -
sample-size-and-cuped.md 4.7 KB
# Sample size, duration, CUPED, and SRM — runnable snippets Versions assumed: **statsmodels 0.14.6** (stable), **scipy 1.17.1**. Install: `pip install "statsmodels>=0.14" "scipy>=1.15"`. All snippets are self-contained and print the numbers they compute. ## 1. Sample size for a conversion rate (proportion) `proportion_effectsize` applies the arcsine (Cohen's h) transform so the normal-approximation power calculation is valid for proportions. ```python import math from statsmodels.stats.power import NormalIndPower from statsmodels.stats.proportion import proportion_effectsize baseline = 0.12 # control conversion mde_abs = 0.015 # minimum detectable effect, absolute (1.5 percentage points) alpha = 0.05 # two-sided power = 0.80 variant = baseline + mde_abs h = proportion_effectsize(baseline, variant) n_per_arm = NormalIndPower().solve_power(effect_size=h, alpha=alpha, power=power, ratio=1.0) n_per_arm = math.ceil(n_per_arm) print(f"n per arm = {n_per_arm}, total = {2 * n_per_arm}") ``` Worked numbers for the example above: h ≈ 0.0449, **n per arm ≈ 7,773**, total ≈ 15,546. ## 2. Sample size for a continuous metric Effect size is Cohen's d = (mean difference you care about) / (pooled standard deviation). Estimate the std from historical data on the same metric. ```python import math from statsmodels.stats.power import TTestIndPower mde_units = 0.50 # e.g. detect a $0.50 lift in revenue per user pooled_std = 4.0 # historical std of revenue per user d = mde_units / pooled_std n_per_arm = math.ceil(TTestIndPower().solve_power(effect_size=d, alpha=0.05, power=0.80, ratio=1.0)) print(f"n per arm = {n_per_arm}") ``` ## 3. n → calendar duration ```python import math n_per_arm = 7773 num_arms = 2 daily_eligible = 1800 # users entering the experiment per day days = math.ceil((n_per_arm * num_arms) / daily_eligible) print(f"raw days = {days}; run full weeks ->", math.ceil(days / 7) * 7) ``` Round **up to whole weeks**. A test that mathematically needs 9 days should run 14: weekday and weekend users are different populations, and stopping mid-week oversamples one slice. ## 4. Analyze a finished proportion test ```python import numpy as np from statsmodels.stats.proportion import proportions_ztest, proportion_confint conv = np.array([540, 590]) # successes: control, variant nobs = np.array([10000, 10000]) # exposures per arm stat, pval = proportions_ztest(conv, nobs) p_c, p_v = conv / nobs lift_abs = p_v - p_c lift_rel = lift_abs / p_c # CI on the difference of two proportions (Wald) se = ((p_c*(1-p_c)/nobs[0]) + (p_v*(1-p_v)/nobs[1])) ** 0.5 lo, hi = lift_abs - 1.96*se, lift_abs + 1.96*se print(f"control={p_c:.3%} variant={p_v:.3%} abs lift={lift_abs:+.3%} ({lift_rel:+.1%} rel)") print(f"p={pval:.4f} 95% CI on abs lift=[{lo:+.3%}, {hi:+.3%}]") ``` Read it as a triple: lift, CI, p. Compare the CI to your MDE — if the interval includes effects below the MDE, you have not shown a *practically* meaningful result even if p < 0.05. ## 5. CUPED — θ via OLS, then test the adjusted metric ```python import numpy as np import statsmodels.api as sm from scipy import stats # y = in-experiment metric per user (e.g. spend during the test) # x = SAME metric measured in the pre-experiment window (must be pre-assignment) # arm = 0 control, 1 variant X = sm.add_constant(x) theta = sm.OLS(y, X).fit().params[1] # slope = Cov(y,x)/Var(x) y_cuped = y - theta * (x - x.mean()) # variance-reduced metric t, p = stats.ttest_ind(y_cuped[arm == 1], y_cuped[arm == 0], equal_var=False) var_reduction = 1 - y_cuped.var() / y.var() print(f"theta={theta:.3f} variance reduction={var_reduction:.1%} t={t:.2f} p={p:.4f}") ``` `var_reduction` is the fraction of variance the covariate explained — that is the power you bought for free. Zero or negative means the covariate is useless here (new users, or `x` uncorrelated with `y`): drop CUPED, it cannot help and a post-assignment covariate would actively bias the estimate. ## 6. Sample ratio mismatch (SRM) gate Run this *before* trusting any result. A significant deviation from the intended split means the experiment is broken at the assignment/logging layer. ```python from scipy import stats observed = [10000, 9430] # users actually seen in each arm expected_ratio = [0.5, 0.5] # intended split total = sum(observed) expected = [total * r for r in expected_ratio] chi2, p = stats.chisquare(f_obs=observed, f_exp=expected) print(f"chi2={chi2:.1f} p={p:.6f}", "-> SRM, do not trust results" if p < 0.001 else "-> split OK") ``` If `p < 0.001`: stop, find why one arm is short (bot filtering, redirect, caching, a broken bucketing hash), fix instrumentation, and rerun. Never reweight your way out of an SRM.
-
-
scripts
-
verify.sh 3.8 KB
#!/usr/bin/env bash set -euo pipefail # verify.sh — ab-testing skill gate. Run from your PROJECT root. # # What it does (read-only, idempotent, NEVER runs anything destructive): # 1. Discovers experiment artifacts produced by this skill: # - Python sizing/analysis scripts whose names hint at experiments # (*sample*size*.py, *ab*test*.py, *experiment*.py, *power*.py, *cuped*.py, *srm*.py) # - experiment-design docs (*experiment*.md, *ab*test*.md, *design*.md) # 2. For each Python artifact: checks scipy + statsmodels import, then executes it under # python3 and asserts it prints at least one integer (a sample size / count). [warn-only] # 3. For each design doc: greps for a primary metric, an MDE, and power/alpha. [warn-only] # # Exit code: non-zero ONLY when a discovered Python artifact crashes (non-zero exit) under an # environment that HAS python3 + scipy + statsmodels. Missing interpreter/libs -> skip, never fail. # Missing fields in a doc -> advisory warn. An empty or clean target exits 0 (no false failure). # Stock macOS bash 3.2 compatible (no mapfile, no associative arrays). YELLOW=$'\033[33m'; GREEN=$'\033[32m'; RED=$'\033[31m'; NC=$'\033[0m' EXIT=0 skip() { printf '%s[skip]%s %s\n' "$YELLOW" "$NC" "$*"; } note() { printf '%s[warn]%s %s\n' "$YELLOW" "$NC" "$*"; } ok() { printf '%s[ok]%s %s\n' "$GREEN" "$NC" "$*"; } err() { printf '%s[fail]%s %s\n' "$RED" "$NC" "$*"; EXIT=1; } ROOT="$(pwd)" # --- discover python sizing/analysis artifacts --- PY_FILES=() while IFS= read -r -d '' f; do PY_FILES+=("$f") done < <( find "$ROOT" \ \( -path '*/node_modules/*' -o -path '*/.git/*' -o -path '*/.venv/*' -o -path '*/venv/*' -o -path '*/site-packages/*' \) -prune -o \ -type f \( -iname '*sample*size*.py' -o -iname '*ab*test*.py' -o -iname '*experiment*.py' \ -o -iname '*power*.py' -o -iname '*cuped*.py' -o -iname '*srm*.py' \) -print0 2>/dev/null ) # --- discover experiment-design docs --- DOC_FILES=() while IFS= read -r -d '' f; do DOC_FILES+=("$f") done < <( find "$ROOT" \ \( -path '*/node_modules/*' -o -path '*/.git/*' -o -path '*/.venv/*' \) -prune -o \ -type f \( -iname '*experiment*.md' -o -iname '*ab*test*.md' -o -iname '*design*.md' \) -print0 2>/dev/null ) if [ "${#PY_FILES[@]}" -eq 0 ] && [ "${#DOC_FILES[@]}" -eq 0 ]; then skip "no experiment scripts or design docs found under $ROOT — design-only conversation" ok "verify.sh passed (empty target)" exit 0 fi # --- python artifacts --- if [ "${#PY_FILES[@]}" -gt 0 ]; then if ! command -v python3 >/dev/null 2>&1; then skip "python3 not installed — cannot execute sizing scripts" elif ! python3 -c 'import scipy, statsmodels' >/dev/null 2>&1; then skip "scipy/statsmodels not importable — install with: pip install scipy statsmodels" else for f in "${PY_FILES[@]}"; do out="$(python3 "$f" 2>&1)" || { err "$f: crashed under python3"; printf '%s\n' "$out" | sed 's/^/ /'; continue; } if printf '%s' "$out" | grep -Eq '[0-9]+'; then ok "$f: executed and printed a number" else note "$f: ran but printed no numeric output — a sizing/analysis script should print an n or p-value" fi done fi fi # --- design docs --- for f in "${DOC_FILES[@]}"; do miss="" grep -Eiq 'primary[[:space:]]+metric|metric[[:space:]]*:' "$f" || miss="$miss primary-metric" grep -Eiq 'mde|minimum[[:space:]]+detectable[[:space:]]+effect' "$f" || miss="$miss MDE" grep -Eiq 'power|alpha|significance[[:space:]]+level' "$f" || miss="$miss power/alpha" if [ -n "$miss" ]; then note "$f: experiment-design doc missing:$miss" else ok "$f: names a primary metric, an MDE, and power/alpha" fi done printf '\n' if [ "$EXIT" -eq 0 ]; then ok "verify.sh passed"; else err "verify.sh found failures"; fi exit "$EXIT"
-
-
SKILL.md 9.6 KB
--- name: ab-testing description: "Use when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go significant. NOT recurring metric tracking (that is `analytics`), NOT north-star/KPI trees (that is `kpi-framework`), NOT projecting metrics forward (that is `forecasting`)." tags: [ab-testing, experimentation, statistics, cuped, sample-size, hypothesis-testing] recommends: [analytics, kpi-framework, forecasting, data-cleaning, python, reporting] origin: risco --- # A/B testing — design and read a defensible experiment An experiment without a pre-committed sample size and a single primary metric is not an experiment. It is a dashboard you stare at until it tells you what you wanted to hear. The discipline lives almost entirely *before* traffic ships: a falsifiable hypothesis, one primary metric, a sample size derived from the smallest effect worth detecting, and a stop rule you cannot renegotiate at 2pm on day four. ## Pre-test checklist — every line true before any traffic Each one is a place experiments die silently. - [ ] A **falsifiable hypothesis** — names the change, the direction, and the metric it moves. - [ ] Exactly **ONE primary metric**. More than one primary = multiple comparisons = inflated false positives. - [ ] **Guardrail metrics** — what you refuse to harm (latency, refunds, unsubscribes) even for a win. - [ ] The **randomization unit = the analysis unit** (usually the user). Mixing them is pseudoreplication. - [ ] An **MDE** — the smallest lift that would change a decision. Not "any difference." - [ ] A **computed sample size** and the **duration** it implies at your real daily eligible traffic. - [ ] A **fixed stop rule** — a date or an n you commit to before launch. No "we'll see how it looks." ## Step 1 — Hypothesis and metrics State a null you can reject. "The new checkout button changes purchase conversion" with H0: conversion equal across arms, H1: it differs. Vague aspirations ("improve the funnel") have no rejection region. Pick one primary metric and freeze it. Why: every extra primary metric is another coin flip at α, so three "primary" metrics turn a 5% false-positive rate into roughly 14%. Demote the rest to secondary. Randomize on the same unit you analyze on. If a user sees the variant on every visit, randomize by user, not by session — analyzing 50k sessions from 8k users treats correlated observations as independent and fabricates significance. ```text Bad: "We think the redesign will improve engagement and revenue and retention." (no null, 3 primaries, no number) Good: "H0: 30-day purchase conversion is equal between control and the new one-click button. H1: it differs. Primary: purchase conversion. Guardrails: refund rate, p95 checkout latency. Randomize by user_id. MDE: +1.5pp absolute on a 12% baseline." ``` ## Step 2 — Sample size from MDE, baseline, and power Defaults: power 0.80, α 0.05 (two-sided). The MDE is yours to choose — it is the smallest effect that would actually change what you do. Rule: required n scales with ~1/MDE². Why: halving the smallest effect you care to detect roughly **quadruples** the traffic and time. This is the single most expensive decision in the design, so set the MDE to a business threshold, never to "whatever is small." For a conversion rate (proportion): ```python from statsmodels.stats.power import NormalIndPower from statsmodels.stats.proportion import proportion_effectsize p1, p2 = 0.12, 0.135 # baseline, baseline + MDE (1.5pp) h = proportion_effectsize(p1, p2) # Cohen's h (arcsine transform) n = NormalIndPower().solve_power(effect_size=h, alpha=0.05, power=0.80, ratio=1.0) print(int(-(-n // 1))) # n PER ARM, rounded up ``` For a continuous metric (revenue per user, time on page) use Welch-style sizing: ```python from statsmodels.stats.power import TTestIndPower effect = mde_in_units / pooled_std # Cohen's d n = TTestIndPower().solve_power(effect_size=effect, alpha=0.05, power=0.80, ratio=1.0) ``` Then convert n to a calendar plan: `days = ceil((n_per_arm * num_arms) / daily_eligible_users)`. If that is 9 days, run a clean **two full weeks** anyway — weekday/weekend mix is part of the population, and a 6-day test oversamples whoever shows up Tuesday. Full worked example (12% baseline, +1.5pp MDE, 80% power) plus runnable sizing, n→duration, CUPED θ and SRM snippets: `references/sample-size-and-cuped.md`. ## Step 3 — Run discipline **Fixed horizon is the default.** Commit to the n/date from Step 2 and read the result once, at the end. **Do not peek and stop at first significance.** Why: checking repeatedly and stopping the moment p < 0.05 inflates the Type-I error far above 5% — with enough looks, a null test crosses 0.05 most of the time. If you genuinely need to stop early, use a *sequential / always-valid* method (confidence sequences, e.g. Netflix's anytime-valid CIs) that holds Type-I error under continuous monitoring. Sequential is strong for **killing losers early** and weak for **calling winners early** — for a confident win, the fixed-horizon read is tighter. **Gate on SRM before you trust anything.** Compute a chi-square test on the observed split versus the intended ratio. If p < 0.001 the assignment or logging is broken — a bot filter dropping one arm, a redirect, a caching bug. Fix the instrumentation and rerun; do not "adjust for it." The peeking Type-I math, sequential/always-valid options, SRM diagnosis, novelty/primacy effects, Simpson's paradox in segments and HARKing all live in `references/pitfalls.md`. ## Step 4 — Analyze Pick the test by metric type: | Metric type | Test | |---|---| | Binary conversion (proportion) | Two-proportion z-test (`statsmodels.stats.proportion.proportions_ztest`) | | Continuous, roughly normal / large n | Welch's t-test (`scipy.stats.ttest_ind(..., equal_var=False)`) | | Continuous, heavy-tailed / skewed (revenue) | Mann-Whitney U, or t-test on a log/winsorized metric | Report **lift + confidence interval + p-value together**. Never p alone. Why: p < 0.05 with a CI of [+0.1pp, +5pp] is "statistically there, practically a coin toss" — the CI tells you the size, p only tells you it is not exactly zero. **Practical significance** = compare the CI to your MDE: if the whole interval sits above the MDE, ship; if it straddles the MDE, you detected *something* too small to matter. **Multiple comparisons.** Two regimes: - Small set of pre-declared **decision** metrics → **Bonferroni** (divide α by the count). Conservative, simple. - Large **exploratory** scan of many metrics/segments → **Benjamini-Hochberg (FDR)**. It keeps far more power than Bonferroni on big scans (in a 20-effect example, ~17 detected vs ~12 under Bonferroni). ## Step 5 — CUPED variance reduction CUPED (Controlled-experiment Using Pre-Experiment Data) subtracts predictable pre-period noise so the same traffic buys more power — or the same power needs less traffic. The adjusted metric: ```text Y_cuped = Y − θ · (X − E[X]) where θ = Cov(Y, X) / Var(X) ``` Estimate θ by regressing the in-experiment metric `Y` on the **pre-experiment** covariate `X` (e.g. each user's spend in the 4 weeks before the test), then analyze `Y_cuped` with the same test as Step 4. When it pays: recurring users with a strong pre-period signal. Reported wins — Netflix ~40% variance reduction on engagement, Statsig 50%+ on common metrics → significance in roughly half the time/traffic. When it does **nothing** — do not bother: brand-new users (no pre-period data), a covariate uncorrelated with the outcome, or — the cardinal sin — a covariate measured *after* assignment, which biases the estimate. The covariate MUST be pre-treatment and independent of which arm a user lands in. Runnable θ-via-OLS snippet in `references/sample-size-and-cuped.md`. ## Anti-patterns | Bad | Why it is wrong | Do instead | |---|---|---| | Peek daily, stop the day p < 0.05 | Repeated looks inflate Type-I error far above α | Fix n/date up front; or a sequential method that holds α | | No sample size set before launch | You will stop on noise and call it a win | Compute n from MDE/baseline/power in Step 2 | | Several "primary" metrics | Each is a coin flip at α; 3 metrics ≈ 14% false-positive | One frozen primary; the rest are secondary | | Ignore the observed split | An SRM means assignment/logging is broken; results are garbage | Chi-square SRM gate before reading anything | | Report only the p-value | Hides effect size — p < 0.05 can be practically zero | Always lift + CI + p; compare CI to MDE | | CUPED on a post-assignment covariate | Covariate correlated with the arm biases θ | Use only pre-treatment, assignment-independent covariates | | Call a winner from an underpowered test | "Not significant" then ≠ "no effect"; you lacked power | Reach planned n, or report the CI and say "inconclusive, here is the range" | | Decide the hypothesis after seeing results (HARKing) | Turns the whole analysis into a fishing expedition | Pre-register hypothesis + primary metric before launch | | Run 6 days because it "looks significant" | Oversamples one weekday slice of the population | Run full weeks; honor the fixed horizon | ## Checkable artifact When this skill emits a Python sizing/analysis script or an experiment-design doc, run `scripts/verify.sh` from your project root. It confirms the script executes under `python3` and prints a numeric sample size, and that any design doc names a primary metric, an MDE, and power/alpha. It is read-only and soft-passes when no artifact is present (a design-only conversation).
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.