product-analytics-experimentation-review
Review frontend analytics instrumentation and A/B or multivariate experiment configurations for event-schema correctness, sample-ratio-mismatch risk, statistically valid stopping rules, and consent-gated privacy compliance before shipping a tracking or experiment change.
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/frontend/product-analytics-experimentation-review
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Product Analytics & Experimentation Review
Purpose
Analytics and experimentation code looks low-risk (it doesn't change what users see) but is exactly where silent, expensive failures accumulate: schema drift that zeroes a KPI dashboard for weeks, sample-ratio mismatch that invalidates a whole test, and unconsented tracking that creates real compliance exposure. This skill exists to apply a measurement-integrity and privacy review before ship, not after a stakeholder notices the dashboard looks wrong.
When to use
Use this skill when the user asks to:
- review new or changed analytics event instrumentation for schema correctness,
- validate an A/B or multivariate experiment's bucketing logic and statistical plan before launch,
- audit whether tracking calls are properly consent-gated for privacy compliance,
- diagnose a suspected sample-ratio mismatch or an experiment result that looks statistically implausible.
When NOT to use
- Reviewing Core Web Vitals or general RUM/tracing instrumentation with no experiment or event-schema angle — hand off to
frontend-observability-rum-instrumentation. - Reviewing generic security/XSS/CSP posture of a page — hand off to
frontend-dom-xss-csp-review; this skill only reviews the analytics/experimentation-specific privacy surface (consent gating and PII-in-events), not the broader page security model. - Choosing a state-management or component-architecture pattern with no analytics or experiment angle — out of scope.
Context7 Documentation Protocol
Analytics vendor SDKs (GA4/gtag.js, feature-flag/experimentation platforms) change event-parameter names, consent-mode signal shapes, and SDK method signatures across versions, and memorized snippets go stale fast. Before making any platform-specific claim:
- Call
ToolSearchwith query"context7"(or"select:mcp__Context7__resolve-library-id,mcp__Context7__query-docs") to load the Context7 tools if not already loaded this session. - Call
mcp__Context7__resolve-library-idfor the specific analytics/experimentation library actually imported in the code under review (e.g. the GA4/gtag.jsSDK, a specific feature-flag/experiment SDK) before describing its event or bucketing API. Do not assume GA4 by default — verify the platform from the actualimport/script-tag evidence first. - Call
mcp__Context7__query-docsfor the specific mechanism in scope — e.g. "GA4 recommended event parameters for purchase event", "Google Consent Mode v2 signal defaults", "GA4 measurement protocol event schema limits" — before ruling on it. Verified library ID for web.dev platform guidance as of this skill'supdateddate:/websites/web_dev_articles. - Known facts verified via Context7/web.dev as of this skill's
updateddate: web.dev documents sendingweb-vitalsmetrics to GA4 viagtag('event', 'web_vitals', {...})withname/value/delta/id/labelfields, and shows GA4 BigQuery-export event tables keyed byevent_name/event_params/event_timestamp/user_pseudo_id— treat any schema claim about GA4's exported event shape as needing this structure, not an invented one. web.dev's Permissions API guidance (permissions-best-practices) documentsnavigator.permissions.query({name: ...})returning astateofgranted/denied/prompt, and states permission grants are scoped per-origin (a grant on one origin does not transfer to a subdomain/different origin) — apply the same non-transferability logic when reasoning about consent scope across subdomains. - If Context7 is unavailable or returns no relevant match, fall back to
official_docs/references/*.mdand mark the claimdocumentation-based (Context7 unavailable)rather than presenting it as freshly verified. - Never invent an analytics event-parameter name, consent-mode signal name, or experimentation-platform API that no queried source confirms.
Lean operating rules
- Identify the actual analytics/experimentation platform in use from the imported SDK or script tag before citing platform-specific behavior — do not assume GA4 or any specific vendor by default.
- Verify that bucketing/assignment logic is deterministic per user (stable hash/seed keyed to a persistent identifier), not re-randomized on refresh, session change, or page reload — this is the single most common cause of invalid experiment results and a leading cause of sample-ratio mismatch.
- Verify consent gating is enforced at the call site of the tracking function itself, not merely present somewhere else on the page — a consent banner existing does not mean a specific event call respects it; trace the actual conditional guarding the SDK call.
- Flag any event schema field that could carry PII (free-text fields, email, precise geo/lat-long, payment data, raw URLs with query strings, user-typed search terms) for hashing/redaction before approving.
- Require a pre-registered primary metric and minimum detectable effect (MDE) for any experiment reviewed; treat their absence as a blocking finding, not a nice-to-have — an experiment analyzed after the fact against whichever metric moved is not a valid test.
- Treat any observed sample split materially off the configured ratio (e.g. configured 50/50 showing as 46/54 or further at meaningful volume) as a sample-ratio-mismatch candidate requiring a chi-squared check, not a rounding artifact to wave off.
- Load
references/srm-and-bucketing-integrity.mdonly when auditing assignment/bucketing logic or diagnosing a suspected sample-ratio mismatch. - Load
references/consent-and-pii-in-events.mdonly when reviewing privacy/consent compliance of tracking calls or event payload PII exposure. - Load
references/stopping-rules-and-peeking.mdonly when evaluating whether an experiment's statistical significance claim or stop/continue decision is valid. - This skill performs static review only; it does not execute experiment code, query a live analytics backend, or flip a feature flag / experiment configuration in production.
Privacy & Consent Depth for Analytics
The generic consent/PII posture above (banner presence is not compliance, check the call site) covers the baseline. Some tracking changes need standard-specific depth: IAB TCF v2.2 purpose-granular consent, Google Consent Mode v2's default-denied timing requirement, and adjacent surfaces (Global Privacy Control/Do-Not-Track, cookie categorization, analytics-endpoint data residency) that a generic consent check can miss.
- Consent Mode v2 defaults must be synchronous and denied-by-default (doc-based, Google Consent Mode v2 docs):
gtag('consent', 'default', {analytics_storage: 'denied', ad_storage: 'denied', ad_user_data: 'denied', ad_personalization: 'denied'})must be set at the top of the page, before the gtag.js/GTM snippet loads and before anygtag('event', ...)call — not inside a CMP callback or an async-loaded script. A latergtag('consent', 'update', ...)call once the user answers the CMP does not retroactively fix a missing or async default; tags may already have fired under an undefined/permissive state. - IAB TCF v2.2 consent is granular per purpose and per vendor, not a single flag (standard-based inference from the TCF spec): a compliant check validates a specific
(vendorId, purposeId)grant from the decoded TC string — "a TC string cookie exists" is presence, not scope. A vendor consented for one purpose (e.g. measurement) is not automatically consented for another (e.g. personalized ads). - Any tracking call — pixel,
gtag,sendBeacon,fetch— that fires before a consent signal exists is unrecoverable exposure: unlike a suppressed JS event, an HTTP request (and any PII in its query string or body) cannot be un-sent once it leaves the client. - PII in event properties is a violation independent of consent state: consent governs whether tracking may happen, not what may be sent once it does — a consented-but-PII-laden event (raw email, full name, unhashed user ID) is still a data-minimization failure.
- Cookies set without a declared category or an explicit expiry cannot be honored by a CMP — flag any analytics/marketing cookie missing
Max-Age/Expiresor a category mapping. - Global Privacy Control (
navigator.globalPrivacyControl) and Do-Not-Track (navigator.doNotTrack) must be checked as an opt-out signal alongside explicit CMP consent, not replaced by it. - Analytics-endpoint data residency must be identified, not assumed — note the destination host/region for each analytics call and flag payloads sent to a default/global endpoint when a residency-scoped endpoint is expected.
- Load
references/privacy-consent-depth-for-analytics.mdwhen a review needs this standard-specific depth (Consent Mode v2 timing, TCF purpose/vendor granularity, GPC/DNT, cookie categorization, data residency) rather than the general consent/PII check alone.
References
Load these only when needed:
- SRM and bucketing integrity — use to verify deterministic, unbiased user assignment and to diagnose a suspected sample-ratio mismatch.
- Consent and PII in events — use to verify tracking calls are consent-gated and event payloads do not leak PII.
- Stopping rules and peeking — use to evaluate whether an experiment's significance claim is valid given its actual monitoring/stopping behavior.
- Privacy and consent depth for analytics — use for IAB TCF v2.2 purpose/vendor granularity, Google Consent Mode v2 default-timing requirements, GPC/Do-Not-Track honoring, cookie categorization/expiry, and analytics-endpoint data residency.
Response minimum
Return, at minimum:
- the analytics/experimentation platform identified and the docs used to verify its behavior,
- schema-correctness verdict against the documented data contract,
- SRM/bucketing-integrity verdict,
- consent-gate and PII findings,
- statistical-validity verdict (pre-registered metric/MDE present, stopping rule sound) with evidence level.
Files (vanguard-frontier-agentic)
-
references
-
consent-and-pii-in-events.md 6 KB
# Consent and PII in Events Use this reference when reviewing whether tracking calls are properly consent-gated and whether event payloads leak personally identifiable information (PII). ## What people get wrong The naive story is: > "We have a cookie consent banner, so our tracking is compliant." Wrong. A consent banner existing somewhere on the page proves nothing about whether any *specific* tracking call actually respects the user's choice. Compliance is a property of the call site, not the page. ## Officially grounded shape Per web.dev's Permissions API guidance (`permissions-best-practices`), browser permission state is queried per capability via `navigator.permissions.query({name: ...})`, returns a `state` of `granted` / `denied` / `prompt`, and — critically — **permission grants are scoped per-origin**: a grant on one origin does not transfer to a different origin or subdomain. This same non-transferability logic applies to consent state in analytics/marketing tooling: a consent decision recorded for one property/origin should not be silently assumed to apply to another origin or a cross-subdomain deployment without an explicit, documented consent-propagation mechanism (e.g. a shared consent cookie deliberately scoped for that purpose). Google's Consent Mode (referenced by GA4 tooling) works by adjusting or blocking specific measurement signals based on granted/denied consent state, rather than by a single all-or-nothing toggle — treat "consent" as a set of distinct signals (e.g. analytics storage, ad storage, ad personalization) each with independent gating, not one binary flag, unless the reviewed implementation is verified via Context7/current docs to behave otherwise. ## Non-negotiable design rules 1. **Consent must be checked at the call site of the tracking function**, immediately before or as a guard condition on the SDK call itself — not only at page load in a separate initialization block that the actual event-firing code doesn't reference. If the tracking call does not read live consent state (or a consent-derived flag) at the point it fires, treat the call as unverified for compliance regardless of banner presence. 2. **Default state before consent is granted must be "no non-essential tracking"**, not "track now and revoke later" — later revocation does not undo data already collected and transmitted. 3. **Consent scope is per-origin/per-property by default.** Any claim that a consent decision on one domain covers a subdomain, related property, or downstream data-sharing partner must point to an explicit propagation mechanism (documented shared cookie, server-side consent state sync) — never assume it transfers implicitly. 4. **Every event-schema field must be classified**: essential/functional (may be exempt from consent gating under most frameworks), or analytics/marketing (must be consent-gated). Do not let a field's classification be assumed from its name alone — check what it actually contains. 5. **PII must never appear in event payloads in plaintext** — this includes free-text search/input fields, full email addresses, exact geolocation (lat/long or address-level), payment instrument details, and raw URLs containing query-string parameters that carry any of the above. Hash, truncate, bucket (e.g. city-level geo instead of lat/long), or drop these fields before the event is considered shippable. 6. **A schema field added for a new feature is not "safe by default."** Every new or changed field must be re-evaluated against rules 4 and 5, not grandfathered in because prior fields passed review. ## Minimal safe review flow 1. Enumerate every distinct tracking/event-firing call site touched by the change (not just the SDK initialization). 2. For each call site, trace backwards to the actual conditional (if any) gating that call on consent state — confirm it reads a live consent value, not a stale/default-true flag. 3. For each event's payload fields, classify each field as essential, analytics, or marketing, and flag any field whose content could plausibly contain PII regardless of its name. 4. For any field flagged as potential PII, confirm it is hashed, truncated, generalized (e.g. geo bucketed to region), or removed before the event ships — do not accept "we'll filter it downstream" as sufficient, since the payload has already left the client by then. 5. If the deployment spans multiple origins/subdomains, confirm the consent-propagation mechanism explicitly, rather than assuming a shared consent state. ## Adversarial checklist Before approving an instrumentation change as consent/PII compliant, answer these: - Does the exact line of code that fires this tracking call check a live consent value, or does it fire unconditionally and rely on something else (a tag-manager rule, a server-side filter) to suppress it later? - What is the default consent state before the user makes any choice — is non-essential tracking off by default? - If this property has multiple subdomains or related origins, is there an explicit, documented mechanism sharing consent state, or is that being assumed? - For every new field in this event's payload, what raw value does it actually carry at runtime — has anyone printed a sample payload, or is the field's safety being assumed from its name? - If a field carries free text (search box, comment, form input), is there any scrubbing before it reaches the event payload, or does it pass through verbatim? If any of these cannot be answered, the compliance posture is unverified, not confirmed. ## When to push back Push back if the user asks to: - fire analytics/marketing events before consent is granted "just to see the data, we'll filter it later," - treat a consent banner's presence on the page as sufficient evidence that a specific new tracking call is compliant, - add a free-text or geolocation field to an event schema without a scrubbing/generalization step, - assume a consent decision made on the main domain automatically applies to a newly added subdomain or partner property. Those are not shortcuts. They are compliance and privacy exposure shipped as a schema change. -
privacy-consent-depth-for-analytics.md 10.7 KB
# Privacy & Consent Depth for Analytics Use this reference when the review needs standard-specific depth beyond the general consent/PII posture covered in `consent-and-pii-in-events.md` — specifically IAB TCF v2.2 purpose-granular consent, Google Consent Mode v2's default-denied timing requirement, Global Privacy Control / Do-Not-Track honoring, cookie categorization/expiry, and analytics-endpoint data residency. This is static-review guidance only; there is no detection corpus behind it (this skill is not a security-category skill). ## What people get wrong The naive story is: > "We call `gtag('consent', 'update', ...)` when the CMP loads, so we're Consent Mode compliant." Wrong. Consent Mode v2 distinguishes **defaults** from **updates**. The *default* consent state — set before the CMP has rendered or the user has answered anything — must itself be `denied` for the relevant signals, and it must be set **synchronously**, before any `gtag`/GTM tag fragment can fire. An `update` call arriving later is the correct mechanism for what happens *after* the user answers; it does nothing to undo tags that already fired under a missing or async-set default. A second naive story: > "We have a TCF consent string cookie (`euconsent-v2`), so every downstream vendor call is covered." Wrong. TCF v2.2 consent is granular **per purpose ID**, not a single yes/no. A vendor may be consented for "measurement" (Purpose 1/analytics-adjacent purposes) but not for "personalized ads" (Purpose 4) or cross-context targeting. A tracking call that fires because *a* TC string cookie exists, without checking whether *that specific vendor and purpose combination* is granted, is not actually gated — it is gated by presence, not by scope. ## Officially grounded shape **Google Consent Mode v2 (doc-based, via Context7 `/websites/developers_google_tag-platform_security_guides`):** - The default consent call must set at minimum all four parameters — `ad_storage`, `analytics_storage`, `ad_user_data`, `ad_personalization` — and per the consent-debugging guide, **"Ensure that the default consent statuses are not set asynchronously"** and the default block "should be placed at the top of your page, before any tag fragments or other code that might use consent settings." - The documented pattern is: `gtag('consent', 'default', {ad_storage: 'denied', ad_user_data: 'denied', ad_personalization: 'denied', analytics_storage: 'denied'})` executed synchronously, before the gtag.js/GTM snippet loads — then `gtag('consent', 'update', {...})` later once the user answers the CMP. - An optional `wait_for_update` parameter (e.g. `500` ms) exists specifically to give an asynchronous CMP banner time to call `update` before tags evaluate the default state — its presence in an implementation is a signal the team is aware of the timing requirement, not a substitute for a synchronous default block. - Treat any implementation that sets defaults inside an async-loaded script, inside a CMP callback, or after the main gtag/GTM snippet as **non-compliant with the documented requirement**, regardless of whether an `update` call exists later. **IAB TCF v2.2 (standard-based — no TCF library resolved in Context7 at time of review; treat as standard-based inference, not Context7-verified):** - Consent and legitimate-interest signals are structured **per purpose ID and per vendor**, encoded into a single TC string. A compliant integration checks a specific `(vendorId, purposeId)` pair — commonly via a decoded consent object exposing something equivalent to `vendorConsents.has(purposeId)` — not merely "does a TC string exist." - The TC string is the standard mechanism for **communicating** consent state to downstream vendors/partners. A server-side or third-party tracking call that omits the TC string (or an equivalent consent signal) gives the receiving party no way to honor purpose-specific restriction, even if the client-side call itself was gated correctly. - Purpose-level gating and vendor-level gating are independent: a vendor consented for one purpose is not automatically consented for another. Do not treat "CMP shows a green banner" as equivalent to "this specific vendor/purpose pair is granted." ## Non-negotiable design rules 1. **Consent Mode defaults must be synchronous and denied-by-default**, set in the page `<head>` before the gtag.js/GTM snippet and before any `gtag('event', ...)` call — not inside a CMP callback, not behind a `DOMContentLoaded`/async script boundary. 2. **No tracking call — pixel, `gtag`, `sendBeacon`, `fetch`, image-tag — fires before a consent signal exists**, whether that signal is a Consent Mode default/update pair or a decoded TCF TC string. "We'll gate it with CSS/display:none" or "we'll filter server-side" does not prevent the HTTP request from having already left the client with its payload. 3. **PII must never ride in event properties**, regardless of consent state: raw email address, full legal name, unhashed user ID, phone number, precise address. Consent governs *whether tracking may happen*, not *what may be sent once it does* — a consented-but-PII-laden event is still a data-minimization violation. 4. **Every cookie set by analytics/marketing code must be categorized (essential / functional / analytics / marketing) and carry an explicit, bounded expiry.** An uncategorized cookie cannot be selectively cleared by a CMP, and an unbounded/session-spanning expiry on a cookie that should be short-lived is itself a flag. 5. **Global Privacy Control (`navigator.globalPrivacyControl === true`) and Do-Not-Track (`navigator.doNotTrack === "1"`) must be checked and honored as an opt-out signal** for non-essential tracking, alongside — not instead of — explicit CMP consent state. Treat code that reads a TC string or Consent Mode flag but never checks `navigator.globalPrivacyControl`/`doNotTrack` as an incomplete opt-out surface. 6. **Data residency of the analytics endpoint must be identified, not assumed.** Note whether the destination (vendor domain, region-specific ingestion endpoint, self-hosted collector) is documented against any residency commitment made to users (e.g. "EU data stays in the EU"); flag payloads sent to a default/global endpoint when a residency-scoped endpoint is available and expected. ## Minimal safe review flow 1. Find the Consent Mode default-consent call (if Google tooling is in use). Confirm it (a) sets all relevant parameters to `denied`, (b) executes synchronously in `<head>`, before the gtag.js/GTM snippet — not inside a CMP callback or async script. 2. Find every tracking call site (pixel `<img>`/`new Image()`, `gtag('event', ...)`, `sendBeacon`, `fetch` to an analytics/marketing endpoint). For each, confirm it is temporally *after* a consent signal is available — not merely after the page has "loaded." 3. If TCF is in use, confirm the guarding logic checks a specific vendor/purpose pair from the decoded TC string, not just TC-string presence. 4. For each event's payload, scan for PII fields regardless of consent state (see `consent-and-pii-in-events.md` rule 5 for the field list). 5. List every cookie set by the code under review; confirm each has a declared category and an explicit `Max-Age`/`Expires` — flag any cookie set without either. 6. Confirm the code checks `navigator.globalPrivacyControl` and/or `navigator.doNotTrack` and suppresses non-essential tracking when either signals opt-out, independent of CMP state. 7. Identify the destination host/endpoint of each analytics call and note whether it matches any stated data-residency commitment; flag mismatches as unverified, not assumed-fine. ## Concrete sinks to grep for - **Unguarded pixel/image tracking**: `<img src="https://` or `new Image().src =` pointed at a tracking/analytics domain, with no preceding consent/TC-string/Consent-Mode check in the same code path. Risk: the HTTP GET — and any query-string PII on it — is unrecoverable once sent, unlike a suppressed JS event. - **`gtag('event', ...)` before `gtag('consent', 'default', ...)`**: search for `gtag('event'` occurrences and confirm a `gtag('consent', 'default'` call precedes them in load order (ideally same `<head>` block, before the gtag.js `src` script tag). Risk: Consent Mode has no default state yet, so Google's tags apply undefined/permissive behavior. - **PII literals in event payload construction**: `email`, `user.email`, `fullName`, `userId` (raw, unhashed) passed directly into an analytics `track()`/`gtag('event', ...)`/`sendBeacon()` call's properties object. Risk: PII leaves the client in plaintext regardless of consent state. - **Cookie set without `Max-Age`/`Expires` and without a category comment/constant**: `document.cookie = "<name>=<value>"` with no `Max-Age=`/`Expires=` attribute, or a cookie name with no mapping in a consent-category config. Risk: cannot be selectively honored, cleared, or excluded by a CMP. ## Adversarial checklist - Is the Consent Mode (or equivalent) default set synchronously in `<head>`, or does it live behind an async script/CMP callback where tags could fire first? - Does the code check a *specific* vendor/purpose grant from the TC string, or only "a TC string cookie is present"? - Does any tracking call — pixel, beacon, fetch — execute before any consent signal (default, TCF string, or otherwise) is available at all? - Does any event payload carry a raw email, name, phone number, or unhashed user ID, independent of whether consent was granted? - Does every cookie set by this code have both a declared category and an explicit expiry, or are any "just set and never revisited"? - Does the code check `navigator.globalPrivacyControl` / `navigator.doNotTrack` and suppress non-essential tracking on either signal, or does it only look at CMP-recorded consent? - Is the destination endpoint for this analytics call known and checked against any data-residency commitment, or is "it's probably fine" being assumed? If any of these cannot be answered from the code itself, treat the privacy/consent posture as **unverified**, not compliant. ## When to push back Push back if the user asks to: - ship a Consent Mode default that is set asynchronously or after the gtag.js/GTM snippet, "since we'll call update quickly anyway," - gate a tracking call on "a TC string cookie exists" rather than the specific vendor/purpose grant it needs, - add a raw email/name/unhashed user ID to an event's properties because "it's useful for support lookups," - set a new analytics/marketing cookie without an expiry or category because "we'll clean it up later," - skip checking Global Privacy Control/Do-Not-Track because "the CMP already covers consent," - send analytics payloads to a default global endpoint when a region-scoped endpoint exists and residency commitments apply. Those are not shortcuts. They are unrecoverable data exposure and standards non-compliance shipped as an analytics change. -
srm-and-bucketing-integrity.md 6.6 KB
# SRM and Bucketing Integrity Use this reference when auditing user-assignment/bucketing logic for an A/B or multivariate experiment, or when diagnosing a suspected sample-ratio mismatch (SRM). ## What people get wrong The naive story is: > "We split traffic 50/50 in the config, so the experiment is randomized correctly." Wrong. A correct config value does not guarantee correct *realized* assignment. SRM — where the observed split between variants diverges from the configured split by more than chance — is one of the most common silent killers of experiment validity, and it is caused by implementation bugs, not by the experimentation platform's randomization algorithm. ## Common root causes of SRM - **Non-deterministic bucketing**: assignment computed from `Math.random()` or an unseeded RNG on each page load/render instead of a stable hash of a persistent identifier (user ID, device ID, or a consistently-stored anonymous ID). This causes the same user to bounce between variants across requests, corrupting both the split and the per-user experience. - **Redirect/loading-time asymmetry**: one variant triggers a client-side redirect or has a slower paint path, so users on slow connections/devices disproportionately bounce before the exposure event fires, undercounting that variant. - **Bot/crawler contamination**: bot traffic is not filtered before assignment and is unevenly distributed across variants (e.g. one variant's URL gets crawled more). - **Caching bypass**: a CDN or browser cache serves a cached response for one variant to users who should have been freshly assigned, without going through the bucketing code path at all. - **Post-assignment filtering that isn't symmetric**: an error-handling or feature-guard path silently excludes users from one variant's exposure logging but not the other's (e.g. a try/catch around only the treatment code path). - **Multiple exposure events per user**: a user is bucketed once but the exposure/"experiment viewed" event fires multiple times (e.g. once per re-render), inflating one variant's logged count without inflating actual unique users. ## Non-negotiable design rules 1. **Assignment must be deterministic per stable identifier.** Given the same user identifier and experiment ID, the bucketing function must return the same variant every time, computed via a consistent hash (not re-rolled per request). If the review cannot point to the exact hash/seed input, treat this as a blocking finding. 2. **The identifier used for bucketing must survive the user's session.** An identifier that resets on page reload, tab close, or unauthenticated-to-authenticated transition will fragment the same real user across variants — verify what the identifier actually is (cookie, local storage, logged-in user ID) and its persistence. 3. **The exposure/assignment log event must fire exactly once per unique assigned user**, at the point the user is committed to see the variant-specific experience — not on every render, not before the variant is actually applied, and not skipped on an error path. 4. **Bot and internal-traffic filtering must run before or as part of assignment**, and must be applied identically to all variants — an asymmetric filter is itself an SRM source. 5. **The chi-squared goodness-of-fit test is the standard tool for detecting SRM** on the ratio between logged variant counts and the configured allocation ratio; a low p-value (commonly p < 0.001 is used as a conservative SRM-alert threshold given how much SRM checks are run) on that test at meaningful sample volume is a signal to halt and debug before trusting any of the experiment's results, not something to explain away as "probably fine." ## Minimal safe review flow 1. Identify the exact code path that computes variant assignment. Confirm it is a pure function of `(stable_user_id, experiment_id)` via a documented hash — not `Math.random()`, not time-based, not re-evaluated per render. 2. Confirm the identifier's lifetime matches or exceeds the experiment's intended duration and survives the user journeys the experiment covers (e.g. anonymous → logged-in transition, cross-device if applicable). 3. Trace the exposure-logging call site: confirm it fires once per assigned user, only after the variant is actually rendered/applied (not merely "assigned" in memory before an early return/error). 4. Check for asymmetric guards: any `try/catch`, feature flag, cache rule, or redirect that could apply to one variant's code path but not the equivalent point in the other variant's path. 5. If actual counts are available, compute the chi-squared statistic against the configured ratio; treat a materially significant deviation as a blocking finding requiring root-cause before results are trusted. ## Adversarial checklist Before approving an experiment's bucketing as sound, answer these: - What exact value is hashed to produce the variant assignment, and where is it read from? - Does that value exist and stay stable before the user has logged in, accepted cookies, or completed onboarding? - Is there any code path (error handler, feature gate, redirect, cache rule) that can prevent the exposure event from firing for one variant but not the other? - If a user's assignment identifier changes mid-experiment (e.g. anon ID merges into logged-in ID), what happens to their variant — do they get reassigned, and if so, does that reassignment apply identically across variants? - Has anyone actually looked at the realized split of exposure-event counts, or is "50/50" only ever been checked in the config file? If any of these cannot be answered, the SRM risk is unverified, not absent. ## High-risk assumptions to kill - "The experimentation platform's randomization is trustworthy, so the split will be fine" — the platform's RNG is rarely the source of real-world SRM; the surrounding application code is. - "We looked at the config and it says 50/50" — config intent is not realized assignment; only logged exposure counts prove realized assignment. - "A small split deviation doesn't matter" — at low absolute counts a deviation may be noise, but the correct response is to run the chi-squared test at the actual observed volume, not to eyeball it. ## When to push back Push back if the user asks to: - launch an experiment where bucketing logic cannot be traced to a deterministic, persistent-identifier hash, - treat an experiment's result as valid without ever having checked the realized split against the configured ratio, - ship an experiment where the exposure-logging event fires before the variant is actually applied (this inflates "assigned" counts independent of any real exposure). Those are not shortcuts; they are how invalid experiment results ship as product decisions. -
stopping-rules-and-peeking.md 6.3 KB
# Stopping Rules and Peeking Use this reference when evaluating whether an experiment's statistical significance claim or stop/continue decision is actually valid, given how it was monitored. ## What people get wrong The naive story is: > "The dashboard shows p < 0.05, so the result is significant — let's ship it." Wrong. A p-value computed under a fixed-sample-size assumption is only valid if the experiment was actually analyzed that way — once, at a pre-determined sample size or duration. Checking the dashboard daily and stopping the moment it crosses a significance threshold ("peeking") inflates the true false-positive rate far above the nominal 5%, often dramatically, because each look is an additional opportunity for noise to cross the threshold by chance. ## Officially grounded shape There is no single universally-mandated statistical framework across experimentation platforms — some use fixed-horizon frequentist testing, some use sequential testing frameworks (e.g. always-valid p-values, mSPRT-based methods) or Bayesian frameworks specifically designed to tolerate continuous monitoring. The review's job is not to mandate one specific framework, but to verify that **the stopping behavior actually used matches the statistical framework the experiment claims to be using.** A fixed-horizon test monitored continuously and stopped early on a significant read is invalid regardless of which platform produced the number. ## Non-negotiable design rules 1. **A primary metric and a minimum detectable effect (MDE) must be pre-registered before the experiment starts.** If the "significant" result was found by scanning a dashboard of many metrics after the fact and picking whichever moved, that is not a valid experiment result — it is multiple-comparisons-inflated noise. Treat the absence of a pre-registered primary metric as a blocking finding. 2. **The required sample size/duration must be computed from the pre-registered MDE, baseline rate, and desired power *before* the experiment launches**, not decided retroactively based on how the data looks partway through. 3. **If the experiment uses a fixed-horizon test, it must not be stopped early based on an interim significant read** — continuous or frequent peeking with a fixed-horizon test invalidates the significance claim. Verify the actual monitoring cadence against the analysis method actually used. 4. **If continuous monitoring genuinely is required (e.g. to catch a severe regression fast), the experiment must use a sequential-testing-aware method** built for that purpose — confirm the platform/analysis actually implements one rather than assuming a standard p-value naturally tolerates repeated looks. 5. **A minimum runtime should be enforced even for a positive early read** to cover known cyclical effects (day-of-week traffic mix, novelty effects that fade, weekly business cycles) — a result significant after two days is not the same evidence as the same result sustained across a full business cycle. 6. **Secondary/guardrail metrics must be declared in advance too**, distinct from the primary metric, so that a regression on a guardrail metric (e.g. latency, error rate, unsubscribe rate) is caught by design rather than discovered only if someone happens to check. ## Minimal safe review flow 1. Confirm a primary metric and MDE were declared before the experiment launched — ask for the pre-registration artifact (ticket, doc, experiment-platform config) rather than accepting a post-hoc description. 2. Identify which statistical framework is actually in use (fixed-horizon frequentist, sequential/always-valid, Bayesian) from the platform's documented behavior — verify via Context7/current docs rather than assuming. 3. Check the actual monitoring/stop history: was the experiment checked once at the planned endpoint, or repeatedly with a stop-on-significance pattern? If the platform's dashboard was checked daily and the experiment stopped as soon as it crossed significance, and the underlying method is fixed-horizon, flag this as invalid regardless of the reported p-value. 4. Confirm a minimum runtime (commonly at least one to two full business cycles, e.g. two weeks, adjusted for the app's actual traffic pattern) was respected even if an early read looked positive. 5. Confirm guardrail/secondary metrics were declared and checked, not just the metric that happened to move favorably. ## Adversarial checklist Before accepting a "statistically significant" result as ship-worthy, answer these: - What was the pre-registered primary metric, and does the "significant" result match that metric — or is it a different metric that happened to move? - What analysis method is this platform's significance number actually based on (fixed-horizon, sequential, Bayesian), and was the monitoring behavior (single look vs. continuous dashboard checking) consistent with that method's assumptions? - Was there a pre-declared minimum sample size or runtime, and was the experiment stopped before, at, or after that point? - Were guardrail metrics (latency, errors, revenue-adjacent metrics not directly targeted) checked, or only the primary metric that moved favorably? - If the same experiment had been stopped a few days earlier or later, is there reason to believe the result would look materially different (day-of-week effects, novelty effects)? If any of these cannot be answered, the significance claim is unverified, not confirmed. ## High-risk assumptions to kill - "It says p < 0.05 so it's real" — without knowing the analysis method and monitoring history, a p-value alone proves nothing. - "We checked it every day and stopped as soon as it looked good" — this is the textbook peeking failure mode, not diligence. - "We'll just ship whichever metric moved" — post-hoc metric selection without pre-registration is multiple-comparisons noise, not a finding. ## When to push back Push back if the user asks to: - ship a result based on a metric that was not the pre-registered primary metric, - stop a fixed-horizon experiment early because an interim dashboard check crossed a significance threshold, - skip a minimum-runtime requirement because an early read "already looks clearly positive," - run an experiment with no declared guardrail metrics on latency, errors, or adjacent business metrics. Those are not efficiency gains. They are how false positives become shipped product decisions.
-
-
metadata.json 1.6 KB
{ "id": "product-analytics-experimentation-review", "name": "Product Analytics & Experimentation Review", "type": "skill", "provider": "frontend", "harnesses": [ "claude-code", "cursor", "codex", "gemini", "kiro", "other" ], "summary": "Reviews frontend analytics instrumentation and A/B/multivariate experiment setups for event-schema correctness, sample-ratio-mismatch risk, valid statistical stopping rules, and consent-gated privacy compliance before an experiment or tracking change ships.", "source_type": "adapted", "official_docs": [ "https://web.dev/articles/vitals-business-impact", "https://developers.google.com/analytics/devguides/collection/ga4", "https://gdpr.eu/cookies/", "https://www.w3.org/TR/permissions/", "https://developers.google.com/tag-platform/security/guides/consent", "https://iabeurope.eu/iab-europe-transparency-consent-framework/" ], "security_notes": "Require consent-gate verification before any non-essential tracking call is approved; require PII scrubbing/hashing review on every new event schema field before it is marked shippable. Also check IAB TCF v2.2 purpose/vendor-granular consent, Google Consent Mode v2 synchronous default-denied timing, GPC/Do-Not-Track honoring, cookie categorization/expiry, and analytics-endpoint data residency. Read-only static review; does not execute experiment code, query a live analytics backend, or change a live feature-flag/experiment configuration.", "last_verified": "2026-07-03", "path": "skills/frontend/product-analytics-experimentation-review", "author": "github: VincentChuWaiChow", "version": "0.1.0" } -
SKILL.md 10.4 KB
--- name: product-analytics-experimentation-review description: Review frontend analytics instrumentation and A/B or multivariate experiment configurations for event-schema correctness, sample-ratio-mismatch risk, statistically valid stopping rules, and consent-gated privacy compliance before shipping a tracking or experiment change. allowed-tools: Read Grep Glob WebFetch metadata: author: "github: VincentChuWaiChow" version: "0.1.0" updated: "2026-07-02" category: data --- # Product Analytics & Experimentation Review ## Purpose Analytics and experimentation code looks low-risk (it doesn't change what users see) but is exactly where silent, expensive failures accumulate: schema drift that zeroes a KPI dashboard for weeks, sample-ratio mismatch that invalidates a whole test, and unconsented tracking that creates real compliance exposure. This skill exists to apply a measurement-integrity and privacy review before ship, not after a stakeholder notices the dashboard looks wrong. ## When to use Use this skill when the user asks to: - review new or changed analytics event instrumentation for schema correctness, - validate an A/B or multivariate experiment's bucketing logic and statistical plan before launch, - audit whether tracking calls are properly consent-gated for privacy compliance, - diagnose a suspected sample-ratio mismatch or an experiment result that looks statistically implausible. ## When NOT to use - Reviewing Core Web Vitals or general RUM/tracing instrumentation with no experiment or event-schema angle — hand off to `frontend-observability-rum-instrumentation`. - Reviewing generic security/XSS/CSP posture of a page — hand off to `frontend-dom-xss-csp-review`; this skill only reviews the analytics/experimentation-specific privacy surface (consent gating and PII-in-events), not the broader page security model. - Choosing a state-management or component-architecture pattern with no analytics or experiment angle — out of scope. ## Context7 Documentation Protocol Analytics vendor SDKs (GA4/gtag.js, feature-flag/experimentation platforms) change event-parameter names, consent-mode signal shapes, and SDK method signatures across versions, and memorized snippets go stale fast. Before making any platform-specific claim: 1. Call `ToolSearch` with query `"context7"` (or `"select:mcp__Context7__resolve-library-id,mcp__Context7__query-docs"`) to load the Context7 tools if not already loaded this session. 2. Call `mcp__Context7__resolve-library-id` for the specific analytics/experimentation library actually imported in the code under review (e.g. the GA4/`gtag.js` SDK, a specific feature-flag/experiment SDK) before describing its event or bucketing API. Do not assume GA4 by default — verify the platform from the actual `import`/script-tag evidence first. 3. Call `mcp__Context7__query-docs` for the specific mechanism in scope — e.g. "GA4 recommended event parameters for purchase event", "Google Consent Mode v2 signal defaults", "GA4 measurement protocol event schema limits" — before ruling on it. Verified library ID for web.dev platform guidance as of this skill's `updated` date: `/websites/web_dev_articles`. 4. Known facts verified via Context7/web.dev as of this skill's `updated` date: web.dev documents sending `web-vitals` metrics to GA4 via `gtag('event', 'web_vitals', {...})` with `name`/`value`/`delta`/`id`/`label` fields, and shows GA4 BigQuery-export event tables keyed by `event_name`/`event_params`/`event_timestamp`/`user_pseudo_id` — treat any schema claim about GA4's exported event shape as needing this structure, not an invented one. web.dev's Permissions API guidance (`permissions-best-practices`) documents `navigator.permissions.query({name: ...})` returning a `state` of `granted`/`denied`/`prompt`, and states permission grants are scoped per-origin (a grant on one origin does not transfer to a subdomain/different origin) — apply the same non-transferability logic when reasoning about consent scope across subdomains. 5. If Context7 is unavailable or returns no relevant match, fall back to `official_docs` / `references/*.md` and mark the claim `documentation-based (Context7 unavailable)` rather than presenting it as freshly verified. 6. Never invent an analytics event-parameter name, consent-mode signal name, or experimentation-platform API that no queried source confirms. ## Lean operating rules - Identify the actual analytics/experimentation platform in use from the imported SDK or script tag before citing platform-specific behavior — do not assume GA4 or any specific vendor by default. - Verify that bucketing/assignment logic is deterministic per user (stable hash/seed keyed to a persistent identifier), not re-randomized on refresh, session change, or page reload — this is the single most common cause of invalid experiment results and a leading cause of sample-ratio mismatch. - Verify consent gating is enforced at the call site of the tracking function itself, not merely present somewhere else on the page — a consent banner existing does not mean a specific event call respects it; trace the actual conditional guarding the SDK call. - Flag any event schema field that could carry PII (free-text fields, email, precise geo/lat-long, payment data, raw URLs with query strings, user-typed search terms) for hashing/redaction before approving. - Require a pre-registered primary metric and minimum detectable effect (MDE) for any experiment reviewed; treat their absence as a blocking finding, not a nice-to-have — an experiment analyzed after the fact against whichever metric moved is not a valid test. - Treat any observed sample split materially off the configured ratio (e.g. configured 50/50 showing as 46/54 or further at meaningful volume) as a sample-ratio-mismatch candidate requiring a chi-squared check, not a rounding artifact to wave off. - Load `references/srm-and-bucketing-integrity.md` only when auditing assignment/bucketing logic or diagnosing a suspected sample-ratio mismatch. - Load `references/consent-and-pii-in-events.md` only when reviewing privacy/consent compliance of tracking calls or event payload PII exposure. - Load `references/stopping-rules-and-peeking.md` only when evaluating whether an experiment's statistical significance claim or stop/continue decision is valid. - This skill performs static review only; it does not execute experiment code, query a live analytics backend, or flip a feature flag / experiment configuration in production. ## Privacy & Consent Depth for Analytics The generic consent/PII posture above (banner presence is not compliance, check the call site) covers the baseline. Some tracking changes need standard-specific depth: IAB TCF v2.2 purpose-granular consent, Google Consent Mode v2's default-denied timing requirement, and adjacent surfaces (Global Privacy Control/Do-Not-Track, cookie categorization, analytics-endpoint data residency) that a generic consent check can miss. - **Consent Mode v2 defaults must be synchronous and denied-by-default** (doc-based, Google Consent Mode v2 docs): `gtag('consent', 'default', {analytics_storage: 'denied', ad_storage: 'denied', ad_user_data: 'denied', ad_personalization: 'denied'})` must be set at the top of the page, before the gtag.js/GTM snippet loads and before any `gtag('event', ...)` call — not inside a CMP callback or an async-loaded script. A later `gtag('consent', 'update', ...)` call once the user answers the CMP does not retroactively fix a missing or async default; tags may already have fired under an undefined/permissive state. - **IAB TCF v2.2 consent is granular per purpose and per vendor, not a single flag** (standard-based inference from the TCF spec): a compliant check validates a specific `(vendorId, purposeId)` grant from the decoded TC string — "a TC string cookie exists" is presence, not scope. A vendor consented for one purpose (e.g. measurement) is not automatically consented for another (e.g. personalized ads). - **Any tracking call — pixel, `gtag`, `sendBeacon`, `fetch` — that fires before a consent signal exists is unrecoverable exposure**: unlike a suppressed JS event, an HTTP request (and any PII in its query string or body) cannot be un-sent once it leaves the client. - **PII in event properties is a violation independent of consent state**: consent governs whether tracking may happen, not what may be sent once it does — a consented-but-PII-laden event (raw email, full name, unhashed user ID) is still a data-minimization failure. - **Cookies set without a declared category or an explicit expiry cannot be honored by a CMP** — flag any analytics/marketing cookie missing `Max-Age`/`Expires` or a category mapping. - **Global Privacy Control (`navigator.globalPrivacyControl`) and Do-Not-Track (`navigator.doNotTrack`) must be checked as an opt-out signal alongside explicit CMP consent**, not replaced by it. - **Analytics-endpoint data residency must be identified, not assumed** — note the destination host/region for each analytics call and flag payloads sent to a default/global endpoint when a residency-scoped endpoint is expected. - Load `references/privacy-consent-depth-for-analytics.md` when a review needs this standard-specific depth (Consent Mode v2 timing, TCF purpose/vendor granularity, GPC/DNT, cookie categorization, data residency) rather than the general consent/PII check alone. ## References Load these only when needed: - [SRM and bucketing integrity](references/srm-and-bucketing-integrity.md) — use to verify deterministic, unbiased user assignment and to diagnose a suspected sample-ratio mismatch. - [Consent and PII in events](references/consent-and-pii-in-events.md) — use to verify tracking calls are consent-gated and event payloads do not leak PII. - [Stopping rules and peeking](references/stopping-rules-and-peeking.md) — use to evaluate whether an experiment's significance claim is valid given its actual monitoring/stopping behavior. - [Privacy and consent depth for analytics](references/privacy-consent-depth-for-analytics.md) — use for IAB TCF v2.2 purpose/vendor granularity, Google Consent Mode v2 default-timing requirements, GPC/Do-Not-Track honoring, cookie categorization/expiry, and analytics-endpoint data residency. ## Response minimum Return, at minimum: - the analytics/experimentation platform identified and the docs used to verify its behavior, - schema-correctness verdict against the documented data contract, - SRM/bucketing-integrity verdict, - consent-gate and PII findings, - statistical-validity verdict (pre-registered metric/MDE present, stopping rule sound) with evidence level.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.