e2e-testing-playwright-review
Reviews Playwright end-to-end test configuration -- fixtures, storageState/auth setup, CI sharding and parallelism, and toHaveScreenshot visual-assertion options -- for reliability and correct gating, grounded in current, version-specific Playwright API docs.
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/frontend/e2e-testing-playwright-review
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
E2E Testing (Playwright) Review
Purpose
Playwright E2E suites fail in two directions: they're flaky enough that teams disable them, or they're so under-configured (no sharding, no masking, no stable waits) that they're slow and noisy without adding confidence. This skill reviews Playwright-specific configuration -- fixtures, auth state, parallelism, and screenshot assertions -- against current official API behavior rather than remembered API shapes that may be stale across majors.
When to use
Use this skill when the user asks to:
- review or configure Playwright test fixtures,
storageState, or auth setup, - diagnose Playwright test flakiness (timing, animation, non-deterministic content),
- configure CI sharding/parallelism for a Playwright suite,
- review or tune
toHaveScreenshotvisual-assertion options (maxDiffPixelRatio,mask,animations).
Context7 Documentation Protocol
Playwright's config shape, CLI flags, and assertion option names change across majors and are documented, not folklore -- never assert a flag, option, or "best practice" from memory.
- Call
ToolSearchwith query"context7"(or"select:mcp__Context7__resolve-library-id,mcp__Context7__query-docs") to load the Context7 tools if not already loaded in this session. - Call
mcp__Context7__resolve-library-idwith library namePlaywrightto obtain the current Context7-compatible ID (/microsoft/playwright); prefer the resolved ID over guessing. - Call
mcp__Context7__query-docsfor the specific claim in question -- e.g. "toHaveScreenshot maxDiffPixelRatio and mask options", "shard CLI flag and blob reporter merge", "storageState project dependencies setup" -- before stating it as fact. Do this per review, not once from a prior session's memory. - Prefer the official docs URLs in
official_docsfor primary normative statements (exact CLI flags, exact config shape); use Context7 to ground and cross-check the claim before writing it into a finding. - If Context7 is unavailable or returns no relevant match, fall back to the
official_docsURLs and mark the claimdocumentation-based (Context7 unavailable)rather than presenting it as freshly verified. - Never invent a config key, CLI flag, assertion option, or fixture API that no queried source confirms.
Lean operating rules
- Always confirm the installed Playwright version before asserting on API option names;
toHaveScreenshotoptions and CLI--shardsyntax are stable but still verify against the project'spackage.jsonversion, not assumption. - Distinguish flakiness caused by real non-determinism (animation, dynamic content, network timing) from flakiness caused by weak locators or missing waits; the fix differs, and misdiagnosis just hides a real timing bug behind a wider tolerance.
- Recommend
animations: 'disabled'and explicitmask/stylePathfor non-deterministic regions before recommending a loosermaxDiffPixelRatio/maxDiffPixels/threshold-- widening tolerance first papers over the actual source of visual noise. - Treat
storageState.jsonfixtures as sensitive: they must come from a dedicated test account, never a real user session, and should not be committed if they contain live tokens (see security notes below). - Recommend CI sharding (
--shard=N/Mwith a matrix strategy plusblobreporter andmerge-reports) only after confirming the suite's actual wall-clock time in CI justifies the added job complexity and report-merge step; sharding without a merge step silently drops report coverage. - Prefer the setup-project/
dependenciespattern (a dedicatedsetupproject producingstorageState, consumed viadependencies: ['setup']) over ad hocglobalSetupfor auth when the project already uses Playwright's project model; both are documented, but they compose differently with sharding and per-project storage state. - Do not conflate
fullyParallel(parallelizes tests within a single file, in addition to across files) with CI-level sharding (--shard, distributes files across separate CI jobs/machines) -- they solve different bottlenecks and a suite can need one, both, or neither. - Load the design-token/visual-regression skill instead of this one when the question is about baseline-approval workflow or a third-party visual-review service (e.g. Chromatic), not Playwright's own
toHaveScreenshotconfig.
References
Load these only when needed:
- Fixtures, auth setup, and storageState security -- use when reviewing or designing
storageState/auth fixtures,setupproject dependencies,globalSetup, or handling ofstorageState.json/HAR files as sensitive artifacts. - CI sharding and parallelism -- use when configuring or reviewing
--shard, matrix CI strategy, blob-reporter merge,fullyParallel, orworkerstuning. - Visual assertion tuning (toHaveScreenshot) -- use when reviewing or tuning
toHaveScreenshot/toMatchSnapshotoptions (animations,mask,maxDiffPixelRatio,maxDiffPixels,threshold,stylePath) or diagnosing visual-diff flakiness.
Response minimum
Return, at minimum:
- the Playwright feature/config area in scope (fixtures, sharding, screenshot assertion),
- evidence level and the exact Playwright version the guidance targets,
- root cause of any flakiness identified (not just a threshold-widening patch),
- proposed config diff (not applied) with the option names verified against docs,
- security caveat on any
storageState/HAR fixture reviewed.
Files (vanguard-frontier-agentic)
-
references
-
ci-sharding-and-parallelism.md 4.1 KB
# CI Sharding and Parallelism Use this reference when configuring or reviewing Playwright's `--shard` CLI flag, a CI matrix strategy, blob-reporter merging, `fullyParallel`, or `workers` tuning. For auth-state interactions across sharded/parallel projects, see `fixtures-and-auth-setup.md`. > Version note: the `blob` reporter and `npx playwright merge-reports` command are the documented mechanism for combining sharded results into one report. Verify exact reporter name and CLI flags against the installed version via Context7 (`/microsoft/playwright`) or the official docs before citing them in a report. ## Officially grounded design points - **`fullyParallel` and sharding solve different bottlenecks.** `fullyParallel: true` in `playwright.config.ts` parallelizes tests *within a single file* in addition to across files (by default, Playwright already parallelizes across files using `workers`). CI-level sharding (`--shard=<index>/<total>`) instead splits the *set of test files* across separate CI jobs/machines. A suite bottlenecked by one huge file with many tests needs `fullyParallel`; a suite bottlenecked by total file count/CI wall-clock time needs sharding (or both). - **`test.describe.parallel()`** is the documented, file-scoped alternative to global `fullyParallel` -- it opts a specific `describe` block into within-file parallel execution without changing the setting for the whole suite. - **Sharding requires a report-merge step to remain useful.** The documented CI pattern is: a matrix job runs `npx playwright test --shard=${{ matrix.shardIndex }}/${{ matrix.shardTotal }}` with `reporter: 'blob'` configured, each shard uploads its `blob-report` directory as a CI artifact, and a separate downstream job downloads all shard artifacts and runs `npx playwright merge-reports --reporter html ./all-blob-reports` to produce one combined report. Sharding without this merge step means each shard produces an isolated, easy-to-miss report instead of one gating artifact. - **`workers`** controls the number of parallel worker processes *within* a single Playwright invocation (i.e., within one shard/CI job); it is a separate lever from both `fullyParallel` and `--shard` and is commonly constrained by the CI runner's available CPU/memory, not just desired speed. ## Non-negotiable design rules 1. **Do not recommend sharding to fix flakiness.** Sharding distributes work across CI jobs; it does not change whether an individual test is flaky and can even change failure order/timing in ways that mask a root cause. If the underlying request is "the suite is unreliable," diagnose flakiness first (locators, waits, isolation) and treat sharding purely as a wall-clock-time lever. 2. **Never introduce `--shard` without also configuring `reporter: 'blob'` and a merge job.** A sharding change that lands without the corresponding merge-reports job in CI produces N disconnected reports and typically means nobody actually reviews the combined result -- treat this as an incomplete change, not a matter of taste. 3. **Justify sharding with a measured wall-clock baseline, not a guess.** Require the current CI suite duration before recommending a specific shard count; splitting a suite that already finishes in 3 minutes into 8 shards adds CI queueing/setup overhead without meaningful benefit and increases the merge-job's own failure surface. 4. **Tune `workers` to the CI runner's actual resources, not a fixed "more is faster" assumption.** Over-provisioning workers on a memory-constrained runner is a documented source of resource contention that manifests as flaky, not slow, tests -- which is easy to misdiagnose as an application bug. 5. **`fail-fast: false` in a CI matrix is the documented choice for shard jobs** so that one failing shard doesn't cancel the others and hide additional failures that would otherwise surface in the same run. ## Response discipline When reviewing a sharding/parallelism setup, state the current shard count and worker configuration as found (repo/CI-config evidence), the measured baseline duration if available (or explicitly flag it as missing), and whether a `merge-reports` step exists -- an incomplete shard-without-merge config is a hard finding, not a style note. -
fixtures-and-auth-setup.md 4.3 KB
# Fixtures, Auth Setup, and storageState Security Use this reference when reviewing or designing Playwright `storageState`/auth fixtures, a `setup` project with `dependencies`, `globalSetup`, or when a `storageState.json`/HAR file needs a security review before it's committed or shared. For sharding/parallelism interactions with per-project storage state, see `ci-sharding-and-parallelism.md`. > Version note: `storageState()` option shapes (e.g. the `indexedDB` capture flag) have been added across Playwright releases. Verify exact option availability against the installed version via Context7 (`/microsoft/playwright`) or the official docs before citing a capability as available. ## Officially grounded patterns Playwright's own docs describe two supported, non-exclusive ways to produce a reusable auth state: - **Setup project + `dependencies`.** A dedicated project (e.g. `{ name: 'setup', testMatch: /.*\.setup\.ts/ }`) performs authentication once and writes `storageState` to a file (e.g. `playwright/.auth/user.json`); browser-specific projects (`chromium`, `firefox`, ...) declare `dependencies: ['setup']` and set `use: { storageState: 'playwright/.auth/user.json' }` to consume it. This is the documented pattern for reusing one authenticated session across multiple projects/browsers without re-authenticating per project. - **`globalSetup`.** A `globalSetup` function (referenced from `playwright.config.ts`) launches a browser, performs the login flow, and calls `page.context().storageState({ path: storageState })` once before the whole run. Both the setup-project pattern and `globalSetup` are documented; the setup-project pattern is the one that composes with `dependencies` and per-project overrides, so prefer it when the project already uses Playwright's project model. - **API-based auth**, skipping the UI entirely: a `setup` test can call `request.post(...)` against a login endpoint and then `request.storageState({ path: authFile })` to persist cookies/tokens without driving a browser -- faster and less flaky than a UI-driven login, when the app's auth flow supports it. - `storageState({ path, indexedDB: true })` is the documented way to also persist IndexedDB-backed session data (relevant for apps that store tokens in IndexedDB rather than cookies/localStorage); omitting `indexedDB: true` for such an app silently drops part of the session and can cause the "auth setup ran but tests still see a logged-out state" failure mode. ## Non-negotiable design rules 1. **Never generate `storageState.json` against a real user's live session.** It must come from a dedicated, disposable test account with no access to production customer data. A file captured from a real session contains live cookies/tokens that are functionally equivalent to that user's credentials. 2. **Treat `storageState.json` as a secret artifact, not a build output.** Do not commit it to a public repository. If it must persist across CI runs, store it as a short-lived CI artifact/secret with restricted access, not a tracked file in version control. 3. **Scrub HAR-file fixtures before committing them.** HAR recordings used for network-mocking (route interception replay) capture full request/response headers, which commonly include `Authorization` headers, API keys, or session cookies. A HAR fixture that hasn't been reviewed for these is a credential-leak risk equivalent to committing a raw token. 4. **Re-authenticate per test-account tier, not per test.** If the suite exercises multiple roles/permission levels, use multiple named `storageState` files (one per role) produced by distinct `setup` tests, rather than one shared elevated-privilege session reused everywhere -- reusing an admin session for tests that only need a standard-user role overstates the privilege the test actually needs and can mask authorization bugs. 5. **Expire and rotate test-account credentials used to produce fixtures on a defined cadence.** A `storageState` fixture generated from a test account whose password/token never rotates is a long-lived credential sitting in CI infrastructure. ## Response discipline When reviewing fixtures, state explicitly whether the `storageState`/HAR file was confirmed to originate from a dedicated test account (repo evidence: setup-test source, CI secret config) or whether that could not be confirmed from available evidence -- do not assume test-only provenance without a citation. -
visual-assertion-tuning.md 4.6 KB
# Visual Assertion Tuning (toHaveScreenshot) Use this reference when reviewing or tuning `toHaveScreenshot`/`toMatchSnapshot` options (`animations`, `mask`, `maskColor`, `stylePath`, `maxDiffPixelRatio`, `maxDiffPixels`, `threshold`) or diagnosing visual-diff flakiness. For flakiness caused by locators/waits rather than visual noise, that is a fixtures/timing question, not a visual-assertion-tuning one. > Version note: `toHaveScreenshot` option names below are drawn from the current `PageAssertions`/`LocatorAssertions` API reference. Verify exact defaults (e.g. whether `animations` defaults to `'disabled'` in the installed version) against Context7 (`/microsoft/playwright`) or the official API reference before citing a default value as fact. ## Officially grounded option set Per the official `PageAssertions.toHaveScreenshot` / `LocatorAssertions.toHaveScreenshot` API reference, the documented tuning options include: - **`animations`** -- whether to allow or disable animations in the screenshot; the framework's internal screenshot implementation passes `animations: helper.options.animations ?? 'disabled'`, i.e. animations are suppressed by default unless explicitly overridden. - **`mask`** -- an array of `Locator`/`ElementHandle` values to mask out of the comparison (paired with `maskColor` to control the mask's fill color). - **`stylePath`** -- a path to a CSS file injected into the page before the screenshot is taken, for style-level suppression of non-deterministic elements (e.g. hiding a live clock or a carousel via CSS) that isn't practical to mask element-by-element. - **`clip`** / **`fullPage`** -- region-of-capture controls, independent of the diff-tolerance controls below. - **Diff-tolerance controls**: `maxDiffPixels` (absolute pixel count allowed to differ), `maxDiffPixelRatio` (ratio of differing pixels, 0-1), and `threshold` (a per-pixel color-difference threshold for the comparator). These can be set per-assertion call or globally in `playwright.config.ts` under `expect.toHaveScreenshot` / `expect.toMatchSnapshot`. - Both `toHaveScreenshot` (web-first, waits for two consecutive identical screenshots before comparing -- i.e. it has its own visual-stability wait built in) and `toMatchSnapshot` (compares an already-captured `page.screenshot()` buffer) support the diff-tolerance options; they differ in whether the wait-for-stability behavior is built in. ## Non-negotiable design rules 1. **Fix the non-determinism before widening tolerance.** If a diff is caused by a specific known-dynamic region (a timestamp, a live counter, an animated element, ad/embed content), the documented, targeted fix is `mask` (or `stylePath` for CSS-level suppression) on that region -- not a global increase to `maxDiffPixelRatio`/`maxDiffPixels`/`threshold`. Widening global tolerance first hides real regressions in the rest of the page along with the intended noise. 2. **Do not disable `animations` masking as a blanket policy without checking whether the default is already `'disabled'`.** Since the built-in implementation defaults `animations` to `'disabled'` unless overridden, an explicit `animations: 'allow'` (or equivalent override) in a config/test is itself worth flagging and asking why -- it re-introduces a documented source of visual flakiness. 3. **A widened `maxDiffPixelRatio`/`maxDiffPixels` needs a stated reason tied to a specific region or rendering variance (e.g. font anti-aliasing differences across OS/CI runners), not a round number picked to make CI green.** An unexplained tolerance value is a signal the underlying flake was never diagnosed. 4. **Global `expect.toHaveScreenshot`/`expect.toMatchSnapshot` config in `playwright.config.ts` sets the default for the whole suite; a per-call override should be justified by why that specific assertion needs different tolerance than the suite default**, not applied broadly as a workaround for one flaky test that then silently loosens every other screenshot assertion using the same call site pattern. 5. **`omitBackground` and `scale` affect what the baseline image actually contains** (transparency handling, DPI scaling); changing either after a baseline was captured invalidates the existing baseline and requires a deliberate baseline-update pass, not an incidental config tweak. ## Response discipline When reviewing visual-assertion config, cite the specific option and value found (with file reference), state whether a wider tolerance is targeted (paired with `mask`/`stylePath` for a named dynamic region) or global/unexplained, and label the API-shape claim as `documentation-based` (grounded via Context7/official API reference this session) versus `inference`.
-
-
metadata.json 1.3 KB
{ "id": "e2e-testing-playwright-review", "name": "E2E Testing (Playwright) Review", "type": "skill", "provider": "frontend", "harnesses": [ "claude-code", "cursor", "codex", "gemini", "kiro", "other" ], "summary": "Reviews Playwright end-to-end test configuration, fixtures, sharding, and screenshot-assertion setup for reliability, CI parallelism, and accurate visual/behavioral gating, grounded in current Playwright API docs.", "source_type": "original", "official_docs": [ "https://playwright.dev/docs/best-practices", "https://playwright.dev/docs/ci", "https://playwright.dev/docs/test-snapshots", "https://playwright.dev/docs/api/class-pageassertions#page-assertions-to-have-screenshot-1", "https://playwright.dev/docs/test-parallel" ], "security_notes": "Playwright storageState.json fixtures can capture real session cookies/auth tokens if generated against a live authenticated session; require these be generated against a dedicated test account and treated as sensitive (not committed to a public repo). HAR-file recordings used for network mocking can capture real API keys in request headers -- scrub before committing.", "last_verified": "2026-07-02", "path": "skills/frontend/e2e-testing-playwright-review", "author": "github: VincentChuWaiChow", "version": "0.1.0" } -
SKILL.md 5.9 KB
--- name: e2e-testing-playwright-review description: Reviews Playwright end-to-end test configuration -- fixtures, storageState/auth setup, CI sharding and parallelism, and toHaveScreenshot visual-assertion options -- for reliability and correct gating, grounded in current, version-specific Playwright API docs. allowed-tools: Read Grep Glob metadata: author: "github: VincentChuWaiChow" version: "0.1.0" updated: "2026-07-02" category: delivery --- # E2E Testing (Playwright) Review ## Purpose Playwright E2E suites fail in two directions: they're flaky enough that teams disable them, or they're so under-configured (no sharding, no masking, no stable waits) that they're slow and noisy without adding confidence. This skill reviews Playwright-specific configuration -- fixtures, auth state, parallelism, and screenshot assertions -- against current official API behavior rather than remembered API shapes that may be stale across majors. ## When to use Use this skill when the user asks to: - review or configure Playwright test fixtures, `storageState`, or auth setup, - diagnose Playwright test flakiness (timing, animation, non-deterministic content), - configure CI sharding/parallelism for a Playwright suite, - review or tune `toHaveScreenshot` visual-assertion options (`maxDiffPixelRatio`, `mask`, `animations`). ## Context7 Documentation Protocol Playwright's config shape, CLI flags, and assertion option names change across majors and are documented, not folklore -- never assert a flag, option, or "best practice" from memory. 1. Call `ToolSearch` with query `"context7"` (or `"select:mcp__Context7__resolve-library-id,mcp__Context7__query-docs"`) to load the Context7 tools if not already loaded in this session. 2. Call `mcp__Context7__resolve-library-id` with library name `Playwright` to obtain the current Context7-compatible ID (`/microsoft/playwright`); prefer the resolved ID over guessing. 3. Call `mcp__Context7__query-docs` for the specific claim in question -- e.g. "toHaveScreenshot maxDiffPixelRatio and mask options", "shard CLI flag and blob reporter merge", "storageState project dependencies setup" -- before stating it as fact. Do this per review, not once from a prior session's memory. 4. Prefer the official docs URLs in `official_docs` for primary normative statements (exact CLI flags, exact config shape); use Context7 to ground and cross-check the claim before writing it into a finding. 5. If Context7 is unavailable or returns no relevant match, fall back to the `official_docs` URLs and mark the claim `documentation-based (Context7 unavailable)` rather than presenting it as freshly verified. 6. Never invent a config key, CLI flag, assertion option, or fixture API that no queried source confirms. ## Lean operating rules - Always confirm the installed Playwright version before asserting on API option names; `toHaveScreenshot` options and CLI `--shard` syntax are stable but still verify against the project's `package.json` version, not assumption. - Distinguish flakiness caused by real non-determinism (animation, dynamic content, network timing) from flakiness caused by weak locators or missing waits; the fix differs, and misdiagnosis just hides a real timing bug behind a wider tolerance. - Recommend `animations: 'disabled'` and explicit `mask`/`stylePath` for non-deterministic regions before recommending a looser `maxDiffPixelRatio`/`maxDiffPixels`/`threshold` -- widening tolerance first papers over the actual source of visual noise. - Treat `storageState.json` fixtures as sensitive: they must come from a dedicated test account, never a real user session, and should not be committed if they contain live tokens (see security notes below). - Recommend CI sharding (`--shard=N/M` with a matrix strategy plus `blob` reporter and `merge-reports`) only after confirming the suite's actual wall-clock time in CI justifies the added job complexity and report-merge step; sharding without a merge step silently drops report coverage. - Prefer the setup-project/`dependencies` pattern (a dedicated `setup` project producing `storageState`, consumed via `dependencies: ['setup']`) over ad hoc `globalSetup` for auth when the project already uses Playwright's project model; both are documented, but they compose differently with sharding and per-project storage state. - Do not conflate `fullyParallel` (parallelizes tests within a single file, in addition to across files) with CI-level sharding (`--shard`, distributes files across separate CI jobs/machines) -- they solve different bottlenecks and a suite can need one, both, or neither. - Load the design-token/visual-regression skill instead of this one when the question is about baseline-approval workflow or a third-party visual-review service (e.g. Chromatic), not Playwright's own `toHaveScreenshot` config. ## References Load these only when needed: - [Fixtures, auth setup, and storageState security](references/fixtures-and-auth-setup.md) -- use when reviewing or designing `storageState`/auth fixtures, `setup` project dependencies, `globalSetup`, or handling of `storageState.json`/HAR files as sensitive artifacts. - [CI sharding and parallelism](references/ci-sharding-and-parallelism.md) -- use when configuring or reviewing `--shard`, matrix CI strategy, blob-reporter merge, `fullyParallel`, or `workers` tuning. - [Visual assertion tuning (toHaveScreenshot)](references/visual-assertion-tuning.md) -- use when reviewing or tuning `toHaveScreenshot`/`toMatchSnapshot` options (`animations`, `mask`, `maxDiffPixelRatio`, `maxDiffPixels`, `threshold`, `stylePath`) or diagnosing visual-diff flakiness. ## Response minimum Return, at minimum: - the Playwright feature/config area in scope (fixtures, sharding, screenshot assertion), - evidence level and the exact Playwright version the guidance targets, - root cause of any flakiness identified (not just a threshold-widening patch), - proposed config diff (not applied) with the option names verified against docs, - security caveat on any `storageState`/HAR fixture reviewed.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.