Claude Cursor GitHub Copilot Skill

e2e-testing-playwright-review

Reviews Playwright end-to-end test configuration -- fixtures, storageState/auth setup, CI sharding and parallelism, and toHaveScreenshot visual-assertion options -- for reliability and correct gating, grounded in current, version-specific Playwright API docs.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download vincentchuwaichow-vanguard-frontier-agentic-skills_frontend_e2e-testing-playwright-review-febe32a.zip · 9 KB
Part of vincentchuwaichow/vanguard-frontier-agentic — 293 skills

Install

skills CLI npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/frontend/e2e-testing-playwright-review
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
Git git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

E2E Testing (Playwright) Review

Purpose

Playwright E2E suites fail in two directions: they're flaky enough that teams disable them, or they're so under-configured (no sharding, no masking, no stable waits) that they're slow and noisy without adding confidence. This skill reviews Playwright-specific configuration -- fixtures, auth state, parallelism, and screenshot assertions -- against current official API behavior rather than remembered API shapes that may be stale across majors.

When to use

Use this skill when the user asks to:

  • review or configure Playwright test fixtures, storageState, or auth setup,
  • diagnose Playwright test flakiness (timing, animation, non-deterministic content),
  • configure CI sharding/parallelism for a Playwright suite,
  • review or tune toHaveScreenshot visual-assertion options (maxDiffPixelRatio, mask, animations).

Context7 Documentation Protocol

Playwright's config shape, CLI flags, and assertion option names change across majors and are documented, not folklore -- never assert a flag, option, or "best practice" from memory.

  1. Call ToolSearch with query "context7" (or "select:mcp__Context7__resolve-library-id,mcp__Context7__query-docs") to load the Context7 tools if not already loaded in this session.
  2. Call mcp__Context7__resolve-library-id with library name Playwright to obtain the current Context7-compatible ID (/microsoft/playwright); prefer the resolved ID over guessing.
  3. Call mcp__Context7__query-docs for the specific claim in question -- e.g. "toHaveScreenshot maxDiffPixelRatio and mask options", "shard CLI flag and blob reporter merge", "storageState project dependencies setup" -- before stating it as fact. Do this per review, not once from a prior session's memory.
  4. Prefer the official docs URLs in official_docs for primary normative statements (exact CLI flags, exact config shape); use Context7 to ground and cross-check the claim before writing it into a finding.
  5. If Context7 is unavailable or returns no relevant match, fall back to the official_docs URLs and mark the claim documentation-based (Context7 unavailable) rather than presenting it as freshly verified.
  6. Never invent a config key, CLI flag, assertion option, or fixture API that no queried source confirms.

Lean operating rules

  • Always confirm the installed Playwright version before asserting on API option names; toHaveScreenshot options and CLI --shard syntax are stable but still verify against the project's package.json version, not assumption.
  • Distinguish flakiness caused by real non-determinism (animation, dynamic content, network timing) from flakiness caused by weak locators or missing waits; the fix differs, and misdiagnosis just hides a real timing bug behind a wider tolerance.
  • Recommend animations: 'disabled' and explicit mask/stylePath for non-deterministic regions before recommending a looser maxDiffPixelRatio/maxDiffPixels/threshold -- widening tolerance first papers over the actual source of visual noise.
  • Treat storageState.json fixtures as sensitive: they must come from a dedicated test account, never a real user session, and should not be committed if they contain live tokens (see security notes below).
  • Recommend CI sharding (--shard=N/M with a matrix strategy plus blob reporter and merge-reports) only after confirming the suite's actual wall-clock time in CI justifies the added job complexity and report-merge step; sharding without a merge step silently drops report coverage.
  • Prefer the setup-project/dependencies pattern (a dedicated setup project producing storageState, consumed via dependencies: ['setup']) over ad hoc globalSetup for auth when the project already uses Playwright's project model; both are documented, but they compose differently with sharding and per-project storage state.
  • Do not conflate fullyParallel (parallelizes tests within a single file, in addition to across files) with CI-level sharding (--shard, distributes files across separate CI jobs/machines) -- they solve different bottlenecks and a suite can need one, both, or neither.
  • Load the design-token/visual-regression skill instead of this one when the question is about baseline-approval workflow or a third-party visual-review service (e.g. Chromatic), not Playwright's own toHaveScreenshot config.

References

Load these only when needed:

  • Fixtures, auth setup, and storageState security -- use when reviewing or designing storageState/auth fixtures, setup project dependencies, globalSetup, or handling of storageState.json/HAR files as sensitive artifacts.
  • CI sharding and parallelism -- use when configuring or reviewing --shard, matrix CI strategy, blob-reporter merge, fullyParallel, or workers tuning.
  • Visual assertion tuning (toHaveScreenshot) -- use when reviewing or tuning toHaveScreenshot/toMatchSnapshot options (animations, mask, maxDiffPixelRatio, maxDiffPixels, threshold, stylePath) or diagnosing visual-diff flakiness.

Response minimum

Return, at minimum:

  • the Playwright feature/config area in scope (fixtures, sharding, screenshot assertion),
  • evidence level and the exact Playwright version the guidance targets,
  • root cause of any flakiness identified (not just a threshold-widening patch),
  • proposed config diff (not applied) with the option names verified against docs,
  • security caveat on any storageState/HAR fixture reviewed.
Files (vanguard-frontier-agentic)
  • references
    • ci-sharding-and-parallelism.md 4.1 KB
      # CI Sharding and Parallelism
      
      Use this reference when configuring or reviewing Playwright's `--shard` CLI flag, a CI matrix strategy, blob-reporter merging, `fullyParallel`, or `workers` tuning. For auth-state interactions across sharded/parallel projects, see `fixtures-and-auth-setup.md`.
      
      > Version note: the `blob` reporter and `npx playwright merge-reports` command are the documented mechanism for combining sharded results into one report. Verify exact reporter name and CLI flags against the installed version via Context7 (`/microsoft/playwright`) or the official docs before citing them in a report.
      
      ## Officially grounded design points
      
      - **`fullyParallel` and sharding solve different bottlenecks.** `fullyParallel: true` in `playwright.config.ts` parallelizes tests *within a single file* in addition to across files (by default, Playwright already parallelizes across files using `workers`). CI-level sharding (`--shard=<index>/<total>`) instead splits the *set of test files* across separate CI jobs/machines. A suite bottlenecked by one huge file with many tests needs `fullyParallel`; a suite bottlenecked by total file count/CI wall-clock time needs sharding (or both).
      - **`test.describe.parallel()`** is the documented, file-scoped alternative to global `fullyParallel` -- it opts a specific `describe` block into within-file parallel execution without changing the setting for the whole suite.
      - **Sharding requires a report-merge step to remain useful.** The documented CI pattern is: a matrix job runs `npx playwright test --shard=${{ matrix.shardIndex }}/${{ matrix.shardTotal }}` with `reporter: 'blob'` configured, each shard uploads its `blob-report` directory as a CI artifact, and a separate downstream job downloads all shard artifacts and runs `npx playwright merge-reports --reporter html ./all-blob-reports` to produce one combined report. Sharding without this merge step means each shard produces an isolated, easy-to-miss report instead of one gating artifact.
      - **`workers`** controls the number of parallel worker processes *within* a single Playwright invocation (i.e., within one shard/CI job); it is a separate lever from both `fullyParallel` and `--shard` and is commonly constrained by the CI runner's available CPU/memory, not just desired speed.
      
      ## Non-negotiable design rules
      
      1. **Do not recommend sharding to fix flakiness.** Sharding distributes work across CI jobs; it does not change whether an individual test is flaky and can even change failure order/timing in ways that mask a root cause. If the underlying request is "the suite is unreliable," diagnose flakiness first (locators, waits, isolation) and treat sharding purely as a wall-clock-time lever.
      2. **Never introduce `--shard` without also configuring `reporter: 'blob'` and a merge job.** A sharding change that lands without the corresponding merge-reports job in CI produces N disconnected reports and typically means nobody actually reviews the combined result -- treat this as an incomplete change, not a matter of taste.
      3. **Justify sharding with a measured wall-clock baseline, not a guess.** Require the current CI suite duration before recommending a specific shard count; splitting a suite that already finishes in 3 minutes into 8 shards adds CI queueing/setup overhead without meaningful benefit and increases the merge-job's own failure surface.
      4. **Tune `workers` to the CI runner's actual resources, not a fixed "more is faster" assumption.** Over-provisioning workers on a memory-constrained runner is a documented source of resource contention that manifests as flaky, not slow, tests -- which is easy to misdiagnose as an application bug.
      5. **`fail-fast: false` in a CI matrix is the documented choice for shard jobs** so that one failing shard doesn't cancel the others and hide additional failures that would otherwise surface in the same run.
      
      ## Response discipline
      
      When reviewing a sharding/parallelism setup, state the current shard count and worker configuration as found (repo/CI-config evidence), the measured baseline duration if available (or explicitly flag it as missing), and whether a `merge-reports` step exists -- an incomplete shard-without-merge config is a hard finding, not a style note.
      
    • fixtures-and-auth-setup.md 4.3 KB
      # Fixtures, Auth Setup, and storageState Security
      
      Use this reference when reviewing or designing Playwright `storageState`/auth fixtures, a `setup` project with `dependencies`, `globalSetup`, or when a `storageState.json`/HAR file needs a security review before it's committed or shared. For sharding/parallelism interactions with per-project storage state, see `ci-sharding-and-parallelism.md`.
      
      > Version note: `storageState()` option shapes (e.g. the `indexedDB` capture flag) have been added across Playwright releases. Verify exact option availability against the installed version via Context7 (`/microsoft/playwright`) or the official docs before citing a capability as available.
      
      ## Officially grounded patterns
      
      Playwright's own docs describe two supported, non-exclusive ways to produce a reusable auth state:
      
      - **Setup project + `dependencies`.** A dedicated project (e.g. `{ name: 'setup', testMatch: /.*\.setup\.ts/ }`) performs authentication once and writes `storageState` to a file (e.g. `playwright/.auth/user.json`); browser-specific projects (`chromium`, `firefox`, ...) declare `dependencies: ['setup']` and set `use: { storageState: 'playwright/.auth/user.json' }` to consume it. This is the documented pattern for reusing one authenticated session across multiple projects/browsers without re-authenticating per project.
      - **`globalSetup`.** A `globalSetup` function (referenced from `playwright.config.ts`) launches a browser, performs the login flow, and calls `page.context().storageState({ path: storageState })` once before the whole run. Both the setup-project pattern and `globalSetup` are documented; the setup-project pattern is the one that composes with `dependencies` and per-project overrides, so prefer it when the project already uses Playwright's project model.
      - **API-based auth**, skipping the UI entirely: a `setup` test can call `request.post(...)` against a login endpoint and then `request.storageState({ path: authFile })` to persist cookies/tokens without driving a browser -- faster and less flaky than a UI-driven login, when the app's auth flow supports it.
      - `storageState({ path, indexedDB: true })` is the documented way to also persist IndexedDB-backed session data (relevant for apps that store tokens in IndexedDB rather than cookies/localStorage); omitting `indexedDB: true` for such an app silently drops part of the session and can cause the "auth setup ran but tests still see a logged-out state" failure mode.
      
      ## Non-negotiable design rules
      
      1. **Never generate `storageState.json` against a real user's live session.** It must come from a dedicated, disposable test account with no access to production customer data. A file captured from a real session contains live cookies/tokens that are functionally equivalent to that user's credentials.
      2. **Treat `storageState.json` as a secret artifact, not a build output.** Do not commit it to a public repository. If it must persist across CI runs, store it as a short-lived CI artifact/secret with restricted access, not a tracked file in version control.
      3. **Scrub HAR-file fixtures before committing them.** HAR recordings used for network-mocking (route interception replay) capture full request/response headers, which commonly include `Authorization` headers, API keys, or session cookies. A HAR fixture that hasn't been reviewed for these is a credential-leak risk equivalent to committing a raw token.
      4. **Re-authenticate per test-account tier, not per test.** If the suite exercises multiple roles/permission levels, use multiple named `storageState` files (one per role) produced by distinct `setup` tests, rather than one shared elevated-privilege session reused everywhere -- reusing an admin session for tests that only need a standard-user role overstates the privilege the test actually needs and can mask authorization bugs.
      5. **Expire and rotate test-account credentials used to produce fixtures on a defined cadence.** A `storageState` fixture generated from a test account whose password/token never rotates is a long-lived credential sitting in CI infrastructure.
      
      ## Response discipline
      
      When reviewing fixtures, state explicitly whether the `storageState`/HAR file was confirmed to originate from a dedicated test account (repo evidence: setup-test source, CI secret config) or whether that could not be confirmed from available evidence -- do not assume test-only provenance without a citation.
      
    • visual-assertion-tuning.md 4.6 KB
      # Visual Assertion Tuning (toHaveScreenshot)
      
      Use this reference when reviewing or tuning `toHaveScreenshot`/`toMatchSnapshot` options (`animations`, `mask`, `maskColor`, `stylePath`, `maxDiffPixelRatio`, `maxDiffPixels`, `threshold`) or diagnosing visual-diff flakiness. For flakiness caused by locators/waits rather than visual noise, that is a fixtures/timing question, not a visual-assertion-tuning one.
      
      > Version note: `toHaveScreenshot` option names below are drawn from the current `PageAssertions`/`LocatorAssertions` API reference. Verify exact defaults (e.g. whether `animations` defaults to `'disabled'` in the installed version) against Context7 (`/microsoft/playwright`) or the official API reference before citing a default value as fact.
      
      ## Officially grounded option set
      
      Per the official `PageAssertions.toHaveScreenshot` / `LocatorAssertions.toHaveScreenshot` API reference, the documented tuning options include:
      
      - **`animations`** -- whether to allow or disable animations in the screenshot; the framework's internal screenshot implementation passes `animations: helper.options.animations ?? 'disabled'`, i.e. animations are suppressed by default unless explicitly overridden.
      - **`mask`** -- an array of `Locator`/`ElementHandle` values to mask out of the comparison (paired with `maskColor` to control the mask's fill color).
      - **`stylePath`** -- a path to a CSS file injected into the page before the screenshot is taken, for style-level suppression of non-deterministic elements (e.g. hiding a live clock or a carousel via CSS) that isn't practical to mask element-by-element.
      - **`clip`** / **`fullPage`** -- region-of-capture controls, independent of the diff-tolerance controls below.
      - **Diff-tolerance controls**: `maxDiffPixels` (absolute pixel count allowed to differ), `maxDiffPixelRatio` (ratio of differing pixels, 0-1), and `threshold` (a per-pixel color-difference threshold for the comparator). These can be set per-assertion call or globally in `playwright.config.ts` under `expect.toHaveScreenshot` / `expect.toMatchSnapshot`.
      - Both `toHaveScreenshot` (web-first, waits for two consecutive identical screenshots before comparing -- i.e. it has its own visual-stability wait built in) and `toMatchSnapshot` (compares an already-captured `page.screenshot()` buffer) support the diff-tolerance options; they differ in whether the wait-for-stability behavior is built in.
      
      ## Non-negotiable design rules
      
      1. **Fix the non-determinism before widening tolerance.** If a diff is caused by a specific known-dynamic region (a timestamp, a live counter, an animated element, ad/embed content), the documented, targeted fix is `mask` (or `stylePath` for CSS-level suppression) on that region -- not a global increase to `maxDiffPixelRatio`/`maxDiffPixels`/`threshold`. Widening global tolerance first hides real regressions in the rest of the page along with the intended noise.
      2. **Do not disable `animations` masking as a blanket policy without checking whether the default is already `'disabled'`.** Since the built-in implementation defaults `animations` to `'disabled'` unless overridden, an explicit `animations: 'allow'` (or equivalent override) in a config/test is itself worth flagging and asking why -- it re-introduces a documented source of visual flakiness.
      3. **A widened `maxDiffPixelRatio`/`maxDiffPixels` needs a stated reason tied to a specific region or rendering variance (e.g. font anti-aliasing differences across OS/CI runners), not a round number picked to make CI green.** An unexplained tolerance value is a signal the underlying flake was never diagnosed.
      4. **Global `expect.toHaveScreenshot`/`expect.toMatchSnapshot` config in `playwright.config.ts` sets the default for the whole suite; a per-call override should be justified by why that specific assertion needs different tolerance than the suite default**, not applied broadly as a workaround for one flaky test that then silently loosens every other screenshot assertion using the same call site pattern.
      5. **`omitBackground` and `scale` affect what the baseline image actually contains** (transparency handling, DPI scaling); changing either after a baseline was captured invalidates the existing baseline and requires a deliberate baseline-update pass, not an incidental config tweak.
      
      ## Response discipline
      
      When reviewing visual-assertion config, cite the specific option and value found (with file reference), state whether a wider tolerance is targeted (paired with `mask`/`stylePath` for a named dynamic region) or global/unexplained, and label the API-shape claim as `documentation-based` (grounded via Context7/official API reference this session) versus `inference`.
      
  • metadata.json 1.3 KB
    {
      "id": "e2e-testing-playwright-review",
      "name": "E2E Testing (Playwright) Review",
      "type": "skill",
      "provider": "frontend",
      "harnesses": [
        "claude-code",
        "cursor",
        "codex",
        "gemini",
        "kiro",
        "other"
      ],
      "summary": "Reviews Playwright end-to-end test configuration, fixtures, sharding, and screenshot-assertion setup for reliability, CI parallelism, and accurate visual/behavioral gating, grounded in current Playwright API docs.",
      "source_type": "original",
      "official_docs": [
        "https://playwright.dev/docs/best-practices",
        "https://playwright.dev/docs/ci",
        "https://playwright.dev/docs/test-snapshots",
        "https://playwright.dev/docs/api/class-pageassertions#page-assertions-to-have-screenshot-1",
        "https://playwright.dev/docs/test-parallel"
      ],
      "security_notes": "Playwright storageState.json fixtures can capture real session cookies/auth tokens if generated against a live authenticated session; require these be generated against a dedicated test account and treated as sensitive (not committed to a public repo). HAR-file recordings used for network mocking can capture real API keys in request headers -- scrub before committing.",
      "last_verified": "2026-07-02",
      "path": "skills/frontend/e2e-testing-playwright-review",
      "author": "github: VincentChuWaiChow",
      "version": "0.1.0"
    }
    
  • SKILL.md 5.9 KB
    ---
    name: e2e-testing-playwright-review
    description: Reviews Playwright end-to-end test configuration -- fixtures, storageState/auth setup, CI sharding and parallelism, and toHaveScreenshot visual-assertion options -- for reliability and correct gating, grounded in current, version-specific Playwright API docs.
    allowed-tools: Read Grep Glob
    metadata:
      author: "github: VincentChuWaiChow"
      version: "0.1.0"
      updated: "2026-07-02"
      category: delivery
    ---
    
    # E2E Testing (Playwright) Review
    
    ## Purpose
    
    Playwright E2E suites fail in two directions: they're flaky enough that teams disable them, or they're so under-configured (no sharding, no masking, no stable waits) that they're slow and noisy without adding confidence. This skill reviews Playwright-specific configuration -- fixtures, auth state, parallelism, and screenshot assertions -- against current official API behavior rather than remembered API shapes that may be stale across majors.
    
    ## When to use
    
    Use this skill when the user asks to:
    
    - review or configure Playwright test fixtures, `storageState`, or auth setup,
    - diagnose Playwright test flakiness (timing, animation, non-deterministic content),
    - configure CI sharding/parallelism for a Playwright suite,
    - review or tune `toHaveScreenshot` visual-assertion options (`maxDiffPixelRatio`, `mask`, `animations`).
    
    ## Context7 Documentation Protocol
    
    Playwright's config shape, CLI flags, and assertion option names change across majors and are documented, not folklore -- never assert a flag, option, or "best practice" from memory.
    
    1. Call `ToolSearch` with query `"context7"` (or `"select:mcp__Context7__resolve-library-id,mcp__Context7__query-docs"`) to load the Context7 tools if not already loaded in this session.
    2. Call `mcp__Context7__resolve-library-id` with library name `Playwright` to obtain the current Context7-compatible ID (`/microsoft/playwright`); prefer the resolved ID over guessing.
    3. Call `mcp__Context7__query-docs` for the specific claim in question -- e.g. "toHaveScreenshot maxDiffPixelRatio and mask options", "shard CLI flag and blob reporter merge", "storageState project dependencies setup" -- before stating it as fact. Do this per review, not once from a prior session's memory.
    4. Prefer the official docs URLs in `official_docs` for primary normative statements (exact CLI flags, exact config shape); use Context7 to ground and cross-check the claim before writing it into a finding.
    5. If Context7 is unavailable or returns no relevant match, fall back to the `official_docs` URLs and mark the claim `documentation-based (Context7 unavailable)` rather than presenting it as freshly verified.
    6. Never invent a config key, CLI flag, assertion option, or fixture API that no queried source confirms.
    
    ## Lean operating rules
    
    - Always confirm the installed Playwright version before asserting on API option names; `toHaveScreenshot` options and CLI `--shard` syntax are stable but still verify against the project's `package.json` version, not assumption.
    - Distinguish flakiness caused by real non-determinism (animation, dynamic content, network timing) from flakiness caused by weak locators or missing waits; the fix differs, and misdiagnosis just hides a real timing bug behind a wider tolerance.
    - Recommend `animations: 'disabled'` and explicit `mask`/`stylePath` for non-deterministic regions before recommending a looser `maxDiffPixelRatio`/`maxDiffPixels`/`threshold` -- widening tolerance first papers over the actual source of visual noise.
    - Treat `storageState.json` fixtures as sensitive: they must come from a dedicated test account, never a real user session, and should not be committed if they contain live tokens (see security notes below).
    - Recommend CI sharding (`--shard=N/M` with a matrix strategy plus `blob` reporter and `merge-reports`) only after confirming the suite's actual wall-clock time in CI justifies the added job complexity and report-merge step; sharding without a merge step silently drops report coverage.
    - Prefer the setup-project/`dependencies` pattern (a dedicated `setup` project producing `storageState`, consumed via `dependencies: ['setup']`) over ad hoc `globalSetup` for auth when the project already uses Playwright's project model; both are documented, but they compose differently with sharding and per-project storage state.
    - Do not conflate `fullyParallel` (parallelizes tests within a single file, in addition to across files) with CI-level sharding (`--shard`, distributes files across separate CI jobs/machines) -- they solve different bottlenecks and a suite can need one, both, or neither.
    - Load the design-token/visual-regression skill instead of this one when the question is about baseline-approval workflow or a third-party visual-review service (e.g. Chromatic), not Playwright's own `toHaveScreenshot` config.
    
    ## References
    
    Load these only when needed:
    
    - [Fixtures, auth setup, and storageState security](references/fixtures-and-auth-setup.md) -- use when reviewing or designing `storageState`/auth fixtures, `setup` project dependencies, `globalSetup`, or handling of `storageState.json`/HAR files as sensitive artifacts.
    - [CI sharding and parallelism](references/ci-sharding-and-parallelism.md) -- use when configuring or reviewing `--shard`, matrix CI strategy, blob-reporter merge, `fullyParallel`, or `workers` tuning.
    - [Visual assertion tuning (toHaveScreenshot)](references/visual-assertion-tuning.md) -- use when reviewing or tuning `toHaveScreenshot`/`toMatchSnapshot` options (`animations`, `mask`, `maxDiffPixelRatio`, `maxDiffPixels`, `threshold`, `stylePath`) or diagnosing visual-diff flakiness.
    
    ## Response minimum
    
    Return, at minimum:
    
    - the Playwright feature/config area in scope (fixtures, sharding, screenshot assertion),
    - evidence level and the exact Playwright version the guidance targets,
    - root cause of any flakiness identified (not just a threshold-widening patch),
    - proposed config diff (not applied) with the option names verified against docs,
    - security caveat on any `storageState`/HAR fixture reviewed.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related