Claude Skill

playwright-ops

Playwright end-to-end testing operations - selectors, fixtures, network mocking, auth, parallelism, CI, visual regression, flake hunting. Use for: playwright, e2e test, end-to-end testing, browser test, getByRole, page object, storageState, trace viewer, flaky test, test sharding

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download 0xdarkmatter-claude-mods-skills_playwright-ops-3dfaf0b.zip · 24 KB
Part of 0xdarkmatter/claude-mods — 94 skills

Install

skills CLI npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/playwright-ops
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
Git git clone https://github.com/0xDarkMatter/claude-mods.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Playwright Operations

End-to-end testing with Playwright Test (@playwright/test, TS/JS). A Python flavor (pytest-playwright) exists with the same browser API but pytest-style fixtures — patterns here translate directly; runner config does not.

Quick Start

npm init playwright@latest          # scaffold config + example test + GH Actions workflow
npx playwright test                 # run all tests, all projects
npx playwright test --project=chromium --grep "@smoke"
npx playwright test --ui            # interactive UI mode (watch, time-travel)
npx playwright codegen https://app.local   # record actions -> generated locators
npx playwright show-report          # open last HTML report
npx playwright show-trace trace.zip # inspect a trace

Selector Strategy

Hierarchy — always prefer the highest tier that uniquely matches:

Tier Locator When
1 page.getByRole('button', { name: 'Submit' }) Anything with an ARIA role — buttons, links, headings, textboxes. Tests a11y for free
2 page.getByLabel('Password') Form fields with labels
3 page.getByPlaceholder('name@example.com') Inputs without labels (fix the label instead, when you can)
4 page.getByText('Welcome back') Non-interactive text content
5 page.getByTestId('cart-total') Stable hook when semantics don't disambiguate. Configure attribute via testIdAttribute
6 page.locator('css=...') / xpath= Last resort. Coupled to DOM structure; breaks on refactor

Why: tiers 1–4 locate the way a user perceives the page — resilient to markup changes, and getByRole fails loudly when accessibility regresses. CSS/XPath encode implementation detail.

Narrowing without CSS:

page.getByRole('listitem')
    .filter({ hasText: 'Product 2' })
    .getByRole('button', { name: 'Add to cart' });

page.getByRole('row').filter({ has: page.getByRole('cell', { name: 'Alice' }) });

Web-First Assertions (no manual waits, ever)

// BAD — checks once, races the render; sleeps are flake factories
expect(await page.getByText('welcome').isVisible()).toBe(true);
await page.waitForTimeout(2000);

// GOOD — auto-retries until pass or timeout
await expect(page.getByText('welcome')).toBeVisible();
await expect(page.getByRole('list')).toHaveCount(3);
await expect(page).toHaveURL(/\/dashboard/);
await expect.soft(page.getByTestId('status')).toHaveText('Active'); // don't stop test on failure

Actions (click, fill) auto-wait for actionability (visible, stable, enabled). If you feel the need for waitForTimeout, you're missing an assertion or an await expect(...) on a state change. For async non-DOM conditions use expect.poll(() => fn()) or expect(async () => {...}).toPass().

Lint guard: enable @typescript-eslint/no-floating-promises — a missing await on an assertion is the most common silent-pass bug.

Config Skeleton

Full production template with comments: assets/playwright.config.template.ts

import { defineConfig, devices } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  fullyParallel: true,
  forbidOnly: !!process.env.CI,
  retries: process.env.CI ? 2 : 0,
  workers: process.env.CI ? 1 : undefined,
  reporter: process.env.CI ? 'blob' : 'html',
  use: {
    baseURL: process.env.BASE_URL ?? 'http://localhost:3000',
    trace: 'on-first-retry',
    testIdAttribute: 'data-testid',
  },
  projects: [
    { name: 'setup', testMatch: /.*\.setup\.ts/ },
    {
      name: 'chromium',
      use: { ...devices['Desktop Chrome'], storageState: 'playwright/.auth/user.json' },
      dependencies: ['setup'],
    },
  ],
  webServer: {
    command: 'npm run dev',
    url: 'http://localhost:3000',
    reuseExistingServer: !process.env.CI,
  },
});

Fixtures Decision Tree

What do I need to share/setup?
│
├─ Per-test object (page object, seeded record)
│  └─ test.extend() test-scoped fixture — setup, await use(x), teardown
│
├─ Expensive, safe-to-share resource (DB pool, test account)
│  └─ Worker-scoped: [fn, { scope: 'worker' }] — once per worker process
│
├─ Side effect every test needs (log capture, network stub)
│  └─ Automatic: [fn, { auto: true }] — runs without being referenced
│
├─ Config-tunable value (locale, default item)
│  └─ Option: ['default', { option: true }] — override in projects[].use
│
├─ Fixtures from several modules
│  └─ mergeTests(testA, testB)
│
└─ Auth state per test file/role
   └─ test.use({ storageState: 'playwright/.auth/admin.json' })

POM-as-fixture (modern recommendation) — page objects are fine; instantiating them by hand in every test is not. Inject via fixture:

// fixtures.ts
import { test as base } from '@playwright/test';
import { TodoPage } from './pages/todo-page';

export const test = base.extend<{ todoPage: TodoPage }>({
  todoPage: async ({ page }, use) => {
    const todoPage = new TodoPage(page);
    await todoPage.goto();
    await use(todoPage);          // test body runs here
  },
});
export { expect } from '@playwright/test';

Page objects should expose locators and actions, not assertions wrapped in try/catch, and never store element handles. Details: references/fixtures-and-pom.md

Network & API

Network need?
│
├─ Stub a third-party API           → page.route('**/api/**', r => r.fulfill({ json }))
├─ Tweak a real response            → const res = await route.fetch(); route.fulfill({ response: res, json })
├─ Simulate failure / offline       → route.abort() / route.fulfill({ status: 500 })
├─ Many endpoints, real shapes      → HAR record + replay (page.routeFromHAR, update: true to record)
├─ Pure API test (no browser)       → request fixture / APIRequestContext
├─ Seed data fast, assert via UI    → hybrid: create via request, verify via page
└─ WebSocket traffic                → page.routeWebSocket(url, ws => ws.onMessage(...))

Hybrid seed-via-API, assert-via-UI — the single biggest speed win in most suites:

test('shows new project', async ({ request, page }) => {
  const res = await request.post('/api/projects', { data: { name: 'Apollo' } });
  expect(res.ok()).toBeTruthy();
  await page.goto('/projects');
  await expect(page.getByRole('link', { name: 'Apollo' })).toBeVisible();
});

Rule of thumb: mock third-party dependencies you don't own; exercise your own backend for real (or mock it deliberately in a separate "frontend-isolated" project). Details: references/network-and-api.md

Authentication

Standard pattern — login once in a setup project, reuse storageState everywhere:

// tests/auth.setup.ts
import { test as setup, expect } from '@playwright/test';

setup('authenticate', async ({ page }) => {
  await page.goto('/login');
  await page.getByLabel('Username').fill(process.env.E2E_USER!);
  await page.getByLabel('Password').fill(process.env.E2E_PASS!);
  await page.getByRole('button', { name: 'Sign in' }).click();
  await expect(page.getByTestId('user-menu')).toBeVisible();   // wait for auth to settle!
  await page.context().storageState({ path: 'playwright/.auth/user.json' });
});
Pattern Use when
One setup project + storageState in use One shared account, tests don't mutate server-side user state
Per-role files (admin.json, user.json) + test.use({ storageState }) Role-based behavior under test
Worker-scoped account fixture (testInfo.parallelIndex) Parallel tests mutate user state — one account per worker
API login (request.post + request.storageState) Login endpoint exists; 10x faster than UI login

Gotchas: add playwright/.auth/ to .gitignore. storageState captures cookies + localStorage — not sessionStorage (persist that manually via page.evaluate + init script). Always assert a logged-in signal before saving state, or you save a half-logged-in race.

Parallelism, Retries, Isolation

Knob Setting Notes
Workers workers: process.env.CI ? 1 : undefined Local: half the logical CPU cores. CI runners are small — shard machines instead of oversubscribing
File-level parallel fullyParallel: true Also makes sharding split per-test, not per-file
Sharding npx playwright test --shard=1/4 One shard per CI machine; merge blob reports after
Retries retries: process.env.CI ? 2 : 0 Pair with trace: 'on-first-retry'; treat "flaky" status as a bug queue, not a fix
Serial test.describe.configure({ mode: 'serial' }) Smell — usually means hidden inter-test coupling

Isolation discipline: every test gets a fresh context/page (cookies, storage) — keep it that way. No test reads state written by another test; shared server-side state is reset via API in beforeEach or scoped per worker (test.info().parallelIndex in usernames/tenant IDs). A suite that only passes single-worker is broken, not "sensitive".

Flake diagnosis: trace: 'on-first-retry' → npx playwright show-trace (DOM snapshots, network, console per action). Local: npx playwright test --ui or PWDEBUG=1 / page.pause(). Repro: --repeat-each=20 --workers=4. Playbook: references/flake-hunting.md

Triage a whole run without eyeballing the report — generate the JSON reporter output, then rank the offenders with the bundled triage tool (scripts/triage-flakes.py):

npx playwright test --reporter=json > results.json   # or reporter: [['json', { outputFile: 'results.json' }]]
scripts/triage-flakes.py results.json                # flaky tests first, then hard fails

It emits a ranked TSV (or --json envelope, schema claude-mods.playwright-ops.flake-triage/v1): flaky tests (passed only on retry) first — ordered by retry count then duration — followed by unexpected hard failures, each with file:line, the status sequence (failed->passed), and total duration. Exit 10 means flakes/fails were found (the triage signal — go fix them); exit 0 means a clean suite. --outcome all includes the passing tests for context; -n N caps rows.

CI (GitHub Actions)

- uses: actions/checkout@v5
- uses: actions/setup-node@v5
  with: { node-version: lts/* }
- run: npm ci
- run: npx playwright install --with-deps chromium   # only browsers you test
- run: npx playwright test
- uses: actions/upload-artifact@v4
  if: ${{ !cancelled() }}
  with: { name: playwright-report, path: playwright-report/, retention-days: 30 }
Decision Guidance
Container vs install-deps mcr.microsoft.com/playwright:vX.Y.Z-jammy image pins browser+OS (best for visual tests); install --with-deps is simpler and fine otherwise. Pin image tag to your @playwright/test version
Browser caching Cache ~/.cache/ms-playwright keyed on Playwright version; skip when using the container
Sharded reports reporter: 'blob' on shards → upload blob-report/ → merge job: npx playwright merge-reports --reporter html ./all-blob-reports
Fail-fast vs full suite PRs: fail-fast: false + --max-failures=10 per shard — see all failures in one round-trip. Smoke gates: fail fast

Full workflows (sharding matrix, merge job, caching): references/ci-patterns.md

Visual Testing

await expect(page).toHaveScreenshot('landing.png', {
  maxDiffPixels: 100,                       // or maxDiffPixelRatio / threshold
  mask: [page.getByTestId('ad-banner')],    // black-box dynamic regions
  fullPage: true,
});
  • First run generates the baseline (test fails); update with npx playwright test --update-snapshots
  • Snapshots are named per browser and platform (landing-chromium-darwin.png) — baselines generated on macOS will not match Linux CI. Fix: generate baselines inside the same Docker image CI uses, or run visual tests only in the container
  • Disable animations: toHaveScreenshot defaults animations: 'disabled'; hide dynamic bits with mask or stylePath (CSS applied at capture time)
  • Global defaults: expect: { toHaveScreenshot: { maxDiffPixels: 100 } } in config
  • toMatchSnapshot() for non-image data (text/buffers)

Component Testing & When to Prefer Cypress

@playwright/experimental-ct-react (also vue/svelte) mounts components in a real browser — still experimental; for component-level work, Vitest browser mode or Testing Library are the safer default, with Playwright covering E2E.

Factor Playwright Cypress
Browsers Chromium, Firefox, WebKit (real Safari engine) Chrome-family, Firefox; WebKit experimental
Parallelism Free, built-in, shardable Paid Cloud for parallel orchestration
Multi-tab / multi-origin / iframes Native Historically constrained
API testing Built-in request context Via cy.request, less ergonomic
Component testing Experimental Mature, first-class
In-browser interactive DX UI mode (excellent) The original benchmark; some teams still prefer it

Reach for Cypress when component testing maturity or an existing Cypress investment dominates; otherwise Playwright is the default for new E2E suites. (Repo also has a sibling cypress-ops skill.)

Debugging & Codegen

Tool Command Use
UI mode npx playwright test --ui Watch mode, time-travel, pick locators
Inspector PWDEBUG=1 npx playwright test or page.pause() Step through actions live
Codegen npx playwright codegen <url> Records actions, emits role-based locators — treat output as a draft, refactor into POMs/fixtures
Trace viewer npx playwright show-trace trace.zip Post-mortem: snapshots, network, console
Headed + slow --headed --debug Eyeball a single test
VS Code extension — Run/debug tests, pick locators in-editor

An official Playwright MCP server (@playwright/mcp) also exists for agent-driven browser automation — distinct from the test runner; don't conflate browsing automation with the test suite.

References

File Contents
references/fixtures-and-pom.md Fixture scopes/options/merging, POM-as-fixture architecture, anti-patterns
references/network-and-api.md route/fulfill/abort, HAR replay, API testing, hybrid seeding, WebSocket
references/ci-patterns.md Full GH Actions workflows: basic, sharded+merge, container, caching, reporters
references/flake-hunting.md Systematic flake diagnosis: traces, repro loops, common causes + fixes
scripts/triage-flakes.py Parse a Playwright JSON report and rank flaky/failing tests (exit 10 = findings); see Flake diagnosis above
assets/playwright.config.template.ts Commented production config template
Files (claude-mods)
  • assets
    • playwright.config.template.ts 4.6 KB
      /**
       * Production Playwright config template.
       *
       * Copy to playwright.config.ts and adjust the marked sections.
       * Conventions baked in:
       *   - blob reporter on CI (shard-mergeable), html locally
       *   - trace on first retry (flake forensics at near-zero cost)
       *   - auth via a `setup` project + storageState (login once, reuse everywhere)
       *   - webServer boots the app and waits for it — no sleeps in CI scripts
       */
      import { defineConfig, devices } from '@playwright/test';
      
      export default defineConfig({
        testDir: './tests',
      
        // Run tests within files in parallel too (also gives per-test shard balancing).
        // Set false only if tests within a file are intentionally ordered.
        fullyParallel: true,
      
        // A stray `test.only` committed to CI silently skips the suite — make it a build failure.
        forbidOnly: !!process.env.CI,
      
        // Retries are flake telemetry, not a fix: retried-then-passed tests show as
        // "flaky" in the report. Keep 0 locally so you feel flakes immediately.
        retries: process.env.CI ? 2 : 0,
      
        // Small CI runners (2-core GitHub hosted) thrash with parallel browser workers.
        // Scale horizontally with --shard instead. Locally, default = ~half the cores.
        workers: process.env.CI ? 1 : undefined,
      
        // blob -> uploaded per shard, merged with `npx playwright merge-reports`.
        reporter: process.env.CI
          ? [['blob'], ['github']]
          : [['html', { open: 'on-failure' }]],
      
        // Per-action timeout defaults are usually fine; raise the global test timeout
        // only for genuinely long flows (or per-test via test.slow()).
        timeout: 30_000,
        expect: {
          timeout: 5_000,
          // Global visual-comparison tolerances; override per assertion when needed.
          toHaveScreenshot: {
            maxDiffPixels: 100,
            // animations: 'disabled' is already the default for screenshots
          },
        },
      
        use: {
          // All page.goto('/relative') and request.get('/api/...') resolve against this.
          baseURL: process.env.BASE_URL ?? 'http://localhost:3000',
      
          // Trace on first retry: the failing run gets full DOM snapshots + network log.
          // Use 'retain-on-failure' instead if you run with retries: 0.
          trace: 'on-first-retry',
          screenshot: 'only-on-failure',
          // video is mostly redundant with traces; enable only if your team skims videos.
          // video: 'retain-on-failure',
      
          // Determinism: pin what the OS would otherwise decide for you.
          locale: 'en-US',
          timezoneId: 'UTC',
          // viewport comes from the device preset per project below.
      
          // Attribute used by page.getByTestId(); align with your frontend convention.
          testIdAttribute: 'data-testid',
      
          // Uncomment if a service worker swallows your route() mocks:
          // serviceWorkers: 'block',
        },
      
        projects: [
          // --- Auth setup: runs first, saves storage state for the browser projects ---
          // tests/auth.setup.ts logs in and calls
          //   page.context().storageState({ path: 'playwright/.auth/user.json' })
          // Keep playwright/.auth/ in .gitignore.
          { name: 'setup', testMatch: /.*\.setup\.ts/ },
      
          {
            name: 'chromium',
            use: {
              ...devices['Desktop Chrome'],
              storageState: 'playwright/.auth/user.json',
            },
            dependencies: ['setup'],
          },
      
          // Enable per browser-support matrix. Remember: each enabled project must be
          // installed in CI (npx playwright install --with-deps firefox webkit).
          // {
          //   name: 'firefox',
          //   use: { ...devices['Desktop Firefox'], storageState: 'playwright/.auth/user.json' },
          //   dependencies: ['setup'],
          // },
          // {
          //   name: 'webkit',
          //   use: { ...devices['Desktop Safari'], storageState: 'playwright/.auth/user.json' },
          //   dependencies: ['setup'],
          // },
      
          // Mobile viewport smoke pass.
          // {
          //   name: 'mobile-chrome',
          //   use: { ...devices['Pixel 7'], storageState: 'playwright/.auth/user.json' },
          //   dependencies: ['setup'],
          //   grep: /@smoke/,
          // },
      
          // Unauthenticated flows (login page itself, public pages) — no storageState.
          // {
          //   name: 'chromium-no-auth',
          //   use: { ...devices['Desktop Chrome'] },
          //   testMatch: /.*\.public\.spec\.ts/,
          // },
        ],
      
        // Playwright boots your app and polls `url` until it responds — replaces
        // "npm start & sleep 15" hacks in CI scripts.
        webServer: {
          command: process.env.CI ? 'npm run build && npm run start' : 'npm run dev',
          url: 'http://localhost:3000',
          reuseExistingServer: !process.env.CI,
          timeout: 120_000,
          stdout: 'pipe', // surface app logs in CI output when boot fails
        },
        // Multiple servers? webServer also accepts an array: [{ api }, { frontend }].
      });
      
  • references
    • ci-patterns.md 6.4 KB
      # CI Patterns (GitHub Actions)
      
      Runnable workflows for Playwright in CI, from single-job to sharded fleets. Adapt paths/commands
      for other CI providers — the shape is identical.
      
      ## Baseline Workflow
      
      ```yaml
      # .github/workflows/playwright.yml
      name: Playwright Tests
      on:
        push:
          branches: [main]
        pull_request:
          branches: [main]
      jobs:
        test:
          timeout-minutes: 60
          runs-on: ubuntu-latest
          steps:
            - uses: actions/checkout@v5
            - uses: actions/setup-node@v5
              with:
                node-version: lts/*
            - name: Install dependencies
              run: npm ci
            - name: Install Playwright browsers
              run: npx playwright install --with-deps chromium
            - name: Run Playwright tests
              run: npx playwright test
            - uses: actions/upload-artifact@v4
              if: ${{ !cancelled() }}          # upload report on failure too
              with:
                name: playwright-report
                path: playwright-report/
                retention-days: 30
      ```
      
      Notes:
      
      - Install only the browsers your projects use (`chromium` above) — saves minutes per run.
      - `if: ${{ !cancelled() }}` keeps the report when tests fail; that's when you need it.
      - Secrets via `env:` on the test step (`E2E_USER: ${{ secrets.E2E_USER }}`), never committed.
      
      ## Container vs install-deps
      
      | Approach | Pros | Cons |
      |----------|------|------|
      | `npx playwright install --with-deps` on the runner | Simple; matches local dev | OS-level rendering drifts with runner image updates — visual baselines can churn |
      | `container: mcr.microsoft.com/playwright:v1.52.0-jammy` | Pinned browser + OS rendering; reproducible visual tests; no install step | Slightly slower job start; must bump tag with `@playwright/test` |
      
      **Always pin the container tag to your exact `@playwright/test` version** — a mismatch produces
      "Executable doesn't exist" or subtle behavior skew.
      
      ```yaml
      jobs:
        test:
          runs-on: ubuntu-latest
          container:
            image: mcr.microsoft.com/playwright:v1.52.0-jammy
          steps:
            - uses: actions/checkout@v5
            - run: npm ci
            - run: npx playwright test
              env:
                HOME: /root      # workaround for firefox in containers
      ```
      
      ## Caching Browsers (non-container path)
      
      ```yaml
      - name: Get Playwright version
        id: pw-version
        run: echo "version=$(node -p "require('@playwright/test/package.json').version")" >> "$GITHUB_OUTPUT"
      - uses: actions/cache@v4
        id: pw-cache
        with:
          path: ~/.cache/ms-playwright
          key: playwright-${{ runner.os }}-${{ steps.pw-version.outputs.version }}
      - run: npx playwright install --with-deps chromium
        if: steps.pw-cache.outputs.cache-hit != 'true'
      - run: npx playwright install-deps chromium     # OS deps aren't cached
        if: steps.pw-cache.outputs.cache-hit == 'true'
      ```
      
      ## Sharding with Blob Reports + Merge
      
      Config side — blob on CI shards, html locally:
      
      ```ts
      // playwright.config.ts
      reporter: process.env.CI ? 'blob' : 'html',
      fullyParallel: true,   // shards split per-test instead of per-file -> better balance
      ```
      
      ```yaml
      jobs:
        playwright-tests:
          runs-on: ubuntu-latest
          strategy:
            fail-fast: false                  # let every shard finish; see ALL failures
            matrix:
              shardIndex: [1, 2, 3, 4]
              shardTotal: [4]
          steps:
            - uses: actions/checkout@v5
            - uses: actions/setup-node@v5
              with: { node-version: lts/* }
            - run: npm ci
            - run: npx playwright install --with-deps chromium
            - run: npx playwright test --shard=${{ matrix.shardIndex }}/${{ matrix.shardTotal }}
            - uses: actions/upload-artifact@v4
              if: ${{ !cancelled() }}
              with:
                name: blob-report-${{ matrix.shardIndex }}
                path: blob-report
                retention-days: 1
      
        merge-reports:
          if: ${{ !cancelled() }}
          needs: [playwright-tests]
          runs-on: ubuntu-latest
          steps:
            - uses: actions/checkout@v5
            - uses: actions/setup-node@v5
              with: { node-version: lts/* }
            - run: npm ci
            - uses: actions/download-artifact@v4
              with:
                path: all-blob-reports
                pattern: blob-report-*
                merge-multiple: true
            - run: npx playwright merge-reports --reporter html ./all-blob-reports
            - uses: actions/upload-artifact@v4
              with:
                name: html-report--attempt-${{ github.run_attempt }}
                path: playwright-report
                retention-days: 14
      ```
      
      `merge-reports` accepts multiple reporters: `--reporter html,github` annotates the PR while also
      producing the browsable report.
      
      ## Fail-Fast vs Full-Suite
      
      | Context | Strategy |
      |---------|----------|
      | PR validation | `fail-fast: false` on the matrix + `maxFailures: 10` (or `--max-failures`) per shard. Developers fix everything in one round-trip instead of whack-a-mole |
      | Smoke gate before deploy | Fail fast — `--grep @smoke`, no retries, abort pipeline on first failure |
      | Nightly full regression | Full suite, retries on, no fail-fast; route the merged report to the team channel |
      
      ## Reporters
      
      | Reporter | Use |
      |----------|-----|
      | `html` | Local + merged CI artifact — the daily driver |
      | `blob` | Shard intermediate; only input for `merge-reports` |
      | `junit` | Test-management ingestion (Jenkins, Azure DevOps, TestRail): `['junit', { outputFile: 'results.xml' }]` |
      | `github` | Inline PR annotations on failures |
      | `list` / `dot` / `line` | Console verbosity choices |
      
      Multiple at once:
      
      ```ts
      reporter: process.env.CI
        ? [['blob'], ['github']]
        : [['html', { open: 'on-failure' }]],
      ```
      
      ## webServer in CI
      
      ```ts
      webServer: {
        command: 'npm run build && npm run start',
        url: 'http://localhost:3000',
        reuseExistingServer: !process.env.CI,   // CI always boots fresh
        timeout: 120_000,
        stdout: 'pipe',                          // surface server logs in CI output
      },
      ```
      
      Playwright waits for `url` to respond before running tests — no `sleep 10` hacks. Multiple
      servers (API + frontend) can be given as an array.
      
      ## CI Hardening Checklist
      
      - [ ] `forbidOnly: !!process.env.CI` — a stray `test.only` fails the build instead of silently skipping the suite
      - [ ] `retries: 2` on CI + `trace: 'on-first-retry'`
      - [ ] `workers: 1` per shard on small runners (2-core GitHub runners thrash above that); scale via shards
      - [ ] Report artifacts uploaded with `if: ${{ !cancelled() }}`
      - [ ] Browser install scoped to actual projects
      - [ ] Container tag or browser cache keyed to the Playwright version
      - [ ] Visual-test baselines generated in the same environment CI runs (see SKILL.md Visual Testing)
      
    • fixtures-and-pom.md 6.6 KB
      # Fixtures and Page Object Architecture
      
      How to structure Playwright Test suites with fixtures as the composition mechanism and page
      objects as thin locator/action wrappers.
      
      ## Built-in Fixtures
      
      | Fixture | Type | Scope | Notes |
      |---------|------|-------|-------|
      | `page` | `Page` | test | Fresh isolated page per test |
      | `context` | `BrowserContext` | test | Fresh context per test — cookies/storage isolated |
      | `browser` | `Browser` | worker | Shared across tests in a worker |
      | `browserName` | `string` | worker | `'chromium' \| 'firefox' \| 'webkit'` |
      | `request` | `APIRequestContext` | test | HTTP client honoring `baseURL` / `extraHTTPHeaders` |
      
      ## Custom Fixtures: the Full Shape
      
      ```ts
      import { test as base } from '@playwright/test';
      
      type TestFixtures = {
        todoPage: TodoPage;        // test-scoped
        defaultItem: string;       // option
      };
      type WorkerFixtures = {
        account: { username: string; password: string };  // worker-scoped
      };
      
      export const test = base.extend<TestFixtures, WorkerFixtures>({
        // Option — overridable per project via projects[].use
        defaultItem: ['Something nice', { option: true }],
      
        // Test-scoped fixture with setup + teardown
        todoPage: async ({ page, defaultItem }, use) => {
          const todoPage = new TodoPage(page);
          await todoPage.goto();
          await todoPage.addToDo(defaultItem);
          await use(todoPage);              // <-- test body executes here
          await todoPage.removeAll();       // teardown runs even if test fails
        },
      
        // Worker-scoped — once per worker process; second generic param
        account: [async ({ browser }, use, workerInfo) => {
          const username = 'user-' + workerInfo.workerIndex;
          const password = await createAccount(username);   // expensive, do once
          await use({ username, password });
          await deleteAccount(username);
        }, { scope: 'worker' }],
      });
      
      export { expect } from '@playwright/test';
      ```
      
      Key mechanics:
      
      - **Lazy**: a fixture only runs if the test (or another fixture) references it.
      - **Composable**: fixtures depend on other fixtures by destructuring them.
      - **Teardown order**: reverse of setup, runs even on failure — replaces brittle `afterEach` chains.
      - The two generic params of `extend<TestFixtures, WorkerFixtures>` map to test scope and worker
        scope respectively. Worker fixtures cannot depend on test fixtures.
      
      ## Fixture Options Reference
      
      | Option | Effect |
      |--------|--------|
      | `{ scope: 'worker' }` | One instance per worker process |
      | `{ auto: true }` | Runs for every test without being referenced — global hooks |
      | `{ option: true }` | Value is a project-configurable option |
      | `{ timeout: 60_000 }` | Separate timeout for slow fixture setup |
      | `{ box: true }` | Hide fixture from report/errors (or `box: 'self'` to hide just its step) |
      | `{ title: 'my fixture' }` | Custom name in reports |
      
      ### Automatic fixtures as global hooks
      
      ```ts
      export const test = base.extend<{ forEachTest: void }, { forEachWorker: void }>({
        // beforeEach/afterEach equivalent, but reusable across files
        forEachTest: [async ({ page }, use) => {
          await page.goto('/');         // before each test
          await use();
          // after each test
        }, { auto: true }],
      
        // once per worker
        forEachWorker: [async ({}, use) => {
          console.log(`Worker ${test.info().workerIndex} starting`);
          await use();
        }, { scope: 'worker', auto: true }],
      });
      ```
      
      ### Overriding built-ins
      
      ```ts
      export const test = base.extend({
        page: async ({ page }, use) => {
          await page.goto('/dashboard');   // every test starts on dashboard
          await use(page);
        },
        // Override storageState to come from a worker fixture (per-worker auth)
        storageState: ({ workerStorageState }, use) => use(workerStorageState),
      });
      ```
      
      ## mergeTests: Composing Fixture Modules
      
      Keep fixture concerns in separate modules and merge at the edge:
      
      ```ts
      // fixtures/db.ts        -> export const test = base.extend<{ db: Db }>({...})
      // fixtures/a11y.ts      -> export const test = base.extend<{ axe: Axe }>({...})
      
      // fixtures/index.ts
      import { mergeTests, mergeExpects } from '@playwright/test';
      import { test as dbTest } from './db';
      import { test as a11yTest } from './a11y';
      
      export const test = mergeTests(dbTest, a11yTest);
      export { expect } from '@playwright/test';
      ```
      
      Tests import `test`/`expect` from your fixtures module, never from `@playwright/test` directly —
      one import path to rule them all.
      
      ## Page Objects: Modern Recommendation
      
      ### POM-as-fixture (preferred)
      
      The page object is a plain class; the **fixture** owns construction and navigation:
      
      ```ts
      // pages/checkout-page.ts
      import { type Page, type Locator, expect } from '@playwright/test';
      
      export class CheckoutPage {
        readonly page: Page;
        readonly cardNumber: Locator;
        readonly payButton: Locator;
      
        constructor(page: Page) {
          this.page = page;
          this.cardNumber = page.getByLabel('Card number');
          this.payButton = page.getByRole('button', { name: 'Pay now' });
        }
      
        async goto() {
          await this.page.goto('/checkout');
        }
      
        async pay(card: string) {
          await this.cardNumber.fill(card);
          await this.payButton.click();
        }
      }
      ```
      
      ```ts
      // a test
      test('pays with valid card', async ({ checkoutPage, page }) => {
        await checkoutPage.pay('4242 4242 4242 4242');
        await expect(page.getByRole('heading', { name: 'Order confirmed' })).toBeVisible();
      });
      ```
      
      ### Rules for healthy page objects
      
      | Rule | Why |
      |------|-----|
      | Store `Locator`s, never element handles | Locators are lazy + auto-retrying; handles go stale |
      | Expose actions + locators; keep assertions in tests (or custom `expect` matchers) | Tests stay readable as specs; POMs stay reusable |
      | No `waitForTimeout` / try-catch flow control inside POMs | Hides flake; actions already auto-wait |
      | Constructor takes `Page` (or a `Locator` root for component objects) only | Keeps them trivially fixture-injectable |
      | Prefer small per-screen objects over one God object | Cheap to compose via fixtures |
      
      ### When to skip POMs entirely
      
      Small suites (< ~20 tests) over stable UIs often read better with raw `getByRole` calls inline.
      POMs earn their keep when the same screen appears in many tests or locators churn. Don't build the
      abstraction before the duplication exists.
      
      ## Hooks vs Fixtures
      
      | Need | Use |
      |------|-----|
      | Shared setup local to one file | `test.beforeEach` is fine |
      | Shared setup across files | Fixture (auto or named) |
      | Expensive once-per-run setup | Project dependencies (setup project) — not `globalSetup`, which skips fixtures/tracing |
      | Once-per-worker setup | Worker-scoped fixture |
      
      `test.beforeAll` runs **once per worker**, not once per run — a classic surprise. For true
      once-per-run work, use a setup project with `dependencies`.
      
    • flake-hunting.md 5.8 KB
      # Flake Hunting
      
      Systematic diagnosis of flaky Playwright tests. A flaky test is a bug — in the test, the app, or
      the environment. Retries buy time to fix it; they are not the fix.
      
      ## Triage Workflow
      
      ```
      Flaky test reported
      │
      1. Get the evidence
      │  └─ trace: 'on-first-retry' in config → download trace from CI artifact
      │     └─ npx playwright show-trace path/to/trace.zip
      │        (per-action DOM snapshots, network, console, timing)
      │
      2. Reproduce locally
      │  └─ npx playwright test failing.spec.ts --repeat-each=20 --workers=4
      │     ├─ Fails alone, repeated        → timing/race within the test
      │     ├─ Fails only with --workers>1  → cross-test state leakage
      │     └─ Fails only in CI             → environment delta (speed, viewport, headless, locale, TZ)
      │
      3. Classify against the table below, fix the CAUSE
      │
      4. Prove the fix
         └─ --repeat-each=50 clean, then watch the "flaky" count in CI reports trend to zero
      ```
      
      CI flake visibility: HTML report marks retried-then-passed tests as **flaky** — review that list
      weekly; it's your queue.
      
      ## Common Causes and Fixes
      
      | Symptom | Root cause | Fix |
      |---------|-----------|-----|
      | Click "worked" but nothing happened | Element re-rendered between locate and click (hydration, list re-sort) | Assert the settled state first: `await expect(row).toBeVisible()` then act; prefer role/text locators that target the final element |
      | Assertion passes locally, times out in CI | CI is slower; manual check raced the render | Replace any non-retrying check with web-first `await expect(...)`; raise `expect.timeout` only if the app is legitimately slow |
      | `waitForTimeout` sprinkled around | Sleeping instead of waiting for a condition | Delete; wait on the observable effect: `expect(locator)`, `page.waitForURL()`, `page.waitForResponse()` |
      | Fails only with multiple workers | Tests share an account/record; one mutates what another reads | Per-worker data: suffix usernames/tenants with `test.info().parallelIndex`; or worker-scoped account fixture |
      | First test after auth flaky | storageState saved before login finished | In auth.setup, assert a logged-in signal (`await expect(page.getByTestId('user-menu')).toBeVisible()`) before `storageState({ path })` |
      | Animation mid-flight in screenshots/clicks | CSS transitions | `toHaveScreenshot` disables animations by default; for actions, assert post-animation state or set `reducedMotion: 'reduce'` in `use` |
      | Time-dependent failures (midnight, month-end, TZ) | Real clock | `await page.clock.setFixedTime(new Date('2026-01-15T10:00:00'))`; pin `timezoneId` and `locale` in `use` |
      | Random data collisions | Shared fixtures with hardcoded names | Unique-per-test names: `` `proj-${test.info().testId}` `` |
      | Network nondeterminism from third parties | Live external calls | Mock them (`route.fulfill` / HAR replay) — see network-and-api.md |
      | Passes in `--headed`, fails headless | Viewport/focus/rendering differences | Pin `viewport` in config; debug headless with traces, not by switching to headed |
      | Fails only on retry / second run | Leftover server-side state from first attempt | Make setup idempotent (upsert, not create); clean up in fixture teardown, which runs on failure too |
      
      ## Tools Reference
      
      | Tool | Invocation | What it gives you |
      |------|-----------|-------------------|
      | Trace viewer | `trace: 'on-first-retry'` → `npx playwright show-trace trace.zip` | Time-travel DOM snapshots, network log, console, action timeline — the primary CI forensic tool |
      | UI mode | `npx playwright test --ui` | Watch mode + live trace while iterating on a fix |
      | Inspector | `PWDEBUG=1 npx playwright test foo.spec.ts` or `await page.pause()` | Step through actions, try locators live |
      | Repeat | `--repeat-each=20` | Statistical reproduction |
      | Stress | `--workers=4` (or more than usual) | Surfaces isolation bugs |
      | Single worker | `--workers=1` | If this "fixes" it, you have cross-test coupling — that's the bug |
      | Verbose API log | `DEBUG=pw:api npx playwright test` | Every Playwright call with timing |
      | Video | `video: 'retain-on-failure'` | Cheaper than trace to skim; less data |
      
      `trace: 'on'` everywhere is expensive — `'on-first-retry'` is the right default; use
      `'retain-on-failure'` if you run without retries.
      
      ## Retrying Non-DOM Conditions
      
      Web-first assertions only retry on locators/page. For everything else:
      
      ```ts
      // Poll an arbitrary async value
      await expect.poll(async () => {
        const res = await request.get(`/api/jobs/${id}`);
        return (await res.json()).status;
      }, { timeout: 30_000, intervals: [1_000] }).toBe('done');
      
      // Retry a block of assertions/actions together
      await expect(async () => {
        const res = await request.get('/health');
        expect(res.status()).toBe(200);
      }).toPass({ timeout: 60_000 });
      ```
      
      Use these for eventual consistency (queues, search indexing, emails) instead of sleep loops.
      
      ## Isolation Discipline Checklist
      
      - [ ] No test depends on another test having run (`test.describe.configure({ mode: 'serial' })` is a red flag, not a tool of first resort)
      - [ ] Server-side state is created per test (API seeding) or per worker (`parallelIndex`-scoped accounts)
      - [ ] Teardown lives in fixtures (runs on failure), not at the end of test bodies
      - [ ] `storageState` files saved only after asserting login completed
      - [ ] `forbidOnly` on CI; `--repeat-each` smoke before merging new specs
      - [ ] Suite passes with `--workers=8 --repeat-each=3` locally before you blame CI
      
      ## Quarantine Pattern
      
      While a flake is being fixed, tag it instead of deleting or `.skip`-ing silently:
      
      ```ts
      test('checkout under load @quarantine', async ({ page }) => { ... });
      ```
      
      ```bash
      npx playwright test --grep-invert @quarantine        # main gate
      npx playwright test --grep @quarantine               # nightly, non-blocking
      ```
      
      Track quarantined tests with an issue each; a quarantine list that only grows is a suite dying in
      slow motion.
      
    • network-and-api.md 5.9 KB
      # Network Mocking and API Testing
      
      Intercepting browser traffic, replaying HAR recordings, testing APIs directly, and the hybrid
      seed-via-API / assert-via-UI pattern.
      
      ## Route Interception: page.route()
      
      ```ts
      // Stub an endpoint entirely
      await page.route('*/**/api/v1/fruits', async route => {
        await route.fulfill({ json: [{ name: 'Strawberry', id: 21 }] });
      });
      
      // Must be registered BEFORE the navigation/action that triggers the request
      await page.goto('/');
      ```
      
      | Method | Effect |
      |--------|--------|
      | `route.fulfill({ json, status, headers, body, path })` | Respond without hitting the network |
      | `route.fetch()` | Execute the real request, get the response for modification |
      | `route.continue({ headers, postData, url })` | Pass through, optionally modified |
      | `route.abort('failed')` | Simulate network failure |
      | `route.fallback()` | Defer to the next matching handler (handlers run last-registered-first) |
      
      ### Modify a real response
      
      ```ts
      await page.route('*/**/api/v1/fruits', async route => {
        const response = await route.fetch();
        const json = await response.json();
        json.push({ name: 'Loquat', id: 100 });
        await route.fulfill({ response, json });   // real status/headers, patched body
      });
      ```
      
      ### Failure-mode tests
      
      ```ts
      await page.route('**/api/orders', route => route.fulfill({ status: 500 }));
      await page.route('**/*.{png,jpg,jpeg}', route => route.abort());   // block heavy assets
      await context.setOffline(true);                                     // whole-context offline
      ```
      
      ### Scope and ordering gotchas
      
      - `page.route` applies to that page; `context.route` to every page in the context (use in a
        fixture for suite-wide stubs).
      - Patterns: glob (`**/api/**`), RegExp, or predicate function. Glob matches the **full URL**.
      - `await page.unroute(pattern)` removes handlers; `page.unrouteAll()` clears them.
      - Service workers can bypass routing — set `serviceWorkers: 'block'` in `use` if your app
        registers one and mocks mysteriously don't fire.
      
      ## HAR Record and Replay
      
      Best for "many endpoints, realistic payloads" — record once against the real backend, replay
      hermetically.
      
      ```ts
      // Record (update: true hits the real network and refreshes the file)
      await page.routeFromHAR('./hars/fruits.har', {
        url: '*/**/api/v1/**',
        update: true,
      });
      
      // Replay (default update: false serves from the file; unmatched requests are aborted)
      await page.routeFromHAR('./hars/fruits.har', { url: '*/**/api/v1/**' });
      ```
      
      CLI recording:
      
      ```bash
      npx playwright open --save-har=example.har --save-har-glob="**/api/**" https://example.com
      ```
      
      Workflow: re-run recording tests with `update: true` whenever the API contract changes, commit the
      HAR + extracted bodies (`.txt`/`.json` sidecars are editable by hand for edge cases).
      
      ## API Testing: request / APIRequestContext
      
      The `request` fixture is an HTTP client honoring config `baseURL` and `extraHTTPHeaders` — no
      browser involved, so it's fast.
      
      ```ts
      // playwright.config.ts
      use: {
        baseURL: 'https://api.github.com',
        extraHTTPHeaders: {
          'Accept': 'application/vnd.github.v3+json',
          'Authorization': `token ${process.env.API_TOKEN}`,
        },
      },
      ```
      
      ```ts
      test('creates a bug report', async ({ request }) => {
        const newIssue = await request.post(`/repos/${USER}/${REPO}/issues`, {
          data: { title: '[Bug] report 1', body: 'Bug description' },
        });
        expect(newIssue.ok()).toBeTruthy();
      
        const issues = await request.get(`/repos/${USER}/${REPO}/issues`);
        expect(await issues.json()).toContainEqual(
          expect.objectContaining({ title: '[Bug] report 1' }),
        );
      });
      ```
      
      Standalone context (different base URL, custom auth, use in setup scripts):
      
      ```ts
      import { request } from '@playwright/test';
      
      const api = await request.newContext({ baseURL: 'https://api.example.com' });
      await api.post('/seed', { data: {...} });
      await api.dispose();   // always dispose manually created contexts
      ```
      
      `request.post` options: `data` (JSON), `form` (urlencoded), `multipart` (file upload),
      `params` (query), `headers`, `failOnStatusCode`.
      
      ## Hybrid: Seed via API, Assert via UI
      
      UI-driven setup is the slowest, flakiest part of most suites. Replace it:
      
      ```ts
      test('renders the new project card', async ({ request, page }) => {
        // Arrange — fast, deterministic, server-side
        const res = await request.post('/api/projects', { data: { name: 'Apollo' } });
        expect(res.ok()).toBeTruthy();
        const { id } = await res.json();
      
        // Act + Assert — the only part that needs a browser
        await page.goto(`/projects/${id}`);
        await expect(page.getByRole('heading', { name: 'Apollo' })).toBeVisible();
      });
      ```
      
      Notes:
      
      - The `request` fixture shares `storageState` with the browser context, so an authenticated UI
        session usually authenticates API calls too. For a different principal, create a separate
        `request.newContext({ storageState: 'playwright/.auth/admin.json' })`.
      - Postcondition checks invert it: act in the UI, verify via `request.get` that the server really
        persisted the thing.
      - Cleanup belongs in fixtures (teardown after `use()`) or `afterAll` API calls — not in the test
        body where a failure skips it.
      
      ## What to Mock vs Exercise
      
      | Dependency | Default |
      |------------|---------|
      | Third-party SaaS (payments, analytics, maps) | **Mock** (`route.fulfill` / HAR). You can't control their data or uptime, and you don't want test purchases |
      | Your own backend | **Real** — that's the integration you're paying E2E tests to verify |
      | Your backend, in a frontend-only project | Mock deliberately and label the project (`name: 'ui-isolated'`) so coverage claims stay honest |
      | Time / randomness | `page.clock.install()` / `page.clock.setFixedTime(...)` for time-dependent UI |
      
      ## WebSocket Mocking
      
      ```ts
      await page.routeWebSocket('wss://example.com/ws', ws => {
        ws.onMessage(message => {
          if (message === 'request') ws.send('response');
        });
      });
      ```
      
      By default the intercepted socket never reaches the server; call `ws.connectToServer()` inside the
      handler to proxy with selective message rewriting.
      
  • scripts
    • .gitkeep 0 B · in bundle
    • triage-flakes.py 10.2 KB
      #!/usr/bin/env python3
      # Rank Playwright tests by flakiness from a JSON report so the agent triages, not eyeballs.
      #
      # Parses a Playwright JSON report (`--reporter=json`) and surfaces the tests
      # worth a human's attention: flaky tests (passed only on retry) first, then
      # hard "unexpected" failures. Flaky tests are ranked by retry count desc, then
      # total duration desc, because the most-retried, slowest test is the worst
      # offender in your queue.
      #
      # Usage:   triage-flakes.py [OPTIONS] [REPORT]
      # Input:   REPORT = path to a Playwright JSON report (positional, default ./results.json)
      # Output:  stdout = ranked findings (TSV, or JSON envelope with --json)
      # Stderr:  headers, summary, progress, errors
      # Exit:    0 parsed fine, no flaky/unexpected tests (clean suite)
      #          2 usage, 3 file not found, 4 malformed/not a Playwright report,
      #          10 DOMAIN SIGNAL: flaky/unexpected tests present (the thing being triaged)
      #
      # Examples:
      #   npx playwright test --reporter=json > results.json
      #   triage-flakes.py results.json
      #   triage-flakes.py --outcome all -n 50 results.json
      #   triage-flakes.py --json results.json | jq '.data[] | select(.outcome=="flaky")'
      
      import argparse
      import json
      import os
      import sys
      from pathlib import Path
      
      # Windows consoles default to cp1252; force UTF-8 so glyphs in framing don't raise
      # UnicodeEncodeError (the repo's standard fix).
      for _stream in (sys.stdout, sys.stderr):
          try:
              _stream.reconfigure(encoding="utf-8")  # type: ignore[attr-defined]
          except (AttributeError, ValueError):
              pass
      
      
      class Term:
          """Tiny ANSI helper mirroring skills/_lib/term.sh (bash-only; per
          TERMINAL-DESIGN.md §9 the Python port is inline). Honors FORCE_COLOR /
          NO_COLOR / TERM_ASCII; ASCII glyph fallback on TERM_ASCII or a non-UTF stream."""
      
          _C = {"green": "\033[32m", "yellow": "\033[33m", "orange": "\033[38;5;208m",
                "red": "\033[31m", "cyan": "\033[36m", "dim": "\033[2m", "off": "\033[0m"}
          _GLYPH = {"ok": "✓", "bad": "✗", "warn": "▲", "skip": "—", "na": "—", "unknown": "?"}
          _ASCII = {"ok": "+", "bad": "x", "warn": "!", "skip": "-", "na": "-", "unknown": "?"}
          _MARK_COLOR = {"ok": "green", "bad": "red", "warn": "orange", "skip": "dim",
                         "na": "dim", "unknown": "yellow"}
      
          def __init__(self, stream=sys.stderr):
              enc = (getattr(stream, "encoding", "") or "").lower()
              self.ascii = (os.environ.get("TERM_ASCII") == "1"
                            or os.environ.get("FLEET_ASCII") == "1" or "utf" not in enc)
              if os.environ.get("FORCE_COLOR"):
                  self.color = True
              elif (os.environ.get("NO_COLOR") is not None or os.environ.get("TERM") == "dumb"
                    or not getattr(stream, "isatty", lambda: False)()):
                  self.color = False
              else:
                  self.color = True
      
          def c(self, name, text):
              return f"{self._C.get(name, '')}{text}{self._C['off']}" if self.color else text
      
          def mark(self, state):
              return self.c(self._MARK_COLOR.get(state, ""),
                            (self._ASCII if self.ascii else self._GLYPH).get(state, "."))
      
          def hdr(self, text):
              return self.c("cyan", f"=== {text} ===")
      
      
      TERM = Term(sys.stderr)
      
      SCHEMA = "claude-mods.playwright-ops.flake-triage/v1"
      
      EXIT_OK = 0
      EXIT_USAGE = 2
      EXIT_NOT_FOUND = 3
      EXIT_VALIDATION = 4
      EXIT_FINDINGS = 10
      
      # Rank order for outcomes: flaky always sorts before unexpected.
      OUTCOME_RANK = {"flaky": 0, "unexpected": 1}
      
      
      def err(msg):
          print(msg, file=sys.stderr)
      
      
      def walk_suites(suites, finds, file_hint=""):
          """Recursively descend the suites tree collecting spec/test results."""
          for suite in suites or []:
              # A suite's file is on the suite node; specs inherit it.
              sfile = suite.get("file") or file_hint
              for spec in suite.get("specs", []) or []:
                  collect_spec(spec, finds, sfile)
              walk_suites(suite.get("suites"), finds, sfile)
      
      
      def collect_spec(spec, finds, sfile):
          title = spec.get("title", "<untitled>")
          sline = spec.get("line", 0)
          sfile = spec.get("file") or sfile
          for test in spec.get("tests", []) or []:
              outcome = test.get("status") or test.get("outcome") or "unknown"
              results = test.get("results", []) or []
              # status sequence ordered by retry index; duration summed across attempts
              ordered = sorted(results, key=lambda r: r.get("retry", 0))
              statuses = [r.get("status", "unknown") for r in ordered]
              duration = sum(int(r.get("duration", 0) or 0) for r in ordered)
              retries = max((r.get("retry", 0) for r in ordered), default=0)
              location = f"{sfile}:{sline}" if sfile else f"?:{sline}"
              finds.append(
                  {
                      "title": title,
                      "location": location,
                      "outcome": outcome,
                      "retries": retries,
                      "statuses": statuses,
                      "durationMs": duration,
                  }
              )
      
      
      def load_report(path):
          """Return parsed Playwright report dict, or raise ValueError if not one."""
          try:
              raw = path.read_text(encoding="utf-8")
          except OSError as e:
              raise FileNotFoundError(str(e))
          try:
              data = json.loads(raw)
          except json.JSONDecodeError as e:
              raise ValueError(f"not valid JSON: {e}")
          if not isinstance(data, dict) or "suites" not in data:
              raise ValueError("missing top-level 'suites' key - not a Playwright JSON report")
          if not isinstance(data["suites"], list):
              raise ValueError("'suites' is not a list — not a Playwright JSON report")
          return data
      
      
      def main(argv=None):
          p = argparse.ArgumentParser(
              prog="triage-flakes.py",
              description="Rank Playwright tests by flakiness from a JSON report.",
              formatter_class=argparse.RawDescriptionHelpFormatter,
              epilog=(
                  "EXAMPLES:\n"
                  "  npx playwright test --reporter=json > results.json\n"
                  "  triage-flakes.py results.json\n"
                  "  triage-flakes.py --outcome all -n 50 results.json\n"
                  "  triage-flakes.py --json results.json | jq '.data[] | select(.outcome==\"flaky\")'\n"
                  "\n"
                  "EXIT CODES:\n"
                  "  0  parsed fine, no flaky/unexpected tests (clean suite)\n"
                  "  2  usage   3  file not found   4  malformed report\n"
                  "  10 flaky/unexpected tests present (the triage signal)\n"
              ),
          )
          p.add_argument(
              "report",
              nargs="?",
              default="results.json",
              help="path to Playwright JSON report (default: ./results.json)",
          )
          p.add_argument("--json", action="store_true", help="emit a JSON envelope instead of TSV")
          p.add_argument("-q", "--quiet", action="store_true",
                         help="suppress the stderr summary header (errors still print)")
          p.add_argument(
              "-n",
              "--limit",
              type=int,
              default=20,
              metavar="N",
              help="cap rows printed (default 20)",
          )
          p.add_argument(
              "--outcome",
              default="flaky,unexpected",
              help="which outcomes to include: flaky | unexpected | all (default flaky,unexpected)",
          )
          args = p.parse_args(argv)
      
          if args.limit < 0:
              err("ERROR: --limit must be >= 0")
              return EXIT_USAGE
      
          sel = args.outcome.strip().lower()
          if sel == "all":
              wanted = None  # all outcomes
          else:
              wanted = {x.strip() for x in sel.split(",") if x.strip()}
              unknown = wanted - {"flaky", "unexpected", "expected", "skipped"}
              if unknown:
                  err(f"ERROR: unknown outcome(s): {', '.join(sorted(unknown))} (use flaky|unexpected|all)")
                  return EXIT_USAGE
      
          path = Path(args.report).resolve()
          if not path.exists():
              err(f"ERROR: report not found: {path}")
              if args.json:
                  print(json.dumps({"error": {"code": "NOT_FOUND", "message": f"report not found: {path}"}}))
              return EXIT_NOT_FOUND
          if not path.is_file():
              err(f"ERROR: not a file: {path}")
              return EXIT_NOT_FOUND
      
          try:
              data = load_report(path)
          except FileNotFoundError as e:
              err(f"ERROR: cannot read report: {e}")
              return EXIT_NOT_FOUND
          except ValueError as e:
              err(f"ERROR: malformed report: {e}")
              if args.json:
                  print(json.dumps({"error": {"code": "VALIDATION", "message": str(e)}}))
              return EXIT_VALIDATION
      
          finds = []
          walk_suites(data.get("suites"), finds)
      
          # The domain signal is computed over ALL findings, regardless of the display
          # filter — a clean suite means zero flaky AND zero unexpected, full stop.
          signal_present = any(f["outcome"] in ("flaky", "unexpected") for f in finds)
      
          if wanted is None:
              shown = list(finds)
          else:
              shown = [f for f in finds if f["outcome"] in wanted]
      
          # Rank: flaky before unexpected (OUTCOME_RANK), then retries desc, duration desc.
          shown.sort(
              key=lambda f: (
                  OUTCOME_RANK.get(f["outcome"], 99),
                  -f["retries"],
                  -f["durationMs"],
              )
          )
      
          capped = shown[: args.limit] if args.limit else shown
      
          total = len(finds)
          flaky_n = sum(1 for f in finds if f["outcome"] == "flaky")
          unexp_n = sum(1 for f in finds if f["outcome"] == "unexpected")
          if not args.quiet:
              err(TERM.hdr(f"Flake triage: {path.name}"))
              flaky_txt = TERM.c("orange", f"{flaky_n} flaky") if flaky_n else "0 flaky"
              unexp_txt = TERM.c("red", f"{unexp_n} unexpected") if unexp_n else "0 unexpected"
              err(f"  {total} tests | {flaky_txt} | {unexp_txt} | showing {len(capped)} of {len(shown)}")
      
          if args.json:
              envelope = {
                  "data": capped,
                  "meta": {
                      "count": len(capped),
                      "total_matched": len(shown),
                      "flaky": flaky_n,
                      "unexpected": unexp_n,
                      "schema": SCHEMA,
                  },
              }
              print(json.dumps(envelope, indent=2))
          else:
              print("outcome\tretries\tstatuses\tduration_ms\tlocation\ttitle")
              for f in capped:
                  print(
                      f"{f['outcome']}\t{f['retries']}\t{'->'.join(f['statuses'])}\t"
                      f"{f['durationMs']}\t{f['location']}\t{f['title']}"
                  )
      
          return EXIT_FINDINGS if signal_present else EXIT_OK
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
  • SKILL.md 15.5 KB
    ---
    name: playwright-ops
    description: "Playwright end-to-end testing operations - selectors, fixtures, network mocking, auth, parallelism, CI, visual regression, flake hunting. Use for: playwright, e2e test, end-to-end testing, browser test, getByRole, page object, storageState, trace viewer, flaky test, test sharding, visual regression, toHaveScreenshot, playwright config, codegen."
    license: MIT
    allowed-tools: "Read Write Bash"
    metadata:
      author: claude-mods
      related-skills: testing-ops, ci-cd-ops
    ---
    
    # Playwright Operations
    
    End-to-end testing with Playwright Test (`@playwright/test`, TS/JS). A Python flavor
    (`pytest-playwright`) exists with the same browser API but pytest-style fixtures — patterns here
    translate directly; runner config does not.
    
    ## Quick Start
    
    ```bash
    npm init playwright@latest          # scaffold config + example test + GH Actions workflow
    npx playwright test                 # run all tests, all projects
    npx playwright test --project=chromium --grep "@smoke"
    npx playwright test --ui            # interactive UI mode (watch, time-travel)
    npx playwright codegen https://app.local   # record actions -> generated locators
    npx playwright show-report          # open last HTML report
    npx playwright show-trace trace.zip # inspect a trace
    ```
    
    ## Selector Strategy
    
    **Hierarchy — always prefer the highest tier that uniquely matches:**
    
    | Tier | Locator | When |
    |------|---------|------|
    | 1 | `page.getByRole('button', { name: 'Submit' })` | Anything with an ARIA role — buttons, links, headings, textboxes. Tests a11y for free |
    | 2 | `page.getByLabel('Password')` | Form fields with labels |
    | 3 | `page.getByPlaceholder('name@example.com')` | Inputs without labels (fix the label instead, when you can) |
    | 4 | `page.getByText('Welcome back')` | Non-interactive text content |
    | 5 | `page.getByTestId('cart-total')` | Stable hook when semantics don't disambiguate. Configure attribute via `testIdAttribute` |
    | 6 | `page.locator('css=...')` / `xpath=` | **Last resort.** Coupled to DOM structure; breaks on refactor |
    
    Why: tiers 1–4 locate the way a user perceives the page — resilient to markup changes, and
    `getByRole` fails loudly when accessibility regresses. CSS/XPath encode implementation detail.
    
    **Narrowing without CSS:**
    
    ```ts
    page.getByRole('listitem')
        .filter({ hasText: 'Product 2' })
        .getByRole('button', { name: 'Add to cart' });
    
    page.getByRole('row').filter({ has: page.getByRole('cell', { name: 'Alice' }) });
    ```
    
    ### Web-First Assertions (no manual waits, ever)
    
    ```ts
    // BAD — checks once, races the render; sleeps are flake factories
    expect(await page.getByText('welcome').isVisible()).toBe(true);
    await page.waitForTimeout(2000);
    
    // GOOD — auto-retries until pass or timeout
    await expect(page.getByText('welcome')).toBeVisible();
    await expect(page.getByRole('list')).toHaveCount(3);
    await expect(page).toHaveURL(/\/dashboard/);
    await expect.soft(page.getByTestId('status')).toHaveText('Active'); // don't stop test on failure
    ```
    
    Actions (`click`, `fill`) auto-wait for actionability (visible, stable, enabled). If you feel the
    need for `waitForTimeout`, you're missing an assertion or an `await expect(...)` on a state change.
    For async non-DOM conditions use `expect.poll(() => fn())` or `expect(async () => {...}).toPass()`.
    
    Lint guard: enable `@typescript-eslint/no-floating-promises` — a missing `await` on an assertion is
    the most common silent-pass bug.
    
    ## Config Skeleton
    
    Full production template with comments: [assets/playwright.config.template.ts](assets/playwright.config.template.ts)
    
    ```ts
    import { defineConfig, devices } from '@playwright/test';
    
    export default defineConfig({
      testDir: './tests',
      fullyParallel: true,
      forbidOnly: !!process.env.CI,
      retries: process.env.CI ? 2 : 0,
      workers: process.env.CI ? 1 : undefined,
      reporter: process.env.CI ? 'blob' : 'html',
      use: {
        baseURL: process.env.BASE_URL ?? 'http://localhost:3000',
        trace: 'on-first-retry',
        testIdAttribute: 'data-testid',
      },
      projects: [
        { name: 'setup', testMatch: /.*\.setup\.ts/ },
        {
          name: 'chromium',
          use: { ...devices['Desktop Chrome'], storageState: 'playwright/.auth/user.json' },
          dependencies: ['setup'],
        },
      ],
      webServer: {
        command: 'npm run dev',
        url: 'http://localhost:3000',
        reuseExistingServer: !process.env.CI,
      },
    });
    ```
    
    ## Fixtures Decision Tree
    
    ```
    What do I need to share/setup?
    │
    ├─ Per-test object (page object, seeded record)
    │  └─ test.extend() test-scoped fixture — setup, await use(x), teardown
    │
    ├─ Expensive, safe-to-share resource (DB pool, test account)
    │  └─ Worker-scoped: [fn, { scope: 'worker' }] — once per worker process
    │
    ├─ Side effect every test needs (log capture, network stub)
    │  └─ Automatic: [fn, { auto: true }] — runs without being referenced
    │
    ├─ Config-tunable value (locale, default item)
    │  └─ Option: ['default', { option: true }] — override in projects[].use
    │
    ├─ Fixtures from several modules
    │  └─ mergeTests(testA, testB)
    │
    └─ Auth state per test file/role
       └─ test.use({ storageState: 'playwright/.auth/admin.json' })
    ```
    
    **POM-as-fixture (modern recommendation)** — page objects are fine; *instantiating them by hand in
    every test* is not. Inject via fixture:
    
    ```ts
    // fixtures.ts
    import { test as base } from '@playwright/test';
    import { TodoPage } from './pages/todo-page';
    
    export const test = base.extend<{ todoPage: TodoPage }>({
      todoPage: async ({ page }, use) => {
        const todoPage = new TodoPage(page);
        await todoPage.goto();
        await use(todoPage);          // test body runs here
      },
    });
    export { expect } from '@playwright/test';
    ```
    
    Page objects should expose **locators and actions**, not assertions wrapped in try/catch, and never
    store element handles. Details: [references/fixtures-and-pom.md](references/fixtures-and-pom.md)
    
    ## Network & API
    
    ```
    Network need?
    │
    ├─ Stub a third-party API           → page.route('**/api/**', r => r.fulfill({ json }))
    ├─ Tweak a real response            → const res = await route.fetch(); route.fulfill({ response: res, json })
    ├─ Simulate failure / offline       → route.abort() / route.fulfill({ status: 500 })
    ├─ Many endpoints, real shapes      → HAR record + replay (page.routeFromHAR, update: true to record)
    ├─ Pure API test (no browser)       → request fixture / APIRequestContext
    ├─ Seed data fast, assert via UI    → hybrid: create via request, verify via page
    └─ WebSocket traffic                → page.routeWebSocket(url, ws => ws.onMessage(...))
    ```
    
    **Hybrid seed-via-API, assert-via-UI** — the single biggest speed win in most suites:
    
    ```ts
    test('shows new project', async ({ request, page }) => {
      const res = await request.post('/api/projects', { data: { name: 'Apollo' } });
      expect(res.ok()).toBeTruthy();
      await page.goto('/projects');
      await expect(page.getByRole('link', { name: 'Apollo' })).toBeVisible();
    });
    ```
    
    Rule of thumb: **mock third-party dependencies you don't own; exercise your own backend for real**
    (or mock it deliberately in a separate "frontend-isolated" project).
    Details: [references/network-and-api.md](references/network-and-api.md)
    
    ## Authentication
    
    Standard pattern — login once in a setup project, reuse `storageState` everywhere:
    
    ```ts
    // tests/auth.setup.ts
    import { test as setup, expect } from '@playwright/test';
    
    setup('authenticate', async ({ page }) => {
      await page.goto('/login');
      await page.getByLabel('Username').fill(process.env.E2E_USER!);
      await page.getByLabel('Password').fill(process.env.E2E_PASS!);
      await page.getByRole('button', { name: 'Sign in' }).click();
      await expect(page.getByTestId('user-menu')).toBeVisible();   // wait for auth to settle!
      await page.context().storageState({ path: 'playwright/.auth/user.json' });
    });
    ```
    
    | Pattern | Use when |
    |---------|----------|
    | One setup project + `storageState` in `use` | One shared account, tests don't mutate server-side user state |
    | Per-role files (`admin.json`, `user.json`) + `test.use({ storageState })` | Role-based behavior under test |
    | Worker-scoped account fixture (`testInfo.parallelIndex`) | Parallel tests mutate user state — one account per worker |
    | API login (`request.post` + `request.storageState`) | Login endpoint exists; 10x faster than UI login |
    
    Gotchas: add `playwright/.auth/` to `.gitignore`. `storageState` captures cookies +
    localStorage — **not sessionStorage** (persist that manually via `page.evaluate` + init script).
    Always assert a logged-in signal before saving state, or you save a half-logged-in race.
    
    ## Parallelism, Retries, Isolation
    
    | Knob | Setting | Notes |
    |------|---------|-------|
    | Workers | `workers: process.env.CI ? 1 : undefined` | Local: half the logical CPU cores. CI runners are small — shard machines instead of oversubscribing |
    | File-level parallel | `fullyParallel: true` | Also makes sharding split per-test, not per-file |
    | Sharding | `npx playwright test --shard=1/4` | One shard per CI machine; merge blob reports after |
    | Retries | `retries: process.env.CI ? 2 : 0` | Pair with `trace: 'on-first-retry'`; treat "flaky" status as a bug queue, not a fix |
    | Serial | `test.describe.configure({ mode: 'serial' })` | Smell — usually means hidden inter-test coupling |
    
    **Isolation discipline:** every test gets a fresh `context`/`page` (cookies, storage) — keep it
    that way. No test reads state written by another test; shared server-side state is reset via API in
    `beforeEach` or scoped per worker (`test.info().parallelIndex` in usernames/tenant IDs). A suite
    that only passes single-worker is broken, not "sensitive".
    
    **Flake diagnosis:** `trace: 'on-first-retry'` → `npx playwright show-trace` (DOM snapshots,
    network, console per action). Local: `npx playwright test --ui` or `PWDEBUG=1` / `page.pause()`.
    Repro: `--repeat-each=20 --workers=4`. Playbook: [references/flake-hunting.md](references/flake-hunting.md)
    
    **Triage a whole run without eyeballing the report** — generate the JSON reporter output, then
    rank the offenders with the bundled triage tool ([scripts/triage-flakes.py](scripts/triage-flakes.py)):
    
    ```bash
    npx playwright test --reporter=json > results.json   # or reporter: [['json', { outputFile: 'results.json' }]]
    scripts/triage-flakes.py results.json                # flaky tests first, then hard fails
    ```
    
    It emits a ranked TSV (or `--json` envelope, schema `claude-mods.playwright-ops.flake-triage/v1`):
    flaky tests (passed only on retry) first — ordered by retry count then duration — followed by
    `unexpected` hard failures, each with `file:line`, the status sequence (`failed->passed`), and total
    duration. **Exit 10 means flakes/fails were found** (the triage signal — go fix them); exit 0 means a
    clean suite. `--outcome all` includes the passing tests for context; `-n N` caps rows.
    
    ## CI (GitHub Actions)
    
    ```yaml
    - uses: actions/checkout@v5
    - uses: actions/setup-node@v5
      with: { node-version: lts/* }
    - run: npm ci
    - run: npx playwright install --with-deps chromium   # only browsers you test
    - run: npx playwright test
    - uses: actions/upload-artifact@v4
      if: ${{ !cancelled() }}
      with: { name: playwright-report, path: playwright-report/, retention-days: 30 }
    ```
    
    | Decision | Guidance |
    |----------|----------|
    | Container vs install-deps | `mcr.microsoft.com/playwright:vX.Y.Z-jammy` image pins browser+OS (best for visual tests); `install --with-deps` is simpler and fine otherwise. **Pin image tag to your `@playwright/test` version** |
    | Browser caching | Cache `~/.cache/ms-playwright` keyed on Playwright version; skip when using the container |
    | Sharded reports | `reporter: 'blob'` on shards → upload `blob-report/` → merge job: `npx playwright merge-reports --reporter html ./all-blob-reports` |
    | Fail-fast vs full suite | PRs: `fail-fast: false` + `--max-failures=10` per shard — see *all* failures in one round-trip. Smoke gates: fail fast |
    
    Full workflows (sharding matrix, merge job, caching): [references/ci-patterns.md](references/ci-patterns.md)
    
    ## Visual Testing
    
    ```ts
    await expect(page).toHaveScreenshot('landing.png', {
      maxDiffPixels: 100,                       // or maxDiffPixelRatio / threshold
      mask: [page.getByTestId('ad-banner')],    // black-box dynamic regions
      fullPage: true,
    });
    ```
    
    - First run generates the baseline (test fails); update with `npx playwright test --update-snapshots`
    - Snapshots are named per browser **and platform** (`landing-chromium-darwin.png`) — baselines
      generated on macOS will not match Linux CI. Fix: generate baselines inside the same Docker image
      CI uses, or run visual tests only in the container
    - Disable animations: `toHaveScreenshot` defaults `animations: 'disabled'`; hide dynamic bits with
      `mask` or `stylePath` (CSS applied at capture time)
    - Global defaults: `expect: { toHaveScreenshot: { maxDiffPixels: 100 } }` in config
    - `toMatchSnapshot()` for non-image data (text/buffers)
    
    ## Component Testing & When to Prefer Cypress
    
    `@playwright/experimental-ct-react` (also vue/svelte) mounts components in a real browser —
    **still experimental**; for component-level work, Vitest browser mode or Testing Library are the
    safer default, with Playwright covering E2E.
    
    | Factor | Playwright | Cypress |
    |--------|-----------|---------|
    | Browsers | Chromium, Firefox, WebKit (real Safari engine) | Chrome-family, Firefox; WebKit experimental |
    | Parallelism | Free, built-in, shardable | Paid Cloud for parallel orchestration |
    | Multi-tab / multi-origin / iframes | Native | Historically constrained |
    | API testing | Built-in `request` context | Via `cy.request`, less ergonomic |
    | Component testing | Experimental | Mature, first-class |
    | In-browser interactive DX | UI mode (excellent) | The original benchmark; some teams still prefer it |
    
    Reach for Cypress when component testing maturity or an existing Cypress investment dominates;
    otherwise Playwright is the default for new E2E suites. (Repo also has a sibling `cypress-ops` skill.)
    
    ## Debugging & Codegen
    
    | Tool | Command | Use |
    |------|---------|-----|
    | UI mode | `npx playwright test --ui` | Watch mode, time-travel, pick locators |
    | Inspector | `PWDEBUG=1 npx playwright test` or `page.pause()` | Step through actions live |
    | Codegen | `npx playwright codegen <url>` | Records actions, emits role-based locators — treat output as a draft, refactor into POMs/fixtures |
    | Trace viewer | `npx playwright show-trace trace.zip` | Post-mortem: snapshots, network, console |
    | Headed + slow | `--headed --debug` | Eyeball a single test |
    | VS Code extension | — | Run/debug tests, pick locators in-editor |
    
    An official Playwright MCP server (`@playwright/mcp`) also exists for agent-driven browser
    automation — distinct from the test runner; don't conflate browsing automation with the test suite.
    
    ## References
    
    | File | Contents |
    |------|----------|
    | [references/fixtures-and-pom.md](references/fixtures-and-pom.md) | Fixture scopes/options/merging, POM-as-fixture architecture, anti-patterns |
    | [references/network-and-api.md](references/network-and-api.md) | route/fulfill/abort, HAR replay, API testing, hybrid seeding, WebSocket |
    | [references/ci-patterns.md](references/ci-patterns.md) | Full GH Actions workflows: basic, sharded+merge, container, caching, reporters |
    | [references/flake-hunting.md](references/flake-hunting.md) | Systematic flake diagnosis: traces, repro loops, common causes + fixes |
    | [scripts/triage-flakes.py](scripts/triage-flakes.py) | Parse a Playwright JSON report and rank flaky/failing tests (exit 10 = findings); see Flake diagnosis above |
    | [assets/playwright.config.template.ts](assets/playwright.config.template.ts) | Commented production config template |
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related