playwright-ops
Playwright end-to-end testing operations - selectors, fixtures, network mocking, auth, parallelism, CI, visual regression, flake hunting. Use for: playwright, e2e test, end-to-end testing, browser test, getByRole, page object, storageState, trace viewer, flaky test, test sharding
Install
npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/playwright-ops
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
git clone https://github.com/0xDarkMatter/claude-mods.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Playwright Operations
End-to-end testing with Playwright Test (@playwright/test, TS/JS). A Python flavor
(pytest-playwright) exists with the same browser API but pytest-style fixtures — patterns here
translate directly; runner config does not.
Quick Start
npm init playwright@latest # scaffold config + example test + GH Actions workflow
npx playwright test # run all tests, all projects
npx playwright test --project=chromium --grep "@smoke"
npx playwright test --ui # interactive UI mode (watch, time-travel)
npx playwright codegen https://app.local # record actions -> generated locators
npx playwright show-report # open last HTML report
npx playwright show-trace trace.zip # inspect a trace
Selector Strategy
Hierarchy — always prefer the highest tier that uniquely matches:
| Tier | Locator | When |
|---|---|---|
| 1 | page.getByRole('button', { name: 'Submit' }) |
Anything with an ARIA role — buttons, links, headings, textboxes. Tests a11y for free |
| 2 | page.getByLabel('Password') |
Form fields with labels |
| 3 | page.getByPlaceholder('name@example.com') |
Inputs without labels (fix the label instead, when you can) |
| 4 | page.getByText('Welcome back') |
Non-interactive text content |
| 5 | page.getByTestId('cart-total') |
Stable hook when semantics don't disambiguate. Configure attribute via testIdAttribute |
| 6 | page.locator('css=...') / xpath= |
Last resort. Coupled to DOM structure; breaks on refactor |
Why: tiers 1–4 locate the way a user perceives the page — resilient to markup changes, and
getByRole fails loudly when accessibility regresses. CSS/XPath encode implementation detail.
Narrowing without CSS:
page.getByRole('listitem')
.filter({ hasText: 'Product 2' })
.getByRole('button', { name: 'Add to cart' });
page.getByRole('row').filter({ has: page.getByRole('cell', { name: 'Alice' }) });
Web-First Assertions (no manual waits, ever)
// BAD — checks once, races the render; sleeps are flake factories
expect(await page.getByText('welcome').isVisible()).toBe(true);
await page.waitForTimeout(2000);
// GOOD — auto-retries until pass or timeout
await expect(page.getByText('welcome')).toBeVisible();
await expect(page.getByRole('list')).toHaveCount(3);
await expect(page).toHaveURL(/\/dashboard/);
await expect.soft(page.getByTestId('status')).toHaveText('Active'); // don't stop test on failure
Actions (click, fill) auto-wait for actionability (visible, stable, enabled). If you feel the
need for waitForTimeout, you're missing an assertion or an await expect(...) on a state change.
For async non-DOM conditions use expect.poll(() => fn()) or expect(async () => {...}).toPass().
Lint guard: enable @typescript-eslint/no-floating-promises — a missing await on an assertion is
the most common silent-pass bug.
Config Skeleton
Full production template with comments: assets/playwright.config.template.ts
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
testDir: './tests',
fullyParallel: true,
forbidOnly: !!process.env.CI,
retries: process.env.CI ? 2 : 0,
workers: process.env.CI ? 1 : undefined,
reporter: process.env.CI ? 'blob' : 'html',
use: {
baseURL: process.env.BASE_URL ?? 'http://localhost:3000',
trace: 'on-first-retry',
testIdAttribute: 'data-testid',
},
projects: [
{ name: 'setup', testMatch: /.*\.setup\.ts/ },
{
name: 'chromium',
use: { ...devices['Desktop Chrome'], storageState: 'playwright/.auth/user.json' },
dependencies: ['setup'],
},
],
webServer: {
command: 'npm run dev',
url: 'http://localhost:3000',
reuseExistingServer: !process.env.CI,
},
});
Fixtures Decision Tree
What do I need to share/setup?
│
├─ Per-test object (page object, seeded record)
│ └─ test.extend() test-scoped fixture — setup, await use(x), teardown
│
├─ Expensive, safe-to-share resource (DB pool, test account)
│ └─ Worker-scoped: [fn, { scope: 'worker' }] — once per worker process
│
├─ Side effect every test needs (log capture, network stub)
│ └─ Automatic: [fn, { auto: true }] — runs without being referenced
│
├─ Config-tunable value (locale, default item)
│ └─ Option: ['default', { option: true }] — override in projects[].use
│
├─ Fixtures from several modules
│ └─ mergeTests(testA, testB)
│
└─ Auth state per test file/role
└─ test.use({ storageState: 'playwright/.auth/admin.json' })
POM-as-fixture (modern recommendation) — page objects are fine; instantiating them by hand in every test is not. Inject via fixture:
// fixtures.ts
import { test as base } from '@playwright/test';
import { TodoPage } from './pages/todo-page';
export const test = base.extend<{ todoPage: TodoPage }>({
todoPage: async ({ page }, use) => {
const todoPage = new TodoPage(page);
await todoPage.goto();
await use(todoPage); // test body runs here
},
});
export { expect } from '@playwright/test';
Page objects should expose locators and actions, not assertions wrapped in try/catch, and never store element handles. Details: references/fixtures-and-pom.md
Network & API
Network need?
│
├─ Stub a third-party API → page.route('**/api/**', r => r.fulfill({ json }))
├─ Tweak a real response → const res = await route.fetch(); route.fulfill({ response: res, json })
├─ Simulate failure / offline → route.abort() / route.fulfill({ status: 500 })
├─ Many endpoints, real shapes → HAR record + replay (page.routeFromHAR, update: true to record)
├─ Pure API test (no browser) → request fixture / APIRequestContext
├─ Seed data fast, assert via UI → hybrid: create via request, verify via page
└─ WebSocket traffic → page.routeWebSocket(url, ws => ws.onMessage(...))
Hybrid seed-via-API, assert-via-UI — the single biggest speed win in most suites:
test('shows new project', async ({ request, page }) => {
const res = await request.post('/api/projects', { data: { name: 'Apollo' } });
expect(res.ok()).toBeTruthy();
await page.goto('/projects');
await expect(page.getByRole('link', { name: 'Apollo' })).toBeVisible();
});
Rule of thumb: mock third-party dependencies you don't own; exercise your own backend for real (or mock it deliberately in a separate "frontend-isolated" project). Details: references/network-and-api.md
Authentication
Standard pattern — login once in a setup project, reuse storageState everywhere:
// tests/auth.setup.ts
import { test as setup, expect } from '@playwright/test';
setup('authenticate', async ({ page }) => {
await page.goto('/login');
await page.getByLabel('Username').fill(process.env.E2E_USER!);
await page.getByLabel('Password').fill(process.env.E2E_PASS!);
await page.getByRole('button', { name: 'Sign in' }).click();
await expect(page.getByTestId('user-menu')).toBeVisible(); // wait for auth to settle!
await page.context().storageState({ path: 'playwright/.auth/user.json' });
});
| Pattern | Use when |
|---|---|
One setup project + storageState in use |
One shared account, tests don't mutate server-side user state |
Per-role files (admin.json, user.json) + test.use({ storageState }) |
Role-based behavior under test |
Worker-scoped account fixture (testInfo.parallelIndex) |
Parallel tests mutate user state — one account per worker |
API login (request.post + request.storageState) |
Login endpoint exists; 10x faster than UI login |
Gotchas: add playwright/.auth/ to .gitignore. storageState captures cookies +
localStorage — not sessionStorage (persist that manually via page.evaluate + init script).
Always assert a logged-in signal before saving state, or you save a half-logged-in race.
Parallelism, Retries, Isolation
| Knob | Setting | Notes |
|---|---|---|
| Workers | workers: process.env.CI ? 1 : undefined |
Local: half the logical CPU cores. CI runners are small — shard machines instead of oversubscribing |
| File-level parallel | fullyParallel: true |
Also makes sharding split per-test, not per-file |
| Sharding | npx playwright test --shard=1/4 |
One shard per CI machine; merge blob reports after |
| Retries | retries: process.env.CI ? 2 : 0 |
Pair with trace: 'on-first-retry'; treat "flaky" status as a bug queue, not a fix |
| Serial | test.describe.configure({ mode: 'serial' }) |
Smell — usually means hidden inter-test coupling |
Isolation discipline: every test gets a fresh context/page (cookies, storage) — keep it
that way. No test reads state written by another test; shared server-side state is reset via API in
beforeEach or scoped per worker (test.info().parallelIndex in usernames/tenant IDs). A suite
that only passes single-worker is broken, not "sensitive".
Flake diagnosis: trace: 'on-first-retry' → npx playwright show-trace (DOM snapshots,
network, console per action). Local: npx playwright test --ui or PWDEBUG=1 / page.pause().
Repro: --repeat-each=20 --workers=4. Playbook: references/flake-hunting.md
Triage a whole run without eyeballing the report — generate the JSON reporter output, then rank the offenders with the bundled triage tool (scripts/triage-flakes.py):
npx playwright test --reporter=json > results.json # or reporter: [['json', { outputFile: 'results.json' }]]
scripts/triage-flakes.py results.json # flaky tests first, then hard fails
It emits a ranked TSV (or --json envelope, schema claude-mods.playwright-ops.flake-triage/v1):
flaky tests (passed only on retry) first — ordered by retry count then duration — followed by
unexpected hard failures, each with file:line, the status sequence (failed->passed), and total
duration. Exit 10 means flakes/fails were found (the triage signal — go fix them); exit 0 means a
clean suite. --outcome all includes the passing tests for context; -n N caps rows.
CI (GitHub Actions)
- uses: actions/checkout@v5
- uses: actions/setup-node@v5
with: { node-version: lts/* }
- run: npm ci
- run: npx playwright install --with-deps chromium # only browsers you test
- run: npx playwright test
- uses: actions/upload-artifact@v4
if: ${{ !cancelled() }}
with: { name: playwright-report, path: playwright-report/, retention-days: 30 }
| Decision | Guidance |
|---|---|
| Container vs install-deps | mcr.microsoft.com/playwright:vX.Y.Z-jammy image pins browser+OS (best for visual tests); install --with-deps is simpler and fine otherwise. Pin image tag to your @playwright/test version |
| Browser caching | Cache ~/.cache/ms-playwright keyed on Playwright version; skip when using the container |
| Sharded reports | reporter: 'blob' on shards → upload blob-report/ → merge job: npx playwright merge-reports --reporter html ./all-blob-reports |
| Fail-fast vs full suite | PRs: fail-fast: false + --max-failures=10 per shard — see all failures in one round-trip. Smoke gates: fail fast |
Full workflows (sharding matrix, merge job, caching): references/ci-patterns.md
Visual Testing
await expect(page).toHaveScreenshot('landing.png', {
maxDiffPixels: 100, // or maxDiffPixelRatio / threshold
mask: [page.getByTestId('ad-banner')], // black-box dynamic regions
fullPage: true,
});
- First run generates the baseline (test fails); update with
npx playwright test --update-snapshots - Snapshots are named per browser and platform (
landing-chromium-darwin.png) — baselines generated on macOS will not match Linux CI. Fix: generate baselines inside the same Docker image CI uses, or run visual tests only in the container - Disable animations:
toHaveScreenshotdefaultsanimations: 'disabled'; hide dynamic bits withmaskorstylePath(CSS applied at capture time) - Global defaults:
expect: { toHaveScreenshot: { maxDiffPixels: 100 } }in config toMatchSnapshot()for non-image data (text/buffers)
Component Testing & When to Prefer Cypress
@playwright/experimental-ct-react (also vue/svelte) mounts components in a real browser —
still experimental; for component-level work, Vitest browser mode or Testing Library are the
safer default, with Playwright covering E2E.
| Factor | Playwright | Cypress |
|---|---|---|
| Browsers | Chromium, Firefox, WebKit (real Safari engine) | Chrome-family, Firefox; WebKit experimental |
| Parallelism | Free, built-in, shardable | Paid Cloud for parallel orchestration |
| Multi-tab / multi-origin / iframes | Native | Historically constrained |
| API testing | Built-in request context |
Via cy.request, less ergonomic |
| Component testing | Experimental | Mature, first-class |
| In-browser interactive DX | UI mode (excellent) | The original benchmark; some teams still prefer it |
Reach for Cypress when component testing maturity or an existing Cypress investment dominates;
otherwise Playwright is the default for new E2E suites. (Repo also has a sibling cypress-ops skill.)
Debugging & Codegen
| Tool | Command | Use |
|---|---|---|
| UI mode | npx playwright test --ui |
Watch mode, time-travel, pick locators |
| Inspector | PWDEBUG=1 npx playwright test or page.pause() |
Step through actions live |
| Codegen | npx playwright codegen <url> |
Records actions, emits role-based locators — treat output as a draft, refactor into POMs/fixtures |
| Trace viewer | npx playwright show-trace trace.zip |
Post-mortem: snapshots, network, console |
| Headed + slow | --headed --debug |
Eyeball a single test |
| VS Code extension | — | Run/debug tests, pick locators in-editor |
An official Playwright MCP server (@playwright/mcp) also exists for agent-driven browser
automation — distinct from the test runner; don't conflate browsing automation with the test suite.
References
| File | Contents |
|---|---|
| references/fixtures-and-pom.md | Fixture scopes/options/merging, POM-as-fixture architecture, anti-patterns |
| references/network-and-api.md | route/fulfill/abort, HAR replay, API testing, hybrid seeding, WebSocket |
| references/ci-patterns.md | Full GH Actions workflows: basic, sharded+merge, container, caching, reporters |
| references/flake-hunting.md | Systematic flake diagnosis: traces, repro loops, common causes + fixes |
| scripts/triage-flakes.py | Parse a Playwright JSON report and rank flaky/failing tests (exit 10 = findings); see Flake diagnosis above |
| assets/playwright.config.template.ts | Commented production config template |
Files (claude-mods)
-
assets
-
playwright.config.template.ts 4.6 KB
/** * Production Playwright config template. * * Copy to playwright.config.ts and adjust the marked sections. * Conventions baked in: * - blob reporter on CI (shard-mergeable), html locally * - trace on first retry (flake forensics at near-zero cost) * - auth via a `setup` project + storageState (login once, reuse everywhere) * - webServer boots the app and waits for it — no sleeps in CI scripts */ import { defineConfig, devices } from '@playwright/test'; export default defineConfig({ testDir: './tests', // Run tests within files in parallel too (also gives per-test shard balancing). // Set false only if tests within a file are intentionally ordered. fullyParallel: true, // A stray `test.only` committed to CI silently skips the suite — make it a build failure. forbidOnly: !!process.env.CI, // Retries are flake telemetry, not a fix: retried-then-passed tests show as // "flaky" in the report. Keep 0 locally so you feel flakes immediately. retries: process.env.CI ? 2 : 0, // Small CI runners (2-core GitHub hosted) thrash with parallel browser workers. // Scale horizontally with --shard instead. Locally, default = ~half the cores. workers: process.env.CI ? 1 : undefined, // blob -> uploaded per shard, merged with `npx playwright merge-reports`. reporter: process.env.CI ? [['blob'], ['github']] : [['html', { open: 'on-failure' }]], // Per-action timeout defaults are usually fine; raise the global test timeout // only for genuinely long flows (or per-test via test.slow()). timeout: 30_000, expect: { timeout: 5_000, // Global visual-comparison tolerances; override per assertion when needed. toHaveScreenshot: { maxDiffPixels: 100, // animations: 'disabled' is already the default for screenshots }, }, use: { // All page.goto('/relative') and request.get('/api/...') resolve against this. baseURL: process.env.BASE_URL ?? 'http://localhost:3000', // Trace on first retry: the failing run gets full DOM snapshots + network log. // Use 'retain-on-failure' instead if you run with retries: 0. trace: 'on-first-retry', screenshot: 'only-on-failure', // video is mostly redundant with traces; enable only if your team skims videos. // video: 'retain-on-failure', // Determinism: pin what the OS would otherwise decide for you. locale: 'en-US', timezoneId: 'UTC', // viewport comes from the device preset per project below. // Attribute used by page.getByTestId(); align with your frontend convention. testIdAttribute: 'data-testid', // Uncomment if a service worker swallows your route() mocks: // serviceWorkers: 'block', }, projects: [ // --- Auth setup: runs first, saves storage state for the browser projects --- // tests/auth.setup.ts logs in and calls // page.context().storageState({ path: 'playwright/.auth/user.json' }) // Keep playwright/.auth/ in .gitignore. { name: 'setup', testMatch: /.*\.setup\.ts/ }, { name: 'chromium', use: { ...devices['Desktop Chrome'], storageState: 'playwright/.auth/user.json', }, dependencies: ['setup'], }, // Enable per browser-support matrix. Remember: each enabled project must be // installed in CI (npx playwright install --with-deps firefox webkit). // { // name: 'firefox', // use: { ...devices['Desktop Firefox'], storageState: 'playwright/.auth/user.json' }, // dependencies: ['setup'], // }, // { // name: 'webkit', // use: { ...devices['Desktop Safari'], storageState: 'playwright/.auth/user.json' }, // dependencies: ['setup'], // }, // Mobile viewport smoke pass. // { // name: 'mobile-chrome', // use: { ...devices['Pixel 7'], storageState: 'playwright/.auth/user.json' }, // dependencies: ['setup'], // grep: /@smoke/, // }, // Unauthenticated flows (login page itself, public pages) — no storageState. // { // name: 'chromium-no-auth', // use: { ...devices['Desktop Chrome'] }, // testMatch: /.*\.public\.spec\.ts/, // }, ], // Playwright boots your app and polls `url` until it responds — replaces // "npm start & sleep 15" hacks in CI scripts. webServer: { command: process.env.CI ? 'npm run build && npm run start' : 'npm run dev', url: 'http://localhost:3000', reuseExistingServer: !process.env.CI, timeout: 120_000, stdout: 'pipe', // surface app logs in CI output when boot fails }, // Multiple servers? webServer also accepts an array: [{ api }, { frontend }]. });
-
-
references
-
ci-patterns.md 6.4 KB
# CI Patterns (GitHub Actions) Runnable workflows for Playwright in CI, from single-job to sharded fleets. Adapt paths/commands for other CI providers — the shape is identical. ## Baseline Workflow ```yaml # .github/workflows/playwright.yml name: Playwright Tests on: push: branches: [main] pull_request: branches: [main] jobs: test: timeout-minutes: 60 runs-on: ubuntu-latest steps: - uses: actions/checkout@v5 - uses: actions/setup-node@v5 with: node-version: lts/* - name: Install dependencies run: npm ci - name: Install Playwright browsers run: npx playwright install --with-deps chromium - name: Run Playwright tests run: npx playwright test - uses: actions/upload-artifact@v4 if: ${{ !cancelled() }} # upload report on failure too with: name: playwright-report path: playwright-report/ retention-days: 30 ``` Notes: - Install only the browsers your projects use (`chromium` above) — saves minutes per run. - `if: ${{ !cancelled() }}` keeps the report when tests fail; that's when you need it. - Secrets via `env:` on the test step (`E2E_USER: ${{ secrets.E2E_USER }}`), never committed. ## Container vs install-deps | Approach | Pros | Cons | |----------|------|------| | `npx playwright install --with-deps` on the runner | Simple; matches local dev | OS-level rendering drifts with runner image updates — visual baselines can churn | | `container: mcr.microsoft.com/playwright:v1.52.0-jammy` | Pinned browser + OS rendering; reproducible visual tests; no install step | Slightly slower job start; must bump tag with `@playwright/test` | **Always pin the container tag to your exact `@playwright/test` version** — a mismatch produces "Executable doesn't exist" or subtle behavior skew. ```yaml jobs: test: runs-on: ubuntu-latest container: image: mcr.microsoft.com/playwright:v1.52.0-jammy steps: - uses: actions/checkout@v5 - run: npm ci - run: npx playwright test env: HOME: /root # workaround for firefox in containers ``` ## Caching Browsers (non-container path) ```yaml - name: Get Playwright version id: pw-version run: echo "version=$(node -p "require('@playwright/test/package.json').version")" >> "$GITHUB_OUTPUT" - uses: actions/cache@v4 id: pw-cache with: path: ~/.cache/ms-playwright key: playwright-${{ runner.os }}-${{ steps.pw-version.outputs.version }} - run: npx playwright install --with-deps chromium if: steps.pw-cache.outputs.cache-hit != 'true' - run: npx playwright install-deps chromium # OS deps aren't cached if: steps.pw-cache.outputs.cache-hit == 'true' ``` ## Sharding with Blob Reports + Merge Config side — blob on CI shards, html locally: ```ts // playwright.config.ts reporter: process.env.CI ? 'blob' : 'html', fullyParallel: true, // shards split per-test instead of per-file -> better balance ``` ```yaml jobs: playwright-tests: runs-on: ubuntu-latest strategy: fail-fast: false # let every shard finish; see ALL failures matrix: shardIndex: [1, 2, 3, 4] shardTotal: [4] steps: - uses: actions/checkout@v5 - uses: actions/setup-node@v5 with: { node-version: lts/* } - run: npm ci - run: npx playwright install --with-deps chromium - run: npx playwright test --shard=${{ matrix.shardIndex }}/${{ matrix.shardTotal }} - uses: actions/upload-artifact@v4 if: ${{ !cancelled() }} with: name: blob-report-${{ matrix.shardIndex }} path: blob-report retention-days: 1 merge-reports: if: ${{ !cancelled() }} needs: [playwright-tests] runs-on: ubuntu-latest steps: - uses: actions/checkout@v5 - uses: actions/setup-node@v5 with: { node-version: lts/* } - run: npm ci - uses: actions/download-artifact@v4 with: path: all-blob-reports pattern: blob-report-* merge-multiple: true - run: npx playwright merge-reports --reporter html ./all-blob-reports - uses: actions/upload-artifact@v4 with: name: html-report--attempt-${{ github.run_attempt }} path: playwright-report retention-days: 14 ``` `merge-reports` accepts multiple reporters: `--reporter html,github` annotates the PR while also producing the browsable report. ## Fail-Fast vs Full-Suite | Context | Strategy | |---------|----------| | PR validation | `fail-fast: false` on the matrix + `maxFailures: 10` (or `--max-failures`) per shard. Developers fix everything in one round-trip instead of whack-a-mole | | Smoke gate before deploy | Fail fast — `--grep @smoke`, no retries, abort pipeline on first failure | | Nightly full regression | Full suite, retries on, no fail-fast; route the merged report to the team channel | ## Reporters | Reporter | Use | |----------|-----| | `html` | Local + merged CI artifact — the daily driver | | `blob` | Shard intermediate; only input for `merge-reports` | | `junit` | Test-management ingestion (Jenkins, Azure DevOps, TestRail): `['junit', { outputFile: 'results.xml' }]` | | `github` | Inline PR annotations on failures | | `list` / `dot` / `line` | Console verbosity choices | Multiple at once: ```ts reporter: process.env.CI ? [['blob'], ['github']] : [['html', { open: 'on-failure' }]], ``` ## webServer in CI ```ts webServer: { command: 'npm run build && npm run start', url: 'http://localhost:3000', reuseExistingServer: !process.env.CI, // CI always boots fresh timeout: 120_000, stdout: 'pipe', // surface server logs in CI output }, ``` Playwright waits for `url` to respond before running tests — no `sleep 10` hacks. Multiple servers (API + frontend) can be given as an array. ## CI Hardening Checklist - [ ] `forbidOnly: !!process.env.CI` — a stray `test.only` fails the build instead of silently skipping the suite - [ ] `retries: 2` on CI + `trace: 'on-first-retry'` - [ ] `workers: 1` per shard on small runners (2-core GitHub runners thrash above that); scale via shards - [ ] Report artifacts uploaded with `if: ${{ !cancelled() }}` - [ ] Browser install scoped to actual projects - [ ] Container tag or browser cache keyed to the Playwright version - [ ] Visual-test baselines generated in the same environment CI runs (see SKILL.md Visual Testing) -
fixtures-and-pom.md 6.6 KB
# Fixtures and Page Object Architecture How to structure Playwright Test suites with fixtures as the composition mechanism and page objects as thin locator/action wrappers. ## Built-in Fixtures | Fixture | Type | Scope | Notes | |---------|------|-------|-------| | `page` | `Page` | test | Fresh isolated page per test | | `context` | `BrowserContext` | test | Fresh context per test — cookies/storage isolated | | `browser` | `Browser` | worker | Shared across tests in a worker | | `browserName` | `string` | worker | `'chromium' \| 'firefox' \| 'webkit'` | | `request` | `APIRequestContext` | test | HTTP client honoring `baseURL` / `extraHTTPHeaders` | ## Custom Fixtures: the Full Shape ```ts import { test as base } from '@playwright/test'; type TestFixtures = { todoPage: TodoPage; // test-scoped defaultItem: string; // option }; type WorkerFixtures = { account: { username: string; password: string }; // worker-scoped }; export const test = base.extend<TestFixtures, WorkerFixtures>({ // Option — overridable per project via projects[].use defaultItem: ['Something nice', { option: true }], // Test-scoped fixture with setup + teardown todoPage: async ({ page, defaultItem }, use) => { const todoPage = new TodoPage(page); await todoPage.goto(); await todoPage.addToDo(defaultItem); await use(todoPage); // <-- test body executes here await todoPage.removeAll(); // teardown runs even if test fails }, // Worker-scoped — once per worker process; second generic param account: [async ({ browser }, use, workerInfo) => { const username = 'user-' + workerInfo.workerIndex; const password = await createAccount(username); // expensive, do once await use({ username, password }); await deleteAccount(username); }, { scope: 'worker' }], }); export { expect } from '@playwright/test'; ``` Key mechanics: - **Lazy**: a fixture only runs if the test (or another fixture) references it. - **Composable**: fixtures depend on other fixtures by destructuring them. - **Teardown order**: reverse of setup, runs even on failure — replaces brittle `afterEach` chains. - The two generic params of `extend<TestFixtures, WorkerFixtures>` map to test scope and worker scope respectively. Worker fixtures cannot depend on test fixtures. ## Fixture Options Reference | Option | Effect | |--------|--------| | `{ scope: 'worker' }` | One instance per worker process | | `{ auto: true }` | Runs for every test without being referenced — global hooks | | `{ option: true }` | Value is a project-configurable option | | `{ timeout: 60_000 }` | Separate timeout for slow fixture setup | | `{ box: true }` | Hide fixture from report/errors (or `box: 'self'` to hide just its step) | | `{ title: 'my fixture' }` | Custom name in reports | ### Automatic fixtures as global hooks ```ts export const test = base.extend<{ forEachTest: void }, { forEachWorker: void }>({ // beforeEach/afterEach equivalent, but reusable across files forEachTest: [async ({ page }, use) => { await page.goto('/'); // before each test await use(); // after each test }, { auto: true }], // once per worker forEachWorker: [async ({}, use) => { console.log(`Worker ${test.info().workerIndex} starting`); await use(); }, { scope: 'worker', auto: true }], }); ``` ### Overriding built-ins ```ts export const test = base.extend({ page: async ({ page }, use) => { await page.goto('/dashboard'); // every test starts on dashboard await use(page); }, // Override storageState to come from a worker fixture (per-worker auth) storageState: ({ workerStorageState }, use) => use(workerStorageState), }); ``` ## mergeTests: Composing Fixture Modules Keep fixture concerns in separate modules and merge at the edge: ```ts // fixtures/db.ts -> export const test = base.extend<{ db: Db }>({...}) // fixtures/a11y.ts -> export const test = base.extend<{ axe: Axe }>({...}) // fixtures/index.ts import { mergeTests, mergeExpects } from '@playwright/test'; import { test as dbTest } from './db'; import { test as a11yTest } from './a11y'; export const test = mergeTests(dbTest, a11yTest); export { expect } from '@playwright/test'; ``` Tests import `test`/`expect` from your fixtures module, never from `@playwright/test` directly — one import path to rule them all. ## Page Objects: Modern Recommendation ### POM-as-fixture (preferred) The page object is a plain class; the **fixture** owns construction and navigation: ```ts // pages/checkout-page.ts import { type Page, type Locator, expect } from '@playwright/test'; export class CheckoutPage { readonly page: Page; readonly cardNumber: Locator; readonly payButton: Locator; constructor(page: Page) { this.page = page; this.cardNumber = page.getByLabel('Card number'); this.payButton = page.getByRole('button', { name: 'Pay now' }); } async goto() { await this.page.goto('/checkout'); } async pay(card: string) { await this.cardNumber.fill(card); await this.payButton.click(); } } ``` ```ts // a test test('pays with valid card', async ({ checkoutPage, page }) => { await checkoutPage.pay('4242 4242 4242 4242'); await expect(page.getByRole('heading', { name: 'Order confirmed' })).toBeVisible(); }); ``` ### Rules for healthy page objects | Rule | Why | |------|-----| | Store `Locator`s, never element handles | Locators are lazy + auto-retrying; handles go stale | | Expose actions + locators; keep assertions in tests (or custom `expect` matchers) | Tests stay readable as specs; POMs stay reusable | | No `waitForTimeout` / try-catch flow control inside POMs | Hides flake; actions already auto-wait | | Constructor takes `Page` (or a `Locator` root for component objects) only | Keeps them trivially fixture-injectable | | Prefer small per-screen objects over one God object | Cheap to compose via fixtures | ### When to skip POMs entirely Small suites (< ~20 tests) over stable UIs often read better with raw `getByRole` calls inline. POMs earn their keep when the same screen appears in many tests or locators churn. Don't build the abstraction before the duplication exists. ## Hooks vs Fixtures | Need | Use | |------|-----| | Shared setup local to one file | `test.beforeEach` is fine | | Shared setup across files | Fixture (auto or named) | | Expensive once-per-run setup | Project dependencies (setup project) — not `globalSetup`, which skips fixtures/tracing | | Once-per-worker setup | Worker-scoped fixture | `test.beforeAll` runs **once per worker**, not once per run — a classic surprise. For true once-per-run work, use a setup project with `dependencies`. -
flake-hunting.md 5.8 KB
# Flake Hunting Systematic diagnosis of flaky Playwright tests. A flaky test is a bug — in the test, the app, or the environment. Retries buy time to fix it; they are not the fix. ## Triage Workflow ``` Flaky test reported │ 1. Get the evidence │ └─ trace: 'on-first-retry' in config → download trace from CI artifact │ └─ npx playwright show-trace path/to/trace.zip │ (per-action DOM snapshots, network, console, timing) │ 2. Reproduce locally │ └─ npx playwright test failing.spec.ts --repeat-each=20 --workers=4 │ ├─ Fails alone, repeated → timing/race within the test │ ├─ Fails only with --workers>1 → cross-test state leakage │ └─ Fails only in CI → environment delta (speed, viewport, headless, locale, TZ) │ 3. Classify against the table below, fix the CAUSE │ 4. Prove the fix └─ --repeat-each=50 clean, then watch the "flaky" count in CI reports trend to zero ``` CI flake visibility: HTML report marks retried-then-passed tests as **flaky** — review that list weekly; it's your queue. ## Common Causes and Fixes | Symptom | Root cause | Fix | |---------|-----------|-----| | Click "worked" but nothing happened | Element re-rendered between locate and click (hydration, list re-sort) | Assert the settled state first: `await expect(row).toBeVisible()` then act; prefer role/text locators that target the final element | | Assertion passes locally, times out in CI | CI is slower; manual check raced the render | Replace any non-retrying check with web-first `await expect(...)`; raise `expect.timeout` only if the app is legitimately slow | | `waitForTimeout` sprinkled around | Sleeping instead of waiting for a condition | Delete; wait on the observable effect: `expect(locator)`, `page.waitForURL()`, `page.waitForResponse()` | | Fails only with multiple workers | Tests share an account/record; one mutates what another reads | Per-worker data: suffix usernames/tenants with `test.info().parallelIndex`; or worker-scoped account fixture | | First test after auth flaky | storageState saved before login finished | In auth.setup, assert a logged-in signal (`await expect(page.getByTestId('user-menu')).toBeVisible()`) before `storageState({ path })` | | Animation mid-flight in screenshots/clicks | CSS transitions | `toHaveScreenshot` disables animations by default; for actions, assert post-animation state or set `reducedMotion: 'reduce'` in `use` | | Time-dependent failures (midnight, month-end, TZ) | Real clock | `await page.clock.setFixedTime(new Date('2026-01-15T10:00:00'))`; pin `timezoneId` and `locale` in `use` | | Random data collisions | Shared fixtures with hardcoded names | Unique-per-test names: `` `proj-${test.info().testId}` `` | | Network nondeterminism from third parties | Live external calls | Mock them (`route.fulfill` / HAR replay) — see network-and-api.md | | Passes in `--headed`, fails headless | Viewport/focus/rendering differences | Pin `viewport` in config; debug headless with traces, not by switching to headed | | Fails only on retry / second run | Leftover server-side state from first attempt | Make setup idempotent (upsert, not create); clean up in fixture teardown, which runs on failure too | ## Tools Reference | Tool | Invocation | What it gives you | |------|-----------|-------------------| | Trace viewer | `trace: 'on-first-retry'` → `npx playwright show-trace trace.zip` | Time-travel DOM snapshots, network log, console, action timeline — the primary CI forensic tool | | UI mode | `npx playwright test --ui` | Watch mode + live trace while iterating on a fix | | Inspector | `PWDEBUG=1 npx playwright test foo.spec.ts` or `await page.pause()` | Step through actions, try locators live | | Repeat | `--repeat-each=20` | Statistical reproduction | | Stress | `--workers=4` (or more than usual) | Surfaces isolation bugs | | Single worker | `--workers=1` | If this "fixes" it, you have cross-test coupling — that's the bug | | Verbose API log | `DEBUG=pw:api npx playwright test` | Every Playwright call with timing | | Video | `video: 'retain-on-failure'` | Cheaper than trace to skim; less data | `trace: 'on'` everywhere is expensive — `'on-first-retry'` is the right default; use `'retain-on-failure'` if you run without retries. ## Retrying Non-DOM Conditions Web-first assertions only retry on locators/page. For everything else: ```ts // Poll an arbitrary async value await expect.poll(async () => { const res = await request.get(`/api/jobs/${id}`); return (await res.json()).status; }, { timeout: 30_000, intervals: [1_000] }).toBe('done'); // Retry a block of assertions/actions together await expect(async () => { const res = await request.get('/health'); expect(res.status()).toBe(200); }).toPass({ timeout: 60_000 }); ``` Use these for eventual consistency (queues, search indexing, emails) instead of sleep loops. ## Isolation Discipline Checklist - [ ] No test depends on another test having run (`test.describe.configure({ mode: 'serial' })` is a red flag, not a tool of first resort) - [ ] Server-side state is created per test (API seeding) or per worker (`parallelIndex`-scoped accounts) - [ ] Teardown lives in fixtures (runs on failure), not at the end of test bodies - [ ] `storageState` files saved only after asserting login completed - [ ] `forbidOnly` on CI; `--repeat-each` smoke before merging new specs - [ ] Suite passes with `--workers=8 --repeat-each=3` locally before you blame CI ## Quarantine Pattern While a flake is being fixed, tag it instead of deleting or `.skip`-ing silently: ```ts test('checkout under load @quarantine', async ({ page }) => { ... }); ``` ```bash npx playwright test --grep-invert @quarantine # main gate npx playwright test --grep @quarantine # nightly, non-blocking ``` Track quarantined tests with an issue each; a quarantine list that only grows is a suite dying in slow motion. -
network-and-api.md 5.9 KB
# Network Mocking and API Testing Intercepting browser traffic, replaying HAR recordings, testing APIs directly, and the hybrid seed-via-API / assert-via-UI pattern. ## Route Interception: page.route() ```ts // Stub an endpoint entirely await page.route('*/**/api/v1/fruits', async route => { await route.fulfill({ json: [{ name: 'Strawberry', id: 21 }] }); }); // Must be registered BEFORE the navigation/action that triggers the request await page.goto('/'); ``` | Method | Effect | |--------|--------| | `route.fulfill({ json, status, headers, body, path })` | Respond without hitting the network | | `route.fetch()` | Execute the real request, get the response for modification | | `route.continue({ headers, postData, url })` | Pass through, optionally modified | | `route.abort('failed')` | Simulate network failure | | `route.fallback()` | Defer to the next matching handler (handlers run last-registered-first) | ### Modify a real response ```ts await page.route('*/**/api/v1/fruits', async route => { const response = await route.fetch(); const json = await response.json(); json.push({ name: 'Loquat', id: 100 }); await route.fulfill({ response, json }); // real status/headers, patched body }); ``` ### Failure-mode tests ```ts await page.route('**/api/orders', route => route.fulfill({ status: 500 })); await page.route('**/*.{png,jpg,jpeg}', route => route.abort()); // block heavy assets await context.setOffline(true); // whole-context offline ``` ### Scope and ordering gotchas - `page.route` applies to that page; `context.route` to every page in the context (use in a fixture for suite-wide stubs). - Patterns: glob (`**/api/**`), RegExp, or predicate function. Glob matches the **full URL**. - `await page.unroute(pattern)` removes handlers; `page.unrouteAll()` clears them. - Service workers can bypass routing — set `serviceWorkers: 'block'` in `use` if your app registers one and mocks mysteriously don't fire. ## HAR Record and Replay Best for "many endpoints, realistic payloads" — record once against the real backend, replay hermetically. ```ts // Record (update: true hits the real network and refreshes the file) await page.routeFromHAR('./hars/fruits.har', { url: '*/**/api/v1/**', update: true, }); // Replay (default update: false serves from the file; unmatched requests are aborted) await page.routeFromHAR('./hars/fruits.har', { url: '*/**/api/v1/**' }); ``` CLI recording: ```bash npx playwright open --save-har=example.har --save-har-glob="**/api/**" https://example.com ``` Workflow: re-run recording tests with `update: true` whenever the API contract changes, commit the HAR + extracted bodies (`.txt`/`.json` sidecars are editable by hand for edge cases). ## API Testing: request / APIRequestContext The `request` fixture is an HTTP client honoring config `baseURL` and `extraHTTPHeaders` — no browser involved, so it's fast. ```ts // playwright.config.ts use: { baseURL: 'https://api.github.com', extraHTTPHeaders: { 'Accept': 'application/vnd.github.v3+json', 'Authorization': `token ${process.env.API_TOKEN}`, }, }, ``` ```ts test('creates a bug report', async ({ request }) => { const newIssue = await request.post(`/repos/${USER}/${REPO}/issues`, { data: { title: '[Bug] report 1', body: 'Bug description' }, }); expect(newIssue.ok()).toBeTruthy(); const issues = await request.get(`/repos/${USER}/${REPO}/issues`); expect(await issues.json()).toContainEqual( expect.objectContaining({ title: '[Bug] report 1' }), ); }); ``` Standalone context (different base URL, custom auth, use in setup scripts): ```ts import { request } from '@playwright/test'; const api = await request.newContext({ baseURL: 'https://api.example.com' }); await api.post('/seed', { data: {...} }); await api.dispose(); // always dispose manually created contexts ``` `request.post` options: `data` (JSON), `form` (urlencoded), `multipart` (file upload), `params` (query), `headers`, `failOnStatusCode`. ## Hybrid: Seed via API, Assert via UI UI-driven setup is the slowest, flakiest part of most suites. Replace it: ```ts test('renders the new project card', async ({ request, page }) => { // Arrange — fast, deterministic, server-side const res = await request.post('/api/projects', { data: { name: 'Apollo' } }); expect(res.ok()).toBeTruthy(); const { id } = await res.json(); // Act + Assert — the only part that needs a browser await page.goto(`/projects/${id}`); await expect(page.getByRole('heading', { name: 'Apollo' })).toBeVisible(); }); ``` Notes: - The `request` fixture shares `storageState` with the browser context, so an authenticated UI session usually authenticates API calls too. For a different principal, create a separate `request.newContext({ storageState: 'playwright/.auth/admin.json' })`. - Postcondition checks invert it: act in the UI, verify via `request.get` that the server really persisted the thing. - Cleanup belongs in fixtures (teardown after `use()`) or `afterAll` API calls — not in the test body where a failure skips it. ## What to Mock vs Exercise | Dependency | Default | |------------|---------| | Third-party SaaS (payments, analytics, maps) | **Mock** (`route.fulfill` / HAR). You can't control their data or uptime, and you don't want test purchases | | Your own backend | **Real** — that's the integration you're paying E2E tests to verify | | Your backend, in a frontend-only project | Mock deliberately and label the project (`name: 'ui-isolated'`) so coverage claims stay honest | | Time / randomness | `page.clock.install()` / `page.clock.setFixedTime(...)` for time-dependent UI | ## WebSocket Mocking ```ts await page.routeWebSocket('wss://example.com/ws', ws => { ws.onMessage(message => { if (message === 'request') ws.send('response'); }); }); ``` By default the intercepted socket never reaches the server; call `ws.connectToServer()` inside the handler to proxy with selective message rewriting.
-
-
scripts
-
.gitkeep 0 B · in bundle
-
triage-flakes.py 10.2 KB
#!/usr/bin/env python3 # Rank Playwright tests by flakiness from a JSON report so the agent triages, not eyeballs. # # Parses a Playwright JSON report (`--reporter=json`) and surfaces the tests # worth a human's attention: flaky tests (passed only on retry) first, then # hard "unexpected" failures. Flaky tests are ranked by retry count desc, then # total duration desc, because the most-retried, slowest test is the worst # offender in your queue. # # Usage: triage-flakes.py [OPTIONS] [REPORT] # Input: REPORT = path to a Playwright JSON report (positional, default ./results.json) # Output: stdout = ranked findings (TSV, or JSON envelope with --json) # Stderr: headers, summary, progress, errors # Exit: 0 parsed fine, no flaky/unexpected tests (clean suite) # 2 usage, 3 file not found, 4 malformed/not a Playwright report, # 10 DOMAIN SIGNAL: flaky/unexpected tests present (the thing being triaged) # # Examples: # npx playwright test --reporter=json > results.json # triage-flakes.py results.json # triage-flakes.py --outcome all -n 50 results.json # triage-flakes.py --json results.json | jq '.data[] | select(.outcome=="flaky")' import argparse import json import os import sys from pathlib import Path # Windows consoles default to cp1252; force UTF-8 so glyphs in framing don't raise # UnicodeEncodeError (the repo's standard fix). for _stream in (sys.stdout, sys.stderr): try: _stream.reconfigure(encoding="utf-8") # type: ignore[attr-defined] except (AttributeError, ValueError): pass class Term: """Tiny ANSI helper mirroring skills/_lib/term.sh (bash-only; per TERMINAL-DESIGN.md §9 the Python port is inline). Honors FORCE_COLOR / NO_COLOR / TERM_ASCII; ASCII glyph fallback on TERM_ASCII or a non-UTF stream.""" _C = {"green": "\033[32m", "yellow": "\033[33m", "orange": "\033[38;5;208m", "red": "\033[31m", "cyan": "\033[36m", "dim": "\033[2m", "off": "\033[0m"} _GLYPH = {"ok": "✓", "bad": "✗", "warn": "▲", "skip": "—", "na": "—", "unknown": "?"} _ASCII = {"ok": "+", "bad": "x", "warn": "!", "skip": "-", "na": "-", "unknown": "?"} _MARK_COLOR = {"ok": "green", "bad": "red", "warn": "orange", "skip": "dim", "na": "dim", "unknown": "yellow"} def __init__(self, stream=sys.stderr): enc = (getattr(stream, "encoding", "") or "").lower() self.ascii = (os.environ.get("TERM_ASCII") == "1" or os.environ.get("FLEET_ASCII") == "1" or "utf" not in enc) if os.environ.get("FORCE_COLOR"): self.color = True elif (os.environ.get("NO_COLOR") is not None or os.environ.get("TERM") == "dumb" or not getattr(stream, "isatty", lambda: False)()): self.color = False else: self.color = True def c(self, name, text): return f"{self._C.get(name, '')}{text}{self._C['off']}" if self.color else text def mark(self, state): return self.c(self._MARK_COLOR.get(state, ""), (self._ASCII if self.ascii else self._GLYPH).get(state, ".")) def hdr(self, text): return self.c("cyan", f"=== {text} ===") TERM = Term(sys.stderr) SCHEMA = "claude-mods.playwright-ops.flake-triage/v1" EXIT_OK = 0 EXIT_USAGE = 2 EXIT_NOT_FOUND = 3 EXIT_VALIDATION = 4 EXIT_FINDINGS = 10 # Rank order for outcomes: flaky always sorts before unexpected. OUTCOME_RANK = {"flaky": 0, "unexpected": 1} def err(msg): print(msg, file=sys.stderr) def walk_suites(suites, finds, file_hint=""): """Recursively descend the suites tree collecting spec/test results.""" for suite in suites or []: # A suite's file is on the suite node; specs inherit it. sfile = suite.get("file") or file_hint for spec in suite.get("specs", []) or []: collect_spec(spec, finds, sfile) walk_suites(suite.get("suites"), finds, sfile) def collect_spec(spec, finds, sfile): title = spec.get("title", "<untitled>") sline = spec.get("line", 0) sfile = spec.get("file") or sfile for test in spec.get("tests", []) or []: outcome = test.get("status") or test.get("outcome") or "unknown" results = test.get("results", []) or [] # status sequence ordered by retry index; duration summed across attempts ordered = sorted(results, key=lambda r: r.get("retry", 0)) statuses = [r.get("status", "unknown") for r in ordered] duration = sum(int(r.get("duration", 0) or 0) for r in ordered) retries = max((r.get("retry", 0) for r in ordered), default=0) location = f"{sfile}:{sline}" if sfile else f"?:{sline}" finds.append( { "title": title, "location": location, "outcome": outcome, "retries": retries, "statuses": statuses, "durationMs": duration, } ) def load_report(path): """Return parsed Playwright report dict, or raise ValueError if not one.""" try: raw = path.read_text(encoding="utf-8") except OSError as e: raise FileNotFoundError(str(e)) try: data = json.loads(raw) except json.JSONDecodeError as e: raise ValueError(f"not valid JSON: {e}") if not isinstance(data, dict) or "suites" not in data: raise ValueError("missing top-level 'suites' key - not a Playwright JSON report") if not isinstance(data["suites"], list): raise ValueError("'suites' is not a list — not a Playwright JSON report") return data def main(argv=None): p = argparse.ArgumentParser( prog="triage-flakes.py", description="Rank Playwright tests by flakiness from a JSON report.", formatter_class=argparse.RawDescriptionHelpFormatter, epilog=( "EXAMPLES:\n" " npx playwright test --reporter=json > results.json\n" " triage-flakes.py results.json\n" " triage-flakes.py --outcome all -n 50 results.json\n" " triage-flakes.py --json results.json | jq '.data[] | select(.outcome==\"flaky\")'\n" "\n" "EXIT CODES:\n" " 0 parsed fine, no flaky/unexpected tests (clean suite)\n" " 2 usage 3 file not found 4 malformed report\n" " 10 flaky/unexpected tests present (the triage signal)\n" ), ) p.add_argument( "report", nargs="?", default="results.json", help="path to Playwright JSON report (default: ./results.json)", ) p.add_argument("--json", action="store_true", help="emit a JSON envelope instead of TSV") p.add_argument("-q", "--quiet", action="store_true", help="suppress the stderr summary header (errors still print)") p.add_argument( "-n", "--limit", type=int, default=20, metavar="N", help="cap rows printed (default 20)", ) p.add_argument( "--outcome", default="flaky,unexpected", help="which outcomes to include: flaky | unexpected | all (default flaky,unexpected)", ) args = p.parse_args(argv) if args.limit < 0: err("ERROR: --limit must be >= 0") return EXIT_USAGE sel = args.outcome.strip().lower() if sel == "all": wanted = None # all outcomes else: wanted = {x.strip() for x in sel.split(",") if x.strip()} unknown = wanted - {"flaky", "unexpected", "expected", "skipped"} if unknown: err(f"ERROR: unknown outcome(s): {', '.join(sorted(unknown))} (use flaky|unexpected|all)") return EXIT_USAGE path = Path(args.report).resolve() if not path.exists(): err(f"ERROR: report not found: {path}") if args.json: print(json.dumps({"error": {"code": "NOT_FOUND", "message": f"report not found: {path}"}})) return EXIT_NOT_FOUND if not path.is_file(): err(f"ERROR: not a file: {path}") return EXIT_NOT_FOUND try: data = load_report(path) except FileNotFoundError as e: err(f"ERROR: cannot read report: {e}") return EXIT_NOT_FOUND except ValueError as e: err(f"ERROR: malformed report: {e}") if args.json: print(json.dumps({"error": {"code": "VALIDATION", "message": str(e)}})) return EXIT_VALIDATION finds = [] walk_suites(data.get("suites"), finds) # The domain signal is computed over ALL findings, regardless of the display # filter — a clean suite means zero flaky AND zero unexpected, full stop. signal_present = any(f["outcome"] in ("flaky", "unexpected") for f in finds) if wanted is None: shown = list(finds) else: shown = [f for f in finds if f["outcome"] in wanted] # Rank: flaky before unexpected (OUTCOME_RANK), then retries desc, duration desc. shown.sort( key=lambda f: ( OUTCOME_RANK.get(f["outcome"], 99), -f["retries"], -f["durationMs"], ) ) capped = shown[: args.limit] if args.limit else shown total = len(finds) flaky_n = sum(1 for f in finds if f["outcome"] == "flaky") unexp_n = sum(1 for f in finds if f["outcome"] == "unexpected") if not args.quiet: err(TERM.hdr(f"Flake triage: {path.name}")) flaky_txt = TERM.c("orange", f"{flaky_n} flaky") if flaky_n else "0 flaky" unexp_txt = TERM.c("red", f"{unexp_n} unexpected") if unexp_n else "0 unexpected" err(f" {total} tests | {flaky_txt} | {unexp_txt} | showing {len(capped)} of {len(shown)}") if args.json: envelope = { "data": capped, "meta": { "count": len(capped), "total_matched": len(shown), "flaky": flaky_n, "unexpected": unexp_n, "schema": SCHEMA, }, } print(json.dumps(envelope, indent=2)) else: print("outcome\tretries\tstatuses\tduration_ms\tlocation\ttitle") for f in capped: print( f"{f['outcome']}\t{f['retries']}\t{'->'.join(f['statuses'])}\t" f"{f['durationMs']}\t{f['location']}\t{f['title']}" ) return EXIT_FINDINGS if signal_present else EXIT_OK if __name__ == "__main__": sys.exit(main())
-
-
SKILL.md 15.5 KB
--- name: playwright-ops description: "Playwright end-to-end testing operations - selectors, fixtures, network mocking, auth, parallelism, CI, visual regression, flake hunting. Use for: playwright, e2e test, end-to-end testing, browser test, getByRole, page object, storageState, trace viewer, flaky test, test sharding, visual regression, toHaveScreenshot, playwright config, codegen." license: MIT allowed-tools: "Read Write Bash" metadata: author: claude-mods related-skills: testing-ops, ci-cd-ops --- # Playwright Operations End-to-end testing with Playwright Test (`@playwright/test`, TS/JS). A Python flavor (`pytest-playwright`) exists with the same browser API but pytest-style fixtures — patterns here translate directly; runner config does not. ## Quick Start ```bash npm init playwright@latest # scaffold config + example test + GH Actions workflow npx playwright test # run all tests, all projects npx playwright test --project=chromium --grep "@smoke" npx playwright test --ui # interactive UI mode (watch, time-travel) npx playwright codegen https://app.local # record actions -> generated locators npx playwright show-report # open last HTML report npx playwright show-trace trace.zip # inspect a trace ``` ## Selector Strategy **Hierarchy — always prefer the highest tier that uniquely matches:** | Tier | Locator | When | |------|---------|------| | 1 | `page.getByRole('button', { name: 'Submit' })` | Anything with an ARIA role — buttons, links, headings, textboxes. Tests a11y for free | | 2 | `page.getByLabel('Password')` | Form fields with labels | | 3 | `page.getByPlaceholder('name@example.com')` | Inputs without labels (fix the label instead, when you can) | | 4 | `page.getByText('Welcome back')` | Non-interactive text content | | 5 | `page.getByTestId('cart-total')` | Stable hook when semantics don't disambiguate. Configure attribute via `testIdAttribute` | | 6 | `page.locator('css=...')` / `xpath=` | **Last resort.** Coupled to DOM structure; breaks on refactor | Why: tiers 1–4 locate the way a user perceives the page — resilient to markup changes, and `getByRole` fails loudly when accessibility regresses. CSS/XPath encode implementation detail. **Narrowing without CSS:** ```ts page.getByRole('listitem') .filter({ hasText: 'Product 2' }) .getByRole('button', { name: 'Add to cart' }); page.getByRole('row').filter({ has: page.getByRole('cell', { name: 'Alice' }) }); ``` ### Web-First Assertions (no manual waits, ever) ```ts // BAD — checks once, races the render; sleeps are flake factories expect(await page.getByText('welcome').isVisible()).toBe(true); await page.waitForTimeout(2000); // GOOD — auto-retries until pass or timeout await expect(page.getByText('welcome')).toBeVisible(); await expect(page.getByRole('list')).toHaveCount(3); await expect(page).toHaveURL(/\/dashboard/); await expect.soft(page.getByTestId('status')).toHaveText('Active'); // don't stop test on failure ``` Actions (`click`, `fill`) auto-wait for actionability (visible, stable, enabled). If you feel the need for `waitForTimeout`, you're missing an assertion or an `await expect(...)` on a state change. For async non-DOM conditions use `expect.poll(() => fn())` or `expect(async () => {...}).toPass()`. Lint guard: enable `@typescript-eslint/no-floating-promises` — a missing `await` on an assertion is the most common silent-pass bug. ## Config Skeleton Full production template with comments: [assets/playwright.config.template.ts](assets/playwright.config.template.ts) ```ts import { defineConfig, devices } from '@playwright/test'; export default defineConfig({ testDir: './tests', fullyParallel: true, forbidOnly: !!process.env.CI, retries: process.env.CI ? 2 : 0, workers: process.env.CI ? 1 : undefined, reporter: process.env.CI ? 'blob' : 'html', use: { baseURL: process.env.BASE_URL ?? 'http://localhost:3000', trace: 'on-first-retry', testIdAttribute: 'data-testid', }, projects: [ { name: 'setup', testMatch: /.*\.setup\.ts/ }, { name: 'chromium', use: { ...devices['Desktop Chrome'], storageState: 'playwright/.auth/user.json' }, dependencies: ['setup'], }, ], webServer: { command: 'npm run dev', url: 'http://localhost:3000', reuseExistingServer: !process.env.CI, }, }); ``` ## Fixtures Decision Tree ``` What do I need to share/setup? │ ├─ Per-test object (page object, seeded record) │ └─ test.extend() test-scoped fixture — setup, await use(x), teardown │ ├─ Expensive, safe-to-share resource (DB pool, test account) │ └─ Worker-scoped: [fn, { scope: 'worker' }] — once per worker process │ ├─ Side effect every test needs (log capture, network stub) │ └─ Automatic: [fn, { auto: true }] — runs without being referenced │ ├─ Config-tunable value (locale, default item) │ └─ Option: ['default', { option: true }] — override in projects[].use │ ├─ Fixtures from several modules │ └─ mergeTests(testA, testB) │ └─ Auth state per test file/role └─ test.use({ storageState: 'playwright/.auth/admin.json' }) ``` **POM-as-fixture (modern recommendation)** — page objects are fine; *instantiating them by hand in every test* is not. Inject via fixture: ```ts // fixtures.ts import { test as base } from '@playwright/test'; import { TodoPage } from './pages/todo-page'; export const test = base.extend<{ todoPage: TodoPage }>({ todoPage: async ({ page }, use) => { const todoPage = new TodoPage(page); await todoPage.goto(); await use(todoPage); // test body runs here }, }); export { expect } from '@playwright/test'; ``` Page objects should expose **locators and actions**, not assertions wrapped in try/catch, and never store element handles. Details: [references/fixtures-and-pom.md](references/fixtures-and-pom.md) ## Network & API ``` Network need? │ ├─ Stub a third-party API → page.route('**/api/**', r => r.fulfill({ json })) ├─ Tweak a real response → const res = await route.fetch(); route.fulfill({ response: res, json }) ├─ Simulate failure / offline → route.abort() / route.fulfill({ status: 500 }) ├─ Many endpoints, real shapes → HAR record + replay (page.routeFromHAR, update: true to record) ├─ Pure API test (no browser) → request fixture / APIRequestContext ├─ Seed data fast, assert via UI → hybrid: create via request, verify via page └─ WebSocket traffic → page.routeWebSocket(url, ws => ws.onMessage(...)) ``` **Hybrid seed-via-API, assert-via-UI** — the single biggest speed win in most suites: ```ts test('shows new project', async ({ request, page }) => { const res = await request.post('/api/projects', { data: { name: 'Apollo' } }); expect(res.ok()).toBeTruthy(); await page.goto('/projects'); await expect(page.getByRole('link', { name: 'Apollo' })).toBeVisible(); }); ``` Rule of thumb: **mock third-party dependencies you don't own; exercise your own backend for real** (or mock it deliberately in a separate "frontend-isolated" project). Details: [references/network-and-api.md](references/network-and-api.md) ## Authentication Standard pattern — login once in a setup project, reuse `storageState` everywhere: ```ts // tests/auth.setup.ts import { test as setup, expect } from '@playwright/test'; setup('authenticate', async ({ page }) => { await page.goto('/login'); await page.getByLabel('Username').fill(process.env.E2E_USER!); await page.getByLabel('Password').fill(process.env.E2E_PASS!); await page.getByRole('button', { name: 'Sign in' }).click(); await expect(page.getByTestId('user-menu')).toBeVisible(); // wait for auth to settle! await page.context().storageState({ path: 'playwright/.auth/user.json' }); }); ``` | Pattern | Use when | |---------|----------| | One setup project + `storageState` in `use` | One shared account, tests don't mutate server-side user state | | Per-role files (`admin.json`, `user.json`) + `test.use({ storageState })` | Role-based behavior under test | | Worker-scoped account fixture (`testInfo.parallelIndex`) | Parallel tests mutate user state — one account per worker | | API login (`request.post` + `request.storageState`) | Login endpoint exists; 10x faster than UI login | Gotchas: add `playwright/.auth/` to `.gitignore`. `storageState` captures cookies + localStorage — **not sessionStorage** (persist that manually via `page.evaluate` + init script). Always assert a logged-in signal before saving state, or you save a half-logged-in race. ## Parallelism, Retries, Isolation | Knob | Setting | Notes | |------|---------|-------| | Workers | `workers: process.env.CI ? 1 : undefined` | Local: half the logical CPU cores. CI runners are small — shard machines instead of oversubscribing | | File-level parallel | `fullyParallel: true` | Also makes sharding split per-test, not per-file | | Sharding | `npx playwright test --shard=1/4` | One shard per CI machine; merge blob reports after | | Retries | `retries: process.env.CI ? 2 : 0` | Pair with `trace: 'on-first-retry'`; treat "flaky" status as a bug queue, not a fix | | Serial | `test.describe.configure({ mode: 'serial' })` | Smell — usually means hidden inter-test coupling | **Isolation discipline:** every test gets a fresh `context`/`page` (cookies, storage) — keep it that way. No test reads state written by another test; shared server-side state is reset via API in `beforeEach` or scoped per worker (`test.info().parallelIndex` in usernames/tenant IDs). A suite that only passes single-worker is broken, not "sensitive". **Flake diagnosis:** `trace: 'on-first-retry'` → `npx playwright show-trace` (DOM snapshots, network, console per action). Local: `npx playwright test --ui` or `PWDEBUG=1` / `page.pause()`. Repro: `--repeat-each=20 --workers=4`. Playbook: [references/flake-hunting.md](references/flake-hunting.md) **Triage a whole run without eyeballing the report** — generate the JSON reporter output, then rank the offenders with the bundled triage tool ([scripts/triage-flakes.py](scripts/triage-flakes.py)): ```bash npx playwright test --reporter=json > results.json # or reporter: [['json', { outputFile: 'results.json' }]] scripts/triage-flakes.py results.json # flaky tests first, then hard fails ``` It emits a ranked TSV (or `--json` envelope, schema `claude-mods.playwright-ops.flake-triage/v1`): flaky tests (passed only on retry) first — ordered by retry count then duration — followed by `unexpected` hard failures, each with `file:line`, the status sequence (`failed->passed`), and total duration. **Exit 10 means flakes/fails were found** (the triage signal — go fix them); exit 0 means a clean suite. `--outcome all` includes the passing tests for context; `-n N` caps rows. ## CI (GitHub Actions) ```yaml - uses: actions/checkout@v5 - uses: actions/setup-node@v5 with: { node-version: lts/* } - run: npm ci - run: npx playwright install --with-deps chromium # only browsers you test - run: npx playwright test - uses: actions/upload-artifact@v4 if: ${{ !cancelled() }} with: { name: playwright-report, path: playwright-report/, retention-days: 30 } ``` | Decision | Guidance | |----------|----------| | Container vs install-deps | `mcr.microsoft.com/playwright:vX.Y.Z-jammy` image pins browser+OS (best for visual tests); `install --with-deps` is simpler and fine otherwise. **Pin image tag to your `@playwright/test` version** | | Browser caching | Cache `~/.cache/ms-playwright` keyed on Playwright version; skip when using the container | | Sharded reports | `reporter: 'blob'` on shards → upload `blob-report/` → merge job: `npx playwright merge-reports --reporter html ./all-blob-reports` | | Fail-fast vs full suite | PRs: `fail-fast: false` + `--max-failures=10` per shard — see *all* failures in one round-trip. Smoke gates: fail fast | Full workflows (sharding matrix, merge job, caching): [references/ci-patterns.md](references/ci-patterns.md) ## Visual Testing ```ts await expect(page).toHaveScreenshot('landing.png', { maxDiffPixels: 100, // or maxDiffPixelRatio / threshold mask: [page.getByTestId('ad-banner')], // black-box dynamic regions fullPage: true, }); ``` - First run generates the baseline (test fails); update with `npx playwright test --update-snapshots` - Snapshots are named per browser **and platform** (`landing-chromium-darwin.png`) — baselines generated on macOS will not match Linux CI. Fix: generate baselines inside the same Docker image CI uses, or run visual tests only in the container - Disable animations: `toHaveScreenshot` defaults `animations: 'disabled'`; hide dynamic bits with `mask` or `stylePath` (CSS applied at capture time) - Global defaults: `expect: { toHaveScreenshot: { maxDiffPixels: 100 } }` in config - `toMatchSnapshot()` for non-image data (text/buffers) ## Component Testing & When to Prefer Cypress `@playwright/experimental-ct-react` (also vue/svelte) mounts components in a real browser — **still experimental**; for component-level work, Vitest browser mode or Testing Library are the safer default, with Playwright covering E2E. | Factor | Playwright | Cypress | |--------|-----------|---------| | Browsers | Chromium, Firefox, WebKit (real Safari engine) | Chrome-family, Firefox; WebKit experimental | | Parallelism | Free, built-in, shardable | Paid Cloud for parallel orchestration | | Multi-tab / multi-origin / iframes | Native | Historically constrained | | API testing | Built-in `request` context | Via `cy.request`, less ergonomic | | Component testing | Experimental | Mature, first-class | | In-browser interactive DX | UI mode (excellent) | The original benchmark; some teams still prefer it | Reach for Cypress when component testing maturity or an existing Cypress investment dominates; otherwise Playwright is the default for new E2E suites. (Repo also has a sibling `cypress-ops` skill.) ## Debugging & Codegen | Tool | Command | Use | |------|---------|-----| | UI mode | `npx playwright test --ui` | Watch mode, time-travel, pick locators | | Inspector | `PWDEBUG=1 npx playwright test` or `page.pause()` | Step through actions live | | Codegen | `npx playwright codegen <url>` | Records actions, emits role-based locators — treat output as a draft, refactor into POMs/fixtures | | Trace viewer | `npx playwright show-trace trace.zip` | Post-mortem: snapshots, network, console | | Headed + slow | `--headed --debug` | Eyeball a single test | | VS Code extension | — | Run/debug tests, pick locators in-editor | An official Playwright MCP server (`@playwright/mcp`) also exists for agent-driven browser automation — distinct from the test runner; don't conflate browsing automation with the test suite. ## References | File | Contents | |------|----------| | [references/fixtures-and-pom.md](references/fixtures-and-pom.md) | Fixture scopes/options/merging, POM-as-fixture architecture, anti-patterns | | [references/network-and-api.md](references/network-and-api.md) | route/fulfill/abort, HAR replay, API testing, hybrid seeding, WebSocket | | [references/ci-patterns.md](references/ci-patterns.md) | Full GH Actions workflows: basic, sharded+merge, container, caching, reporters | | [references/flake-hunting.md](references/flake-hunting.md) | Systematic flake diagnosis: traces, repro loops, common causes + fixes | | [scripts/triage-flakes.py](scripts/triage-flakes.py) | Parse a Playwright JSON report and rank flaky/failing tests (exit 10 = findings); see Flake diagnosis above | | [assets/playwright.config.template.ts](assets/playwright.config.template.ts) | Commented production config template |
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.