Claude Cursor GitHub Copilot Skill

frontend-testing-strategy-review

Reviews frontend test-pyramid shape, critical-path coverage, and flaky-test governance across unit, component, integration, and E2E layers (Vitest/Jest, Testing Library, Playwright/Cypress), loading framework references only when the task needs them.

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download vincentchuwaichow-vanguard-frontier-agentic-skills_frontend_frontend-testing-strategy-review-febe32a.zip · 13 KB
Part of vincentchuwaichow/vanguard-frontier-agentic — 293 skills

Install

skills CLI npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/frontend/frontend-testing-strategy-review
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
Git git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Frontend Testing Strategy Review

Purpose

A green CI badge does not prove a critical user journey works. This skill exists to separate "tests exist" from "tests exercise the actual failure modes that matter" — test-pyramid shape, critical-path coverage, and flaky-test governance — without dumping every testing-library's full API surface into every review.

When to use

Use this skill when the user asks to:

  • review or design a frontend test strategy across unit, component, integration, and E2E layers,
  • diagnose why a test suite is slow, flaky, or not catching regressions,
  • decide whether a given assertion belongs at the unit, component, or E2E layer,
  • audit critical-user-journey coverage (checkout, auth, forms) before a release,
  • evaluate a proposed test-framework migration (Jest to Vitest, Cypress to Playwright).

Context7 Documentation Protocol

Test-runner and testing-library APIs (retry semantics, coverage-threshold config, query priority, selector strategy) change across majors and are documented, not folklore — never assert a flag, config shape, or "best practice" from memory.

  1. Call ToolSearch with query "context7" (or "select:mcp__Context7__resolve-library-id,mcp__Context7__query-docs") to load the Context7 tools if not already loaded in this session.
  2. Call mcp__Context7__resolve-library-id for the framework in question: /microsoft/playwright for Playwright, /vitest-dev/vitest for Vitest, /testing-library/testing-library-docs for Testing Library, /cypress-io/cypress-documentation for Cypress. Prefer these resolved IDs over guessing a library name.
  3. Call mcp__Context7__query-docs for the specific claim in question — e.g. "coverage threshold configuration", "web-first assertion retry behavior", "getByRole query priority", "avoid fixed waits" — before stating it as fact. Do this per review, not once from a prior session's memory.
  4. Prefer the official docs URLs in official_docs for primary normative statements (e.g. exact CLI flags, exact config shape); use Context7 to ground and cross-check the claim before writing it into a finding.
  5. If Context7 is unavailable or returns no relevant match, fall back to the official_docs URLs and mark the claim documentation-based (Context7 unavailable) rather than presenting it as freshly verified.
  6. Never invent a config key, CLI flag, matcher name, or API method that no queried source confirms.

Lean operating rules

  • Classify the pyramid shape first (unit:component:E2E ratio, counted from actual test files/CI artifacts, not the user's description of it) before recommending any specific test; a shape problem needs a rebalancing plan, not one more test.
  • Treat coverage percentage as a weak signal; require evidence the suite asserts on error states, loading states, and accessibility tree for critical paths, not just the happy-path DOM. A file at 100% line coverage with no error-path assertion is not "well tested."
  • Treat flaky-test quarantine (.skip, test.fixme, it.skip, describe.skip, Cypress {retries: N} used to paper over root cause) as a tracked liability with an owner and expiry, never a silent, permanent state.
  • Prefer official, version-specific docs (via Context7) over memory for any framework API claim — Playwright, Vitest, Cypress, and Testing Library APIs and retry/wait semantics change across majors.
  • Never request or accept real user credentials, session cookies, or production API keys as test fixtures; require synthetic data or network mocking (MSW, cy.intercept, Playwright route interception).
  • Do not recommend a framework migration (Jest→Vitest, Cypress→Playwright) without a measured baseline (suite duration, ESM/native-TS support gap, current flake rate) and a rollback path; framework preference alone is not a migration justification.
  • Distinguish "flaky because of the test" (missing await, race on network mock, unstable selector) from "flaky because of the app" (real timing bug, unhandled async state) — the fix differs and misdiagnosis just hides a production bug behind a retry.
  • Load references only for the layer/framework in scope; do not load the Playwright reference for a pure Vitest unit-coverage question, and do not load the pyramid-shape reference for a single flaky-test triage.

References

Load these only when needed:

Response minimum

Return, at minimum:

  • the test-pyramid shape observed (unit/component/E2E counts or ratio, with the file/CI-artifact evidence it came from, not an estimate),
  • critical-user-journey coverage gaps, cited to a specific file/line or CI-artifact,
  • flaky-test inventory status for any .skip/fixme/retry-suppressed test found (owner, expiry, or "untracked liability"),
  • a minimal, prioritized diff-level test plan (not a full-suite rewrite),
  • evidence level (live evidence, repo evidence, documentation-based, inference) for every claim, and explicit flag when a claim is documentation-based (Context7 unavailable).
Files (vanguard-frontier-agentic)
  • references
    • e2e-framework-review.md 5.6 KB
      # Playwright and Cypress E2E Review
      
      Use this reference for E2E-layer specifics: locator/selector strategy, web-first assertions and retry semantics, test isolation, and evidence requirements for a Cypress↔Playwright migration decision. For where an assertion belongs in the pyramid at all, see `pyramid-shape-and-coverage.md`. For root-causing an already-flaky E2E test, see `flaky-test-governance.md`.
      
      > Version note: Playwright and Cypress config APIs and CLI flags evolve across majors. Verify exact option names/defaults against the installed version via Context7 (`/microsoft/playwright`, `/cypress-io/cypress-documentation`) or the official docs before citing a config shape in a report.
      
      ## Officially grounded design points
      
      **Playwright:**
      - Web-first assertions (`expect(locator).toHaveText(...)`, `.toBeVisible()`, etc.) auto-retry against the locator until the condition is met or a timeout is reached — this is the documented mechanism that removes the need for manual waits/sleeps. A test using a bare `page.locator(...)` value comparison instead of an `expect(locator).to...` assertion is not benefiting from this retry behavior.
      - `TestConfig.retries` (default `0`) controls CI-level retry-on-failure; `TestConfig.repeatEach` reruns a test N times and is documented specifically as a flaky-test debugging aid, not a production resilience setting.
      - Test isolation is a core design property — each test gets a fresh browser context by default; a suite that shares page/context state across tests (e.g. via a module-level `let page` reused without a `beforeEach` reset) is fighting the framework's isolation model and is a common source of order-dependent flake.
      
      **Cypress:**
      - `cy.get(...)` and other commands automatically retry/query until the element appears or a timeout is hit — this is native, not something the test author needs to add.
      - Cypress's own best-practices docs explicitly name `cy.wait(<fixed-ms>)` before an assertion as an anti-pattern to avoid; the documented fix is to let the trailing `.should(...)`/assertion retry instead of inserting a fixed delay.
      - Cypress's own best-practices docs recommend a dedicated `data-cy` (or equivalent `data-*`) attribute for element selection over tag/class/id selectors, specifically because tag/class/id are coupled to styling/behavior and churn independently of test intent.
      
      ## Non-negotiable design rules
      
      1. **Selector strategy is accessibility-first, then a stable test hook — never CSS structure.** Prefer role/label/text queries (see `unit-component-framework-review.md` for the Testing Library query-priority list, which applies to E2E-adjacent component checks too); for pure E2E flows where role/label queries are impractical, a dedicated, styling-independent test attribute (`data-testid` for Playwright, `data-cy` for Cypress, per each tool's own convention) is the documented fallback — not `nth-child`, generated class names, or brittle text matches on copy that changes with content/locale.
      2. **No fixed waits before an assertion.** Any `page.waitForTimeout(...)` or `cy.wait(<ms>)` placed immediately before an assertion (rather than waiting on a genuine external event — a specific network response, a route change) is a flake-risk anti-pattern per both tools' own guidance. Replace with the framework's auto-retrying assertion.
      3. **Mock or intercept the network at a defined boundary; do not silently mix real and mocked calls.** `cy.intercept()` / Playwright route interception should be applied consistently for a given test — a test that mocks the primary API call but lets a background analytics/tracking call hit the real network (or vice versa) introduces nondeterminism unrelated to the behavior under test.
      4. **Reserve E2E for the small set of genuinely cross-system critical journeys.** Do not let E2E become the default place to add "one more assertion" because it's easiest to write against a running app — see `pyramid-shape-and-coverage.md` for the layer-selection rule.
      5. **State isolation between tests is the framework's job — don't fight it.** Don't reuse a `page`/session across tests for speed in a way that reintroduces order dependency; use the framework's fixture/hook lifecycle (Playwright fixtures, Cypress `beforeEach`) for setup/teardown instead of manual sharing.
      
      ## Cypress → Playwright (or the reverse) migration evidence bar
      
      Do not recommend a framework switch on preference alone. Require, at minimum:
      
      - a **measured baseline**: current suite wall-clock duration in CI, current flake rate (failures not reproducible on rerun, over a stated window), and any hard blocker (e.g. multi-tab/multi-origin testing needs, which the tools support differently) driving the request,
      - an explicit statement of **what does not migrate automatically** — custom commands/plugins, CI parallelization config, visual-regression tooling integration, and existing flaky-test workarounds all need to be re-authored, not copy-pasted,
      - a **rollback path** — can the old suite keep running in parallel (even read-only, non-blocking) until the new one has proven equivalent coverage over some number of CI runs, or is this a hard cutover with no fallback if the new suite has gaps.
      
      A migration proposal missing a measured baseline or a rollback path is a rewrite driven by framework fanboyism, not evidence — push back and ask for the baseline before endorsing it.
      
      ## Response discipline
      
      When reviewing E2E tests, cite the specific selector/wait/isolation pattern found (with file reference) rather than a general "selectors could be better" statement, and label whether the claim about framework retry/isolation behavior is `documentation-based` (grounded via Context7/official docs this session) or `inference`.
      
    • flaky-test-governance.md 4.9 KB
      # Flaky-Test Diagnosis and Governance
      
      Use this reference when triaging a flaky or slow suite, reviewing quarantine (`.skip`/`fixme`/retry) usage, or setting a flake-tracking policy.
      
      ## What people get wrong
      
      The naive response to a flaky test is:
      
      > "Add a retry / bump the timeout / skip it for now, we'll fix it later."
      
      That treats the symptom and destroys the signal. A retry that makes a flaky test "pass" does not make the underlying condition go away — it either hides a real race condition in the app, or hides a bad test that will keep costing CI time and eroding trust in red builds ("it's probably just flaky" is how real regressions ship).
      
      ## Two root-cause categories — diagnose which one before prescribing a fix
      
      **1. Flaky because of the test:**
      - missing `await` on an async assertion or interaction,
      - a fixed sleep/wait (`cy.wait(3000)`, `await new Promise(r => setTimeout(r, 500))`) racing against real timing instead of waiting on a condition,
      - an unstable selector (nth-child, generated class name, text that changes with content/locale) that matches the wrong element or none, intermittently,
      - test-order dependency (shared mutable fixture/module state leaking between tests),
      - unmocked network call racing against a mocked one, or a mock that isn't reset between tests.
      
      **Fix:** correct the test. Playwright and Cypress both provide auto-retrying, web-first assertions (e.g. `expect(locator).toHaveText(...)`, `cy.get(...).should(...)`) specifically so a fixed wait is never necessary — Cypress's own best-practices guidance explicitly frames `cy.wait(<ms>)` before an assertion as the anti-pattern to replace with a retrying assertion. Prefer a stable, semantic selector (role/label/text, or a dedicated test-id attribute intentionally isolated from styling/behavior, e.g. Cypress's documented `data-cy` convention) over structural/CSS selectors.
      
      **2. Flaky because of the app:**
      - a genuine race condition (UI renders before data resolves, double-submit possible, state update ordering depends on network timing),
      - a timing-dependent bug that only manifests under CI load/latency, not locally.
      
      **Fix:** this is a production bug wearing a test-flake costume. The test correctly caught it. Do not "fix" this by adding a retry to the test — that ships the bug and just makes the test stop catching it. Escalate as an app defect finding, not a test-infra finding.
      
      Do not accept "just retry it" as a fix without first determining which category applies. A retry that happens to make a test-side-race pass is masking a bug in the test; a retry that makes an app-side race pass is masking a bug in the product.
      
      ## Quarantine governance
      
      Treat every `.skip`, `it.skip`, `describe.skip`, `test.fixme`, `xit`, or Cypress config-level `retries` override used to suppress a known-flaky test as a **tracked liability**, not a resolved state. For each instance found, the finding must record:
      
      - **owner** — who is responsible for un-quarantining it,
      - **reason** — why it's flaky (from the root-cause categories above, if known; "unknown" is an acceptable but explicit answer),
      - **expiry/tracking** — a linked issue or a stated re-review date; a skip with no linked issue and no date is an **untracked liability** and should be flagged as such, at elevated severity if the skipped test covers a critical user journey (see `pyramid-shape-and-coverage.md`).
      
      A quarantine with none of the above is worse than an honestly-failing test: it produces false confidence in a currently-green suite.
      
      ## Playwright-specific retry configuration
      
      Playwright's `TestConfig.retries` sets a maximum retry count for failed tests (default `0`, no retries) and `TestConfig.repeatEach` reruns each test N times, which the official docs frame explicitly as a *debugging* tool for flaky tests, not a permanent fix. A production CI config with `retries` set above 0 is a legitimate resilience buffer against genuine infra noise (network blips in a real CI runner) — but retries silently masking a *reproducible* failure (fails 100% of the time without the retry) is a different problem and should be flagged separately from noise-tolerance retries.
      
      > Verify the exact current default and config shape for `retries`/`repeatEach` against the installed Playwright version via Context7 or `official_docs` before citing a specific number in a report — these are config surface, not folklore.
      
      ## Response discipline
      
      When reporting a flaky-test finding, state:
      
      - which root-cause category it falls into (test-side vs app-side), with the specific evidence (missing await, fixed wait, unstable selector, or a described race condition),
      - current quarantine status and whether it is tracked (owner + expiry) or untracked,
      - whether the flaky test covers a critical user journey (elevates severity if so — see `pyramid-shape-and-coverage.md`),
      - the minimal fix (stabilize the test) vs the escalation (file an app-defect finding) — do not conflate the two into a single "add retries" recommendation.
      
    • pyramid-shape-and-coverage.md 4.9 KB
      # Test Pyramid Shape and Critical-Path Coverage
      
      Use this reference when classifying the unit:component:E2E ratio of a suite, deciding which layer a given assertion belongs at, or auditing whether critical user journeys actually have coverage.
      
      ## What people get wrong
      
      The naive story is:
      
      > "We have 2,000 tests and 85% coverage, so the suite is solid."
      
      That number says nothing about shape or what those tests actually assert. Two failure patterns hide behind a healthy-looking coverage number:
      
      1. **Ice-cream-cone shape** — few unit tests, a moderate component layer, and a huge, slow, brittle E2E layer that re-proves logic the unit layer should own. This is slow CI, high flake surface, and expensive maintenance for marginal confidence gain.
      2. **Hollow pyramid** — lots of unit tests, but they assert on implementation details (internal state, private methods, snapshot diffs of unrelated markup) instead of behavior, so a real regression in user-facing behavior can still ship green.
      
      Coverage percentage cannot distinguish either failure from a healthy suite. Only reading the tests can.
      
      ## Classifying the shape
      
      Count actual test files/cases per layer, don't accept a verbal estimate:
      
      - **Unit** — pure functions, reducers, utility/hook logic in isolation, no DOM rendering.
      - **Component** — a single component rendered with Testing Library (or framework-equivalent), asserting on rendered output/accessibility tree and user interaction, network mocked.
      - **Integration** — multiple components/modules wired together (e.g. a form + its validation + a mocked API layer), still no real browser.
      - **E2E** — real (or real-enough) browser, real routing, hits a running app (locally or staging), via Playwright/Cypress.
      
      A commonly cited healthy shape is unit-heavy, E2E-light — many more fast, isolated tests than slow, full-stack ones — but do not apply a fixed ratio as a hard gate; the right shape depends on the codebase's actual risk surface (a form-heavy app has a legitimately larger component layer than a data-pipeline-heavy one). Flag a shape as a *finding* when:
      
      - E2E test count is a large fraction of total tests but the app's core logic is stateless/computational (should be unit-tested instead), or
      - there are numerous integration/E2E tests re-asserting something a unit test already covers with no additional confidence gained, or
      - the unit layer is large but assertion inspection shows most assertions target internal state/props rather than rendered/observable behavior.
      
      ## Deciding which layer an assertion belongs at
      
      - **Business logic, formatting, validation rules, reducers/selectors** → unit. If it doesn't need a DOM, don't give it one.
      - **"Does this component render correctly given these props/state, and can a user interact with it?"** → component, with network/API mocked.
      - **"Do these pieces work together" (form + validation + submit handler + mocked API)** → integration.
      - **"Does the real critical journey work end-to-end through real routing/auth/persistence?"** → E2E, reserved for a *small number* of the highest-value flows, not every page.
      
      Push back when a user proposes writing an E2E test for something a unit or component test can cover — the E2E test costs more (runtime, flakiness surface, infra) for the same assertion.
      
      ## Critical-user-journey coverage audit
      
      A "critical user journey" is a flow whose failure has direct revenue, compliance, or trust impact — auth (login/signup/logout/password reset), checkout/payment, core conversion action (the thing the product exists to let users do), and any flow with a regulatory obligation (consent capture, data export/deletion).
      
      For each critical journey, verify presence of assertions on:
      
      - **the happy path**, obviously, but that alone is not sufficient;
      - **error states** — network failure, validation rejection, expired session, rate limiting — does the suite prove the user sees a recoverable error, or does it stop at the happy path and assume errors "probably work"?
      - **loading/pending states** — is there an assertion the UI shows a loading indicator and doesn't allow a double-submit, or is this untested?
      - **accessibility tree for the interactive path** — for a critical flow, are the interactive elements queried via role/label (proving they're actually reachable via assistive tech), or only via test-id/class (which proves nothing about usability)?
      
      A journey with only a happy-path E2E test and no error-state coverage is a **coverage gap**, not "covered." State this explicitly in the finding rather than crediting partial coverage as complete.
      
      ## Response discipline
      
      When reporting pyramid shape or coverage gaps, cite the evidence:
      
      - file/directory counts (e.g. "14 `*.spec.ts` under `e2e/`, 6 under `src/**/*.test.ts`"), or
      - a CI coverage/test-report artifact if the user provided one.
      
      Do not report a ratio or gap you inferred from the user's description alone without inspecting the actual test files — label such a claim `inference` and say so.
      
    • unit-component-framework-review.md 4.9 KB
      # Vitest and Testing Library Unit/Component Review
      
      Use this reference for unit/component-layer specifics: query priority (`getByRole` over test IDs), mocking boundaries, and coverage-threshold configuration. For where an assertion belongs in the pyramid at all, see `pyramid-shape-and-coverage.md`.
      
      > Version note: Vitest and Testing Library config/query APIs evolve across majors (e.g. Vitest's coverage-threshold `perFile` inheritance behavior changed across versions). Verify exact option names/defaults against the installed version via Context7 (`/vitest-dev/vitest`, `/testing-library/testing-library-docs`) or the official docs before citing a config shape or migration-breaking change in a report.
      
      ## Officially grounded design points
      
      **Testing Library query priority (documented, not a style opinion):**
      
      Testing Library's own docs define an explicit priority order, and a review should measure component tests against it directly:
      
      1. **Queries accessible to everyone** — top priority, because they reflect both visual and assistive-technology user experience: `getByRole` (the primary choice — it exposes name/role/state exactly as the accessibility tree does), `getByLabelText` (form fields), `getByPlaceholderText` (secondary, for inputs), `getByText` (non-interactive elements), `getByDisplayValue` (current form value).
      2. **Semantic queries** — lower priority than accessible queries; use when an accessible query isn't practical.
      3. **Test IDs** — documented as the *last resort*, specifically because a `data-testid` proves nothing about whether a real user (including one using assistive technology) can actually find/use the element.
      
      A component test suite that predominantly uses `getByTestId`/`container.querySelector` where `getByRole`/`getByLabelText` would work is a **finding**, not a style nit — it means the tests are not verifying the thing that actually matters (can a user, including an AT user, interact with this), and it silently permits an accessibility regression (e.g. a missing accessible name) to pass a test that only checked a test-id existed.
      
      Custom test-id queries built via `queryHelpers.queryByAttribute` are supported for a project's own attribute convention (e.g. `data-test-id` instead of `data-testid`), but the priority order still applies — a custom test-id query is still a last resort, not a replacement for role/label queries.
      
      **Vitest coverage thresholds:**
      
      Vitest's `coverage.thresholds` config supports global thresholds (`lines`, `functions`, `branches`, `statements`), per-glob-pattern thresholds (each pattern requiring `perFile: true` explicitly to apply per-file, since patterns do not inherit the top-level `perFile` setting), and negative-number thresholds meaning "no more than N uncovered items" rather than a percentage. A reviewed config claiming a threshold is enforced should be checked against these actual semantics — e.g. a glob-pattern threshold block without its own `perFile: true` will NOT enforce per-file coverage even if the top-level config has `perFile: true`, because glob patterns do not inherit it.
      
      ## Non-negotiable design rules
      
      1. **Assert on rendered/observable behavior, not implementation detail.** A component test that reaches into instance state, calls a private method, or snapshot-diffs unrelated markup churn (e.g. a full-component snapshot that breaks on any unrelated style change) is brittle and doesn't prove user-facing correctness — see the "hollow pyramid" failure mode in `pyramid-shape-and-coverage.md`.
      2. **Query by role/label first; test-id is the last resort, not the default.** Apply the Testing Library priority order above as a hard review criterion, not a suggestion.
      3. **Mock at the network boundary, not inside application logic.** Component/integration tests should mock the API layer (MSW, or an injected fetch client) rather than mocking internal application functions, so the test still exercises real component logic/wiring and doesn't silently pass after a refactor that keeps the mocked function's shape but breaks real integration.
      4. **Coverage thresholds are a floor, not a target.** A threshold catches *regression below a floor*; it does not prove the newly-added code at 100% coverage actually asserts on error/loading states (see `pyramid-shape-and-coverage.md`). Do not treat "coverage threshold met" as equivalent to "critical path covered."
      5. **Watch-mode/local dev ergonomics should not leak into CI config.** Settings tuned for fast local iteration (e.g. skipping type-checking, relaxed isolation) need an explicit, separate CI config — verify the CI config path, don't assume local config is what runs in the pipeline.
      
      ## Response discipline
      
      When reviewing unit/component tests, cite the specific query/mock/assertion pattern found (with file reference), state which Testing Library priority tier the primary queries fall into, and label whether a coverage-threshold or config-semantics claim is `documentation-based` (grounded via Context7/official docs this session) or `inference`.
      
  • metadata.json 1.2 KB
    {
      "id": "frontend-testing-strategy-review",
      "name": "Frontend Testing Strategy Review",
      "type": "skill",
      "provider": "frontend",
      "harnesses": [
        "claude-code",
        "cursor",
        "codex",
        "gemini",
        "kiro",
        "other"
      ],
      "summary": "Reviews unit/component/integration/E2E test-pyramid shape, coverage of critical user journeys, and flaky-test governance for a frontend codebase using Vitest/Jest, Testing Library, and Playwright/Cypress, loading framework-specific reference guidance only when needed.",
      "source_type": "original",
      "official_docs": [
        "https://playwright.dev/docs/best-practices",
        "https://vitest.dev/guide/",
        "https://testing-library.com/docs/queries/about/#priority",
        "https://docs.cypress.io/app/core-concepts/best-practices"
      ],
      "security_notes": "Never accept or request real credentials/session tokens for test fixtures; require synthetic data or mocked network layers (MSW, cy.intercept, Playwright route interception). Flag any test-only auth/CSRF bypass flag that could ship enabled in a production build as a hard-gate finding.",
      "last_verified": "2026-07-02",
      "path": "skills/frontend/frontend-testing-strategy-review",
      "author": "github: VincentChuWaiChow",
      "version": "0.1.0"
    }
    
  • SKILL.md 6.3 KB
    ---
    name: frontend-testing-strategy-review
    description: Reviews frontend test-pyramid shape, critical-path coverage, and flaky-test governance across unit, component, integration, and E2E layers (Vitest/Jest, Testing Library, Playwright/Cypress), loading framework references only when the task needs them.
    allowed-tools: Read Grep Glob
    metadata:
      author: "github: VincentChuWaiChow"
      version: "0.1.0"
      updated: "2026-07-02"
      category: delivery
    ---
    
    # Frontend Testing Strategy Review
    
    ## Purpose
    
    A green CI badge does not prove a critical user journey works. This skill exists to separate "tests exist" from "tests exercise the actual failure modes that matter" — test-pyramid shape, critical-path coverage, and flaky-test governance — without dumping every testing-library's full API surface into every review.
    
    ## When to use
    
    Use this skill when the user asks to:
    
    - review or design a frontend test strategy across unit, component, integration, and E2E layers,
    - diagnose why a test suite is slow, flaky, or not catching regressions,
    - decide whether a given assertion belongs at the unit, component, or E2E layer,
    - audit critical-user-journey coverage (checkout, auth, forms) before a release,
    - evaluate a proposed test-framework migration (Jest to Vitest, Cypress to Playwright).
    
    ## Context7 Documentation Protocol
    
    Test-runner and testing-library APIs (retry semantics, coverage-threshold config, query priority, selector strategy) change across majors and are documented, not folklore — never assert a flag, config shape, or "best practice" from memory.
    
    1. Call `ToolSearch` with query `"context7"` (or `"select:mcp__Context7__resolve-library-id,mcp__Context7__query-docs"`) to load the Context7 tools if not already loaded in this session.
    2. Call `mcp__Context7__resolve-library-id` for the framework in question: `/microsoft/playwright` for Playwright, `/vitest-dev/vitest` for Vitest, `/testing-library/testing-library-docs` for Testing Library, `/cypress-io/cypress-documentation` for Cypress. Prefer these resolved IDs over guessing a library name.
    3. Call `mcp__Context7__query-docs` for the specific claim in question — e.g. "coverage threshold configuration", "web-first assertion retry behavior", "getByRole query priority", "avoid fixed waits" — before stating it as fact. Do this per review, not once from a prior session's memory.
    4. Prefer the official docs URLs in `official_docs` for primary normative statements (e.g. exact CLI flags, exact config shape); use Context7 to ground and cross-check the claim before writing it into a finding.
    5. If Context7 is unavailable or returns no relevant match, fall back to the `official_docs` URLs and mark the claim `documentation-based (Context7 unavailable)` rather than presenting it as freshly verified.
    6. Never invent a config key, CLI flag, matcher name, or API method that no queried source confirms.
    
    ## Lean operating rules
    
    - Classify the pyramid shape first (unit:component:E2E ratio, counted from actual test files/CI artifacts, not the user's description of it) before recommending any specific test; a shape problem needs a rebalancing plan, not one more test.
    - Treat coverage percentage as a weak signal; require evidence the suite asserts on error states, loading states, and accessibility tree for critical paths, not just the happy-path DOM. A file at 100% line coverage with no error-path assertion is not "well tested."
    - Treat flaky-test quarantine (`.skip`, `test.fixme`, `it.skip`, `describe.skip`, Cypress `{retries: N}` used to paper over root cause) as a tracked liability with an owner and expiry, never a silent, permanent state.
    - Prefer official, version-specific docs (via Context7) over memory for any framework API claim — Playwright, Vitest, Cypress, and Testing Library APIs and retry/wait semantics change across majors.
    - Never request or accept real user credentials, session cookies, or production API keys as test fixtures; require synthetic data or network mocking (MSW, `cy.intercept`, Playwright route interception).
    - Do not recommend a framework migration (Jest→Vitest, Cypress→Playwright) without a measured baseline (suite duration, ESM/native-TS support gap, current flake rate) and a rollback path; framework preference alone is not a migration justification.
    - Distinguish "flaky because of the test" (missing await, race on network mock, unstable selector) from "flaky because of the app" (real timing bug, unhandled async state) — the fix differs and misdiagnosis just hides a production bug behind a retry.
    - Load references only for the layer/framework in scope; do not load the Playwright reference for a pure Vitest unit-coverage question, and do not load the pyramid-shape reference for a single flaky-test triage.
    
    ## References
    
    Load these only when needed:
    
    - [Test pyramid shape and critical-path coverage](references/pyramid-shape-and-coverage.md) — use when classifying unit:component:E2E ratio, deciding which layer an assertion belongs at, or auditing critical-user-journey coverage.
    - [Flaky-test diagnosis and governance](references/flaky-test-governance.md) — use when triaging a flaky or slow suite, reviewing quarantine (`.skip`/`fixme`) usage, or setting a flake-tracking policy.
    - [Playwright and Cypress E2E review](references/e2e-framework-review.md) — use for E2E-layer specifics: locator/selector strategy, web-first assertions and retry semantics, test isolation, and Jest/Cypress-to-Playwright migration evidence requirements.
    - [Vitest and Testing Library unit/component review](references/unit-component-framework-review.md) — use for unit/component-layer specifics: query priority (`getByRole` over test IDs), mocking boundaries, and coverage-threshold configuration.
    
    ## Response minimum
    
    Return, at minimum:
    
    - the test-pyramid shape observed (unit/component/E2E counts or ratio, with the file/CI-artifact evidence it came from, not an estimate),
    - critical-user-journey coverage gaps, cited to a specific file/line or CI-artifact,
    - flaky-test inventory status for any `.skip`/`fixme`/retry-suppressed test found (owner, expiry, or "untracked liability"),
    - a minimal, prioritized diff-level test plan (not a full-suite rewrite),
    - evidence level (`live evidence`, `repo evidence`, `documentation-based`, `inference`) for every claim, and explicit flag when a claim is `documentation-based (Context7 unavailable)`.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related