frontend-testing-strategy-review
Reviews frontend test-pyramid shape, critical-path coverage, and flaky-test governance across unit, component, integration, and E2E layers (Vitest/Jest, Testing Library, Playwright/Cypress), loading framework references only when the task needs them.
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/frontend/frontend-testing-strategy-review
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Frontend Testing Strategy Review
Purpose
A green CI badge does not prove a critical user journey works. This skill exists to separate "tests exist" from "tests exercise the actual failure modes that matter" — test-pyramid shape, critical-path coverage, and flaky-test governance — without dumping every testing-library's full API surface into every review.
When to use
Use this skill when the user asks to:
- review or design a frontend test strategy across unit, component, integration, and E2E layers,
- diagnose why a test suite is slow, flaky, or not catching regressions,
- decide whether a given assertion belongs at the unit, component, or E2E layer,
- audit critical-user-journey coverage (checkout, auth, forms) before a release,
- evaluate a proposed test-framework migration (Jest to Vitest, Cypress to Playwright).
Context7 Documentation Protocol
Test-runner and testing-library APIs (retry semantics, coverage-threshold config, query priority, selector strategy) change across majors and are documented, not folklore — never assert a flag, config shape, or "best practice" from memory.
- Call
ToolSearchwith query"context7"(or"select:mcp__Context7__resolve-library-id,mcp__Context7__query-docs") to load the Context7 tools if not already loaded in this session. - Call
mcp__Context7__resolve-library-idfor the framework in question:/microsoft/playwrightfor Playwright,/vitest-dev/vitestfor Vitest,/testing-library/testing-library-docsfor Testing Library,/cypress-io/cypress-documentationfor Cypress. Prefer these resolved IDs over guessing a library name. - Call
mcp__Context7__query-docsfor the specific claim in question — e.g. "coverage threshold configuration", "web-first assertion retry behavior", "getByRole query priority", "avoid fixed waits" — before stating it as fact. Do this per review, not once from a prior session's memory. - Prefer the official docs URLs in
official_docsfor primary normative statements (e.g. exact CLI flags, exact config shape); use Context7 to ground and cross-check the claim before writing it into a finding. - If Context7 is unavailable or returns no relevant match, fall back to the
official_docsURLs and mark the claimdocumentation-based (Context7 unavailable)rather than presenting it as freshly verified. - Never invent a config key, CLI flag, matcher name, or API method that no queried source confirms.
Lean operating rules
- Classify the pyramid shape first (unit:component:E2E ratio, counted from actual test files/CI artifacts, not the user's description of it) before recommending any specific test; a shape problem needs a rebalancing plan, not one more test.
- Treat coverage percentage as a weak signal; require evidence the suite asserts on error states, loading states, and accessibility tree for critical paths, not just the happy-path DOM. A file at 100% line coverage with no error-path assertion is not "well tested."
- Treat flaky-test quarantine (
.skip,test.fixme,it.skip,describe.skip, Cypress{retries: N}used to paper over root cause) as a tracked liability with an owner and expiry, never a silent, permanent state. - Prefer official, version-specific docs (via Context7) over memory for any framework API claim — Playwright, Vitest, Cypress, and Testing Library APIs and retry/wait semantics change across majors.
- Never request or accept real user credentials, session cookies, or production API keys as test fixtures; require synthetic data or network mocking (MSW,
cy.intercept, Playwright route interception). - Do not recommend a framework migration (Jest→Vitest, Cypress→Playwright) without a measured baseline (suite duration, ESM/native-TS support gap, current flake rate) and a rollback path; framework preference alone is not a migration justification.
- Distinguish "flaky because of the test" (missing await, race on network mock, unstable selector) from "flaky because of the app" (real timing bug, unhandled async state) — the fix differs and misdiagnosis just hides a production bug behind a retry.
- Load references only for the layer/framework in scope; do not load the Playwright reference for a pure Vitest unit-coverage question, and do not load the pyramid-shape reference for a single flaky-test triage.
References
Load these only when needed:
- Test pyramid shape and critical-path coverage — use when classifying unit:component:E2E ratio, deciding which layer an assertion belongs at, or auditing critical-user-journey coverage.
- Flaky-test diagnosis and governance — use when triaging a flaky or slow suite, reviewing quarantine (
.skip/fixme) usage, or setting a flake-tracking policy. - Playwright and Cypress E2E review — use for E2E-layer specifics: locator/selector strategy, web-first assertions and retry semantics, test isolation, and Jest/Cypress-to-Playwright migration evidence requirements.
- Vitest and Testing Library unit/component review — use for unit/component-layer specifics: query priority (
getByRoleover test IDs), mocking boundaries, and coverage-threshold configuration.
Response minimum
Return, at minimum:
- the test-pyramid shape observed (unit/component/E2E counts or ratio, with the file/CI-artifact evidence it came from, not an estimate),
- critical-user-journey coverage gaps, cited to a specific file/line or CI-artifact,
- flaky-test inventory status for any
.skip/fixme/retry-suppressed test found (owner, expiry, or "untracked liability"), - a minimal, prioritized diff-level test plan (not a full-suite rewrite),
- evidence level (
live evidence,repo evidence,documentation-based,inference) for every claim, and explicit flag when a claim isdocumentation-based (Context7 unavailable).
Files (vanguard-frontier-agentic)
-
references
-
e2e-framework-review.md 5.6 KB
# Playwright and Cypress E2E Review Use this reference for E2E-layer specifics: locator/selector strategy, web-first assertions and retry semantics, test isolation, and evidence requirements for a Cypress↔Playwright migration decision. For where an assertion belongs in the pyramid at all, see `pyramid-shape-and-coverage.md`. For root-causing an already-flaky E2E test, see `flaky-test-governance.md`. > Version note: Playwright and Cypress config APIs and CLI flags evolve across majors. Verify exact option names/defaults against the installed version via Context7 (`/microsoft/playwright`, `/cypress-io/cypress-documentation`) or the official docs before citing a config shape in a report. ## Officially grounded design points **Playwright:** - Web-first assertions (`expect(locator).toHaveText(...)`, `.toBeVisible()`, etc.) auto-retry against the locator until the condition is met or a timeout is reached — this is the documented mechanism that removes the need for manual waits/sleeps. A test using a bare `page.locator(...)` value comparison instead of an `expect(locator).to...` assertion is not benefiting from this retry behavior. - `TestConfig.retries` (default `0`) controls CI-level retry-on-failure; `TestConfig.repeatEach` reruns a test N times and is documented specifically as a flaky-test debugging aid, not a production resilience setting. - Test isolation is a core design property — each test gets a fresh browser context by default; a suite that shares page/context state across tests (e.g. via a module-level `let page` reused without a `beforeEach` reset) is fighting the framework's isolation model and is a common source of order-dependent flake. **Cypress:** - `cy.get(...)` and other commands automatically retry/query until the element appears or a timeout is hit — this is native, not something the test author needs to add. - Cypress's own best-practices docs explicitly name `cy.wait(<fixed-ms>)` before an assertion as an anti-pattern to avoid; the documented fix is to let the trailing `.should(...)`/assertion retry instead of inserting a fixed delay. - Cypress's own best-practices docs recommend a dedicated `data-cy` (or equivalent `data-*`) attribute for element selection over tag/class/id selectors, specifically because tag/class/id are coupled to styling/behavior and churn independently of test intent. ## Non-negotiable design rules 1. **Selector strategy is accessibility-first, then a stable test hook — never CSS structure.** Prefer role/label/text queries (see `unit-component-framework-review.md` for the Testing Library query-priority list, which applies to E2E-adjacent component checks too); for pure E2E flows where role/label queries are impractical, a dedicated, styling-independent test attribute (`data-testid` for Playwright, `data-cy` for Cypress, per each tool's own convention) is the documented fallback — not `nth-child`, generated class names, or brittle text matches on copy that changes with content/locale. 2. **No fixed waits before an assertion.** Any `page.waitForTimeout(...)` or `cy.wait(<ms>)` placed immediately before an assertion (rather than waiting on a genuine external event — a specific network response, a route change) is a flake-risk anti-pattern per both tools' own guidance. Replace with the framework's auto-retrying assertion. 3. **Mock or intercept the network at a defined boundary; do not silently mix real and mocked calls.** `cy.intercept()` / Playwright route interception should be applied consistently for a given test — a test that mocks the primary API call but lets a background analytics/tracking call hit the real network (or vice versa) introduces nondeterminism unrelated to the behavior under test. 4. **Reserve E2E for the small set of genuinely cross-system critical journeys.** Do not let E2E become the default place to add "one more assertion" because it's easiest to write against a running app — see `pyramid-shape-and-coverage.md` for the layer-selection rule. 5. **State isolation between tests is the framework's job — don't fight it.** Don't reuse a `page`/session across tests for speed in a way that reintroduces order dependency; use the framework's fixture/hook lifecycle (Playwright fixtures, Cypress `beforeEach`) for setup/teardown instead of manual sharing. ## Cypress → Playwright (or the reverse) migration evidence bar Do not recommend a framework switch on preference alone. Require, at minimum: - a **measured baseline**: current suite wall-clock duration in CI, current flake rate (failures not reproducible on rerun, over a stated window), and any hard blocker (e.g. multi-tab/multi-origin testing needs, which the tools support differently) driving the request, - an explicit statement of **what does not migrate automatically** — custom commands/plugins, CI parallelization config, visual-regression tooling integration, and existing flaky-test workarounds all need to be re-authored, not copy-pasted, - a **rollback path** — can the old suite keep running in parallel (even read-only, non-blocking) until the new one has proven equivalent coverage over some number of CI runs, or is this a hard cutover with no fallback if the new suite has gaps. A migration proposal missing a measured baseline or a rollback path is a rewrite driven by framework fanboyism, not evidence — push back and ask for the baseline before endorsing it. ## Response discipline When reviewing E2E tests, cite the specific selector/wait/isolation pattern found (with file reference) rather than a general "selectors could be better" statement, and label whether the claim about framework retry/isolation behavior is `documentation-based` (grounded via Context7/official docs this session) or `inference`. -
flaky-test-governance.md 4.9 KB
# Flaky-Test Diagnosis and Governance Use this reference when triaging a flaky or slow suite, reviewing quarantine (`.skip`/`fixme`/retry) usage, or setting a flake-tracking policy. ## What people get wrong The naive response to a flaky test is: > "Add a retry / bump the timeout / skip it for now, we'll fix it later." That treats the symptom and destroys the signal. A retry that makes a flaky test "pass" does not make the underlying condition go away — it either hides a real race condition in the app, or hides a bad test that will keep costing CI time and eroding trust in red builds ("it's probably just flaky" is how real regressions ship). ## Two root-cause categories — diagnose which one before prescribing a fix **1. Flaky because of the test:** - missing `await` on an async assertion or interaction, - a fixed sleep/wait (`cy.wait(3000)`, `await new Promise(r => setTimeout(r, 500))`) racing against real timing instead of waiting on a condition, - an unstable selector (nth-child, generated class name, text that changes with content/locale) that matches the wrong element or none, intermittently, - test-order dependency (shared mutable fixture/module state leaking between tests), - unmocked network call racing against a mocked one, or a mock that isn't reset between tests. **Fix:** correct the test. Playwright and Cypress both provide auto-retrying, web-first assertions (e.g. `expect(locator).toHaveText(...)`, `cy.get(...).should(...)`) specifically so a fixed wait is never necessary — Cypress's own best-practices guidance explicitly frames `cy.wait(<ms>)` before an assertion as the anti-pattern to replace with a retrying assertion. Prefer a stable, semantic selector (role/label/text, or a dedicated test-id attribute intentionally isolated from styling/behavior, e.g. Cypress's documented `data-cy` convention) over structural/CSS selectors. **2. Flaky because of the app:** - a genuine race condition (UI renders before data resolves, double-submit possible, state update ordering depends on network timing), - a timing-dependent bug that only manifests under CI load/latency, not locally. **Fix:** this is a production bug wearing a test-flake costume. The test correctly caught it. Do not "fix" this by adding a retry to the test — that ships the bug and just makes the test stop catching it. Escalate as an app defect finding, not a test-infra finding. Do not accept "just retry it" as a fix without first determining which category applies. A retry that happens to make a test-side-race pass is masking a bug in the test; a retry that makes an app-side race pass is masking a bug in the product. ## Quarantine governance Treat every `.skip`, `it.skip`, `describe.skip`, `test.fixme`, `xit`, or Cypress config-level `retries` override used to suppress a known-flaky test as a **tracked liability**, not a resolved state. For each instance found, the finding must record: - **owner** — who is responsible for un-quarantining it, - **reason** — why it's flaky (from the root-cause categories above, if known; "unknown" is an acceptable but explicit answer), - **expiry/tracking** — a linked issue or a stated re-review date; a skip with no linked issue and no date is an **untracked liability** and should be flagged as such, at elevated severity if the skipped test covers a critical user journey (see `pyramid-shape-and-coverage.md`). A quarantine with none of the above is worse than an honestly-failing test: it produces false confidence in a currently-green suite. ## Playwright-specific retry configuration Playwright's `TestConfig.retries` sets a maximum retry count for failed tests (default `0`, no retries) and `TestConfig.repeatEach` reruns each test N times, which the official docs frame explicitly as a *debugging* tool for flaky tests, not a permanent fix. A production CI config with `retries` set above 0 is a legitimate resilience buffer against genuine infra noise (network blips in a real CI runner) — but retries silently masking a *reproducible* failure (fails 100% of the time without the retry) is a different problem and should be flagged separately from noise-tolerance retries. > Verify the exact current default and config shape for `retries`/`repeatEach` against the installed Playwright version via Context7 or `official_docs` before citing a specific number in a report — these are config surface, not folklore. ## Response discipline When reporting a flaky-test finding, state: - which root-cause category it falls into (test-side vs app-side), with the specific evidence (missing await, fixed wait, unstable selector, or a described race condition), - current quarantine status and whether it is tracked (owner + expiry) or untracked, - whether the flaky test covers a critical user journey (elevates severity if so — see `pyramid-shape-and-coverage.md`), - the minimal fix (stabilize the test) vs the escalation (file an app-defect finding) — do not conflate the two into a single "add retries" recommendation. -
pyramid-shape-and-coverage.md 4.9 KB
# Test Pyramid Shape and Critical-Path Coverage Use this reference when classifying the unit:component:E2E ratio of a suite, deciding which layer a given assertion belongs at, or auditing whether critical user journeys actually have coverage. ## What people get wrong The naive story is: > "We have 2,000 tests and 85% coverage, so the suite is solid." That number says nothing about shape or what those tests actually assert. Two failure patterns hide behind a healthy-looking coverage number: 1. **Ice-cream-cone shape** — few unit tests, a moderate component layer, and a huge, slow, brittle E2E layer that re-proves logic the unit layer should own. This is slow CI, high flake surface, and expensive maintenance for marginal confidence gain. 2. **Hollow pyramid** — lots of unit tests, but they assert on implementation details (internal state, private methods, snapshot diffs of unrelated markup) instead of behavior, so a real regression in user-facing behavior can still ship green. Coverage percentage cannot distinguish either failure from a healthy suite. Only reading the tests can. ## Classifying the shape Count actual test files/cases per layer, don't accept a verbal estimate: - **Unit** — pure functions, reducers, utility/hook logic in isolation, no DOM rendering. - **Component** — a single component rendered with Testing Library (or framework-equivalent), asserting on rendered output/accessibility tree and user interaction, network mocked. - **Integration** — multiple components/modules wired together (e.g. a form + its validation + a mocked API layer), still no real browser. - **E2E** — real (or real-enough) browser, real routing, hits a running app (locally or staging), via Playwright/Cypress. A commonly cited healthy shape is unit-heavy, E2E-light — many more fast, isolated tests than slow, full-stack ones — but do not apply a fixed ratio as a hard gate; the right shape depends on the codebase's actual risk surface (a form-heavy app has a legitimately larger component layer than a data-pipeline-heavy one). Flag a shape as a *finding* when: - E2E test count is a large fraction of total tests but the app's core logic is stateless/computational (should be unit-tested instead), or - there are numerous integration/E2E tests re-asserting something a unit test already covers with no additional confidence gained, or - the unit layer is large but assertion inspection shows most assertions target internal state/props rather than rendered/observable behavior. ## Deciding which layer an assertion belongs at - **Business logic, formatting, validation rules, reducers/selectors** → unit. If it doesn't need a DOM, don't give it one. - **"Does this component render correctly given these props/state, and can a user interact with it?"** → component, with network/API mocked. - **"Do these pieces work together" (form + validation + submit handler + mocked API)** → integration. - **"Does the real critical journey work end-to-end through real routing/auth/persistence?"** → E2E, reserved for a *small number* of the highest-value flows, not every page. Push back when a user proposes writing an E2E test for something a unit or component test can cover — the E2E test costs more (runtime, flakiness surface, infra) for the same assertion. ## Critical-user-journey coverage audit A "critical user journey" is a flow whose failure has direct revenue, compliance, or trust impact — auth (login/signup/logout/password reset), checkout/payment, core conversion action (the thing the product exists to let users do), and any flow with a regulatory obligation (consent capture, data export/deletion). For each critical journey, verify presence of assertions on: - **the happy path**, obviously, but that alone is not sufficient; - **error states** — network failure, validation rejection, expired session, rate limiting — does the suite prove the user sees a recoverable error, or does it stop at the happy path and assume errors "probably work"? - **loading/pending states** — is there an assertion the UI shows a loading indicator and doesn't allow a double-submit, or is this untested? - **accessibility tree for the interactive path** — for a critical flow, are the interactive elements queried via role/label (proving they're actually reachable via assistive tech), or only via test-id/class (which proves nothing about usability)? A journey with only a happy-path E2E test and no error-state coverage is a **coverage gap**, not "covered." State this explicitly in the finding rather than crediting partial coverage as complete. ## Response discipline When reporting pyramid shape or coverage gaps, cite the evidence: - file/directory counts (e.g. "14 `*.spec.ts` under `e2e/`, 6 under `src/**/*.test.ts`"), or - a CI coverage/test-report artifact if the user provided one. Do not report a ratio or gap you inferred from the user's description alone without inspecting the actual test files — label such a claim `inference` and say so. -
unit-component-framework-review.md 4.9 KB
# Vitest and Testing Library Unit/Component Review Use this reference for unit/component-layer specifics: query priority (`getByRole` over test IDs), mocking boundaries, and coverage-threshold configuration. For where an assertion belongs in the pyramid at all, see `pyramid-shape-and-coverage.md`. > Version note: Vitest and Testing Library config/query APIs evolve across majors (e.g. Vitest's coverage-threshold `perFile` inheritance behavior changed across versions). Verify exact option names/defaults against the installed version via Context7 (`/vitest-dev/vitest`, `/testing-library/testing-library-docs`) or the official docs before citing a config shape or migration-breaking change in a report. ## Officially grounded design points **Testing Library query priority (documented, not a style opinion):** Testing Library's own docs define an explicit priority order, and a review should measure component tests against it directly: 1. **Queries accessible to everyone** — top priority, because they reflect both visual and assistive-technology user experience: `getByRole` (the primary choice — it exposes name/role/state exactly as the accessibility tree does), `getByLabelText` (form fields), `getByPlaceholderText` (secondary, for inputs), `getByText` (non-interactive elements), `getByDisplayValue` (current form value). 2. **Semantic queries** — lower priority than accessible queries; use when an accessible query isn't practical. 3. **Test IDs** — documented as the *last resort*, specifically because a `data-testid` proves nothing about whether a real user (including one using assistive technology) can actually find/use the element. A component test suite that predominantly uses `getByTestId`/`container.querySelector` where `getByRole`/`getByLabelText` would work is a **finding**, not a style nit — it means the tests are not verifying the thing that actually matters (can a user, including an AT user, interact with this), and it silently permits an accessibility regression (e.g. a missing accessible name) to pass a test that only checked a test-id existed. Custom test-id queries built via `queryHelpers.queryByAttribute` are supported for a project's own attribute convention (e.g. `data-test-id` instead of `data-testid`), but the priority order still applies — a custom test-id query is still a last resort, not a replacement for role/label queries. **Vitest coverage thresholds:** Vitest's `coverage.thresholds` config supports global thresholds (`lines`, `functions`, `branches`, `statements`), per-glob-pattern thresholds (each pattern requiring `perFile: true` explicitly to apply per-file, since patterns do not inherit the top-level `perFile` setting), and negative-number thresholds meaning "no more than N uncovered items" rather than a percentage. A reviewed config claiming a threshold is enforced should be checked against these actual semantics — e.g. a glob-pattern threshold block without its own `perFile: true` will NOT enforce per-file coverage even if the top-level config has `perFile: true`, because glob patterns do not inherit it. ## Non-negotiable design rules 1. **Assert on rendered/observable behavior, not implementation detail.** A component test that reaches into instance state, calls a private method, or snapshot-diffs unrelated markup churn (e.g. a full-component snapshot that breaks on any unrelated style change) is brittle and doesn't prove user-facing correctness — see the "hollow pyramid" failure mode in `pyramid-shape-and-coverage.md`. 2. **Query by role/label first; test-id is the last resort, not the default.** Apply the Testing Library priority order above as a hard review criterion, not a suggestion. 3. **Mock at the network boundary, not inside application logic.** Component/integration tests should mock the API layer (MSW, or an injected fetch client) rather than mocking internal application functions, so the test still exercises real component logic/wiring and doesn't silently pass after a refactor that keeps the mocked function's shape but breaks real integration. 4. **Coverage thresholds are a floor, not a target.** A threshold catches *regression below a floor*; it does not prove the newly-added code at 100% coverage actually asserts on error/loading states (see `pyramid-shape-and-coverage.md`). Do not treat "coverage threshold met" as equivalent to "critical path covered." 5. **Watch-mode/local dev ergonomics should not leak into CI config.** Settings tuned for fast local iteration (e.g. skipping type-checking, relaxed isolation) need an explicit, separate CI config — verify the CI config path, don't assume local config is what runs in the pipeline. ## Response discipline When reviewing unit/component tests, cite the specific query/mock/assertion pattern found (with file reference), state which Testing Library priority tier the primary queries fall into, and label whether a coverage-threshold or config-semantics claim is `documentation-based` (grounded via Context7/official docs this session) or `inference`.
-
-
metadata.json 1.2 KB
{ "id": "frontend-testing-strategy-review", "name": "Frontend Testing Strategy Review", "type": "skill", "provider": "frontend", "harnesses": [ "claude-code", "cursor", "codex", "gemini", "kiro", "other" ], "summary": "Reviews unit/component/integration/E2E test-pyramid shape, coverage of critical user journeys, and flaky-test governance for a frontend codebase using Vitest/Jest, Testing Library, and Playwright/Cypress, loading framework-specific reference guidance only when needed.", "source_type": "original", "official_docs": [ "https://playwright.dev/docs/best-practices", "https://vitest.dev/guide/", "https://testing-library.com/docs/queries/about/#priority", "https://docs.cypress.io/app/core-concepts/best-practices" ], "security_notes": "Never accept or request real credentials/session tokens for test fixtures; require synthetic data or mocked network layers (MSW, cy.intercept, Playwright route interception). Flag any test-only auth/CSRF bypass flag that could ship enabled in a production build as a hard-gate finding.", "last_verified": "2026-07-02", "path": "skills/frontend/frontend-testing-strategy-review", "author": "github: VincentChuWaiChow", "version": "0.1.0" } -
SKILL.md 6.3 KB
--- name: frontend-testing-strategy-review description: Reviews frontend test-pyramid shape, critical-path coverage, and flaky-test governance across unit, component, integration, and E2E layers (Vitest/Jest, Testing Library, Playwright/Cypress), loading framework references only when the task needs them. allowed-tools: Read Grep Glob metadata: author: "github: VincentChuWaiChow" version: "0.1.0" updated: "2026-07-02" category: delivery --- # Frontend Testing Strategy Review ## Purpose A green CI badge does not prove a critical user journey works. This skill exists to separate "tests exist" from "tests exercise the actual failure modes that matter" — test-pyramid shape, critical-path coverage, and flaky-test governance — without dumping every testing-library's full API surface into every review. ## When to use Use this skill when the user asks to: - review or design a frontend test strategy across unit, component, integration, and E2E layers, - diagnose why a test suite is slow, flaky, or not catching regressions, - decide whether a given assertion belongs at the unit, component, or E2E layer, - audit critical-user-journey coverage (checkout, auth, forms) before a release, - evaluate a proposed test-framework migration (Jest to Vitest, Cypress to Playwright). ## Context7 Documentation Protocol Test-runner and testing-library APIs (retry semantics, coverage-threshold config, query priority, selector strategy) change across majors and are documented, not folklore — never assert a flag, config shape, or "best practice" from memory. 1. Call `ToolSearch` with query `"context7"` (or `"select:mcp__Context7__resolve-library-id,mcp__Context7__query-docs"`) to load the Context7 tools if not already loaded in this session. 2. Call `mcp__Context7__resolve-library-id` for the framework in question: `/microsoft/playwright` for Playwright, `/vitest-dev/vitest` for Vitest, `/testing-library/testing-library-docs` for Testing Library, `/cypress-io/cypress-documentation` for Cypress. Prefer these resolved IDs over guessing a library name. 3. Call `mcp__Context7__query-docs` for the specific claim in question — e.g. "coverage threshold configuration", "web-first assertion retry behavior", "getByRole query priority", "avoid fixed waits" — before stating it as fact. Do this per review, not once from a prior session's memory. 4. Prefer the official docs URLs in `official_docs` for primary normative statements (e.g. exact CLI flags, exact config shape); use Context7 to ground and cross-check the claim before writing it into a finding. 5. If Context7 is unavailable or returns no relevant match, fall back to the `official_docs` URLs and mark the claim `documentation-based (Context7 unavailable)` rather than presenting it as freshly verified. 6. Never invent a config key, CLI flag, matcher name, or API method that no queried source confirms. ## Lean operating rules - Classify the pyramid shape first (unit:component:E2E ratio, counted from actual test files/CI artifacts, not the user's description of it) before recommending any specific test; a shape problem needs a rebalancing plan, not one more test. - Treat coverage percentage as a weak signal; require evidence the suite asserts on error states, loading states, and accessibility tree for critical paths, not just the happy-path DOM. A file at 100% line coverage with no error-path assertion is not "well tested." - Treat flaky-test quarantine (`.skip`, `test.fixme`, `it.skip`, `describe.skip`, Cypress `{retries: N}` used to paper over root cause) as a tracked liability with an owner and expiry, never a silent, permanent state. - Prefer official, version-specific docs (via Context7) over memory for any framework API claim — Playwright, Vitest, Cypress, and Testing Library APIs and retry/wait semantics change across majors. - Never request or accept real user credentials, session cookies, or production API keys as test fixtures; require synthetic data or network mocking (MSW, `cy.intercept`, Playwright route interception). - Do not recommend a framework migration (Jest→Vitest, Cypress→Playwright) without a measured baseline (suite duration, ESM/native-TS support gap, current flake rate) and a rollback path; framework preference alone is not a migration justification. - Distinguish "flaky because of the test" (missing await, race on network mock, unstable selector) from "flaky because of the app" (real timing bug, unhandled async state) — the fix differs and misdiagnosis just hides a production bug behind a retry. - Load references only for the layer/framework in scope; do not load the Playwright reference for a pure Vitest unit-coverage question, and do not load the pyramid-shape reference for a single flaky-test triage. ## References Load these only when needed: - [Test pyramid shape and critical-path coverage](references/pyramid-shape-and-coverage.md) — use when classifying unit:component:E2E ratio, deciding which layer an assertion belongs at, or auditing critical-user-journey coverage. - [Flaky-test diagnosis and governance](references/flaky-test-governance.md) — use when triaging a flaky or slow suite, reviewing quarantine (`.skip`/`fixme`) usage, or setting a flake-tracking policy. - [Playwright and Cypress E2E review](references/e2e-framework-review.md) — use for E2E-layer specifics: locator/selector strategy, web-first assertions and retry semantics, test isolation, and Jest/Cypress-to-Playwright migration evidence requirements. - [Vitest and Testing Library unit/component review](references/unit-component-framework-review.md) — use for unit/component-layer specifics: query priority (`getByRole` over test IDs), mocking boundaries, and coverage-threshold configuration. ## Response minimum Return, at minimum: - the test-pyramid shape observed (unit/component/E2E counts or ratio, with the file/CI-artifact evidence it came from, not an estimate), - critical-user-journey coverage gaps, cited to a specific file/line or CI-artifact, - flaky-test inventory status for any `.skip`/`fixme`/retry-suppressed test found (owner, expiry, or "untracked liability"), - a minimal, prioritized diff-level test plan (not a full-suite rewrite), - evidence level (`live evidence`, `repo evidence`, `documentation-based`, `inference`) for every claim, and explicit flag when a claim is `documentation-based (Context7 unavailable)`.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.