Claude Cursor Skill

improving-tests

Improve test design, speed, and coverage with behavior-focused tests,

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download alexei-led-cc-thingz-src_skills_improving-tests-ce56bb4.zip · 8 KB
Part of alexei-led/cc-thingz — 91 skills

Install

skills CLI npx skills add https://github.com/alexei-led/cc-thingz/tree/master/src/skills/improving-tests
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install alexei-led-cc-thingz@llmmart
Git git clone https://github.com/alexei-led/cc-thingz.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole alexei-led/cc-thingz collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Test Improvement

Make tests catch real behavior regressions without blocking safe changes. Suite latency is a quality attribute; coverage is a signal, not the goal.

Without write access, return proposed changes (file, change, reason) instead of applying them.

Read the reference for each language in scope; it covers test patterns and suite speed. Use the matching writing-<lang> skill for toolchain commands.

  • C#: references/csharp.md
  • Go: references/go.md
  • Java/Kotlin: references/java-kotlin.md
  • Python: references/python.md
  • Rust: references/rust.md
  • TypeScript/JavaScript: references/typescript.md
  • Browser, Playwright, HTMX: references/web.md

Modes

If the mode is missing, ask which one:

  • review: find weak, duplicate, brittle, missing, slow, or flaky tests.
  • refactor: simplify tests without changing covered behavior.
  • coverage: add tests for uncovered business behavior and error paths.
  • tdd: one red-green-refactor slice at a time.
  • performance: cut suite latency without weakening behavior.
  • full: all of the above.

Rules

  • Ask before adding a test framework, runner plugin, or tool. Learn the framework, helpers, and conventions from nearby tests and follow them first.
  • Change production code only inside an approved TDD slice. If no safe behavior seam exists, stop and report what production change would create one.
  • Test through the contract users or adjacent modules rely on: public module, API, CLI, component, or service. Use an integration seam when behavior depends on real wiring (database, filesystem, HTTP, serialization, config); a unit seam when behavior is pure and cheap.
  • Mock only system boundaries (network, clock, randomness, filesystem, subprocesses, external services). Prefer real collaborators or in-memory fakes for domain code.
  • Assert behavior, not private helpers, call counts, or layout. Business-critical arguments get exact matches.
  • Parameterize cases that share setup and assertions; keep separate tests when that reads clearer.
  • Delete shallow or duplicate tests once stronger boundary tests cover the behavior.
  • Characterization tests before risky changes to legacy code capture current visible behavior, quirks included, at the public boundary.
  • TDD: each slice has one test that failed for the expected reason before the smallest passing code, and refactoring happens only while green. No bulk suites for imagined behavior.
  • If a code-graph tool (GitNexus, codegraph) is installed and fresh, use it to find affected flows and high fan-in code that needs regression coverage.

Speed

Performance work records the same command's wall time before and after, names the bottleneck (discovery, import or compile, setup, test body, external boundary, runner config, or parallel balance), and leaves a guard against regression (durations output, per-test ceiling, slow marker, or focused command).

  • Remove waste before removing checks: real sleeps, real external I/O, repeated setup, expensive imports, broad discovery, coverage-on-default.
  • Keep coverage, race, mutation, browser, live, and end-to-end modes off the default loop; run them in their own tier.
  • Parallelism needs isolated state: per-test or per-worker ports, temp dirs, databases, and filenames. Treat parallel-only failures as isolation bugs.
  • Never make a number look better by hiding failures, deleting edge cases, or skipping fast tests.

Done when the relevant build/test/lint checks pass on what you changed, or you name each check that did not run and why.

Report

Mode, tests changed, key changes (path:line — change), coverage and timing before → after when measured, and each check with pass, fail, or skipped and why.

Platform additions

No target-specific additions.

Files (cc-thingz)
  • .agentbundler
    • targets
      • claude.json 1.5 KB
        {
          "bodyPatch": {
            "mode": "sections",
            "sections": [
              {
                "headingPath": ["Test Improvement", "Platform additions"],
                "body": "\n### Arguments\n\n`$ARGUMENTS` may name the mode: `review`, `refactor`, `coverage`, `tdd`, `performance`, or `full`.\n\n### Claude tools\n\n- Missing mode, or approval to add a framework or tool: ask with `AskUserQuestion`.\n- Track sessions longer than two steps with `TaskCreate` and `TaskUpdate`.\n- Search small scopes directly; spawn read-only `reviewer` agents only for broad or mixed-language audits.\n"
              }
            ]
          },
          "frontmatterPatch": {
            "allowed-tools": [
              "Task",
              "TaskOutput",
              "TaskCreate",
              "TaskUpdate",
              "TaskList",
              "Read",
              "Grep",
              "Glob",
              "LS",
              "Edit",
              "Write",
              "AskUserQuestion",
              "Bash(go test *)",
              "Bash(go tool *)",
              "Bash(golangci-lint *)",
              "Bash(pytest *)",
              "Bash(uv run pytest *)",
              "Bash(bun test *)",
              "Bash(bun run *)",
              "Bash(npm test *)",
              "Bash(pnpm test *)",
              "Bash(yarn test *)",
              "Bash(vitest *)",
              "Bash(jest *)",
              "Bash(npx vitest *)",
              "Bash(npx jest *)",
              "Bash(node --test *)",
              "Bash(npx playwright *)",
              "Bash(bunx playwright *)",
              "Bash(cargo test *)",
              "Bash(cargo nextest *)",
              "Bash(dotnet test *)",
              "Bash(./gradlew *)",
              "Bash(./mvnw *)"
            ],
            "argument-hint": "[review|refactor|coverage|tdd|performance|full]",
            "context": "fork",
            "user-invocable": true
          }
        }
        
  • references
    • csharp.md 733 B
      # C# /.NET Tests
      
      Use writing-csharp for toolchain commands. Follow the project's xUnit, NUnit, or
      MSTest style.
      
      - Iterate on the nearest test project; run the containing solution when the change crosses projects or no test project is obvious.
      - Async tests return `async Task`; never `async void`.
      - ASP.NET Core: use the project's existing host or `WebApplicationFactory` harness before adding one.
      - EF-backed code: test through the real persistence seam (SQLite in-memory or a test container) instead of mocking LINQ providers.
      - Cover nullability, cancellation, and permission paths when the code handles them.
      - Replace sleeps with deterministic synchronization, a controllable `TimeProvider`, or polling with a hard timeout.
      
    • go.md 1.3 KB
      # Go Tests
      
      Use writing-go for toolchain commands.
      
      ## Patterns
      
      - Table-driven tests with a descriptive `name` and `t.Run(tc.name, ...)`; split large or unrelated matrices.
      - With testify, `require` for preconditions that must stop the test, `assert` for independent checks. Check errors before dereferencing results.
      - `mock.Anything` only for contexts, loggers, and true don't-care values. Use mockery only if the repo already does.
      - Setup that needs many mocks or globals is a design smell to report, not to paper over.
      
      ## Speed
      
      - Package-list mode (`go test ./pkg/...`) caches passing results; bare `go test` does not. Avoid `-count=1` unless you are chasing a flake or side effect. `-run`, `-short`, `-parallel`, `-failfast`, `-timeout`, and `-v` stay cacheable.
      - Tests that read env vars or module files miss the cache when those change.
      - `t.Parallel()` only with isolated state. `t.Setenv` is process-wide and panics under parallel tests or parallel ancestors.
      - Tune `-parallel` only after measuring; above CPU count it often slows down.
      - Replace sleeps with channels, contexts, fake clocks, or `testing/synctest` when the module's Go version supports it.
      - Gate real databases, networks, and containers behind `testing.Short()` or build tags.
      - `-race`, coverage profiles, and benchmarks are separate tiers, not the edit loop. Race failures are blocking.
      
    • java-kotlin.md 981 B
      # Java and Kotlin Tests
      
      Use writing-java-kotlin for toolchain commands. Identify the framework (JUnit 5,
      JUnit 4, TestNG, Kotest), assertion library, and mock library from nearby tests.
      
      - Filter to one module and class in the edit loop (`:module:test --tests`, `-pl module -Dtest=`); run the full build only at the end or for cross-module changes.
      - No full `@SpringBootTest` for pure logic. Use a slice (`@WebMvcTest`, `@DataJpaTest`) or no context.
      - Testcontainers, real databases, and browser tests belong in a separate Gradle or Maven profile.
      - Coroutines: `runTest` with a test scheduler; not `runBlocking` in new tests.
      - Do not mix Mockito and MockK in one test class. `verify` only side effects that matter; no `any()` on business-critical arguments.
      - `@ParameterizedTest`/`@MethodSource` or Kotest `forAll` for case matrices; AssertJ-style `assertThat` for clearer failures.
      - Shared static caches or singleton state across tests cause flakes; reset or isolate them.
      
    • python.md 2 KB
      # Python Tests
      
      Use writing-python for toolchain commands. Read `conftest.py` before changing
      fixtures.
      
      ## Patterns
      
      - `@pytest.mark.parametrize` with readable IDs (`pytest.param(..., id="empty-input")`).
      - Fixture scope as narrow as correctness needs. Reuse fixtures to hide noise, not behavior.
      - Patch where the object is used, not where it is defined. Use `autospec`/`create_autospec` on important boundaries and `AsyncMock` for async.
      - Test async code through its async public behavior with the project's configured async plugin.
      - Import errors and collection failures are blocking.
      
      ## Speed
      
      - Find the bottleneck with `--durations=20 --durations-min=0.5`; when collection dominates, `python -X importtime -m pytest --collect-only`.
      - `pytest-cov` instruments every line; run it only for coverage work.
      - `pytest-xdist` only if configured or approved. Start with `-n auto --dist worksteal` for uneven durations; `loadscope`/`loadfile`/`loadgroup` only when grouped fixtures save more than the imbalance costs. More workers do not fix time spent waiting on sleeps or I/O.
      - Key shared real resources (ports, databases, tmux sessions, fixed filenames, queues) by `PYTEST_XDIST_WORKER`.
      - Replace `time.sleep`/`asyncio.sleep` with poll-until-condition helpers (small interval, hard timeout), or freeze time with the project's clock tool. Lower test-only timeouts so failure paths are fast too.
      - Session or module fixtures only for immutable setup; for costly mutable artifacts (git repos, populated databases), build once and copy per test.
      - Move heavy import-time work into fixtures. New pytest config: explicit `testpaths` and `--import-mode=importlib`.
      - Stub expensive autouse behavior in unit tests; opt in to the real thing where it is under test.
      - Mark slow tiers (`integration`, `e2e`, `live`, `llm`) and keep the default command on the fast tier. Use `--last-failed` in the edit loop.
      - Mature suites can add a hook that fails non-slow tests above an env-tunable wall-time ceiling.
      
    • rust.md 1 KB
      # Rust Tests
      
      Use writing-rust for toolchain commands.
      
      - Unit tests in `#[cfg(test)] mod tests` for private logic worth covering; `tests/*.rs` for public crate or CLI behavior; doctests for copyable public examples.
      - nextest (when the project uses it) does not run doctests; add `cargo test --doc`.
      - Filter by package and test name in the edit loop; keep full workspace runs, coverage, Miri, and benchmarks as separate tiers. Never `cargo clean` as a routine step.
      - Assert error variants or messages only when callers depend on them.
      - Prefer small fakes behind local traits over mock frameworks. Add crates such as `assert_cmd`, `tempfile`, `insta`, `proptest`, `mockall`, or `wiremock` only when already used or justified.
      - Async tests use the project's macro (`#[tokio::test]`); no blocking mutex guard across `.await`; channels, barriers, or paused time instead of sleeps.
      - Normalize volatile snapshot output (paths, timestamps, ordering).
      - For unsafe or concurrency-sensitive code, run configured Miri, loom, or sanitizers and report when unavailable.
      
    • typescript.md 1.6 KB
      # TypeScript and JavaScript Tests
      
      Use writing-typescript for toolchain commands and the project's runner.
      
      ## Patterns
      
      - `it.each`/`test.each` with object cases for multi-field inputs.
      - React: test user-visible behavior with Testing Library. Query by role, then label, then text; test ID last. Prefer `user-event`; use async queries for async state. Cover loading, error, empty, and permission states that matter.
      - Snapshots only for complex, stable output.
      - Network: MSW or the project's HTTP fakes. Type mocks (`vi.mocked()`) so contract drift fails to compile. Restore mocks per project convention.
      
      ## Speed
      
      - Use the runner's related or changed-file mode in the edit loop.
      - Coverage and Jest `--detectOpenHandles` (which forces serial execution) are diagnostics, not the default loop.
      - Tune worker count by measurement: transform cost, memory, DOM environments, and database limits often make fewer workers faster.
      - Disable per-file isolation only when tests clean their globals. Never trade determinism for speed.
      - When the runner reports transform, import, setup, or environment time, cut the largest bucket first. Keep global setup and preloads minimal; they run for every focused test.
      - Explicit test roots and ignores, so discovery skips build output, fixtures, vendor trees, and e2e suites.
      - Pure logic runs in the Node environment, not a DOM one. Scope DOM environments to the files that need them; do not swap DOM implementations unless the project accepts the fidelity loss.
      - Fake timers or poll-until-condition helpers instead of sleeps. Reset only the timers, mocks, and modules a test changed; blanket resets dominate tiny tests.
      
    • web.md 839 B
      # Browser and HTMX Tests
      
      Use browser-automation for running flows interactively. If Playwright is not
      configured, say so instead of guessing.
      
      - `playwright test --list` is safe discovery; run browser tests only when needed. An empty list means no coverage, not unknown coverage.
      - Locators: `getByRole`, then `getByLabel`, then `getByText`; `getByTestId` only when semantics cannot express it. No XPath, deep CSS, or generated IDs.
      - Playwright auto-waits: assert with `expect(locator).toBeVisible()` instead of fixed timeouts.
      - One user-visible flow per test. Delete page-load smoke tests once flow tests cover the page.
      - HTMX: assert the swapped, added, or removed element, not internal events.
      - Route-mock external services instead of calling them.
      - Classify failures as locator, timing, app behavior, test logic, or environment.
      
  • SKILL.md 4.1 KB
    ---
    description: Improve test design, speed, and coverage with behavior-focused tests,
      useful seams, characterization tests, TDD, and test refactoring. Use when improving
      tests, optimizing slow suites, adding coverage, refactoring brittle tests, removing
      test waste, or working test-first. NOT for fixing production bugs (use fixing-code),
      or reviewing non-test code quality (use reviewing-code).
    name: improving-tests
    ---
    
    # Test Improvement
    
    Make tests catch real behavior regressions without blocking safe changes. Suite
    latency is a quality attribute; coverage is a signal, not the goal.
    
    Without write access, return proposed changes (file, change, reason) instead of
    applying them.
    
    Read the reference for each language in scope; it covers test patterns and
    suite speed. Use the matching `writing-<lang>` skill for toolchain commands.
    
    - C#: `references/csharp.md`
    - Go: `references/go.md`
    - Java/Kotlin: `references/java-kotlin.md`
    - Python: `references/python.md`
    - Rust: `references/rust.md`
    - TypeScript/JavaScript: `references/typescript.md`
    - Browser, Playwright, HTMX: `references/web.md`
    
    ## Modes
    
    If the mode is missing, ask which one:
    
    - `review`: find weak, duplicate, brittle, missing, slow, or flaky tests.
    - `refactor`: simplify tests without changing covered behavior.
    - `coverage`: add tests for uncovered business behavior and error paths.
    - `tdd`: one red-green-refactor slice at a time.
    - `performance`: cut suite latency without weakening behavior.
    - `full`: all of the above.
    
    ## Rules
    
    - Ask before adding a test framework, runner plugin, or tool. Learn the framework, helpers, and conventions from nearby tests and follow them first.
    - Change production code only inside an approved TDD slice. If no safe behavior seam exists, stop and report what production change would create one.
    - Test through the contract users or adjacent modules rely on: public module, API, CLI, component, or service. Use an integration seam when behavior depends on real wiring (database, filesystem, HTTP, serialization, config); a unit seam when behavior is pure and cheap.
    - Mock only system boundaries (network, clock, randomness, filesystem, subprocesses, external services). Prefer real collaborators or in-memory fakes for domain code.
    - Assert behavior, not private helpers, call counts, or layout. Business-critical arguments get exact matches.
    - Parameterize cases that share setup and assertions; keep separate tests when that reads clearer.
    - Delete shallow or duplicate tests once stronger boundary tests cover the behavior.
    - Characterization tests before risky changes to legacy code capture current visible behavior, quirks included, at the public boundary.
    - TDD: each slice has one test that failed for the expected reason before the smallest passing code, and refactoring happens only while green. No bulk suites for imagined behavior.
    - If a code-graph tool (GitNexus, codegraph) is installed and fresh, use it to find affected flows and high fan-in code that needs regression coverage.
    
    ## Speed
    
    Performance work records the same command's wall time before and after, names
    the bottleneck (discovery, import or compile, setup, test body, external
    boundary, runner config, or parallel balance), and leaves a guard against
    regression (durations output, per-test ceiling, slow marker, or focused command).
    
    - Remove waste before removing checks: real sleeps, real external I/O, repeated setup, expensive imports, broad discovery, coverage-on-default.
    - Keep coverage, race, mutation, browser, live, and end-to-end modes off the default loop; run them in their own tier.
    - Parallelism needs isolated state: per-test or per-worker ports, temp dirs, databases, and filenames. Treat parallel-only failures as isolation bugs.
    - Never make a number look better by hiding failures, deleting edge cases, or skipping fast tests.
    
    Done when the relevant build/test/lint checks pass on what you changed, or you
    name each check that did not run and why.
    
    ## Report
    
    Mode, tests changed, key changes (`path:line — change`), coverage and timing
    before → after when measured, and each check with pass, fail, or skipped and why.
    
    ## Platform additions
    
    No target-specific additions.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related