Claude
Skill
testing-ops
Cross-language testing strategies and patterns. Triggers on: test pyramid, unit test, integration test, e2e test, TDD, BDD, test coverage, mocking strategy, test doubles, test isolation.
Virus-scanned
Reviewed automatically before listing.
Download
0xdarkmatter-claude-mods-skills_testing-ops-3dfaf0b.zip · 14 KB
Install
skills CLI
npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/testing-ops
Claude Code
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
Git
git clone https://github.com/0xDarkMatter/claude-mods.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Testing Patterns
Universal testing strategies and patterns applicable across languages.
The Test Pyramid
/\
/ \ E2E Tests (few, slow, expensive)
/ \ - Full system tests
/------\ - Real browser/API calls
/ \
/ Integ \ Integration Tests (some)
/ Tests \ - Service boundaries
/--------------\ - Database, APIs
/ \
/ Unit Tests \ Unit Tests (many, fast, cheap)
------------------ - Single function/class
- Mocked dependencies
Test Types
Unit Tests
Scope: Single function/method/class
Speed: Milliseconds
Dependencies: All mocked
When: Every code change
Coverage: 80%+ of codebase
Integration Tests
Scope: Multiple components together
Speed: Seconds
Dependencies: Real databases, mocked external APIs
When: PR/merge, critical paths
Coverage: Key integration points
End-to-End Tests
Scope: Full user journey
Speed: Minutes
Dependencies: Real system (or staging)
When: Pre-deploy, nightly
Coverage: Critical user flows only
Test Naming Convention
test_<unit>_<scenario>_<expected>
Examples:
- test_calculate_total_with_discount_returns_reduced_price
- test_user_login_with_invalid_password_returns_401
- test_order_submit_when_out_of_stock_raises_error
Arrange-Act-Assert (AAA)
def test_calculate_discount():
# Arrange - Set up test data and dependencies
cart = Cart()
cart.add_item(Item(price=100))
discount = Discount(percent=10)
# Act - Execute the code under test
total = cart.calculate_total(discount)
# Assert - Verify the results
assert total == 90
Test Doubles
| Type | Purpose | Example |
|---|---|---|
| Stub | Returns canned data | stub.get_user.returns(fake_user) |
| Mock | Verifies interactions | mock.send_email.assert_called_once() |
| Spy | Records calls, uses real impl | spy.on(service, 'save') |
| Fake | Working simplified impl | FakeDatabase() instead of real DB |
| Dummy | Placeholder, never used | null object for required param |
Test Isolation Strategies
Database Isolation
Option 1: Transaction rollback (fast)
- Start transaction before test
- Rollback after test
Option 2: Truncate tables (medium)
- Clear all data between tests
Option 3: Separate database (slow)
- Each test gets fresh database
External Service Isolation
Option 1: Mock at boundary
- Replace HTTP client with mock
Option 2: Fake server
- WireMock, MSW, VCR cassettes
Option 3: Contract testing
- Pact, consumer-driven contracts
What to Test
MUST Test
- Business logic and calculations
- Input validation and error handling
- Security-sensitive code (auth, permissions)
- Edge cases and boundary conditions
SHOULD Test
- Integration points (DB, APIs)
- State transitions
- Configuration handling
AVOID Testing
- Framework internals
- Third-party library behavior
- Simple getters/setters
- Private implementation details
Test Quality Checklist
- Tests are independent (no order dependency)
- Tests are deterministic (no flaky tests)
- Tests are fast (unit < 100ms, integration < 5s)
- Tests have clear names describing behavior
- Tests cover happy path AND error cases
- Tests don't repeat production logic
- Mocks are minimal (only external boundaries)
Additional Resources
./references/tdd-workflow.md- Test-Driven Development cycle./references/mocking-strategies.md- When and how to mock./references/test-data-patterns.md- Fixtures, factories, builders./references/ci-testing.md- Testing in CI/CD pipelines
Scripts
./scripts/coverage-check.sh- Run coverage and fail if below threshold
Files (claude-mods)
-
assets
-
.gitkeep 0 B · in bundle
-
-
references
-
ci-testing.md 6.9 KB
# CI/CD Testing Patterns Testing strategies for continuous integration pipelines. ## Test Pipeline Stages ``` ┌─────────────────────────────────────────────────────────────────┐ │ CI Pipeline │ │ │ │ ┌──────┐ ┌──────┐ ┌───────┐ ┌─────┐ ┌──────┐ ┌────┐│ │ │ Lint │ → │ Unit │ → │ Build │ → │Integ│ → │ E2E │ → │Dep.││ │ │ │ │Tests │ │ │ │Tests│ │Tests │ │ ││ │ └──────┘ └──────┘ └───────┘ └─────┘ └──────┘ └────┘│ │ 1m 2-5m 1-3m 5-10m 10-30m - │ │ │ │ ◄─────── Fast Feedback ───────► ◄─── Comprehensive ──────► │ └─────────────────────────────────────────────────────────────────┘ ``` ## GitHub Actions Example ```yaml name: CI on: push: branches: [main] pull_request: branches: [main] jobs: lint: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: actions/setup-python@v5 with: python-version: '3.11' - name: Lint run: | pip install ruff ruff check . unit-tests: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: actions/setup-python@v5 with: python-version: '3.11' - name: Install dependencies run: pip install -e .[test] - name: Run unit tests run: pytest tests/unit -v --cov=src --cov-report=xml - name: Upload coverage uses: codecov/codecov-action@v4 integration-tests: needs: unit-tests runs-on: ubuntu-latest services: postgres: image: postgres:15 env: POSTGRES_PASSWORD: postgres ports: - 5432:5432 options: >- --health-cmd pg_isready --health-interval 10s --health-timeout 5s --health-retries 5 steps: - uses: actions/checkout@v4 - uses: actions/setup-python@v5 with: python-version: '3.11' - name: Run integration tests env: DATABASE_URL: postgres://postgres:postgres@localhost:5432/test run: pytest tests/integration -v e2e-tests: needs: integration-tests runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - name: Run E2E tests run: | docker-compose up -d pytest tests/e2e -v docker-compose down ``` ## Test Parallelization ### pytest-xdist ```yaml - name: Run tests in parallel run: pytest -n auto # Use all available CPUs - name: Run with specific workers run: pytest -n 4 # 4 parallel workers ``` ### Matrix Testing ```yaml jobs: test: strategy: matrix: python-version: ['3.9', '3.10', '3.11'] os: [ubuntu-latest, macos-latest] runs-on: ${{ matrix.os }} steps: - uses: actions/setup-python@v5 with: python-version: ${{ matrix.python-version }} - run: pytest ``` ### Sharded Tests ```yaml jobs: test: strategy: matrix: shard: [1, 2, 3, 4] steps: - name: Run test shard run: pytest --shard-id=${{ matrix.shard }} --num-shards=4 ``` ## Caching for Speed ```yaml - name: Cache pip packages uses: actions/cache@v4 with: path: ~/.cache/pip key: ${{ runner.os }}-pip-${{ hashFiles('**/requirements*.txt') }} restore-keys: | ${{ runner.os }}-pip- - name: Cache pytest uses: actions/cache@v4 with: path: .pytest_cache key: pytest-${{ github.sha }} restore-keys: pytest- ``` ## Flaky Test Handling ### Retry Mechanism ```yaml - name: Run tests with retry uses: nick-fields/retry@v3 with: timeout_minutes: 10 max_attempts: 3 command: pytest tests/e2e ``` ### pytest-rerunfailures ```bash # Rerun failed tests up to 3 times pytest --reruns 3 --reruns-delay 1 ``` ### Quarantine Flaky Tests ```python @pytest.mark.flaky(reruns=3, reruns_delay=2) def test_sometimes_fails(): # This test is known to be flaky pass @pytest.mark.skip(reason="Flaky - investigating") def test_quarantined(): pass ``` ## Test Reports ### JUnit XML ```yaml - name: Run tests run: pytest --junitxml=results.xml - name: Publish Test Results uses: dorny/test-reporter@v1 if: always() with: name: Test Results path: results.xml reporter: java-junit ``` ### Coverage Reports ```yaml - name: Run with coverage run: pytest --cov=src --cov-report=xml --cov-report=html - name: Upload coverage to Codecov uses: codecov/codecov-action@v4 with: files: ./coverage.xml fail_ci_if_error: true - name: Coverage comment on PR uses: py-cov-action/python-coverage-comment-action@v3 ``` ## Branch Protection Rules ```yaml # Require tests to pass before merge # Settings → Branches → Branch protection rules Required status checks: - lint - unit-tests - integration-tests Require branches to be up to date: Yes ``` ## Test Selection ### Changed Files Only ```yaml - name: Get changed files id: changed uses: tj-actions/changed-files@v41 with: files: | src/** tests/** - name: Run affected tests if: steps.changed.outputs.any_changed == 'true' run: pytest tests/ -v ``` ### Skip Expensive Tests ```yaml - name: Quick tests on PR if: github.event_name == 'pull_request' run: pytest -m "not slow and not e2e" - name: Full tests on main if: github.ref == 'refs/heads/main' run: pytest ``` ## Secrets in Tests ```yaml - name: Run tests with secrets env: API_KEY: ${{ secrets.TEST_API_KEY }} DATABASE_URL: ${{ secrets.TEST_DATABASE_URL }} run: pytest tests/integration # Use environment for sensitive tests jobs: integration: environment: testing # Requires approval steps: - run: pytest tests/integration ``` ## Best Practices 1. **Fast feedback first** - Run linting and unit tests before slow tests 2. **Fail fast** - Stop pipeline on first failure (`pytest -x`) 3. **Parallel when possible** - Use matrix builds and xdist 4. **Cache aggressively** - Pip, node_modules, docker layers 5. **Keep tests deterministic** - No reliance on external state 6. **Isolate flaky tests** - Quarantine or fix, don't ignore 7. **Report clearly** - Use test reporters and coverage comments 8. **Secure secrets** - Never log, use GitHub secrets -
mocking-strategies.md 7 KB
# Mocking Strategies When, what, and how to mock effectively. ## When to Mock ### ALWAYS Mock - External HTTP APIs - Databases in unit tests - File system in unit tests - Time-dependent operations - Random number generators - Email/SMS services ### SOMETIMES Mock - Internal services (depends on test type) - Caches - Message queues ### NEVER Mock - The code under test itself - Simple value objects - Pure functions without side effects ## The Testing Boundary ``` ┌─────────────────────────────────────────────────────┐ │ Your Application │ │ │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ │ Business │ -> │ Service │ -> │Repository│ │ │ │ Logic │ │ Layer │ │ Layer │ │ │ └──────────┘ └──────────┘ └──────────┘ │ │ │ │ │ ▼ │ │ ┌─────────────────┐ │ │ │ BOUNDARY │ │ │ │ (Mock Here!) │ │ │ └─────────────────┘ │ │ │ │ └────────────────────────────────────────│────────────┘ ▼ ┌─────────────────┐ │External Services│ │ - Database │ │ - APIs │ │ - File System │ └─────────────────┘ ``` ## Mock Patterns ### Stub Pattern (Canned Responses) ```python # Use when you need predictable return values def test_get_user_returns_user_data(mocker): mock_db = mocker.patch("app.database.get_user") mock_db.return_value = {"id": 1, "name": "Alice"} result = user_service.get_user(1) assert result["name"] == "Alice" ``` ### Mock Pattern (Verify Interactions) ```python # Use when you need to verify calls were made def test_order_sends_confirmation_email(mocker): mock_email = mocker.patch("app.email.send") order_service.place_order(user_id=1, items=[...]) mock_email.assert_called_once_with( to="user@example.com", subject="Order Confirmation", body=mocker.ANY ) ``` ### Spy Pattern (Record + Real Implementation) ```python # Use when you want real behavior but need to track calls def test_caching_reduces_db_calls(mocker): spy = mocker.spy(database, "query") # First call hits database result1 = cached_service.get_data("key") # Second call should use cache result2 = cached_service.get_data("key") assert spy.call_count == 1 # Only called once assert result1 == result2 ``` ### Fake Pattern (Simplified Implementation) ```python # Use for complex dependencies that need real behavior class FakeEmailService: def __init__(self): self.sent_emails = [] def send(self, to, subject, body): self.sent_emails.append({ "to": to, "subject": subject, "body": body }) def test_order_workflow(fake_email): order_service = OrderService(email_service=fake_email) order_service.place_order(user_id=1, items=[...]) assert len(fake_email.sent_emails) == 1 assert "Order Confirmation" in fake_email.sent_emails[0]["subject"] ``` ## Mock Anti-Patterns ### Over-Mocking ```python # BAD - Mocking everything def test_order_total(mocker): mock_cart = mocker.Mock() mock_item = mocker.Mock() mock_item.price = 100 mock_cart.items = [mock_item] mock_cart.calculate_total.return_value = 100 # ?! # This tests nothing - we mocked the thing we're testing! assert mock_cart.calculate_total() == 100 # GOOD - Only mock boundaries def test_order_total(): cart = Cart() cart.add_item(Item(price=100)) assert cart.calculate_total() == 100 ``` ### Mocking Too Deep ```python # BAD - Mocking internal implementation def test_process_order(mocker): mocker.patch("app.order.Order._validate_inventory") mocker.patch("app.order.Order._calculate_tax") mocker.patch("app.order.Order._apply_discount") # Now coupled to internal implementation! # GOOD - Mock at the boundary def test_process_order(mocker): mocker.patch("app.inventory_service.check_availability") mocker.patch("app.tax_service.calculate") # External services, not internal methods ``` ### Mock Setup Longer Than Test ```python # BAD - Test is mostly setup def test_user_registration(mocker): mock_db = mocker.patch("app.db") mock_email = mocker.patch("app.email") mock_validator = mocker.patch("app.validator") mock_logger = mocker.patch("app.logger") mock_db.create_user.return_value = {"id": 1} mock_email.send.return_value = True mock_validator.validate.return_value = [] # ... 20 more lines of setup # The actual test is tiny result = register_user("test@example.com") assert result.success # GOOD - Use fixtures and factories @pytest.fixture def registration_mocks(mocker): return RegistrationMocks(mocker) # Encapsulate setup def test_user_registration(registration_mocks): result = register_user("test@example.com") assert result.success ``` ## Dependency Injection for Testability ```python # Hard to test - creates own dependencies class OrderService: def __init__(self): self.db = Database() # Can't mock! self.email = EmailService() # Easy to test - dependencies injected class OrderService: def __init__(self, db: Database, email: EmailService): self.db = db self.email = email # Test with mocks def test_order_service(mocker): mock_db = mocker.Mock() mock_email = mocker.Mock() service = OrderService(db=mock_db, email=mock_email) ``` ## Contract Testing When mocking external services, verify your mocks match reality: ```python # Record real responses @pytest.fixture(scope="session") def vcr_config(): return {"record_mode": "once"} @pytest.mark.vcr() def test_github_api(): response = github_client.get_user("octocat") assert response["login"] == "octocat" # Or use contract testing (Pact) def test_user_service_contract(): pact.given("user exists").upon_receiving( "a request for user" ).with_request( method="GET", path="/users/1" ).will_respond_with( status=200, body={"id": 1, "name": Like("string")} ) ``` -
tdd-workflow.md 5.2 KB
# Test-Driven Development (TDD) Red-Green-Refactor cycle for quality code. ## The TDD Cycle ``` ┌─────────────────────────────────────────┐ │ │ │ ┌─────┐ ┌─────┐ ┌─────────┐ │ │ │ RED │ -> │GREEN│ -> │REFACTOR │ ──┐│ │ └─────┘ └─────┘ └─────────┘ ││ │ ^ │ │ │ └───────────────────────────────┘ │ │ │ └─────────────────────────────────────────┘ ``` ### 1. RED: Write a Failing Test ```python # Start with the test def test_calculate_discount_applies_percentage(): cart = Cart() cart.add_item(Item(price=100)) total = cart.calculate_total(discount_percent=10) assert total == 90 # Fails - function doesn't exist yet ``` ### 2. GREEN: Make It Pass (Minimal Code) ```python # Write minimal code to pass class Cart: def __init__(self): self.items = [] def add_item(self, item): self.items.append(item) def calculate_total(self, discount_percent=0): total = sum(item.price for item in self.items) return total * (1 - discount_percent / 100) ``` ### 3. REFACTOR: Improve the Code ```python # Clean up while tests pass class Cart: def __init__(self): self._items: list[Item] = [] def add_item(self, item: Item) -> None: self._items.append(item) @property def subtotal(self) -> Decimal: return sum(item.price for item in self._items) def calculate_total(self, discount_percent: int = 0) -> Decimal: discount_multiplier = Decimal(100 - discount_percent) / 100 return self.subtotal * discount_multiplier ``` ## TDD Rules ### Three Laws of TDD 1. **Don't write production code** until you have a failing test 2. **Write only enough test** to fail (compilation counts) 3. **Write only enough production code** to pass the test ### Test Size Rules ``` Tests should be: - Fast (< 100ms each) - Isolated (no shared state) - Repeatable (same result every time) - Self-validating (pass/fail, no manual inspection) - Timely (written before production code) ``` ## TDD Workflow Example ### Step 1: List Test Cases ``` Feature: Shopping Cart Discount Test cases: [ ] Empty cart returns 0 [ ] Single item returns item price [ ] Multiple items returns sum [ ] Percentage discount applied correctly [ ] Maximum discount capped at 50% [ ] Negative discount treated as 0 ``` ### Step 2: Start with Simplest Test ```python def test_empty_cart_returns_zero(): cart = Cart() assert cart.calculate_total() == 0 ``` ### Step 3: Implement and Move to Next ```python # After passing, add next test def test_single_item_returns_price(): cart = Cart() cart.add_item(Item(price=50)) assert cart.calculate_total() == 50 ``` ### Step 4: Build Up Complexity ```python def test_discount_capped_at_50_percent(): cart = Cart() cart.add_item(Item(price=100)) total = cart.calculate_total(discount_percent=75) assert total == 50 # Capped at 50% max discount ``` ## When to Use TDD ### Good For: - Business logic - Algorithms - Data transformations - API contracts - Complex conditionals ### Less Suitable For: - UI/visual elements - Exploratory prototyping - One-off scripts - Integration with external systems ## TDD Anti-Patterns ### Testing Implementation Details ```python # BAD - Tests internal state def test_cart_has_items_list(): cart = Cart() cart.add_item(Item(price=10)) assert len(cart._items) == 1 # Tests implementation! # GOOD - Tests behavior def test_cart_counts_items(): cart = Cart() cart.add_item(Item(price=10)) assert cart.item_count == 1 # Tests public interface ``` ### Tests That Mirror Code ```python # BAD - Test duplicates implementation def test_calculate_total(): cart = Cart() cart.add_item(Item(price=10)) cart.add_item(Item(price=20)) # This is just reimplementing the function expected = 10 + 20 assert cart.calculate_total() == expected # GOOD - Tests expected outcome def test_calculate_total(): cart = Cart() cart.add_item(Item(price=10)) cart.add_item(Item(price=20)) assert cart.calculate_total() == 30 ``` ## Kata Practice ### String Calculator ``` Create a calculator that takes a string of numbers and returns their sum. Step 1: "" returns 0 Step 2: "1" returns 1 Step 3: "1,2" returns 3 Step 4: Handle unknown number of numbers Step 5: Handle newlines as delimiters: "1\n2,3" returns 6 Step 6: Support custom delimiters: "//;\n1;2" returns 3 Step 7: Throw on negative numbers with message including all negatives ``` ### FizzBuzz ``` Step 1: Return "1" for 1 Step 2: Return "2" for 2 Step 3: Return "Fizz" for 3 Step 4: Return "Buzz" for 5 Step 5: Return "Fizz" for 6 (multiple of 3) Step 6: Return "Buzz" for 10 (multiple of 5) Step 7: Return "FizzBuzz" for 15 (multiple of both) ``` -
test-data-patterns.md 6.4 KB
# Test Data Patterns Strategies for managing test data effectively. ## Fixtures ### Basic Fixture ```python import pytest @pytest.fixture def user(): return User(id=1, name="Test User", email="test@example.com") def test_user_greeting(user): assert user.greeting() == "Hello, Test User!" ``` ### Fixture with Cleanup ```python @pytest.fixture def temp_database(): db = create_test_database() yield db db.drop() # Cleanup after test ``` ### Shared Fixtures (conftest.py) ```python # tests/conftest.py @pytest.fixture(scope="session") def app(): """Application shared across all tests.""" return create_app(testing=True) @pytest.fixture(scope="function") def client(app): """Fresh client for each test.""" return app.test_client() ``` ## Factory Pattern ### Simple Factory ```python def make_user(**overrides): """Factory function for creating test users.""" defaults = { "id": 1, "name": "Test User", "email": "test@example.com", "active": True, } return User(**{**defaults, **overrides}) def test_inactive_user(): user = make_user(active=False) assert not user.can_login() ``` ### Factory Fixture ```python @pytest.fixture def user_factory(): """Factory fixture for creating multiple users.""" created = [] def _create(**overrides): user = make_user(**overrides) created.append(user) return user yield _create # Cleanup for user in created: user.delete() def test_user_comparison(user_factory): user1 = user_factory(name="Alice") user2 = user_factory(name="Bob") assert user1 != user2 ``` ### Factory Boy (Python) ```python import factory from factory import Faker class UserFactory(factory.Factory): class Meta: model = User id = factory.Sequence(lambda n: n + 1) name = Faker("name") email = Faker("email") created_at = Faker("date_time_this_year") # Usage def test_users(): user = UserFactory() admin = UserFactory(role="admin") users = UserFactory.create_batch(10) ``` ## Builder Pattern ```python class UserBuilder: """Fluent builder for test users.""" def __init__(self): self._data = { "id": 1, "name": "Test User", "email": "test@example.com", "role": "user", "active": True, } def with_name(self, name: str) -> "UserBuilder": self._data["name"] = name return self def as_admin(self) -> "UserBuilder": self._data["role"] = "admin" return self def inactive(self) -> "UserBuilder": self._data["active"] = False return self def build(self) -> User: return User(**self._data) # Usage def test_admin_access(): admin = UserBuilder().as_admin().build() assert admin.can_access_admin_panel() def test_inactive_user(): user = UserBuilder().inactive().build() assert not user.can_login() ``` ## Mother Pattern ```python class ObjectMother: """Pre-configured test objects for common scenarios.""" @staticmethod def valid_user() -> User: return User( id=1, name="Valid User", email="valid@example.com", active=True ) @staticmethod def admin_user() -> User: return User( id=2, name="Admin User", email="admin@example.com", role="admin", active=True ) @staticmethod def expired_subscription() -> Subscription: return Subscription( user_id=1, expires_at=datetime.now() - timedelta(days=30), plan="basic" ) # Usage def test_admin_permissions(): admin = ObjectMother.admin_user() assert admin.can_delete_users() ``` ## Fixture Composition ```python @pytest.fixture def address(): return Address(street="123 Main St", city="Test City") @pytest.fixture def user(address): return User(name="Test User", address=address) @pytest.fixture def order(user): return Order(user=user, items=[]) def test_order_address(order): assert order.shipping_address.city == "Test City" ``` ## Data Files ### JSON Fixtures ```python # tests/fixtures/users.json [ {"id": 1, "name": "Alice", "role": "admin"}, {"id": 2, "name": "Bob", "role": "user"} ] # tests/conftest.py @pytest.fixture def sample_users(): with open("tests/fixtures/users.json") as f: return json.load(f) ``` ### YAML Fixtures ```yaml # tests/fixtures/config.yaml database: host: localhost port: 5432 name: test_db users: - id: 1 name: Alice - id: 2 name: Bob ``` ```python @pytest.fixture def config(): with open("tests/fixtures/config.yaml") as f: return yaml.safe_load(f) ``` ## Randomized Data ```python from faker import Faker fake = Faker() def test_user_email_validation(): # Random but valid email email = fake.email() user = User(email=email) assert user.is_valid_email() def test_with_seed(): # Reproducible random data Faker.seed(12345) user = make_user(name=fake.name()) # Same name every time with seed 12345 ``` ## Best Practices ### 1. Keep Fixtures Close to Tests ``` tests/ ├── conftest.py # Shared fixtures ├── unit/ │ ├── conftest.py # Unit test fixtures │ └── test_user.py └── integration/ ├── conftest.py # Integration fixtures └── test_api.py ``` ### 2. Use Descriptive Names ```python # BAD @pytest.fixture def data(): return {...} # GOOD @pytest.fixture def user_with_expired_subscription(): return {...} ``` ### 3. Minimize Fixture Scope ```python # Use function scope (default) unless you have a reason @pytest.fixture(scope="function") # Default def user(): ... # Session scope only for expensive, read-only fixtures @pytest.fixture(scope="session") def database_schema(): ... ``` ### 4. Avoid Test Data Dependencies ```python # BAD - Tests depend on each other def test_create_user(): user = create_user("test@example.com") # User exists in DB after this test def test_get_user(): user = get_user("test@example.com") # Depends on previous test! # GOOD - Each test is independent def test_create_user(db): user = create_user("test@example.com") assert user.email == "test@example.com" def test_get_user(db, user_factory): user_factory(email="test@example.com") # Create own data found = get_user("test@example.com") assert found is not None ```
-
-
scripts
-
coverage-check.sh 3.9 KB
#!/usr/bin/env bash # Run pytest with coverage and exit non-zero when coverage is below a threshold. # # Usage: coverage-check.sh [--threshold N] [--cov TARGET] [pytest-args...] # Input: a pytest project in the current directory; OR CM_COVERAGE_OVERRIDE=PCT # to judge a given percentage offline (test seam — skips pytest entirely) # Output: one verdict line on stdout, e.g. "coverage-check<TAB>pass<TAB>90<TAB>80" # Stderr: status banners and the full pytest run (progress, coverage report, errors) # Exit: 0 pass, 1 below threshold, 2 usage, 5 pytest missing # # Examples: # coverage-check.sh # coverage-check.sh --threshold 90 --cov mypkg # CM_COVERAGE_OVERRIDE=72 coverage-check.sh --threshold 80 # offline test mode set -uo pipefail THRESHOLD=80 COV_TARGET="src" PYTEST_ARGS=() OVERRIDE="${CM_COVERAGE_OVERRIDE:-}" usage() { cat <<'EOF' Usage: coverage-check.sh [--threshold N] [--cov TARGET] [pytest-args...] Run pytest with coverage; exit non-zero when coverage is below --threshold. The verdict is one line on stdout; pytest output and status banners go to stderr. Options: --threshold N minimum coverage percent (default 80) --cov TARGET package to measure (default src) Offline test seam: CM_COVERAGE_OVERRIDE=PCT skip pytest, judge this percentage instead. Exit codes: 0 pass, 1 below threshold, 2 usage, 5 pytest missing. Examples: coverage-check.sh coverage-check.sh --threshold 90 --cov mypkg CM_COVERAGE_OVERRIDE=72 coverage-check.sh --threshold 80 EOF } while [[ $# -gt 0 ]]; do case "$1" in -h|--help) usage; exit 0 ;; --threshold) [[ $# -ge 2 ]] || { echo "coverage-check.sh: --threshold needs a value" >&2; exit 2; } THRESHOLD="$2"; shift 2 ;; --threshold=*) THRESHOLD="${1#--threshold=}"; shift ;; --cov) [[ $# -ge 2 ]] || { echo "coverage-check.sh: --cov needs a value" >&2; exit 2; } COV_TARGET="$2"; shift 2 ;; --cov=*) COV_TARGET="${1#--cov=}"; shift ;; --) shift; while [[ $# -gt 0 ]]; do PYTEST_ARGS+=("$1"); shift; done ;; -*) echo "coverage-check.sh: unknown option: $1" >&2; usage >&2; exit 2 ;; *) PYTEST_ARGS+=("$1"); shift ;; esac done # threshold must be a number if ! [[ "$THRESHOLD" =~ ^[0-9]+([.][0-9]+)?$ ]]; then echo "coverage-check.sh: --threshold must be a number, got '$THRESHOLD'" >&2 exit 2 fi # ge PCT THRESHOLD -> returns 0 if PCT >= THRESHOLD (float-safe via awk) ge() { awk -v a="$1" -v b="$2" 'BEGIN { exit !(a+0 >= b+0) }'; } # Offline test seam: judge a supplied percentage without running pytest. This is # what lets the threshold logic be exercised offline, deterministically, with no # slow suite and no network — the same pattern check-ytdlp-version.sh uses. if [[ -n "$OVERRIDE" ]]; then if ! [[ "$OVERRIDE" =~ ^[0-9]+([.][0-9]+)?$ ]]; then echo "coverage-check.sh: CM_COVERAGE_OVERRIDE must be a number, got '$OVERRIDE'" >&2 exit 2 fi if ge "$OVERRIDE" "$THRESHOLD"; then printf 'coverage-check\tpass\t%s\t%s\n' "$OVERRIDE" "$THRESHOLD"; exit 0 else printf 'coverage-check\tfail\t%s\t%s\n' "$OVERRIDE" "$THRESHOLD"; exit 1 fi fi # Live path: run pytest with coverage. All of its output (progress + coverage # report) is data for a human, not for this script's caller, so it goes to # stderr; the one-line verdict on stdout is the machine-readable product. command -v pytest >/dev/null 2>&1 || { echo "coverage-check.sh: pytest not found (pip install pytest pytest-cov)" >&2 exit 5 } printf '=== Running tests with coverage (threshold %s%%) ===\n' "$THRESHOLD" >&2 pytest \ --cov="$COV_TARGET" \ --cov-report=term-missing \ --cov-report=html \ --cov-fail-under="$THRESHOLD" \ "${PYTEST_ARGS[@]}" >&2 rc=$? case "$rc" in 0) printf 'coverage-check\tpass\t-\t%s\n' "$THRESHOLD"; exit 0 ;; 2) printf 'coverage-check\tfail\t-\t%s\n' "$THRESHOLD"; exit 1 ;; # pytest-cov below-threshold signal *) printf 'coverage-check.sh: pytest exited %s\n' "$rc" >&2; exit "$rc" ;; esac
-
-
tests
-
run.sh 5.2 KB
#!/usr/bin/env bash # Self-test for testing-ops — fully offline, deterministic, Linux-safe. # # coverage-check.sh gates a project on a coverage threshold. To exercise its # threshold logic WITHOUT running a slow real pytest suite, the script exposes # an offline test seam (CM_COVERAGE_OVERRIDE=PCT) that substitutes a given # percentage for the live measurement — the same pattern check-ytdlp-version.sh # uses for its installed/latest versions. Asserts the --help contract, semantic # exit codes (0 pass / 1 below / 2 usage / 5 missing-dep), and stream separation # (the verdict on stdout; banners + pytest on stderr). Never invokes real pytest. # # Usage: bash tests/run.sh # Exit: 0 all pass, 1 one or more failures set -uo pipefail HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SKILL="$(dirname "$HERE")" V="$SKILL/scripts/coverage-check.sh" SB="$(mktemp -d)"; trap 'rm -rf "$SB"' EXIT PASS=0; FAIL=0 ok() { PASS=$((PASS+1)); printf ' PASS %s\n' "$1"; } no() { FAIL=$((FAIL+1)); printf ' FAIL %s\n' "$1"; } expect_exit() { [[ "$2" == "$3" ]] && ok "$1 (exit $3)" || no "$1 (want $2 got $3)"; } expect_has() { case "$3" in *"$2"*) ok "$1";; *) no "$1 (missing '$2')";; esac; } echo "=== testing-ops self-test ===" # ── contract ────────────────────────────────────────────────────────────────── echo "-- contract --" bash -n "$V" 2>/dev/null && ok "bash -n coverage-check.sh" || no "bash -n coverage-check.sh" bash "$V" --help >/dev/null 2>&1; expect_exit "--help exits 0" 0 $? bash "$V" -h >/dev/null 2>&1; expect_exit "-h exits 0" 0 $? out="$(bash "$V" --help 2>/dev/null)" expect_has "--help has Examples" "xamples" "$out" expect_has "--help documents exit 1 (below)" "below threshold" "$out" expect_has "--help documents exit 2 (usage)" "usage" "$out" expect_has "--help documents exit 5 (missing-dep)" "pytest missing" "$out" expect_has "--help names the test seam" "CM_COVERAGE_OVERRIDE" "$out" bash "$V" --bogus >/dev/null 2>&1; expect_exit "unknown flag -> 2" 2 $? bash "$V" --threshold >/dev/null 2>&1; expect_exit "--threshold needs value -> 2" 2 $? bash "$V" --threshold notnum >/dev/null 2>&1; expect_exit "--threshold non-numeric -> 2" 2 $? CM_COVERAGE_OVERRIDE=bogus bash "$V" --threshold 80 >/dev/null 2>&1; expect_exit "bad override -> 2" 2 $? # ── threshold logic (offline seam; no real pytest) ─────────────────────────── echo "-- threshold logic (seamed) --" CM_COVERAGE_OVERRIDE=90 bash "$V" --threshold 80 >/dev/null 2>&1; expect_exit "90>=80 -> pass 0" 0 $? CM_COVERAGE_OVERRIDE=80 bash "$V" --threshold 80 >/dev/null 2>&1; expect_exit "80>=80 boundary -> pass 0" 0 $? CM_COVERAGE_OVERRIDE=79 bash "$V" --threshold 80 >/dev/null 2>&1; expect_exit "79<80 -> below 1" 1 $? CM_COVERAGE_OVERRIDE=72 bash "$V" --threshold 80 >/dev/null 2>&1; expect_exit "72<80 -> below 1" 1 $? CM_COVERAGE_OVERRIDE=100 bash "$V" --threshold 99.5 >/dev/null 2>&1; expect_exit "100>=99.5 float -> pass 0" 0 $? CM_COVERAGE_OVERRIDE=99.4 bash "$V" --threshold 99.5 >/dev/null 2>&1; expect_exit "99.4<99.5 float -> below 1" 1 $? # default threshold is 80 when --threshold is omitted CM_COVERAGE_OVERRIDE=79 bash "$V" >/dev/null 2>&1; expect_exit "default threshold 80: 79 -> below 1" 1 $? CM_COVERAGE_OVERRIDE=81 bash "$V" >/dev/null 2>&1; expect_exit "default threshold 80: 81 -> pass 0" 0 $? # ── stream separation (verdict on stdout; no banner leaks) ─────────────────── echo "-- stream separation --" out="$(CM_COVERAGE_OVERRIDE=90 bash "$V" --threshold 80 2>/dev/null)"; rc=$? expect_exit "seamed pass exit 0" 0 "$rc" expect_has "stdout verdict is pass" "pass" "$out" expect_has "stdout verdict carries threshold" "80" "$out" case "$out" in *$'\n'*) no "stdout has more than one line";; *) ok "stdout is a single verdict line";; esac case "$out" in *"==="*) no "banner leaked onto stdout";; *) ok "no banner on stdout";; esac out="$(CM_COVERAGE_OVERRIDE=72 bash "$V" --threshold 80 2>/dev/null)"; rc=$? expect_exit "seamed fail exit 1" 1 "$rc" expect_has "stdout verdict is fail" "fail" "$out" case "$out" in *"==="*) no "banner leaked onto stdout (fail)";; *) ok "no banner on stdout (fail)";; esac # ── missing-dep path: no seam AND no pytest -> exit 5 with install hint ─────── echo "-- missing-dep --" # Scrub PATH so a host-installed pytest is not resolvable, keeping the suite # offline and deterministic regardless of host tooling (bash itself stays # resolvable under /usr/bin:/bin). If pytest somehow survives the scrub, SKIP # rather than ever launching a real suite. if PATH=/usr/bin:/bin command -v pytest >/dev/null 2>&1; then echo " SKIP missing-dep exit 5 (pytest resolvable even under scrubbed PATH)" else out="$(PATH=/usr/bin:/bin bash "$V" 2>&1)"; rc=$? expect_exit "no pytest, no seam -> 5" 5 "$rc" expect_has "names pytest install hint" "pytest" "$out" case "$out" in *"pip install pytest"*) ok "hint suggests install";; *) no "hint missing install suggestion";; esac fi echo "" echo "=== $PASS passed, $FAIL failed ===" [[ "$FAIL" -eq 0 ]] || exit 1 exit 0
-
-
SKILL.md 4.1 KB
--- name: testing-ops description: "Cross-language testing strategies and patterns. Triggers on: test pyramid, unit test, integration test, e2e test, TDD, BDD, test coverage, mocking strategy, test doubles, test isolation." license: MIT compatibility: "Language-agnostic patterns. Framework-specific details in references." allowed-tools: "Read Write Bash" metadata: author: claude-mods --- # Testing Patterns Universal testing strategies and patterns applicable across languages. ## The Test Pyramid ``` /\ / \ E2E Tests (few, slow, expensive) / \ - Full system tests /------\ - Real browser/API calls / \ / Integ \ Integration Tests (some) / Tests \ - Service boundaries /--------------\ - Database, APIs / \ / Unit Tests \ Unit Tests (many, fast, cheap) ------------------ - Single function/class - Mocked dependencies ``` ## Test Types ### Unit Tests ``` Scope: Single function/method/class Speed: Milliseconds Dependencies: All mocked When: Every code change Coverage: 80%+ of codebase ``` ### Integration Tests ``` Scope: Multiple components together Speed: Seconds Dependencies: Real databases, mocked external APIs When: PR/merge, critical paths Coverage: Key integration points ``` ### End-to-End Tests ``` Scope: Full user journey Speed: Minutes Dependencies: Real system (or staging) When: Pre-deploy, nightly Coverage: Critical user flows only ``` ## Test Naming Convention ``` test_<unit>_<scenario>_<expected> Examples: - test_calculate_total_with_discount_returns_reduced_price - test_user_login_with_invalid_password_returns_401 - test_order_submit_when_out_of_stock_raises_error ``` ## Arrange-Act-Assert (AAA) ```python def test_calculate_discount(): # Arrange - Set up test data and dependencies cart = Cart() cart.add_item(Item(price=100)) discount = Discount(percent=10) # Act - Execute the code under test total = cart.calculate_total(discount) # Assert - Verify the results assert total == 90 ``` ## Test Doubles | Type | Purpose | Example | |------|---------|---------| | **Stub** | Returns canned data | `stub.get_user.returns(fake_user)` | | **Mock** | Verifies interactions | `mock.send_email.assert_called_once()` | | **Spy** | Records calls, uses real impl | `spy.on(service, 'save')` | | **Fake** | Working simplified impl | `FakeDatabase()` instead of real DB | | **Dummy** | Placeholder, never used | `null` object for required param | ## Test Isolation Strategies ### Database Isolation ``` Option 1: Transaction rollback (fast) - Start transaction before test - Rollback after test Option 2: Truncate tables (medium) - Clear all data between tests Option 3: Separate database (slow) - Each test gets fresh database ``` ### External Service Isolation ``` Option 1: Mock at boundary - Replace HTTP client with mock Option 2: Fake server - WireMock, MSW, VCR cassettes Option 3: Contract testing - Pact, consumer-driven contracts ``` ## What to Test ### MUST Test - Business logic and calculations - Input validation and error handling - Security-sensitive code (auth, permissions) - Edge cases and boundary conditions ### SHOULD Test - Integration points (DB, APIs) - State transitions - Configuration handling ### AVOID Testing - Framework internals - Third-party library behavior - Simple getters/setters - Private implementation details ## Test Quality Checklist - [ ] Tests are independent (no order dependency) - [ ] Tests are deterministic (no flaky tests) - [ ] Tests are fast (unit < 100ms, integration < 5s) - [ ] Tests have clear names describing behavior - [ ] Tests cover happy path AND error cases - [ ] Tests don't repeat production logic - [ ] Mocks are minimal (only external boundaries) ## Additional Resources - `./references/tdd-workflow.md` - Test-Driven Development cycle - `./references/mocking-strategies.md` - When and how to mock - `./references/test-data-patterns.md` - Fixtures, factories, builders - `./references/ci-testing.md` - Testing in CI/CD pipelines ## Scripts - `./scripts/coverage-check.sh` - Run coverage and fail if below threshold
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.