Reproduce-first bug hunt for Cursor
Forces a failing test that reproduces the bug before any fix — so you know it's actually fixed.
#debugging #coding
What vetted this — trust report
The problem this solves
Paste a bug report into an agent and it will produce a fix in about twelve seconds. The fix will look right. It will often make the symptom disappear. And roughly a third of the time it will be addressing a different cause than the one you have, because a plausible fix for a plausible cause is exactly what a model is best at generating.
You find out three weeks later, when the bug comes back under slightly different conditions.
A red-to-green test is the only cheap way to tell the difference. If a test failed before the change and passes after it, and it fails for the reason you believe, you have evidence. Without that, you have a guess with good grammar.
Recipe
Paste the bug report, then drive Cursor through these steps in order. Don't accept a fix until step 3 has gone green-from-red.
1. Locate, don't fix
Using the codebase, find the 1–2 functions most likely responsible for this behaviour. Quote them with their file paths. Explain what makes each a candidate. Don't change anything yet.
The "don't change anything yet" is load-bearing. Without it you get a diff, and once there's a diff in front of you the conversation is about that diff rather than about the bug.
If the candidates look wrong, say so and give it the extra context — a stack trace, the log line, the
call path — rather than letting it start editing. @file the files you already suspect.
2. Reproduce
Write the smallest failing test that reproduces this bug. Run it and show me it fails — and confirm it fails for the expected reason, not for a setup or import error.
This is the step people skip and the step that carries the whole recipe. Two checks before you move on:
- Did it actually run? A test that errors on a missing fixture is not a reproduction.
- Did it fail for the right reason? "Expected 3, got 4" is a reproduction. "NullReferenceException in the test harness" is a broken test.
If it can't reproduce the bug, that is now the task. An unreproducible bug isn't ready to be fixed — see When you can't reproduce below.
3. Fix minimally
Make the smallest change that turns that test green. Don't refactor unrelated code, don't rename anything, and don't modify the test.
That last clause matters more with an agent than with a person. Faced with a red test, agents will readily "fix" it by adjusting the assertion — which converts your evidence into a rubber stamp silently.
4. Guard
Add one more test for the nearest edge case this bug hints at.
Bugs come in families. An off-by-one at the upper bound usually has a sibling at the lower bound; a null that slipped through one path usually has a second path. This step costs a minute and catches the sequel.
5. Explain
In two sentences: the root cause, and why this fix addresses the cause rather than the symptom. Then: what else in this codebase has the same shape of bug?
The second question is the one worth asking. Most bugs are instances of a class, and the search that finds four more is the highest-value minute in the whole exercise.
Why it works
A fix without a red→green test is a hypothesis someone stopped testing. This recipe makes the reproduction the contract: the test is the definition of the bug, and it stays in the suite as permanent proof the bug does not come back.
Cursor specifics
- Use
Cmd/Ctrl+Kinline edit for step 3, not the agent. You want a small diff you can read. - The agent is the right tool for step 1 — searching the codebase is what it's good at.
- Add this to
.cursor/rules/so you don't retype it every time:When fixing a bug: write a failing test first, show it fail, then fix. Never edit an existing test to make it pass — report the failure instead. - Turn off auto-run for terminal commands during this workflow. You want to see each test run and read its output yourself.
When you can't reproduce
That's the task now, and the same discipline applies. Generate hypotheses about why it doesn't reproduce and test those:
- Environment — config, versions, locale, timezone, OS
- Data — the row shape that only exists in production
- Timing — a race your test runs too fast (or too slow) to hit
- Scale — only appears past a volume threshold
- Accumulated state — a fresh test database has none
Add logging that discriminates between those hypotheses, ship it, and wait. Untargeted logging just produces more haystack.
Failure modes
- The test passes on the first run. You reproduced something else. Re-read the bug report; usually a precondition was missed.
- The fix is large. A large fix for a small bug means the diagnosis is wrong, or you're refactoring under cover of a bug fix. Split it.
- The agent fixes the test. Revert, restate the rule, redo step 3.
- Green test, bug still reported. Your reproduction didn't match the user's scenario. Go back to step 2 with their exact inputs — not your approximation of them.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.