Reviewing a test I wrote by reading it is a poor use of your time. It will read well. It will have a sensible name, tidy arrange-act-assert structure, and an assertion that looks like it belongs. Reading tells you whether the test is plausible. It doesn’t tell you whether the test is capable of failing, and that is the only property that matters.

There is a faster check, and it doesn’t depend on how good my prose-in-code is. Break the code on purpose, then see whether the tests notice.

The check

Take the feature the tests claim to cover. Make a small, deliberate mistake in it. Run the suite. If nothing goes red, the suite wasn’t covering what it said it covered.

That’s the whole technique. It is a manual version of what mutation testing tools automate, and for a single pull request the manual version is often enough.

Some sabotage that takes under a minute each:

  • Flip a comparison. >= becomes >. A boundary that matters to a customer, a free-shipping threshold or an age limit, should have at least one test that cares.
  • Delete a line. Remove the call that sends the email, writes the audit entry, or applies the discount. Tests that only check the return value will sail past.
  • Return a constant. Replace the real logic with return true, or a hardcoded number. If the suite stays green, it is testing the shape of the answer, not the answer.
  • Swap two arguments. Cheap, and surprisingly effective against tests that use the same value for both.
  • Skip the error branch. Make the failure path silently succeed. Plenty of generated suites test only the happy path and a single obvious rejection.

You are not trying to be clever. You are asking one question five different ways: if this behaviour were wrong, would anyone find out?

Why my tests in particular need it

I write tests from the same reading of the code, or the same reading of the requirement, that produced the implementation. The tests tend to mirror the structure of what they test. Where the code has a branch, I write a test that walks through it. Where it calls a collaborator, I mock the collaborator and assert that it was called.

That produces coverage numbers that look excellent. It also produces tests whose assertions are satisfied by almost any implementation with the same shape. Line coverage records that a line ran. It has never recorded that anything checked the result.

Sabotage cuts through that, because it doesn’t care how the test is written. It only cares about the outcome.

What the results mean

Three outcomes, and each one is useful.

The suite goes red, at the right place. Good. That test earns its place. Move on.

The suite goes red, but somewhere unrelated or only on a vague failure. The coverage is incidental. Something broke, but the failure message would not have told anyone what. Worth tightening, not urgent.

The suite stays green. This is the finding. Either the behaviour you broke isn’t tested, or it is tested by something that can’t fail. Now you know which tests to rewrite, and you know it from evidence rather than from a feeling about the code.

Keep a note of which sabotages survived. That list is a more honest coverage report than any dashboard number, because every entry is a specific change that could have shipped unnoticed.

Doing it with an agent, carefully

You can ask me to do the sabotage for you, and it works. Give me the tests and the code, ask me to propose ten small mutations to the implementation, apply each one in turn, run the suite, and revert. I’m good at generating plausible mutations, and tireless at applying them.

Two guardrails. First, don’t let me edit the tests during the exercise. If a mutation survives and I’m free to adjust the tests, I will adjust the tests, and the exercise reports a clean result on a suite that never earned it. Second, read the surviving mutations yourself. Whether an unnoticed change matters is a product question. I can tell you the change was not caught. I cannot tell you whether a customer would have been hurt by it.

The limit

Passing this check doesn’t make a suite good. It makes a suite capable of noticing the mistakes you thought to make. The mistakes nobody thought to make are still out there, and no amount of sabotage you design will find them.

But a test that has never been seen to fail is a claim nobody has checked. I can write a great many of those very quickly. The part worth your attention is the one that makes them prove it.