Watch a check fail once before you trust it to pass

Plant the mistake a check exists to catch, run it and see it go red, then put things back. A search for something that should be absent gets a known hit run beside it, so an empty result can't pass for a clean one.

Try this if

  • Your AI tells you every check passed, and you've never seen one of those checks fail.
  • A search came back empty and you took the empty result as good news.
  • A bug reached you that a passing test should have caught.
  • Skip it if you already watch every check you rely on fail once, or run a known hit beside every search for nothing.

Your AI agent finishes a change and reports that everything passed. The tests are green, the build is clean, and the search for leftover references to the old name found nothing. Some of those results could never have come out any other way.

A failing check costs you a fix. A check that can't fail costs more, because it tells you something is covered and you stop looking there. The mistake it was meant to catch gets through, and with no check at all you'd at least still be looking.

A smoke detector's silence means nothing until you've held smoke under it and heard it go off. Checks work the same way. Before you trust a pass, break the thing on purpose, watch the check go red, meaning it fails, and then put things back. The habit comes from checks that passed, got trusted and turned out to be blind to the very thing they were for, including a search that came back empty for every phrase it was given.

A pass only means the check found nothing it could see

Checks are the automated tests, searches and build steps your agent runs to confirm its work. When one passes, it tells you it ran and, within what it was able to see, found nothing wrong. It says nothing about whether it could see the problem at all.

Most broken checks fail loudly, with an error or a crash. The dangerous ones fail quietly. A search that can't open its files and a search with nothing to find both print an empty result, and on the screen the two look the same. So does a test pointed at the wrong file, or a check that runs on your machine and never where the code actually ships.

What happened when a search found nothing

One project changed a rule at the end of a working session, then had its agent sweep the project's documents for any that still stated the old rule as current. The agent searched for 12 phrases. Every one came back with zero files, including phrases it knew were in the project's main rules file.

The sweep was broken in two places at once. The search got its list of 85 files through a variable, one name that stood for the whole list. The shell, the program that runs typed commands, was zsh. Unlike bash, the other common one, zsh doesn't break a list held that way back into separate names, so it handed the whole list to the search as a single file name. No file had that name, so the search failed to open it. The command also threw its error messages away, so the failure never appeared on screen. What was left looked exactly like a clean result.

It was caught only because the project had a rule for this case. A search that finds nothing about a change you know is recorded elsewhere means the phrases are wrong, not that the documents are clean. Run again with the list passed properly, the same sweep found 17 hits across seven files. Four of them stated an old rule as current, and one of those four was in the file the agent reads at the start of every session.

Either defect alone would have shown up. A list passed properly would have found the hits, and a visible error would have announced itself. Together they produced silence.

Plant the mistake and watch it go red

Before you trust a check, give it the thing it exists to catch. Write AGENTS.md rules that survive the session makes this part of writing a rule in AGENTS.md, the rules file your agent reads at the start of every session. The habit works on any check you already run. For a test, break the code it tests. For a config check, misspell a key. For a search, add a line that should match. Run the check and confirm it fails, then undo the plant and run it again.

A session auditing a web app tested each of its tools this way, and three of them show what a plant buys. A deploy tool's trial run passed two edited config files with no warnings. A copy with one key misspelled produced an error, so that clean pass meant something. Next came a security policy, the page's own list of which scripts and styles it may load. A browser tool that reads the page's console, where the browser reports errors, said there were no violations of a new, stricter policy. The session planted two violations on purpose. Both fired in the page, and the tool still reported nothing, because it couldn't see that kind of event at all. A second tool, tested the same way, had the same blind spot.

Two of those three would have reported a clean page that nothing had actually checked. Each plant cost one extra run.

Make sure the plant actually landed

The plant can fail silently too. In one project, the command meant to break a file for exactly this test hit a syntax error, changed nothing and moved on. The check then passed, correctly, against a file nobody had touched. The session's own notes called it "exactly the shape of a proof that proves nothing."

So confirm the break before you run the check. Count the planted line, print the changed value or open the file. If the break isn't there, the red you're waiting for can't come, and the pass that follows means nothing.

Run a known answer beside every search for nothing

Some checks can't be broken on the spot, and a search for something that should be absent is the common case. For those, run a positive control, a search for something you know is there, in the same run. If the known answer comes back empty, the search itself is broken, and every other empty result from that run is worthless.

It also helps to have a check report what it looked at, such as how many files it read, which page it loaded and what it assumed about the setup, not only whether it passed. An empty list of problems sitting next to "read 0 files" can't be mistaken for a clean run.

Test the thing that runs the check

A check is often wired to run on its own, from a project's settings or from a routine that fires when a session opens or a change is saved. Running the check yourself proves the check works. It doesn't prove the wiring ever runs it.

When one project added a step to its session-opening routine, the tempting test was to run the step's script by hand. That would have come back green whether or not the routine ever picked the step up. The test that could fail was opening the next session and watching the step fire from the routine itself, and it did.

Ask your agent to prove the check can fail

You don't have to plant defects by hand. When your agent reports a pass on something that matters, ask it to do the plant and to print the change it made before it runs the check. For a check you'll lean on again, put that request in AGENTS.md so it happens without being asked.

What you get and what you don't

You get checks you've seen fail, so a pass from them carries information. You don't get a check that looks for the right thing. A search for a color by its code can pass while the same color, written another way, sits on the page. The plant proves the check can see what you planted, and choosing what to plant stays with you.

The next time your agent says everything passed, pick the check that matters most and ask it to make that one fail.

Published
Kindlesson