Judge your AI agent's work by using it, not by reading the code

Try every change yourself before your agent saves it, and leave the rest to automatic checks. A report says what was written, and only using the thing shows what it does.

Try this if

  • Something your AI agent reported as done broke the first time you used it.
  • You find out whether a change works by hearing about it, not by trying it.
  • Skip it if you read every change the agent makes, line by line, before it's saved.

Your AI agent finishes a change and reports on it. A careful report says what was written, what builds and what passed its checks. None of that says whether the thing works when a person uses it, and you can't read the code to find out.

A feature can build, pass every check the agent wrote and still do the wrong thing the moment someone touches it. If the report is your only evidence, the first real test is a user, and whatever got built in the meantime sits on top of the problem.

The answer is to use the thing, and to make the agent wait for you before it commits any change you can see. Dispatch, a Mac app for starting voice dictation from a hotkey, was built this way, and each session's closing note splits what was checked into two lists. In one session's note, the checked list included recording a new hotkey, which the app's owner had tried in the running app, and the menu's appearance, which the owner signed off on after four rounds. It also said the app built cleanly at every step. The unchecked list included a setting that makes the trigger start and stop recording, which was written and built but never tried from start to finish.

The menu took four rounds because every workaround introduced a new defect, and each one was caught by looking at the menu, not by reading about it. That setting builds, and nobody knows yet whether it works. A report can tell you what was written and checked, and only using the thing tells you what happens when someone touches it.

The agent can't commit what you haven't seen

A commit is a save point in version control, the history of your project's files. The rule goes in AGENTS.md, the rules file your agent reads at the start of every session, worded so it fires at one exact step. That app's AGENTS.md says no change to what you see or touch gets committed until the owner has confirmed it in the running app.

The rule ties your check to the commit, and it works best with a procedure written into it. The agent makes the change, starts the project so you can use it, tells you exactly what to look at and where, and waits. You try it. Only then does the change get saved.

The part to insist on is "exactly what to look at." A request to "check it looks right" gets a glance. A request to open the settings page, record a new hotkey and confirm it starts dictation in the app you use gets a test.

Using the thing is slow, so save it for what only a person can judge, like whether a screen makes sense or an interaction feels right. Everything a machine can check goes to a small program that passes or fails, such as a link that resolves or a file that exists, and Write AGENTS.md rules that survive the session covers how to write one.

What you get and what you don't

You get a project where nothing you can see reaches your saved history until a person has tried it.

You don't get code you understand. If the project ever needs someone who can read it, a security review or a handover to an engineer, that's a separate job, and this method doesn't do it for you. The files that carry the project's memory between sessions are in Stop your AI agent from forgetting your project between sessions, and the jobs a team splits between people are in Treat AI agents like a product team and give them roles.

Add the confirmation rule to your AGENTS.md today.

Published
Kindlesson