Try this if
- You have an AI agent plan work before it builds, and you can't read the code well enough to tell whether its plan is right.
- You've carried a plan to a second AI for a critique and back again, more than once, for the same change.
- You work in Claude Code. The recipe needs its plan mode and a separate agent that Claude Code starts.
- The recipe hands you what a second, read-only agent found wrong in the plan and whether each finding was fixed, before anything is built.
- Skip it for a one-line change. The recipe offers to bow out when the task is that small.
Ask an AI agent for a plan and it hands you a confident one. Somewhere in it is a claim it never checked. It might be a file it says is missing, a count it never took or a guess about how a tool behaves, and it's written in the same sure voice as everything it did check.
You can't tell which is which by reading, especially if you don't read code. So you find out when the build breaks, or you run the loop by hand. You carry the plan to a second AI for a critique, bring the critique back, get a revision and carry that out again, three or four rounds until the rounds stop finding anything. An audit of 83 planning sessions found about 57 percent of plans drew a revision, and the biggest single cause was plans asserting things they never verified. Every one of those was caught by a separate reviewer, never by the agent that wrote the plan.
The answer is to put the round trip inside the planning, so the plan reaches you after one attack and one revision, and at least one of those manual rounds is gone. What you read first is the short list of what the attack found and what was done about each finding.
The recipe came out of that loop, and out of a log of every finding both reviewers made, the outside one and a fresh agent the planner started inside the same session. When the log was closed, 21 of its findings were claims made without reading the file, 42 were things the planner read and still got wrong, and 219 were judgment calls, the kind no check can settle and that stay yours. The in-session reviewer was catching most of the first two kinds before the outside review ever saw the plan, so the round trip was dropped. The mistakes that kept coming back had already been written into the checks.
The planner checks its own claims first
The planner works in plan mode, where it can read and plan but can't edit files. Before the plan reaches anyone, it runs six checks on its own draft, and a plan that fails one goes back for the missing read instead of forward to you.
Every claim about the project has to trace to a file it opened in this session, never memory or a guess from a file name. The read has to be as wide as the claim, so "all the templates" means every template was read. Every verification step has to be able to fail. Anything assumed about a tool or a library becomes the first step, a probe that's allowed to change the plan. Every real choice names the option that lost and why. And every verification step names the goal it proves, so a goal nothing tests shows up as a gap.
The ledger is the part the reviewer attacks
The plan ends with a ledger, a list of its load-bearing claims. Each entry says what was claimed, quotes the lines that were read to check it and says what came back. When the planner couldn't confirm something, the entry says unverified.
An honest unverified is worth more than a made-up verified, because a false verified hides the one bug a reviewer could have caught. The ledger's last line either says there are no open decisions or lists every one, so a real fork never gets quietly settled on your behalf. Each open decision reaches you ready to answer, with the stakes, what was already checked, a recommendation and the options.
The reviewer shares none of the planner's context
The reviewer is a separate agent that can read the project but can't edit it, on the strongest model available and never a weaker one than the planner. It gets the plan and the goals of the work, and nothing else. It never gets a claim to check or a problem the planner suspects, because a reviewer handed the answer agrees with it.
Its job is to reopen every source the ledger cites and try to prove the entry wrong, running the ledger's commands again where it can. Then it checks for scope creep, logic that won't work, anything irreversible, missing edge cases and assumptions nobody stated. Every finding has to point at a file and a line, or it goes under could not verify. In a test run on a small cleanup script, it proved one of the planner's ledger entries false and caught a change the plan would have made to how the script behaves when run normally, which the plan claimed to leave alone.
The planner fixes what the reviewer found and logs each finding with what it did about it. Then it stops. A second review of the revision drifts toward agreement.
The mistakes that repeat become checks
Reread the review logs now and then. In testing, Claude Code saved each plan in ~/.claude/plans/, a folder in your home directory, and the recipe ends every plan with its review log. When the same kind of finding turns up in plan after plan, write it into the project's rules file, AGENTS.md, or its decisions log, so the planner reads it at the start of every session and stops making it.
The recipe's own checks were earned the same way. One clause in the first check exists because a plan twice claimed the next free decision number from a read that had gone stale by the time it wrote the entry, so now the planner rereads both sides at the moment it writes. A single review catches this plan's errors, and a rule written from a repeated finding stops the planner making that one again.
What you get and what you don't
You get a plan whose claims were checked twice, and nothing gets built until you approve it.
You don't get a guarantee. The reviewer catches claims it can check against the project, and the judgment calls, whether the work is worth doing at all, stay yours. It also needs Claude Code, and it stops and says so anywhere else.
Take this spell for a spin
The block below is written for your AI rather than for you. Paste it into Claude Code with the task underneath, or paste it on its own right after you and the agent agree on a next step, and it plans that. It enters plan mode, runs the checks and the review, and shows you the review log and any open decisions first, with the plan below them. It builds nothing until you approve, and at the end it offers to keep itself as a skill, a saved command you call by name, and saves nothing unless you say yes.
Run it on the next change you would have told your agent to build straight away.
Paste this into your AI agent or a new chat. The post above says what it does. Some prompts set something up in your project that keeps working after, and others run once, right where you paste them. None of them pushes anything to a remote.
You are planning a piece of work before any of it is built, and having the plan checked by a second agent that shares none of your context. This needs Claude Code, because it uses plan mode and a separate agent you start yourself. If you cannot enter plan mode or cannot start a separate agent here, say so in one sentence and stop, rather than doing half of it.
The task is: $ARGUMENTS. If that slot is empty or still reads as a placeholder, the task is whatever I wrote after these instructions. If there is nothing there either, the task is the next step we just agreed on in this conversation, and you do not need me to restate it. If nothing has been agreed, ask me what to plan and wait.
Enter plan mode before you do anything else, before you read a file or restate the task. Everything below assumes the plan is written inside plan mode from the first step.
Then decide whether the task needs this at all. If it is a small change, roughly one file with nothing uncertain about how the tools behave and no earlier decision in play, say so and offer to skip the checks and just do it. If I agree, stop here.
Otherwise write the plan as you normally would, with the acceptance criteria stated plainly: what must be true when the work is done. Then, before you show it to me, run these six checks on your own draft. If a check fails, do the reading or the probe now and fix the plan. A plan that fails any check does not get shown.
1. Every claim about the project traces to a read you ran in this session. Every count, file name, "this exists" and "this is missing" must come from a file you opened or a search you ran, never from memory or a guess from a file name. When a claim says two strings match, or that a number is the next free one, re-read both sides at the moment you write it, because files change while you work.
2. The read has to be as wide as the claim. A claim about all of something needs all of it read. A claim about one value needs that value read, not the file skimmed. A claim about what a change will cause needs the things that depend on it read, not just the place being edited.
3. Every verification step can fail. For each one, say what you would see if it came out red, and cover the unchanged path and the failure path as well as the one where everything works. A check that would pass even when the thing it tests is broken gets replaced.
4. Anything the plan assumes about a tool, a platform or a library that you have not confirmed in this session becomes step one, a probe whose result is allowed to change the plan. Never put it late in a plan that collapses if it fails.
5. For every choice with more than one reasonable answer, write the option chosen and at least one rejected, each with its reason. If the project keeps a log of decisions, name every recorded decision the plan touches, and flag any reversal as something for me to sign off, never a silent change. Confirm nothing we agreed earlier in this conversation was dropped.
6. Every verification step ends by naming the acceptance criterion it proves. A criterion nothing verifies, or a step that verifies nothing, is a finding.
Append a section to the plan headed Verification ledger. One entry per load-bearing claim, in this shape:
- Claim: the claim.
Checked: the exact read, search or probe, quoting the lines you read rather than saying you read them, or "could not confirm".
Result: what you found, and which step it changed if it changed one.
The ledger's last line is either NO UNRESOLVED DECISIONS, written exactly like that, or a numbered list of every decision the plan could not settle. A fork the plan could not settle never gets quietly resolved to its most likely answer. A made-up "verified" is worse than an honest "unverified", because it hides the one bug a reviewer could have caught, so when you could not confirm something, write Result: UNVERIFIED and either make it a probe step or leave it for the reviewer.
For each decision you cannot settle, do all the work you can first, then put it to me ready to answer: what is at stake in plain words, what you already checked, what you would pick and why, and the exact options.
Then have the plan checked, before you show it to me. The plan is already saved in the file plan mode gave you; use that exact path and never guess it. Start one separate agent that can read but cannot edit anything, on the strongest model available to you and never a weaker one than you are running on. Give it only the message below, with the plan's path and the acceptance criteria filled in. The acceptance criteria are the goals of the work and nothing else. Never name a claim to check, a line to look at or a problem you already suspect, because that hands the reviewer the answer and it will agree with you. Give it nothing from this conversation.
REVIEWER MESSAGE STARTS
Plan file: [path]
What this plan must accomplish, in plain language: [acceptance criteria]
You are reviewing this plan before anything in it is built. Edit nothing. Read the plan, then read the project's own instruction files if it has them, because they often state limits the plan has to respect.
First confirm that every file, function and path the plan names exists as described. If something is missing, say so before going further, because the plan may have been written against a different state of the project.
The plan ends with a verification ledger of claims its author says were checked. Treat it as a list of claims to attack, never a list of facts. Reopen each cited source and try to prove the entry wrong, running its commands again where you can. A ledger entry you cannot reconfirm is not verified.
Then check the plan for: scope, meaning work nobody asked for; correctness, meaning logic that will not work or contradicts the code; blast radius, meaning everything it really touches and anything irreversible, flagged loudly; missing edge cases; verification that could not fail; assumptions it never states; and whether it solves the problem that was asked or a slightly different one.
Anchor every finding to a file and the line or function you read. A finding with no anchor is not a finding and goes under could not verify. Do not pad. If the plan is sound, say so plainly.
Report under four headings. Must fix: bugs, scope violations and irreversible risks, each as an instruction naming the file or step. Should consider: likely improvements. Could not verify: everything you could not confirm, each as an instruction to confirm it, and this section is required; if it is empty, say you checked everything against the code. Open decisions: real choices the plan leaves open, each with the stakes in plain words, what you checked, a recommendation and the exact options, for the author to put to the person who owns the work.
End with two lines. First, exactly one of: no problems found in what I could check, address the must-fixes first, or needs rework. Second, what you read and confirmed and what you could not.
REVIEWER MESSAGE ENDS
When the reviewer returns, fix the plan for every must-fix and for each should-consider you accept. Append a section headed Review log with each finding word for word and what you did about it, fixed or declined and why. Then run git status in the project and confirm the review changed nothing, since the reviewer cannot edit and the plan file lives outside the project. If something did change, stop and tell me. If the project is not a git repository, say that this check could not run.
One review, one revision. Do not send the revised plan back for another review, because a reviewer checking your answers to its own findings drifts toward agreeing with you.
Then show me the plan and leave plan mode, opening with the review log and any open decisions, because those are what I read first, with the plan itself below them. Build nothing until I approve it. Do not commit anything and do not push.
At the end, tell me in one line that I can have this on a keystroke instead of a paste, and create no file unless I answer yes, in which case save these instructions as a skill wherever Claude Code keeps them, with one line at the top saying it only runs when I call it by name. If a skill with that name is already there, tell me and ask before you replace it.
