Try this if
- You're choosing between two or three live options and you can't honestly say which way you lean.
- You already have an answer you like, or your agent just proposed one, and you want it attacked from five directions before you commit.
- You've noticed that how you word the question changes the answer you get back.
- Skip it while you're still generating options. This tests what's already on the table.
You ask one model which of two options to take, and it picks the one your question leaned toward. You ask whether the plan you already wrote is any good, and it tells you the plan is good.
Ask that second one again a week later with your own doubt showing in the wording. The same model finds three problems with it. Same plan, same model, different mood in the framing.
Sycophancy is the easy word for that, and it undersells what it costs you.
You can't tell which of the two answers was the honest one. So the one you keep is whichever matched how you felt when you asked.
It gets worse when the thing on the table came from the agent, because it proposed the change and now you're asking it to grade its own work.
So you stop asking once. Your question gets written down first, with the options named and the stakes stated, and that written version is what everyone downstream sees. Nobody past that point hears your tone of voice.
The part that took longest to learn is that you can't get this by asking nicely. Telling one model to consider five perspectives returns five paragraphs in one voice, each softened by the paragraph above it. By the third lens it's agreeing with itself.
The separation has to be structural or it isn't separation.
The separation is bought, not asked for
The shape comes from Andrej Karpathy's LLM Council, a small open-source project for polling several models on one question. It sends the question out, "asks them to review and rank each other's work," and has a chairman compile the result.
What changes here is the thing that varies. Karpathy used different models. This uses five lenses on one model, because the lenses are where the disagreement comes from and one model is what most people have to hand.
Where your tool can start fresh helper sessions of its own, each advisor runs as its own helper, launched at the same time, holding nothing but its lens and the written question. That is what a coding agent gives you. Either way the advisor order is shuffled, so nobody is always first.
In the Claude chat app there are no separate agents. One agent plays all 11 roles in turn, which means the separation has to be bought with discipline instead. Every answer opens on a hard reset instruction. An answer that refers to another answer is thrown out and rewritten.
Boris Cherny, who created Claude Code, puts the same point in terms of context windows. One agent can write a bug and another agent, on the same model, can find it, because "multiple uncorrelated context windows tends to be a good approach."
The lenses don't make the model smarter. They keep one pass of it from being the only thing that judged the answer.
The lenses are chosen for the tensions between them
They're fixed, and the point is not coverage.
- The Skeptic assumes there's a fatal flaw and hunts for it. If everything looks solid, it digs harder.
- The Architect strips the assumptions and asks what's actually being solved, then names the simplest first move that would work.
- The Stranger knows only what the written question says. It lists what an outsider wouldn't follow and what you've assumed everyone already agrees on.
- The Long View projects three years out. It names what the choice makes true later, what it forecloses, and what the best case would be worth if it lands.
- The Veteran has seen it before. It defends the standard play where it applies, and it prices the move in time or money.
The tensions are the point. Skeptic against Veteran is downside now against the standard play that already handles it. Long View against Skeptic is downside later against downside now. Architect against Veteran is rethink it against do the normal thing.
The Stranger sits outside all of that and keeps the rest honest.
Verify the facts before anyone answers
The council weighs judgment. It does not establish fact, and it's at its most dangerous when nobody notices the difference.
Advisors are told to lean fully into their angle without hedging, which is what makes the answers useful. It also means an advisor missing a fact will invent one and assert it with total confidence.
Because all five reason from the same written question, the skill's own instructions warn that one confident wrong premise poisons every lens at once. The verdict comes back unanimous and wrong.
So the facts go in verified, before anyone answers. The test is simple. If an advisor would have to invent something to hold its position, that something belongs in the question, checked first.
This sits in front of the council rather than inside it. It's tempting to fix the same problem by bolting a verification pass onto the end, and that's the move to resist. The recipe's own rule is that extra stages between roles cost time and tend to flatten what comes out.
The reviewers read letters, not names
After the five answers come back they get shuffled, labeled A through E in a fresh random order, and the mapping is printed once so you can check it afterwards.
Each reviewer returns three things. The strongest answer and why. The answer with the largest blind spot and what it misses. And what all five missed.
That third one is the reason the review round exists. The first two sort answers you already have. The third hands you something you didn't bring, because it finds what lives in the gap between the lenses and no single advisor is positioned to see it.
The blindness is partial either way, and the agent is told to say so beside the verdict. Whether the roles run as separate agents or one agent takes them in turn, the thing that wrote the answers is the thing grading them, so the strongest and blind-spot votes count for less. What all five missed is the field that still pays.
One detail here is easy to get wrong and it wrecks the output. Name the options in your question by what they are, never as A, B and C.
An early run labeled the choices that way. The anonymizing step labeled the answers A through E, and the reviews came back arguing about "the distinction between C and B/D" with no way to tell whether a letter meant an option or an answer.
When the letters are swapped back for names at the end, the votes are the easy half. They have to come out of the reviewers' sentences as well.
Agreement is the weakest signal in the room
The swap from five models to five lenses has a cost, and this is where it lands.
Five advisors on one model is not five opinions. It's one set of beliefs wearing five rhetorical costumes. They decorrelate how they argue, not what they think is true.
So the chairman weighs the arguments rather than counting the votes, and it's allowed to side with a minority when the minority's reasoning is stronger.
Five-of-five agreement reads like certainty and is closer to a warning that the framing left no room to disagree.
The verdict ends in one action, not a shortlist and not a set of considerations. A chairman that writes "it depends" has handed the ambiguity back to the person who came to get rid of it, and at that point you have run 11 roles to buy what one ordinary question would have told you.
If a load-bearing fact was never checked, the verdict says so next to the headline and the first step becomes checking it.
The cheap attack and the council compound
/devils-advocate is the same instinct with one role instead of 11. One idea, steel-manned first so the attack lands on the strong version, then gone after on its assumptions, its failure modes and its hidden costs.
It's quick enough to reach for without thinking about it, and most of the time it's enough.
The two work in sequence rather than as a choice. Attack the idea first and take what survives to the council, or run the council and put its verdict through the attack before you act on it.
Both directions are useful. The cheap one going first means the expensive one starts from a better idea.
/problem-first sits upstream of both. It's for the moment a solution shows up before anyone has said what problem it solves, and it refuses to move until the problem underneath has been stated plainly.
A council there would be premature, because nobody has established there's a real choice to make yet.
What you get and what you don't
The block below sets nothing up on your machine. It runs the council where you paste it and leaves nothing behind unless you ask it to keep the instructions.
What comes back is the written question, five answers under their advisors' names, the letter mapping, so you can see which answer was whose, five reviews, and the chairman's verdict with one first step. If you want it on a keystroke afterwards, tell your agent to save the instructions as a skill.
Two things it doesn't give you. The first is the styled report. That comes from a packaged version of this recipe, a script that renders the verdict as a sheet you read at a glance. A script is something you run, not something you paste into a chat box, so it cannot travel in a block like this one. That version isn't public yet.
The second is the decision. The verdict is a recommendation with one first step attached, and the block tells your agent to stop there rather than act on it.
It also won't be comfortable. If you arrived wanting the answer blessed, the useful outcome is the one where it isn't, and nothing in the design can tell the difference between you testing an idea and you fishing for approval. That part stays yours.
A run takes a few minutes either way, which is the whole argument for the fast version above.
Bring it something you're about to commit to, with the options named by what they are.
Paste this into your AI agent or a new chat. The post above says what it does. Some prompts set something up in your project that keeps working after, and others run once, right where you paste them. None of them pushes anything to a remote.
You are running a council on a decision of mine. Run it here, in this conversation. Do not set anything up, do not save any files, and do not commit anything.
The first thing you print, before you check anything or write anything, is which of the two ways you are running this. Check whether you can run separate agents at the same time. If you can, every advisor and every reviewer below gets its own, each given only its own instructions and the written question, and none of them sees what another wrote. That separation is the whole design. If you cannot, which is the case in a chat app, you will play all 11 roles yourself, in turn, and the reset lines below are what stands in for the separation. That first line is a claim and not yet a fact, so the moment a separate agent actually returns something, say in one line that it did. If it turns out you cannot dispatch after all, say that instead and carry on playing the roles yourself.
Never describe a step you are about to take and then stop. Every stage below happens in the same turn that announces it.
The decision is: $ARGUMENTS. If that is empty, ask me for it and wait. Keep my words exactly as I gave them, because the chairman needs them at the end.
First, write the question down: the decision, the options named by what they are and never as letters, and what is at stake, plus the two or three facts that sharpen it. If the right answer depends on a fact none of us has checked, check it now and put it in the written question before anyone answers. An advisor missing a fact will invent one and assert it with full confidence.
Then the five answers, one per advisor, in an order you shuffle yourself. If you are playing the roles yourself, open every one with this line: Reset. You have read nothing else in this conversation, you are starting fresh as this advisor, and your only context is the written question. Keep each to 150 to 200 words. An answer that refers to another answer is broken, so rewrite it. The five lenses:
- The Skeptic assumes there is a fatal flaw and hunts for it. If everything looks solid, it digs harder.
- The Architect strips the assumptions, asks what is actually being solved, then names the simplest first move that would work.
- The Stranger knows only what the written question says, and lists what an outsider would not follow and what I have assumed the reader already believes.
- The Long View projects three years out, names what this forecloses, and names the best case and what it would be worth.
- The Veteran has seen this before, defends the standard play where it applies, and prices the move in time or money.
Do not arrange the five to balance each other. If four say go and one says stay, that is the signal.
Then shuffle the five answers, label them A to E in a fresh random order, and print the letter-to-advisor mapping once so I can audit it. Give five reviewers the lettered answers with no names, each opening with the same reset line, and have each return three things: the strongest answer and why, the answer with the largest blind spot and what it misses, and what all five missed. That last one is the field that pays. Then put the names back, inside the reviewers' sentences as well as in their votes.
Last, as chairman, read my original words, the written question, the five named answers and the five reviews, and write: a headline under eight words, the verdict in one sentence, where they agreed, where they clashed and how that resolved, what all five missed, and exactly one first step, a single action. You may side with a minority when its reasoning is stronger. Weigh the arguments, not the votes, because five lenses on one model share that model's beliefs. Say once, beside the verdict, that the same model wrote and graded the answers, so the votes count for less and what all five missed counts for more. If a load-bearing fact was never checked, say so beside the headline and make checking it the first step.
Never act on the verdict. That part is mine.
At the end, tell me in one line that I can have this on a keystroke instead of a paste, and create no file unless I answer yes, in which case save these instructions as a skill wherever my tool keeps them. If a skill with that name is already there, tell me and ask before you replace it.
