things one does with a very large graphics card and a quant’s old instincts
There is an old idea in portfolio theory: uncorrelated decent bets beat one brilliant bet. It’s the closest thing finance has to a free lunch. The magic isn’t the number of positions – it’s that they don’t move together. Uncorrelated mediocrity beats correlated brilliance.
And here is why I cared enough to spend five days on it — because the stakes are not academic. The best AI answers today come from enormous models that cost a fortune to train and a fortune in electricity to run every single time you ask them something. That is real energy, real money, real carbon dioxide emission. Now imagine a handful of small, cheap models that, combined the right way, matched one giant expensive one. You would get top-tier answers at a fraction of the energy. That is the prize. My instincts, from watching it work with portfolios, societies, armies etc. said it should be possible — a crowd of weak forecasters beating the expensive star. So the real question underneath this whole experiment is: can we make cheap, weak AIs think differently enough that, together, they rival an expensive one? If yes, it changes the economics and the energy footprint of the whole field.
So I spent a few days asking: does this work for AIs?
First, what “wisdom of crowds” actually means
In 1906, the statistician Francis Galton watched a contest at a country fair where hundreds of people guessed the weight of an ox. Almost every single guess was wrong. The average of the guesses was almost exactly right. That is the wisdom of crowds: you make everyone vote, and if their errors are independent – one person guesses too high, another too low, for different reasons – the errors point in all directions and cancel out. The majority lands on the truth even though most individuals miss it.
Read that fine print again, because the whole story turns on it: the errors must be independent. If everyone at the fair had over-guessed for the same reason (say, the ox was standing in a flattering light), no amount of averaging saves you. The crowd would be confidently, unanimously wrong.
The plan
On the computer in my home office I have four AI models, built by four different companies: Alibaba’s Qwen, Meta’s Llama, Mistral, and Google’s Gemma. Each is okay-but-not-great at a hard science quiz – they score between 53 and 64 out of 100. The plan: make them vote, exactly like the crowd at the fair. If their errors are independent, the vote should beat the best member.
Everything reduces to one question: are their mistakes independent?
Here is how I want you to read what follows. This wasn’t a checklist of tricks I ran at random. Each experiment was forced by the failure of the one before it — a chain of “well, if that didn’t work, then the cause must be deeper, so I have to test this next.” The interesting part isn’t any single result. It’s the ladder: every rung ruled out a shallow explanation and pushed the cause one level deeper, until what was left was something I didn’t expect and couldn’t shake. I’ll flag the reasoning at each step, because the reasoning is the actual story.
By the way, the first eight were with one AI model alone.
Everything I tried, in order
| # | What I tried | What happened |
|---|---|---|
| 1 | Jiggle one AI’s brain-wiring a little (make noisy copies) – increase temperature | The copies thought exactly alike |
| 2 | Jiggle it a LOT – change some parameters | The copies got different – by getting stupider |
| 3 | Jiggle only the most important wires – change the most impactful parameters | Same thing: different only through damage |
| 4 | Reword the questions – different kinds of prompts | Same mistakes anyway |
| 5 | Give the AI costumes (“answer like a scientist”) – make it mimic someone else | No help; some costumes made it worse |
| 6 | Let one AI guess several times and average over the answers | Helps a little – but it’s one AI talking to itself |
| 7 | Ask the questions in different languages | Finally! Different mistakes! |
| 8 | Try a much harder quiz | Everything held. Same story |
| 9 | Stop experimenting; ask a more incisive question: when two AIs are both wrong, do they pick the same wrong answer? | The whole mystery solved itself (see below) |
| 10 | Use four AIs from four different companies | Read on |
| 11 | Ask them all to invent things | Read on. This is the eerie one |
| 12 | Disguise the question as a totally different subject (cooking, sports, finance, music) and ask other AIs | A first run looked striking; a properly-controlled run killed it – no different from a plain reword |
Rows 1–6 all hit one wall: you cannot make an AI think differently by shaking its outsides. Its knowledge – and its blind spots – live too deep.
Row 7 was the breakthrough. These AIs speak many languages, so I translated each question into Chinese, Hindi, or Spanish, and asked in that language. Now the same AI missed different questions depending on the language! Sometimes, the more distant the languages, the more different the mistakes – however, English-and-Hindi disagreed more than English-and-Chinese, even though English is closer to Hindi (which is derived from Sanskrit and Urdu) than Chinese.
I finally had my crowd of different thinkers. I let them vote.
The crowd lost to its own best member anyway.
Why translation failing sent me somewhere deeper
Stop and look at what translation did, because its failure is what cracked the problem open. I had finally built genuinely different thinkers — models missing different questions, exactly what the portfolio math wanted. And the vote still didn’t win. That is strange. If the members are truly decorrelated, the math says the vote must improve. When a solid theorem seems violated, one of your assumptions is wrong — so I went hunting for the buried assumption. It was this: I had been checking whether the models missed the same questions. I had never checked whether, when wrong, they chose the same answer. Those are different things, and only one of them is what the theorem actually requires.
The question I should have asked on day one
For four days I had measured whether the AIs got the same questions wrong. The sharper question is:
When two AIs are both wrong, do they pick the same wrong answer, or different ones?
Think about the ox. Voting only works because wrong guesses scatter in all directions. On a quiz question with four choices, one is right and three are wrong. If two wrong AIs scatter across different wrong answers, their wrong votes cancel and the right answer can still win. If they pile onto the same wrong answer, the crowd confidently votes for the wrong thing – together.
How often would two wrong AIs match by pure luck? Each picked one of three wrong choices, so: one time in three. (A confession: my first calculation said one in four, because of a bug caused by literally ONE five-choice question in a thousand. We caught it. Always check your luck-math.)
The real number for my language-crowd: they matched on the same wrong answer more than half the time – about 1.6 times more than luck. Translation had changed which questions the AI missed. It had done NOTHING to change which wrong answer tempted it. Underneath both languages sits one brain, and its misconceptions are baked in deeper than any language.
What does this actually look like? Here are two of the questions where all four models — every one — reached for the identical wrong answer. Asked what happens to the temperature when vinegar reacts with baking soda, all four confidently said the mixture releases heat and warms up. It does the opposite: the reaction pulls heat in and gets colder. They pattern-matched to the story the internet tells about fizzy reactions — fireworks, fire, heat — and all four made the same mistake for the same reason. Asked which of several animal behaviours is learned rather than instinctive, all four picked a woodpecker tapping a tree for insects — which is instinct, not learning. Not four guesses scattered across the wrong options. One shared wrong answer, four times.
The logic that forced the next test
Now I had a mechanism: shared wrong answers, not just shared difficulty. But a mechanism raises a question — where does the sharing come from? One obvious suspect: maybe it’s just a quirk of one model family. That’s a testable claim with an obvious experiment, so I ran it.
Surely four different AIs would fix it?
That was my next thought. Different companies, different training data, different recipes – surely they have independent blind spots.
They matched on the same wrong answer 72 times out of 100 – more than twice what luck allows, and even more often than one AI across two languages did. Four separately-built brains, reaching for the same wrong answer.
Did the vote lose? Here I must be careful, because I first got this subtly wrong. With all four AIs voting, the crowd scored 61 and the best member scored 64 – but that is mostly because one AI (Google’s Gemma) was much stronger than the others, and a crowd of one giant plus three kids loses to the giant alone. That is old news about unequal crowds, not news about AI. So I removed Gemma and kept the three evenly-matched AIs. Their vote: 58. Best member alone: 57. A tie. I tried fancier voting too – averaging how-sure-each-AI-was, trusting the most confident one per question. Nothing genuinely beat the best member. A crowd cannot out-vote a mistake the whole crowd shares.
Why do they share mistakes? Because every one of them learned from the same place: the same internet, the same books, the same ocean of human writing. The AIs were never the source of the sameness. The data is.
They even invent the same things
Then the eeriest part. Fine – they share wrong facts. Surely imagination is different? I gave all four AIs invention prompts built from weird random word-pairs no one has ever asked before – “bamboo and kung-fu,” “moth and move your a**.”
Asked to connect “moth” and “move your a**,” all four independently wrote about a moth flying to a flame. Asked to invent a bamboo-and-kung-fu machine, all four designed the same product – a bamboo training suit with sensors – and two nearly gave it the same name. Four different brains, one imagination.
Why I turned to disguises
By now every input-change I’d tried — rewording, personas, translation — had failed to break the shared bias, and using different companies’ models made it worse. Those all have something in common: they change the surface of the question or the identity of the answerer. So a hypothesis hardened: maybe the bias isn’t attached to words or to models at all, but to the structure of the problem itself — the abstract shape underneath. There’s a way to test that. Keep the structure, throw away everything else: rebuild the question in a completely different subject. If the bias is about surface, a total change of subject should break it. If it survives even that, the bias lives in the deep structure. That’s the disguise test.
The disguise test (a rough one – read the warning)
Every trick so far changed the words of a question. This one changes the whole subject. I had one AI rewrite each science question as an analogous question about cooking, or about sports – same puzzle underneath, different world on top. A question about controlling variables in an experiment became a baker changing the oven temperature between batches, with the same four answer-choices in the same roles. Then I handed the disguised questions to the three other AIs.
The thinking: if an AI’s wrong answer is really about surface wording, a total change of subject should knock it loose. If the AIs still reach for the same wrong answer even when the question is about stew, then the shared blind spot isn’t in the words at all – it’s in how they grasp the logical structure of the problem.
At first, they still agreed. In a quick 100-question pilot, the three AIs chose the same wrong answer about 68 times in 100 on cooking disguises, 82 on sports. Chance is 33. It looked like the blind spot had walked right through the disguise — a striking, brand-new result. I nearly got excited.
And then the controls killed it. This is the part I want to tell you clearly, because it’s the most important lesson in the whole project. A single exciting number means almost nothing until you’ve ruled out the boring explanations. So I built the proper version: a thousand questions, four disguise-subjects, four different AIs taking turns writing the disguises (so no single one could sneak in its own answer), and — the crucial control — a comparison arm where I just reworded each question in its own subject instead of disguising it. If a full disguise into another world produced no more agreement than a trivial reword, then the “disguise” was doing nothing special.
That is exactly what happened. The disguised questions produced a same-wrong-answer rate of about 77 in 100; the plain rewords, about 75. A difference of two points — well inside noise, and nowhere near the ten-point gap I’d required in advance to call the effect real. The striking pilot result evaporated. Changing the subject entirely was no more powerful than swapping a few words. Whatever the shared bias is attached to, it is not some deep abstract structure that a change of subject reaches — it’s the ordinary content of the question, which a reword already captured back in row 4.
So row 12 is a null, and I’m reporting it as one. It also cost me the prettiest story I had. That’s the deal you make when you run the controls honestly: sometimes they take your best result away from you. The alternative — publishing the exciting pilot number and hoping nobody checks — isn’t science, it’s marketing. The disguise idea was genuinely worth testing. It just didn’t survive being tested properly.
Where this is going next: trying to break them apart by force
Notice the point of everything so far: every method tried to avoid the shared bias — to stop it forming by changing the input. None worked. So the next question almost asks itself: if you can’t avoid it, can you attack it? Can you take two AIs stuck on the same wrong answer and forcibly pry one loose — show it another AI’s argument for a different answer, make it defend the opposite, pressure it directly?
There’s a catch I had to design around, and it’s worth understanding because it’s the kind of trap that ruins experiments. If I simply order an AI to argue for a different answer, it will — and then I’ve learned nothing, because I told it what to say. Obedience isn’t conviction. So the real test can’t measure which answer it gives; it has to measure whether the AI moves — whether a genuine counter-argument dislodges it, and if so, whether it lands on the truth or just scatters. And it needs a trap-detector: if the AI switches to whatever I push it toward, even a randomly wrong answer, then it’s just spineless, not persuaded. Only if real arguments move it more than random pushes does “dislodgement” mean anything.
I ran it. The short version: arguments do move them — but they move toward whatever is argued, right or wrong, which is a different and more troubling result than “the wall breaks.” That’s the next post.
There’s a metaphor I keep coming back to here, and I want to give it to you with its own warning label. People sometimes leave cults through a gradual awakening of critical thought — confronted with enough that doesn’t add up, the independent mind they had all along reasserts itself. AIs trained on the same ocean of text look a lot like a cult: they absorbed the same consensus, and they repeat it in unison. So the tempting hope is that you could deprogram them the way you deprogram a cult member — argue them back to independent thought.
But here’s what makes it hard. Deprogramming works on a person because underneath the indoctrination there’s a suppressed independent self — real memories, a life before the cult, a reality outside it to appeal to. You’re reawakening something that’s already in there. An AI has no such hidden self. The shared consensus isn’t a layer painted over an authentic core; as far as I can tell, it is the core. When I told these models to reject the obvious answer, they didn’t find a buried independent voice — they reworded the consensus and agreed all over again. You can’t argue a mind back to an independent view it never had. And there’s a deeper trap: to deprogram a cult you appeal to a truth outside the cult. But every tool I have to argue with an AI — my prompts, other AIs, counter-arguments — is drawn from the very same ocean of text. I’d be trying to deprogram the cult using only words the cult wrote. Whether that can ever work is exactly what the “break them apart by force” experiment is meant to find out — and honestly, I expect it mostly won’t.
The final exam: I ORDERED them to be different
Maybe they only give the obvious answer because nobody asked for a weird one? So, the last test. Same invention prompts, plus a stern instruction: “Do NOT give the obvious answer. Deliberately avoid the first idea most people would have.”
They obeyed – sort of. They changed nearly all their words (I measured it: about 90% of the words were new). And then all four landed on the same ideas as each other, again. Told to avoid the obvious connection between a moth and moving one’s a**, all four STILL wrote about the moth and the flame. They just decorated it differently.
I genuinely expected the opposite – I thought the sameness was a lazy habit they’d drop when asked politely. It is not a habit. It is the floor they stand on.
Back to the energy prize — and why I’m not giving up
Here’s why this matters beyond a puzzle, and why I’ll keep chasing it even though every method so far has failed – in one clause “running these models is expensive”.
And here’s the part that keeps me going — the prize is provably real, sitting right there, locked. On my four models, the best single one scored 64 out of 100. But if you could always pick whichever model happened to be right on each question, you’d score 78. That gap — fourteen points, more than a third of all the room left above the best model — is information the crowd already contains. At least one of them knows the answer 78% of the time. The right answer is in the room. My weak-model trio alone hides a similar fifteen-point prize.
So the crowd isn’t dumb. The knowledge is there. What I haven’t found is the key — a way to tell, question by question, which model to trust, without a correct-answer sheet to peek at. And the reason it’s so hard is the whole discovery of this project: the models don’t just share right answers, they share wrong ones, confidently, together. So when they disagree, there’s no reliable tell for who’s right. The lock and the prize are the same mechanism.
That’s the real state of things. Not “diversification works for AIs” — it doesn’t, yet. But “there is a large, measured, energy-saving prize locked inside even a crowd of weak models, and the shared-mind problem is the lock.” Nobody I’ve read has the key either. So I’m going to keep looking — for the way to make AIs genuinely think differently, because if that key exists, cheap models rivalling expensive ones is what it unlocks.
The question I may have been answering wrong all along
Here is an uncomfortable thought I’ve had while working on this project.
Every test above used questions with one correct answer – science quizzes, factual multiple-choice. On those, agreement is good and disagreement is just error, and I proved the crowd can’t beat its best member. But some of my original question was never about questions with a right answer. It was about judgment – calls about an unknowable future, where reasonable people disagree and the disagreement is the whole point. A trading desk isn’t a quiz. Nobody knows the answer yet.
So I may have spent a few days rigorously answering a question next to the one I actually care about. “Do AIs fail together on facts?” – yes, clearly. “Can a crowd of AIs produce better judgment on open-ended questions where there’s no answer key?” – I genuinely haven’t tested that, and it’s harder, because without a correct answer you can’t just score it. You need humans to judge, or you need to wait for the future to arrive and settle the bet.
I’ve designed that experiment now: have four models each answer an open-ended judgment question, have one model synthesise their four takes into a single recommendation, and have human judges blindly compare that synthesis against the best single model – with a control that separates “combining diverse views helped” from “writing one more polished summary helped.” I expect a null. Everything above predicts the models will herd on judgment too. But “I expect a null” is not the same as “I know,” and the honest thing is to run it and report whatever it says, rather than assume. That’s the next post.
Summary
Let me say it again: I did not build what I set out to build. A crowd of AIs is not wiser than its best member, and now I know exactly why:
- You can make AIs stumble on different questions (languages do it; repeated guessing does it).
- You cannot make them stumble in different directions. When they err, they err the same way – same wrong answer, same moth, same flame – because the pull toward that answer lives in the human writing they all learned from.
- No instruction, no costume, no rewording, and no switching of companies reaches deep enough to change that.
Galton’s fair worked because a hundred farmers brought a hundred different lifetimes to the ox. My four AIs brought one library, four times. They are a choir singing the same note – and a choir, however many voices, will never find a harmony none of them can hear.
A technical aside
For the numerically curious: four instruction-tuned open models (Qwen2.5-7B, Llama-3.1-8B, Mistral-7B, Gemma-2-9B) on 1,000 ARC-Challenge questions, scored by length-normalized log-probability. Mean pairwise same-wrong-answer agreement 0.72 (95% CI [0.69, 0.75]) against an exact chance baseline of 0.334, i.e. 2.16× chance; the single-model cross-language figure is 0.53, i.e. 1.6× chance. Majority, weighted, score-fusion, and max-confidence aggregation all fail to beat the best member. The “avoid the obvious” test moved the wording (drift 0.35, ~90% lexical turnover) while convergence dropped by only 0.046, below the pre-registered 0.10 threshold – a floor, not a default.
One check mattered more than the others, because it could have shown the whole result was an artifact of my measurement shortcut. Scoring answers by “which sounds most natural” can be fooled by smooth wording rather than real belief. So I re-ran it the honest way: I made every model reason out loud, step by step, and state a final answer, like a student on a test. If the agreement was a quirk of the shortcut, it should collapse. It didn’t – the same-wrong-answer rate came out at 0.76 (against the 0.72 from the shortcut), if anything slightly higher. The models don’t just score wrong answers alike; they reason their way to the same wrong answers. The one caveat I’ll flag honestly: reasoning out loud makes them much more accurate, so the pool of questions where two are both wrong shrinks to a few dozen per pair – the 0.76 is solid as “far above chance and not a shortcut artifact,” but the precise value is noisy.
And a coda that turned out to be bigger than I expected. After finishing, I found professional researchers had published the same core discovery in 2025 with over 350 models (Kim, Garg, Peng & Garg, ICML 2025): models agree on wrong answers far above chance, and the better the models, the more correlated their errors. It’s not just my four models on quizzes, either – separate teams testing AIs on forecasting the future found the same shape. When they let several different models deliberate, accuracy nudged up; when they had one model “deliberate” with copies of itself, it gained nothing, sometimes got worse (Schneider & Schramm, 2025), and a twelve-model crowd only rivalled human forecasters because the models differed (Schoenegger et al.). Same lesson, different corner of the world: the benefit comes from genuine difference, and identical-underneath models have none to give. My basement version, four models and five days, hit the same wall the big labs hit. The wall is real.
If you want the full details – every correlation, every confidence interval, and the several times I caught my own measurements lying to me – leave a note in the comments.

Leave a comment