Blog
Why multiple AI models beat one, explained through the actual research.
Why a Single AI Grading Its Own Work Can't Be Fully Trusted
A 2025 study found GPT-4 rates its own answers far more favorably than humans do — and the reason has nothing to do with being right.
August 22, 2026 · 6 min read
The Research Behind Letting AI Models Argue With Each Other
MIT and Google Brain researchers found that language models debating each other lift math accuracy from 67% to 82% and cut factual errors substantially — but only certain kinds of mistakes get fixed this way.
August 22, 2026 · 7 min read
More AI Answers Can Beat a Bigger AI Model, Research Shows
Tencent researchers found that a smaller model asked repeatedly and voted on its own answers can match — or beat — a much larger model asked just once.
August 22, 2026 · 6 min read
Why Group Brainstorming Fails — and AI Doesn't Have the Same Problem
A 1987 study found that brainstorming together, rather than alone, cuts the number of ideas nearly in half — for a reason that AI participants don't experience at all.
August 22, 2026 · 7 min read
Stanford's AI Research Team Included a Dedicated Skeptic. It Worked.
A Stanford/Chan Zuckerberg Biohub team built specialized AI agents — including one whose only job was to criticize the others — to design real, lab-tested COVID nanobodies. Over 90% worked.
August 22, 2026 · 6 min read
When More AI Agents Make Things Worse — and How to Avoid It
A 2026 study found that giving AI agents distinct expert personas can produce the least diverse ideas of any setup tested — and pinpoints exactly which structures cause it, and which prevent it.
August 22, 2026 · 8 min read