
Anthropic gave thirty AI agents the same coding task. Eighteen of them named their git branch exactly the same thing.
That result opens one of the strangest findings of 2026: when teams of AI agents face the classic “hidden profile” experiment, where the shared evidence points to the wrong answer and the decisive facts are scattered across individuals, they succeed only 17 to 36 percent of the time. Human groups in the original 1985 study? 18 percent. Forty years, a completely different kind of mind, the same failure.
In this episode we dig into why. We trace the mathematics of groupthink from information cascades to Condorcet juries, then follow the trail into territory embedded engineers know well: the 1986 Knight and Leveson experiment that shattered the independence assumption in N-version software, Airbus’s dissimilar redundancy, and the Lufthansa flight where two frozen sensors outvoted the one telling the truth. Along the way, honeybees show us a working reference design: a two-milligram brain that refuses to repeat a rumor.
We close with the fixes: engineered dissenters, forced disclosure protocols, reputation infrastructure for agents, and the case for “keeping the weirdness alive” through a genuinely diverse AI ecosystem, including heterogeneous fleets of small models at the edge.

Leave a Reply