
On 20 May 2026, an unreleased OpenAI model produced a counterexample to a conjecture Paul Erdős made in 1946. Nine world-class mathematicians reviewed the construction and found it sound. On 8 September, ten thousand agents working for 88 hours produced a singularity in the three-dimensional Navier-Stokes equations, formally verified in Lean, resolving a Millennium Prize problem. Neither result was in any training corpus, because neither existed.
This episode asks what actually happened. If a model generates a proposition that was never in its training data, is that discovery or hallucination? The uncomfortable answer is that at the moment of generation they are the same operation. Nothing in the sampling process distinguishes a true novel claim from a false one. The partition is imposed afterwards, from outside, by a verifier the model does not run.
We work through the 2026 evidence on both sides. Yan and colleagues at ACL measured the unverifiable output space across 32,400 generations and found that only 4.7 percent of it qualifies as creative synthesis rather than groundless fabrication. Kalai and colleagues at OpenAI argue that models hallucinate because training rewards guessing over admitting uncertainty, which means the disposition to conjecture is optimised for rather than accidental. Meanwhile the computability literature argues that some irreducible error rate is mathematically necessary for any model family we can actually build.
The conclusion is that the ceiling on machine knowledge is not the training corpus. Every confirmed case of machine-originated knowledge this year paired a generative model with a verifier that was not a language model: a proof assistant, a code executor, a cell viability assay. The generator supplies variation and something outside it supplies selection, which is the structure of evolution by natural selection. That makes the event horizon domain-shaped rather than knowledge-shaped. Mathematics has a perfect verifier and is falling quickly. Fields without an oracle will not move by this route, regardless of how important they are.
We close on the edge angle. A model deployed on a device is a proposer operating without a verifier. A robot’s grasp either holds the object or drops it, and that is a selector. This is the strongest available argument that embodiment is not decorative.
This episode is a sequel to “Large Language Monkeys: Why Noise Yields No Knowledge” (S6E7), which argued that a random source contains no knowledge because nothing selects the signal. Here we ask what happens when a selector exists.
Referenced in this episode: Quanta Magazine on the Erdős problems and on Navier-Stokes; Yan et al., ACL 2026; Kalai, Nachum, Vempala and Zhang, arXiv 2509.04664; AlphaEvolve and FunSearch; Shumailov et al., Nature 2024 on model collapse.

The content is really great