AlphaProof Nexus: DeepMind AI Solves 56-Year Erdős Problem
Google DeepMind introduces AlphaProof Nexus, combining Gemini 3.1 Pro and Lean language to solve a 56-year-old Erdős math problem at low cost.

Stock photo for illustration only, not from the actual event
- DeepMind launches AlphaProof Nexus pairing Gemini 3.1 Pro with the Lean language
- Successfully solved 9 out of 353 open Erdős math problems including a 56-year-old riddle
- All proofs are rigorously checked by a computer compiler with zero blind trust
- Processing cost runs at just a few hundred dollars per problem or roughly $60 per TPU
Mathematical problems left behind by Hungarian mathematician Paul Erdős that remained unanswered for up to 56 years have finally been closed by an artificial intelligence system operating at a cost of just a few hundred dollars per problem. More important than the answers themselves is the fact that every single line of the proof has been machine-verified, eliminating the need to blindly trust that the AI did not hallucinate.
Google DeepMind recently published a research paper titled Advancing Mathematics Research with AI-Driven Formal Proof Search on arXiv to introduce AlphaProof Nexus, a framework pairing the Large Language Model Gemini 3.1 Pro with Lean, a programming language for computer-checkable mathematical proofs. The system successfully resolved 9 out of 353 open Erdős problems, proved 44 out of 492 conjectures from the Online Encyclopedia of Integer Sequences (OEIS), and closed algebraic geometry questions that had stalled for roughly 15 years.
A major obstacle in utilizing LLMs for mathematical research has been reliability, since natural language proofs often contain subtle errors that require manual expert review. To counter this, DeepMind chose to have the AI write formal proofs using the Lean language, structuring proofs like computer code so that the Lean compiler can verify every logical step. Any compiler error messages are immediately fed back to the AI to refine its output in the subsequent iteration.

Stock photo for illustration only, not from the actual event
AlphaProof Nexus operates by taking a proof sketch in Lean with gaps left intentionally blank by researchers, allowing AI agents to fill in those gaps. The team experimented with four agent configurations: a baseline running multiple independent Gemini 3.1 Pro instances, a second version integrating reinforcement learning-trained AlphaProof, a third incorporating an evolutionary mechanism ranked by Elo scores, and a fully equipped agent combining all features used for open-ended problems.
"From specialized systems requiring custom training to simple agent loops."
Google DeepMind Research Team
Integrating the Lean language with LLMs represents a pivotal shift from mere probabilistic text generation to rigorous formal logic verification. Compiler feedback loops directly constrain the AI against hallucinations, allowing the mathematical research community to place genuine confidence in theoretical computer-generated outputs.
The tested Erdős problems originated from ErdosProblems.com, including problem number 125 left open since 1996 concerning base-3 and base-4 number sets, which the AI proved could not exceed zero density. Computationally, each problem consumed about 27.5 hours on a Tensor Processing Unit, amounting to roughly 60 US dollars in processing expenses.
Source: Techsauce
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment