Google's Gemini Agents Produce New Proofs for Five Unsolved Math Problems
Key Takeaways
- •AlphaProof Nexus, a DeepMind framework powered by Gemini 3.1 Pro, runs multiple independent prover subagents that reason and revise iteratively until a proof passes automated machine-checkable verification.
- •The framework resolved 9 of 353 previously unsolved Erdős problems, including two that had resisted mathematicians for 56 years, and proved 44 of 492 conjectures listed in the Online Encyclopedia of Integer Sequences.
- •Each solution came at a computational cost of only a few hundred dollars, and domain experts who reviewed the mid-2026 arXiv paper confirmed both the proofs and the faithful formalization of the original conjectures.
- •Aletheia, a semi-autonomous research assistant built on Gemini Deep Think, has evaluated roughly 700 open Erdős problems since its December 2025 deployment and produced solutions for 13, four of which were novel autonomous resolutions.
- •Gemini Deep Think scored 35 of 42 points at the 2025 International Mathematical Olympiad, reaching gold-medal standard by solving five of six problems within the standard time limits.

Google's Gemini-powered mathematics agents have produced proofs for five previously unsolved math problems, using teams of AI agents whose work is graded by unforgiving automated checkers.
The best-documented implementation of this approach is a framework DeepMind calls AlphaProof Nexus. It runs multiple independent prover subagents in multi-turn reasoning loops, powered by Gemini 3.1 Pro. The agents do not produce a single answer and stop there; they reason, revise, and retry over several rounds until a proof either passes the automated checkers or fails. That design means acceptance hinges on machine-checkable verification rather than a model's confidence in its own output.
The numbers behind the results
AlphaProof Nexus resolved 9 of 353 previously unsolved Erdős problems — open questions posed by Paul Erdős, the prolific 20th-century Hungarian mathematician known for attaching cash prizes to his problems — two of which had stumped mathematicians for 56 years. The same framework proved 44 of 492 conjectures listed in the Online Encyclopedia of Integer Sequences, known as OEIS, a long-running reference where mathematicians catalog integer patterns. The results came at a computational cost of only a few hundred dollars per problem.
A detailed arXiv paper describing the work was posted around mid-2026. Domain experts reviewed the results and confirmed both the proofs and the faithful formalization of the original conjectures. Formalization refers to translating a human-written math problem into precise machine-checkable language, and an imprecise translation can mean proving the wrong statement entirely. That expert sign-off matters: it addresses the recurring question of whether AI-generated mathematics is genuinely correct or merely persuasive.
Aletheia and the Olympiad record
A separate system, Aletheia, built on Gemini Deep Think, operates as a semi-autonomous research assistant. Since its deployment in December 2025, Aletheia has evaluated approximately 700 open Erdős problems and identified or created solutions for 13 of them, earning praise from expert reviewers. Four of those 13 were novel autonomous resolutions.
Gemini Deep Think also holds a competition track record. At the 2025 International Mathematical Olympiad, the annual contest widely regarded as the toughest mathematics competition for pre-university students, it scored 35 out of 42 points, reaching gold-medal standard, and solved 5 of 6 problems within the standard time limits that human contestants face.
Why the agentic approach is drawing attention
Pushmeet Kohli and other DeepMind researchers have emphasized that the agentic approach could extend beyond pure mathematics. They identified fields such as combinatorics and quantum optics as potential targets, and the research also highlights applications in algebraic geometry. The common thread is that these are domains where a claimed result can be tested against objective criteria — exactly the conditions under which this agent-plus-verifier setup has demonstrated measurable output.
What it means for mathematics and AI
The findings indicate that pairing agent-based reasoning with automated verification may have applications well beyond the problems studied, though the research notes that current AI systems still require human oversight for novelty assessment, meaning people must confirm whether a result is genuinely new. That human-in-the-loop requirement frames what to watch from here: how the framework fares in the fields DeepMind has flagged, and whether expert reviewers continue to certify its output, will shape how much of this momentum carries into wider research practice.