Epoch AI's Game Puzzle Benchmarks Leave Top AI Models Below 60%, Exposing Reasoning Gaps
Key Takeaways
- •Epoch AI released two benchmarks comprising 100 programmatically generated puzzles each, designed to prevent memorization and measure authentic reasoning without permitting agent scaffolding, chain-of-thought prompting, or tool use.
- •On Mystery Game Puzzles, the highest-performing model achieved only 59% while open-weight models plateaued at 38%, revealing a significant gap between closed-source and open-weight systems on reasoning-intensive tasks.
- •Chess Puzzle performance improved from 37% with GPT-5 at launch in December 2025 to 54% with subsequent model releases, indicating measurable but still incomplete progress on independent reasoning.
- •A related evaluation called EBR-bench showed limited AI improvement in learning from repeated interactions as of July 2026, reinforcing that interactive reasoning remains a persistent challenge.
- •The performance differential between open-weight and closed-source models raises questions about whether architecture improvements alone can close the reasoning gap or whether proprietary training methodologies play a decisive role.

Epoch AI, a nonprofit research institute, has introduced two game-based evaluation tools — Mystery Game Puzzles and Chess Puzzles — aimed at rigorously testing the reasoning capabilities of leading artificial intelligence models. The initial findings reveal significant limitations in how well today's AI systems handle novel problem-solving scenarios, a result that arrives as multiple frontier labs increasingly market their latest models around reasoning and planning proficiency.
The standout result: the highest score achieved on the Mystery Game Puzzles benchmark is just 59%, while open-weight models reach no more than 38%.
How the Benchmarks Work
Each benchmark comprises 100 programmatically generated puzzles. The Chess Puzzles follow a relatively conventional format — given a specific board position, the model must identify the best move.
The Mystery Game Puzzles take a fundamentally different approach. In this variant, the AI is not told which game it is playing. The game's identity is intentionally concealed, eliminating any reliance on memorized patterns or training data shortcuts. This design directly addresses what Epoch AI identifies as a contamination problem in traditional AI benchmarks, where models often train on internet datasets that include the very tests used to evaluate them.
That contamination concern has grown more pressing as widely used benchmarks such as MMLU and GSM8K have seen top scores cluster near ceiling, narrowing their ability to differentiate between leading models. Evaluations that resist memorization have become scarce, making programmatically generated tests increasingly valuable for honest capability assessment.
By programmatically generating puzzles and obscuring the game format, Epoch AI removes the safety net of pattern recognition, compelling models to exhibit authentic spatial reasoning and planning capabilities.
Responses are evaluated using normalized move notation measured against an established answer key. The benchmarks do not permit agent scaffolding, chain-of-thought prompting techniques, or tool use.
Current Performance Results
On the Mystery Game Puzzles, no model has exceeded 59%. Open-weight models plateau at 38%, underscoring a notable performance gap between frontier closed-source systems and their open-weight counterparts.
The Chess Puzzles present a somewhat more encouraging picture. GPT-5 recorded a score of 37% when the benchmark was introduced in December 2025. Subsequent model releases have since raised that figure to 54%, reflecting measurable improvement within a relatively compressed timeframe.
However, even the 54% chess puzzle result underscores a critical distinction: proficiency at chess with access to search-based tools is fundamentally different from reasoning through unfamiliar positions independently. Traditional chess engines such as Stockfish achieve superhuman play through exhaustive search trees and evaluation functions, not through the kind of generalized reasoning that language models are expected to demonstrate. A model scoring 54% without tools is being measured against a fundamentally different standard than engine-assisted play.
Neither benchmark appears to be approaching saturation — a notable property at a time when many established AI evaluations can no longer meaningfully separate frontier models from one another.
Broader Implications
Epoch AI operates a wider benchmarking hub that includes the Epoch Capabilities Index (ECI), which aggregates model evaluations across mathematics, coding, and gameplay. The ECI is structured to deliver rapid capability estimates without depending on any single test.
A related evaluation called EBR-bench, based on the game Earthborne Rangers, showed limited AI improvement in learning from repeated interactions as of July 2026. This result reinforces the observation that interactive reasoning remains a persistent challenge for AI systems.
The climb from 37% to 54% on chess puzzles over several months indicates that progress continues. Yet the mystery game results suggest that generalized reasoning is advancing at a slower pace than promotional materials from AI companies might imply.
The 38% versus 59% gap between open-weight and closed-source models on reasoning-intensive tasks highlights a tangible competitive disparity. For decentralized AI projects that depend on open models, this performance differential represents a structural disadvantage that additional compute infrastructure alone is unlikely to resolve. The gap also raises questions about whether open-weight development can close the reasoning deficit through architecture improvements alone, or whether proprietary training methodologies and data curation play a decisive role that current open licensing cannot replicate.