A benchmark built from open problems
> ArXiv CS.CL researchers introduced FrontierMath Erdős (FME), a benchmark of 68 Erdős problems that were still open as of August 2026.> To count as solved, an AI must prove or disprove the conjecture in the Lean proof assistant.> The problems were chosen by the paper's second author from 652 open problems listed on erdosproblems.com, selected for mathematical interest and difficulty.> The paper argues that recent AI resolutions of open problems are isolated demonstrations rather than a systematic study of capability — FME is meant to close that gap by evaluating every model on the same fixed problems, autonomously and under the same budget. [1]
Results under a $300-per-problem budget
> Five AIs were evaluated with a budget of $300 per problem.> One model, GPT-6 Astra, scored 3%; every other model scored 0%.> Caveat: this is the paper's own evaluation protocol and scoring, and only five models were tested, so the 3% vs. 0% spread is a narrow snapshot rather than a broad ranking of AI math ability. The abstract does not break down per-problem results or error modes. [1]
Sources
- FrontierMath Erd\H{o}s
ArXiv CS.CL (Computation and Language) · Reporting ·