Epoch grows open-math benchmark as AI solves two more
Epoch AI expanded its FrontierMath Open Problems set to 50 research questions and logged two fresh AI solutions, both in the middle "Moderately Interesting" tier. Three problems were pruned in the same update.
Epoch AI has updated FrontierMath: Open Problems, its benchmark of mathematics questions that remain genuinely unsolved by human researchers, expanding the roster to 50 problems and recording two new solutions credited to AI systems. Both fall in what Epoch calls the "Moderately Interesting" tier, the middle band of a significance scale maintained with input from an editorial board of mathematicians.
One of the newly cleared problems asked for superpermutations shorter than the best known constructions, over 8, 9 and 10 symbols. Epoch says an AI system achieved this during the past week and that the improved construction currently stands. The second asked for a genus 2 curve over the rationals carrying a rational torsion point of prime order at least 31; Epoch notes that such a curve had in fact already been found by human mathematicians, which limits how much the AI result should be read as novel discovery.
Across the full set, Epoch's tracker now marks three problems solved: the two above, plus an entry on a presentation of the absolute Galois group of Q₂, which sits one tier higher under "Solid result". The tiers above that — "Major advance" and "Breakthrough" — remain untouched.
Epoch also pruned the collection in the same update, removing three problems. Two were dropped because their verifiers could not reliably confirm whether a submitted solution was correct — one concerning a surface with a high number of singularities, the other an algorithm for deciding whether a knot has unknotting number 1. The third was removed for failing to clear the minimum notability threshold Epoch wants the set to enforce. It is housekeeping that doubles as an admission of how hard it is to keep a fixed benchmark of "open" problems both open, meaningful, and reliably gradeable once capable solvers are pointed at it.
The design tension is the point of the exercise. Unlike FrontierMath's graded tiers of unpublished-but-solved problems, Open Problems has no reference answers; a solution is scored by whether the mathematical community accepts it. That makes the set a slow, low-throughput signal, but one that cannot be gamed by memorisation or contaminated by training data, and one where each resolved item is permanently consumed.
Why it matters
Most claims of AI mathematical discovery arrive as vendor announcements without independent adjudication. Epoch's set supplies a neutral, mathematician-refereed ledger of which research questions have actually fallen and how notable each was — including the deflationary detail that one of this batch had a prior human solution. As labs increasingly market autonomous research capability, a public register that distinguishes real firsts from rediscoveries becomes the thing buyers and policymakers can check against.