Tech • AI • Robotics • Game

VIDEO
ENFR

Daily Podcast full article

OpenProblemBench targets 82 unknowns

OpenProblemBench proposes a sharper test for scientific AI: not another exam with hidden answers, but 82 unresolved problems from mathematics and theoretical physics, scored through model-based review of complete solutions and partial progress. The promise is a benchmark closer to research; the risk is mistaking automated judgment for discovery.

Generated October 9, 2026 at 12:18 PM1370 words
AI-generated illustration

A benchmark aimed at the frontier, not the textbook

OpenProblemBench arrives with a deliberately uncomfortable question: if AI systems are increasingly described as scientific collaborators, should they be evaluated only on problems whose answers are already known? The new benchmark, submitted to arXiv on October 8, 2026, collects 82 unresolved problems in foundational theoretical sciences and asks models to make progress where no reference solution is available . Current research trackers have already placed the paper inside the AI-for-science and scientific-reasoning stream, reflecting how directly it speaks to the debate over whether language models can move beyond recall and recombination .

The subject title is exact: OpenProblemBench targets 82 unknowns. The unknowns are not casual riddles or synthetic puzzles. They are drawn from mathematics and theoretical physics, with each task framed by research context, assumptions, known progress and references . That design tries to avoid a common weakness in “hard” AI tests: a model can be impressive at reconstructing a theorem or solving a contest problem while still giving little evidence that it can operate at the edge of current knowledge.

What is inside the 82-problem suite

The benchmark is split between 46 mathematics problems and 36 physics problems . Its mathematics side includes areas such as number theory, algebraic geometry, topology, analysis, dynamics, combinatorics, optimization and logic; its physics side includes statistical mechanics, integrable systems, quantum information, field theory, gauge theory, optics and mathematical physics . The paper’s problem list stretches from famous long-running questions, including a Goldbach-related item, to much newer research questions .

That range matters. OpenProblemBench is not one uniform test with one difficulty scale. It is closer to a curated map of difficult terrain: some questions demand proof, some demand classification, some require a counterexample, and others require computational or constructive evidence that can survive mathematical inspection. The authors say they selected problems whose proposed solutions admit comparatively clear checks of decisive mathematical or computational claims . In other words, the tasks are open, but the evaluation target is not meant to be entirely subjective.

The benchmark’s construction also signals a shift in what “AI evaluation” is trying to measure. Earlier tests often ask whether a model knows what science already knows. OpenProblemBench asks whether a model can identify what remains unproved, choose a productive line of attack, and avoid turning suggestive evidence into an overconfident conclusion. That is a research skill, not just a question-answering skill.

How do you score a problem with no answer key?

The central methodological problem is obvious: if nobody knows the answer, what counts as a score? OpenProblemBench answers with a multi-judge protocol. Four evaluator models independently assess each submission for correctness, completeness and degree of progress, without using reference solutions . The Cool Papers listing reproduces the same core abstract and publish timestamp, confirming the benchmark’s headline setup: 82 unresolved problems, four evaluator models, seven solver configurations and a top mean judged solve rate for GPT-6-Astra .

The scoring is not binary. Submissions can be judged solved, partial or unsolved, while partial progress is further divided into trivial, nontrivial and breakthrough categories . Nontrivial progress may include a meaningful special family, an improved bound or a useful reduction. Breakthrough progress is reserved for work that removes a central obstacle without fully closing the problem . The benchmark also tracks overclaiming, such as treating a conditional argument as unconditional or extrapolating finite evidence to an infinite class .

That last point is crucial. In open-problem research, a plausible-looking derivation is often not the achievement; the achievement is knowing exactly what it proves. A model that produces ideas but systematically overstates them may still help exploration, but it also increases the verification burden on human scientists.

The first leaderboard is impressive, but not a verdict

Across seven evaluated configurations, the paper reports GPT-6-Astra at the highest mean judged solve rate, 14.0% in the abstract and 14.02% in the detailed results . The evaluated full-size open models are reported at 5.5% to 6.7%, while Flash models are reported at 2.4% to 3.7% . ArXiv Troller’s paper page also lists the paper as submitted on October 8 and last updated on October 9, 2026, keeping the public record within the current 72-hour reporting window .

The raw solve-rate headline should not be read as “AI has solved 14% of a set of famous open problems.” The paper is more careful than that. Its results are judge-assessed outcomes, not journal-accepted discoveries. The authors explicitly frame OpenProblemBench as a way to investigate AI capabilities and limitations as contributors to theoretical science, not as a substitute for expert validation .

The more interesting finding may be the amount of partial progress. The study evaluates 574 final submissions and produces 2,296 reviews . It reports that solved judgments remain rare outside the leading configuration, but that several models receive substantial numbers of nontrivial or breakthrough partial judgments . That is closer to how research actually works: a useful reduction, a corrected representation or a new computational certificate may matter even when the original problem remains open.

Why representation, proof gaps and overclaiming matter

The paper’s case comparisons are valuable because they move beyond the scoreboard. Stronger outcomes are associated with changes in problem representation, arguments that generalize beyond finite evidence, and proof steps that close the obligations needed for a full solution . Weaker outcomes often leave exactly the kinds of gaps mathematicians and theoretical physicists care about: unproved lemmas, exceptional parameter regimes, finite checks presented as general proof, or computational pipelines that agree with each other while sharing the same error .

This is where OpenProblemBench feels most distinct from conventional exams. In a known-answer benchmark, a system can sometimes stumble into the right final response. In an open-problem benchmark, the path matters because there is no authoritative answer to compare against. The evaluator must ask whether the argument actually covers the assumptions, quantifiers and edge cases in the problem statement.

That makes the benchmark both ambitious and fragile. It is ambitious because it tries to score the texture of research progress. It is fragile because automated judges can share blind spots, especially when evaluating advanced mathematics or theoretical physics. A consensus among model judges is useful evidence, but it is not peer review.

The contamination question

OpenProblemBench also enters a world where benchmark contamination is no longer a theoretical worry. If the problems are drawn from the literature, then models may have seen the source papers, related conjectures or partial solutions during training. The benchmark tries to address this by focusing on unresolved questions rather than known targets, and by packaging context for each problem . But contamination can still appear in subtler forms: a model may reproduce a known partial route, cite familiar terminology, or lean on memorized fragments without actually advancing the argument.

The authors’ emphasis on checkable decisive claims helps, but it does not eliminate the issue. For future releases, the benchmark’s credibility will depend on transparent literature cutoffs, frozen task versions, clear scoring rubrics and, ideally, independent expert audits. Without those controls, a high score could reflect a mixture of genuine reasoning, lucky reconstruction and evaluator generosity.

The final boss is still peer review

OpenProblemBench’s strongest contribution is not the claim that today’s AI systems can solve open science at scale. It is the insistence that research-grade AI should be tested against research-grade uncertainty. A benchmark built around 82 unknowns forces models to confront the difference between evidence and proof, between a special case and a theorem, and between a persuasive sketch and a result that can survive expert scrutiny.

For now, the correct reading is measured: OpenProblemBench raises the bar for machine-discovery claims, shows a possible way to grade partial scientific progress, and reports that leading systems can sometimes make judge-assessed headway on unresolved theoretical problems . But the final boss is not the leaderboard. It is independent verification by the communities that own these problems. There is no walkthrough available.

Comments

Be the first to comment.

Sources from the last 72 hours

  1. [1]OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical SciencesOct 8, 2026, 4:43 AM
  2. [2]OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences | Cool Papers - Immersive Paper DiscoveryOct 8, 2026, 4:43 AM
  3. [3]OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences - arXiv TrollerOct 9, 2026, 2:00 AM
  4. [4]AI for Science research | The Latest in AIOct 8, 2026, 2:00 AM

AI-generated article based on recent web research, then preserved as a dated editorial snapshot.