Full article — scored 10/10
OpenProblemBench benchmark assesses AI progress on 82 unresolved scientific problems
A new arXiv preprint introduces OpenProblemBench, a benchmark designed to test whether AI systems can make meaningful progress on scientific questions that do not yet have known answers. The suite spans 82 unresolved problems in mathematics and theoretical physics, uses multiple AI evaluators rather than reference solutions, and reports that the strongest tested configuration leads on judged solve rate while other models often show partial but incomplete research progress [1].
A benchmark aimed beyond the answer key
OpenProblemBench enters the AI-evaluation landscape with a deliberately uncomfortable premise: if the most consequential scientific problems are unsolved, then a benchmark for research-capable AI cannot rely only on questions whose answers are already known. The paper, submitted to arXiv on October 8, 2026, frames this as a test of whether AI can contribute at the boundary of established knowledge rather than merely reconstruct it .
The benchmark contains 82 unresolved research problems drawn from mathematics and theoretical physics. Each item is packaged with background, assumptions, prior progress, references, and a statement of what a resolution would have to establish . That structure matters because the task is not simply to produce a final number or a short proof. A system must identify what remains open, choose a productive route, distinguish evidence from proof, and avoid overstating what its argument actually shows .
The subject is narrow but strategically chosen. The authors say they selected problems whose proposed solutions can be checked with comparatively clear mathematical or computational criteria: proof steps can be inspected, constructions can be tested against assumptions, and consequential calculations can be reproduced . In other words, OpenProblemBench tries to occupy a difficult middle ground: problems remain open, but claims of progress should still be auditable.
What is inside OpenProblemBench
The dataset is split between 46 mathematics problems and 36 physics problems . The mathematics side includes areas such as arithmetic, algebraic geometry, topology, analysis, dynamics, combinatorics, optimization, and logic, while the physics side includes statistical mechanics, integrable systems, quantum information, gauge and field theory, optics, and mathematical physics .
The time span of the questions is also wide. The paper says one item traces back to Goldbach’s 1742 problem, while another comes from work first posted in 2025 and published in 2026, meaning the benchmark ranges from relatively recent conjectures to problems with centuries of history . That breadth is important for interpretation: the benchmark is not a uniform exam. It is a collection of research-frontier targets whose difficulty, disciplinary style, and verification burden vary sharply.
The authors describe the problem construction as literature-grounded. Candidate questions were identified from literature-derived information, then checked against source papers and subsequent literature to assess whether they remained open at the time of curation . The retained statements were reorganized into self-contained task descriptions, with solvers receiving the problem background, significance, current progress, and references .
That design reflects a central tension in evaluating AI for science. A model cannot be expected to solve an open problem from a one-line prompt, but giving too much context risks turning the task into a guided exercise. OpenProblemBench’s compromise is to supply the scientific frame while leaving the system responsible for the research move: the proof, counterexample, classification, reduction, computation, or argument that would actually advance the question.
How the benchmark judges progress without known answers
Because these are unresolved problems, OpenProblemBench cannot score outputs against a conventional answer key. Instead, four evaluator models independently assess each frozen submission for correctness, completeness, and degree of progress, without access to a reference solution . The four evaluators are GPT-5.6-Sol, Kimi-K3, GLM-5.3, and Qwen3.8-Max .
The evaluation protocol asks each evaluator to identify the obligations imposed by the problem, check assumptions and quantifiers, inspect decisive proof steps, and assess supporting computations . The evaluator must judge the submitted argument as written, rather than filling in missing lemmas or importing another reviewer’s conclusion . This is a crucial safeguard because open-problem work often fails in precisely the gap between a plausible idea and a complete proof.
The output categories are also more nuanced than right or wrong. A review can classify a submission as solved, partial, or unsolved; partial results are then divided into trivial, nontrivial, and breakthrough levels . “Nontrivial” progress can mean a meaningful family, improved bound, or useful reduction, while “breakthrough” progress is reserved for work that removes a central obstacle and substantially changes the remaining task . The benchmark also tracks overclaiming, such as treating a conditional result as unconditional or extrapolating finite evidence to an infinite family .
This scoring scheme makes OpenProblemBench less like a standardized test and more like a structured peer-review simulation. It asks not only whether an AI system reaches a solution, but whether it can make honest, scoped, reusable progress when complete resolution is unavailable.
Reported model results
The initial study evaluates seven solver configurations across all 82 problems, producing 574 final submissions and 2,296 reviews . The tested configurations include GPT-6-Astra, GPT-5.6-Sol, Kimi-K3, GLM-5.3, Qwen3.8-Max, GLM-5.3-Flash, and DeepSeek-V4.1-Flash, run through different execution environments described in the paper .
GPT-6-Astra leads the benchmark by mean judged solve rate. The arXiv abstract reports a 14.0% mean judged solve rate for GPT-6-Astra, compared with 5.5% to 6.7% for the evaluated full-size open models and 2.4% to 3.7% for Flash models . In the detailed HTML version, the mean judged solve rate is given as 14.02%, and Astra also has the largest average share of outcomes rated solved, breakthrough, or nontrivial, at 84.45% .
The picture below the leader is more complicated. The paper reports that solved judgments remain rare for every other configuration, with mean rates from 2.44% to 7.01% . Yet solve rate alone hides useful differences: Qwen3.8-Max and Kimi-K3 receive solved or substantive partial judgments in 73.48% and 74.09% of reviews, compared with 63.11% for GPT-5.6-Sol . Kimi-K3 also has the largest breakthrough-partial segment, at 6.10%, suggesting that some of its best attempts may remove major obstacles without completing the full problem .
The Flash-model results are also not simply a story of weaker performance. GLM-5.3-Flash has fewer solved judgments than GLM-5.3 on average, 3.66% versus 5.49%, but their combined solved, breakthrough, and nontrivial shares are close, at 63.72% and 62.80% . That suggests smaller or faster configurations may still generate valuable intermediate work, even if they less often close a full argument.
Why partial progress matters
For research evaluation, partial progress is not a consolation prize. In mathematics and theoretical physics, a useful reduction, exact special case, counterexample route, or computational certificate can reshape a problem even when it does not finish it. OpenProblemBench explicitly tries to measure this intermediate layer.
The paper reports 2,288 completed outcome assessments out of 2,296 primary reviews . Among the 2,095 partial judgments, 768 are trivial, 1,252 are nontrivial, and 75 are breakthrough outcomes . These are review-level counts, so the same submission may be judged differently by different evaluators, but the distribution shows why a pure solve-rate leaderboard would be too blunt for this domain.
The case studies emphasize recurring distinctions between strong and incomplete attempts. In some examples, successful systems reformulate a problem in a way that directly addresses the required conclusion, while other systems build computational pipelines or simulations that stop short of proof . In others, a model’s finite evidence must be converted into a general argument covering all admissible cases . A third pattern is gap closure: a submission may appear convincing until an unproved normalization, injectivity condition, exceptional case, or correspondence is required to complete the solution .
These are not minor editorial issues. They are the core of scientific reasoning. A system that can propose ideas but cannot state their scope accurately may accelerate exploration while also increasing the burden on human reviewers. By measuring overclaims, OpenProblemBench treats reliability of self-reporting as part of research competence .
The role of tools, web access, and execution settings
OpenProblemBench also shows that model performance is not just a property of model weights. The paper includes a comparison of Qwen3.8-Max under different harness and web-access settings on the same 82 statements, with GLM-5.3 used as evaluator .
In that comparison, Qwen3.8-Max in OpenCode without web access receives 7 solved judgments and 68 nontrivial or breakthrough partial results . In Claude Code without web access, it receives 5 solved judgments and 59 nontrivial or breakthrough partial results . Enabling web access in Claude Code leaves the solved count at 5 but raises substantive partial results back to 68, including more breakthrough outcomes .
The implication is not that web access automatically solves open problems. In this specific comparison, web-enabled solving did not increase the overall judged solve rate, but it did support more substantive progress . The result underscores a broader point for AI science benchmarks: agents, tools, retrieval policies, time budgets, and execution environments are part of the system being evaluated.
Limits of the claim
The authors are explicit that OpenProblemBench’s results are model judgments, not final scientific validation. They say independent expert review is needed to establish correctness and novelty, and warn that evaluator models may share errors . Agreement among AI judges therefore should not be mistaken for a proof that a problem has been solved.
The paper also notes that the benchmark’s 82 questions are uneven across subfields, partly because the checkability criterion favors certain kinds of mathematics and mathematically intensive theoretical physics . The authors call for future releases to specify literature cutoffs, add human-researcher baselines, extend coverage to other sciences, and use repeated runs under common budgets .
Those caveats are not incidental. They define the benchmark’s current status. OpenProblemBench is best read as a new instrument for studying AI research behavior, not as a certified list of machine-made discoveries.
Why it matters
OpenProblemBench reflects a shift in AI evaluation from knowledge recall and contest-style problem solving toward research-frontier assessment. Its central contribution is not only the 82-problem dataset, but the evaluation model: score complete solutions, grade partial progress, diagnose overclaiming, and compare how systems approach the same unresolved question.
If the benchmark holds up under expert scrutiny, it could help separate models that merely sound research-capable from systems that can produce checkable advances. Just as importantly, it could reveal where AI fails: confusing finite evidence with proof, missing exceptional cases, relying on shared implementation errors, or overstating conditional claims.
For now, the headline result is measured but significant. OpenProblemBench reports that current AI systems can sometimes make judged progress on unresolved mathematics and theoretical-physics problems, with GPT-6-Astra leading the tested configurations, but complete solves remain uncommon and require independent verification . That is not the end of human science. It is a sharper way to ask how far AI has really moved toward participating in it.
Sources from the last 72 hours
- [1][2610.11118] OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical SciencesOct 8, 2026, 4:43 AM
- [2]OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical SciencesOct 8, 2026, 2:00 AM
- [3]OpenProblemBench: Benchmarking AI on Open Problem… - arXivOct 8, 2026, 2:00 AM
AI-generated article based on recent web research, then preserved as a dated editorial snapshot.
