The rapid progress of large language models (LLMs) raises a fundamental question: how do we determine whether a new model is genuinely more capable than previous versions? As AI systems continue to improve, evaluating their abilities has become increasingly difficult because many traditional benchmarks no longer measure advanced reasoning effectively. This is similar to evaluating a university student using only high school level questions. Traditional benchmarks have been useful for comparing existing models, but newer systems require more demanding evaluation criteria. FrontierScience, introduced by OpenAI, addresses this gap by testing models on challenging scientific reasoning and research-related problems that better represent expert level tasks.
Two Levels of Scientific Reasoning
FrontierScience evaluates AI systems through two different types of scientific tasks: Olympiad level problems and research oriented problems.
The Olympiad section focuses on advanced problems from physics, chemistry, and biology similar in difficulty to international Olympiad competitions. These problems usually have clear solutions that can be expressed through a numerical answer, equation, or short explanation. They test whether a model can apply scientific knowledge and complex reasoning under strict constraints.
The Research section focuses on a different type of intelligence. Instead of searching for a single correct answer, these problems represent situations scientists may encounter during real research. They require models to interpret information, evaluate possibilities, and make scientifically justified decisions.

Building Problems That Challenge Advanced AI
A major difficulty in evaluating powerful AI systems is creating problems that can reveal their limitations. If a benchmark becomes too easy, it stops showing meaningful differences between models. To avoid this, scientific experts and former Olympiad participants created FrontierScience problems designed to reflect genuine scientific reasoning. Olympiad questions capture the difficulty of competition-level challenges, while research problems are inspired by real scientific investigations. The questions also went through careful review to maintain the intended difficulty. Problems that models solved too easily were reconsidered, helping the benchmark remain useful as AI systems continue to improve.
Why Scientific Reasoning Is Harder to Measure
Evaluating scientific answers is more complicated than checking whether a number or expression is correct. In many research situations, several approaches can lead to meaningful solutions, and the reasoning behind an answer matters as much as the final result.
To address this issue, FrontierScience uses a rubric-based evaluation system for research problems. Instead of only checking whether the final answer matches a reference solution, the evaluation examines different parts of the response, including reasoning quality and scientific accuracy.
This raises an important question about AI capability: can models truly reason through unfamiliar scientific problems, or are they mainly reproducing existing knowledge?
Where Current AI Systems Excel and Struggle
The evaluation compared several advanced AI models from different organizations. The results showed that current models perform strongly on structured scientific tasks, especially in chemistry and problems with clearly defined answers.
However, their performance decreases when they face open-ended research problems. This difference reveals an important limitation of modern AI systems. Solving a carefully designed problem is different from conducting scientific investigation, where scientists must work with uncertainty, incomplete information, and unexpected outcomes.
FrontierScience therefore provides a more demanding way to measure AI progress by testing not only what models know, but how effectively they apply knowledge in complex scientific situations.
Conclusion
As artificial intelligence continues to advance, measuring the true capabilities of AI models has become increasingly important. Traditional benchmarks have helped compare models, but many focus on structured questions that measure recognition and information retrieval rather than deeper scientific reasoning.
FrontierScience addresses this limitation by evaluating models through Olympiad level and research level scientific tasks that require complex reasoning, understanding, and problem solving. The results show that while current AI systems perform well on structured scientific problems, they still struggle with open ended research scenarios that require flexibility and judgment.
This presents the importance of developing evaluation methods that measure genuine problem-solving abilities. Future benchmarks should focus not only on what AI systems know, but also on how effectively they can reason, adapt, and support scientific discovery.