Recent research reveals that existing methods for evaluating AI's physics skills were fundamentally flawed. When experts re-graded the tests manually, it turned out AI had already pushed its physics problem-solving abilities to the maximum.
Imagine this: you took a physics exam when you were a student, and even though your derivation, formulas, and final answers were perfect, your teacher gave you a zero, calling it “wrong.” You would feel aggrieved, right? Recently, something exactly like this has been happening in the AI industry. We have long believed that “latest AI models still struggle with physics,” but it turns out the evaluation criteria themselves were “broken.”
Why Does This Matter?
| Whether AI solves physics problems well or not carries significance beyond mere test scores. Physics problems are on a different level than simple word alignment. To solve them, one needs complex reasoning: grasping what is happening (identifying physical models), weighing necessary prerequisites (setting assumptions), and choosing appropriate mathematical formulas to calculate accurately ([Source: HowGoodAreFrontierModelsatPhysics? | alphaXiv](https://www.alphaxiv.org/abs/2609.13009)). |
In other words, physics is an excellent benchmark for judging whether AI is truly “thinking” rather than simply “copying” data. If AI can solve physics problems perfectly, we can expect that it understands the principles of the real world to some degree.
Making It Easy: The ‘Broken’ Answer Key
Why did we think AI was bad at physics until now? To put it simply, it’s like having AI solve a very difficult puzzle, but the way that puzzle was graded was a mess.
| John Sous of Yale University pointed out that existing physics benchmarks (performance evaluation tools) consistently marked AI as incorrect even when it provided the right answer ([Source: Howgoodarefrontiermodelsatphysics? | Hacker News](https://news.ycombinator.com/item?id=49731620)). It is as if an answer key required writing “5,” but penalized you for writing “five.” |
| Upon discovering this issue, the research team had humans manually re-grade the AI’s answers. The results were startling. The latest AI, known as “frontier models” (the cutting edge of current technology), had already achieved the maximum possible scores on existing physics evaluation tools (saturation) ([Source: 2609.13009v1 | How Good Are Frontier Models at Physics?](https://arxiv.org/abs/2609.13009), [Source: How Good Are Frontier Models at Physics? | arXiv](https://arxiv.org/abs/2609.13009)). |
To use an analogy, we thought we had a child who could only solve simple elementary school math problems, but it turned out they already possessed the skills to solve university-level problems.
Where Do We Stand Now?
Does this mean AI has mastered all the physics in the world? Not necessarily. According to the research, AI models are still adept at “recognition,” such as looking at an object and identifying where it is, but they can still show limitations in “physical reasoning,” such as predicting how that object will move in the future (Source: Does Physics Knowledge Emerge in Frontier Models? — Lacuna).
In other words, knowing an apple is there is one thing, but calculating the exact trajectory it will take and where it will land when dropped is another. High recognition accuracy does not immediately translate into deep physical logic (Source: Does Physics Knowledge Emerge in Frontier Models? — Lacuna).
What Happens Next?
| John Sous described this phenomenon as “a bit scary” ([Source: Howgoodarefrontiermodelsatphysics? | Hacker News](https://news.ycombinator.com/item?id=49731620)). While we fail to create proper tools to measure AI’s capabilities, AI might already be far ahead of what we imagine. |
In the future, we must move beyond simple problem-solving-oriented evaluations and introduce more sophisticated methods that measure how logically AI understands and applies laws of physics. Now that the speed of technological development has completely outpaced our evaluation systems, we need a precise “magnifying glass” to look into AI as much as we need the technology to handle it.
References
- [2609.13009v1] How Good Are Frontier Models at Physics? (https://arxiv.org/abs/2609.13009v1)
-
How Good Are Frontier Models at Physics? alphaXiv (https://www.alphaxiv.org/abs/2609.13009) -
How good are frontier models at physics? Hacker News (https://news.ycombinator.com/item?id=49731620) - [2609.13009] How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks (https://arxiv.org/abs/2609.13009)
- [Paper Note] How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks (https://github.com/AkihikoWatanabe/paper_notes/issues/6574)
- How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks (https://arxiv.org/html/2609.13009)
- Does Physics Knowledge Emerge in Frontier Models? — Lacuna (https://lacuna.tiptreesystems.com/work/does-physics-knowledge-emerge-in-frontier-models/wrk_5011f7d6acc6fe82d5198eb25de3c61d)
- AI finds them too easy and gets bored
- They grade correct answers as incorrect
- They do not include physics formulas at all
- To simply test memorization skills
- Because it requires complex reasoning skills such as selecting physical models, setting assumptions, and deriving equations
- Because the goal is for AI to become a physicist
- They remain at a basic level
- They show average performance
- They are excellent enough to reach the maximum score of existing benchmarks