Summary
Domain experts re-grade six physics benchmarks for GPT-5.6 Sol, Fable 5 and Gemini 3.1 Pro, separating model errors from broken questions, reference solutions and graders. On audited public-source cases initially marked wrong, 148 of 152 are benchmark or grader errors. After repair or exclusion, GPT-5.6 Sol rises from 47.3% to 78.7% mean@4 on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark; corrected pass@4 on retained CritPt tasks reaches 94.4%. The corrected scores use retained or repaired subsets, so they are not a direct like-for-like leaderboard replacement.
Why it matters
For teams using scientific benchmarks to route models or claim capability gaps, evaluator and reference quality are now large enough to dominate the measured error. High-stakes benchmark suites need expert adjudication, equivalent-answer handling and retained-set reporting. Near-saturation of repaired closed-ended tasks also means new evaluations must test open-ended modeling and experimental judgment.
Technical details
| Benchmarks | HLE-Physics, CMT-Benchmark, CritPt, PHYBench, PRISM-Physics and UGPhysics |
|---|---|
| Models | GPT-5.6 Sol, Claude Fable 5 and Gemini 3.1 Pro |
| Public Source Audit | 148 of 152 initially rejected cases (97.37%) attributed to benchmark or grader errors |
| Gpt 5 6 Sol | HLE-Physics mean@4 47.28% to 78.66%; CMT 61.00% to 87.24%; retained CritPt mean@4 87.50% and pass@4 94.44% |
| Evaluation Warning | pre-audit and corrected scores often use different question sets after repair or exclusion |
| Availability | paper and detailed audit tables; no code repository linked |
Tags
benchmark-auditscientific-reasoningevaluationphysicsllm-as-a-judge