Summary

Domain experts re-grade six physics benchmarks for GPT-5.6 Sol, Fable 5 and Gemini 3.1 Pro, separating model errors from broken questions, reference solutions and graders. On audited public-source cases initially marked wrong, 148 of 152 are benchmark or grader errors. After repair or exclusion, GPT-5.6 Sol rises from 47.3% to 78.7% mean@4 on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark; corrected pass@4 on retained CritPt tasks reaches 94.4%. The corrected scores use retained or repaired subsets, so they are not a direct like-for-like leaderboard replacement.

Why it matters
For teams using scientific benchmarks to route models or claim capability gaps, evaluator and reference quality are now large enough to dominate the measured error. High-stakes benchmark suites need expert adjudication, equivalent-answer handling and retained-set reporting. Near-saturation of repaired closed-ended tasks also means new evaluations must test open-ended modeling and experimental judgment.
Technical details
Benchmarks HLE-Physics, CMT-Benchmark, CritPt, PHYBench, PRISM-Physics and UGPhysics
Models GPT-5.6 Sol, Claude Fable 5 and Gemini 3.1 Pro
Public Source Audit 148 of 152 initially rejected cases (97.37%) attributed to benchmark or grader errors
Gpt 5 6 Sol HLE-Physics mean@4 47.28% to 78.66%; CMT 61.00% to 87.24%; retained CritPt mean@4 87.50% and pass@4 94.44%
Evaluation Warning pre-audit and corrected scores often use different question sets after repair or exclusion
Availability paper and detailed audit tables; no code repository linked
Tags
benchmark-auditscientific-reasoningevaluationphysicsllm-as-a-judge