Summary

Researchers audit 254 public SWE-bench submissions across four splits using their per-instance verdicts. On Verified, none of the 29 adjacent pairs in the top 30 is distinguishable by an exact paired test. The top ten share 285 successes and 51 failures, leaving only 164 informative tasks. Holding the model fixed, observed scaffold ranges reach 29.8 points, larger than the 8.8-point spread across the top 30.

Why it matters
For teams selecting a coding agent, small leaderboard gaps do not justify a procurement decision. Compare model-and-scaffold pairs on an internal workload, report paired uncertainty, and use tiers instead of a strict rank. The released audit pipeline makes that check reproducible without rerunning models.
Technical details
Dataset 254 public submissions across Verified, Lite, Test and Multimodal
Verified Top 10 285 shared successes; 51 shared failures; 164 discriminating instances
Paired Test 0 of 29 adjacent top-30 Verified pairs separable at alpha 0.05
Scaffold Effect within-model observed range up to 29.8 percentage points
Artifact frozen inputs, scripts, normalized design and tier membership released
Limitation observational audit; non-rejection does not prove equivalence
Tags
coding-agentSWE-benchevaluationbenchmarkscaffold