Summary

IBIB (arXiv 2609.10494, Thu 10 Sep digest) starts from a measurement error: enterprises deploy systems, not checkpoints — usable capability depends jointly on weights, serving route, precision, output contract, and harness — yet all 18 audited benchmarks score advertised model identifiers. The paper's IB2 protocol makes the difference reportable in three parts: a gold-blind capability-binding preflight verifying a route can execute the evaluation contract before any task reaches it; a reliability-inclusive first-pass scoring rule that keeps failure in the score while keeping unsupported capability out; and structurally score-blind adjudication. The reference instantiation is 128 locked tasks with 987 assertions over document, spreadsheet, chart, tool and database work, and stays sealed — the procedure is the artifact, not the corpus. Across eleven systems, capability availability proved measurable, with single-route systems dropping out on entire task classes.

Why it matters
This is the eval methodology enterprise procurement actually needs: score the route you will run (your gateway, your quantization, your harness), not the vendor's model ID. The sealed-corpus design (publish the protocol, keep tasks locked) is a workable answer to benchmark contamination, and the preflight step catches the 'benchmark passes but your deployment can't execute the contract' failure before you buy.
Technical details
Arxiv 2609.10494, announced in the Thu 10 Sep 2026 digest
Motivation 18 audited benchmarks all score advertised model identifiers; enterprise capability depends on weights + route + precision + output contract + harness
Protocol gold-blind capability-binding preflight; reliability-inclusive first-pass scoring; structurally score-blind adjudication; algorithms, classification tables, request contract and manifest schemas released
Instantiation 128 locked tasks / 987 assertions (document, spreadsheet, chart, tool, database), sealed; 11 systems evaluated; capability availability measurable, single-route systems drop out on entire task classes
Tags
evaluationbenchmark-methodologyenterpriseprocurementserving-route