Summary

Harbor Adapters and Harbor-Index (arXiv 2609.04298, Mon 7 Sep digest) is unified evaluation infrastructure for agentic benchmarks: adapters port more than 80 benchmarks to evaluate arbitrary agents, validated through code review and parity experiments. Using it, the authors run 8 models across capability tiers on 54 benchmarks, each with Terminus-2 and one of 3 native harnesses — enabling broader analysis of agent capabilities and failure modes than previously possible. Harbor-Index distills the adapted suite into 82 difficult, diverse, high-quality tasks spanning 29 benchmarks (difficulty filtering, AI and human audit, audit-and-fix loop): it preserves the challenge and breadth of large-scale agentic evaluation while being affordable to run — no evaluated model-harness configuration saturates it.

Why it matters
The agent-eval world's fragmentation tax is real: every benchmark ships its own environment and integration. 80+ adapters with parity validation is the largest standardization effort yet, and the 82-task Harbor-Index is calibrated to stay discriminative at frontier level while cheap enough for routine runs — exactly the artifact a platform team needs as a standing regression suite. Notably, terminal-bench's unified dataset already lives under the harborframework org on GitHub (seen in GLM-5.3's eval-metadata PRs), so the ecosystem adoption has started.
Technical details
Adapters 80+ benchmarks ported to evaluate arbitrary agents; validated via code review and parity experiments
Evaluation 8 models across capability tiers x 54 benchmarks; each model run with Terminus-2 and one of 3 native harnesses
Harbor Index 82 curated tasks spanning 29 benchmarks; difficulty filtering + AI/human audit + audit-and-fix loop; affordable to run; no evaluated model-harness configuration saturates it
Ecosystem terminal-bench unified dataset lives under the harborframework GitHub org (referenced in GLM-5.3 eval-metadata PRs)
Tags
agent-evaluationbenchmarkevaluation-infrastructureharnessharbor