Summary

VLoc Bench pairs vulnerable and patched snapshots for 500 real vulnerabilities across 290 repositories. Agents receive a weakness description and read-only terminal access, then identify affected files or say the patch removed the flaw. Across 27 models and four static analyzers, the best File F1 reaches only 0.229; in 38.4% of tasks no tested model finds a correct location. Even stronger locators can report unsupported files after a fix.

Why it matters
For teams deploying security agents, file localization and abstention on patched code need separate acceptance tests. A high-level vulnerability verdict alone does not establish reliable repository navigation.
Technical details
Dataset 500 vulnerabilities; 290 repositories; six ecosystems; 147 CWE classes
Evaluation 27 models and four static analyzers
Best File F1 0.229
Tasks Without Correct Location 38.4%
Code no repository linked on abstract page
Tags
agent-evaluationsecuritycoding-agents