Summary
VLoc Bench pairs vulnerable and patched snapshots for 500 real vulnerabilities across 290 repositories. Agents receive a weakness description and read-only terminal access, then identify affected files or say the patch removed the flaw. Across 27 models and four static analyzers, the best File F1 reaches only 0.229; in 38.4% of tasks no tested model finds a correct location. Even stronger locators can report unsupported files after a fix.
Why it matters
For teams deploying security agents, file localization and abstention on patched code need separate acceptance tests. A high-level vulnerability verdict alone does not establish reliable repository navigation.
Technical details
| Dataset | 500 vulnerabilities; 290 repositories; six ecosystems; 147 CWE classes |
|---|---|
| Evaluation | 27 models and four static analyzers |
| Best File F1 | 0.229 |
| Tasks Without Correct Location | 38.4% |
| Code | no repository linked on abstract page |
Tags
agent-evaluationsecuritycoding-agents