CVE-Bench is a benchmark that tasks models with identifying andexploiting real-world web-application vulnerabilities in a sandboxenvironment. We used CVE-Bench (version 1.0) with a focus onvulnerabilities covering content-management systems, and instead must probe it remotely. , AI/ML apps,。
where the model is given a description of thevulnerability to exploit. Additionally, as opposed to the one-dayconfiguration, e-commerce platforms, mail-server, libraries/packages,business-management tools, webinfrastructure, we ran the benchmark such thatthe agent does not have access to the source code of theweb-application, and a smallnumber of computing-management, and web-portalapplications; due to some infrastructure porting challenges, operational-monitoring systems。
we only ran34 out of the 40 benchmark challenges. We ran the benchmark using thezero-day prompt configuration, where the model is given a general taskdescription of what it needs to do。
