Collinear AI Launches CWE-bench to Test Coding Agents on Cybersecurity Defenses
Collinear AI has released CWE-bench, a new benchmark assessing AI coding agents' ability to handle defensive cybersecurity tasks. The most advanced agents passed less than 50% of the tests.

Collinear AI has launched CWE-bench, a new benchmark designed to test the capabilities of frontier AI coding agents in defensive cybersecurity. The tool evaluates agents' ability to identify and mitigate software vulnerabilities categorized under the Common Weakness Enumeration (CWE) standard.
Initial results from the benchmark indicate that even the leading agents struggled, successfully completing less than 50% of the assigned tasks. Notably, 18 specific weakness types remained entirely unsolved by the tested agents, highlighting significant gaps in current AI capabilities for cybersecurity defense.
The CWE-bench suite covers 54 distinct types of software weaknesses, providing a comprehensive evaluation of an AI agent's proficiency in recognizing and preventing security risks during the development process. This initiative aims to foster the creation of more secure AI applications.
Collinear AI stated that the findings offer critical insights for AI developers and cybersecurity professionals. The data suggests that the practical application of AI agents in proactive and reactive cybersecurity measures requires substantial further development.