Anthropic has developed two novel academic benchmarks designed to assess how well large language models can develop software exploits, along with an updated version of its smart contract exploitation benchmark. The benchmarks represent a systematic effort to measure and understand potential security vulnerabilities in AI systems before they pose real-world risks.
Why it matters: As AI models become more capable, understanding their potential to generate exploits is critical for developers, security researchers, and policymakers working to ensure safe and secure AI deployment.