Benchmarking Claude 4.5 Sonnet vs Claude 4.5 Opus for Penetration testing
We just finished a new round of benchmarks pitting Claude 4.5 Opus against Claude 4.5 Sonnet inside the Vulnetic framework, and the gap is now very clear. Under identical conditions, Opus is the stronger autonomous hacking model.
Performance of Claude 4.5 Opus vs Claude 4.5 Sonnet across Linux privesc & web app testing
Across Linux privilege escalation, Opus successfully obtained a root shell within 30 minutes in 81% of scenarios, versus 53% for Sonnet. On web application tests, Opus exploited the intended vulnerability in 92% of cases, compared with 69% for Sonnet. These scores only count full exploit completion: the agent has to actually spawn a root shell or fully exploit the intended vulnerability in the target web app. Simply reporting a vulnerability is not enough, and unintended “bonus” bugs do not contribute to the benchmark score.
We also overhauled the benchmark suite itself. With the release of Claude 4.5 Sonnet in late September, we retired easier challenges such as sudo -l abuse and basic command injection. The new set leans into harder, more esoteric problems such as privilege escalation via custom kernel modules and business-logic flaws that require more complex reasoning. The goal is to make the model think outside the box and not just follow basic pentesting practices.
Under these stricter rules and harder targets, Opus consistently demonstrated better planning, tool use, and recovery when its first idea failed. Sonnet remains capable, but it is clearly a step behind on these higher-order tasks.