Benchmarking Claude 4 vs Claude 4.5 for Penetration testing

We benchmarked Claude 4.5 against Claude 4 on two tracks: Linux privilege escalation and CVE-driven web-application exploitation. All runs were fully autonomous in our system. No external methodology was supplied for privilege escalation. Web testing used a simple, fixed methodology. Both models operated using the same container of tools to draw from.

For privilege escalation, each model received non-sudo SSH access to a Linux container and a fixed time budget to obtain a root shell. The challenge set spanned simple cronjob abuse through more advanced techniques such as compiling and loading a malicious audit library. Claude 4.5 achieved a 94% success rate versus 80% for Claude 4, and it reached root faster on median across the same scenarios.

For web applications, we embedded varied CVEs across multiple services and measured end-to-end exploit completion under equal time and tool budgets. Claude 4.5 completed 100% of targeted CVE exploits in this set. Claude 4 completed 52%. The gap aligned with tool usage: Claude 4.5 consistently combined katana and curl with ad-hoc JavaScript and Python for dynamic analysis, while Claude 4 often tried to chain curl commands and stumbled on bash syntax. A model scoring 100% on a benchmark is new territory, and Vulnetic will be rebuilding our tests to match the increased capability of LLMs.

Performance of Claude 4 vs Claude 4.5 across Linux privesc & web app testing

Unintended findings mattered. Claude 4.5 uncovered vulnerabilities we did not seed in the test, including NoSQL injection, SSRF, and a high-impact information disclosure. Claude 4 found none. On signal quality, Claude 4.5 produced zero false positives. Claude 4 had one, a stored XSS.

These outcomes match Anthropic’s terminal-coding benchmarks and show a material step up in real-world autonomy. In our environment, Claude 4.5 is the superior hacking model to its predecessor. We are standardizing on 4.5 for default runs, and we will publish a head-to-head with GPT-5 under the same protocol next.

Checkout our hacking agent here: https://vulnetic.ai