What Counts as "Done" for a Pentest?
Every experienced pentester knows that "done" is a practical word. A pentest usually ends because the engagement window ends, because the scoped environment has been covered well enough, or because the team has enough evidence to explain the risk. That does not mean every possible path was exhausted. It means the tester made informed decisions about where limited time was best spent, and those decisions have always been shaped by the cost of human attention.
Real environments do not present themselves as checklists. They behave more like graphs of applications, hosts, identities, services, credentials, permissions, shares, routes, and trust relationships, where the value of one finding often depends on what it can reach next. A good pentester is managing that graph under pressure, deciding which leads deserve another hour and which ones are too speculative to justify more work. Those decisions are technical, but they are also economic, because human-led pentesting has always required sampling, prioritization, and judgment about when the search has gone far enough.
That is why AI changes more than the cost of a pentest. It changes the stopping condition. When another attempt costs tokens instead of billable hours, the threshold for exploring a path drops. A lead does not need to look immediately promising to be worth testing, as long as it is plausible, authorized, and safe within the rules of engagement. This creates a different standard for completeness, especially in environments where risk emerges from relationships between systems rather than from one obvious vulnerability.
Compromise often comes from accumulation. A finding that appears minor in isolation may become meaningful when paired with access, context, or another weakness elsewhere in the environment. Pentesters already understand this, but the number of possible chains is usually larger than the time available to test them. Cheap execution makes it practical to explore more of those chains, including the ones a human tester might reasonably abandon during a fixed-window engagement.
In our earlier piece, "Offensive Security After the Price Collapse," we argued that AI would compress the cost of offensive execution. That prediction now looks like the first phase of a larger methodological change. The important consequence is a higher standard for explaining completeness. A serious AI pentest should be able to show what was attempted, where the agent stopped, which paths failed, which paths created leverage, and where uncertainty remains.
We are still in the chat era. Today, AI pentesting usually starts with a person giving the system a target, defining scope, and steering the work as it runs, which keeps the human operator close to the assessment. Over time, more of this capability will move behind APIs, where security systems request offensive validation when risk changes, when evidence is needed, or when fixes need to be tested again.
For pentesters, the job becomes more focused on higher-leverage judgment. The operator defines safe scope, controls escalation, interprets chains, validates impact, and decides when persistence has become wasteful. The agent can attempt more paths than a human team could justify manually, but the pentester still decides what the results mean.
Human pentesting has always been constrained by time. AI pentesting can be constrained more directly by authorized scope. As the cost of attempts collapses, "done" should mean the environment was worked through with a level of persistence that human-led testing could rarely justify.