Ever since LLMs and AI agents have been able to run tools on their own and interpret their own results, a narrative has emerged in the market that I keep encountering more and more in customer conversations – and above all on LinkedIn: autonomous AI pentesters. Systems that are supposedly able to carry out an entire penetration testing phase without human supervision: from enumeration to exploitation, autonomous, consistent, traceable.
That is exactly why we put Strix and PentAGI through a field test. Not in the lab, not in a simple demo environment, but held to the yardstick we apply to every pentest we run here at binsec: Is the result consistent, reproducible, and reliable?
The answer is nuanced, but clear: OSINT works well, and to some extent exploitation works too. Full autonomy, in practice, is inconsistent and completely unreliable. The potential, therefore, does not lie in replacing the pentester, but in supporting them – and in automating individual, small sub-steps of a pentest.
What We Tested
For the field test we took two fully autonomous systems that are regularly referred to as “autonomous pentesters” in the community and in the market:
- Strix
- PentAGI
Both systems are based on the same basic principle: an LLM-driven agent system carries out tasks autonomously – enumeration, analysis, partial exploitation – without a pentester steering every step. The promise is clear: a human defines the scope and the objective, and the AI does the rest.
The Big Picture: Sub-steps Yes, Full Autonomy No
In our tests, both systems were able to perform individual sub-steps of a penetration test. Enumeration of services and endpoints works in principle. To some extent, exploitation is also possible.
The problem is not that the systems can’t do anything. The problem is that they deliver neither reliably nor consistently. An autonomous agent that produces a solid result on one run with identical scope and finds none of it on the next is unusable in practice. A pentest is not a one-shot affair. It is a controlled process – and that is precisely what the current generation of autonomous agents fundamentally lacks.
Strix in the Test: A Working Core Feature, a Production-Ready Interface Not in Sight
Strix has shown in our tests that, under certain conditions, it can find vulnerabilities in code or on web servers. It’s not a given, but the basic functionality is there.
However, the result was neither reliable nor consistent. Whether a known vulnerability was found or not depended more on the individual run than is acceptable for a credible test process.
The problem became much more visible on the interface side, though. The user environment of Strix is not production-ready:
- The interface contains bugs.
- Strix crashes from time to time, in particular when it encounters unexpected AI output.
- The only output is a Markdown file with no fixed structure.
For a professional pentesting process, this is a decisive criterion. A test result delivered as unstructured Markdown text with no fixed layout can’t be turned into a reliable report, and it is of no use to clients, management, or IT operations. A good pentest report follows a clear structure, prioritizes findings, assesses risks, and derives actionable measures. Strix’s output delivers none of that.
PentAGI in the Test: More Mature on the Surface, but with Clear Limits
PentAGI appears more production-ready from the surface than Strix. The task organization is better, and the structure of the workflows is easier to follow. Anyone who compares the two systems directly will immediately notice that PentAGI is further along in the question of how the work is organized.
But on closer inspection, the same fundamental problem as with Strix emerges: inconsistency.
Over the course of our testing, we observed several concrete problems:
- Flows sometimes pause when they are supposed to run in the background.
- Agents often do more than the current task, causing a lot of work to be executed twice.
- Scope issues: PentAGI sometimes scanned IPs and ports that had not been specified.
The last point is the dealbreaker for us. An autonomous agent that operates outside the defined scope is not deployable for a professional penetration test. A pentest always runs within a clearly defined, contractually delineated scope. Anyone who acts outside that scope leaves the framework within which the test is authorized. This is not just a methodological weakness – it is a critical problem!
Why Real Environments Are More Than Typical Test Cases
Even more important than the question of the consistency of the results is the question of where autonomous AI agents can sensibly be deployed at all. Because the reality of a pentest is far more complex than any of the systems mentioned can capture.
Most AI pentesting runs against external infrastructure or public web applications. That is the simplest case: the attack surface is clearly delineated, the system boundaries are technically visible, and if the agent goes rogue, it remains relatively controllable, because it can only move in places where you can control it in the first place.
An internal deployment poses significantly higher challenges. Internally, there is no technical edge that automatically stops an agent. Boundaries and conditions can only be upheld contractually and organizationally – and that is precisely where the current generation of autonomous agents fails, as we saw with PentAGI. The larger the internal segment, the greater the potential damage that an agent that doesn’t hold to its limits can cause.
And the environment itself is rarely as simple as it appears in test cases. The interplay between different systems – e.g. consisting of an API, a web app, a mobile app, and third-party systems – can be many times more complex than any typical, isolated test case. A vulnerability path that spans multiple systems requires an understanding of context, dependencies, and permission models that autonomous agents cannot reliably provide today.
Hardware-centric testing is not possible in principle. Bluetooth interfaces, medical devices, IoT systems, industrial control systems. Everywhere a pentest enters the physical domain, we need measurement instruments, specialized tools, and a manual approach that a software agent cannot replicate. This is not a shortcoming of the current AI – it is a fundamental limit.
Why Full Autonomy (Still) Doesn’t Deliver
The two systems share a pattern we know well from practice: they can solve sub-tasks, but they cannot carry a process.
A penetration test is not a chain of independent tasks. It is a controlled process with a defined scope, clear objectives, controlled execution, and traceable evaluation. What autonomous agents deliver today are fragments: sometimes good, sometimes bad, sometimes in scope, sometimes out, sometimes documented, sometimes reproducible, sometimes not.
That is the real problem. It is not that AI is weaker today than an experienced pentester – in most cases, it would be anyway, even if it may simply be faster on some simple tasks. The real problem is that consistency, traceability, and scope discipline are missing. And it is precisely these three things that turn a penetration test into a reliable result rather than an AI pentesting lottery.
Where the Potential Actually Lies
This does not mean that AI has no role in pentesting. Quite the opposite. But the role is different from the one the market is currently promising.
The real potential does not lie in the autonomous agent that replaces the pentester. It lies in supporting the pentester and in the automation of individual, small sub-steps:
- Automated enumeration: Recurring, time-consuming tasks such as capturing endpoints, services, or configurations can be automated reliably.
- Result analysis and preparation: Reviewing large volumes of data, filtering data, generating your own API documentation for client applications – here, AI can provide excellent support.
- Documentation support: Turning raw data into structured reports, performing quality control on the report for logical errors, spelling and grammar mistakes. This takes the load off the pentester in the phase that is the least fun – but that still consumes time even with tools like PTDoc.
That is exactly where the value lies. Not in the “autonomous pentest”, but in the more efficient, structured pentest in which the human keeps control and the AI takes over the tedious drudge work.
