Autonomous Pen Testing: It’s not what’s vulnerable. It’s what’s exploitable.

Written by Gregg Meade

  • Cyber Security
  • Insights
  • News

Your penetration test gave you 4,000 findings. Your team has time for 40. An attacker’s-eye view is how you choose. 

Every security team knows the feeling of the vulnerability backlog. Vulnerability scans are good at finding weaknesses but less effective at telling you which ones matter. The result is a long, colour-coded list, and the quiet pressure of deciding what to ignore. 

The discipline lies in choosing what not to do. The below explores how autonomous penetration testing changes that decision. 

More vulnerabilities; More noise

CISA’s guidelines for Known Exploited Vulnerabilities (KEVs) have become increasingly strict. Under Binding Operational Directive (BOD) 26-04, issued in June 2026, the standard 14-day patching window has been dramatically accelerated for high-risk assets. Instead of five days, CISA now mandates a three-day remediation deadline for actively exploited, automatable vulnerabilities affecting internet-facing systems. This ultimately means that, in practice, defenders are off the pace, and the main reason is a lack of prioritisation, in the face of a tsunami of vulnerabilities. AI has made this worse by adding more vulnerabilities and more noise. 

Vulnerability scanners do a useful job, and they will show a handful of criticals – some highs and a scattering of mediums and lows. But the real question is whether a finding is exploitable, and what an attacker could do once they had exploited it. 

When a ‘medium risk’ becomes ‘high’

KHIPU’s Red Team observes this pattern regularly. A scan reports a critical vulnerability. If exploited, it would give access to one system, but the attacker would be contained and unable to go further. Meanwhile, a medium-rated finding sits lower down the list, where a team working top-down would reach it late. 

An autonomous test exploits that medium and finds it lets the attacker compromise the system and go on to the domain controller, the crown jewels of the environment. 

Insecure SMB configurations often carry a CVSS score of three or four, and in isolation that is fair. If that weakness sits in a chain that ends in domain compromise and full control of the network, he argued, it becomes a ten. It is the thing you must deal with now. 

CVSS remains useful as a general hypothesis about how a weakness might behave. It can’t see your environment or understand your business, and that is the gap that attack-path testing fills. Autonomous penetration testing adjusts priority dynamically, based on what a proven attack chain does in your network. 

Think in chains, not lists

As was explained in a recent session by Dan Bird, Field CTO at Horizon3, “Defenders, tend to think in lists and layers. Attackers think in chains. Each step in a chain is visible, along with the conditions that had to be true for the attacker to progress.” 

A single vulnerability can sit at the head of tens, or even hundreds, of potential exploits. Identify that one point, fix it, and you may secure the environment many times over. The message to a stretched team is that you may not need to fix 100 things and might only need to fix one. 

Evolving Penetration Test Outcomes 

From “we have a problem” to “here’s the business risk” 

Each automated pen test also generates a picture of risk along three dimensions: what is provably exploitable, which threat actors are known to use that technique, and the precise business impact if it is used against you. 

Armed with this level of detail, IT teams have more context to take to senior leadership. Instead of telling the board about an ActiveMQ problem, they can say that the organisation is exploitable through a specific remote code execution flaw, that this is the route state-linked actors have used to take operational control of networks, that he knows how to fix it, and that he needs resources to do it. 

The KHIPU services also provides reporting that flexes to the audience. There is an executive summary for the C-suite, a full penetration test report, and a break-fix report for remediation engineers that says how the attack was done and how to close it. 

What this looks like in practice – Two anonymised engagements

Case file Alpha

Was an authenticated internal test across about 100 hosts. The customer supplied credentials to simulate an attacker who had already got a foothold. The platform reached Active Directory and retrieved the credential database. It then cracked the hashes and compared them with known breached passwords. It flagged weak credentials and found users sharing the same password across accounts. One engineer used the same password for a standard account and a privileged one. Compromise the first, and you have the second. The tooling also searched locally stored files and found plain-text passwords in firewall configuration files and backups. The customer had known about none of it. They reset credentials, changed how the directory database was accessed, and received a break-fix report. 

Case file Bravo

Was an unauthenticated external test of 20 hosts. It turned up open ports on several sites and IP addresses that the ISPs’ records suggested did not belong to the customer at all. KHIPU paused and asked the customer to confirm ownership before testing further. The customer took the query to their ISPs, closed ports, and added access controls where ports had to stay open. A one-click verify then confirmed the change. 

Closing the loop 

That last step is where many programmes fall over. Once an engineer says a fix is done, the platform can run a targeted re-test, dubbed a ‘one-click verify’. It takes a fraction of the time of a full test and returns evidence that the hole is closed, so the change is proven and not merely reported. 

Run continuously, that produces a steady improvement that can be shown. 

Try it on your own estate: