Top AI pentester agents
The top AI pentester agents on The Agent Benchmark are XBOW (9.4 of 10), CodeAnt AI (6.6) and Escape (4.9), of 12 ranked. Data as of 3 Oct 2026.
An AI pentester attacks running applications, networks, or systems to find exploitable vulnerabilities, as a penetration tester or red team would.
- Ranked agents
- 12
- in Security
- Median score
- 3.0
- of 10
- From YC
- 11
- 1 from outside YC
Ranked pentester agents
All 12, best score first. The score adds public proof, scale, momentum and autonomy; it is not a hands-on test. How scoring works
- 1
XBOWXBOW is an autonomous offensive security platform that simulates real attacks to find and validate vulnerabilities in your applications.Proof9Scale10Momentum10Autonomy8Score9.4 - 2
CodeAnt AIW24Autonomous offensive and defensive cybersecurity platformProof7Scale6Momentum7Autonomy6Score6.6 - 3
EscapeW23Offensive security for the teams that are 100x outnumberedProof3Scale7Momentum4Autonomy6Score4.9 - 4
ParameterW26Security builds trust. Strengthen both, all on one platform.Proof4Scale4Momentum3Autonomy6Score4.1 - 5
CascoSp25Autonomous security testing for web apps, APIs, cloud, and AI systemsProof2Scale4Momentum3Autonomy6Score3.5 - 6
FabraixS26The world's frontier hacker of AI agents.Proof2Scale1Momentum6Autonomy6Score3.3 - 7
MindFortSp25Autonomous Security AgentsProof0Scale3Momentum2Autonomy8Score2.6 - 8
TridentS26Autonomous offensive security agents. The best defense is offenseProof0Scale2Momentum4Autonomy6Score2.5 - 9
AntigenF25Continuous offensive security for the enterprise.Proof0Scale3Momentum2Autonomy6Score2.3 - 10
MetalwareS23Firmware cybersecurityProof0Scale2Momentum2Autonomy6Score2.0 - 11
GhostEyeS25Your always-on red team.Proof0Scale3Momentum0Autonomy6Score1.8 - 12
Veria LabsF25Continuous AI pentesting that finds and fixes vulnerabilitiesProof0Scale2Momentum0Autonomy6Score1.5
Can agents do this job?
Nearly all tasksin security
Public benchmarks cover 35% of the work in this job family (7 benchmarks). About the job, not one agent.
Source: Can Agents Work?
Data as of 3 Oct 2026. The score uses public evidence only. Read the methodology · Missing an agent? Submit it