Top AI QA engineer agents
The top AI QA engineer agents on The Agent Benchmark are Momentic (7.1 of 10), Spur (4.8) and TesterArmy (3.9), of 15 ranked. Data as of 3 Oct 2026.
An AI QA engineer tests software by running the application (web, mobile, or API), finding and reporting bugs, or writing and maintaining automated tests.
- Ranked agents
- 15
- 1 more not ranked
- Median score
- 2.4
- of 10
- From YC
- 15
- none from outside YC
Ranked QA engineer agents
All 15, best score first. The score adds public proof, scale, momentum and autonomy; it is not a hands-on test. How scoring works
- 1
MomenticW24Mo, the AI QA engineer that bug-bashes your app before every releaseProof7Scale6Momentum9Autonomy6Score7.1 - 2
SpurS24Spur is your AI QA Engineer. Test your websites with natural language.Proof3Scale4Momentum7Autonomy6Score4.8 - 3
TesterArmySp26Test your app with AI, catch bugs before users doProof2Scale3Momentum6Autonomy6Score3.9 - 4
QualGentSp25AI Mobile App Quality Assurance TesterProof2Scale4Momentum2Autonomy8Score3.5 - 5
NarrativeW23AI agents for end-to-end testingProof3Scale3Momentum2Autonomy6Score3.2 - 6
Decipher AIW24QA agents that write tests 10x faster with zero maintenanceProof5Scale2Momentum0Autonomy6Score3.0 - 7
AutosanaS25AI agents for E2E testing across mobile & webProof0Scale3Momentum3Autonomy6Score2.6 - 8
nunu.aiW23Building the first multimodal agents to play and test games.Proof0Scale5Momentum0Autonomy6Score2.4 - 9
LarkS25The E2E testing layer for AI-driven developmentProof0Scale3Momentum0Autonomy6Score1.8 - 10
ApproximaW26Your software should build itself.Proof0Scale1Momentum1Autonomy8Score1.8 - 11
CanaryW26Adversarial AI that breaks your AIProof0Scale1Momentum1Autonomy6Score1.5 - 12
LucentW26AI that automatically improves products from user behaviorProof0Scale1Momentum1Autonomy6Score1.5 - 13
SimulithicF26User simulations for production monitoring.Proof0Scale1Momentum1Autonomy6Score1.5 - 14
DocketSp25AI agents for web testingProof0Scale1Momentum0Autonomy6Score1.2 - 15
HaystackS24Replay real customer journeys against every change before you shipProof0Scale1Momentum0Autonomy6Score1.2
Not ranked
Inactive or acquired. They keep a profile and a score, without a rank.
Can agents do this job?
Most tasksin software engineering
Public benchmarks cover 47% of the work in this job family (9 benchmarks). About the job, not one agent.
Source: Can Agents Work?
Data as of 3 Oct 2026. The score uses public evidence only. Read the methodology · Missing an agent? Submit it
