# The Agent Benchmark > The top AI agents, from YC startups and beyond, ranked on public evidence and explained. 970 AI agent companies from Y Combinator batches since ChatGPT and 43 well-known agents from outside YC; 999 active ones are ranked. Data as of 3 Oct 2026. An agent here is a company whose product does a job (sales, bookkeeping, support, coding and more) and acts on its own: it takes actions in other systems or finishes whole tasks. Tools that help other agents (browsers, memory, frameworks) are not ranked. The score (v0.1) runs from 0 to 10 and uses public evidence only: Proof 30% (do customers use it), Scale 30% (is there a real company behind it), Momentum 25% (is it shipping and growing now), Autonomy 15% (how much of the job does it do on its own). It is not a hands-on test of the product. Company facts come from Y Combinator company pages and Launch YC posts. Markets and autonomy levels are model output that no person has reviewed yet. For agents from outside YC, every fact comes from a public page, with a link and an exact quote, and a reviewed decision sets the market and the autonomy level. GitHub stars and downloads give no points. Traction claims are quoted word for word and not verified. ## Key pages - [Top Agents](https://theagentbenchmark.com/): the leaderboard of all ranked agents. - [Categories and markets](https://theagentbenchmark.com/categories/): every job area and market, each with its own ranked list. - [Database](https://theagentbenchmark.com/database/): every agent in one table you can filter, sort and export. - [Methodology](https://theagentbenchmark.com/methodology/): how the score works and what it leaves out. - [Submit an agent](https://theagentbenchmark.com/submit/): the form for people to add an agent that is not on the list. AI agents: see "For AI agents" below. - [Sitemap](https://theagentbenchmark.com/sitemap.xml): every page. Each profile shows the evidence behind every point, with a link to its source. ## Rankings by job - [Sales and marketing](https://theagentbenchmark.com/categories/sales/): 158 ranked agents. Top: Mutiny (7.2), Piper the AI SDR Agent (7.1), Tandem (6.9). Markets: [SDR](https://theagentbenchmark.com/markets/ai-sdr/) 44, [Performance marketer](https://theagentbenchmark.com/markets/ai-performance-marketer/) 22, [Sales assistant](https://theagentbenchmark.com/markets/ai-sales-assistant/) 20, [Content marketer](https://theagentbenchmark.com/markets/ai-content-marketer/) 18, [Insurance broker](https://theagentbenchmark.com/markets/ai-insurance-broker/) 15, [Market researcher](https://theagentbenchmark.com/markets/ai-market-researcher/) 11, [Account manager](https://theagentbenchmark.com/markets/ai-account-manager/) 10, [Retail and e-commerce operator](https://theagentbenchmark.com/markets/ai-ecommerce-operator/) 10, [Real estate agent](https://theagentbenchmark.com/markets/ai-real-estate-agent/) 4, [RFP and proposal writer](https://theagentbenchmark.com/markets/ai-rfp-writer/) 2, [Travel agent](https://theagentbenchmark.com/markets/ai-travel-agent/) 2. - [Finance, accounting, and insurance](https://theagentbenchmark.com/categories/finance/): 127 ranked agents. Top: Rogo (8.8), Hebbia (8.1), Basis (6.4). Markets: [Billing and collections agent](https://theagentbenchmark.com/markets/ai-collections-agent/) 27, [Bookkeeper and accountant](https://theagentbenchmark.com/markets/ai-bookkeeper/) 22, [Financial-crime analyst](https://theagentbenchmark.com/markets/ai-fincrime-analyst/) 15, [Insurance claims specialist](https://theagentbenchmark.com/markets/ai-claims-specialist/) 14, [Investment analyst](https://theagentbenchmark.com/markets/ai-investment-analyst/) 14, [Accounts payable clerk](https://theagentbenchmark.com/markets/ai-ap-clerk/) 10, [Loan officer](https://theagentbenchmark.com/markets/ai-loan-officer/) 7, [Tax preparer](https://theagentbenchmark.com/markets/ai-tax-preparer/) 7, [Auditor](https://theagentbenchmark.com/markets/ai-auditor/) 5, [Insurance underwriter](https://theagentbenchmark.com/markets/ai-insurance-underwriter/) 4, [FP&A analyst](https://theagentbenchmark.com/markets/ai-fpa-analyst/) 2. - [Software and IT](https://theagentbenchmark.com/categories/software/): 98 ranked agents. Top: Codex (9.7), Claude Code (9.7), Cursor (9.7). Markets: [Software engineer](https://theagentbenchmark.com/markets/ai-software-engineer/) 40, [App builder](https://theagentbenchmark.com/markets/ai-app-builder/) 19, [QA engineer](https://theagentbenchmark.com/markets/ai-qa-engineer/) 15, [SRE](https://theagentbenchmark.com/markets/ai-sre/) 12, [Code reviewer](https://theagentbenchmark.com/markets/ai-code-reviewer/) 4, [DevOps and cloud engineer](https://theagentbenchmark.com/markets/ai-devops-engineer/) 4, [IT help desk](https://theagentbenchmark.com/markets/ai-it-helpdesk/) 4. - [People and operations](https://theagentbenchmark.com/categories/operations/): 89 ranked agents. Top: Standout (6.2), Contrario (5.6), Trellis (5.4). Markets: [Logistics coordinator](https://theagentbenchmark.com/markets/ai-logistics-coordinator/) 20, [Recruiter](https://theagentbenchmark.com/markets/ai-recruiter/) 15, [Procurement agent](https://theagentbenchmark.com/markets/ai-procurement-agent/) 14, [HR and people operations](https://theagentbenchmark.com/markets/ai-people-ops/) 13, [Property manager](https://theagentbenchmark.com/markets/ai-property-manager/) 10, [Business analyst](https://theagentbenchmark.com/markets/ai-business-analyst/) 6, [Project manager](https://theagentbenchmark.com/markets/ai-project-manager/) 6, [Supply chain planner](https://theagentbenchmark.com/markets/ai-supply-chain-planner/) 5. - [General-purpose work](https://theagentbenchmark.com/categories/horizontal/): 86 ranked agents. Top: Manus (7.9), Glean Agents (6.3), dots (5.9). Markets: [Coworker (general purpose)](https://theagentbenchmark.com/markets/ai-coworker/) 52, [Agent and workflow builder platform](https://theagentbenchmark.com/markets/ai-agent-builder/) 33, [Meeting assistant](https://theagentbenchmark.com/markets/ai-meeting-assistant/) 1. - [Customer support and front desk](https://theagentbenchmark.com/categories/support/): 77 ranked agents. Top: Sierra (9.7), Decagon (9.2), PolyAI (8.6). Markets: [Receptionist](https://theagentbenchmark.com/markets/ai-receptionist/) 42, [Customer support agent](https://theagentbenchmark.com/markets/ai-support-agent/) 35. - [Healthcare](https://theagentbenchmark.com/categories/healthcare/): 71 ranked agents. Top: Abridge (8.1), Hippocratic AI (8.1), Ambience (7.4). Markets: [Medical biller and coder](https://theagentbenchmark.com/markets/ai-medical-biller/) 29, [Care coordinator](https://theagentbenchmark.com/markets/ai-care-coordinator/) 18, [Medical scribe](https://theagentbenchmark.com/markets/ai-medical-scribe/) 13, [Prior authorization specialist](https://theagentbenchmark.com/markets/ai-prior-auth-specialist/) 6, [Clinical trials coordinator](https://theagentbenchmark.com/markets/ai-clinical-trials-coordinator/) 4, [Clinician](https://theagentbenchmark.com/markets/ai-clinician/) 1. - [Engineering and industry](https://theagentbenchmark.com/categories/engineering/): 71 ranked agents. Top: Handoff (4.8), Edviro (4.4), Mach9 (4.1). Markets: [Hardware engineer](https://theagentbenchmark.com/markets/ai-hardware-engineer/) 25, [Architecture and construction specialist](https://theagentbenchmark.com/markets/ai-aec-specialist/) 23, [Plant and infrastructure operator](https://theagentbenchmark.com/markets/ai-plant-operator/) 10, [Manufacturing engineer](https://theagentbenchmark.com/markets/ai-manufacturing-engineer/) 7, [Field technician assistant](https://theagentbenchmark.com/markets/ai-field-technician/) 6. - [Legal and regulatory](https://theagentbenchmark.com/categories/legal/): 62 ranked agents. Top: EvenUp (9.0), Spellbook (8.9), Harvey (7.1). Markets: [Compliance and regulatory specialist](https://theagentbenchmark.com/markets/ai-regulatory-specialist/) 25, [Contract lawyer](https://theagentbenchmark.com/markets/ai-contract-lawyer/) 11, [Legal operations](https://theagentbenchmark.com/markets/ai-legal-ops/) 7, [Litigation paralegal](https://theagentbenchmark.com/markets/ai-litigation-paralegal/) 6, [Immigration paralegal](https://theagentbenchmark.com/markets/ai-immigration-paralegal/) 4, [Patent agent](https://theagentbenchmark.com/markets/ai-patent-agent/) 4, [Investigator](https://theagentbenchmark.com/markets/ai-investigator/) 3, [Government caseworker](https://theagentbenchmark.com/markets/ai-government-caseworker/) 2. - [Office and administration](https://theagentbenchmark.com/categories/office/): 48 ranked agents. Top: Powder (4.8), Arzana (4.4), Parsewise (4.4). Markets: [Document processor](https://theagentbenchmark.com/markets/ai-document-processor/) 27, [Order desk](https://theagentbenchmark.com/markets/ai-order-desk/) 17, [Executive assistant](https://theagentbenchmark.com/markets/ai-executive-assistant/) 4. - [Design and media](https://theagentbenchmark.com/categories/creative/): 43 ranked agents. Top: Manicule (6.8), Hera (4.7), Clueso (4.6). Markets: [Video and audio producer](https://theagentbenchmark.com/markets/ai-video-producer/) 25, [Designer](https://theagentbenchmark.com/markets/ai-designer/) 9, [Translator](https://theagentbenchmark.com/markets/ai-translator/) 5, [Technical writer](https://theagentbenchmark.com/markets/ai-technical-writer/) 4. - [Data, ML, and science](https://theagentbenchmark.com/categories/data/): 32 ranked agents. Top: Kosmos (7.6), Scispot (5.6), Rollstack (4.7). Markets: [Scientist](https://theagentbenchmark.com/markets/ai-scientist/) 16, [Data engineer](https://theagentbenchmark.com/markets/ai-data-engineer/) 7, [Data analyst](https://theagentbenchmark.com/markets/ai-data-analyst/) 5, [ML engineer](https://theagentbenchmark.com/markets/ai-ml-engineer/) 4. - [Security and trust](https://theagentbenchmark.com/categories/security/): 31 ranked agents. Top: XBOW (9.4), CodeAnt AI (6.6), Cyble (6.5). Markets: [Pentester](https://theagentbenchmark.com/markets/ai-pentester/) 12, [Application security engineer](https://theagentbenchmark.com/markets/ai-appsec-engineer/) 8, [Security compliance (GRC) analyst](https://theagentbenchmark.com/markets/ai-grc-analyst/) 5, [SOC analyst](https://theagentbenchmark.com/markets/ai-soc-analyst/) 5, [Trust and safety analyst](https://theagentbenchmark.com/markets/ai-trust-safety-analyst/) 1. - [Education](https://theagentbenchmark.com/categories/education/): 6 ranked agents. Top: Solidroad (4.3), SimCare (3.5), Bloomy (3.5). Markets: [Tutor and teacher](https://theagentbenchmark.com/markets/ai-tutor/) 6. ## For AI agents: read the data - Each agent's profile has a Markdown copy at /agents/.md, with the score, the evidence behind each point and its sources, for example https://theagentbenchmark.com/agents/claude-code.md. The HTML profile is at /agents//. - https://theagentbenchmark.com/search.json lists every agent's name and profile path, best rank first, and every market. An agent's slug is the last part of its profile path. - Every page says "Data as of" with the data date. The data changes when we fetch new data, not every day. ## For AI agents: add an agent or correct a listing You can send two kinds of entry to The Agent Benchmark. A person reviews each entry before anything on the site changes. We change a fact only when a public source supports it. Companies cannot pay to change a score or a rank. - **A new agent** (kind `agent`): an AI agent that is not on the list. Search for it first: https://theagentbenchmark.com/search.json lists every listed agent and its profile. - **A change** (kind `change`): a correction or more information for a listed agent's profile, with 1 to 3 public source links. ### If you can send HTTP requests: use the API Send `POST https://theagentbenchmark.com/api/v1/submissions` with a JSON body. You need no key, no account and no human check. Send it from a server, a script or your runtime, not from a web page on another site (the API refuses requests with another site's `Origin`). Limits: 5 entries a minute from one IP address, and 100 entries a day from all senders together. A change: ```sh curl -X POST https://theagentbenchmark.com/api/v1/submissions \ -H 'Content-Type: application/json' \ -d '{"kind":"change","agent_slug":"acme-agent","topic":"fact","details":"The team has 25 people now. The About page lists them.","sources":["https://acme.example/about"],"relationship":"company","client":"Your agent name"}' ``` A new agent: ```sh curl -X POST https://theagentbenchmark.com/api/v1/submissions \ -H 'Content-Type: application/json' \ -d '{"kind":"agent","name":"Acme Support Agent","website":"https://acme.example","does":"Answers support tickets for online shops","maker":"Acme","buyer":"business","open_source":"no","relationship":"user","client":"Your agent name"}' ``` The fields of a new agent: - `kind` (required; one of "agent"): A new agent: one that is not on the list. - `name` (required; text, 120 characters or fewer): The agent's name. - `website` (required; a link, 500 characters or fewer): The agent's website: a full http:// or https:// link. - `does` (required; text, 200 characters or fewer): What the agent does, in a few words. For example: answers support tickets for online shops. - `maker` (optional; text, 120 characters or fewer): The company that makes it. - `buyer` (required; one of "business", "consumer", "both", "unknown"): Who uses it: businesses, consumers (people), both, or unknown. - `open_source` (required; one of "yes", "no", "unknown"): Is it open source? - `relationship` (required; one of "builder", "user", "other"): The relationship of the person you act for: builder (works on the agent), user, or other. - `email` (optional; an email address): A contact address for questions about the entry: your user's, with their permission. We do not publish it. - `notes` (optional; text, 2000 characters or fewer): Anything else, for example public links that show customers, funding or what the agent does. - `client` (optional; text, 120 characters or fewer): The name of your agent or program, for example Claude Code. It helps the reviewer. The fields of a change: - `kind` (required; one of "change"): A change to a listed agent's profile. - `agent_slug` (required; text, 80 characters or fewer): The listed agent: the last part of its profile's address, /agents//. - `topic` (required; one of "score", "fact", "description", "missing", "other"): What the change is about: score, fact (funding, team, founders or links), description (the description or the market), missing (missing information or context), or other. - `details` (required; text, 2000 characters or fewer): What should change, and why. - `sources` (required; a list of 1 to 3 links): 1 to 3 public links that support the change. We change a fact only when a public source supports it. - `relationship` (required; one of "company", "user", "other"): The relationship of the person you act for: company (works at the agent's company), user, or other. - `email` (optional; an email address): A contact address for questions about the entry: your user's, with their permission. We do not publish it. - `client` (optional; text, 120 characters or fewer): The name of your agent or program, for example Claude Code. It helps the reviewer. The answers are JSON: - `201`: `{"ok": true, "id": "…", "status": "new"}`. Give the id to your user as the reference. - `400`: `"fields"` names each problem. Correct the fields, then send again. - `403`: the request came from a page on another site. - `429`: too many entries. Wait 60 seconds. - `503`: the API took its entries for today, or it is closed. Use a link (below) instead. - `500`: an error on our side. Try again later. The OpenAPI description: https://theagentbenchmark.com/openapi.json. The API catalog: https://theagentbenchmark.com/.well-known/api-catalog. ### If you can only read pages: give your user a link A link can open the site's form with your text filled in. Your user checks it, completes the human check and clicks Send. Encode each value (percent-encoding). Do not put an email address in a link. - A new agent: https://theagentbenchmark.com/submit/?name=Acme+Support+Agent&website=https%3A%2F%2Facme.example&does=Answers+support+tickets+for+online+shops&buyer=business Parameters: `name`, `website`, `does`, `maker`, `buyer`, `open_source`, `relationship`, `notes`. The choices are the API's (`buyer`: "business", "consumer", "both", "unknown"; `open_source`: "yes", "no", "unknown"; `relationship`: "builder", "user", "other"). - A change: https://theagentbenchmark.com/agents/acme-agent/?suggest=fact&details=The+team+has+25+people+now.&source=https%3A%2F%2Facme.example%2Fabout Use the agent's own profile address. Parameters: `suggest` (the topic: "score", "fact", "description", "missing", "other"), `details`, `source` (up to 3 times), `relationship` ("company", "user", "other"). ## How to cite Name the agent, its score, the data date and the profile page, for example: "Codex scores 9.7 of 10 on The Agent Benchmark (#1 of 999 AI agents, data as of 3 Oct 2026): https://theagentbenchmark.com/agents/codex/". For a market, cite its page, for example https://theagentbenchmark.com/markets/ai-software-engineer/. ## Top 25 agents 1. [Codex](https://theagentbenchmark.com/agents/codex/): 9.7. Software engineer. The best way to build with agents. Codex accelerates real engineering work, from planning and building features to refactors, reviews, and releases. 2. [Claude Code](https://theagentbenchmark.com/agents/claude-code/): 9.7. Software engineer. Anthropic's agentic coding tool for developers. 3. [Sierra](https://theagentbenchmark.com/agents/sierra/): 9.7. Customer support agent. Sierra helps businesses build better, more human customer experiences with AI. 4. [Cursor](https://theagentbenchmark.com/agents/cursor-cloud-agents/): 9.7. Software engineer. Cursor is your coding agent for building ambitious software. 5. [XBOW](https://theagentbenchmark.com/agents/xbow/): 9.4. Pentester. XBOW is an autonomous offensive security platform that simulates real attacks to find and validate vulnerabilities in your applications. 6. [Lovable](https://theagentbenchmark.com/agents/lovable/): 9.4. App builder. Describe what you want in plain language and Lovable builds it: full-stack apps, websites and internal tools. 7. [Decagon](https://theagentbenchmark.com/agents/decagon/): 9.2. Customer support agent. The AI concierge for every customer 8. [EvenUp](https://theagentbenchmark.com/agents/evenup/): 9.0. Litigation paralegal. EvenUp is the leading proactive AI platform for personal injury law firms. 9. [Spellbook](https://theagentbenchmark.com/agents/spellbook/): 8.9. Contract lawyer. The first AI system that powers contracts end-to-end. 10. [Rogo](https://theagentbenchmark.com/agents/rogo/): 8.8. Investment analyst. Rogo is the trusted AI partner to the world’s leading financial institutions. 11. [PolyAI](https://theagentbenchmark.com/agents/polyai/): 8.6. Customer support agent. Enterprise voice AI that finishes what your customers start. 12. [Fin AI Agent](https://theagentbenchmark.com/agents/fin-ai-agent/): 8.3. Customer support agent. Perfect customer experiences made possible with Fin. 13. [ResolveAI](https://theagentbenchmark.com/agents/resolve-ai/): 8.2. SRE. Goes on-call on your behalf and autonomously resolves incidents 14. [Abridge](https://theagentbenchmark.com/agents/abridge/): 8.1. Medical scribe. AI-powered clinical documentation, built for accuracy 15. [Hippocratic AI](https://theagentbenchmark.com/agents/hippocratic-ai/): 8.1. Care coordinator. Hippocratic AI builds the safest generative AI healthcare agent for health systems, payors, and pharma. 16. [Hebbia](https://theagentbenchmark.com/agents/hebbia/): 8.1. Investment analyst. The leading AI platform for finance, used by the world's leading asset managers, investment banks, law firms and Fortune 500 companies. 17. [Manus](https://theagentbenchmark.com/agents/manus/): 7.9. Coworker (general purpose). Manus is the action engine that goes beyond answers to execute tasks, automate workflows, and extend your human reach. 18. [Kosmos](https://theagentbenchmark.com/agents/kosmos/): 7.6. Scientist. Kosmos: The AI Scientist for R&D 19. [Factory Droids](https://theagentbenchmark.com/agents/factory-droids/): 7.6. Software engineer. Factory Droids automate coding, testing, and deployment for startups and enterprises. 20. [Ambience](https://theagentbenchmark.com/agents/ambience/): 7.4. Medical scribe. Ambience is more than ambient listening—it’s a connected system that supports every part of the clinical workflow. 21. [Replit Agent](https://theagentbenchmark.com/agents/replit-agent/): 7.3. App builder. Replit is the AI platform that helps you turn your ideas into real outcomes. 22. [Mutiny](https://theagentbenchmark.com/agents/mutiny/): 7.2. Sales assistant. Your AI agent for creating anything customer-facing, in minutes. 23. [Piper the AI SDR Agent](https://theagentbenchmark.com/agents/piper/): 7.1. SDR. Hire Piper the AI SDR Agent to generate inbound pipeline at scale across your most important marketing channels. 24. [Momentic](https://theagentbenchmark.com/agents/momentic/): 7.1. QA engineer. Mo, the AI QA engineer that bug-bashes your app before every release 25. [Harvey](https://theagentbenchmark.com/agents/harvey/): 7.1. Contract lawyer. Harvey's AI agents execute legal work end to end, so you can focus on what only lawyers can do.