Methodology

Which companies we rank, how the score works, and what it leaves out. Every rule is on this page.

Which agents we rank

We start from 2,701 Y Combinator companies: every company in a batch from Winter 2023 on, plus older companies with a Launch YC post on or after 2022-11-30. An agent here is a product that does a job for a business and acts on its own. A company is in when all three are true:

  1. Its main market is a job, such as sales outreach, bookkeeping, support or code review. Tools that other agents use (browsers, inboxes, memory, frameworks, evals) are not ranked: 461 companies in 15 tool markets are out.
  2. Businesses buy it for work. Consumer apps and products for personal life are out. Outside YC, a personal agent that acts for the person is in (see below).
  3. It acts. Its autonomy level is 2 or more: it takes actions in other systems or finishes whole tasks. Copilots and AI tools that only draft or advise are out. An AI-native service, which sells the finished work, is in at any level.

Rules apply these tests to model labels of each company's YC page. For 2,495 companies, an automated evidence review decides instead: two AI models from different families read the company's website and YC page, and adjudicator models resolve the cases where they disagree. When the adjudicators do not agree, the company stays under review and is not listed. Each decision cites quotes that code finds on the stored pages. The review takes out companies whose customers buy goods, insurance, loans, care, transport or access to a marketplace, also when the company uses AI to run its own operations. It took out 1,525 of the 2,495. No person has checked these decisions.

That gives 970 agents from YC and 43 from outside YC, in 78 markets. Of the 1,417 YC job companies, 519 are not agents. 14 agents are inactive or acquired: they keep a profile and a score, but no rank. 999 agents are ranked.

Agents from outside YC

Many well-known agents come from companies that did not go through Y Combinator, such as Devin, Harvey and OpenClaw. We add them one at a time: 43 so far. For now we pick agents that are well known. An agent that someone submits joins after a person checks it.

  • Every fact has a source. Each fact comes from a public page: the company's site, a release post, a careers page or a funding announcement. We keep the link and the exact quote, and code checks that the quote is on the stored page. Each point on the profile links to its source.
  • The same rules. The score uses the same parts, points and weights. A release post counts as a Launch YC post, and the careers page counts as the YC jobs page. When a maker has more than one product in our data, such as OpenAI or Anthropic, a release post counts only for the product that its title names.
  • The maker's team. When a larger company makes the agent, team size and funding are the maker's: Scale asks if a real company stands behind the product. Proof is about the product only: it counts claims about the listed product, not claims about the whole company, a platform or a family of products under one brand. The profile names the maker, and the company that owns the maker when there is one.
  • Big companies: one product at a time. A big company makes many products. We list one when all three are true:
    1. It is a named product that a customer can get by itself, with its own page, plan or install. A mode inside a larger app is out, for example ChatGPT Work, Claude Cowork, Copilot Cowork and ServiceNow AI Agents.
    2. Customers buy it to get work done. A model, an SDK or a harness alone is a building block, not an agent.
    3. Its proof is about the product (see the maker's team above).
  • The harness. The harness is the software that runs the agent's model, tools and steps. The profile shows it only when a public page states it, with the quote and a link. An agent can have more than one: the one that it runs on, and others that the user can switch to. The harness gives no points. A model, such as GPT or Claude, is not a harness.
  • No batch. The new-company point in Momentum goes to a company founded in 2025 or later, instead of a recent YC batch.
  • Reviewed labels. The market, the autonomy level and whether it counts as an agent come from a reviewed decision that cites the product's own pages, not from a classifier's probabilities. The profile says who set them: an AI agent that drafts them for a maintainer, or a maintainer.
  • Open source and personal agents. They are in the AI coworker (general purpose) market, with an “Open source” or “Consumer” tag. GitHub stars and downloads give no points.

Views by maker. The home page shows all agents, or one of three views. YC: the agent's company went through Y Combinator. Big companies: companies with several separate businesses, and the companies that one of them owns. Our list is Alphabet, Anthropic, Meta, Microsoft, OpenAI, Salesforce, ServiceNow, SpaceX and Workday. Startups: all other agents. A view changes only which agents show and their rank within it; it gives no points.

The score

The score adds four parts by weight. Each part is 0 to 10 and comes from public evidence only. A part with no evidence is 0, so a quiet company scores low even if its product is good.

Proof 30%

Do customers use it? Public claims of revenue, named customers, results and usage, in the company's own words.

Revenue or growth+4
Named customers+3
A customer result+2
A customer count+2
Work done at scale+2

Each kind counts once. The total stops at 10. A model picks which sentences state traction; the words are the company's own and not verified. When a larger company makes the agent, only claims about the listed product count, not claims about the whole company, a platform or a family of products under one brand.

Scale 30%

Is there a real company behind it? Team size on its YC page and the largest round in its funding news.

Team of 100 or more+7
Team of 50 to 99+6
Team of 20 to 49+5
Team of 10 to 19+4
Team of 5 to 9+3
Team of 3 or 4+2
Team of 2+1
Largest round in its news: $20M or more+3
Largest round in its news: $5M to $19M+2
Largest round in its news: under $5M+1

Team size is from its YC page (outside YC: a public source about the maker). Rounds come from the news headlines on its YC page, so many are missing. The company must be the one that raised: a valuation, a total raised to date or another company's round is not a round. In an AI check of 271 headlines, 229 of the 230 rounds found were correct, and 21 rounds were missed.

Momentum 25%

Is it shipping and growing now? Launches, hiring, new claims and funding news in the last 6 to 12 months.

A Launch YC post in the last 6 months+3
Open roles on YC now (5 or more: +3)+2
A new traction claim in the last 6 months+2
Funding news in the last year+2
In one of the last 4 YC batches+1

The Rising view sorts by this part. Outside YC, a release post counts as a launch (only for the product that it names, when the maker has more than one), the careers page counts for hiring, and a company founded in 2025 or later gets the batch point.

Autonomy 15%

How much of the job does it do on its own? Does it act in other systems and finish whole tasks? Read by a model from the company's own pages.

Level 2: Supervised agent. Takes actions in other systems; a person approves key steps.6
Level 3: Autonomous worker. Completes whole tasks; people handle exceptions.8

A model reads the company's own pages. Level 3 needs “finishes whole tasks”; level 2 needs that or “acts in other systems”. An AI-native service gets its level by the same rules: selling the finished work does not raise it. When the evidence review includes a company, it has found with quotes that the AI acts in the work, so its level is at least 2. Outside YC, a reviewed decision with quotes sets the level.

An example: Codex, #1

Proof 10 × 0.30 + Scale 10 × 0.30 + Momentum 10 × 0.25 + Autonomy 8 × 0.15 = 9.7

See why Codex ranks first → Every profile shows the points behind each part, with links to the sources.

Reading the pills

  • 8 7 or more
  • 6 5.5 to 6.9
  • 4.5 4 to 5.4
  • 3 2.5 to 3.9
  • 1 under 2.5
  • 0 nothing found

Categories and markets

Each agent has one market: the job it does, such as SDR or bookkeeper (78 markets with agents). Markets roll up into 14 categories. The taxonomy links each market to a job family in the Can Agents Work? atlas, so a market can be compared with public benchmarks and U.S. wage data.

  • Sales & marketing 158SDR, Performance marketer, Sales assistant, Content marketer, …
  • Finance 127Billing and collections agent, Bookkeeper and accountant, Financial-crime analyst, Insurance claims specialist, …
  • Software 98Software engineer, App builder, QA engineer, SRE, …
  • People & ops 89Logistics coordinator, Recruiter, Procurement agent, HR and people operations, …
  • General 86Coworker (general purpose), Agent and workflow builder platform, Meeting assistant
  • Support 77Receptionist, Customer support agent
  • Healthcare 71Medical biller and coder, Care coordinator, Medical scribe, Prior authorization specialist, …
  • Industry 71Hardware engineer, Architecture and construction specialist, Plant and infrastructure operator, Manufacturing engineer, …
  • Legal 62Compliance and regulatory specialist, Contract lawyer, Legal operations, Litigation paralegal, …
  • Office 48Document processor, Order desk, Executive assistant
  • Design & media 43Video and audio producer, Designer, Translator, Technical writer
  • Data & science 32Scientist, Data engineer, Data analyst, ML engineer
  • Security 31Pentester, Application security engineer, Security compliance (GRC) analyst, SOC analyst, …
  • Education 6Tutor and teacher

Can agents do this job?

Each profile shows how far public benchmarks say agents can do the agent's job family, from the Can Agents Work? atlas (24 Sep 2026). The steps run:

Not yetSome tasksMost tasksNearly all tasksSome projectsMost projects

This is about the job, not the company. It does not change the score: a strong agent in a hard job keeps its rank.

Sources

  • Companies: Y Combinator company pages and Launch YC posts, fetched 3 Oct 2026: one-liners, descriptions, team size, jobs, founders, logos, news and launches.
  • Agents from outside YC: public pages of each company (its site, release posts, careers page and funding announcements), each fact with a link and an exact quote.
  • Labels: a decision model, Jev 1.13 (TypeSafe), picks each company's market, autonomy level and traction claims from options we define. It does not write text. Each question is asked twice with the options reversed, and the answers are averaged.
  • Evidence review: for 2,495 companies, an OpenAI GPT model and a second model of another family (Anthropic Claude; DeepSeek for 50 development cases) each read the company's website and YC page and record what customers pay for, who the customer is and what the AI does, with quotes. Adjudicators decide the disagreements and check a sample of agreements: Claude for the development cases; for the other companies, a second GPT model and a Claude model, which must agree. Code checks every quote against the stored page.
  • Benchmarks and job families: the Can Agents Work? atlas, CC BY 4.0, commit b51fedc.
  • Jobs and wages: U.S. Bureau of Labor Statistics (OEWS), through the atlas.

Limits

  • Public evidence, not a product test. The score shows what a company has shown in public. Nobody has used these agents for this ranking yet.
  • Companies' own words. Claims are quoted word for word and not checked. A company that overstates gets an overstated score.
  • Model labels. No person has reviewed the markets, autonomy levels or eligibility decisions. In a spot check of 40 companies on an earlier version, 36 had the right market. We tested the evidence review on 150 companies that have separate AI reference labels: it listed 98, and the reference calls all 98 agents. It kept 12 of the 110 reference agents under review, and its market matched the reference for 89 of the 98. The reference is also made by AI models, two of them the same as in the review, so this measures agreement, not truth. The rules cannot see a company that uses AI only to run its own operations; the evidence review covers 2,495 companies so far.
  • Visibility counts. Launches, hiring and news raise the score. A quiet company with happy customers ranks lower than it should.
  • Funding is partial. It comes from headlines on YC pages and, outside YC, from announcements we found, not a funding database. Amounts in other currencies are not shown.
  • Mostly YC. Outside YC we cover only some well-known agents so far. Many good agents are not on the list yet.

Corrections

  • Anyone can suggest a change. Use “Suggest a change” on an agent's profile, with a link to a public source. To add a missing agent, submit it.
  • A person reviews each one. A fact changes only when a public source supports it.
  • No paid changes. Companies cannot pay to change a score or a rank.

What's next

  • Agent trials. Give each agent the same objective in a sandboxed computer, record what it does, and score the result. Trial results would join the public evidence.
  • Reviews. Ratings from teams that buy and use the agents.
  • Open and closed agents. A view for each, and how they compare.

Back to the leaderboard →

↑ ↓ to move · Enter to open