Top AI scientist agents
The top AI scientist agents on The Agent Benchmark are Kosmos (7.6 of 10), Scispot (5.6) and Delineate (3.6), of 16 ranked. Data as of 3 Oct 2026.
An AI scientist does scientific research work, for example literature review, forming hypotheses, designing or running lab experiments, or analyzing experimental results.
- Ranked agents
- 16
- in Data & science
- Median score
- 2.5
- of 10
- From YC
- 15
- 1 from outside YC
Ranked scientist agents
All 16, best score first. The score adds public proof, scale, momentum and autonomy; it is not a hands-on test. How scoring works
- 1
Kosmosby Edison ScientificKosmos: The AI Scientist for R&DProof5Scale8Momentum10Autonomy8Score7.6 - 2
ScispotS21The Best Data Infrastructure for BiotechsProof2Scale6Momentum9Autonomy6Score5.6 - 3
DelineateW25Agents for Accelerated Clinical Trial DesignProof4Scale5Momentum0Autonomy6Score3.6 - 4
83 SciencesS26AI-native materials discovery powered by unpublished experimental dataProof0Scale2Momentum6Autonomy6Score3.0 - 5
OhmW23Ohm helps engineering teams accelerate hardware development & testing.Proof0Scale4Momentum3Autonomy6Score2.8 - 6
RasynS26Making chemicals 100,000x fasterProof0Scale2Momentum4Autonomy8Score2.8 - 7
Diffuse BioW23Generative AI for protein designProof0Scale4Momentum2Autonomy6Score2.6 - 8
10x ScienceW26The AI-native platform for next-generation protein characterization.Proof0Scale3Momentum3Autonomy6Score2.6 - 9
InferaSp26Control lab instruments with natural language.Proof0Scale1Momentum4Autonomy8Score2.5 - 10
The Synthesis CompanyS24100x faster scientific evidence synthesisProof5Scale0Momentum0Autonomy6Score2.4 - 11
NovaflowS25The AI data analyst for biology labsProof0Scale2Momentum3Autonomy6Score2.3 - 12
AemonW26The Forward-Deployed AI Research EngineerProof0Scale2Momentum3Autonomy6Score2.3 - 13
Yoneda LabsW24Foundation Model for Chemical ManufacturingProof0Scale3Momentum0Autonomy6Score1.8 - 14
Synthetic SciencesW26Infrastructure for Autonomous ScienceProof0Scale1Momentum1Autonomy8Score1.8 - 15
Strand AIW26Multimodal foundation models to predict uncollected patient biologyProof0Scale1Momentum1Autonomy6Score1.5 - 16
FrekilSp25RWE Generation in minutesProof0Scale1Momentum0Autonomy6Score1.2
Can agents do this job?
Some tasksin science and research
Public benchmarks cover 0% of the work in this job family (1 benchmark). About the job, not one agent.
Source: Can Agents Work?
Data as of 3 Oct 2026. The score uses public evidence only. Read the methodology · Missing an agent? Submit it