Pull to refresh
Logo
OpenEvidence launches Darwin, first medical AI to ace US licensing benchmark

OpenEvidence launches Darwin, first medical AI to ace US licensing benchmark

New Capabilities

Four-model family spans five-second answers to deep research reports

7 days ago: OpenEvidence launches four medical AI models

Overview

Updated 6 days ago

OpenEvidence released four medical AI models on September 5. Darwin, the flagship, is the first to score 100% on MedQA, a benchmark drawn from US medical licensing exam questions.

Osler, Sackett, and Snow answer in five seconds to five minutes and are free to verified US clinicians. OpenEvidence says more US physicians use its platform than all rivals combined, with 1.12 million verified clinicians. Darwin is in research preview; the company published its benchmark annotations and Darwin's full responses.

Why it matters

The medical AI more US physicians use than any rival now has a perfect exam-scorer, with its reasoning bound for the exam room.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

100%
Darwin score on MedQA
First perfect score on the independent medical AI benchmark.
660/660
MedQA questions answered correctly
Darwin answered all 660 questions on the physician-refined evaluation set correctly.
Under 5 seconds
Osler response time
Fastest production model, built for point-of-care questions.
~30 seconds
Sackett response time
Mid-tier model with additional rounds of searching.
~5 minutes
Snow response time
Deep-investigation model for complex cases and unsettled evidence.
4
Models released
Darwin in research preview; Osler, Sackett, Snow open to all clinicians.
1.12M
Verified US clinicians
Licensed-verified clinicians using OpenEvidence as of September 2026, including physicians, nurses, nurse practitioners, and physician assistants.

Voices

Curated perspectives — historical figures and your fellow readers.

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

People Involved

Organizations Involved

Timeline

June 2026 September 2026

3 events Latest: 7 days ago
  1. OpenEvidence launches four medical AI models

    Latest Product Launch

    Darwin tops MedQA with the first perfect score; Osler, Sackett, and Snow open free to all clinicians on web and mobile.

  2. STAT previews the model family

    Media Coverage

    STAT Health Tech reports on the upcoming release and generative AI medical devices reaching the market quickly.

  3. Nature Medicine study questions specialized clinical tools

    Research

    Independent study found general-purpose frontier models outperformed an earlier OpenEvidence tool and UpToDate Expert AI on benchmarks and real clinical queries. It does not evaluate Darwin.

Scenarios

1

Darwin graduates to the exam room

Likely Resolves by Sep 5, 2027

Discussed by: OpenEvidence's stated roadmap; industry analysts at Fierce Healthcare

Darwin's research preview expands from institutional partners like NORD and academic researchers to general availability for all verified clinicians. OpenEvidence has said Darwin's capabilities will flow into Osler, Sackett, and Snow as safeguards are validated, which could push the perfect benchmark score into daily point-of-care use.

2

Independent evaluation tempers the perfect score

Possible Resolves by Sep 5, 2027

Discussed by: Clinical AI researchers; the June 2026 Nature Medicine precedent

A peer-reviewed evaluation tests Darwin at real clinical sites with clinicians in the loop and finds benchmark gains do not translate to better decisions. The June 2026 Nature Medicine study set this precedent for an earlier OpenEvidence tool, and the company itself concedes single-model benchmarks do not match how tools are used.

3

Rivals match the perfect score

Possible Resolves by Jun 5, 2027

Discussed by: Frontier labs shipping medical models: Anthropic (Claude Fable 5), OpenAI (GPT-5.6 Sol), Google (Gemini 3.7 Flash)

A competitor publishes a model scoring 100% on MedQA, erasing Darwin's benchmark edge. Rivals sit close already: Claude Fable 5 at 99.7%, Gemini 3.7 Flash at 99.2%, GPT-5.6 Sol at 99.1%. Darwin answers 660 of 660 questions correctly, so the gap is narrowest at the very top.

Historical Context

3 moments from history that rhyme with this story — and how they unfolded.

2013-2022

IBM Watson for Oncology (2013-2022)

IBM's medical AI, built after Watson won Jeopardy, was sold to hospitals as a cancer treatment advisor. MD Anderson cancelled a $62 million project in 2017 after internal documents showed the system produced unsafe treatment recommendations.

Then

Maimonides Medical Center stopped using it; IBM sold off Watson Health in 2022.

Now

Became the cautionary tale for medical AI that dazzles in demos but fails in real clinical settings.

Why this matters now

Darwin faces the same question IBM Watson could not answer: whether benchmark mastery survives contact with real patients and real clinician judgment.

July 2021

DeepMind's AlphaFold (2021)

DeepMind's specialized model predicted protein structures from amino acid sequences, solving a 50-year biology problem that general approaches could not crack, validated independently at CASP14.

Then

The prediction database became a standard research tool used by hundreds of thousands of scientists.

Now

Proved a single-purpose model can beat general ones at a hard scientific task.

Why this matters now

Supports the argument that a dedicated medical reasoning model like Darwin can outperform general frontier models in clinical domains.

March 2023

GPT-4 tops the USMLE (2023)

OpenAI's GPT-4 scored in roughly the 90th percentile on USMLE-style questions in a widely cited study, showing general-purpose models could approach medical exam competence without medical training.

Then

Spurred a wave of medical AI startups and regulatory review of AI clinical tools.

Now

Set the benchmark arms race Darwin now resets with a perfect score.

Why this matters now

Darwin's 100% on MedQA is the direct next step in a benchmark progression general models began in 2023.

Sources

(11)