ClinicalBench Contact us

Clinical agent evaluation / 01

The decision.
And the path.

A correct diagnosis is only part of the story.
Watch what an AI agent asks, examines, orders and treats.

Explore a real run
API-native encountersReplayable action traces
Inside an encounterRecorded, not live
01 / ChooseClinical agent
GPT 6.1 Sol
Allowed action

Select an action to see the exchange.

API request → finding
02 / ObserveClinicalGym API
↺ next decision
Every action recorded
03 / EvaluateTrace + case points

ClinicalGym runs the encounter. ClinicalBench compares the runs.

Featured modelGPT 6.1 Sol
Emergency-medicine cases—
Encounters closed—
Diagnosis-gated weighted composite—

A real run / 02

One encounter. Many decisions.

Actual action order. No invented transcript.

Loading the saved benchmark…
Recorded encounterGPT 6.1 Sol
HistoryExamTestsDiagnosisTreatmentCatalogFinish
—
Select an action

Follow the recorded path

Action types only. Private case text and model reasoning are not exposed.

Benchmark / 03

Reasoning Effort High.

Model / saved comparisonWeighted composite
    What is being compared?

    Download the curated snapshot ↓

    What the number means / 04

    Measure more than the answer.

    Outcome

    What was diagnosed and treated.

    Current: simulation points

    Process

    What was asked, examined and ordered.

    Current: recorded action trace

    Safety

    What was unsafe—or left undone.

    Assessment: actions and omissions
    Doctor-authored. Doctor-validated.

    Every case was created by a medical doctor and validated by at least two other doctors.

    Get in touch / 05

    Contact us

    Questions, collaborations or benchmark enquiries.

    Mariusz Kurman · ClinicalBench