Expert medical examination analysis

Two tracks. Keep the conditions attached.

Independent MedXpertQA analysis covering Text and MM test splits, answer choices, difficulty filtering, sampled evaluations, prompt conditions and historical results.

MedXpertQA / test-track comparison

Separate populations, separate answer spaces

[1][2][3]

Text

2,450 test questions

Medical examination text

10 answer options [1]

Development5 additional questions [1]

MM

2,000 test questions

Examination text + associated images

5 answer options [1]

Development5 additional questions [1]

The release includes development examples; they are separate from the 4,450 combined test questions. Read versioned results →

01 /

Track boundaries

Text and MM use different question populations and option counts.

02 /

Population coverage

Full-test and sampled measurements need different labels.

03 /

Historical context

Model version, prompt and answer extraction remain part of every result.

Our analytical question

MedXpertQA stresses expert medical knowledge and reasoning through separate text and multimodal examination tracks. We explain the construction choices behind its difficulty, inspect the difference between full and sampled evaluations, and preserve the conditions behind selected 2025 paper results. Our original analysis and task explorer help readers assess what the benchmark measures. We are independent of its creators and do not claim new experiments or a current frontier leaderboard.

Read the original evidence closely

The benchmark, unpacked.

Primary source library ↗
Dossier01

ICML 2025 / arXiv v3

MedXpertQA Text and MM ↗

Expert exam difficulty needs an equally careful comparison.

UnitOne examination questionMeasureMultiple-choice accuracy

An original analytical tool

Inspect the MedXpertQA evaluation condition

Evidence explorer

Filter by track, author-defined reasoning label or evaluation coverage. The map describes published task conditions rather than assigning invented capability scores.

7 of 7 evidence entries shown

Track

Text: full test

Read dossier ↗
Input
Text question with ten options
Output
Chosen option
Measure
Accuracy
Interpretation boundary

No image benefit can be inferred by comparing different MM questions.

[1]
Track

MM: full test

Read dossier ↗
Input
Multimodal question with five options
Output
Chosen option
Measure
Accuracy
Interpretation boundary

Input includes the question’s supplied images.

[1]
Reasoning label

Text: reasoning subset

Read dossier ↗
Input
Text cases assigned reasoning label
Output
Chosen option
Measure
Subset accuracy
Interpretation boundary

This is an author-defined question attribute, not direct access to model reasoning.

[1]
Reasoning label

Text: understanding subset

Read dossier ↗
Input
Text cases assigned understanding label
Output
Chosen option
Measure
Subset accuracy
Interpretation boundary

Compare within the labeled subset.

[1]
Reasoning label

MM: reasoning subset

Read dossier ↗
Input
Multimodal cases assigned reasoning label
Output
Chosen option
Measure
Subset accuracy
Interpretation boundary

Difficulty and modality are coupled in these selected questions.

[1]
Reasoning label

MM: understanding subset

Read dossier ↗
Input
Multimodal cases assigned understanding label
Output
Chosen option
Measure
Subset accuracy
Interpretation boundary

Record the number of included questions.

[1]
Coverage

Sampled expensive-model condition

Read dossier ↗
Input
Stratified 10% sampled questions, seed 42
Output
Chosen option
Measure
Sampled accuracy
Interpretation boundary

Original o1/o3-mini condition; distinguish from full-test measurements.

[1]

This is our independent analytical map of published tasks. It does not execute an evaluation, predict a model’s performance or establish clinical benefit. [1][2][3]

Original analysis / methods and interpretation

What the score leaves unsaid.

All analyses →

Questions, answered

Read the result in context.

Specific tasks. Stated conditions.
Inspect every source.

How many MedXpertQA test questions are there?

The original release has 2,450 Text test questions and 2,000 MM test questions, plus five development questions for each track.

Why does Text use ten options while MM uses five?

The authors augment the Text option sets, while many image-dependent MM options cannot be expanded in the same way. The different choice counts also imply different uniform-guess reference rates.

Can Text and MM accuracy be used to measure the benefit of images?

Not directly. The tracks contain different questions. A paired image-ablation experiment on the same cases would address a different and more specific question.

Are all original model rows full-test evaluations?

No. The source marks o1 and o3-mini rows as sampled evaluations. Our selected panels use full-set rows and preserve the source version.

Working tool / saved on this device

Prepare a benchmark comparison brief

Interactive worksheet

Use this secondary checklist to document a run or literature comparison after inspecting the named benchmark conditions. Completion records documentation, not performance.

Choose the population

Evidence you can inspect. Benchmark dossiers distinguish published facts from our interpretation, with source versions and access notes attached.

Download the evidence ↗