Independent MedXpertQA analysis covering Text and MM test splits, answer choices, difficulty filtering, sampled evaluations, prompt conditions and historical results.
The release includes development examples; they are separate from the 4,450 combined test questions. Read versioned results →
01 /
Track boundaries
Text and MM use different question populations and option counts.
02 /
Population coverage
Full-test and sampled measurements need different labels.
03 /
Historical context
Model version, prompt and answer extraction remain part of every result.
Our analytical question
MedXpertQA stresses expert medical knowledge and reasoning through separate text and multimodal examination tracks. We explain the construction choices behind its difficulty, inspect the difference between full and sampled evaluations, and preserve the conditions behind selected 2025 paper results. Our original analysis and task explorer help readers assess what the benchmark measures. We are independent of its creators and do not claim new experiments or a current frontier leaderboard.
Filter by track, author-defined reasoning label or evaluation coverage. The map describes published task conditions rather than assigning invented capability scores.
This is our independent analytical map of published tasks. It does not execute an evaluation, predict a model’s performance or establish clinical benefit. [1][2][3]
A practical audit of full versus sampled evaluations, prompts and historical model rows.
5 min read
Questions, answered
Read the result in context.
Specific tasks. Stated conditions. Inspect every source.
How many MedXpertQA test questions are there?+
The original release has 2,450 Text test questions and 2,000 MM test questions, plus five development questions for each track.
Why does Text use ten options while MM uses five?+
The authors augment the Text option sets, while many image-dependent MM options cannot be expanded in the same way. The different choice counts also imply different uniform-guess reference rates.
Can Text and MM accuracy be used to measure the benefit of images?+
Not directly. The tracks contain different questions. A paired image-ablation experiment on the same cases would address a different and more specific question.
Are all original model rows full-test evaluations?+
No. The source marks o1 and o3-mini rows as sampled evaluations. Our selected panels use full-set rows and preserve the source version.
Working tool / saved on this device
Prepare a benchmark comparison brief
Interactive worksheet
Use this secondary checklist to document a run or literature comparison after inspecting the named benchmark conditions. Completion records documentation, not performance.
Evidence you can inspect. Benchmark dossiers distinguish published facts from our interpretation, with source versions and access notes attached.