Methods & resources

A benchmark designed to be inspected, not just ranked.

RADWORLD keeps model outputs, references, operation-specific metrics, and figure sources separate so every result can be traced to its protocol.

02

Fair comparisons

Patient-disjoint cohorts, frozen model releases, shared prompts, explicit interface eligibility, and operation-specific baselines.

03

Reader validation

Independent radiologist studies test whether report-derived targets remain recognizable under the rendered inputs given to models.

Download and reproduce

The interactive site is generated from the same verified result files used by the manuscript.

Citation

RADWORLD

Do vision–language models have a world model of the patient?

@article{radworld2026, title={RADWORLD: Do Vision--Language Models Have a World Model of the Patient?}, year={2026} }