Patient world model benchmark

Can a model keep one patient in mind?

RADWORLD follows radiological reasoning from a CT examination to the next clinical action. Six connected operations reveal where vision-language models stop describing and start losing the patient state.

latentpatient
state
st
1Assess
2Read
3Compare
4Predict
5Conclude
6Advise
23 models22,866 CT–report pairs179 findings19 organ systems
01

Six operations. One state.

Each operation asks a different question about the same patient. Strong answers should be correct—and agree with one another.

What the benchmark reveals

Anatomy is visible.
The patient state is not.

Models can recognize gross anatomy, yet lose findings as the prompt changes, fail to recover change from paired examinations, and miss new findings on hidden follow-up.

Compare all 22 models →
91.6%best organ visibility
66.5%best Read breadth
39.7%best detect + locate
3 / 40best strict integrated chains
02

Explore beyond the paper

Filter thousands of finding–model results, inspect operation-specific failure modes, and compare model families without flattening incompatible metrics into one score.