Patient world model benchmark
Can a model keep one patient in mind?
RADWORLD follows radiological reasoning from a CT examination to the next clinical action. Six connected operations reveal where vision-language models stop describing and start losing the patient state.
latentpatient
statest
statest
1Assess
2Read
3Compare
4Predict
5Conclude
6Advise
23 models22,866 CT–report pairs179 findings19 organ systems
01
Six operations. One state.
Each operation asks a different question about the same patient. Strong answers should be correct—and agree with one another.
What the benchmark reveals
Anatomy is visible.
The patient state is not.
Models can recognize gross anatomy, yet lose findings as the prompt changes, fail to recover change from paired examinations, and miss new findings on hidden follow-up.
Compare all 22 models →91.6%best organ visibility
66.5%best Read breadth
39.7%best detect + locate
3 / 40best strict integrated chains
02
Explore beyond the paper
Filter thousands of finding–model results, inspect operation-specific failure modes, and compare model families without flattening incompatible metrics into one score.