AI Explanations Improve Diagnostic Accuracy in Radiology
Research from Ludwig-Maximilians-Universität München (LMU) indicates that AI models like ChatGPT can enhance diagnostic accuracy in radiology, provided their explanations are clear and step-by-step.

A study from Ludwig-Maximilians-Universität München (LMU) has found that artificial intelligence (AI) can improve diagnostic accuracy in radiology, but its usefulness is contingent on how the AI explains its recommendations.
Large language models such as ChatGPT are increasingly discussed as support tools in medicine. They can summarize information, suggest diagnoses, and explain their assessments in understandable language. The study's key finding is that the quality of the AI's explanations is crucial for its practical benefit.
A research team from LMU, LMU Klinikum, the Karlsruhe Institute of Technology, and the University of Bayreuth conducted a randomized experiment involving 101 radiologists. Participants assessed real clinical cases with imaging results and were tasked with formulating a diagnosis. The experiment compared a control group without AI to three groups receiving different types of support from a multimodal language model: a diagnosis only, a differential diagnosis, or a step-by-step "chain-of-thought" explanation. The latter transparently detailed image characteristics, clinical clues, and exclusion criteria.
The findings revealed that radiologists achieved the highest diagnostic accuracy when utilizing the step-by-step AI explanations. Their success rate was 12.2 percentage points higher than that of the control group. Simple diagnostic outputs and differential diagnoses yielded poorer results. In cases with erroneous AI suggestions, participants were more likely to follow differential diagnoses, indicating a risk of automation bias. The step-by-step explanations, however, helped in identifying correct clues and recognizing errors.
These results suggest that not only the quality of the diagnosis is important, but also the form of explanation that supports radiologists in their critical evaluation. Step-by-step reasoning makes the model's argumentation visible and facilitates comparison with the physician's own expertise. According to Professor Stefan Feuerriegel, a corresponding author of the study, these conclusions extend beyond medicine, highlighting the importance for users to actively scrutinize and understand AI-generated responses, rather than simply accepting them.