Medical AI assistants are only as trustworthy as their reasoning — and when they hallucinate, the consequences can be life-threatening. Most existing tools for catching hallucinations in medical AI treat errors as a single category, leaving clinicians and developers blind to where reasoning breaks down. ClinHallu addresses this by decomposing the reasoning process into three stages: visual recognition, knowledge recall, and reasoning integration. With over 7,000 validated cases, it enables developers to pinpoint exactly which stage is responsible for an error. Potential applications include building safer radiology AI, clinical decision support systems, and diagnostic tools where traceability and accuracy are paramount.
Authors: Sicheng Yang, Hangjie Yuan, Wenjun Zhang, Jinwang Wang, Yichen Qian, Weihua Chen, Fan Wang, Lei Zhu
Paper: https://arxiv.org/abs/2606.14697v1