Explainable Artificial Intelligence, Clinician Trust and Diagnostic Accuracy in Multi-modal Medical Imaging: A Critical Narrative Review
Ogechi Okoroafor *
Morgan State University, Baltimore, United States of America.
Musa Shamaki Ibrahim
Cell kinetics lab, DLMP, Mayo Clinic, Rochester, United States of America.
Chimuanya Ibecheozor
Norwich University of the Arts, Norwich, United Kingdom.
Chioma Emmanuela Ukatu
Wayne County Community College, Detroit, United States of America.
Emmanuel Arhin
Goldey Beacom University, Wilmington, United States of America.
*Author to whom correspondence should be addressed.
Abstract
Explainable artificial intelligence is widely promoted as the mechanism through which opaque deep learning systems will earn the confidence of clinicians and improve diagnostic performance in medical imaging. The claim has acquired regulatory and commercial momentum faster than it has acquired empirical support, and it has been examined almost entirely in single-modality settings even though contemporary diagnostic models increasingly integrate radiological, pathological, textual and structured clinical data. This review critically evaluates the evidence linking explanation to clinician trust and to diagnostic accuracy, with particular attention to multi-modal imaging contexts. Literature was identified through structured searching of openly accessible scholarly indexes and metadata registries, supplemented by citation tracking and verification of every source against its record of registration. The synthesis distinguishes three largely separate evidence streams that are frequently conflated: technical evaluations of explanation fidelity, human-participant studies of trust and reliance, and normative arguments about transparency. Technical evaluations consistently show that widely deployed attribution maps localise abnormalities less accurately than dedicated detection or segmentation models and than expert annotation, and that several methods are insensitive to the parameters of the models they purport to describe. Human-participant studies are more heterogeneous. Explanations improve accuracy in some reader experiments, particularly for clinicians working outside their task speciality, yet they also increase acceptance of incorrect advice and do not reliably protect against systematically biased models. Effects depend on explanation format, reader expertise, the correctness of the underlying prediction and the way trust is operationalised, which varies markedly across studies. Evidence specific to multi-modal systems is scarce, attribution methods rarely apportion importance across modalities in a clinically meaningful way, and evaluation frameworks derived from single-image tasks transfer poorly. The most defensible conclusion is that explanation is neither a general solution to opacity nor an irrelevance, but a design variable whose effect on reliance is conditional and occasionally harmful.
Keywords: Explainable artificial intelligence, medical image analysis, multimodal data fusion, clinician trust, diagnostic accuracy, automation bias, uncertainty quantification