A TRUST AWARE MULTI MODAL EXPLAINABLE DEEP LEARNING FRAMEWORK FOR MEDICAL DIAGNOSIS USING MEDICAL IMAGES AND ELECTRONIC HEALTH RECORDS
DOI:
https://doi.org/10.4238/eab2bg41Keywords:
Explainable Artificial Intelligence (XAI), Multi-Modal Learning, Medical Diagnosis, Deep Learning, Electronic Health Records (EHR), Medical Image Analysis, Trustworthy AI, Clinical Decision Support Systems.Abstract
Accurate and timely medical diagnosis plays a critical role in improving patient outcomes and supporting effective clinical decision-making. In recent years, Artificial Intelligence (AI) and deep learning techniques have demonstrated remarkable success in a wide range of healthcare applications, including disease detection, risk prediction, and medical image analysis. Despite their high predictive performance, many deep learning models operate as black-box systems, making it difficult for healthcare professionals to understand how diagnostic decisions are generated. This lack of transparency presents significant challenges for clinical adoption, particularly in environments where trust, accountability, and regulatory compliance are essential. In addition, most existing diagnostic models rely on a single source of information and do not fully utilize the complementary insights available from medical images and electronic health records (EHRs).To overcome these limitations, this paper presents MedXTrust, a trust-aware multi-modal explainable deep learning framework for medical diagnosis. The proposed framework combines medical imaging data and structured clinical information within a unified architecture to support more accurate and interpretable diagnostic predictions. A Vision Transformer (ViT) is employed to learn representative features from medical images, while a Clinical Transformer is used to model patient information extracted from EHRs. These feature representations are integrated through a cross-modal attention mechanism that captures complex relationships between imaging findings and clinical variables. To enhance transparency and support clinical interpretation, the framework incorporates a hybrid explainability module consisting of SHAP-based feature attribution, Grad-CAM++ visual explanations, and counterfactual reasoning. This combination enables both global and patient-specific explanations, helping clinicians understand the factors influencing model predictions. Furthermore, a trust evaluation module is introduced to assess explanation fidelity, stability, clinician agreement, and overall system reliability.The proposed framework is evaluated using publicly available medical imaging and EHR datasets, with performance assessed through diagnostic accuracy, explainability quality, and clinician-centered trust metrics. The results demonstrate the potential of MedXTrust to deliver highly accurate predictions while providing meaningful and actionable explanations. By integrating multi-modal learning, explainable AI, and trust assessment within a single framework, this study contributes to the development of transparent, reliable, and clinically applicable AI systems that align with emerging requirements for trustworthy healthcare technologies.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

