ARTIFICIAL INTELLIGENCE-BASED MULTIMODAL DEEP LEARNING FRAMEWORK FOR EARLY DETECTION OF LARYNGEAL CANCER USING VOICE SIGNALS AND LARYNGOSCOPIC IMAGES
DOI:
https://doi.org/10.4238/afwm9t31Keywords:
Laryngeal Cancer; Artificial Intelligence; Multimodal Deep Learning; Voice Analysis; Laryngoscopic Imaging; Medical Image Analysis; Computer-Aided Diagnosis; Head and Neck Oncology.Abstract
Laryngeal cancer is one of the most prevalent malignancies of the head and neck and is associated with substantial morbidity due to impaired voice production, swallowing dysfunction, and reduced quality of life. Early diagnosis is essential for improving treatment outcomes and preserving laryngeal function; however, conventional diagnostic pathways frequently rely on invasive procedures, specialist interpretation, and advanced imaging techniques that may delay clinical decision-making. Artificial intelligence (AI)-based computer-aided diagnostic systems have emerged as promising tools for supporting early disease detection using multimodal clinical data. This study proposes a multimodal deep learning framework for the automated detection of laryngeal cancer by integrating acoustic voice analysis with laryngoscopic image assessment. The proposed framework incorporates preprocessing of voice recordings and laryngoscopic images, extraction of acoustic biomarkers including Mel-Frequency Cepstral Coefficients (MFCCs), Linear Predictive Coding (LPC) coefficients, pitch, formant frequencies, jitter, and shimmer, together with deep image representations obtained using convolutional neural networks. A hybrid CNN–Long Short-Term Memory (LSTM) architecture was employed to fuse complementary acoustic and visual features for binary classification of healthy individuals and patients with confirmed laryngeal cancer. The proposed framework was evaluated using a multimodal dataset comprising 500 participants, including healthy controls and clinically confirmed laryngeal cancer cases. Experimental results achieved an overall classification accuracy of 97.4%, sensitivity of 96.8%, specificity of 97.9%, precision of 97.2%, F1-score of 97.0%, and an area under the receiver operating characteristic curve (AUC) of 0.985. Comparative analysis with conventional machine learning approaches, including Support Vector Machines, Artificial Neural Networks, and Random Forest classifiers, demonstrated improved diagnostic performance of the proposed multimodal framework. These findings indicate that the integration of acoustic voice biomarkers and laryngoscopic imaging enhances the accuracy and reliability of AI-assisted laryngeal cancer detection. The proposed approach has the potential to support early clinical screening, facilitate timely referral for specialist evaluation, and improve decision-making in both hospital-based practice and telemedicine settings.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

