A Multimodal Deep Learning Model for Laryngeal and Hypopharyngeal Lesions Diagnosis: A Multicenter Retrospective Study

Published in SSRN (Preprint), 2025

Key Innovations

  • Multimodal Fusion Architecture: Simultaneously processes laryngoscopic images, clinical features (age and gender) and patient complaint
  • Clinical Validation: Achieved superior results comparing to clinicians with over 3 years of experience
  • Contrastive Learning: Utilizes a contrastive learning objective function that enhances model performance

Recommended citation: Zhang, J., Liang, K., Yao, Q., Zhang, Y., et al. (2025). A Multimodal Deep Learning Model for Laryngeal and Hypopharyngeal Lesions Diagnosis: A Multicenter Retrospective Study. SSRN. https://doi.org/10.2139/ssrn.5216153
Download Paper | Download Bibtex