Multimodal AI for Holistic Oral Cancer Diagnosis and Assessment
摘要
This chapter introduces a novel multimodal AI system designed for holistic oral cancer diagnosis and assessment. By integrating analyses from oral cavity RGB images, histopathology images, and structured patient clinical data, the system addresses limitations of single-modality diagnostics. Convolutional Neural Networks (CNNs) are employed for image feature extraction and classification, while a Large Language Model (LLM) synthesizes insights from all modalities to generate a comprehensive medical report. The system enhances diagnostic accuracy, facilitates early detection, and supports clinical decision-making by providing a richer, interpretable, and clinically relevant assessment. Implementation details, including data preprocessing, modality-specific AI models, feature fusion, and the role of LLMs, are discussed. The chapter also introduces the Multimodal Oral Cancer Assistant (MOCA) agent, an application leveraging this system to empower clinicians with integrated insights. Future directions include expanding datasets, advanced feature fusion techniques, clinical validation, and integration with Electronic Health Records (EHRs).