Extracting Requirements to Fine-Tune Large Language Models in Multimodal Medical Diagnosis
摘要
Large language models constitute a breakthrough state-of-the-art Artificial Intelligence technology which is rapidly evolving and promises to aid in medical diagnosis. In this work, we engineered an evaluation system that incorporates two independent steps of a novel methodology, namely (1)evaluation via structured interactions and (2)follow-up, domain-specific analysis based on data extracted via the previous interactions. Using this paradigm, (1)we evaluate the correctness and accuracy of generated medical diagnosis with publicly available multimodal-multiple-choice-questions in the domain of Pathology and (2) proceed to a systemic and comprehensive analysis of extracted results. These results serve as the basis for further targeted optimization and fine-tuning of the employed Generative model. We used the Generative Pre-trained Transformer version 4 with vision (GPT-4V) as the model that generates responses to complex, medical questions consisting of both images and text, and we explored a wide range of diseases, conditions, chemical compounds, and related entity types that are included in the vast knowledge domain of Pathology. The model scored approximately 84% of correct diagnoses. We further analyzed the findings of our work, following an analytical approach which included Image-Metadata-Analysis, Named-Entity-Recognition and Knowledge-Graphs. Weaknesses of the model were revealed on specific knowledge paths, leading to a further understanding of its shortcomings in specific areas and discovering paths for fine-tuning and improvements. Our methodology and findings transcend the constraints of a singular model or domain, offering broad applicability and generalizability across various contexts.