The Diagnosis of Typical Medical Cases Through Optimized Fine-Tuning of Large Language Models
摘要
Large Language Models (LLMs) have gained widespread attention in both academia and industry due to their exceptional performance in natural language processing (NLP) tasks. While many NLP tasks related to healthcare already have established evaluation benchmarks, there is currently no comprehensive evaluation task or standard for diagnostic consistency based on typical clinical cases. In this paper, we introduce a diagnostic consistency evaluation task based on real-world clinical cases. This task integrates diagnostic information for common diseases and aims to provide a comprehensive and objective assessment of the diagnostic capabilities of medical LLMs by accurately simulating the decision-making process of doctors during disease diagnosis. We fine-tuned a general-purpose LLM using LoRA and QLoRA methods, conducting generative end-to-end evaluations during the training process. Additionally, we enhanced model performance through model ensemble techniques. Our approach achieved the highest ranking in both the preliminary and final rounds of CHIP2024: The Diagnostic Consistency Task in Typical Medical Case Records, validating the effectiveness of our method.