Innovative application of large language model in the prediction of potential depression in the elderly: based on CHARLS dataset
摘要
Screening for late-life depression remains difficult in rapidly ageing societies due to heterogeneous symptoms, uneven resources across regions, and the lack of simple tools that work at scale and over time. We address this by rendering each respondent’s CHARLS record into a concise, templated natural-language summary that a large language model can comprehend, unifying structured health, functional, socioeconomic, and social-support fields (and, where available, interview text) in a single representation. We then fine-tune instruction models (DeepSeek and Qwen) with parameter-efficient adapters (LoRA) on attention projections, alongside explicit handling of missingness and class imbalance, and with de-identification for governance. In our setup, LoRA drastically reduces the number of trainable parameters, helping to curb compute and overfitting while keeping accuracy. On a held-out split of CHARLS, DeepSeek variants achieve macro-F1 around 80% (weighted-F1