Estimating depression and anxiety scores from conversational speech in females with and without comorbidity
摘要
Major depressive disorder and generalized anxiety disorder together account for a substantial share of global disability. They occur more frequently in women and their frequent co-occurrence complicates both diagnosis and treatment. To investigate how comorbidity affects speech-based assessment, we collected conversational speech data covering both negative and positive topics from 130 Turkish women at a psychiatric clinic. Each patient was diagnosed with depression, anxiety, or comorbid depression and anxiety by a psychiatrist. From each recording, we extracted classic acoustic descriptors (MFCCs and the AVEC-2013 feature set) alongside learned embeddings from three pretrained self-supervised models—HuBERT, WavLM, and wav2vec2—each fine-tuned with a lightweight regression head. The LoRA fine-tuned wav2vec2 achieved the best accuracy (depression: RMSE 8.44, MAE 6.84; anxiety: RMSE 12.49, MAE 10.71), surpassing WavLM, HuBERT, and all baseline feature sets while using significantly fewer trainable parameters. To evaluate architectural variants, we experimented with five alternative heads on the wav2vec2 backbone. Recordings in response to negative question consistently performed better. Comorbid depression and anxiety cases were assessed more accurately than depression-only cases (RMSE 7.34 vs. 8.03), underscoring the value of modeling clinical heterogeneity. Together, our research shows that co-occurring anxiety enhances the acoustic salience of depressive symptoms.