<p>Major depressive disorder and generalized anxiety disorder together account for a substantial share of global disability. They occur more frequently in women and their frequent co-occurrence complicates both diagnosis and treatment. To investigate how comorbidity affects speech-based assessment, we collected conversational speech data covering both negative and positive topics from 130 Turkish women at a psychiatric clinic. Each patient was diagnosed with depression, anxiety, or comorbid depression and anxiety by a psychiatrist. From each recording, we extracted classic acoustic descriptors (MFCCs and the AVEC-2013 feature set) alongside learned embeddings from three pretrained self-supervised models—HuBERT, WavLM, and wav2vec2—each fine-tuned with a lightweight regression head. The LoRA fine-tuned wav2vec2 achieved the best accuracy (depression: RMSE 8.44, MAE 6.84; anxiety: RMSE 12.49, MAE 10.71), surpassing WavLM, HuBERT, and all baseline feature sets while using significantly fewer trainable parameters. To evaluate architectural variants, we experimented with five alternative heads on the wav2vec2 backbone. Recordings in response to negative question consistently performed better. Comorbid depression and anxiety cases were assessed more accurately than depression-only cases (RMSE 7.34 vs. 8.03), underscoring the value of modeling clinical heterogeneity. Together, our research shows that co-occurring anxiety enhances the acoustic salience of depressive symptoms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Estimating depression and anxiety scores from conversational speech in females with and without comorbidity

  • Aslı Beşirli,
  • Tuğrahan Karakadıoğlu,
  • Cenk Demiroğlu

摘要

Major depressive disorder and generalized anxiety disorder together account for a substantial share of global disability. They occur more frequently in women and their frequent co-occurrence complicates both diagnosis and treatment. To investigate how comorbidity affects speech-based assessment, we collected conversational speech data covering both negative and positive topics from 130 Turkish women at a psychiatric clinic. Each patient was diagnosed with depression, anxiety, or comorbid depression and anxiety by a psychiatrist. From each recording, we extracted classic acoustic descriptors (MFCCs and the AVEC-2013 feature set) alongside learned embeddings from three pretrained self-supervised models—HuBERT, WavLM, and wav2vec2—each fine-tuned with a lightweight regression head. The LoRA fine-tuned wav2vec2 achieved the best accuracy (depression: RMSE 8.44, MAE 6.84; anxiety: RMSE 12.49, MAE 10.71), surpassing WavLM, HuBERT, and all baseline feature sets while using significantly fewer trainable parameters. To evaluate architectural variants, we experimented with five alternative heads on the wav2vec2 backbone. Recordings in response to negative question consistently performed better. Comorbid depression and anxiety cases were assessed more accurately than depression-only cases (RMSE 7.34 vs. 8.03), underscoring the value of modeling clinical heterogeneity. Together, our research shows that co-occurring anxiety enhances the acoustic salience of depressive symptoms.