<p>Depression is a prevalent mental health disorder with significant consequences, making its early and accurate detection essential. Traditional diagnostic methods relying on self-report questionnaires are subjective, underscoring the need for objective, automated approaches. However, unimodal models often fail to capture the full complexity of depressive symptoms. To address this, we propose VTA-DepressNet, a multimodal deep learning architecture that integrates visual, audio, and textual features through an attention-driven fusion mechanism. The model utilizes a Conv-BiLSTM for visual features, a Conv-BiGRU for audio features, and a Transformer encoder with an attention mechanism for textual data. Experimental results on the DAIC-WOZ dataset with fivefold cross-validation demonstrate that VTA-DepressNet achieves an F1-score of 0.83, significantly outperforming unimodal baselines. To validate generalization, the model was further evaluated on the Extended DAIC (E-DAIC) dataset, achieving a competitive F1-score of 0.77. These results not only affirm the effectiveness of combining behavioral and linguistic cues but also highlight the potential of digital mental-health tools for early screening and intervention in both clinical and educational contexts. In practice, VTA-DepressNet could be integrated into telehealth platforms or school-based mental health monitoring systems to support timely and scalable depression assessment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VTA-DepressNet: a dual-level attention fusion for multimodal depression detection

  • Nguyen Nguyen,
  • Minh Nguyen,
  • Tri Doan,
  • Hau Vo,
  • Nha Tran,
  • Hung Nguyen

摘要

Depression is a prevalent mental health disorder with significant consequences, making its early and accurate detection essential. Traditional diagnostic methods relying on self-report questionnaires are subjective, underscoring the need for objective, automated approaches. However, unimodal models often fail to capture the full complexity of depressive symptoms. To address this, we propose VTA-DepressNet, a multimodal deep learning architecture that integrates visual, audio, and textual features through an attention-driven fusion mechanism. The model utilizes a Conv-BiLSTM for visual features, a Conv-BiGRU for audio features, and a Transformer encoder with an attention mechanism for textual data. Experimental results on the DAIC-WOZ dataset with fivefold cross-validation demonstrate that VTA-DepressNet achieves an F1-score of 0.83, significantly outperforming unimodal baselines. To validate generalization, the model was further evaluated on the Extended DAIC (E-DAIC) dataset, achieving a competitive F1-score of 0.77. These results not only affirm the effectiveness of combining behavioral and linguistic cues but also highlight the potential of digital mental-health tools for early screening and intervention in both clinical and educational contexts. In practice, VTA-DepressNet could be integrated into telehealth platforms or school-based mental health monitoring systems to support timely and scalable depression assessment.