In current clinical practice, effective methods for the automatic diagnosis of depression remain limited. Accurate diagnosis of depression is essential, as it ensures timely and appropriate treatment for depressive patients, thereby reducing long-term health burdens on both individuals and society. Recently, automatically predicting depression severity through audio signals has garnered increasing attention. These methods typically depend on handcrafted features or deep features achieved by deep learning. Given the complementarity between handcrafted and deep features, their combination enables a more effective capture of emotional sensitivity associated with depression. Motivated by this observation, this work proposes an automated attention-based audio depression recognition integrating handcrafted and deep features that leverage the strengths of both handcrafted and deep features to enhance the accuracy of depression detection. The proposed method comprises two primary components: a feature extraction module and an attention-based feature fusion module. Specifically, the feature extraction module extracts both handcrafted LLDs and deep features by using a pre-trained self-supervised model. Moreover, a new strategy of feature selection based on an attention mechanism is designed to adaptively select more effective handcrafted features for downstream tasks. The attention-based feature fusion module employs a cross-attention mechanism to effectively integrate these features, followed by a fully-connected layer to predict the depression scores. Experimental results on the typical AVEC2013 dataset indicate that the proposed method significantly improves the accuracy of audio depression detection analysis over the other methods and demonstrates its substantial potential applications to automatic depression diagnosis in clinical practice.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention-Based Audio Depression Recognition Integrating Handcrafted and Deep Features

  • Chenyu Jin,
  • Shuchang Zhao,
  • Shiqing Zhang,
  • Zhewei Fang,
  • Junjie Xie,
  • Ying Chen

摘要

In current clinical practice, effective methods for the automatic diagnosis of depression remain limited. Accurate diagnosis of depression is essential, as it ensures timely and appropriate treatment for depressive patients, thereby reducing long-term health burdens on both individuals and society. Recently, automatically predicting depression severity through audio signals has garnered increasing attention. These methods typically depend on handcrafted features or deep features achieved by deep learning. Given the complementarity between handcrafted and deep features, their combination enables a more effective capture of emotional sensitivity associated with depression. Motivated by this observation, this work proposes an automated attention-based audio depression recognition integrating handcrafted and deep features that leverage the strengths of both handcrafted and deep features to enhance the accuracy of depression detection. The proposed method comprises two primary components: a feature extraction module and an attention-based feature fusion module. Specifically, the feature extraction module extracts both handcrafted LLDs and deep features by using a pre-trained self-supervised model. Moreover, a new strategy of feature selection based on an attention mechanism is designed to adaptively select more effective handcrafted features for downstream tasks. The attention-based feature fusion module employs a cross-attention mechanism to effectively integrate these features, followed by a fully-connected layer to predict the depression scores. Experimental results on the typical AVEC2013 dataset indicate that the proposed method significantly improves the accuracy of audio depression detection analysis over the other methods and demonstrates its substantial potential applications to automatic depression diagnosis in clinical practice.