Depression Classification Using Token Merging-Based Speech Spectrotemporal Transformer
摘要
This paper introduces a novel approach for depression classification, utilizing multimodal token merging (ToMe) within a speech spectrotemporal transformer framework. The model’s efficacy is evaluated with log-mel spectrograms and autocorrelation tempograms extracted from depressed and non depressed speech. The results demonstrate the effectiveness of ToMe when integrated with attention mechanisms of audio spectrogram transformer (AST) models, such as AST and data efficient image transformer (DeiT) encoders. This underscores the importance of the token pruning mechanism utilized in the study. Additionally, a multimodal dual-channel architecture is introduced, featuring two distinct feature modalities extracted from speech: spectrograms and autocorrelation tempograms. The novel ToMe dual-channel AST and ToMe dual-channel AST with DeiT encoder models demonstrate remarkable performance on two different datasets, namely the EATD-Corpus (Chinese) and DAIC-WoZ (English), providing promising results for depression detection.