Respiratory Sound Classification via Multi-view Feature Fusion with Enhanced Convolutional Neural Network and Audio Spectrogram Transformer
摘要
Respiratory sound recognition is a challenging task. This paper presents an automatic classification model that combines convolutional neural network and audio spectrogram transformer. It utilizes respiratory sound recordings from 418 patients with respiratory diseases. Each respiratory cycle is transformed into a Mel-spectrogram via Fourier transform, which serves as input to the model. Our methodology implements a dual-perspective parallel extraction approach that identifies comprehensive and specific acoustic patterns. These representations merge through contrastive learning techniques and an adaptive fusion mechanism to enhance diagnostic precision. Performance evaluation reveals the framework attains accuracy rates of 69.50% and 92.01% in binary classification (normal/abnormal) on ICBHI 2017 and SPRSound datasets. For multi-class categorization, it reaches 63.27% and 90.63% respectively, surpassing contemporary methodologies. These findings validate the efficacy of our integrated approach for computerized respiratory acoustic analysis.