Interpretable Machine Learning Techniques for Audio Source Separation
摘要
Recent advancements in deep learning have revolutionized audio source separation, addressing challenges like the cocktail party problem. This study explores the efficacy of Convolutional Neural Networks (CNNs), specifically the Conv-TasNet architecture, for separating vocals, drums, bass, and other instruments from mixed audio using spectrogram-based methods, and deep learning models. Leveraging the MUSDB18 dataset, these models were trained and evaluated, and the best model achieved significant performance metrics such as an accuracy of 0.93, Signal-to-Distortion Ratio (SDR) of 13.2 dB, Signal-to-Interference Ratio (SIR) of 14.6 dB, and Signal-to-Artifact Ratio (SAR) of 10.5 dB. Additionally, explainability techniques like SHAP and Integrated Gradient were employed to enhance model interpretability. These findings contribute to advancing more efficient and interpretable audio processing systems.