Exploiting Automatic Source Separation and Music Transcription for Music Emotion Recognition
摘要
We present a set of novel emotionally relevant audio features related to percussion and individual instrument information to help improve the classification of emotions in music. To this end, as intermediate steps, we performed automatic music transcription (using the MT3 framework) and music source separation (using the Demucs framework). Leveraging the outcomes of these two frameworks, we developed a set of novel features that capture information related to musical texture, rhythm, and melody, which are among the least represented musical dimensions in Music Emotion Recognition (MER). To validate our work, we employed the recently created MERGE dataset, which contains over 3000 30-s audio clips annotated in terms of Russell’s emotion quadrants. To assess the impact of the proposed features, we compared the classification results obtained in this dataset with current state-of-the-art features. The conducted experiments show that the novel features improved the music emotion classification results. Moreover, the best-performing approach achieved an F1-score of 74.1% and employed 200 features (after feature selection), of which 62 were novel.