A Comprehensive Exploration of Network-Based Approaches for Singing Voice Separation
摘要
The process of extracting vocals and music accompaniment in the song is termed as singing voice separation (SVS). SVS task enables remixing, karaoke, and music transcription. Signal, salience, and network are the approaches to extract vocals from mixed audio. In recent years, network-based approaches have shown remarkable success in isolating the vocal components of a mixed audio. In the network-based approaches, there are various methodologies to give an efficient result. This paper briefly explains the performance of DenseNet, including MMDenseNet and MMDenseLSTM. Optimized HRNet using strategies like high-resolution long short-term memory (HR-LSTM) and multi-resolution networks. Subsequently, surveys the functionalities of U-Net like Dense U-Net and Wave U-Net. Also, the behavior of Y-Net, ResNet approaches, and time-domain audio separation network (TasNet) like VAT-SNet and Conv-TasNet, and the process of MultiResUNet and the performance of Gated Recurrent Unit (GRU) in RNN are also disclosed. This paper offers a roadmap for fellow researchers in the field of singing voice separation, assisting them in making well-informed decisions regarding the selection and adaptation of methodologies tailored to their specific application.