Audio Source Separation: Advances and Challenges
摘要
One of the most crucial steps in processing audio signals is called audio source separation (ASS). The goal of the ASS is to separate the different parts of a mixed audio signal into their own signals. This job is hard because there may be sound sources that overlap in both time and frequency domains. Conventional ASS approaches, such as non-negative matrix factorization (NMF) and Blind Source Separation (BSS) have not been able to separate signals better than other methods. Recently, new, advanced ways have come out that show promise for ASS. Although they have their limitations, deep learning models have emerged because of their ability to detect the origin of a sound based on the statistical features of the audio data. This paper gives an overview of the different ASS techniques, discussing the different approaches used in this field and the performance evaluation metrics used to measure how well they work.