DiffTimb: Diffusion Models for Many-to-Many Timbre Transfer
摘要
The goal of the timbre transfer task is to apply the auditory characteristics of one instrument to another instrument. Past timbre transfer methods faced numerous challenges due to the limitations of the selected generation model. These challenges include difficulties in training, ambiguous results, poor diversity, and slow generation speed. To address these issues, we propose a diffusion model for timbre transfer, DiffTimb, which can be combined with vocoder to achieve single-note timbre transfer. Unlike traditional models that require training a specific model for each pair of timbre transfers, DiffTimb enables the conversion of various musical instrument timbres received in a single model, namely, many-to-many timbre transfer. We compare two Unet-based denoising frameworks and propose two feature fusion methods to fuse timbre and pitch features. We conduct unconditional and conditional generation pre-experiments to verify the generation ability of DiffTimb. To improve the generation effect, we clean the Nsynth dataset and get two subsets for our task. Experimental results show that DiffTimb has outperformed previous methods in both subjective and objective evaluations.