Computer Auditory Imagination: A Multi-model Approach for Audio to Video Synthesis
摘要
The creation of synchronized videos for audio content, particularly when visual elements are essential for enhancing ambient or abstract audio experiences, poses formidable challenges for content creators. Existing audio-to-video generation systems rely heavily on the presence of speech or lyrics in the audio, limiting their applicability to a broader range of audio types. To address this limitation, this work aims to develop a multi-model approach capable of comprehending non-linguistic audio, such as sound effects or abstract sounds, and generating visually compelling videos in response. The objective of the work is to seamlessly synchronize audio and video components. To achieve this, the focus revolves around the integration of audio-to-text and text-to-image/video models, facilitating the seamless transformation of audio cues into captivating visual representations. This innovative approach holds great promise for expanding the creative possibilities in the realm of audio-visual content generation, offering innovative solutions and opening up new avenues for artistic expression.