In recent years, with the rapid development of AIGC technologies, significant progress has been made in the field of face photo-sketch synthesis, which plays a crucial role in law enforcement and digital entertainment. However, existing research tends to focus solely on facial information while neglecting accompanying audio information, potentially leading to the omission of crucial information about the identity. Due to the modality gap between face photos and sketches, directly applying existing audio-driven video generation approaches usually yeild poor performance. To this end, we propose a novel method for audio-driven face photo-sketch video generation. Our method integrates sketch portrait generation, audio feature extraction, joint optimization of expression and pose networks, and 3D facial rendering, which implements realistic facial expressions and head poses generation sensitive to audio in sketch style. To enhance the naturalness and clarity of the generated face photo-sketch videos, we further design a sketch portrait embedding method that optimally integrates face photo-sketch synthesis into a conventional audio-driven model for sketch video generation. Extensive experiments show that our method outperforms existing methods in both qualitative and quantitative evaluations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Audio-Driven Face Photo-Sketch Video Generation

  • Siyue Zhou,
  • Qun Guan,
  • Chunlei Peng,
  • Decheng Liu,
  • Yu Zheng

摘要

In recent years, with the rapid development of AIGC technologies, significant progress has been made in the field of face photo-sketch synthesis, which plays a crucial role in law enforcement and digital entertainment. However, existing research tends to focus solely on facial information while neglecting accompanying audio information, potentially leading to the omission of crucial information about the identity. Due to the modality gap between face photos and sketches, directly applying existing audio-driven video generation approaches usually yeild poor performance. To this end, we propose a novel method for audio-driven face photo-sketch video generation. Our method integrates sketch portrait generation, audio feature extraction, joint optimization of expression and pose networks, and 3D facial rendering, which implements realistic facial expressions and head poses generation sensitive to audio in sketch style. To enhance the naturalness and clarity of the generated face photo-sketch videos, we further design a sketch portrait embedding method that optimally integrates face photo-sketch synthesis into a conventional audio-driven model for sketch video generation. Extensive experiments show that our method outperforms existing methods in both qualitative and quantitative evaluations.