Creating content with specified identities (ID) has attracted significant interest in the field of image generative models. However, its extension to video generation is not well explored. In this work, we propose a simple yet effective subject identity controllable video generation framework, termed Video Custom Diffusion (VCD). With a specified identity defined by a few images, VCD reinforces the identity information extraction, and injects frame-wise correlation for stable video outputs. We propose three novel components: 1) an ID module to extract ID features; 2) a 3D Gaussian Noise Prior for better inter-frame consistency; and 3) Face VCD and Tiled VCD modules to upscale the video with detailed characteristics. We conducted extensive experiments to show that VCD is able to generate stable and high-quality videos with specific human subjects animated in diverse scenes and motions. We further show that VCD could integrate conditional input and prompt travel to enable more delicate controls.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Magic-Me: Identity-Specific Video Customized Diffusion

  • Ze Ma,
  • Daquan Zhou,
  • Xue-She Wang,
  • Chun-Hsiao Yeh,
  • Xiuyu Li,
  • Huanrui Yang,
  • Zhen Dong,
  • Kurt Keutzer,
  • Jiashi Feng

摘要

Creating content with specified identities (ID) has attracted significant interest in the field of image generative models. However, its extension to video generation is not well explored. In this work, we propose a simple yet effective subject identity controllable video generation framework, termed Video Custom Diffusion (VCD). With a specified identity defined by a few images, VCD reinforces the identity information extraction, and injects frame-wise correlation for stable video outputs. We propose three novel components: 1) an ID module to extract ID features; 2) a 3D Gaussian Noise Prior for better inter-frame consistency; and 3) Face VCD and Tiled VCD modules to upscale the video with detailed characteristics. We conducted extensive experiments to show that VCD is able to generate stable and high-quality videos with specific human subjects animated in diverse scenes and motions. We further show that VCD could integrate conditional input and prompt travel to enable more delicate controls.