The task of video virtual try-on aims to fit the target clothes to a person in the video with spatio-temporal consistency. While image-based virtual try-on has been achieved impressive progress, the challenge of how to efficiently replicate the success of image try-on in the video domain remains, which includes two primary challenges: 1) how to generate high-fidelity warping clothes to the target person in a video; 2) how to generate clothes and non-target body parts (e. g. arms, neck) in harmony in a dynamic process. To address them, we present ClothAnimate in this work, a novel diffusion-based framework that aims to tackle the task of video virtual try-on. Specifically, we first utilize attention methods from diffusion as backbone and design two attention modules: Region-aware Introverted Attention and Cloth-aware Extroverted Attention enhancing realistic and controllable virtual try-on. Secondly, we develop a video diffusion model to encode temporal information, where a Garment Encoder is introduced to extract clothes features and pass them to diffusion process. Moreover, motion modules are injected to boost temporally consistent. Extensive experiments demonstrate that our method achieves superior satisfactory video try-on results and video fidelity compared to previous methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ClothAnimate: Boosting Video Virtual Try-On via Novel Attention-Control Diffusion Model

  • Xin Zhang,
  • Siting Huang,
  • Yifan Xie,
  • Xiangyang Luo,
  • Tao Feng,
  • Fei Ma,
  • Fei Yu

摘要

The task of video virtual try-on aims to fit the target clothes to a person in the video with spatio-temporal consistency. While image-based virtual try-on has been achieved impressive progress, the challenge of how to efficiently replicate the success of image try-on in the video domain remains, which includes two primary challenges: 1) how to generate high-fidelity warping clothes to the target person in a video; 2) how to generate clothes and non-target body parts (e. g. arms, neck) in harmony in a dynamic process. To address them, we present ClothAnimate in this work, a novel diffusion-based framework that aims to tackle the task of video virtual try-on. Specifically, we first utilize attention methods from diffusion as backbone and design two attention modules: Region-aware Introverted Attention and Cloth-aware Extroverted Attention enhancing realistic and controllable virtual try-on. Secondly, we develop a video diffusion model to encode temporal information, where a Garment Encoder is introduced to extract clothes features and pass them to diffusion process. Moreover, motion modules are injected to boost temporally consistent. Extensive experiments demonstrate that our method achieves superior satisfactory video try-on results and video fidelity compared to previous methods.