Image-based virtual try-on aims to transfer a given garment onto a specific person seamlessly. Existing methods utilize a global appearance flow to adjust the spatial distribution of garments to the desired style. However, when receiving challenging human poses or garment textures, existing methods are limited in encoding the local context of each part and in anisotropic modeling. Moreover, these methods fail to explore semantic differences within the global context, so they are still insufficient in maintaining consistency between warped garments and preserved person regions. To address these challenges, we propose a novel virtual try-on framework geared toward Immersive Style Outfitting, termed ISO-VTON. Specifically, we first propose a StyleGAN-based fusion block, which encodes the local context of garment decoupled parts via local style vectors and then guides components to inject global style cues during assembly, thereby achieving fine-grained control over the warped components. Second, we develop a Dual Cross-Attention block to investigate the longer-range semantic correspondence between the garment and the person in latent space. This block enhances key regions of the input data to preserve garment details, achieving a natural integration of garment and body shape. Experiments on the VITON-HD benchmark show that our method outperforms state-of-the-art virtual try-on methods in both qualitative and quantitative evaluations. Code is available at https://github.com/shibashijiu/ISO-VTON .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ISO-VTON: Fine-Grained Style-Local Flows with Dual Cross-Attention for Immersive Outfitting

  • Yuliu Guo,
  • Chao Fang,
  • Zhaojing Wang,
  • Li Li

摘要

Image-based virtual try-on aims to transfer a given garment onto a specific person seamlessly. Existing methods utilize a global appearance flow to adjust the spatial distribution of garments to the desired style. However, when receiving challenging human poses or garment textures, existing methods are limited in encoding the local context of each part and in anisotropic modeling. Moreover, these methods fail to explore semantic differences within the global context, so they are still insufficient in maintaining consistency between warped garments and preserved person regions. To address these challenges, we propose a novel virtual try-on framework geared toward Immersive Style Outfitting, termed ISO-VTON. Specifically, we first propose a StyleGAN-based fusion block, which encodes the local context of garment decoupled parts via local style vectors and then guides components to inject global style cues during assembly, thereby achieving fine-grained control over the warped components. Second, we develop a Dual Cross-Attention block to investigate the longer-range semantic correspondence between the garment and the person in latent space. This block enhances key regions of the input data to preserve garment details, achieving a natural integration of garment and body shape. Experiments on the VITON-HD benchmark show that our method outperforms state-of-the-art virtual try-on methods in both qualitative and quantitative evaluations. Code is available at https://github.com/shibashijiu/ISO-VTON .