CAF-VTON: cross-attention layered fusion based latent diffusion virtual try-On
摘要
Virtual try-on technology has gained significant attention in e-commerce, digital retail, and virtual reality due to its ability to enhance user experience and reduce return rates. However, generating accurate and natural virtual try-on results remains challenging, especially when dealing with complex human poses and clothing deformations. In this paper, we propose CAF-VTON, a novel latent diffusion-based virtual try-on network that introduces cross-attention layered fusion for the first time in virtual try-on tasks. By leveraging a layered cross-attention mechanism, CAF-VTON can progressively extract both local and global features of human poses and clothing, enabling the capture of fine details essential for realistic virtual try-on. Research shows that on high-resolution datasets, CAF-VTON outperforms existing methods in most scenarios. It has achieved good results in terms of realism, detail fidelity, and pose consistency. Our research provides new ideas for optimizing virtual try-on solutions and has certain application potential in the fashion and retail industries.