<p>With the rapid development of artificial intelligence content generation technology, visual image generation, as one of the core directions, has become an important research hotspot in the field of computer vision and deep learning. However, existing text-to-image generation methods commonly suffer from shallow semantic understanding and limited style expression. Therefore, the study proposes a visual image generation method based on Contrastive Language-Image Pre-training (CLIP) semantic guidance and text inversion. The semantic content and artistic style are collaboratively controlled by integrating the cross-modal semantic understanding capabilities of CLIP and the learnable style embedding of the text inversion mechanism. The experimental results show that the CLIP similarity of the research model in the semantic alignment experiment reached 0.59 ± 0.02, which is significantly improved compared to 0.28 of the traditional Generative Adversarial Network (GAN) and 0.35 of the CLIP model. In the style transfer experiment, the style score of the research model reached 92.80 ± 1.5, and the Fréchet inception distance was only 45.12, with the variance of style consistency also dropping to 0.09, achieving an average improvement of more than 30% compared to traditional methods. In manual evaluation, the preference rate of the research model exceeded all comparison methods, reaching more than 51%. Comprehensive analysis shows that the research model can achieve high fidelity and natural transition of styles while ensuring accurate semantic expression, and the results have both structural rationality and artistic expression.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Visual Image Generation Based on CLIP Semantic Guidance and Text Inversion

  • Yue Ba,
  • Pengfei Jing,
  • Sazrinee Zainal Abidin,
  • Nazlina Shaari,
  • Raja Ahmad Azmeer Raja Ahmad Effendi

摘要

With the rapid development of artificial intelligence content generation technology, visual image generation, as one of the core directions, has become an important research hotspot in the field of computer vision and deep learning. However, existing text-to-image generation methods commonly suffer from shallow semantic understanding and limited style expression. Therefore, the study proposes a visual image generation method based on Contrastive Language-Image Pre-training (CLIP) semantic guidance and text inversion. The semantic content and artistic style are collaboratively controlled by integrating the cross-modal semantic understanding capabilities of CLIP and the learnable style embedding of the text inversion mechanism. The experimental results show that the CLIP similarity of the research model in the semantic alignment experiment reached 0.59 ± 0.02, which is significantly improved compared to 0.28 of the traditional Generative Adversarial Network (GAN) and 0.35 of the CLIP model. In the style transfer experiment, the style score of the research model reached 92.80 ± 1.5, and the Fréchet inception distance was only 45.12, with the variance of style consistency also dropping to 0.09, achieving an average improvement of more than 30% compared to traditional methods. In manual evaluation, the preference rate of the research model exceeded all comparison methods, reaching more than 51%. Comprehensive analysis shows that the research model can achieve high fidelity and natural transition of styles while ensuring accurate semantic expression, and the results have both structural rationality and artistic expression.