<p>Text-to-Thangka generation requires preserving both semantic accuracy and textural details. Current methods struggle with fine-grained feature extraction, multi-level feature integration, and discriminator overfitting due to limited Thangka data. We present HST-GAN, a novel framework combining parallel hybrid attention with differentiable symmetric augmentation. The architecture features a Parallel Spatial-Channel Attention module (PSCA) for precise localization of deity facial features and ritual object textures, along with a Hierarchical Feature Fusion Network (HLFN) for multi-scale alignment. The framework’s Differentiable Symmetric Augmentation (DiffAugment) dynamically adjusts discriminator inputs to prevent overfitting while improving generalization. On the T2IThangka dataset, HST-GAN achieves an Inception Score of 2.08 and reduces Fréchet Inception Distance to 87.91, demonstrating superior performance over baselines on the Oxford-102 benchmark.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hierarchical symmetric GAN for Thangka image generation

  • Wenjin Hu,
  • Yan Zhao,
  • Lemei Yin,
  • Guoquan Zhang

摘要

Text-to-Thangka generation requires preserving both semantic accuracy and textural details. Current methods struggle with fine-grained feature extraction, multi-level feature integration, and discriminator overfitting due to limited Thangka data. We present HST-GAN, a novel framework combining parallel hybrid attention with differentiable symmetric augmentation. The architecture features a Parallel Spatial-Channel Attention module (PSCA) for precise localization of deity facial features and ritual object textures, along with a Hierarchical Feature Fusion Network (HLFN) for multi-scale alignment. The framework’s Differentiable Symmetric Augmentation (DiffAugment) dynamically adjusts discriminator inputs to prevent overfitting while improving generalization. On the T2IThangka dataset, HST-GAN achieves an Inception Score of 2.08 and reduces Fréchet Inception Distance to 87.91, demonstrating superior performance over baselines on the Oxford-102 benchmark.