A sketch-and-text guided diffusion framework for Tibetan painting generation for digital heritage preservation
摘要
Tibetan painting is an important form of visual cultural heritage, characterized by complex iconographic structures, dense decorative patterns, and distinctive stylistic conventions. Generating Tibetan painting images from sparse sketches and textual descriptions remains challenging because sparse sketches provide incomplete structural cues, while culturally plausible synthesis requires semantic consistency, structural coherence, and style-aware representation. To address these challenges, we propose STP-Diff, a sketch- and text-guided diffusion framework for Tibetan painting generation. STP-Diff integrates sketch semantic parsing, sparse-to-dense structure enhancement, multi-condition semantic fusion, and line-conditioned style prior learning to guide diffusion with structural and semantic constraints. We further construct an extended HHTP dataset with paired sketch, line drawing, color image, and text annotations. Experiments on HHTP and SketchyCOCO demonstrate that STP-Diff improves structural preservation, semantic alignment, visual plausibility, and generalization, providing a computational approach for digital preservation and creative reuse of Tibetan painting heritage.