Text-to-Image Synthesis: A Comprehensive Review and Analysis
摘要
The primary objective of developing an automated model that can accurately recognise keywords by analysing their visual characteristics and generate the related visual content. The process of converting text into images has, in recent years, attracted significant from both researchers and practitioners owing to its extensive and varied uses in several industries. In this field, the primary objective is to generate visual representations that align with the provided verbal explanation in terms of logical coherence and conceptual authenticity. It continues to encounter various obstacles, mostly concerning the degree of visual authenticity and semantic coherence. In order to tackle these difficulties, several strategies have been suggested, primarily depending on GANs, VAE and others to achieve the best possible results. This paper investigates the existing methodologies and assesses the frequently employed datasets for synthesising T2I, as well as evaluation metrics utilised in this specific field.