AI-Driven Visual Generation: Generative Adversarial Neural Network (GAN) and Diffusion Models
摘要
Artificial intelligence, particularly in the realm of image generation, has witnessed a surge in prominence in recent decades. This research delves into the theoretical foundations and practical applications of visual generation software, exploring its intersection with art and audio-visual production. By analyzing the evolution of text-to-image and image-to-image generative AI, we aim to provide a comprehensive overview of their strengths, weaknesses, and the salient features of contemporary AI tools. Our focus is on three leading diffusion models: Dall-E, MidJourney, and Stable Diffusion. Through an in-depth examination of their communicative capabilities, we identify similarities and differences, offering a user-friendly guide. We emphasize the potential of these AI tools in enhancing spatial and creative design, while also highlighting the need for a nuanced understanding of their iterative and informed communication dynamics. This analysis includes works using generative adversarial neural networks (GANs) and diffusion models from the 2010s to the present. By examining the evolution of these technologies and the resulting creative outputs, we aim to provide a comprehensive understanding of their capabilities and limitations. Through this exploration, we seek to shed light on the evolving landscape of AI-driven visual generation and to identify emerging trends and best practices for creative professionals.