Multimodal Content Generation
摘要
In this chapter, we will review the advances that are being made in this new field of multimodal content generation and also discuss several challenges associated with this emerging technology. First, we will understand the machine learning techniques that drive this technology–most notably, the concept of adversarial learning and diffusion modeling. Then we will learn about how these techniques are applied to several input-to-output mappings, most notably, text-to-image generation, and the current state-of-the-art in image generation under these various input-to-output settings. Finally, we will discuss challenges in the evaluation and benchmarking of various dimensions of multimodal content generation, as well as the risks posed by malicious use of such technology.