In this chapter, we will review the advances that are being made in this new field of multimodal content generation and also discuss several challenges associated with this emerging technology. First, we will understand the machine learning techniques that drive this technology–most notably, the concept of adversarial learning and diffusion modeling. Then we will learn about how these techniques are applied to several input-to-output mappings, most notably, text-to-image generation, and the current state-of-the-art in image generation under these various input-to-output settings. Finally, we will discuss challenges in the evaluation and benchmarking of various dimensions of multimodal content generation, as well as the risks posed by malicious use of such technology.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Content Generation

  • Man Luo,
  • Tejas Gokhale,
  • Neeraj Varshney,
  • Yezhou Yang,
  • Chitta Baral

摘要

In this chapter, we will review the advances that are being made in this new field of multimodal content generation and also discuss several challenges associated with this emerging technology. First, we will understand the machine learning techniques that drive this technology–most notably, the concept of adversarial learning and diffusion modeling. Then we will learn about how these techniques are applied to several input-to-output mappings, most notably, text-to-image generation, and the current state-of-the-art in image generation under these various input-to-output settings. Finally, we will discuss challenges in the evaluation and benchmarking of various dimensions of multimodal content generation, as well as the risks posed by malicious use of such technology.