<p>This survey explores the evolution and application of U-Net in generative AI, highlighting its success across various modalities, including image, text, audio, video, 3D, and pose/action generation. Initially designed for biomedical segmentation, U-Net has been adapted and enhanced with architectural innovations such as normalization techniques, self and cross-attention mechanisms, and residual connections. These advancements have made U-Net a powerful backbone for modern generative models in diffusion-based frameworks, GANs, and autoregressive architectures. The survey comprehensively reviews U-Net’s modality-specific applications, from high-resolution image synthesis and text-to-image generation to speech enhancement, video generation, 3D reconstruction, and pose/action generation. Despite its widespread success, U-Net faces challenges in computational efficiency, contextual understanding, and scalability for multimodal tasks. Future directions focus on optimizing U-Net for lightweight and real-time applications, enhancing its contextual awareness, and improving its integration with emerging architectures like transformers and diffusion models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Revisiting U-Net: a foundational backbone for modern generative AI

  • Marvin John Ignacio,
  • Sangyun Shin,
  • Hulin Jin,
  • Seong Joon Yoo,
  • Dongil Han,
  • Yong-Guk Kim

摘要

This survey explores the evolution and application of U-Net in generative AI, highlighting its success across various modalities, including image, text, audio, video, 3D, and pose/action generation. Initially designed for biomedical segmentation, U-Net has been adapted and enhanced with architectural innovations such as normalization techniques, self and cross-attention mechanisms, and residual connections. These advancements have made U-Net a powerful backbone for modern generative models in diffusion-based frameworks, GANs, and autoregressive architectures. The survey comprehensively reviews U-Net’s modality-specific applications, from high-resolution image synthesis and text-to-image generation to speech enhancement, video generation, 3D reconstruction, and pose/action generation. Despite its widespread success, U-Net faces challenges in computational efficiency, contextual understanding, and scalability for multimodal tasks. Future directions focus on optimizing U-Net for lightweight and real-time applications, enhancing its contextual awareness, and improving its integration with emerging architectures like transformers and diffusion models.