Transformers in Computer Vision
摘要
This chapter delves into the application of transformers in computer vision, focusing on the Vision Transformer (ViT) architecture. It explores the mathematical formulation of ViT, its design principles, and its implementation in tasks such as image classification, object detection, and image generation. Advanced analysis highlights the role of transformers in achieving state-of-the-art performance in visual tasks.