Precisely deciphering human activities from video footage poses a formidable challenge, particularly when confronted with a myriad of confounding factors such as varied body postures, object appearance inconsistencies, occlusions, and inherent disparities within and across activity classes. This study ventures into this intricate domain by harnessing the transformative power of the vision transformer (ViT) model to tackle the demanding task of sports action recognition. Leveraging the UCF sports action dataset as our testing ground, we demonstrate the remarkable capabilities of the ViT model, achieving an astounding training accuracy of 94%. This groundbreaking application of ViT in the realm of computer vision heralds a significant leap forward, showcasing its unparalleled precision and efficiency in deciphering the intricate nuances of sports actions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Vision Transformer in Sports Action: Recognizing Athletic Activities Across Varied Sporting Domains

  • Krunal Maheriya,
  • Bhavya Patel,
  • Mrugendrasinh Rahevar,
  • Martin Parmar

摘要

Precisely deciphering human activities from video footage poses a formidable challenge, particularly when confronted with a myriad of confounding factors such as varied body postures, object appearance inconsistencies, occlusions, and inherent disparities within and across activity classes. This study ventures into this intricate domain by harnessing the transformative power of the vision transformer (ViT) model to tackle the demanding task of sports action recognition. Leveraging the UCF sports action dataset as our testing ground, we demonstrate the remarkable capabilities of the ViT model, achieving an astounding training accuracy of 94%. This groundbreaking application of ViT in the realm of computer vision heralds a significant leap forward, showcasing its unparalleled precision and efficiency in deciphering the intricate nuances of sports actions.