<p>With the rapid development of the smart sports industry, the application of badminton robots in competitive training, event assistance, and other scenarios has become increasingly widespread. However, the imbalance between the real-time demands of high-speed motion scenarios and the model’s computational efficiency and feature expression capability remains a bottleneck that restricts its development. To address this, this paper focuses on the research of balancing speed and accuracy in real-time badminton detection. For the mainstream badminton detection model TrackNetV2, which has frame rate limitations (27.29 FPS) and computational redundancy, we propose a lightweight badminton detection network, EGHT, which is both accurate and fast. EGHT implemented using our proposed EGhostNet (a GhostNet backbone with an embedded ECA attention module) as the backbone network, combined with a channel reduction strategy, and the introduction of a lightweight spatiotemporal fusion module (LiteSTF). EGhostNet enhances the feature focusing ability on high-speed moving objects through an embedded attention mechanism (ECA), while the LiteSTF module uses a separable 3D convolution structure to efficiently model spatiotemporal features. In addition, to address the issue of localization errors in practical applications, a novel loss function is designed that reduces the localization errors rate (LER) by 36.94% at the cost of a certain detection accuracy loss. Experimental results show that EGHT achieves a detection speed of 56.75 FPS on the NVIDIA GTX 1070 platform (the theoretical detection frame rate is 170.25 FPS), a 106.8% improvement over TrackNetV2, while reducing the parameter count to 9.11% of the original model (1.01M). The detection accuracy reaches 86.84% (+0.24%) and the precision is 98.13% (+0.78%), significantly improving the feature expression capability and inference efficiency in high-speed motion scenarios, effectively overcoming the balance problem between computational efficiency and detection accuracy in traditional models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

EGHT: a lightweight spatiotemporal fusion and dynamic attention model for real-time badminton detection

  • Yunkang Nie,
  • Gang Li,
  • Yongqiang Fan,
  • Fengyi Wang

摘要

With the rapid development of the smart sports industry, the application of badminton robots in competitive training, event assistance, and other scenarios has become increasingly widespread. However, the imbalance between the real-time demands of high-speed motion scenarios and the model’s computational efficiency and feature expression capability remains a bottleneck that restricts its development. To address this, this paper focuses on the research of balancing speed and accuracy in real-time badminton detection. For the mainstream badminton detection model TrackNetV2, which has frame rate limitations (27.29 FPS) and computational redundancy, we propose a lightweight badminton detection network, EGHT, which is both accurate and fast. EGHT implemented using our proposed EGhostNet (a GhostNet backbone with an embedded ECA attention module) as the backbone network, combined with a channel reduction strategy, and the introduction of a lightweight spatiotemporal fusion module (LiteSTF). EGhostNet enhances the feature focusing ability on high-speed moving objects through an embedded attention mechanism (ECA), while the LiteSTF module uses a separable 3D convolution structure to efficiently model spatiotemporal features. In addition, to address the issue of localization errors in practical applications, a novel loss function is designed that reduces the localization errors rate (LER) by 36.94% at the cost of a certain detection accuracy loss. Experimental results show that EGHT achieves a detection speed of 56.75 FPS on the NVIDIA GTX 1070 platform (the theoretical detection frame rate is 170.25 FPS), a 106.8% improvement over TrackNetV2, while reducing the parameter count to 9.11% of the original model (1.01M). The detection accuracy reaches 86.84% (+0.24%) and the precision is 98.13% (+0.78%), significantly improving the feature expression capability and inference efficiency in high-speed motion scenarios, effectively overcoming the balance problem between computational efficiency and detection accuracy in traditional models.