<p>Effective gesture recognition in Virtual Reality (VR) and Augmented Reality (AR) faces significant challenges from varying hand angles, postures, and complex backgrounds, limiting real-time application potential. This paper proposes the Dynamic Feature Pyramid Network (DFPN), a novel gesture recognition architecture that combines adaptive dynamic convolutional kernels with multi-scale feature fusion capabilities. We provide theoretical guarantees for our approach through generalization error analysis, demonstrating that our model’s error bound converges to zero with sufficient training data. Extensive experiments show that DFPN achieves 84.88% accuracy in complex gesture recognition scenarios involving multi-angle rotation, while maintaining an inference time of approximately 20&#xa0;ms per sample on a Tesla V100 GPU. This favorable accuracy-efficiency trade-off makes DFPN particularly suitable for resource-constrained environments and real-time applications where response latency is critical. The DFPN model demonstrates superior gradient flow dynamics during training, with loss values converging substantially faster than baseline approaches. Our contribution advances gesture-based human-computer interaction technology for VR/AR systems and establishes both theoretical and practical foundations for future developments in efficient, responsive contactless interaction systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dynamic feature pyramid network for real-time gesture recognition

  • Shiming Ma,
  • Quansheng Liu

摘要

Effective gesture recognition in Virtual Reality (VR) and Augmented Reality (AR) faces significant challenges from varying hand angles, postures, and complex backgrounds, limiting real-time application potential. This paper proposes the Dynamic Feature Pyramid Network (DFPN), a novel gesture recognition architecture that combines adaptive dynamic convolutional kernels with multi-scale feature fusion capabilities. We provide theoretical guarantees for our approach through generalization error analysis, demonstrating that our model’s error bound converges to zero with sufficient training data. Extensive experiments show that DFPN achieves 84.88% accuracy in complex gesture recognition scenarios involving multi-angle rotation, while maintaining an inference time of approximately 20 ms per sample on a Tesla V100 GPU. This favorable accuracy-efficiency trade-off makes DFPN particularly suitable for resource-constrained environments and real-time applications where response latency is critical. The DFPN model demonstrates superior gradient flow dynamics during training, with loss values converging substantially faster than baseline approaches. Our contribution advances gesture-based human-computer interaction technology for VR/AR systems and establishes both theoretical and practical foundations for future developments in efficient, responsive contactless interaction systems.