<p>Gesture recognition, an intuitive and efficient human–robot interaction method, shows great potential in smart living applications. However, deployment in home environments faces challenges from complex backgrounds, hand-like distractors, and varying illumination, particularly for lightweight and real-time implementations. To address these limitations in balancing efficiency, speed, and accuracy, a novel lightweight static gesture recognition network, DCC-DETR, based on RT-DETR, is proposed. First, an improved StarNet backbone is developed to enhance feature extraction efficiency, optimizing performance without compromising accuracy. Second, a Cascaded Group Local Attention Network (CGLAN) is designed to improve the perception and processing of local information significantly. Third, a Context-Guided Spatial Reconstruction Feature Pyramid Network (CSRFPN) enhances multi-scale feature representation and robust feature integration. Finally, the Wise Focaler-ShapeIoU loss function enables robust bounding box regression, leading to improved localization precision. Extensive experiments on self-built, 100 Days of Hands, and EgoHands datasets demonstrate DCC-DETR’s superiority in detection accuracy, inference speed, and model efficiency, highlighting its practical utility across diverse scenarios. Compared to RT-DETR, DCC-DETR improves mAP50 by 1.0% and increases FPS by 105.8%, while reducing computation by 80.0%, parameters by 63.2%, and significantly compressing model size, making it highly suitable for resource-limited environments. When deployed on NVIDIA Jetson AGX Orin with TensorRT acceleration, DCC-DETR achieves an end-to-end GPU inference latency of 6.63 ms, 1.62<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11554_2025_1763_Article_IEq1.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation> faster than RT-DETR (10.71 ms) under identical conditions, demonstrating superior real-time capability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DCC-DETR: a real-time lightweight gesture recognition network for home human–robot interaction

  • Xianyi Chen,
  • Hao Yin

摘要

Gesture recognition, an intuitive and efficient human–robot interaction method, shows great potential in smart living applications. However, deployment in home environments faces challenges from complex backgrounds, hand-like distractors, and varying illumination, particularly for lightweight and real-time implementations. To address these limitations in balancing efficiency, speed, and accuracy, a novel lightweight static gesture recognition network, DCC-DETR, based on RT-DETR, is proposed. First, an improved StarNet backbone is developed to enhance feature extraction efficiency, optimizing performance without compromising accuracy. Second, a Cascaded Group Local Attention Network (CGLAN) is designed to improve the perception and processing of local information significantly. Third, a Context-Guided Spatial Reconstruction Feature Pyramid Network (CSRFPN) enhances multi-scale feature representation and robust feature integration. Finally, the Wise Focaler-ShapeIoU loss function enables robust bounding box regression, leading to improved localization precision. Extensive experiments on self-built, 100 Days of Hands, and EgoHands datasets demonstrate DCC-DETR’s superiority in detection accuracy, inference speed, and model efficiency, highlighting its practical utility across diverse scenarios. Compared to RT-DETR, DCC-DETR improves mAP50 by 1.0% and increases FPS by 105.8%, while reducing computation by 80.0%, parameters by 63.2%, and significantly compressing model size, making it highly suitable for resource-limited environments. When deployed on NVIDIA Jetson AGX Orin with TensorRT acceleration, DCC-DETR achieves an end-to-end GPU inference latency of 6.63 ms, 1.62 \(\times\) × faster than RT-DETR (10.71 ms) under identical conditions, demonstrating superior real-time capability.