<p>This study proposes a novel Multi Curvelet Transformer Network (MCTN) for fine-grained human behavior recognition in dynamic video scenarios. A key challenge in this field lies in accurately identifying human actions under adverse conditions such as motion blur, occlusion, and varying illumination. To address this, we introduce a motion blur restoration module leveraging the curvelet transform to enhance motion image clarity, thereby improving downstream behavior detection. Furthermore, we enhance the Transformer architecture by embedding curvelet-based multi-scale attention mechanisms, which significantly improve the model’s ability to extract spatial-temporal features at different resolutions. The proposed network also adopts a multi-curvelet transform structure to deepen semantic representation. Experimental results on benchmark datasets, including an action recognition dataset and the MSCOCO dataset, demonstrate that MCTN achieves superior performance, reaching a mean average precision (mAP) of 0.822. These results underscore the potential of MCTN in real-time intelligent video analysis and human-computer interaction applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Curvelet-enhanced transformer architecture for blurred action fine-grained detection

  • Yuxiang Ren,
  • Zhetao Guo,
  • Wei Zhang,
  • Yushi Shen,
  • Ying Xing

摘要

This study proposes a novel Multi Curvelet Transformer Network (MCTN) for fine-grained human behavior recognition in dynamic video scenarios. A key challenge in this field lies in accurately identifying human actions under adverse conditions such as motion blur, occlusion, and varying illumination. To address this, we introduce a motion blur restoration module leveraging the curvelet transform to enhance motion image clarity, thereby improving downstream behavior detection. Furthermore, we enhance the Transformer architecture by embedding curvelet-based multi-scale attention mechanisms, which significantly improve the model’s ability to extract spatial-temporal features at different resolutions. The proposed network also adopts a multi-curvelet transform structure to deepen semantic representation. Experimental results on benchmark datasets, including an action recognition dataset and the MSCOCO dataset, demonstrate that MCTN achieves superior performance, reaching a mean average precision (mAP) of 0.822. These results underscore the potential of MCTN in real-time intelligent video analysis and human-computer interaction applications.