<p>Action recognition remains challenging due to background interference, viewpoint variation, and fine-grained visual similarity. Most existing methods rely on single-modal features or simple multi-stream concatenation, limiting the use of pose priors. To address this, we propose PGFiT-Net, a two-stream framework integrating pose-guided multi-layer spatial gating with vector-level FiLM conditioning. Pose-derived heatmaps suppress background noise at multiple stages, while pose vectors produce channel-wise affine parameters for global modulation, forming complementary spatial and channel constraints. Experiments on Stanford40 and UCF-Sports show that PGFiT-Net clearly surpasses single-stream and late-fusion baselines on common evaluation metrics, matching or exceeding recent state-of-the-art methods. Ablation and visualization analyses further verify the effectiveness and complementarity of both modules. PGFiT-Net provides an efficient and scalable solution for single-frame action recognition with pose priors.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PGFiT-Net: a two-stream human action recognition model with pose-gated and FiLM-conditioned fusion

  • Hong Zhang,
  • Bo Yang,
  • Shijin Zhang

摘要

Action recognition remains challenging due to background interference, viewpoint variation, and fine-grained visual similarity. Most existing methods rely on single-modal features or simple multi-stream concatenation, limiting the use of pose priors. To address this, we propose PGFiT-Net, a two-stream framework integrating pose-guided multi-layer spatial gating with vector-level FiLM conditioning. Pose-derived heatmaps suppress background noise at multiple stages, while pose vectors produce channel-wise affine parameters for global modulation, forming complementary spatial and channel constraints. Experiments on Stanford40 and UCF-Sports show that PGFiT-Net clearly surpasses single-stream and late-fusion baselines on common evaluation metrics, matching or exceeding recent state-of-the-art methods. Ablation and visualization analyses further verify the effectiveness and complementarity of both modules. PGFiT-Net provides an efficient and scalable solution for single-frame action recognition with pose priors.