<p>Human-object interaction (HOI) information provides cues or constraints for spatio-temporal action detection (STAD). Existing STAD models that incorporate HOI relationships utilize heavyweight 3D backbones to extract temporal information and require matching many region proposals during the HOI feature generation stage, which results in significant computational overhead. To address these limitations, we propose a multi-level spatio-temporal and interaction feature fusion (MLSTIF) network, whose 2D and 3D backbones are implemented via You Only Look Once (YOLO)v7 and 3D-EfficientNetv2, respectively, to extract multi-level spatial and spatio-temporal features. To exploit the spatio-temporal information within videos, a spatial and spatio-temporal feature fusion (SSTFF) module is designed to fuse motion information acquired from multi-level spatio-temporal features into spatial features at corresponding levels, and to more effectively extract moving target-related interaction content from videos, a decoupled deformable HOI transformer (Decoupled Deformable HOITR) module is designed to produce higher-order classification and regression interaction features, which improves the accuracy of the spatio-temporal locations and action categories of moving targets. Compared with that of all similar methods used in the experiments, the MLSTIF improves the accuracy by 1.4-11.3% on the UCF101-24 dataset and by 0.9-11.4% on the J-HMDB dataset. Among similar methods, some models have at least three times the computational cost of MLSTIF. Compared with that of the models with comparable computational costs, such as the YOWO series and YWOM, the accuracy of MLSTIF is improved by 1.4-7.2% on the AVA v2.2 dataset.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MLSTIF: multi-level spatio-temporal and human-object interaction feature fusion network for spatio-temporal action detection

  • Rui Yang,
  • Hui Zhang,
  • Mulan Qiu,
  • Min Wang

摘要

Human-object interaction (HOI) information provides cues or constraints for spatio-temporal action detection (STAD). Existing STAD models that incorporate HOI relationships utilize heavyweight 3D backbones to extract temporal information and require matching many region proposals during the HOI feature generation stage, which results in significant computational overhead. To address these limitations, we propose a multi-level spatio-temporal and interaction feature fusion (MLSTIF) network, whose 2D and 3D backbones are implemented via You Only Look Once (YOLO)v7 and 3D-EfficientNetv2, respectively, to extract multi-level spatial and spatio-temporal features. To exploit the spatio-temporal information within videos, a spatial and spatio-temporal feature fusion (SSTFF) module is designed to fuse motion information acquired from multi-level spatio-temporal features into spatial features at corresponding levels, and to more effectively extract moving target-related interaction content from videos, a decoupled deformable HOI transformer (Decoupled Deformable HOITR) module is designed to produce higher-order classification and regression interaction features, which improves the accuracy of the spatio-temporal locations and action categories of moving targets. Compared with that of all similar methods used in the experiments, the MLSTIF improves the accuracy by 1.4-11.3% on the UCF101-24 dataset and by 0.9-11.4% on the J-HMDB dataset. Among similar methods, some models have at least three times the computational cost of MLSTIF. Compared with that of the models with comparable computational costs, such as the YOWO series and YWOM, the accuracy of MLSTIF is improved by 1.4-7.2% on the AVA v2.2 dataset.