<p>Human-Object Interaction (HOI) detection, which aims to simultaneously localize humans and objects while identifying their interactive relationships, has conventionally employed dual-branch Transformer architectures that decouple human-object pair detection from interaction classification. However, such designs frequently suffer from inadequate cross-branch contextual interactions, thereby constraining their relational reasoning capabilities. Moreover, the limited quantity of annotated human-object pairs in existing datasets further restricts the set prediction performance of Transformer-based models. To address these challenges, propose an innovative Multi-Relation Three-Branch Fusion Network (MRTBF) that leverages unary (individual entities), binary (pair-wise), and ternary (interaction-aware) features extracted from human, object, and interaction representations. Through our novel Multi-Relation Cross-Fusion (MRCF) module, the network achieves comprehensive contextual interactions among three decoder branches, substantially enhancing relational reasoning. Furthermore, we present an optimized data augmentation technique called Similar Relation Picture Stitching (SRPS), which synthesizes images with semantically coherent backgrounds during training. This approach not only increases the quantity of annotated human-object pairs but also enhances the visual diversity of training samples, leading to improved model performance. We have compared our proposed MRTBF with existing state-of-the-art methods on two public benchmarks, including V-COCO and HICO-DET. The results have showed that MRTBF out performs the existing best-performing methods on both the above two benchmarks, validating its efficacy in detecting human-object interactions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MRTBF-Net: A Multi-Relation Three-Branch Fusion Transformer with Cross-Modal Context for Human-Object Interaction Detection

  • Haifeng Sang,
  • Huabao Li

摘要

Human-Object Interaction (HOI) detection, which aims to simultaneously localize humans and objects while identifying their interactive relationships, has conventionally employed dual-branch Transformer architectures that decouple human-object pair detection from interaction classification. However, such designs frequently suffer from inadequate cross-branch contextual interactions, thereby constraining their relational reasoning capabilities. Moreover, the limited quantity of annotated human-object pairs in existing datasets further restricts the set prediction performance of Transformer-based models. To address these challenges, propose an innovative Multi-Relation Three-Branch Fusion Network (MRTBF) that leverages unary (individual entities), binary (pair-wise), and ternary (interaction-aware) features extracted from human, object, and interaction representations. Through our novel Multi-Relation Cross-Fusion (MRCF) module, the network achieves comprehensive contextual interactions among three decoder branches, substantially enhancing relational reasoning. Furthermore, we present an optimized data augmentation technique called Similar Relation Picture Stitching (SRPS), which synthesizes images with semantically coherent backgrounds during training. This approach not only increases the quantity of annotated human-object pairs but also enhances the visual diversity of training samples, leading to improved model performance. We have compared our proposed MRTBF with existing state-of-the-art methods on two public benchmarks, including V-COCO and HICO-DET. The results have showed that MRTBF out performs the existing best-performing methods on both the above two benchmarks, validating its efficacy in detecting human-object interactions.