<p>In human–object interaction (HOI) detection task, ensuring that interactive pairs receive higher attention weights while reducing the weight of non-interaction pairs is imperative for enhancing HOI detection accuracy. Guiding attention learning is also a key aspect of existing transformer-based algorithms. To tackle this challenge, this study proposes a novel approach termed Interaction Confidence Score Learning Attention (ICSLA), which introduces weakening and augmentation operations into the original attention weight calculation and feature extraction processes. In ICSLA, feature learning is coupled with confidence score learning, simultaneously. Leveraging ICSLA, a new and universal decoder is devised, establishing a transformer-based one-stage HOI detection architecture. Experimental results demonstrate the effectiveness of the proposed method in improving HOI detection accuracy, offering valuable insights for further optimization of attention mechanisms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Interaction Confidence Attention for Human–Object Interaction Detection

  • Hong-Bo Zhang,
  • Wang-Kai Lin,
  • Hang Su,
  • Qing Lei,
  • Jing-Hua Liu,
  • Ji-Xiang Du

摘要

In human–object interaction (HOI) detection task, ensuring that interactive pairs receive higher attention weights while reducing the weight of non-interaction pairs is imperative for enhancing HOI detection accuracy. Guiding attention learning is also a key aspect of existing transformer-based algorithms. To tackle this challenge, this study proposes a novel approach termed Interaction Confidence Score Learning Attention (ICSLA), which introduces weakening and augmentation operations into the original attention weight calculation and feature extraction processes. In ICSLA, feature learning is coupled with confidence score learning, simultaneously. Leveraging ICSLA, a new and universal decoder is devised, establishing a transformer-based one-stage HOI detection architecture. Experimental results demonstrate the effectiveness of the proposed method in improving HOI detection accuracy, offering valuable insights for further optimization of attention mechanisms.