Intent recognition is a crucial task in natural language understanding. In the real world, discerning human intentions requires integrating information from speech, facial expressions, and gestures, among other modalities. Multimodal intent recognition can combine data from diverse modalities to interpret human language and behavior. However, most existing methods struggle to process large volumes of redundant, unaligned multimodal sequential data, leading to inefficiencies in modeling multimodal fusion within such unaligned datasets. The proposed method introduces an Auxiliary Context Module (ACM). In particular, the ACM uses each modality's utterance-level representations as a global multimodal context, which interacts with local unimodal information to improve each other. This method offers better performance than earlier local-local cross-modal interaction strategies while also avoiding the quadratic scaling penalty. The Weighted Multihead Fusion Network (WMF) further refines fusion results through gated neural networks and multi-head attention systems. In experiments, the proposed method compares to the most advanced methods, significant performance improvements have been achieved on two datasets. Additionally, ablation experiments confirm the significant contributions of the ACM module and the WMF method in enhancing modal feature representation and improving intent recognition performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Auxiliary Context Module and Weighted Multihead Fusion for Multimodal Intent Recognition

  • Yichao Xia,
  • Jinmiao Song,
  • Shengwei Tian

摘要

Intent recognition is a crucial task in natural language understanding. In the real world, discerning human intentions requires integrating information from speech, facial expressions, and gestures, among other modalities. Multimodal intent recognition can combine data from diverse modalities to interpret human language and behavior. However, most existing methods struggle to process large volumes of redundant, unaligned multimodal sequential data, leading to inefficiencies in modeling multimodal fusion within such unaligned datasets. The proposed method introduces an Auxiliary Context Module (ACM). In particular, the ACM uses each modality's utterance-level representations as a global multimodal context, which interacts with local unimodal information to improve each other. This method offers better performance than earlier local-local cross-modal interaction strategies while also avoiding the quadratic scaling penalty. The Weighted Multihead Fusion Network (WMF) further refines fusion results through gated neural networks and multi-head attention systems. In experiments, the proposed method compares to the most advanced methods, significant performance improvements have been achieved on two datasets. Additionally, ablation experiments confirm the significant contributions of the ACM module and the WMF method in enhancing modal feature representation and improving intent recognition performance.