This paper presents an enhanced attention-based joint model for intent detection and slot filling, critical tasks in Natural Language Understanding (NLU). The proposed model integrates scalar attention, an association matrix, and a scaling factor to address key challenges, including preserving token order and stabilizing attention scores. Through experiments on the ATIS and SNIPS datasets, the model demonstrated significant improvements over state-of-the-art baselines, achieving a 1.38% gain in accuracy and a 2.03% improvement in F1-score on ATIS, and a 0.94% accuracy gain and 1.06% F1-score improvement on SNIPS. The model’s ability to capture contextual dependencies and maintain training stability is highlighted, particularly for the slot filling task, where the association matrix and scaling factor played crucial roles. Ablation studies further underscored the impact of these enhancements, with results demonstrating optimal performance at specific channel configurations. This study offers a robust framework for advancing joint learning models in linguistically diverse settings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing an Attention-Based Joint Learning for Intent Detection and Slot Filling with Order Preservation and Scaling

  • Yusuf Idris Muhammad,
  • Naomie Salim,
  • Anazida Zainal,
  • Sharin Hazlin Huspi,
  • Ismail Balarabe Isah

摘要

This paper presents an enhanced attention-based joint model for intent detection and slot filling, critical tasks in Natural Language Understanding (NLU). The proposed model integrates scalar attention, an association matrix, and a scaling factor to address key challenges, including preserving token order and stabilizing attention scores. Through experiments on the ATIS and SNIPS datasets, the model demonstrated significant improvements over state-of-the-art baselines, achieving a 1.38% gain in accuracy and a 2.03% improvement in F1-score on ATIS, and a 0.94% accuracy gain and 1.06% F1-score improvement on SNIPS. The model’s ability to capture contextual dependencies and maintain training stability is highlighted, particularly for the slot filling task, where the association matrix and scaling factor played crucial roles. Ablation studies further underscored the impact of these enhancements, with results demonstrating optimal performance at specific channel configurations. This study offers a robust framework for advancing joint learning models in linguistically diverse settings.