<p>How to maximize the user experience via multimodal interactive systems is a popular topic in artificial intelligence (AI) research as a result of the development of Artificial Intelligence Generated Content (AIGC) technology. In order to improve the user experience of multimodal interactive system, this study proposes a Knowledge-Enhanced Multimodal Dialogue Generation (KEMD-G) model. This model constructs a knowledge selection and confidence filtering mechanism and designs a knowledge-enhanced multimodal feature bidirectional refinement fusion module. It achieves collaborative modeling of text, vision and external knowledge in the encoding stage. A cross-modal knowledge weighting mechanism is introduced in the decoding stage to improve the semantic relevance and multimodal consistency of generated content. The model adopts Bidirectional and Auto-Regressive Transformers (BART) for text encoding and decoding, combined with a visual Transformer for visual feature extraction. The results show that the performance of KEMD-G on MMConv and Photochat datasets is better than that of the contrast model, and the performance on Bilingual Evaluation Understudy- 1-gram (BLEU-1) and Bidirectional Encoder Representations from Transformers (BERT)-score is 39.69% and 0.88, respectively, and <i>p</i> &lt; 0.05. The ablation experiments demonstrated that the knowledge confidence screening mechanism, the bidirectional refined feature fusion module, and the multimodal weighted attention mechanism at the decoding end all effectively improved the model performance. To verify the system usability, a controlled user study is designed, and a t-test is used for significance analysis. The results show that KEMD-G is significantly superior to the baseline system in terms of system usability, interface clarity, and dialogue efficiency (<i>p</i> &lt; 0.05). In summary, KEMD-G provides an innovative and practical solution for AIGC multimodal dialogue generation tasks, with certain application value and broad research prospects.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Design and evaluation model of AIGC multimodal interactive system for user experience optimization

  • Yifei Wang,
  • Shanling Shen

摘要

How to maximize the user experience via multimodal interactive systems is a popular topic in artificial intelligence (AI) research as a result of the development of Artificial Intelligence Generated Content (AIGC) technology. In order to improve the user experience of multimodal interactive system, this study proposes a Knowledge-Enhanced Multimodal Dialogue Generation (KEMD-G) model. This model constructs a knowledge selection and confidence filtering mechanism and designs a knowledge-enhanced multimodal feature bidirectional refinement fusion module. It achieves collaborative modeling of text, vision and external knowledge in the encoding stage. A cross-modal knowledge weighting mechanism is introduced in the decoding stage to improve the semantic relevance and multimodal consistency of generated content. The model adopts Bidirectional and Auto-Regressive Transformers (BART) for text encoding and decoding, combined with a visual Transformer for visual feature extraction. The results show that the performance of KEMD-G on MMConv and Photochat datasets is better than that of the contrast model, and the performance on Bilingual Evaluation Understudy- 1-gram (BLEU-1) and Bidirectional Encoder Representations from Transformers (BERT)-score is 39.69% and 0.88, respectively, and p < 0.05. The ablation experiments demonstrated that the knowledge confidence screening mechanism, the bidirectional refined feature fusion module, and the multimodal weighted attention mechanism at the decoding end all effectively improved the model performance. To verify the system usability, a controlled user study is designed, and a t-test is used for significance analysis. The results show that KEMD-G is significantly superior to the baseline system in terms of system usability, interface clarity, and dialogue efficiency (p < 0.05). In summary, KEMD-G provides an innovative and practical solution for AIGC multimodal dialogue generation tasks, with certain application value and broad research prospects.