<p>Structured document understanding aims to read and analyze both textual and structured information in form documents. Existing models lack sufficient integration of cross-modal information, leaving room for improvement in accuracy. To address this issue, this paper proposes a new network GCAF network based on LiLT. In GCAF, a residual gate module is introduced to ensure that important information from each modality is seamlessly inputted into the fusion module. Additionally, Previous methods often rely on simple concatenation for multiple modalities in form documents, which may not facilitate effective cross-modal fusion. this paper introduces a cross-attention mechanism integrated into the layout information side within LiLT. By doing so, the fusion of layout and textual information is achieved more effectively, leading to superior integration of multimodal information. Experimental evaluations on the FUNSD dataset demonstrate that the proposed GCAF model achieves superior performance with an accuracy of 0.8912.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A structured document understanding model based on gate mechanism and cross attention

  • Bin Jiang,
  • Lei Zhang,
  • Yong Wang

摘要

Structured document understanding aims to read and analyze both textual and structured information in form documents. Existing models lack sufficient integration of cross-modal information, leaving room for improvement in accuracy. To address this issue, this paper proposes a new network GCAF network based on LiLT. In GCAF, a residual gate module is introduced to ensure that important information from each modality is seamlessly inputted into the fusion module. Additionally, Previous methods often rely on simple concatenation for multiple modalities in form documents, which may not facilitate effective cross-modal fusion. this paper introduces a cross-attention mechanism integrated into the layout information side within LiLT. By doing so, the fusion of layout and textual information is achieved more effectively, leading to superior integration of multimodal information. Experimental evaluations on the FUNSD dataset demonstrate that the proposed GCAF model achieves superior performance with an accuracy of 0.8912.