<p>Visually Rich Document Understanding (VRDU) requires models that effectively capture spatial layouts and semantic relationships among textual and visual entities. This paper presents a Graph-Augmented Multi-Stage Transformer Model that integrates Graph Neural Networks (GNNs) with 2D positional embeddings to enhance spatial reasoning and contextual representations. The proposed model introduces learnable row–column embeddings and a hierarchical multi-stage transformer architecture for efficient and progressive feature refinement. Comprehensive evaluations on the FUNSD and DocVQA datasets demonstrate consistent performance improvements, achieving 91.35% F1 and 79.91% ANLS scores, respectively, establishing new benchmarks in structured document understanding. Furthermore, evaluation on the DUDE dataset, which comprises complex, multi-page, and heterogeneous documents, illustrates the model’s scalability and robustness to real-world document variations. Comparative analyses with LayoutLMv3, LiLT, and Qwen2-VL confirm that the proposed approach achieves strong generalization with competitive efficiency, making it a robust and practical solution for document layout understanding tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Graph-Augmented Multi-Stage Transformer Model for Document Layout Understanding

  • Aresha Arshad,
  • Momina Moetesum,
  • Adnan ul Hasan,
  • Faisal Shafait

摘要

Visually Rich Document Understanding (VRDU) requires models that effectively capture spatial layouts and semantic relationships among textual and visual entities. This paper presents a Graph-Augmented Multi-Stage Transformer Model that integrates Graph Neural Networks (GNNs) with 2D positional embeddings to enhance spatial reasoning and contextual representations. The proposed model introduces learnable row–column embeddings and a hierarchical multi-stage transformer architecture for efficient and progressive feature refinement. Comprehensive evaluations on the FUNSD and DocVQA datasets demonstrate consistent performance improvements, achieving 91.35% F1 and 79.91% ANLS scores, respectively, establishing new benchmarks in structured document understanding. Furthermore, evaluation on the DUDE dataset, which comprises complex, multi-page, and heterogeneous documents, illustrates the model’s scalability and robustness to real-world document variations. Comparative analyses with LayoutLMv3, LiLT, and Qwen2-VL confirm that the proposed approach achieves strong generalization with competitive efficiency, making it a robust and practical solution for document layout understanding tasks.