A literature review serves as a foundation for current research methodologies and methods to be used, and research is validated by earlier experiments and findings if the same dataset is available. Document layout analysis (DLA) serves as a preliminary phase within document comprehension systems, responsible for identifying and annotating the physical arrangement of documents. It encompasses the automated extraction of both structural and contextual insights from document images, with applications spanning document retrieval, text extraction, and machine translation. DLA’s core objective lies in disassembling a document image into its fundamental components, which encompass text blocks, tables, visuals, and headings. Empirical evaluations demonstrate that advanced deep learning models, including Mask RCNN and LayoutLM, consistently achieve accuracies exceeding 90% on diverse datasets such as PubLayNet and DocLayNet. These models excel in improving precision, recall, and segmentation accuracy, highlighting the transformative potential of deep learning in this domain. Despite notable progress, challenges persist, including handling non-standard layouts, multilingual documents, and the need for large annotated datasets. This survey underscores the importance of advancing DLA techniques to address these challenges, paving the way for more accurate, robust, and accessible document digitization systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Document Layout Analysis: A Review

  • Atul Kumar,
  • Gurpreet Singh Lehal

摘要

A literature review serves as a foundation for current research methodologies and methods to be used, and research is validated by earlier experiments and findings if the same dataset is available. Document layout analysis (DLA) serves as a preliminary phase within document comprehension systems, responsible for identifying and annotating the physical arrangement of documents. It encompasses the automated extraction of both structural and contextual insights from document images, with applications spanning document retrieval, text extraction, and machine translation. DLA’s core objective lies in disassembling a document image into its fundamental components, which encompass text blocks, tables, visuals, and headings. Empirical evaluations demonstrate that advanced deep learning models, including Mask RCNN and LayoutLM, consistently achieve accuracies exceeding 90% on diverse datasets such as PubLayNet and DocLayNet. These models excel in improving precision, recall, and segmentation accuracy, highlighting the transformative potential of deep learning in this domain. Despite notable progress, challenges persist, including handling non-standard layouts, multilingual documents, and the need for large annotated datasets. This survey underscores the importance of advancing DLA techniques to address these challenges, paving the way for more accurate, robust, and accessible document digitization systems.