This paper introduces a comprehensive method for automatically detecting and extracting tables from scanned documents. Using a region-based convolutional neural network (RCNN) with VGG19 as its backbone, our approach achieves an impressive detection accuracy of approximately 93%, capable of identifying both bordered and borderless tables, enhancing its versatility across document types. A split-and-merge technique is used for table structure recognition, producing a weighted average F1 score of 52.3% compared to a popular method, CascadeTabNet, with 23.2%. Our approach is validated using the Marmot dataset, comprising over a thousand English scanned documents, ensuring the robustness and generalisability of our model across diverse document sets. By automating the extraction of tabular data from scanned documents, our methodology streamlines information retrieval processes, boosting efficiency and accuracy in data-driven tasks. As such, our approach represents a significant advancement in document analysis and holds promise for a wide range of real-world applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Table Detection and Tabular Data Extraction from Scanned Documents

  • Yash Amethiya,
  • Garima Bajwa

摘要

This paper introduces a comprehensive method for automatically detecting and extracting tables from scanned documents. Using a region-based convolutional neural network (RCNN) with VGG19 as its backbone, our approach achieves an impressive detection accuracy of approximately 93%, capable of identifying both bordered and borderless tables, enhancing its versatility across document types. A split-and-merge technique is used for table structure recognition, producing a weighted average F1 score of 52.3% compared to a popular method, CascadeTabNet, with 23.2%. Our approach is validated using the Marmot dataset, comprising over a thousand English scanned documents, ensuring the robustness and generalisability of our model across diverse document sets. By automating the extraction of tabular data from scanned documents, our methodology streamlines information retrieval processes, boosting efficiency and accuracy in data-driven tasks. As such, our approach represents a significant advancement in document analysis and holds promise for a wide range of real-world applications.