In recent decades, numerous algorithms have been proposed for extracting tables from PDF documents, with most designed as general-purpose table extractors. This paper introduces a novel algorithm specifically tailored for extracting order tables, addressing specific issues not covered by existing general-purpose extractors. Our approach, named OTE, identifies leading rows in an order table through a clustering algorithm and employs heuristics to recognize additional row lines and annotate columns. Through an evaluation of 115 order documents from customers of a medium-sized company, we demonstrate that: i) OTE surpasses general-purpose extractors, ii) accurately identifies over 95% of order tables in PDF documents, and iii) correctly identifies 81% of all listed article IDs, even when included in the article description.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

OTE: A Tool For Extracting Tabular Purchasing Order Information From PDF Documents

  • Michael Scholz,
  • Jörg Bauer

摘要

In recent decades, numerous algorithms have been proposed for extracting tables from PDF documents, with most designed as general-purpose table extractors. This paper introduces a novel algorithm specifically tailored for extracting order tables, addressing specific issues not covered by existing general-purpose extractors. Our approach, named OTE, identifies leading rows in an order table through a clustering algorithm and employs heuristics to recognize additional row lines and annotate columns. Through an evaluation of 115 order documents from customers of a medium-sized company, we demonstrate that: i) OTE surpasses general-purpose extractors, ii) accurately identifies over 95% of order tables in PDF documents, and iii) correctly identifies 81% of all listed article IDs, even when included in the article description.