<p>Automated table structure recognition from input images faces significant visual and linguistic challenges, including layout variances, table densities, and the presence of empty or multi-line cells. This problem is crucial due to its extensive applications in business and scientific domains. For this purpose, we present <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(TabStruct-Net\ V2\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>T</mi> <mi>a</mi> <mi>b</mi> <mi>S</mi> <mi>t</mi> <mi>r</mi> <mi>u</mi> <mi>c</mi> <mi>t</mi> <mo>-</mo> <mi>N</mi> <mi>e</mi> <mi>t</mi> <mspace width="4pt" /> <mi>V</mi> <mn>2</mn> </mrow> </math></EquationSource> </InlineEquation>, our end-to-end trainable model using a two-step process for table parsing: <i>Top Down</i>, decomposing the table image into cells, and <i>Bottom Up</i>, associating cells to define the complete structure by assigning row and column spans. This method leverages constraints inspired by human cognition, such as cell alignment, continuity, and non-overlapping characteristics, optimizing the objective function to preserve these attributes. We introduce a vision transformer backbone augmented with local attention tailored for high-resolution input images. We also enhance inference speed for row and column adjacency prediction by replacing sampling-based graph neural networks with a single head self attention layer. Furthermore, we propose an effective split-and-merge strategy to handle highly dense tables without altering the network architecture. Additionally, we present problem-specific training strategies that improve performance across diverse visual characteristics and support cross-domain adaptations. Finally, through our robust protocol for comprehensive dataset analysis and result interpretation, we offer clear directions for future work in automated table structure recognition, highlighting its potential for substantial advancements in practical applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From pixels to tables: reconstructing complex tables from document images

  • Sachin Raja,
  • Ajoy Mondal,
  • C. V. Jawahar

摘要

Automated table structure recognition from input images faces significant visual and linguistic challenges, including layout variances, table densities, and the presence of empty or multi-line cells. This problem is crucial due to its extensive applications in business and scientific domains. For this purpose, we present \(TabStruct-Net\ V2\) T a b S t r u c t - N e t V 2 , our end-to-end trainable model using a two-step process for table parsing: Top Down, decomposing the table image into cells, and Bottom Up, associating cells to define the complete structure by assigning row and column spans. This method leverages constraints inspired by human cognition, such as cell alignment, continuity, and non-overlapping characteristics, optimizing the objective function to preserve these attributes. We introduce a vision transformer backbone augmented with local attention tailored for high-resolution input images. We also enhance inference speed for row and column adjacency prediction by replacing sampling-based graph neural networks with a single head self attention layer. Furthermore, we propose an effective split-and-merge strategy to handle highly dense tables without altering the network architecture. Additionally, we present problem-specific training strategies that improve performance across diverse visual characteristics and support cross-domain adaptations. Finally, through our robust protocol for comprehensive dataset analysis and result interpretation, we offer clear directions for future work in automated table structure recognition, highlighting its potential for substantial advancements in practical applications.