Visual semantic entity recognition (visual SER) aims to extract contents that fall in key fields from the given visually-rich document image, and it has been widely applied across diverse scenarios. Most existing visual SER methods employ the BIO tagging schema to extract key entities, necessitating well-organized OCR results at the entity level as prior information. However, meeting this prerequisite is challenging in real-world applications. General OCR engines typically provide disordered line-level results, where entities with multiple text lines are split into several segments. Moreover, some adjacent entities may fall into the same detection box, posing challenges for accurate span detection and text aggregation. To address this issue, this paper introduces a novel framework, ROISER (Real wOrld vIsual Semantic Entity Recognition), integrating entity line span detection, line aggregation, and line classification to achieve visual SER with real-world OCR input. Experiment results demonstrate that our model outperforms existing approaches on various benchmarks, showcasing its effectiveness and compatibility for practical applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ROISER: Towards Real World Semantic Entity Recognition from Visually-Rich Documents

  • Zening Lin,
  • Jiapeng Wang,
  • Wenhui Liao,
  • Weicong Dai,
  • Longfei Xiong,
  • Lianwen Jin

摘要

Visual semantic entity recognition (visual SER) aims to extract contents that fall in key fields from the given visually-rich document image, and it has been widely applied across diverse scenarios. Most existing visual SER methods employ the BIO tagging schema to extract key entities, necessitating well-organized OCR results at the entity level as prior information. However, meeting this prerequisite is challenging in real-world applications. General OCR engines typically provide disordered line-level results, where entities with multiple text lines are split into several segments. Moreover, some adjacent entities may fall into the same detection box, posing challenges for accurate span detection and text aggregation. To address this issue, this paper introduces a novel framework, ROISER (Real wOrld vIsual Semantic Entity Recognition), integrating entity line span detection, line aggregation, and line classification to achieve visual SER with real-world OCR input. Experiment results demonstrate that our model outperforms existing approaches on various benchmarks, showcasing its effectiveness and compatibility for practical applications.