<p>Tibetan document layout analysis (DLA) is a crucial aspect of the digitization of Tibetan documents. However, two significant issues persist: (1) the inadequacy of Tibetan document datasets;(2) insufficient utilization of layout information by existing models. To address these challenges, we have constructed a novel dataset, the Aba Tibetan Newspaper Dataset (AbaTND), consisting of 571 color images of Tibetan newspapers, effectively filling a data gap in this field. Additionally, we propose an advanced model, YOLOv10-CBRC (YOLOv10-CBAM-Re-Upsample-Re-SCDown-CRCV3), along with its variant YOLOv10-RC (YOLOv10-Re-Upsample-Re-SCDown-CRCV3), aimed at enhancing information utilization. Building upon YOLOv10 as the baseline model, the proposed model implements three key improvements: (1) the replacement of the Partial Self-Attention (PSA) module with the Convolutional Block Attention Module (CBAM), which enhances perception of channel-wise information and spatial localization of objects,(2) the adoption of dual-branch Re-upsample and Re-SCDown modules (2Re), which facilitates more effective utilization of multi-scale information,(3) the design of a novel classification feature processor, CIBwithResidualCV3 (CRCV3), which improves performance in classification tasks.Experimental results demonstrate that YOLOv10-CBRC achieves a mAP<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="44443_2025_168_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="29" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{50\text {-}95}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mrow> <mn>50</mn> <mtext>-</mtext> <mn>95</mn> </mrow> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> of 77.6% on the AbaTND, while YOLOv10-RC reaches mAP<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="44443_2025_168_Article_IEq1.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="29" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{50\text {-}95}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mrow> <mn>50</mn> <mtext>-</mtext> <mn>95</mn> </mrow> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> scores of 70.6% and 74.6% on the <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="44443_2025_168_Article_IEq3.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="50" /> </InlineMediaObject> <EquationSource Format="TEX">\(D^4LA\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msup> <mi>D</mi> <mn>4</mn> </msup> <mi>L</mi> <mi>A</mi> </mrow> </math></EquationSource> </InlineEquation> and IIIT-AR-13K datasets, respectively, significantly outperforming baseline model. The dataset and source code are publicly available at: <a href="https://github.com/fengmuyanghua/YOLOv10-CBRC-AbaTND">https://github.com/fengmuyanghua/YOLOv10-CBRC-AbaTND</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

YOLOv10-CBRC: A high-precision document image layout analysis model

  • Zhenjie Wu,
  • Weilan Wang,
  • Hongrui Li

摘要

Tibetan document layout analysis (DLA) is a crucial aspect of the digitization of Tibetan documents. However, two significant issues persist: (1) the inadequacy of Tibetan document datasets;(2) insufficient utilization of layout information by existing models. To address these challenges, we have constructed a novel dataset, the Aba Tibetan Newspaper Dataset (AbaTND), consisting of 571 color images of Tibetan newspapers, effectively filling a data gap in this field. Additionally, we propose an advanced model, YOLOv10-CBRC (YOLOv10-CBAM-Re-Upsample-Re-SCDown-CRCV3), along with its variant YOLOv10-RC (YOLOv10-Re-Upsample-Re-SCDown-CRCV3), aimed at enhancing information utilization. Building upon YOLOv10 as the baseline model, the proposed model implements three key improvements: (1) the replacement of the Partial Self-Attention (PSA) module with the Convolutional Block Attention Module (CBAM), which enhances perception of channel-wise information and spatial localization of objects,(2) the adoption of dual-branch Re-upsample and Re-SCDown modules (2Re), which facilitates more effective utilization of multi-scale information,(3) the design of a novel classification feature processor, CIBwithResidualCV3 (CRCV3), which improves performance in classification tasks.Experimental results demonstrate that YOLOv10-CBRC achieves a mAP \(_{50\text {-}95}\) 50 - 95 of 77.6% on the AbaTND, while YOLOv10-RC reaches mAP \(_{50\text {-}95}\) 50 - 95 scores of 70.6% and 74.6% on the \(D^4LA\) D 4 L A and IIIT-AR-13K datasets, respectively, significantly outperforming baseline model. The dataset and source code are publicly available at: https://github.com/fengmuyanghua/YOLOv10-CBRC-AbaTND.