<p>Automated document classification is a vital research domain, utilizing approaches that employ textual features (e.g., plain text analysis), visual features (e.g., feature extraction from document images), or hybrid techniques that combine both modalities to capture both textual and visual features to enhance performance. Our work introduces two hybrid models, HEADoC<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13748_2025_411_Article_IEq1.gif" Format="GIF" Height="11" Rendition="HTML" Resolution="72" Type="Linedraw" Width="39" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{BASE}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mrow> <mi mathvariant="italic">BASE</mi> </mrow> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> with 27.7 million parameters and HEADoC<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13748_2025_411_Article_IEq2.gif" Format="GIF" Height="11" Rendition="HTML" Resolution="72" Type="Linedraw" Width="47" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{LARGE}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mrow> <mi mathvariant="italic">LARGE</mi> </mrow> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> with 90.58 million parameters. A key innovation is our deep attention mechanism, a streamlined yet efficient method inspired by conventional attention frameworks, which facilitates the smooth integration of the two modalities. During experimentation, we observed that ArcFace loss, a metric learning approach effective for heterogeneous datasets with distinct intra-class characteristics, performed poorly in our task due to the homogeneity of document classes in standard benchmarks. Both models were evaluated on the RVL-CDIP and Tobacco3482 datasets. On Tobacco3482, HEADoC<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13748_2025_411_Article_IEq1.gif" Format="GIF" Height="11" Rendition="HTML" Resolution="72" Type="Linedraw" Width="39" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{BASE}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mrow> <mi mathvariant="italic">BASE</mi> </mrow> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> achieved 95.98% accuracy, while HEADoC<InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13748_2025_411_Article_IEq2.gif" Format="GIF" Height="11" Rendition="HTML" Resolution="72" Type="Linedraw" Width="47" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{LARGE}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mrow> <mi mathvariant="italic">LARGE</mi> </mrow> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> reached 96.66%. For RVL-CDIP, accuracies were 92.95% and 93.62%, respectively. HEADoC<InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="13748_2025_411_Article_IEq1.gif" Format="GIF" Height="11" Rendition="HTML" Resolution="72" Type="Linedraw" Width="39" /> </InlineMediaObject> <EquationSource Format="TEX">\(_{BASE}\)</EquationSource> <EquationSource Format="MATHML"><math> <mmultiscripts> <mrow /> <mrow> <mi mathvariant="italic">BASE</mi> </mrow> <mrow /> </mmultiscripts> </math></EquationSource> </InlineEquation> not only surpassed numerous state-of-the-art models but also proved itself to be the most compact architecture in comparison, highlighting efficiency without sacrificing performance. Having the most compact sizes amongst their competitors, our models can be trained faster than its rival architectures.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HEADoC: Highly Efficient and Accurate Document Classifier Optimized Using Semantic Distances

  • Muhammad Usman Khan,
  • Muhammad Adnan Tariq,
  • Muhammad Shaoor Siddique

摘要

Automated document classification is a vital research domain, utilizing approaches that employ textual features (e.g., plain text analysis), visual features (e.g., feature extraction from document images), or hybrid techniques that combine both modalities to capture both textual and visual features to enhance performance. Our work introduces two hybrid models, HEADoC \(_{BASE}\) BASE with 27.7 million parameters and HEADoC \(_{LARGE}\) LARGE with 90.58 million parameters. A key innovation is our deep attention mechanism, a streamlined yet efficient method inspired by conventional attention frameworks, which facilitates the smooth integration of the two modalities. During experimentation, we observed that ArcFace loss, a metric learning approach effective for heterogeneous datasets with distinct intra-class characteristics, performed poorly in our task due to the homogeneity of document classes in standard benchmarks. Both models were evaluated on the RVL-CDIP and Tobacco3482 datasets. On Tobacco3482, HEADoC \(_{BASE}\) BASE achieved 95.98% accuracy, while HEADoC \(_{LARGE}\) LARGE reached 96.66%. For RVL-CDIP, accuracies were 92.95% and 93.62%, respectively. HEADoC \(_{BASE}\) BASE not only surpassed numerous state-of-the-art models but also proved itself to be the most compact architecture in comparison, highlighting efficiency without sacrificing performance. Having the most compact sizes amongst their competitors, our models can be trained faster than its rival architectures.