This paper presents a novel and robust document image dewarping method, namely DocMamba, based on the idea of selective state space sequence modeling. It consists of three modules, document image augmentation and feature extraction, sequence modeling and contextual information learning, and the robust document image dewarping. In particular, given a distorted document image, we first extract its deep convolution features, outputting a group of down-sampled feature maps. Each feature map is flatten into a vector, and a document sequence is built by all these feature vectors. The contextual information hidden in the sequence are learned by using the Selective State Space Sequence Model. That is, sequence-to-sequence transformations are performed based on the Mamba2 blocks. All sequences are then reshaped to updated feature maps and further encoded by using the dilated convolution layers. Finally, the original feature maps and the final feature maps are adding together and fed into a rectification decoder to estimate a coarse backward mapping. The final rectified image is achieved by performing the up-sampled backward mapping on the original distorted image. Extensive experiments conducted on the DocUNet and DIR300 benchmarks showed the effectiveness of the proposed method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DocMamba: Robust Document Image Dewarping via Selective State Space Sequence Modeling

  • Miaolin Han,
  • Huibin Li

摘要

This paper presents a novel and robust document image dewarping method, namely DocMamba, based on the idea of selective state space sequence modeling. It consists of three modules, document image augmentation and feature extraction, sequence modeling and contextual information learning, and the robust document image dewarping. In particular, given a distorted document image, we first extract its deep convolution features, outputting a group of down-sampled feature maps. Each feature map is flatten into a vector, and a document sequence is built by all these feature vectors. The contextual information hidden in the sequence are learned by using the Selective State Space Sequence Model. That is, sequence-to-sequence transformations are performed based on the Mamba2 blocks. All sequences are then reshaped to updated feature maps and further encoded by using the dilated convolution layers. Finally, the original feature maps and the final feature maps are adding together and fed into a rectification decoder to estimate a coarse backward mapping. The final rectified image is achieved by performing the up-sampled backward mapping on the original distorted image. Extensive experiments conducted on the DocUNet and DIR300 benchmarks showed the effectiveness of the proposed method.