Tibetan newspapers, as important carriers of culture and information, require effective segmentation of their text and images as a crucial step before content recognition. Through precise segmentation technology, we can effectively extract and manage information from Tibetan newspapers, thereby enhancing the accuracy and efficiency of subsequent recognition and reconstruction tasks, and laying a solid foundation for the digital preservation and application of Tibetan newspapers. However, the increasingly innovative layout design of Tibetan newspapers has undoubtedly increased the difficulty for models to understand the content on the page. Especially when dealing with high-resolution and complexly composed images, the model faces challenges in achieving precise overall segmentation, and the segmentation of small objects is also prone to imprecision or failure. Analyzing the relationships between different regions is even more challenging. In this paper, we propose an improved DeepLabv3+ model. By introducing a pyramid structure, it better captures the context and spatial information of the image. At the same time, the backbone network is replaced from Xception to a more lightweight improved version, MobileNetV2, and a feature fusion mechanism is adopted in the backbone network to enhance the model’s ability in feature extraction. The improved model’s segmentation performance has increased from 76% to 79.09%, the F_score from 84.5% to 93.8%, and Accuracy from 93.69% to 93.88%. These improvements significantly enhance the model’s ability to handle segmentation in high-resolution complex scenes and small objects, providing an effective solution for the automatic segmentation of Tibetan newspapers.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on Small Object Segmentation of Tibetan Newspapers Using DeepLabV3+ at High Resolution

  • Hongrui Li,
  • Weilan Wang,
  • Zhenjie Wu,
  • Dazhi Yang

摘要

Tibetan newspapers, as important carriers of culture and information, require effective segmentation of their text and images as a crucial step before content recognition. Through precise segmentation technology, we can effectively extract and manage information from Tibetan newspapers, thereby enhancing the accuracy and efficiency of subsequent recognition and reconstruction tasks, and laying a solid foundation for the digital preservation and application of Tibetan newspapers. However, the increasingly innovative layout design of Tibetan newspapers has undoubtedly increased the difficulty for models to understand the content on the page. Especially when dealing with high-resolution and complexly composed images, the model faces challenges in achieving precise overall segmentation, and the segmentation of small objects is also prone to imprecision or failure. Analyzing the relationships between different regions is even more challenging. In this paper, we propose an improved DeepLabv3+ model. By introducing a pyramid structure, it better captures the context and spatial information of the image. At the same time, the backbone network is replaced from Xception to a more lightweight improved version, MobileNetV2, and a feature fusion mechanism is adopted in the backbone network to enhance the model’s ability in feature extraction. The improved model’s segmentation performance has increased from 76% to 79.09%, the F_score from 84.5% to 93.8%, and Accuracy from 93.69% to 93.88%. These improvements significantly enhance the model’s ability to handle segmentation in high-resolution complex scenes and small objects, providing an effective solution for the automatic segmentation of Tibetan newspapers.