Text line segmentation approach combining deep learning model and traditional image processing techniques - application to transliteration of Cham manuscripts
摘要
Text line segmentation in historical document images is a crucial task in document analysis and understanding. Historical documents often feature degraded and complex layouts, and they may exhibit distinct writing styles and languages, posing challenges for text line segmentation. Our study focuses on the text line segmentation task for Cham manuscript images. Cham manuscripts, written on Chinese paper and composed in Middle Cham, serve as precious sources to understand the daily lives of the inhabitants of the Champa Kingdom and provide insights into Champa possessions and their management. To address the challenges of text line segmentation in Cham manuscript images, this study proposes an approach that combines deep learning methods with traditional image processing techniques. Our approach takes into account the characteristics of Cham manuscripts to effectively tackle the challenges posed by these documents. Specifically, the procedure begins by leveraging pre-trained deep learning models and includes a fine-tuning step to extract the candidate baseline. Then, various image post-processing algorithms are introduced to further improve segmentation accuracy. Extensive experiments on our Cham manuscripts dataset consisting of 627 manuscripts with approximately 8300 segmented text lines using different evaluation criteria demonstrate the significant superiority of our approach over existing methods. Our method outperforms state-of-the-art performance by a considerable margin.