Towards Enhanced Word Spotting in Historical Devanagari Documents Using Deep CNN Architecture
摘要
In this paper, a deep Convolutional Neural Network approach is proposed toward more challenging word spotting in historical text documents that may be affected by variant writing styles, distortions, and noise in the background. The model learns discriminative representations along with comparison with HMM and SVM models for raw images, thereby getting rid of the need for handcrafted features. It outperforms state-of-the-art methods in benchmark datasets, being applied to historical handwritten Devanagari scripts and layouts and having proved to be a powerful method for indexing and retrieving historical archives. Therefore, this work helps move forward the domain of digital preservation and access for efficient exploration of information locked within historical manuscripts.