A novel approach for text prediction and recognition from natural scene images
摘要
In the digital era, texts act as one of the essential means for transmitting information. However, it is very crucial to detect text patterns in an extremely complex background and multilingual environment. Although improvements in Deep Learning (DL)-centric techniques have succeeded in enhancing the accuracy of Text Detection (TD) and recognition from scene images, performing the same in a cluttered background still remains challenging. Therefore, this paper presents a novel nature scene text prediction approach using the Vocabulary-Transformer-Pipeline pruner (VTP)-based Phrasenet. Initially, the Natural Scene (NS) image dataset is collected from publicly available sources. Then, pre-processing is performed for the NS image; here, foreground and background pixel separation, clutter background removal by LMS, image reconstruction, and image pyramid are done. From the pre-processed image, the text region is predicted by using RSA-SEDO-MSER. Next, significant features are extracted from the predicted text region by employing SIFT, IFT, and HOG techniques. Subsequently, based on MEBOA-HF, the text is recognized by extracting the text region’s candidate pixels. Meanwhile, the non-text regions are recognized by placing the rotating anchor boxes. Then, features are extracted from the recognized non-texts for efficient recognition criterion. Thereafter, the extracted features are subjected to Pixel-BERT for efficient mapping of the image pixels with text. Eventually, the text is detected from the mapped text by utilizing VTP-PN. The experimental outcomes proved that the proposed methodology achieved a higher accuracy of 97.597%, thus outperforming prevailing techniques.