Challenges and Gaps in Scene Text Detection and Recognition: A Detailed Survey
摘要
Recently, the amount of accessible real-time scene images has been significantly increased due to a variety of real-time applications, for instance, self-driving automobiles, virtual reality, and smart document processing. This has fostered research activities on text detection and recognition, most especially in scenes that are cluttered and changing frequently and hence difficult for traditional approaches. However, it is also a fact that, with the help of this deep learning technology, it has boosted up the competency of text recognition systems a notch more, which can perfectly work in various cases optimally. Before the development of deep learning models, the OCR system followed more rule-based approaches. Such methods of OCR were typical for printed textual materials and were efficient in the cases when working with real scene images was challenging. The text was of different fonts and backgrounds and the quality of theimages was not quite standard and thus affected the recognition performance. Many traditional, manually designed feature extraction methods, used for feature extraction have been observed to lack the capability of generalizability, especially in complex and congested scenes. After that, with the advent of deep learning, a new perspective on the identification and recognition of text appeared. These, along with the new concepts such as the attention mechanism, transformer architecture, and several others, have resulted in deeper learning models and highly efficient ones that offer high accuracy levels. This manuscript also gives a systematic survey of most of the conventional as well as deep learning-based text detection methods. It reviews the major concerns of text detection and the way deep learning approaches have overcome shortcomings of previous approaches as well as limitations, which include the issue of detecting text at any orientation in different font types and of different languages. The paper also brings into the light the most recent developments in the hierarchy of text recognition where detection and recognition stages are merged into a single model to perform well in real-time applications. Also, this paper discusses the possibilities of Large Language Models (LLMs) in transforming text detection and recognition systems. Models like GPT and BERT used in this paper are considered to be among the best-performing LLMs. It is expected to see great enhancements in text recognition tasks when incorporating the LLMs with conventional vision models in terms of contextual understanding and multilingual understanding and better performance with noisy data. In general, this paper strives to present a comprehensive account of text detection and/or recognition methodology and the prime importance of deep learning and potential future work due to LLMs. Such refinements do not only enhance the likelihood and velocity of scene-based text recognition tools but also direct new approaches to their real-world uses in diverse sectors.