<p>Scene Text Recognition (STR) involves deciphering textual content embedded within complex, natural scene images, often following detection stages or integrated into end-to-end pipelines. Addressing the challenge of STR in noisy target domains, characterized by inter-domain and intra-domain noise, cluttered backgrounds, and irregular text shapes, this study proposes a robust and understandable framework titled Fusion-Based Adaptation for Scene Text Recognition (FASTR). The framework integrates a primary classifier with an epistemically aware auxiliary classifier to model uncertainty, supported by a novel Adaptive Scale Feature Module (ASFM) that enhances localisation through pixel-level mask prediction and multi-scale fusion. A Triple-Level Confidence (TLC) strategy—categorized into high, medium, and low consistency thresholds—is introduced to enforce consistency loss and improve generalisation across domains. Additionally, a pseudo-labelling scheme refines the adaptation process through self-training under structured domain noise. FASTR is trained and evaluated on both synthetic (SynthText, MJSynth) and real-world (ICDAR 2013, SVT, and IIIT5K) datasets. It achieves a word recognition accuracy of 92.4% on IIIT5K, 89.7% on SVT, and 93.1% on ICDAR 2013, outperforming state-of-the-art baselines by an average margin of 2.8%. On cross-domain benchmarks with added noise, FASTR maintains high performance, achieving an average F1-score of 90.5%, with precision and recall values of 91.2% and 89.9%, respectively. Hyperparameters, training configurations, and evaluation metrics are transparently documented to ensure reproducibility. The findings demonstrate superior scale robustness, effective domain adaptation, and resilience to cluttered backgrounds, with explainability preserved through interpretable confidence maps and visual cues.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient Scene Text Recognition in Noisy Environments Using Fusion-Based Adaptation and Triple-Level Confidence Modeling

  • Weam M. Binjumah,
  • Bala Dhandayuthapani Veerasamy,
  • S. Kavitha,
  • Suchi Mishra,
  • Shraddha Viraj Pandit,
  • Omair Ameerbakhsh

摘要

Scene Text Recognition (STR) involves deciphering textual content embedded within complex, natural scene images, often following detection stages or integrated into end-to-end pipelines. Addressing the challenge of STR in noisy target domains, characterized by inter-domain and intra-domain noise, cluttered backgrounds, and irregular text shapes, this study proposes a robust and understandable framework titled Fusion-Based Adaptation for Scene Text Recognition (FASTR). The framework integrates a primary classifier with an epistemically aware auxiliary classifier to model uncertainty, supported by a novel Adaptive Scale Feature Module (ASFM) that enhances localisation through pixel-level mask prediction and multi-scale fusion. A Triple-Level Confidence (TLC) strategy—categorized into high, medium, and low consistency thresholds—is introduced to enforce consistency loss and improve generalisation across domains. Additionally, a pseudo-labelling scheme refines the adaptation process through self-training under structured domain noise. FASTR is trained and evaluated on both synthetic (SynthText, MJSynth) and real-world (ICDAR 2013, SVT, and IIIT5K) datasets. It achieves a word recognition accuracy of 92.4% on IIIT5K, 89.7% on SVT, and 93.1% on ICDAR 2013, outperforming state-of-the-art baselines by an average margin of 2.8%. On cross-domain benchmarks with added noise, FASTR maintains high performance, achieving an average F1-score of 90.5%, with precision and recall values of 91.2% and 89.9%, respectively. Hyperparameters, training configurations, and evaluation metrics are transparently documented to ensure reproducibility. The findings demonstrate superior scale robustness, effective domain adaptation, and resilience to cluttered backgrounds, with explainability preserved through interpretable confidence maps and visual cues.