Integrating Contextual Attention with Deep Neural Layers Enhances Vietnamese Text Identification Accuracy in Outdoor Images and Videos
摘要
This study introduces a robust technique for identifying Vietnamese text in outdoor images and videos by integrating contextual attention mechanisms with deep neural network layers. The method involves three key steps: (i) Feature extraction using the Resnet-50 backbone network; (ii) Feature aggregation using the Contextual Attention Architecture, and (iii) Text segmentation using the Fully Convolutional Network. The proposed approach was evaluated on real-world scene text image datasets, and experimental results indicate that it can accurately identify text in diverse shapes with high stability and precision.