This study introduces a robust technique for identifying Vietnamese text in outdoor images and videos by integrating contextual attention mechanisms with deep neural network layers. The method involves three key steps: (i) Feature extraction using the Resnet-50 backbone network; (ii) Feature aggregation using the Contextual Attention Architecture, and (iii) Text segmentation using the Fully Convolutional Network. The proposed approach was evaluated on real-world scene text image datasets, and experimental results indicate that it can accurately identify text in diverse shapes with high stability and precision.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating Contextual Attention with Deep Neural Layers Enhances Vietnamese Text Identification Accuracy in Outdoor Images and Videos

  • Thi-Thanh-Tan Nguyen

摘要

This study introduces a robust technique for identifying Vietnamese text in outdoor images and videos by integrating contextual attention mechanisms with deep neural network layers. The method involves three key steps: (i) Feature extraction using the Resnet-50 backbone network; (ii) Feature aggregation using the Contextual Attention Architecture, and (iii) Text segmentation using the Fully Convolutional Network. The proposed approach was evaluated on real-world scene text image datasets, and experimental results indicate that it can accurately identify text in diverse shapes with high stability and precision.