A High Accuracy Text CAPTCHA Recognition Approach Through Opertimized Vision Transformer
摘要
CAPTCHA (Completely Automated Public Turing test to tell Computers and Humans Apart) is an automated verification mechanism designed to distinguish between human visitors and automated systems as a vital component for cyber security. However, current CAPTCHA recognition technologies still struggle to keep pace with the sophistication CAPTCHA generation algorithms. This paper introduces a novel approach to CAPTCHA recognition, utilizing the Vision Transformer (ViT) and enhancing it with two key optimizations: the Permutation Visual Model, which improves the model’s spatial understanding, and Transfer Learning for the Vision Transformer, which accelerates the model’s adaptation to new tasks. The method was rigorously tested against a diverse set of practical data, simulating real-world scenarios. The results are promising, with an accuracy rate of over 95% across multiple websites, indicating a significant advancement in CAPTCHA recognition technology.