Evaluating Vision Language Models for Handwritten Text Recognition
摘要
Vision Language Models (VLMs) promise a new generation of end-to-end handwritten text recognisers by unifying visual perception and linguistic reasoning. Motivated by their rapid progress but limited evidence on handwriting, we present a zero-shot evaluation of five current representative VLMs on two public benchmarks that cover both contemporary English and early-modern Spanish manuscripts. Using tailored prompts—without any task-specific fine-tuning—we steer each model toward line-level transcription and assess accuracy with the standard Word and Character Error Rates. Our study reveals that language, script style and historical spelling substantially challenge today’s models; nonetheless, state-of-the-art proprietary architectures already approach the reliability of state of the art dedicated recognisers, while compact open-source variants yield competitive results that could be boosted with light adaptation. These findings suggest the viability of VLMs as a general and easy end to end strategy for handwritten text recognition and highlight the need for targeted fine-tuning, richer historical data and layout-aware prompting to unlock their full potential in large-scale manuscript transcription projects.