Handwriting recognition enables the automatic transcription of large volumes of digitized collections, providing access to the content. However, regardless of the system used, some recognition errors still occur. With the advancement of Large Language Models (LLMs), the question arises whether these models can improve handwriting recognition as a post-processing step. We have developed a method for LLM-based post-correction and evaluated it on three benchmark datasets, namely Washington, Bentham, and IAM. We consistently achieved a character error rate reduction of up to 30%, though we observed significant variability depending on the prompt and the LLM used.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Post-correction of Handwriting Recognition Using Large Language Models

  • Jean Pool Pereyra Principe,
  • Andreas Fischer,
  • Anna Scius-Bertrand

摘要

Handwriting recognition enables the automatic transcription of large volumes of digitized collections, providing access to the content. However, regardless of the system used, some recognition errors still occur. With the advancement of Large Language Models (LLMs), the question arises whether these models can improve handwriting recognition as a post-processing step. We have developed a method for LLM-based post-correction and evaluated it on three benchmark datasets, namely Washington, Bentham, and IAM. We consistently achieved a character error rate reduction of up to 30%, though we observed significant variability depending on the prompt and the LLM used.