Postcorrection of Weak Transcriptions by Large Language Models in the Iterative Process of Handwritten Text Recognition
摘要
The problem of accelerating the construction of accurate editorial annotations for handwritten archival texts within an incremental training cycle based on weak transcription is considered. Unlike previously published results, this work is focused on integrating automatic postcorrection of weak transcriptions using large language models (LLMs). A protocol for applying LLMs at the line level is proposed and implemented in a few-shot setup with carefully designed prompts and strict output format control (preservation of prereform orthography, protection of proper names and numerals, prohibition of structural changes to lines). Experiments have been conducted on the corpus of diaries of A.V. Sukhovo-Kobylin. As the base recognition model, we use the line-level version of the vertical attention network (VAN). The results show that LLM postcorrection (exemplified by the ChatGPT-4o service) substantially improves the readability of weak transcriptions and significantly reduces the word error rate (in our experiments, by about –12 percentage points), without degrading the character error rate. Another service tested, DeepSeek-R1, has demonstrated less stable behavior. Practical prompt engineering and limitations (context length limits, risk of “hallucinations”) are discussed, and recommendations are provided for the safe integration of LLM postcorrection into an iterative annotation pipeline to reduce expert annotators’ workload and speed up the digitization of historical archives.