Developing Datasets for Training OCR/HTR Models for the Late 19th Century Greek Texts
摘要
Historical documents are vital for preserving cultural information. To access their content, OCR and HTR technologies transcribe them into text. Challenges arise with handwritten texts due to varied writing styles. This study develops two corpora to train models for 19th-century Greek texts. The first dataset includes printed texts from the “Ellinomnimon” archive, and the second comprises handwritten documents from the Lasithi Demogerontia archives. Moreover, using the Transkribus platform, two models were iteratively refined, enhancing their ability to transcribe Greek historical documents from 1800 to 1870.