Reweaving the Threads of Korean History: AI-Driven Restoration of the Daegu-bu Household Registers (1681–1876)
摘要
In this study, we have applied advanced masked language models (MLMs)—BERT, DistilBERT, ELECTRA, and RoBERTa—to infer missing and misinterpreted values in comprehensive family register data. Our data compiles Daegu-bu household register books, triennially published from 1392 to 1910 in Joseon (premodern Korea) for listing the demographic characteristics of taxpayers. MLM models primarily detect and infer transcription errors and blanks caused by the historical deterioration of original books. The results show that RoBERTa outperforms the other models using byte-pair encoding tokenization. The accuracy rates of non-RoBERTa models have been found to deteriorate with an increase in the number of variables. The performance of RoBERTa in inferring morphologically rich and sparse data extends the frontiers of historical research by reconstructing primary sources for understanding the past.