Exploring Active Learning Approaches in Treebank Development
摘要
Developing language treebanks is a challenging and time-consuming process, especially for under-resourced languages with limited data. This study proposes the use of active learning approaches to automate aspects of this process, aiming to reduce both annotation duration and cost. We propose practical active annotation schemes in which experts strategically select sentences for the training set, initiating a circular process involving annotation prediction, expert correction, and model retraining. To validate the feasibility of these schemes, we applied them to 300 annotated sentences from the newly created Pomak corpus (an under-resourced language), recently published in Universal Dependencies treebanks, as the result of our efforts. Through experiments involving a simple weighted summation of annotation errors, we identified an optimal strategy. This strategy resulted in a 69% reduction in total annotation duration and an associated 81% decrease in total corresponding cost compared to a typical manual annotation, demonstrating the effectiveness of this approach.