UFLA-FORMS: an academic forms dataset for information extraction in the Portuguese language
摘要
Information Extraction aims to analyze and extract relevant information in document samples. For visual documents, such as academic and commercial forms, key-value pair extraction is capable of extracting and grouping the requested information automatically. The state of the art presents multimodal models capable of analyzing document’s image, text and layout. To study and use these models in downstream tasks, labeled datasets must be available for fine-tuning of models. However, publicly available data is scarce, even more so for the Portuguese language. In this context, this work aims to provide a dataset of academic forms, named UFLA-FORMS, for the task of Information Extraction. The dataset is composed of 200 manually labeled samples, containing 7710 entities with 4442 relationships pairs between them. The labeling obtained a Kappa coefficient of agreement of 0.918 for the labeling of entities and 0.909 for the relationships attributed between them. The dataset was experimentally evaluated through cross-validation with hyper-parameter search in Named Entity Recognition and Relation Extraction tasks, obtaining, respectively, an average