A novel NLP-driven approach for enriching artefact descriptions, provenance, and entities in cultural heritage
摘要
Despite the availability of numerous open datasets on cultural heritage, limited research has focussed on structuring and normalising this type of data, particularly through the extraction of entities from unstructured texts. This step is crucial for enriching, analysing, and understanding these complex datasets. This study presents a procedure designed to streamline the creation of domain-specific datasets for training natural language processing models and evaluates their performance across three distinct datasets generated using this procedure. A zero-shot learning model, the Generalist and Lightweight Model for Named Entity Recognition, was assessed alongside pre-trained spaCy models on three datasets created in the framework of the European Union-funded Research Intelligence Technology for Heritage and Market Security project: one containing provenance information on artefacts from North American museums, another detailing stolen cultural goods in Romania, and a third with structured yet unclassified data on WWII-looted Polish art. Further training of spaCy models on these newly defined datasets revealed that fine-tuned models significantly outperform their non-fine-tuned counterparts, with the best results from the Transformer model fine-tuned on provenance data. This success can be largely attributed to the standardised conventions in provenance research. In contrast, the model fine-tuned on descriptive information performed poorly, likely due to extensive descriptions containing non-essential data that increased model uncertainty. This work highlights the potential of automating entity extraction to build knowledge graphs for cultural object databases, enabling advanced analytical approaches such as Network Analysis.