This paper introduces py-amr2fred, a Python library that converts natural language text into OWL-compliant RDF Knowledge Graphs (KGs) through an advanced pipeline integrating Large Language Models, Abstract Meaning Representation (AMR) parsing and semantic enrichment. Designed for scalability and flexibility, the library addresses limitations in existing solutions and facilitates seamless integration into diverse applications. To demonstrate its effectiveness, we present MusicBO, a domain-specific KG capturing the historical, cultural, and relational aspects of musical heritage. Constructed from a multilingual corpus, MusicBO leverages the pipeline’s capabilities for text processing, AMR parsing, RDF transformation, and quality assurance via back-translation validation. The resulting graph, comprising over 531,000 triples, is publicly accessible and serves as a resource for education, research, and digital storytelling in cultural heritage. Additionally, we propose an intrinsic evaluation method for the quality assessment of the generated KGs, leveraging Open Knowledge Extraction motifs. A manually curated benchmark dataset complements this evaluation framework, providing a valuable resource for future research in text-to-KG construction. The contributions of this work underscore the potential of py-amr2fred in advancing automated, scalable, and domain-independent KG generation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

py-amr2fred: A Python Library for Converting Text into OWL-Compliant RDF KGs

  • Aldo Gangemi,
  • Arianna Graciotti,
  • Antonello Meloni,
  • Andrea G. Nuzzolese,
  • Valentina Presutti,
  • Diego Reforgiato Recupero,
  • Alessandro Russo

摘要

This paper introduces py-amr2fred, a Python library that converts natural language text into OWL-compliant RDF Knowledge Graphs (KGs) through an advanced pipeline integrating Large Language Models, Abstract Meaning Representation (AMR) parsing and semantic enrichment. Designed for scalability and flexibility, the library addresses limitations in existing solutions and facilitates seamless integration into diverse applications. To demonstrate its effectiveness, we present MusicBO, a domain-specific KG capturing the historical, cultural, and relational aspects of musical heritage. Constructed from a multilingual corpus, MusicBO leverages the pipeline’s capabilities for text processing, AMR parsing, RDF transformation, and quality assurance via back-translation validation. The resulting graph, comprising over 531,000 triples, is publicly accessible and serves as a resource for education, research, and digital storytelling in cultural heritage. Additionally, we propose an intrinsic evaluation method for the quality assessment of the generated KGs, leveraging Open Knowledge Extraction motifs. A manually curated benchmark dataset complements this evaluation framework, providing a valuable resource for future research in text-to-KG construction. The contributions of this work underscore the potential of py-amr2fred in advancing automated, scalable, and domain-independent KG generation.