Purpose <p>The goal of our work is to develop a multi-view validation framework for evaluating LLM-generated knowledge graph (KG) triples. The proposed approach aims to address the lack of established validation procedure in the context of LLM-supported KG construction.</p> Methods <p>The proposed framework evaluates the LLM-generated triples across three dimensions: semantic plausibility, ontology-grounded type compatibility, and structural importance. We demonstrate the performance for GPT-4 generated concept-specific (e.g., for medications, diagnosis, procedures) triples in the context of chronic kidney disease (CKD).</p> Results <p>The proposed approach consistently achieves high-quality results across evaluated GPT-4 generated triples, strong semantic plausibility (semantic score mean: 0.79), excellent type compatibility (type score mean: 0.84), and high structural importance of entities within the CKD knowledge domain (ResourceRank mean: 0.94).</p> Conclusion <p>The validation framework offers a reliable and scalable method for evaluating quality and validity of LLM-generated triples across three views: semantic plausibility, type compatibility, and structural importance. The framework demonstrates robust performance in filtering high-quality triples and lays a strong foundation for fast and reliable medical KG construction and validation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A multi-view validation framework for LLM-generated knowledge graphs of chronic kidney disease

  • Aditya Kumar,
  • Dilpreet Singh,
  • Mario Cypko,
  • Oliver Amft

摘要

Purpose

The goal of our work is to develop a multi-view validation framework for evaluating LLM-generated knowledge graph (KG) triples. The proposed approach aims to address the lack of established validation procedure in the context of LLM-supported KG construction.

Methods

The proposed framework evaluates the LLM-generated triples across three dimensions: semantic plausibility, ontology-grounded type compatibility, and structural importance. We demonstrate the performance for GPT-4 generated concept-specific (e.g., for medications, diagnosis, procedures) triples in the context of chronic kidney disease (CKD).

Results

The proposed approach consistently achieves high-quality results across evaluated GPT-4 generated triples, strong semantic plausibility (semantic score mean: 0.79), excellent type compatibility (type score mean: 0.84), and high structural importance of entities within the CKD knowledge domain (ResourceRank mean: 0.94).

Conclusion

The validation framework offers a reliable and scalable method for evaluating quality and validity of LLM-generated triples across three views: semantic plausibility, type compatibility, and structural importance. The framework demonstrates robust performance in filtering high-quality triples and lays a strong foundation for fast and reliable medical KG construction and validation.