K-Bloom: unleashing the power of pre-trained language models in extracting knowledge graph with predefined relations
摘要
Pre-trained language models have become popular in natural language processing tasks, but their inner workings and knowledge acquisition processes remain unclear. To address this issue, we introduce K-Bloom—a refined search-and-score mechanism tailored for seed-guided exploration in pre-trained language models, ensuring both accuracy and efficiency in extracting relevant entity pairs and relationships. Specifically, our crawling procedure is divided into two sub-tasks. Using a few seed entity pairs to minimize the need for extensive manual effort or predefined knowledge, we expand the knowledge graph with new entity pairs around these seeds. To evaluate the effectiveness of our proposed model, we conducted experiments on two datasets that cover the general domain. Our resulting knowledge graphs serve as symbolic representations of the source pre-trained language models, providing valuable insights into their knowledge capacities. Additionally, they enhance our understanding of the pre-trained language models’ capabilities when automatically evaluated on large language models. The experimental results demonstrate that our method outperforms the baseline approach by up to 5.62% in terms of accuracy in various settings of the two benchmarks. We believe that our approach offers a scalable and flexible solution for knowledge graph construction and can be applied to different domains and novel contexts.