With the increasing maturity of deep learning technology, large language models have shown excellent performance in the field of natural language processing, but the performance of information extraction is yet to be further improved. In this paper, we use 7B Llama-2 as a base model training to obtain a large language model capable of using natural language to guide information extraction tasks, which can solve the challenges faced by traditional information extraction methods, and also innovatively use multi-task learning optimization to improve the performance of the large model. We performed large-scale pre-training and instruction tuning on the big model LLama-2. Based on the high-quality and richly typed training data automatically constructed by ChatGPT, remote supervision, and other algorithms, which contains a total of 1 million entities, relations, and events, we designed corresponding English templates for instruction tuning. Second, we performed supervised fine-tuning of the model using manually labeled high-quality training sets. To further improve the model performance, we innovatively adopt a multi-task learning optimization strategy—GradNorm, which can dynamically adjust the weights of different tasks, thus balancing the losses among tasks during the training process and reducing the overall training loss. After information extraction experiments, our model is compared with other models to test the performance of uniform information extraction for large models, and the experimental results show that our model performs well.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LLM for Uniform Information Extraction Using Multi-task Learning Optimization

  • Ying Li,
  • Zhen Tan,
  • Weidong Xiao

摘要

With the increasing maturity of deep learning technology, large language models have shown excellent performance in the field of natural language processing, but the performance of information extraction is yet to be further improved. In this paper, we use 7B Llama-2 as a base model training to obtain a large language model capable of using natural language to guide information extraction tasks, which can solve the challenges faced by traditional information extraction methods, and also innovatively use multi-task learning optimization to improve the performance of the large model. We performed large-scale pre-training and instruction tuning on the big model LLama-2. Based on the high-quality and richly typed training data automatically constructed by ChatGPT, remote supervision, and other algorithms, which contains a total of 1 million entities, relations, and events, we designed corresponding English templates for instruction tuning. Second, we performed supervised fine-tuning of the model using manually labeled high-quality training sets. To further improve the model performance, we innovatively adopt a multi-task learning optimization strategy—GradNorm, which can dynamically adjust the weights of different tasks, thus balancing the losses among tasks during the training process and reducing the overall training loss. After information extraction experiments, our model is compared with other models to test the performance of uniform information extraction for large models, and the experimental results show that our model performs well.