Extracting patent effect words from patents is a crucial component of the patent analysis process, as these words play an essential role in assessing the technological efficacy and innovation potential of patents. Although leveraging large language models (LLMs) for this purpose is currently the mainstream approach, the significant computational resources and costs associated with such models motivate the exploration of smaller, open-source LLMs for similar tasks. This paper presents a knowledge distillation approach, using a dataset of zine battery patent, employing GPT-4 (GPT-4 specifically refers to gpt-4-0125-preview model in this paper.) as the teacher model, Qwen2.5-7B and Qwen2.5-7B-Instruct as the student model. We subsequently compare the ROUGE scores and BERT-Score of teacher model, student model, and the knowledge-distilled student model. The results indicate that the knowledge-distilled model outperforms the original student model in extracting patent effect words. In some scenarios, the knowledge-distilled student models can match or even exceed the performance of the teacher model, while having fewer parameters, lower costs, and faster inference speeds.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Patent Information Extraction Based on Teacher Student Model - A Case Study of Zinc Battery Patent Dataset

  • Lingchen Cai,
  • Chunjiang Liu,
  • Shirui Yu,
  • Yiran Cai,
  • Haiyun Xu,
  • Kaidi Wang

摘要

Extracting patent effect words from patents is a crucial component of the patent analysis process, as these words play an essential role in assessing the technological efficacy and innovation potential of patents. Although leveraging large language models (LLMs) for this purpose is currently the mainstream approach, the significant computational resources and costs associated with such models motivate the exploration of smaller, open-source LLMs for similar tasks. This paper presents a knowledge distillation approach, using a dataset of zine battery patent, employing GPT-4 (GPT-4 specifically refers to gpt-4-0125-preview model in this paper.) as the teacher model, Qwen2.5-7B and Qwen2.5-7B-Instruct as the student model. We subsequently compare the ROUGE scores and BERT-Score of teacher model, student model, and the knowledge-distilled student model. The results indicate that the knowledge-distilled model outperforms the original student model in extracting patent effect words. In some scenarios, the knowledge-distilled student models can match or even exceed the performance of the teacher model, while having fewer parameters, lower costs, and faster inference speeds.