Foundational Language Models perform significantly better on downstream tasks in the biomedical domain upon being further pre-trained on extensive biomedical corpora, but this continual pre-training incurs heavy computational costs. Indeed, some of the most performant biomedical language models incur even more computing costs during domain-specific training than the entire training cost of the foundational models they are initialised from. In this paper, we argue that much of the extended pre-training is redundant, with models seemingly wasting valuable resources re-learning lexical and semantic patterns already well-represented in their foundational models such as BERT, T5 and GPT. Focusing on Masked Language Models, we introduce a novel domain-specific masking strategy that is designed to facilitate continual learning while minimizing the training cost. Using this approach, we train and present a BERT-based model that matches or surpasses traditionally trained biomedical language models in performance across several downstream classification tasks while incurring up to 11 times lower training costs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Choose Your Words Wisely: Domain-Adaptive Masking Makes Language Models Learn Faster

  • Vanshpreet S. Kohli,
  • Aaron Monis,
  • Radhika Mamidi

摘要

Foundational Language Models perform significantly better on downstream tasks in the biomedical domain upon being further pre-trained on extensive biomedical corpora, but this continual pre-training incurs heavy computational costs. Indeed, some of the most performant biomedical language models incur even more computing costs during domain-specific training than the entire training cost of the foundational models they are initialised from. In this paper, we argue that much of the extended pre-training is redundant, with models seemingly wasting valuable resources re-learning lexical and semantic patterns already well-represented in their foundational models such as BERT, T5 and GPT. Focusing on Masked Language Models, we introduce a novel domain-specific masking strategy that is designed to facilitate continual learning while minimizing the training cost. Using this approach, we train and present a BERT-based model that matches or surpasses traditionally trained biomedical language models in performance across several downstream classification tasks while incurring up to 11 times lower training costs.