The applicability of large language models (LLMs) in real-world tasks is largely based on instruction tuning (IT) and in-context learning (ICL). Although IT has demonstrated strong performance across tasks, ICL provides a fast alternative for task adaptation by leveraging examples without requiring explicit gradient updates. However, explicit comparisons between IT and ICL in LLM have not been explored sufficiently. This study evaluates the classification performance of six LLMs employing IT and ICL in five computational social science datasets in a few-shot setting. Experimental results reveal that ICL consistently outperforms IT in the majority of datasets. Further investigation into the impact of training sample size and sample selection strategies reveals that: indiscriminate sample augmentation does not consistently improve LLM performance in ICL or IT, and optimizing in-context example selection strategies proves more effective than increasing sample quantity. We also compare three prompting strategies, demonstrating that ICL is more effective than zero-shot prompting and Chain-of-Thought (CoT).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

In-Context Learning vs Instruction Tuning: A Systematic Comparison of Few-Shot Adaptation Strategies for Large Language Models

  • Taihang Wang,
  • Xiaoman Xu,
  • Yimin Wang,
  • Ye Jiang

摘要

The applicability of large language models (LLMs) in real-world tasks is largely based on instruction tuning (IT) and in-context learning (ICL). Although IT has demonstrated strong performance across tasks, ICL provides a fast alternative for task adaptation by leveraging examples without requiring explicit gradient updates. However, explicit comparisons between IT and ICL in LLM have not been explored sufficiently. This study evaluates the classification performance of six LLMs employing IT and ICL in five computational social science datasets in a few-shot setting. Experimental results reveal that ICL consistently outperforms IT in the majority of datasets. Further investigation into the impact of training sample size and sample selection strategies reveals that: indiscriminate sample augmentation does not consistently improve LLM performance in ICL or IT, and optimizing in-context example selection strategies proves more effective than increasing sample quantity. We also compare three prompting strategies, demonstrating that ICL is more effective than zero-shot prompting and Chain-of-Thought (CoT).