<p>Large language models (LLMs) have gained significant popularity, with various models demonstrating different domains and intelligence. Deep learning methods, especially neural networks, are used to build LLMs, and their training requires enormous volumes of data. They have significantly advanced natural language processing (NLP), enabling applications like chatbots, virtual assistants, language translation, and content generation. However, training LLMs is computationally intensive due to large model sizes, extensive datasets, iterative processes, specialized hardware, and high energy consumption. To address these challenges, quantization has been introduced. This process reduces the precision of numerical values, such as weights and activations, thereby decreasing memory and computational requirements. But this technique can affect model performance negatively, therefore, recent research focuses on minimizing accuracy loss. Techniques like mixed precision training and adaptive quantization have been developed to balance efficiency and performance. This paper surveys the existing 1-bit quantization approaches, providing insights and recommendations for future work. The goal is to enable the efficient and cost-effective deployment of LLMs without compromising their performance, broadening their accessibility and applicability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A survey on 1-bit quantized large language models

  • Kritika Tripathi,
  • Devanshi Malik,
  • Abhi Akshat,
  • Kusum Lata

摘要

Large language models (LLMs) have gained significant popularity, with various models demonstrating different domains and intelligence. Deep learning methods, especially neural networks, are used to build LLMs, and their training requires enormous volumes of data. They have significantly advanced natural language processing (NLP), enabling applications like chatbots, virtual assistants, language translation, and content generation. However, training LLMs is computationally intensive due to large model sizes, extensive datasets, iterative processes, specialized hardware, and high energy consumption. To address these challenges, quantization has been introduced. This process reduces the precision of numerical values, such as weights and activations, thereby decreasing memory and computational requirements. But this technique can affect model performance negatively, therefore, recent research focuses on minimizing accuracy loss. Techniques like mixed precision training and adaptive quantization have been developed to balance efficiency and performance. This paper surveys the existing 1-bit quantization approaches, providing insights and recommendations for future work. The goal is to enable the efficient and cost-effective deployment of LLMs without compromising their performance, broadening their accessibility and applicability.