<p>High-dimensional token embeddings underpin generative language models, as they can capture subtle semantic information and significantly enhance the modelling of complex language patterns. However, this high dimensionality also introduces considerable model parameters and prohibitively high model storage and memory requirements, which are particularly unaffordable for low-end devices. Targeting no extra training data and insufficient computation cases, we propose a training-free model compression approach based on the tensor-train decomposition (TTD), whereby each pre-trained token embedding is converted into a lower-dimensional matrix product state (MPS). We investigate what language capabilities are preserved under training-free compression at different compression ratios, providing insights into the distinct redundancy structures captured by tensor-based versus pruning-based compression methods. We then comprehensively investigate the low-rank structures extracted by this approach, in terms of the compression ratio, the language task performance, and latency on a typical low-end device (i.e., Raspberry Pi). Our approach trades increased inference latency (no more than 0.5 ms/token reconstruction overhead on Raspberry Pi) for substantial memory and storage reduction, making it suitable for deployment scenarios where memory and storage are the primary bottleneck. Taking GPT family, OPT models, Qwen (2.5–0.5 B and 3–0.6 B) and MiniCPM4-0.5B as case studies, our approach for the embedding layer compression consistently achieves a compression factor 0.5× – 2.0×. The extension of our approach for dense layers (feed-forward layers and attention layers) compression, can improve the model language task performance in zero-shot reasoning tasks. Our performance analysis of different tasks reveals that tensor decomposition preserves higher-level logical reasoning capabilities (e.g., BoolQ, ARC-Challenge), while pruning-based methods like SliceGPT maintain advantages for tasks requiring broad lower-level lexical feature coverage (e.g., HellaSwag, WinoGrande), demonstrating that different compression approaches preserve complementary linguistic capabilities.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Surviving Resource-constraint Compression: Capability Retention under Tensor-train Decomposition for Sub-billion Parameter Language Models

  • Mingxue Xu,
  • Yao Lei Xu,
  • Danilo P. Mandic

摘要

High-dimensional token embeddings underpin generative language models, as they can capture subtle semantic information and significantly enhance the modelling of complex language patterns. However, this high dimensionality also introduces considerable model parameters and prohibitively high model storage and memory requirements, which are particularly unaffordable for low-end devices. Targeting no extra training data and insufficient computation cases, we propose a training-free model compression approach based on the tensor-train decomposition (TTD), whereby each pre-trained token embedding is converted into a lower-dimensional matrix product state (MPS). We investigate what language capabilities are preserved under training-free compression at different compression ratios, providing insights into the distinct redundancy structures captured by tensor-based versus pruning-based compression methods. We then comprehensively investigate the low-rank structures extracted by this approach, in terms of the compression ratio, the language task performance, and latency on a typical low-end device (i.e., Raspberry Pi). Our approach trades increased inference latency (no more than 0.5 ms/token reconstruction overhead on Raspberry Pi) for substantial memory and storage reduction, making it suitable for deployment scenarios where memory and storage are the primary bottleneck. Taking GPT family, OPT models, Qwen (2.5–0.5 B and 3–0.6 B) and MiniCPM4-0.5B as case studies, our approach for the embedding layer compression consistently achieves a compression factor 0.5× – 2.0×. The extension of our approach for dense layers (feed-forward layers and attention layers) compression, can improve the model language task performance in zero-shot reasoning tasks. Our performance analysis of different tasks reveals that tensor decomposition preserves higher-level logical reasoning capabilities (e.g., BoolQ, ARC-Challenge), while pruning-based methods like SliceGPT maintain advantages for tasks requiring broad lower-level lexical feature coverage (e.g., HellaSwag, WinoGrande), demonstrating that different compression approaches preserve complementary linguistic capabilities.