Novel Activation Sparsification Approach for Large Language Models
摘要
Large Language Models (LLMs) require a lot of computational resources for inference. That is why the latest advancements in hardware design may offer many possibilities for speeding the LLM up. For example, TPU optimize calculations on data, transformed into the Coordinate sparse tensor format. The SparseCore processing unit that performs the calculations is heavily tailored for the extremely sparse embeddings of Deep Learning Recommendation Models. The other example of the enhanced hardware is Sparse Tensor Cores, that offer support for