Expansion of the memory pyramid in the era of large models: compute-intensive compute-in-memory and memory-intensive compute-in-memory
摘要
The rise of large-scale models has significantly increased the demand for high computational power and throughput in artificial intelligence (AI) chips, presenting challenges for existing architectures. Compute-in-memory (CIM) has emerged as a promising solution to address these bottlenecks. This study redefines the traditional computing architecture pyramid by introducing the CIM pyramid, structured around different storage media. CIM architectures are classified into two categories: compute-intensive CIM, which leverages static random-access memory (SRAM) and embedded dynamic random access memory (eDRAM), and memory-intensive CIM, which utilizes DRAM and non-volatile memory (NVM). The research reviews and analyzes recent advancements in both compute-intensive and memory-intensive CIM, highlighting their development trends and challenges. Furthermore, the study suggests a hybrid, heterogeneous architecture that integrates both CIM types with traditional computing systems, aiming to address the diverse computational needs of large models in the future.