<p>Computing-in-memory (CIM) has been proposed to solve “memory wall” problem for its short data path and high energy efficiency. However, previous CIM methods are hard to maintain the energy efficiency benefits at various accumulation lengths, which is essential for vision transformer and large CNN models. This work presents a lightning-like hybrid CIM macro using (1) a lightning-like analog/digital hybrid structure to maintain energy efficiency and inference accuracy; (2) a digital 4:2 compressor-based adder tree along with a double-regularization training method to reduce area/power cost; and (3) an analog-storage quantizer circuit (ASQC) for a scalable accumulation length, with an output ratio of 1. A fabricated 22-nm 64-kB hybrid-domain SRAM-CIM macro achieved an energy efficiency of 60.8 TOPS/W, for INT8 MAC operations at different accumulation lengths of 128 to 2048.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A 22-nm 64-kB lightning-like hybrid computing-in-memory macro with a compressed adder tree and analog-storage quantizers for transformer and CNNs

  • An Guo,
  • Xi Chen,
  • Fangyuan Dong,
  • Jinwu Chen,
  • Zhihang Yuan,
  • Xing Hu,
  • Guangyu Sun,
  • Xiaomin Li,
  • Arindam Basu,
  • Jun Yang,
  • Xin Si

摘要

Computing-in-memory (CIM) has been proposed to solve “memory wall” problem for its short data path and high energy efficiency. However, previous CIM methods are hard to maintain the energy efficiency benefits at various accumulation lengths, which is essential for vision transformer and large CNN models. This work presents a lightning-like hybrid CIM macro using (1) a lightning-like analog/digital hybrid structure to maintain energy efficiency and inference accuracy; (2) a digital 4:2 compressor-based adder tree along with a double-regularization training method to reduce area/power cost; and (3) an analog-storage quantizer circuit (ASQC) for a scalable accumulation length, with an output ratio of 1. A fabricated 22-nm 64-kB hybrid-domain SRAM-CIM macro achieved an energy efficiency of 60.8 TOPS/W, for INT8 MAC operations at different accumulation lengths of 128 to 2048.