While the potential of in-memory RRAM computation for achieving energy-efficient NNs is recognized, concerns persist about its relative scalability to support modern NNs with billions of parameters. In this context, this paper presents GLoRia, a GPU-RRAM architecture and associated software stack to handle these limitations. We strategically identify the optimal NN layers for RRAM acceleration, enhancing the scalability of RRAMs for complex NN architectures and reducing energy consumption. We validate our approach using practical large CNN and GPT models, showing a 6.4 \(\times \) decrease in energy consumption, without compromising inference accuracy, thanks to the proposed strategy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GLoRia: An Energy-Efficient GPU-RRAM System Stack for Large Neural Networks

  • Rafael Fão de Moura,
  • Michael Jordan,
  • Luigi Carro

摘要

While the potential of in-memory RRAM computation for achieving energy-efficient NNs is recognized, concerns persist about its relative scalability to support modern NNs with billions of parameters. In this context, this paper presents GLoRia, a GPU-RRAM architecture and associated software stack to handle these limitations. We strategically identify the optimal NN layers for RRAM acceleration, enhancing the scalability of RRAMs for complex NN architectures and reducing energy consumption. We validate our approach using practical large CNN and GPT models, showing a 6.4 \(\times \) decrease in energy consumption, without compromising inference accuracy, thanks to the proposed strategy.