The advent of smart devices and applications has been transformative, reshaping how we interact with technology in daily life. All these smart devices incorporate training the machines through different algorithms so that customized output can be delivered for each sequence of input events. These algorithms must be deployed faster to synchronize with the present world’s demands, reducing the training and pre-processing time. This paper explores the synergy between hardware and software in the context of machine learning, emphasizing the integration of dedicated vector data processors into the versatile RISC-V architecture. Through customization of core designs and implementation of the AXI protocol as a reliable interconnect, a comprehensive System on Chip (SoC) is developed to enhance the efficiency of pre-processing stages in machine learning algorithms. The integration of a specialized co-processor not only alleviates the burden on the primary unit but also enhances processing speed. Experimentation across diverse input image configurations, including matrices of varying dimensions, has enabled the determination of an optimal core count to balance hardware demand and processing efficiency.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development of a System on Chip (SoC) for Matrix Multiplication Utilizing RISC-V and Vector Processor

  • Yash Purohit,
  • Devesh Pareek,
  • Vijay Savani

摘要

The advent of smart devices and applications has been transformative, reshaping how we interact with technology in daily life. All these smart devices incorporate training the machines through different algorithms so that customized output can be delivered for each sequence of input events. These algorithms must be deployed faster to synchronize with the present world’s demands, reducing the training and pre-processing time. This paper explores the synergy between hardware and software in the context of machine learning, emphasizing the integration of dedicated vector data processors into the versatile RISC-V architecture. Through customization of core designs and implementation of the AXI protocol as a reliable interconnect, a comprehensive System on Chip (SoC) is developed to enhance the efficiency of pre-processing stages in machine learning algorithms. The integration of a specialized co-processor not only alleviates the burden on the primary unit but also enhances processing speed. Experimentation across diverse input image configurations, including matrices of varying dimensions, has enabled the determination of an optimal core count to balance hardware demand and processing efficiency.