Abstract <p>We present a baseline accuracy of classifying audio recordings of command words in Russian from a recent dataset RuSC using a three-layer convolutional spiking neural network of integrate-and-fire neurons. The network is obtained by transferring weights from a trained network of ReLU neurons of same topology, and then by adjusting neuron thresholds using a same-topology network with the ClipFloor activation function. In order to make the network prospectively deployable to neuromorphic processors, its synaptic weights are quantized to 8-bit integer. When the duration of presenting one input sample is 200 time steps of spiking network, the resulting performance is the f1-micro of 98<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\%\)</EquationSource> <!--BPhysMGU2570290Serenko-m1--> </InlineEquation>, which is just 1<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\%\)</EquationSource> <!--BPhysMGU2570290Serenko-m2--> </InlineEquation> lower than originally reported on that dataset with artificial neural networks. This result might be a starting point against which further spiking network solutions for keyword spotting in Russian could be compared.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Classifying Russian Speech Commands with a Hardware-Deployable Spiking Neural Network Transferred from an Artificial Neural Network

  • A. Serenko,
  • R. Rybka,
  • A. Naumov,
  • A. Gryaznov,
  • A. Sboev

摘要

Abstract

We present a baseline accuracy of classifying audio recordings of command words in Russian from a recent dataset RuSC using a three-layer convolutional spiking neural network of integrate-and-fire neurons. The network is obtained by transferring weights from a trained network of ReLU neurons of same topology, and then by adjusting neuron thresholds using a same-topology network with the ClipFloor activation function. In order to make the network prospectively deployable to neuromorphic processors, its synaptic weights are quantized to 8-bit integer. When the duration of presenting one input sample is 200 time steps of spiking network, the resulting performance is the f1-micro of 98 \(\%\) , which is just 1 \(\%\) lower than originally reported on that dataset with artificial neural networks. This result might be a starting point against which further spiking network solutions for keyword spotting in Russian could be compared.