This chapter presents an integrated approach to accelerating robot audition technologies—Automatic Speech Recognition (ASR), Sound Source Localization (SSL), and Sound Source Separation (SSS)—on edge computing platforms using GPUs and FPGAs. For ASR, we introduce CASENet, a lightweight CNN architecture optimized for speech command classification, achieving state-of-the-art accuracy with reduced model size and computational cost. We also implement a CNN accelerator on an SoC-based edge server to further improve inference latency. For SSL and SSS, we target HARK, a widely used robot audition framework, and explore its efficient realization using parallel computing strategies on both GPU and FPGA platforms. The GPU-based implementation delivers real-time processing capability with large-scale microphone arrays, while the FPGA-based solution offers energy-efficient inference suitable for edge environments. Experimental results demonstrate significant reductions in latency and power consumption while maintaining high calculation accuracy, paving the way for robust and practical robot audition systems in diverse application scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robot Audition

  • Lin Zirui,
  • Haris Gulzar,
  • Kazuhiro Nakadai,
  • Hideharu Amano

摘要

This chapter presents an integrated approach to accelerating robot audition technologies—Automatic Speech Recognition (ASR), Sound Source Localization (SSL), and Sound Source Separation (SSS)—on edge computing platforms using GPUs and FPGAs. For ASR, we introduce CASENet, a lightweight CNN architecture optimized for speech command classification, achieving state-of-the-art accuracy with reduced model size and computational cost. We also implement a CNN accelerator on an SoC-based edge server to further improve inference latency. For SSL and SSS, we target HARK, a widely used robot audition framework, and explore its efficient realization using parallel computing strategies on both GPU and FPGA platforms. The GPU-based implementation delivers real-time processing capability with large-scale microphone arrays, while the FPGA-based solution offers energy-efficient inference suitable for edge environments. Experimental results demonstrate significant reductions in latency and power consumption while maintaining high calculation accuracy, paving the way for robust and practical robot audition systems in diverse application scenarios.