Robot Audition
摘要
This chapter presents an integrated approach to accelerating robot audition technologies—Automatic Speech Recognition (ASR), Sound Source Localization (SSL), and Sound Source Separation (SSS)—on edge computing platforms using GPUs and FPGAs. For ASR, we introduce CASENet, a lightweight CNN architecture optimized for speech command classification, achieving state-of-the-art accuracy with reduced model size and computational cost. We also implement a CNN accelerator on an SoC-based edge server to further improve inference latency. For SSL and SSS, we target HARK, a widely used robot audition framework, and explore its efficient realization using parallel computing strategies on both GPU and FPGA platforms. The GPU-based implementation delivers real-time processing capability with large-scale microphone arrays, while the FPGA-based solution offers energy-efficient inference suitable for edge environments. Experimental results demonstrate significant reductions in latency and power consumption while maintaining high calculation accuracy, paving the way for robust and practical robot audition systems in diverse application scenarios.