<p>Airborne sound source classification has become critical for enhancing situational awareness, surveillance, environmental monitoring, and safety applications. This study proposes a low-resource Lightweight Airborne Acoustic Transformer (LAiT) designed for real-time classification of drones, helicopters, birds, and environmental sounds using minimal computational overhead. The architecture integrates inverted residual blocks with a compact multi-head self-attention mechanism, evaluated to capture both local and global acoustic patterns efficiently. LAiT with four attention head configuration achieves the highest performance among all lightweight baselines, surpassing the ultra-lightweight neural network (ULNN) by 1.18% in accuracy while requiring only 0.23&#xa0;M parameters and a reduced memory footprint of less than 1&#xa0;MB. The model exhibits near-linear computational scaling across the attention head and depth variation and is compatible with low-power ARM-class devices even under reduced input configuration. The model shows as inference times below 1&#xa0;s per audio frame, nearly 0.27–1.05 × faster compared to existing lightweight audio transformer models. The model was found to operate stably even under multiple nonstationary noise-augmented conditions&#xa0;with signal--to-noise ratio between 5-15 dB taken from TAU Urban Acoustic Scenes 2022 and NOISEX92 recordings. LAiT maintains superior robustness compared to models of similar size, reflecting strong generalization in challenging environments with suitability to edge computing applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Low resource multi-head self-attention based transformer for airborne acoustic classification system

  • Anuj Kumar Mishra,
  • Ripul Ghosh

摘要

Airborne sound source classification has become critical for enhancing situational awareness, surveillance, environmental monitoring, and safety applications. This study proposes a low-resource Lightweight Airborne Acoustic Transformer (LAiT) designed for real-time classification of drones, helicopters, birds, and environmental sounds using minimal computational overhead. The architecture integrates inverted residual blocks with a compact multi-head self-attention mechanism, evaluated to capture both local and global acoustic patterns efficiently. LAiT with four attention head configuration achieves the highest performance among all lightweight baselines, surpassing the ultra-lightweight neural network (ULNN) by 1.18% in accuracy while requiring only 0.23 M parameters and a reduced memory footprint of less than 1 MB. The model exhibits near-linear computational scaling across the attention head and depth variation and is compatible with low-power ARM-class devices even under reduced input configuration. The model shows as inference times below 1 s per audio frame, nearly 0.27–1.05 × faster compared to existing lightweight audio transformer models. The model was found to operate stably even under multiple nonstationary noise-augmented conditions with signal--to-noise ratio between 5-15 dB taken from TAU Urban Acoustic Scenes 2022 and NOISEX92 recordings. LAiT maintains superior robustness compared to models of similar size, reflecting strong generalization in challenging environments with suitability to edge computing applications.