Safety-Oriented Voice Dispatch in Underground Mining: A Robust Conformer Framework with Semantic Consistency
摘要
To address the challenge of insufficient recognition accuracy for dispatch voice commands in complex noise environments within underground coal mines, this paper proposes a speech safety information parsing method based on an improved Conformer architecture. Overcoming the limitations of conventional end-to-end models in acoustic-semantic cross-modal alignment and multi-turn dialogue context modeling, we innovatively introduce a sentence-level consistency fusion mechanism to construct a hybrid Conformer encoder with CTC/Attention decoding framework. Through contrastive learning between the encoder’s global speech representations and decoder-generated textual semantic vectors, we design a cross-modal MSE loss function to enforce topological consistency in semantic space. Experimental results on the ZH_MINE_ASR dataset under -5 dB signal-to-noise ratio conditions demonstrate that the improved model reduces CER by 15.11% compared to the baseline Conformer architecture. The sentence-level constrained hybrid architecture achieves an additional 6.75% CER reduction over single CTC/Attention models. Cross-domain transfer experiments validate the model’s strong generalization capability in low-resource scenarios, showing a 18.74% relative CER reduction after domain-specific fine-tuning. This research provides an effective technical pathway for accurate interpretation of safety-critical operation commands, demonstrating particular advantages in key scenarios including hydraulic support initial pressure warnings and conveyor fault diagnostics, thereby significantly enhancing the reliability of underground coal mine safety management systems.