Multi-source information fusion using CNN-LSTM-Attention for bone layer recognition in robotic orthopedic grinding
摘要
Epiphyseal opening requires precise localization for bony bridge resection. Traditional surgery is challenged by unclear bony bridge boundary localization and inaccurate grinding precision, while existing robot-assisted operations predominantly focus on pre-operative localization with limited intraoperative autonomous decision-making capabilities. To address this, we propose a multi-source information fusion framework using a CNN-LSTM-Attention network for real-time bone layer differentiation, specifically identifying idling, cancellous, and cortical bone states, during robotic orthopedic grinding. First, the mapping relationship between acceleration, force and acoustic signals and bone density is analyzed, serving as the basis of bone layer recognition. Second, a three-channel parallel late-feature-fusion CNN-LSTM-Attention network is established. The dataset was constructed from multiple independent grinding trials with trial-wise splitting to prevent data leakage under representative robotic grinding conditions. Over five independent runs, the proposed method achieves a test accuracy of 95.06% ± 0.53%, significantly outperforming comparison models including CNN, CNN-Attention, CNN-LSTM, and non-deep-learning baselines (Random Forest, Extra Trees, SVM, KNN). Ablation studies isolating CNN, LSTM, and Attention contributions are provided, and the learned Squeeze-and-Excitation attention weights are visualized to confirm dynamic cross-modal feature weighting. With approximately 0.32 million parameters, the model introduces an inference latency of 2.8 ms on a desktop CPU and 1.2 ms on an NVIDIA Jetson Orin, well below the 50 ms robot control cycle, confirming real-time feasibility. Additionally, we explore single-signal, dual-signal, and triple-signal fusion settings, demonstrating that tri-modal fusion achieves the best performance.
Graphical Abstract