Adaptive Contextual Embedding for Robust Far-View Borehole Detection
摘要
In industrial quarrying and mining operations, precisely locating densely distributed small-scale boreholes from limited far-view imagery is essential for accurate explosive placement, optimal rock fragmentation, and prevention of dangerous misfires. Misidentified or undetected boreholes lead directly to inefficient blasting outcomes, elevated operational costs, and significant safety hazards for personnel and equipment. However, existing detection methods, including widely-used YOLO-based architectures, struggle with reliably identifying these boreholes due to their extremely small scale, dense arrangements, and limited distinctive visual features. Furthermore, practical constraints related to camera placement and quarry geometry severely restrict the availability of sufficient high-quality annotated data, exacerbating the challenge. To address these combined challenges of visual complexity and limited training data, we propose an adaptive embedding-based detection framework designed for efficient generalization under resource-constrained conditions. Our approach introduces three synergistic components, each leveraging exponential moving averages (EMA) for stable feature learning: (1) adaptive augmentation, dynamically adjusting to illumination and textural variability; (2) embedding stabilization, ensuring consistent spatial and temporal embedding representations despite limited visual distinctions; and (3) contextual refinement, enhancing discrimination of boreholes from visually similar background noise by integrating spatial context. By applying EMA consistently across these components, our model efficiently learns robust, stable embeddings, enabling effective generalization to previously unseen and challenging scenarios. Comprehensive experiments conducted on a proprietary quarry-site dataset demonstrate significant performance improvements over baseline YOLO architectures. Specifically, our integrated approach achieves a notable increase in mean Average Precision (mAP) from 61.5% to 74.9% compared to the baseline YOLOv11 model, clearly validating our method’s practical suitability for robust deployment under constrained real-world conditions.