Jiajun Xu, Fuming Qu, Zhanliang Niu, Yaming Ji, Lingyu Zhao, Weihua Zhou · Processes 2026 · 2026
DOI: 10.3390/pr14193171
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Reliable personnel detection in underground mines is essential for safe intelligent mining, but uneven illumination, dust occlusion, thermal interference, and cross-modal parallax can degrade RGB-T perception. We propose an Environment-Driven Fusion and Optimization Network (EnvFONet), which uses environmental degradation as a unified prior for feature representation, multimodal fusion, and training optimization. First, Dynamic Receptive Field Relaxation (DRFR) adaptively interpolates original and locally smoothed thermal features to balance fine-detail preservation and contextual robustness. Second, an Env-AdaIN-guided spatially gated fusion module applies environment-conditioned residual affine modulation and high-frequency pixel-level weighting to reduce cross-modal statistical drift and parallax-induced boundary artifacts. Third, Environment-Adaptive Relaxation Loss (EARL) adjusts regression supervision for difficult matched samples according to environmental degradation and localization quality. On the self-collected Anshan underground iron-ore mine dataset, EnvFONet achieves 96.8% mAP50, 63.9% mAP75, and 62.8% mAP50:95. On the public LLVIP benchmark, it achieves 94.4% mAP50 and 59.6% mAP50:95. These results show that environment-conditioned fusion and optimization improve RGB-T personnel detection robustness under the evaluated degraded conditions.
No comments yet — start the discussion below.