Yonghan Dai, Suoyang Zhang, Yue Fei, Huimin Zhang, Nan Zhang, Xiaofeng Zhou · AI for Engineering 2026 · 2026
DOI: 10.3390/aieng1030012
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Ultra-low-resolution (ULR) thermal array sensors (e.g., 8×8) offer a privacy-preserving and cost-effective paradigm for intelligent edge perception. However, extracting robust spatial features is severely hindered by extreme spatial aliasing, background thermal interference, and environmental temperature drift. In this paper, we propose VMEMSTNet, an ultra-lightweight framework that tackles the ULR perception bottleneck through a thermodynamics-inspired cross-modal learning strategy. Rather than treating ULR recognition purely as a computer vision task, our approach integrates the underlying sensing physics. To bridge the resolution gap, we leverage a high-resolution visual modality strictly as offline cross-modal guidance. Specifically, a YOLOv8n-seg teacher provides spatial shape priors via Gaussian thermal diffusion—explicitly mimicking physical heat dissipation for spatial knowledge translation—while a MobileNetV2 teacher guides semantic probability distillation. Furthermore, a thermodynamics-inspired preprocessing pipeline fuses dynamic quantile-based normalization with absolute temperature difference (ΔT) to maintain contrast consistency across environmental drifts. Using hand gesture recognition as a representative spatial classification task, experimental results demonstrate that VMEMSTNet achieves 95.7% accuracy. Remarkably, it outperforms heavier generic edge models while utilizing only 20.0 k parameters and 0.135 MFLOPs. By effectively decoupling spatial semantics from thermodynamic noise, this work establishes a robust, hardware-aware deployment pathway for ULR thermal perception on resource-constrained microcontrollers.
No comments yet — start the discussion below.