Bin Xiao, Junjie Yan · Applied Sciences 2026 · 2026
DOI: 10.3390/app16199624
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Small object detection (SOD) in optical remote sensing images (RSI) is essential for aerial perception yet remains challenged by the severe degradation of fine-grained spatial features during downsampling and the high sensitivity of Intersection over Union (IoU) metrics to tiny positional shifts. To address these limitations, we propose YOLO-SPN, an efficient detection network based on YOLOv11 that couples spatial-detail-preserving downsampling, scale adaptation, and stable gradient regression within a single optimization pipeline. Specifically, a Spatial-to-Depth Convolution (SPDConv) module reconstructs shallow-layer downsampling to retain high-frequency features; a high-resolution P2 detection head provides a dedicated shallow receptive field; and a Normalized Wasserstein Distance (NWD) loss models bounding boxes as 2D Gaussian distributions to supply continuous gradients. Experiments on DIOR, AI-TOD, and VisDrone raise mAP50 over the YOLOv11 baseline by 3.0, 0.7, and 2.0 percentage points, respectively, while requiring only 2.66M parameters and running at 136 FPS on an NVIDIA A100 GPU at 640×640 input resolution. The results indicate a favorable accuracy–efficiency trade-off for remote sensing perception under a tight parameter budget.
No comments yet — start the discussion below.