Yang Zhong, Xiuzai ZHANG, Juanjuan Ji, Yunzhong Shen · Remote Sensing 2026 · 2026
DOI: 10.3390/rs18172997
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Small objects in UAV imagery are easily obscured by sensor noise, adverse illumination, and weather-like appearance degradation. To address this problem, this paper proposes YOLO-ROSS, a lightweight detector built on YOLOv11. Its two architectural contributions are a C3k2_DTAB feature-extraction block, which combines grouped channel self-attention and masked-window self-attention to protect weak local evidence while adding non-local context, and an AFPN-P2 neck, which preserves high-resolution geometry and progressively aligns shallow detail with deep semantics. For end-to-end deployment, the detector also adopts the rank-consistent one-to-many/one-to-one assignment used by YOLOv10; this adopted component is evaluated separately but is not claimed as a new label-assignment algorithm. On the mixed-corruption VisDrone benchmark, YOLO-ROSS obtains 0.309 mean mAP@0.5 over eight runs, an absolute improvement of 0.042 over YOLOv11n. Its all-class peak F1 is approximately 0.40, compared with 0.36 for YOLOv11n, although the optimal confidence threshold shifts from 0.146 to 0.186. A direct-transfer evaluation on SODA-D-Robustness provides an additional check under a combined driving-domain and appearance-corruption shift. These results support improved accuracy under the specified controlled RGB corruptions, while not establishing universal real-weather or cross-modal robustness.
No comments yet — start the discussion below.