Bowen Song, Qiaoli Yang · Measurement Science and Technology 2026 · 2026
DOI: 10.1088/1361-6501/aea4df
Measurement Science and TechnologyJournal182 h-indexCounts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
To improve small-object feature preservation and fusion in complex unmanned aerial vehicle (UAV) imagery, this paper proposes MSEV-DETR, a multi-stage detection framework based on the Real-Time Detection Transformer (RT-DETR). The framework consists of three successive stages. Firstly, Cross Stage Partial with Multi-Scale Feature Enhancement (C2f\_MSFE) is embedded into multiple backbone stages to construct multi-scale local difference responses and adaptively aggregate dilated convolution branches. Secondly, the Feature Fusion Block (FFB) performs selective cross-level fusion among the P2-derived feature, the upsampled deeper feature, and the native P3 feature through similarity-guided interaction and channel recalibration. Finally, the Variance Modulation Block (VMB) refines the fused high-resolution representation using local mean--variance modulation and a ConvMLP-based feature transformation. Experiments on the VisDrone2019 dataset show that, compared with YOLOv26-m, MSEV-DETR increases AP from 18.6% to 22.3% and AP50 from 33.2% to 38.8%, corresponding to absolute improvements of 3.7 and 5.6 percentage points, respectively. These gains correspond to relative improvements of 19.89% and 16.87%, while the parameter count is reduced from 20.4 M to 16.5 M. The results demonstrate the effectiveness of MSEV-DETR for UAV small-object detection in complex aerial scenes.
No comments yet — start the discussion below.