Seonghak Kim, Deukryeol Yoon, Young-Hwa Sung · ISPRS Journal of Photogrammetry and Remote Sensing 2026 · 2026
DOI: 10.1016/j.isprsjprs.2026.09.033
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Knowledge Distillation (KD) is a promising method for improving student model performance by transferring knowledge from high-capacity teacher models. However, in remote sensing object detection, where small objects are densely clustered, existing KD techniques suffer from notable limitations. By compressing scale-dependent characteristics into a single feature embedding, these methods often overlook scale-specific information, leading to scale entanglement dominated by large objects and, ultimately, degraded detection performance. We propose Distilling Adaptive Scale Knowledge (DASK), a framework designed to address these challenges by decoupling features into multi-scale representations for knowledge transfer and independent optimization. DASK consists of two key components: (1) Multi-Scale Feature (MSF), which separates features corresponding to small, medium, and large objects with multiple dilation rates; and (2) Scale-Specific Update (SSU), which mitigates inter-scale interference and regulates optimization dynamics. Specifically, SSU assigns a lower momentum to feature losses associated with small objects to induce a careful and precise optimization process, while applying a higher momentum to large object features to ensure stable and accelerated training. This design promotes more balanced optimization across scales. Extensive experiments on DOTA, VisDrone, and a custom dataset demonstrate that DASK consistently outperforms existing KD methods, highlighting its effectiveness for remote sensing applications.
No comments yet — start the discussion below.