Yongbao Ai, Tianxiang Gao, Zhipeng Lin, Longqi Yang, Qingyu Chang · Journal of Imaging 2026 · 2026
DOI: 10.3390/jimaging12080377
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Vision Transformers, especially Swin Transformer, have become default backbones for various vision tasks but suffer from high memory consumption and training costs. This letter proposes MoR–Swin, a novel architecture that integrates Mixture of Recursions (MoR) into Swin Transformer. An adaptive token-level recursion mechanism dynamically allocates computational depth based on semantic complexity. A recursive window attention module and a lightweight router with load balancing loss are introduced. Extensive experiments on ImageNet classification, COCO detection, and ADE20K segmentation show that MoR–Swin reduces parameters by about 50% and accelerates inference up to twofold at a modest accuracy cost (within about 0.5 points of Swin-B on ImageNet-1K). It provides a new technical pathway for optimizing Vision Transformer models, significantly enhancing their applicability in resource-constrained environments.
No comments yet — start the discussion below.