Yuncheng Jiang, Chun-Mei Feng, Jiankun Hu, Lusheng Wang, Le Zhang · Tsinghua Science & Technology 2026 · 2026
DOI: 10.26599/tst.2026.9010095
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Medical multi-modal foundation models (MFMs) have shown strong generalization potential for smart clinical applications. However, conflicting cross-modal priors can cause them to underperform single-modal foundation models on specific downstream tasks. We define this phenomenon as “negative multi-modal transfer”, where the transfer gain of a multi-modal backbone becomes inconsistent or negative for a target modality. Through controlled experiments, we observe that (1) the extent of pretraining and (2) which layers are selected for fine-tuning substantially affect whether knowledge transfers effectively across modalities in downstream tasks. To address this, we propose MedRev, an efficient fine-tuning paradigm developed to reverse the negative transfer effect in medical foundation models. MedRev introduces a mixture-of-pre-trained experts strategy that adaptively lever-ages knowledge from different pre-trained checkpoints based on downstream data. Moreover, a dynamic layer freezing strategy is employed to decrease unnecessary updates and fine-tune those most beneficial for the target modality. Extensive experiments on downstream datasets across ten modalities demonstrate that MedRev consistently mitigates the negative transfer and increases fine-tuning performance.
No comments yet — start the discussion below.